Commit Graph
14935 Commits
Author SHA1 Message Date
Ruben FiszelandClaude Opus 5.5 9f40cdca62 feat: add pull_batch to claim jobs for many waiting workers at once (#11350)
* feat: add pull_batch to claim jobs for many waiting workers at once

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep jobs admitted by earlier batch passes when a re-pull fails

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep suspended flows first on every batch re-pull pass

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 15:24:39 +02:00
hugocasa ede4103b00 feat: share AI guidance with the CLI skills and lint flow groups (#11397)
* feat: lint AI agent tool names and flow groups in wmill lint

* refactor: assemble chat and CLI AI guidance from one topic table

* feat: share flow groups, reuse and pipeline guidance with the CLI skills

* feat: share raw app, data table and secret guidance between chat and CLI

* docs: document the shared AI guidance source for contributors

* fix: keep wmill lint running on flows with malformed collections

* fix: reject skill descriptions that are not plain YAML text

* fix: tighten fence typos, script base scope and app prompt order

* fix: skip tool name checks on agent steps linked to a saved agent

* fix: catch any misspelled prompt fence and soften the tool name claim

* fix: align cli eval harness with the files and steps wmill init adds

* fix: drop cli eval checks that expect unrequested deploy commands

* fix: list ansible as mainless and c# Main in script base guidance
2026-09-29 14:47:42 +02:00
GuilhemandClaude Opus 5.5 d44c901647 replace the AI sessions beta banner with a feedback link (#11401)
* feat: make the AI sessions beta banner dismissible

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: move AI sessions feedback link into assistant settings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor: reduce the sessions beta gate setter to opt-in only

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: size the sessions activate buttons with unifiedSize

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep the sessions page activate button at its previous height

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 14:47:22 +02:00
AlexRV12 c34abf7330 fix: refuse flow preview restarts from runs the caller cannot read (#11407) 2026-09-29 14:46:51 +02:00
8741843d2e feat: external instance cluster for data tables and Ducklake catalogs (#11197)
* fix pg_dump stuck on version 17 on nix

* fix(datatables): refuse a malformed role annotation instead of ignoring it

`-- Role operator`, `-- role operator;` and `-- role operator -- why` all failed
the annotation parser's exact-match rule, so the query fell through to the data
table's default role and ran, silently, under a login the author did not choose.
Naming a role exists precisely to not do that.

A leading comment whose first word is `role` is now an annotation attempt: the
keyword matches case-insensitively, one trailing `;` is tolerated, and anything
else is an error naming the line. Only callers that already know the target is a
`datatable://` reference ever run this, so ordinary SQL keeps its comments.

Also bumps the dev shell's postgres client to 18 — it trailed the server the dev
database runs, which takes out every data table export, clone and fork-with-data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): refuse a malformed role query string instead of ignoring it

`?Role=analytics`, `?role=` and `?x=1&role=…` all fell through the reference
parser's exact-match rule, so the connection resolved to the data table's default
role and ran under a login the caller never asked for — the URI half of the same
trap as a malformed `-- role` annotation.

The key now matches case-insensitively, and anything else in the query string is
an error naming it; `role` is the only parameter a reference takes. Callers that
only need the entry keep a lenient `datatable_ref_name`, since they never act on
the role. The DuckDB `ATTACH` parser propagates it rather than attaching under
the default.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): carry the role annotation into the row_to_json retry

The retry rebuilds its SQL from `pruneComments(code)`, so the leading comment
block never reached the second attempt — and with it the `-- role <name>` line
that decides which login the query runs as. The retry connected as the data
table's default role instead, so a query the first attempt was denied could
succeed on the second, reported as "recovered with the row_to_json fix".

Carry the leading comment block over. The retry itself is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* chore(datatables): don't mount the roles UI until the ACL editor lands

Enforcement ships first. The permissions drawer is what turns roles on, and the
catalog section is what creates them — both are only useful once there is a way
to grant a role the privileges it needs, which arrives with the ACL editor. Left
mounted they would offer a feature whose other half does not exist.

The two components are complete and reviewed; only their call sites here are
commented out, with a note pointing the follow-up PRs at them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): honour `-- role: x`, and fix the DuckDB attach test

Two review findings, both real.

`attach_datatable_parses_name_and_role` never compiled: `parse_attach_datatable`
returns `Result<Option<_>>` now and one call site kept a single `unwrap`. Its
`?Role=analytics` case also asserted a refusal, contradicting the parser in the
same commit, which matches the key case-insensitively. Replaced with the cases
that are genuinely malformed, and a positive one for the cased key.

`-- role: analytics` fell through to the default role — the silent fallback the
strict parser exists to remove, for the spelling most likely to be typed. The
keyword now accepts an optional colon, attached or spaced, while a word that
merely starts with it (`rolebased`) is still not an attempt.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): clone a fork's pointer instead of failing after the copy

Forking a fork with cloning left an orphan database. The preflight resolves the
pointer and sees the governing entry, so both endpoints ran and filled the new
database; `apply_forked_datatable` then refused the inherited pointer and rolled
the fork back, stranding a registered `wm_fork_*` that no entry names and whose
name blocks the retry.

Refusing earlier would have been the smaller change, but forking a fork and
cloning worked before pointers existed, so it would trade an orphan for a
regression. Resolve what the pointer names and write the terminal entry the
clone needs: the whole `database` object rather than a patch of its
`resource_path`, since a pointer has none, and `reference` removed with it.

Also accepts `-- role=x` and `-- Role = x`, two more spellings that fell through
to the default role.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): refuse to roll back the catalog while roles exist

The down migration dropped the table and left every role behind: live Postgres
logins whose passwords only that table carried, so after a revert Windmill could
neither use, disable nor delete them, and re-applying could not recreate them
because the names were taken. Cleaning up here is not possible either — dropping
a role means reassigning what it owns in every instance database, and a
migration runs in one — so it now refuses while the catalog is non-empty and
says to delete the roles through instance settings, which does the cluster work.

Also enforces the instance-only invariant the resolved-pointer clone relies on
rather than only asserting it in a comment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* refactor(datatables): settle clonability in one place, before anything is created

A clone is three stages a workspace apart — `create_pg_database`, then
`import_pg_database`, then `apply_forked_datatable` inside the fork transaction.
Only the third can roll back, and `CREATE DATABASE` is not transactional, so any
refusal that lives there strands a registered `wm_fork_*` that no entry names
and whose name blocks the retry.

That orphan has now been fixed three times, most recently reintroduced by a
guard added one commit ago. Patching each new refusal into the first endpoint is
not the fix; having two places that can refuse is. `ensure_datatable_is_clonable`
now answers every reason a copy can be refused and returns what it resolved, and
the stage that writes the entry only does the work.

Also takes an ACCESS EXCLUSIVE lock before the rollback guard counts, so a role
created concurrently cannot slip between the check and the drop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): let a retried clone reclaim its own leftover database

A clone creates its target database one request before it copies into it, and
the fork that would name it is written a request after that. Any failure in
between — a pg_dump error, a bad restore, a dropped connection, the source's
roles changing mid-flow — left a registered `wm_fork_*` that no entry names,
and every retry then failed on its name. This predates data table roles.

`create_pg_database` now reclaims such a leftover before creating: only a
`wm_fork_*` database Windmill registered as a data table database and that no
data table or ducklake entry names, in any workspace, archived ones included.
The drop never terminates connections, so a clone still copying into it makes
the reclaim fail instead of being cut off. It is limited to callers who
administer the source — reaching it is not enough, since on a data table
without roles every member reaches it — and anyone else gets the refusal an
existing database always got.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Revert "fix(datatables): let a retried clone reclaim its own leftover database"

This reverts commit 7dd3275a10.

The reclaim tied the caller to the source they administer, but not to the
database it dropped. Between another workspace's import and its final fork
request, that workspace's target is full, registered, unnamed and has no open
connection, so an admin of any instance data table could name it and have it
dropped and recreated empty. The victim's fork would then commit pointing at
the empty copy. Safe reclaim needs durable clone ownership and serialization
with the request that names the database; until then the leftover stays, as it
did before this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(datatables): record the stale clone database as a known limitation

A clone is three requests and `CREATE DATABASE` is not transactional, so a
failure after the first leaves a registered `wm_fork_*` behind, as it did
before data table roles. Accepted for this PR: it is harmless to data and goes
away once the clone is a single server-side operation.

The comment also records why the obvious fix is wrong: reclaiming the leftover
on retry, without durable clone ownership, can drop another workspace's fully
copied database between its import and its final fork request.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): bounce the streams reading a data table when it is deleted

Deleting a governing data table, or the workspace that holds it, only collected
the fork pointers it stranded, for the warning. A Postgres trigger or capture
already streaming through one of those pointers kept the replication connection
it opened while the pointer still resolved, so it went on dispatching the
governing database's rows after the fork lost access — until its connection
happened to restart. The governing workspace's own streams on a deleted entry
did the same.

Both deletion paths now bounce the affected listeners inside their own
transaction, through the helper a permission change already uses, so a
listener that reconnects re-resolves the entry and finds it gone. The helper is
split so a caller can pass the (workspace, local name) pairs it already holds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep the fork schema baseline, and bounce streams on every removal

Three fixes from review.

`edit_datatable_config` took `forked_from` wholesale from the stored entry, so
the fork schema diff's save of an advanced baseline was silently discarded and
an applied change was offered again. Whether an entry carries a clone stamp is
still carried from the store, since that is what marks its database droppable,
but the baseline inside it is now taken from the request.

The stranded-pointer warning and the stream bounce ran over the optional
`deleted_datatables` hint, which the settings-sync CLI never sends, so removing
a governing data table through `wmill` bounced nothing. Removals are now derived
from the stored configuration against the saved one.

`delete_workspace` read the pointers to bounce before its transaction, so a fork
committing a pointer during the deletion was missed. The read now happens inside
the transaction, after the workspace row is deleted: a fork's insert key-share
locks that row through its parent foreign key, so it is either seen or fails on
the missing parent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(datatables): keep Postgres triggers and data table roles apart

A replication stream reads every row of every table whatever the data table's
roles grant, and its listener checks access only when it connects. Rather than
chase every way access can change and bounce the streams each one affects, a
data table now carries one or the other:

- a Postgres trigger or capture cannot be created on, or connect to, a data
  table under roles;
- roles cannot be turned on while an enabled trigger or a live capture reads
  the data table, its own or a fork's through its pointer. The refusal names
  each one to disable.

This removes the stream bounces on roles edits and on data table and workspace
deletion, and the trigger gate that admitted admins. The fork schema baseline
fix from the same review round is kept.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): refuse a Postgres trigger on a data table under roles when it is saved

Creating or editing a trigger that points at a data table under roles was
accepted, and its listener then retried the refused connection every 30
seconds forever. The save is now refused, and a trigger that reaches such a
data table anyway (re-enabled, or cloned into a fork) is disabled by its
listener with the reason, as a missing replication slot is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): disable a data table role before deleting it

Deleting a role reassigns and drops what it owns in each registered
database on its own connection, and each of those passes commits as it
goes. A database failing part-way left the role enabled in the catalog and
able to log in, but already stripped in the databases reached before it.
The role is now disabled in its own commit first, so a failed delete
leaves a disabled role to retry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): serialize roles going on with a stream starting

Turning roles on looked for enabled triggers and live captures once,
without a lock anything starting a stream also took. A trigger enabled in
that window could have its listener connect before roles committed, and a
healthy listener never checks again. Both transitions now serialize on one
advisory lock: roles going on hold it exclusive while they look, and
trigger create, edit and enable, and capture setup and ping hold it shared
while they commit. Either the look sees the stream, or the listener
connects after roles are committed and refuses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): wait out live listeners, and resolve stored names containing `?`

Turning roles on counted a trigger as gone once disabled, and a capture
once its client stopped pinging, but the listener keeps its replication
connection until its next heartbeat notices. A trigger or capture whose
listener pinged in the last 15 seconds, the window a server holds a
listener for, now still counts as streaming.

Data table names could contain `?` before they were restricted, and such
entries are still stored. Splitting `?role=` off a reference misread them:
`a?b` became `a` with an unknown parameter, and the clone checks looked at
a different entry than the one copied. An entry stored under the whole
reference is now looked up first, in the Postgres executor, DuckDB ATTACH
and the clone checks. Agent workers cannot read the workspace and keep
the strict parse, which refuses such a name rather than misreading it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): warn when a settings sync strands fork pointers

A settings save reported the fork pointers left resolving to nothing only
for the names in `deleted_datatables`, which `wmill sync push` never sends.
The save now works out what it removed from the locked entries, and the
CLI prints the stranded pointers it returns.

Also correct the replication helper's contract: no role or admin check
makes a replication connection safe, so a data table under roles is
refused outright rather than gated as an admin operation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): refuse a save that drops a data table's roles through an undeclared rename

A data table's roles follow its entry only through a declared rename. A
settings sync sends the whole map and never declares one, so renaming a
data table under roles there read as a delete and a new entry on the same
database: the new entry carried no roles, and every caller connected as
admin. Such a save is now refused, naming both entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): no entry without roles may newly reach a database under roles

The previous guard only caught a new name replacing an entry under roles.
A whole-map save could also repoint an existing entry without roles at
that database, or another workspace could point one there, and every
caller of that entry would connect as admin. The rule is now stated on
the saved entries: one that carries no roles and newly points at an
instance database any entry under roles uses, in this workspace or
another, is refused. A declared rename carries its roles and passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move data table role catalog and resolution to the enterprise edition

Roles are an Enterprise Edition feature. The catalog, the Postgres logins,
CONNECT convergence, tenant evaluation and the role half of connection
resolution move to windmill-ee-private. Every public function keeps its path
and signature and forwards through datatable_roles_oss, which re-exports the
enterprise implementation or, without it, refuses.

Without the enterprise edition a data table under roles, or a caller naming a
role, is refused a connection rather than resolved as admin, and the reach and
admin-access checks refuse one under roles. A data table not under roles
resolves as before in every edition, and an instance database keeps the
CONNECT grants it was created with. The catalog lock, the stream lock, the
tenant cascades and the permissions stripping stay in OSS: they only restrict.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move the data table permissions endpoints to the enterprise edition

The permissions read, save and usable-roles handlers move to
windmill-ee-private; the routes stay registered and, without the enterprise
edition, answer that data table roles are an Enterprise Edition feature.
ensure_governs_datatable and ensure_reaches_datatable keep their paths: the
first refuses, the second passes a data table not under roles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move the data table role catalog endpoints to the enterprise edition

The superadmin list, create, update and delete handlers move to
windmill-ee-private. The routes stay registered and, without the enterprise
edition, refuse after authentication.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* test(datatables): run the roles tests on the enterprise edition, refusals without it

Each test that exercises roles runs with private and enterprise. Two tests run
without them: every roles route answers the Enterprise refusal, and a data
table saved under roles, or a named role, is refused a connection while one
not under roles resolves as before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): gate the roles UI mount sites on an enterprise license

Both mount sites are still commented out; the gate travels with them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* test(datatables): run the tenant matcher test on the enterprise edition

The matcher it covers is enterprise code now, so without the enterprise
edition the test hit the stub and failed the default windmill-common run. It
runs with private and enterprise, and a counterpart without them asserts that
no tenant list covers anyone, the wildcard and a workspace admin included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* chore: update ee-repo-ref to a1873dbb67f2302b85ff5362f8387b48eccdb607

This commit updates the EE repository reference after PR #783 was merged in windmill-ee-private.

Previous ee-repo-ref: 5c853e2c20eca6b748415fc0d6862a6ebfb5fec4

New ee-repo-ref: a1873dbb67f2302b85ff5362f8387b48eccdb607

Automated by sync-ee-ref workflow.

* fix(datatables): refuse roles while a same-workspace alias reaches the database

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): add an ACL editor for data table roles

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): data table roles in the DB manager and raw apps

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: never add a role to the reference of a data table whose name contains '?'

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: read the roles of a data table whose name contains '?'

The generated client leaves a '?' in a path param unencoded, so the lookup
404'd and the raw-app picker blocked Start on such a data table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: take every pooled connection before the ACL apply locks

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: refresh grant options only after the ACL apply validates its plan

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix(datatables): refuse a reference naming both a legacy data table and a role

When a workspace stores both `sales` and a legacy `sales?role=analytics`, the
reference resolved to the legacy entry without a role, so browsing `sales` as
`analytics` reached another data table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: declare the default role in migrations written for a data table whose name contains '?'

Such a data table connects as its default role without naming it, so the
migrations the manager wrote for it declared no role and ran as admin.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): let CE migrations connect as an explicitly named admin

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: add only missing grant options before an ACL apply, never default privileges

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix(datatables): serialize roles going on with aliases saved from other workspaces

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(datatables): note that legacy names with ? cannot be migrated

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: run one data table ACL apply at a time per server before it connects

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* feat(datatables): set up an external instance cluster for data tables

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): send external cluster passwords as SCRAM verifiers

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): scope external cluster credential readers to the crate

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [ee] feat(datatables): external_instance data tables on the external cluster

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: hold the ACL connection to the database that was authorized

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix(datatables): compare the external cluster settings under a row lock before storing setup

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): only drop external databases Windmill marked, and check use under the lock

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: build the ACL connection from the authorized data table entry

Resolving the settings again could land on a resource with the same
database name on another server, which the later entry checks never see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: check ACL read reach against the entry it connects from

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* feat(datatables): put a data table's connection under Postgres roles

A data table backed by the instance database resolved to exactly one Postgres connection,
`custom_instance_user`, for everyone who could reach it at all. There was no way to say
this job reads, that one writes, this one never sees the salaries table.

A data table role is now a real Postgres login on the cluster, defined once for the
instance by a superadmin and named exactly as they named it. A script that declares
`-- role analytics` connects as `analytics`, and Postgres decides what it may touch —
grants are ordinary SQL. Windmill answers only "may this caller ask for this role", from
the tenant lists on the data table entry: `u/alice`, `g/analysts`, `f/finance` or `*`.
A data table with no `permissions` block behaves exactly as before.

Everything that opens a connection on someone's behalf goes through one chokepoint,
`get_datatable_resource_from_db`, which takes the identity explicitly and fails closed when
there is none. The role logs in as itself — never `SET ROLE`, which a script could
`RESET ROLE` its way out of.

A fork's data table entry becomes a pointer at the workspace that governs it rather than a
copy of it. The settings clone used to hand a fork a byte-identical entry naming the
parent's database, which a fork admin could edit to grant themselves `admin` there; a
pointer has nothing local to edit, and its tenants are evaluated as a member of the
governing workspace, by email. `permissions` is stripped from the workspace export and
ignored on import: tenants name principals of one workspace, and a settings push is not
where an access decision should be made.

Operations that see the whole database whatever the roles grant stay with the governing
workspace's admins: editing the roles, a migration that declares none, and opening a
replication stream for a Postgres trigger or capture.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): gate the paths that reach a whole database as admin

Auditing what still resolved through the unchecked resolver turned up three that act for a
caller and hand back the admin connection: `resolve_pg_source_checked` (behind schema
export, the full-schema read, database creation, import and the forked-database drop), the
connection test, and the schema snapshot a fork clone takes of its parent. On a data table
under roles each let any workspace member — or a fork admin who is nobody in the governing
workspace — read or copy the whole database whatever its roles grant.

All three now require admin reach on the governing workspace. A dump taken under a
restricted role would be a silently truncated copy rather than an error, so refusing is the
only right answer for the copy paths.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): confine roles to the instance database, and stop a fork reaching the parent's bookkeeping

A data table role is a login on Windmill's own Postgres. Nothing stopped a workspace admin
putting a *resource-backed* data table under roles, at which point the executor dialled the
host that resource names — one the admin chose — with the role's real cluster password, and
`CONNECT` is granted to every registered instance database. Both ends now refuse: the
permissions endpoint rejects the save, and the chokepoint refuses to substitute credentials
on a non-instance entry rather than trusting the record it read.

Two more places reached the governing database without answering to it. The initial-migration
generator returned a `pg_dump` of the whole schema to any member. And the migration
rename/delete cascade followed a fork's pointer into the parent, so a fork admin renaming or
removing their own local entry relabelled or wiped the parent's `_wm_migrations` — after
which the parent re-runs every migration from zero. The remote half is now skipped when the
entry resolves into another workspace, which is also just correct: a fork renaming what it
calls a data table changes nothing about the data table.

Also: revoking a tenant now bounces the replication streams of every workspace holding an
entry that resolves here, not only the governing one, so a fork's trigger stops rather than
living on inside its open connection; the instance role catalog and the governing workspace's
tenant lists are no longer returned to someone who cannot edit them; and the tenant rename
dedup collapses non-adjacent duplicates, per role rather than once any role changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): fail loudly where a role or a pointer can be left half-recorded

Three ways the feature could end up in a state nobody could see or undo.

Creating a role writes the cluster first and the catalog second, but the catalog write was an
`UPDATE` that matched nothing when the instance Postgres settings row was absent — leaving a
live login with a password nobody recorded: invisible to the catalog, un-recreatable because
the name is taken, and un-deletable because there is no entry to delete. It now errors, so
the operation is retryable once the row is restored.

Deleting a workspace only nulls the fork lineage; the data table entries pointing at it are
left resolving to nothing. Sweeping them is not an option — turning a pointer back into a copy
would hand each fork the database outright — so the delete now names the data tables it
stranded, and resolving one says which workspace is missing rather than reporting a data table
this workspace never had.

`InstanceDatatableRole` derived `Debug` while holding a Postgres password; it is now
hand-written so `{:?}` on the catalog cannot put a live credential in a log line.

Adds the two branches the reviews found unpinned: a caller who is not a member of the
governing workspace at all, and `NoIdentity` — the compatibility path for an agent worker that
predates this and sends no job id.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): unbreak two operator messages and two comments that described other code

The two strings this branch added for states an operator hits once — the catalog write that
matched nothing, and the delete that stranded a pointer — were collapsed from their multi-line
form with the indentation left in, so both rendered with a fourteen-space gap mid-sentence.

`list_datatables` claimed to report a chain it cannot follow and then dropped it; it does drop
it, and the comment now says why that is the right place to stay quiet. The non-superadmin
check in `edit_datatable_config` was introduced as also covering references, which it does not
and need not: `reference` is overwritten from the stored entry for every caller before the
check runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): serialize role catalog mutations, and state each helper's authorization contract

The catalog is one JSON document, so create, rename, enable and delete are all
read-modify-write. Two concurrent creates read the same snapshot, both succeed in the
cluster, and the second write drops the first — leaving a live Postgres login with a password
nobody recorded, which is the exact state the delete path exists to prevent. Every mutation
now runs in one transaction holding an advisory lock across the read, the cluster DDL and the
write, so a lost update cannot happen and a failure rolls the whole thing back. The DDL
helpers take that transaction rather than the pool, which is what makes the lock cover them.

Their statements moved off `sqlx::raw_sql`: the simple protocol is only needed for genuinely
multi-statement SQL, and its future is not `Send`, which an axum handler holding the
transaction requires. Each of these is one statement anyway.

The new cross-crate surface now says what callers must do. `read_role_catalog` returns
plaintext credentials; `create`/`rename`/`set_login`/`drop_instance_role` and
`converge_connect_grants` mutate cluster-wide state; `read_datatable_entry` reads a workspace's
raw config. All of them are superadmin-gated by their current handlers, but nothing said so at
the definition, which is where the next caller looks.

Also: the roles table reloads after a failed login toggle instead of leaving it claiming a flip
that did not land; the rename affordance is the design-system `Button`, not a raw one; and
`resolve_datatable_pg_as_caller` drops a `role` parameter no caller ever filled — browsing
resolves as the data table's default until the database manager grows a picker.

Why role passwords stay a plain `String` while the instance user's password beside them is a
`StringOrSecretRef`, asked three times across reviews: that one is a secret ref because an
operator supplies it and may want it from their own backend, while these are minted here and
never entered by anyone, so there is nothing for a ref to point at. Encrypting generated
secrets at rest is a separate change that would take the replication password with it. Now
said at the field.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): give the role catalog its own row, out of reach of the config machinery

Putting it inside `custom_instance_pg_databases` was the wrong call, and it cost two ways.
The catalog serializes a generated Postgres password per role, and that row is the
operator-facing instance config, so the passwords reached `get_instance_config` and its YAML
editor — a live cluster credential in a response body, a UI field and any log of either.
Worse in the other direction: `to_settings_map` strips the catalog, so a full-row upsert of
that key writes the row back without it and the catalog is gone, while the cluster keeps every
login it described.

`custom_instance_replication_pwd` is the precedent and says exactly why — a generated secret,
written only by the server, never operator-authored, hidden so the config machinery cannot
read, rewrite or drop it. The catalog is the same thing, so it now has the same shape:
`datatable_roles`, in `HIDDEN_SETTINGS`, `PROTECTED_SETTINGS` and the agent-worker denylist.
No redaction to keep in step with three code paths, and no way for a neighbouring write to
take it out.

Two races on the same shared documents. `edit_datatable_config` read the stored data tables
outside its transaction and then wrote the whole `datatable` document, so a permissions save
committing in between was silently rolled back; it now reads under `FOR UPDATE`. And
`set_datatable_permissions` validated role ids against the catalog before opening its
transaction, so a deletion in between let it write a deleted role back — including as the
default, which every later job then fails on; it now holds the catalog lock and the settings
row across validation and write.

Completes the authorization contracts the previous commit claimed but did not finish:
`read_datatable_entry` (which it named and missed), `resolve_governing_datatable`, whose whole
job is to answer for a workspace the caller may not belong to, and
`converge_connect_grants_with`, which had not inherited its wrapper's.

Also the generic Python SDK reference: `_format_py_params` learned the bare `*` last time, but
`extract_py_functions` is a second formatter and still rendered `datatable(name, role)`, so
code written from that page passed a keyword-only argument positionally.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): make the concurrency test pin the handlers, and the contracts describe what is enforced

The concurrency test reimplemented the read-modify-write inline, so deleting the lock from all
three handlers left it green — it pinned Postgres, not the code it was written for. It now
drives `create_datatable_role` twice concurrently and asserts the catalog kept both names.
Checked the way the last one should have been: removing the lock from the handler makes it
fail with "wmtest_a_… is a live cluster login the catalog forgot".

The contracts added last commit were stricter than this PR's own callers, which is worse than
none — the next reader sees a rule already broken and learns to ignore it.
`read_role_catalog` said superadmin-only while two of its four callers are open to any
workspace member, and `converge_connect_grants` said superadmin while
`set_datatable_permissions` reaches it as a workspace admin. Both were fine on substance: the
rule that actually holds is about the credential never reaching a response, log, audit record
or export, not about who may call. They now say that. `read_datatable_entry` gets the same
treatment rather than the one the earlier message claimed for it: it is the primitive every
resolution goes through, so it is deliberately open, and what must not escape is `permissions`
— it names the governing workspace's users, groups and folders.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): close the last ways a role or a pointer can be left pointing at nothing

The raw settings readers hand back whatever is in the row, so moving the catalog into its own
`global_settings` key protected the config machinery and left `GET /settings/global/datatable_roles`
and the settings listing returning every live password. Both now filter that one key. The
neighbouring `custom_instance_replication_pwd` has the same shape and is not touched here: it
predates this and widening the fix to it is a decision about an operator workflow, not a
consequence of this change.

Three ways a save could leave something resolving to nothing:

A permissioned data table could be moved to a PostgreSQL resource. The block was carried across
as a server-owned field, the runtime refuses roles on a resource-backed table, so the save
succeeded and every job afterwards failed. Refused instead — turning roles off first is one step,
and it keeps discarding an access decision something somebody chose.

Renaming a governing data table left every fork pointing at the old name: the data table
disappears from their pickers and their jobs stop, with nothing in the renaming workspace to
suggest why. The rename now follows into the pointers in the same transaction.

Deleting one cannot be followed the same way, so it is reported instead — the response names what
it stranded, the way deleting a workspace does, and the fork's own error already says which
workspace is gone.

Also: `ensure_instance_db_grant_options_unchecked` claimed superadmin while the permissions
handler reaches it as a workspace admin (the same class fixed last commit, one instance missed);
the role entry kept an `instance_config_schema` derive it no longer needs; `write_role_catalog`
was the one writer of that table not stamping `updated_at`; and the concurrency test dropped its
roles only on success — a failing run is exactly the one that creates them without recording them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* refactor(datatables): put the role catalog in its own table, not in global_settings

Five findings across three rounds were all the same choice. A set of live Postgres credentials
was living in `global_settings`, which has generic read, list, write, config-export and CLI
round-trip paths that know nothing about what they carry: the passwords reached the instance
config and its YAML editor, a full-row upsert of a neighbouring key erased the catalog,
`GET /settings/global/{key}` and the settings listing returned them raw, and this round the
redaction that fixed the last two turned `wmill instance push` into something that wipes every
password — a fix breaking the assumption the previous fix made. `POST /settings/global/datatable_roles`
could also empty it outside the lock.

The approved plan offered a table or `global_settings`, so this is the other option it already
allowed rather than a new design. `datatable_role` is a table: no generic settings path can read
it, list it, export it, write it or round-trip it, so none of the five needs a guard. The
redaction, the hidden/protected/agent-denylist entries and the JSON document all go with it.

One row per role also removes the read-modify-write the concurrency work was about: two
concurrent creates are two inserts, and the unique index on `name` is what settles a collision.
The advisory lock stays for the one window rows do not cover — `CREATE ROLE` is invisible to
another transaction until commit, so without it both creates pass their `pg_roles` check.

Also from this round: rename mappings are checked against the configuration they claim to
describe, since fork pointers are rewritten from them — a caller could otherwise submit
`main -> missing` against an unchanged config and repoint every fork of `main` at a name nothing
has, and `A -> B` plus `B -> C` moved what pointed at `A` all the way to `C`. And the warning
naming forks a delete stranded reached the response but not the screen: both the data table
settings save and the workspace delete now show it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): validate a rename against the save it describes, and re-check under the locks

Three from the round, all about deciding on state that could already have moved.

A permission save resolved the data table and checked it was instance-backed before taking any
lock, then wrote under one. A config save committing in between could move the table onto a
PostgreSQL resource — recreating exactly what the transition guard refuses — or rename it, in
which case the write targeted a key that no longer existed and reported success having changed
nothing. It now re-resolves and re-checks on the locked state.

Rename validation checked that the source existed before and the target existed after, which
still accepts `main -> decoy` against a save that keeps both: every fork of `main` then follows
onto a different data table, silently, because it keeps resolving. The rule is now the actual
old-to-new key transition — a source may only survive if another rename took its name, and a
target may only pre-exist if another rename freed it. That also stops two sources sharing one
target, and it admits a swap, which the previous guard refused: `datatables` is keyed by name, so
a swap cannot be done one save at a time, and refusing it was a regression against main. The
pointer cascade now runs in two passes through a temporary name, the way the migration cascade
one layer down already handles the same shape, so `A -> B` with `B -> C` moves each pointer once
from what it named before the save.

The tenant mutators say what they are for: they write an access decision for any workspace named,
with an arbitrary mutation, and exist for the transaction that frees or renames a principal.
Editing a decision on purpose belongs in the permissions endpoint.

Carried in the same change: the stranded-fork list is a field rather than a phrase to grep out of
a success string; the pointer cascade matches with `EXISTS` instead of a `LIKE` over the whole
document, so a workspace whose pointers name something else is not rewritten to a byte-identical
value under an exclusive lock; and `InstanceDatatableRole` drops the serde derives left over from
the JSON document, one of which would emit `pwd`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): cascade on the leave route that is used, gate migrations before the admin connection, and drop a role atomically

The tenant cascade on leaving went onto `/users/leave`. The UI and the generated client call
`/workspaces/leave` — a different handler in a different crate with the same name — which
deleted the membership and left `u/<username>` in the tenant lists. Leaving and rejoining
therefore restored the access the leave was supposed to end, and a later account taking the
username would have inherited it. The regression test drives the route the client actually
calls; without the fix it fails with "leaving kept the tenant".

The migration endpoints authorized too late. `run_datatable_migrations` opened the data table's
admin connection, created `_wm_migrations` and read it before reaching the per-migration role
check — so with nothing pending, nothing was checked at all. Rollback returned before its check
when nothing was applied, and the status endpoint had none. All three now ask, before any
connection is opened, whether the caller can reach the data table as any role at all; which role
a given migration runs as is still decided per migration, and by the executor after that.

Deleting a role committed the cluster drop and the catalog row, then swept the tenant lists in
separate transactions. A sweep failing part-way left workspaces naming a role nothing can connect
as, while the retry answered `NotFound` because the catalog entry was already gone. The sweep now
runs in the same transaction, so the drop, the row and every tenant list commit together.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): refuse to copy a data table that is under roles

pg_dump carries no roles and the import runs with --no-privileges, so a copied
data table arrives owned by the admin connection with no GRANT for any role.
The settings clone brings `permissions` across, so the fork's tenants pass
Windmill's check, connect as the role they were given, and are denied by
Postgres on everything: an entry that reads as configured and answers nothing.

Refuse the copy — in the import endpoint before any data moves, and in the fork
path the CLI takes. Replaying the source's owners and ACLs into the clone is
what lifts this, and is a change of its own. Dropping `permissions` from the
copy instead would be the unsafe half, since the copy holds the parent's rows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): refuse the clone's database too, not only its data

A clone is two endpoints: `create_pg_database` then `import_pg_database`. Only
the second refused a data table under roles, so a fork asking to clone one
created and registered an empty `wm_fork_…` instance database and then failed —
and nothing collects it, since `drop_forked_datatable_databases` only drops
entries carrying `forked_from` and no entry names this one.

Refuse in both, so the clone stops before a database exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* nit worker error msg

* fix pg_dump stuck on version 17 on nix

* fix(datatables): refuse a malformed role annotation instead of ignoring it

`-- Role operator`, `-- role operator;` and `-- role operator -- why` all failed
the annotation parser's exact-match rule, so the query fell through to the data
table's default role and ran, silently, under a login the author did not choose.
Naming a role exists precisely to not do that.

A leading comment whose first word is `role` is now an annotation attempt: the
keyword matches case-insensitively, one trailing `;` is tolerated, and anything
else is an error naming the line. Only callers that already know the target is a
`datatable://` reference ever run this, so ordinary SQL keeps its comments.

Also bumps the dev shell's postgres client to 18 — it trailed the server the dev
database runs, which takes out every data table export, clone and fork-with-data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): refuse a malformed role query string instead of ignoring it

`?Role=analytics`, `?role=` and `?x=1&role=…` all fell through the reference
parser's exact-match rule, so the connection resolved to the data table's default
role and ran under a login the caller never asked for — the URI half of the same
trap as a malformed `-- role` annotation.

The key now matches case-insensitively, and anything else in the query string is
an error naming it; `role` is the only parameter a reference takes. Callers that
only need the entry keep a lenient `datatable_ref_name`, since they never act on
the role. The DuckDB `ATTACH` parser propagates it rather than attaching under
the default.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): carry the role annotation into the row_to_json retry

The retry rebuilds its SQL from `pruneComments(code)`, so the leading comment
block never reached the second attempt — and with it the `-- role <name>` line
that decides which login the query runs as. The retry connected as the data
table's default role instead, so a query the first attempt was denied could
succeed on the second, reported as "recovered with the row_to_json fix".

Carry the leading comment block over. The retry itself is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* chore(datatables): don't mount the roles UI until the ACL editor lands

Enforcement ships first. The permissions drawer is what turns roles on, and the
catalog section is what creates them — both are only useful once there is a way
to grant a role the privileges it needs, which arrives with the ACL editor. Left
mounted they would offer a feature whose other half does not exist.

The two components are complete and reviewed; only their call sites here are
commented out, with a note pointing the follow-up PRs at them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): honour `-- role: x`, and fix the DuckDB attach test

Two review findings, both real.

`attach_datatable_parses_name_and_role` never compiled: `parse_attach_datatable`
returns `Result<Option<_>>` now and one call site kept a single `unwrap`. Its
`?Role=analytics` case also asserted a refusal, contradicting the parser in the
same commit, which matches the key case-insensitively. Replaced with the cases
that are genuinely malformed, and a positive one for the cased key.

`-- role: analytics` fell through to the default role — the silent fallback the
strict parser exists to remove, for the spelling most likely to be typed. The
keyword now accepts an optional colon, attached or spaced, while a word that
merely starts with it (`rolebased`) is still not an attempt.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): clone a fork's pointer instead of failing after the copy

Forking a fork with cloning left an orphan database. The preflight resolves the
pointer and sees the governing entry, so both endpoints ran and filled the new
database; `apply_forked_datatable` then refused the inherited pointer and rolled
the fork back, stranding a registered `wm_fork_*` that no entry names and whose
name blocks the retry.

Refusing earlier would have been the smaller change, but forking a fork and
cloning worked before pointers existed, so it would trade an orphan for a
regression. Resolve what the pointer names and write the terminal entry the
clone needs: the whole `database` object rather than a patch of its
`resource_path`, since a pointer has none, and `reference` removed with it.

Also accepts `-- role=x` and `-- Role = x`, two more spellings that fell through
to the default role.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): refuse to roll back the catalog while roles exist

The down migration dropped the table and left every role behind: live Postgres
logins whose passwords only that table carried, so after a revert Windmill could
neither use, disable nor delete them, and re-applying could not recreate them
because the names were taken. Cleaning up here is not possible either — dropping
a role means reassigning what it owns in every instance database, and a
migration runs in one — so it now refuses while the catalog is non-empty and
says to delete the roles through instance settings, which does the cluster work.

Also enforces the instance-only invariant the resolved-pointer clone relies on
rather than only asserting it in a comment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* refactor(datatables): settle clonability in one place, before anything is created

A clone is three stages a workspace apart — `create_pg_database`, then
`import_pg_database`, then `apply_forked_datatable` inside the fork transaction.
Only the third can roll back, and `CREATE DATABASE` is not transactional, so any
refusal that lives there strands a registered `wm_fork_*` that no entry names
and whose name blocks the retry.

That orphan has now been fixed three times, most recently reintroduced by a
guard added one commit ago. Patching each new refusal into the first endpoint is
not the fix; having two places that can refuse is. `ensure_datatable_is_clonable`
now answers every reason a copy can be refused and returns what it resolved, and
the stage that writes the entry only does the work.

Also takes an ACCESS EXCLUSIVE lock before the rollback guard counts, so a role
created concurrently cannot slip between the check and the drop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): let a retried clone reclaim its own leftover database

A clone creates its target database one request before it copies into it, and
the fork that would name it is written a request after that. Any failure in
between — a pg_dump error, a bad restore, a dropped connection, the source's
roles changing mid-flow — left a registered `wm_fork_*` that no entry names,
and every retry then failed on its name. This predates data table roles.

`create_pg_database` now reclaims such a leftover before creating: only a
`wm_fork_*` database Windmill registered as a data table database and that no
data table or ducklake entry names, in any workspace, archived ones included.
The drop never terminates connections, so a clone still copying into it makes
the reclaim fail instead of being cut off. It is limited to callers who
administer the source — reaching it is not enough, since on a data table
without roles every member reaches it — and anyone else gets the refusal an
existing database always got.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Revert "fix(datatables): let a retried clone reclaim its own leftover database"

This reverts commit 7dd3275a10.

The reclaim tied the caller to the source they administer, but not to the
database it dropped. Between another workspace's import and its final fork
request, that workspace's target is full, registered, unnamed and has no open
connection, so an admin of any instance data table could name it and have it
dropped and recreated empty. The victim's fork would then commit pointing at
the empty copy. Safe reclaim needs durable clone ownership and serialization
with the request that names the database; until then the leftover stays, as it
did before this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(datatables): record the stale clone database as a known limitation

A clone is three requests and `CREATE DATABASE` is not transactional, so a
failure after the first leaves a registered `wm_fork_*` behind, as it did
before data table roles. Accepted for this PR: it is harmless to data and goes
away once the clone is a single server-side operation.

The comment also records why the obvious fix is wrong: reclaiming the leftover
on retry, without durable clone ownership, can drop another workspace's fully
copied database between its import and its final fork request.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): bounce the streams reading a data table when it is deleted

Deleting a governing data table, or the workspace that holds it, only collected
the fork pointers it stranded, for the warning. A Postgres trigger or capture
already streaming through one of those pointers kept the replication connection
it opened while the pointer still resolved, so it went on dispatching the
governing database's rows after the fork lost access — until its connection
happened to restart. The governing workspace's own streams on a deleted entry
did the same.

Both deletion paths now bounce the affected listeners inside their own
transaction, through the helper a permission change already uses, so a
listener that reconnects re-resolves the entry and finds it gone. The helper is
split so a caller can pass the (workspace, local name) pairs it already holds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep the fork schema baseline, and bounce streams on every removal

Three fixes from review.

`edit_datatable_config` took `forked_from` wholesale from the stored entry, so
the fork schema diff's save of an advanced baseline was silently discarded and
an applied change was offered again. Whether an entry carries a clone stamp is
still carried from the store, since that is what marks its database droppable,
but the baseline inside it is now taken from the request.

The stranded-pointer warning and the stream bounce ran over the optional
`deleted_datatables` hint, which the settings-sync CLI never sends, so removing
a governing data table through `wmill` bounced nothing. Removals are now derived
from the stored configuration against the saved one.

`delete_workspace` read the pointers to bounce before its transaction, so a fork
committing a pointer during the deletion was missed. The read now happens inside
the transaction, after the workspace row is deleted: a fork's insert key-share
locks that row through its parent foreign key, so it is either seen or fails on
the missing parent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(datatables): keep Postgres triggers and data table roles apart

A replication stream reads every row of every table whatever the data table's
roles grant, and its listener checks access only when it connects. Rather than
chase every way access can change and bounce the streams each one affects, a
data table now carries one or the other:

- a Postgres trigger or capture cannot be created on, or connect to, a data
  table under roles;
- roles cannot be turned on while an enabled trigger or a live capture reads
  the data table, its own or a fork's through its pointer. The refusal names
  each one to disable.

This removes the stream bounces on roles edits and on data table and workspace
deletion, and the trigger gate that admitted admins. The fork schema baseline
fix from the same review round is kept.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): refuse a Postgres trigger on a data table under roles when it is saved

Creating or editing a trigger that points at a data table under roles was
accepted, and its listener then retried the refused connection every 30
seconds forever. The save is now refused, and a trigger that reaches such a
data table anyway (re-enabled, or cloned into a fork) is disabled by its
listener with the reason, as a missing replication slot is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): disable a data table role before deleting it

Deleting a role reassigns and drops what it owns in each registered
database on its own connection, and each of those passes commits as it
goes. A database failing part-way left the role enabled in the catalog and
able to log in, but already stripped in the databases reached before it.
The role is now disabled in its own commit first, so a failed delete
leaves a disabled role to retry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): serialize roles going on with a stream starting

Turning roles on looked for enabled triggers and live captures once,
without a lock anything starting a stream also took. A trigger enabled in
that window could have its listener connect before roles committed, and a
healthy listener never checks again. Both transitions now serialize on one
advisory lock: roles going on hold it exclusive while they look, and
trigger create, edit and enable, and capture setup and ping hold it shared
while they commit. Either the look sees the stream, or the listener
connects after roles are committed and refuses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): wait out live listeners, and resolve stored names containing `?`

Turning roles on counted a trigger as gone once disabled, and a capture
once its client stopped pinging, but the listener keeps its replication
connection until its next heartbeat notices. A trigger or capture whose
listener pinged in the last 15 seconds, the window a server holds a
listener for, now still counts as streaming.

Data table names could contain `?` before they were restricted, and such
entries are still stored. Splitting `?role=` off a reference misread them:
`a?b` became `a` with an unknown parameter, and the clone checks looked at
a different entry than the one copied. An entry stored under the whole
reference is now looked up first, in the Postgres executor, DuckDB ATTACH
and the clone checks. Agent workers cannot read the workspace and keep
the strict parse, which refuses such a name rather than misreading it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): warn when a settings sync strands fork pointers

A settings save reported the fork pointers left resolving to nothing only
for the names in `deleted_datatables`, which `wmill sync push` never sends.
The save now works out what it removed from the locked entries, and the
CLI prints the stranded pointers it returns.

Also correct the replication helper's contract: no role or admin check
makes a replication connection safe, so a data table under roles is
refused outright rather than gated as an admin operation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): refuse a save that drops a data table's roles through an undeclared rename

A data table's roles follow its entry only through a declared rename. A
settings sync sends the whole map and never declares one, so renaming a
data table under roles there read as a delete and a new entry on the same
database: the new entry carried no roles, and every caller connected as
admin. Such a save is now refused, naming both entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): no entry without roles may newly reach a database under roles

The previous guard only caught a new name replacing an entry under roles.
A whole-map save could also repoint an existing entry without roles at
that database, or another workspace could point one there, and every
caller of that entry would connect as admin. The rule is now stated on
the saved entries: one that carries no roles and newly points at an
instance database any entry under roles uses, in this workspace or
another, is refused. A declared rename carries its roles and passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move data table role catalog and resolution to the enterprise edition

Roles are an Enterprise Edition feature. The catalog, the Postgres logins,
CONNECT convergence, tenant evaluation and the role half of connection
resolution move to windmill-ee-private. Every public function keeps its path
and signature and forwards through datatable_roles_oss, which re-exports the
enterprise implementation or, without it, refuses.

Without the enterprise edition a data table under roles, or a caller naming a
role, is refused a connection rather than resolved as admin, and the reach and
admin-access checks refuse one under roles. A data table not under roles
resolves as before in every edition, and an instance database keeps the
CONNECT grants it was created with. The catalog lock, the stream lock, the
tenant cascades and the permissions stripping stay in OSS: they only restrict.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move the data table permissions endpoints to the enterprise edition

The permissions read, save and usable-roles handlers move to
windmill-ee-private; the routes stay registered and, without the enterprise
edition, answer that data table roles are an Enterprise Edition feature.
ensure_governs_datatable and ensure_reaches_datatable keep their paths: the
first refuses, the second passes a data table not under roles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move the data table role catalog endpoints to the enterprise edition

The superadmin list, create, update and delete handlers move to
windmill-ee-private. The routes stay registered and, without the enterprise
edition, refuse after authentication.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* test(datatables): run the roles tests on the enterprise edition, refusals without it

Each test that exercises roles runs with private and enterprise. Two tests run
without them: every roles route answers the Enterprise refusal, and a data
table saved under roles, or a named role, is refused a connection while one
not under roles resolves as before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): gate the roles UI mount sites on an enterprise license

Both mount sites are still commented out; the gate travels with them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* test(datatables): run the tenant matcher test on the enterprise edition

The matcher it covers is enterprise code now, so without the enterprise
edition the test hit the stub and failed the default windmill-common run. It
runs with private and enterprise, and a counterpart without them asserts that
no tenant list covers anyone, the wildcard and a workspace admin included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* chore: update ee-repo-ref to a1873dbb67f2302b85ff5362f8387b48eccdb607

This commit updates the EE repository reference after PR #783 was merged in windmill-ee-private.

Previous ee-repo-ref: 5c853e2c20eca6b748415fc0d6862a6ebfb5fec4

New ee-repo-ref: a1873dbb67f2302b85ff5362f8387b48eccdb607

Automated by sync-ee-ref workflow.

* fix(datatables): refuse roles while a same-workspace alias reaches the database

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): let CE migrations connect as an explicitly named admin

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): serialize roles going on with aliases saved from other workspaces

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(datatables): note that legacy names with ? cannot be migrated

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): add an ACL editor for data table roles

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: take every pooled connection before the ACL apply locks

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: refresh grant options only after the ACL apply validates its plan

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: add only missing grant options before an ACL apply, never default privileges

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: run one data table ACL apply at a time per server before it connects

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: hold the ACL connection to the database that was authorized

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: build the ACL connection from the authorized data table entry

Resolving the settings again could land on a resource with the same
database name on another server, which the later entry checks never see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: check ACL read reach against the entry it connects from

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* feat(datatables): Ducklake catalogs on the external instance cluster

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): never grant CREATEROLE to custom_instance_user on the external cluster

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(datatables): state the authorization contract of external database usage lookups

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): drop a DuckDB data table secret once its ATTACH has used it

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf(datatables): resolve a workspace's data tables per pointer hop, not per entry

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): hold the parent's settings while a fork points at its data tables

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): refuse repointing the external cluster while it is in use, and keep verify-ca working for pg_dump

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): protect external databases pending fork cleanup, and describe Ducklake usage in the API

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): per-cluster data table role catalogs, with roles on the external instance cluster

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): write the external cluster setting under the lifecycle lock, and check fork targets are registered

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): register external fork catalogs under the lifecycle lock, and keep certificate verification in DuckDB attaches

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep certificate verification when DuckDB attaches an external data table

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep certificate verification when DuckDB attaches an external data table

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): refuse fork cleanup of an external database another workspace uses

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): refuse rolling back while external data tables are under roles, and type external_instance in the CLI

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): stop counting storage-only fork cleanup rows as uses of an external database

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): bind fork database copies to their workspace, and count every use before dropping one

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): authenticate instance database setup before writing its status, and keep a fork reservation across it

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): create external databases only on a cluster setup succeeded on, and document the registry reader

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep only the most recently used DuckDB root certificate files

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): migrate fork reservations on workspace rename, and lock the parent's data tables for the whole fork

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): take the fork data table lock once, before the external cluster's

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(datatables): describe the external instance cluster and how to run one locally

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep fork reservations private, drop a cleaned-up entry with its database, and serialize cleanup with settings saves

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): hold the fork lock across a fork import, and carry the reservation inside the setup write

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): check the external cluster setting on its own transaction, and gate the registry probe

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): add an ACL editor for data table roles

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: take every pooled connection before the ACL apply locks

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: refresh grant options only after the ACL apply validates its plan

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: add only missing grant options before an ACL apply, never default privileges

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: run one data table ACL apply at a time per server before it connects

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: hold the ACL connection to the database that was authorized

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: build the ACL connection from the authorized data table entry

Resolving the settings again could land on a resource with the same
database name on another server, which the later entry checks never see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: check ACL read reach against the entry it connects from

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* chore: update ee-repo-ref to 7e338e4dabf91689bfd7fb0333c6534040b17b59

This commit updates the EE repository reference after PR #787 was merged in windmill-ee-private.

Previous ee-repo-ref: 0edd40979cf36bfba59323f3f6a0811ae1369cf5

New ee-repo-ref: 7e338e4dabf91689bfd7fb0333c6534040b17b59

Automated by sync-ee-ref workflow.

* fix(datatables): keep DuckDB root certificate files in the job directory

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: drop the unused json import from the settings crate

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): drop the serde_json::json import left unused by the fork setup write

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): accept external_instance data tables in the settings form type

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): serialize fork database reservations, cleanup, setup and Ducklake saves on the database

- A fork copy is created and registered under its workspace's fork lock, and refused once the
  workspace is archived; a rename re-migrates reservations under that lock after archiving.
- Ducklake saves lock every instance database they newly name, as data table saves do.
- Instance database setup holds the database's lock until its entry is written.
- Fork import also holds the database's lock across the restore.
- Cleanup re-reads the entry under its locks before dropping anything.
- Non-superadmins no longer see fork copies reserved for workspaces they are not in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): state the authorization contract of the external cluster status and write check

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): lock the fork copy at finalization, and order cleanup's registry write after its settings row

- Fork finalization holds the copy's database lock from its availability check to the commit.
- Cleanup removes the registry entry in its own transaction, in a task of its own, instead of on a
  second connection; a rename migrates reservations after its settings rewrites, so both take the
  settings rows before the registry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): hold the fork lock across a rename's settings copy

Fork cleanup of the old id could otherwise drop a copy the renamed workspace goes on using. Also
states the authorization contract of the instance database drop helpers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): refuse moving the external cluster to another host or port while it holds databases

Every settings writer goes through the same check as removal: the login, TLS and maintenance
database may still change, the cluster may not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): clone a data table under roles with its owners and grants (#11120)

Claude-Session: https://claude.ai/code/session_01UbrtwiYNfayrmqouBJHwGV

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: open the raw app data table drawer when the workspace has none

Selecting the first data table of an empty list passed undefined to the name
check, which threw instead of opening the drawer on no data table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): offer cloning a data table under roles where its grants can be replayed

The server clones such a data table and replays the source's owners and grants,
which only the Enterprise Edition does, so the fork wizard hid both clone
options everywhere instead of on a build that cannot replay them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: name the placeholder the empty raw app data drawer renders

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): replace a job directory file at the DuckDB root certificate path instead of trusting it

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: pin the enterprise refusal the role pickers read as 'not under roles'

The server's sentence and the frontend's copy of it were coupled by nothing,
so rewording either one turned every role picker on a community build into a
failed lookup.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): type Ducklake catalogs on the external instance cluster in the settings form

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): read a roles answer only for the workspace it was asked in

A fork and its parent each have their own roles on a data table of the same
name, so an answer stamped with the name alone settled the role from the
workspace the editor was acting on before.

Also derive the AI table creation flag from the data replaced into the editor:
data naming no data table left the flag on from before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep an instance database a settings save is waiting to name

Cleanup for a database whose setup failed took the lock first, read no user,
and dropped it while a save blocked on that same lock was about to commit a
reference to it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): check the workspace stamp in the default database selector too

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): read what a save racing instance-database cleanup committed

A transaction blocked on the lock may still roll back, so keeping the database
for it stranded one whose name then blocks every retry: it is let through and
its outcome read instead. The waiter query also matches this database's locks
only, since pg_locks spans the cluster.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): tell a waiting request apart from the workspaces using a database

Both callers render what cleanup returns as the workspaces that keep the
database, so a waiting request's pid read as one of them. Each now words that
case itself, and the give-up comment names where the kept name actually goes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): take the cleanup lock on a connection the pool cannot reclaim

A session lock outlives the future holding it, so a cancellation between
taking it and releasing it handed a locked session back to the pool, where
every later settings save waits on it. Detached, the connection closes when it
is dropped and the server releases the lock with the session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): close the cleanup connection on drop instead of detaching it

Detaching released the pool permit while the session stayed alive, so
concurrent cleanups waiting on their locks could open as many connections as
they liked. Closing on drop covers the same cancellation and keeps them
counted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref to 179d2454a063ee818a9387a1eafdd354a217a16c

This commit updates the EE repository reference after PR #802 was merged in windmill-ee-private.

Previous ee-repo-ref: c2e43f5b5ff753d339b70e232fd17b6ffecb054d

New ee-repo-ref: 179d2454a063ee818a9387a1eafdd354a217a16c

Automated by sync-ee-ref workflow.

* chore: update ee-repo-ref to fd5b8af748f2c985b13e18d9ea30894f3bd7e9a3

This commit updates the EE repository reference after PR #798 was merged in windmill-ee-private.

Previous ee-repo-ref: 3145e422d61d580f0a82804f075285c112879da0

New ee-repo-ref: fd5b8af748f2c985b13e18d9ea30894f3bd7e9a3

Automated by sync-ee-ref workflow.

* fix(datatables): classify a fork import's target from the resolution that built its connection

A second read could see the entry flipped to a resource and skip the reservation check while the
connection already built still reached the instance cluster.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): run external database lifecycle on the caller's transaction, one lock order for fork cleanup

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): check a fork copy's reservation on the locked transaction's connection

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): make and drop fork copies of external data tables on the external cluster

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): refuse a fork whose external copy something already names

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): check an external fork copy's uses before removing its entry, so child fork pointers still count

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): migrate external cluster fork reservations on workspace rename

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): configure the external instance cluster and pick its databases from the UI

Instance settings gets an External Postgres tab: the cluster's connection, the setup run
with its report, and the databases Windmill created there. Data table and Ducklake settings
offer the external_instance kind, with a picker over those databases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep the external cluster form out of the instance settings the page bulk-saves

The form seeded its key into the settings store after the page snapshotted them, so merely
opening the tab made the page send a setting the validator refuses, failing an admin's
unrelated save. The form is local state now, and only the setup writes the key.

The database picker also treats a superadmin-only listing as authoritative: a workspace
admin, who cannot list, no longer sees a saved database as missing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(tests): app paths are validated, so a guest app cannot be renamed to one with a space

Path validation on apps landed after this test, which still expected a space to be accepted
and created its fixture on a ':' path. The scopable-path guard it exists for is kept by
planting that path on the row, which is now the only way an app can hold one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): one Managed Postgres tab, and a switch for Windmill's own database

The instance settings tab covers both substrates Windmill administers: the external cluster,
which can now be disabled once nothing sits on it, and Windmill's own database, which an
operator can turn off so the cluster is the only one a workspace may newly name. A save that
names it then refuses, as the create endpoint does.

Workspace settings name them the way an admin meets them: Postgres Resource, and Managed
instance, qualified as Internal or External only while both are on offer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): offer the external cluster in the add-data-table wizard, and give Ducklake its room back

The wizard gets the external cluster as a fourth substrate: pick a database it already holds
or name a new one, which the run creates before writing the entry, so Try again does not trip
on a database the last attempt made.

Ducklake's maintenance column moves into the settings popover next to the extra args, which
the wider catalog select had squeezed the name box out of.

A setting a component saves itself now also moves the page's baseline, or Discard would
restore what it replaced and a later save would send that stale value back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): only a database this wizard run made makes its retry idempotent

A retry skipped the create whenever the name was registered, which let a run adopt a database
another superadmin had made and share their data without the warning the existing-database
branch shows. The run tracks what it created instead, and reports it so Discard can say the
database is still there. Creating one now refreshes the registry the next run's default name
and its validation read, and that default counts the external cluster's databases rather than
the instance's.

Row ownership carries the substrate: "instance" and "external_instance" can hold the same name,
and a row repointed between them while a run probes must not read as that run's own.

Also fixes the merge leftovers in the DuckDB executor's tests, which cargo check never compiles:
the new PgDatabase field and the attach helper's job directory.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): shorten the picked managed kind, and say what an empty database list means

The closed select has room for the qualifier, not the whole name, so a picked managed kind
reads "Managed (Internal)". An external entry gets the tooltip its internal counterpart has,
and a database picker with nothing in it says to type a name rather than "No items found".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep the wizard's external-database bookkeeping optional and parked

The run deps gained a required field, which every existing caller of runSetup -- the sibling
test suite included -- does not pass. It defaults to empty instead, the parked payload carries
it across a Supabase redirect as it claims to, and a resumed run counts it as something left
behind. Two regressions pin the create: made on a first attempt, skipped only for a database
this run made.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): review an external-cluster database as one, not as a Postgres resource

The review step fell through to the resource branch for the new provider, so it announced a
connection already in the workspace for a database that does not exist yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): show the server's refusal instead of an object in the toast

Saving data table settings passed the whole error to the toast, which rendered the request
and response as JSON, and the instance-database wizard replaced it with "check console".
Both surface the message the API sent, through the helper that exists for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: point at the EE commit dropping the unused import

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): define a role on the cluster its data table sits on

The permissions drawer created and listed roles without naming a cluster, so a role defined
from an external data table became a login on Windmill's own, and the refresh then replaced
the correct catalog with the internal one. The drawer passes the data table's cluster to both.

The wizard also read a database it had just created as a name collision, which blocked
returning to the failed attempt and finishing under the same name.

And the merge had left the CONNECT-grant pass and its log in both the caller and the callee.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): manage the external cluster's role catalog, and name the cluster in the copy

The roles drawer offered one catalog, so the external cluster's roles could only be reached
through a data table sitting on it — and became unreachable once the last one was gone, while
the cluster could not be unset until they were dropped. Opened from the page it now offers the
clusters the instance has; a data table's own drawer still pins its cluster, and its copy names
that cluster rather than "the instance".

A DuckDB attach also decides whether to verify certificates the way every other Postgres
connection does, so a resource carrying a root certificate is no longer downgraded to require
here alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): bind the roles list to the cluster it was read for

Switching catalogs left the previous cluster's rows on screen and applied whichever response
landed last, so a slower read could seat one cluster's logins under the other's heading while
every control acted on the wrong id — invisible where a name exists on both. The switch clears
the rows, a token discards a response the selection has moved past, and the controls stay inert
while a catalog is being read.

The drop confirmation also names the cluster whose databases it reaches.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref to dfe7b8b9204d21e0264fbea1c6f6eedf9e738d56

This commit updates the EE repository reference after PR #810 was merged in windmill-ee-private.

Previous ee-repo-ref: 30db1b33bda0446f5fc5dfe353fbb226d57a26d4

New ee-repo-ref: dfe7b8b9204d21e0264fbea1c6f6eedf9e738d56

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
2026-09-29 14:45:47 +02:00
hugocasaandClaude Opus 5.5 797147ea7a fix: report a worker's last job when it ran under one poll interval (#11399)
* fix: report a worker's last job when it ran under one poll interval

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: share the unreported job slot with the interactive worker shell

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: pin that a main-loop ping without a job keeps the last one

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 14:27:23 +02:00
hugocasaandClaude Opus 5.5 97fa55b719 feat: detect and alert when a schedule skips occurrences (#10917)
* docs: plan for detecting skipped schedule occurrences

Design plan only, no implementation. Records the scheduler's re-anchoring
behaviour, the measurements behind it, and the three-piece design that came
out of reviewing the alternatives.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* docs: state the user-facing outcome in the schedule plan

The plan described the mechanism but never what a user would see, which made it
hard to judge what the work is worth.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* docs: state which cause the schedule plan catches, and correct its scope

Records which of the two causes each piece covers, and corrects the overrun
scope: a script schedule carrying retry or dynamic_skip is pushed as a
SingleStepFlow, so it re-arms at step 0 entry and its occurrences overlap like
a flow's. Resolves the no_flow_overlap question, splits the read-time work into
bounded detection and editor-only counting behind measured croner costs, and
fixes the delivery order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: count the occurrences a schedule skipped

A schedule that overruns its interval, or waits for a worker, silently loses
the occurrences in between: the scheduler keeps one queued occurrence and
re-anchors on the clock, so nothing records that a run was due and never
happened.

Recovers the sequence from rows that already exist rather than writing per
occurrence. `push_scheduled_job` anchors on `now_from_db` inside the
transaction that inserts the job, and `v2_job.created_at` defaults to that
same transaction timestamp, so `scheduled_for = find_next(created_at)` holds
exactly and the whole occurrence history is derivable.

The schedules list reports how many of the recent runs were followed by a lost
occurrence, and a new occurrences endpoint carries the per-run wait and
duration behind it. Detection is one `find_next` per gap, which stays bounded
on a full page; counting walks the gap and runs only for a single schedule.

The one write is `occurrence_baseline_at`, advanced at create, edit,
re-enable and re-arm. Gaps older than it span a pause, a cron change, a
re-enable or a reconciler re-arm, none of which mean runs were lost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: show the wait and run time behind a schedule's skipped occurrences

The list badge says a schedule is losing runs; this says which of the two
causes did it. A large wait means not enough workers, a long run means the job
outgrew its interval, and the pair is what tells them apart. Sits under the
existing upcoming-events panel, so due and overdue read together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: flag a schedule that is running late right now

Reconstruction is retrospective: a gap only appears once the next occurrence
has a row, which needs the current one to finish. A schedule wedged mid-run
shows nothing until it moves, which is the case an operator most wants to see.

An occurrence still in flight past the time its own successor was due will
cost that successor, so `now > find_next(scheduled_for)` is the signal, needing
no threshold and self-calibrating across a daily and a per-minute schedule. It
applies only where occurrences serialize; an overlapping schedule starts its
successor on time and would flag constantly while healthy.

The queue is read in one aggregating pass keyed on (trigger, runnable_path)
rather than a subquery per schedule, and an overlapping schedule holds more
than one root row, hence the aggregate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: run the schedule overrun alert from the monitor pass

Wires `schedule_overrun_alerts` in next to `jobs_waiting_alerts`, every 30
iterations (~5 min). Its Enterprise implementation lives in
windmill-labs/windmill-ee-private#772; only the wiring and the OSS stub are
here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* chore: refresh the sqlx offline cache

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: record and alert when a schedule skips occurrences

push_scheduled_job compares each chained occurrence with the slot after
the previous one. A gap is written to schedule.skipped_occurrences off the
push transaction, alerts once when a clean schedule starts skipping, and
recovers on the next clean chain. The schedules list shows a badge.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* feat: alert only on a streak of skipping runs, keep a recent skip visible

The skip state now describes the current streak and is written in the push
transaction, so it commits or rolls back with the push. The alert fires
once when 3 runs in a row skipped, and the list keeps a muted badge for 7
days after the latest skip. Editing or toggling a schedule resets it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* fix: name the missed-occurrence state after what it counts, alert only once committed

Renames the columns to late_run_streak, missed_occurrences and
last_missed_at, keeps the missed count after a streak ends so the muted
badge can show it, and rewords both badges. The alert task now reads the
streak FOR SHARE, which waits for the push transaction, so a push that
rolls back and retries alerts once. A failed slot count leaves the streak
untouched, and a schedule deleted mid-push no longer fails it. Adds an
integration test for the streak and its reset.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* fix: recover the late run alert, store the missed slot, name it missed throughout

The alert now recovers (and so acknowledges itself) when a streak that
alerted ends on a run on time, under the schedule:{path} resource used by
the other trigger alerts. last_missed_at records the last missed cron slot
rather than when the late run chained, and the counting helpers say
missed, since skipped already names occurrences queued and not run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* fix: scope the late run alert to its workspace, acknowledge it on edit, toggle and delete

Recovery acknowledges alerts by resource alone, so the resource now
carries the workspace. Editing, toggling or deleting a schedule clears
its streak and a disabled or deleted one never chains a run on time, so
those handlers acknowledge its open alert after committing. Past the
1000-slot cap, last_missed_at falls back to the detection time.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

* refactor: raise the late run alert like the other critical alerts

Drops the recovery, the workspace-scoped resource and the acknowledgement
on edit, toggle and delete: the alert now fires once per streak with no
resource and is acknowledged from the alerts feed, as the trigger and job
failure alerts are. The FOR SHARE read stays, so a push that rolls back
across the flow path's retries still alerts once. Notes in openapi that
past 1000 misses in one late run the count is a floor and the time
approximate.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-29 12:40:25 +02:00
17448c97d3 count trigger suspend, resume and discard (#11405)
* feat(telemetry): count trigger suspend, resume and discard

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor: shorten the telemetry disclosure to one line per category

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: correct the resume branch comments and note the pre-commit fire count

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: name feature adoption in the telemetry disclosure

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: update ee-repo-ref to 9855e1b7a43a0a33e04f8accf1c497af3fd9b139

This commit updates the EE repository reference after PR #834 was merged in windmill-ee-private.

Previous ee-repo-ref: 1d5b128ec956c156fe549cf099ba0dbc1b6467bf

New ee-repo-ref: 9855e1b7a43a0a33e04f8accf1c497af3fd9b139

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-09-29 12:26:11 +02:00
GuilhemandClaude Opus 5 f367eaf6d0 feat: run turns in several flow chat conversations at once (#11202)
* feat: run turns in several flow chat conversations at once

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep finished turns finished and cached chats current in the flow chat pool

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: attribute a turn's rows by job id as well as sequence

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: count only real stream updates and retry the job-id read

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep a chat that holds an unsent draft, and take one back when its first turn is withdrawn

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: ignore a stale running-turn snapshot, keep a withdrawn chat's draft, poll after clean stream ends

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: follow the turn running now when the listing named one already over

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep replacement turns and SSE fallback moving

* fix: keep replacement turn handoffs active

* fix: preserve unread badge line height

* fix: settle local fallback handoffs

* fix: settle refused turn handoffs

* fix: scope turn handoffs to conversation

* fix: drop stale turn handoffs

* refactor: move the queued message and 409 handling into per-conversation turns

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address cubic's review of the parallel flow chat turns

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: clear a stale failure on refresh, and tighten the docs and test waits

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: recover running rows past the first page, and drop the failure a re-read disproves

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: check the running-turn query at compile time, and narrow what a refresh clears

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: settle a failed turn only from an answer that turn wrote

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: settle a failed turn from its own answer, and only while it is still the failure shown

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: drop a failure whose answer arrived even when a newer turn owns the error

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: free an answered failure whatever the turn that started meanwhile is doing

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: drop a rows read that a turn outran, rather than merging it under newer messages

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: drop a rows read whose conversation was left and opened again

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: hand over a file still being read when its composer goes

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: count a drop's routing as work in flight, so its file is handed over too

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: hold the send until every file a conversation is owed has landed

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: keep a panel mounted per conversation instead of handing its draft over

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep the withdrawn chat whose composer was written in, not the empty one

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep the chat in front of the reader when both withdrawn composers were written in

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep a retry's own run arguments when a turn elsewhere refuses it

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: name panels apart across pools, and read a flow's inputs when its chat is built

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-29 11:36:30 +02:00
GuilhemandClaude Opus 5.5 e68ff0d969 feat: add tree view to the schedules page (#11360)
* feat: add tree view to the schedules page

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep schedule job previews visible inside tree folders

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep job preview loading while hovered and close it on scroll

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep schedules outside u/ and f/ in the tree view

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: add tree view to the trigger list pages (#11400)

* feat: add tree view to the trigger list pages

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor: let a TreeViewState own the tree view setting

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: drop the doubled bottom border at the end of a tree

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 10:54:46 +02:00
Ruben Fiszelandrubenfiszel 651a6b01f5 chore(main): release 1.819.0 (#11364)
* chore(main): release 1.819.0

* Apply automatic changes

---------

Co-authored-by: rubenfiszel <275584+rubenfiszel@users.noreply.github.com>
v1.819.0
2026-09-29 08:33:39 +02:00
Ruben FiszelandClaude Opus 5.5 cedd6dc901 fix: hold interpolated references and captures to the token path scopes (#11391)
* fix: hold interpolated references and captures to the token path scopes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: let a resource read cover its own linked secret variable

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: resolve policy-granted app upload resources on the viewer's rls

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: cover multi-secret linked variables and keep capture paths out of refusals

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 18:34:55 +02:00
Ruben FiszelandClaude Opus 5.5 163a4ffa4e fix: scope flow resume to its workspace and minting to the job's run (#11392)
* fix: scope flow resume to its workspace and minting to the job's run

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: name the lineage columns resume minting checks

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 18:31:44 +02:00
Ruben FiszelandClaude Opus 5.5 14a2619ad2 fix: gate batch rerun on job read access, scope started_at to workspace (#11387)
* fix: gate batch rerun on job read access and scope started_at lookup to the workspace

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: assert batch rerun denial comes from the read gate

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 18:29:58 +02:00
Ruben FiszelandClaude Opus 5.5 3eaf2888c0 fix: scope workspace dependencies create to the path workspace (#11385)
* fix: scope workspace dependencies create to the path workspace

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: state the workspace_id must-match contract in the spec and struct

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 18:15:10 +02:00
AlexRV12andClaude Opus 5.5 5e59cefd1f fix(frontend): apply operator write locks from the session's operating workspace (#11395)
Claude-Session: https://claude.ai/code/session_01VAj4mmm2YThZLVkivrgsbb

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 18:13:06 +02:00
Ruben FiszelandClaude Opus 5.5 f4dcaf3e45 fix: list only the paths the caller can read in path autocomplete (#11388)
* fix: list only the paths the caller can read in path autocomplete

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: bound the path autocomplete cache by total path count

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 17:55:16 +02:00
Ruben FiszelandClaude Opus 5.5 ec6ec1b06f fix: judge IPv4 embedded in IPv6 and pin the object storage test connect (#11389)
* fix: judge IPv4 embedded in IPv6 and pin the object storage test connect

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: let the public-only object store client reach the egress proxy

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: refuse private IP literals and the proxy host in the public-only store client

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: refuse the egress proxy as a target whether named or an IP literal

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 17:43:15 +02:00
Diego Imbert 22b5a1cd62 fix: keep test panel controls off the args form in debug mode (#11382) 2026-09-28 17:28:40 +02:00
Diego ImbertandClaude Opus 5.5 d76a962331 feat: add hub sync button to the resource types tab (#11375)
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 17:24:20 +02:00
Diego ImbertandClaude Opus 5.5 d9c7d71f07 fix: redesign run not found page and fix switching to the right workspace (#11374)
* fix: redesign run not found page and clear stale not-found on workspace switch

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor: use design-system Button for workspace rows on run not found page

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 17:23:59 +02:00
hugocasaandClaude Opus 5.5 60ef82196e fix: keep smtp_clicktracking_off when syncing instance config (#11372)
* fix: keep smtp_clicktracking_off when syncing instance config

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: name smtp regression test after what it guards

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 17:18:48 +02:00
990a726409 feat: add provenance claims to job OIDC tokens (#11369)
* feat: add provenance claims to job OIDC tokens and mark preview sub

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: require a flow or script job's version to belong to its path for deployed

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: derive app script paths server-side and test job provenance in CE

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: count an app script as deployed only when a deployed app run stamped it

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: state the deployed condition for the preview sub prefix

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: keep the plain OIDC sub for previews by users who can write the path

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: refuse OIDC tokens to previews by users who cannot write the path

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor: keep OIDC token issuance unchanged, leaving provenance to the claims

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor: name the root job's trigger claim root_trigger_kind

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref to 421cf2a8b4f98b421e93c0fc7c1c378314a66e50

This commit updates the EE repository reference after PR #831 was merged in windmill-ee-private.

Previous ee-repo-ref: 7acd384875deba4b01a502e628a153b11c82eecb

New ee-repo-ref: 421cf2a8b4f98b421e93c0fc7c1c378314a66e50

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-09-28 11:42:58 +02:00
Diego ImbertandClaude Opus 5.5 3838cd6ee0 fix: stop the schedule enabled toggle from showing unsaved changes (#11390)
* fix: keep schedule enabled toggle from reading as unsaved changes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: only fold the enabled toggle into the baseline when it is deployed

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: skip the enabled revert once the drawer moved to another schedule

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 19:50:02 +02:00
Ruben FiszelandClaude Opus 5.5 893e64f630 fix: only restart a flow on a version of its own path and workspace (#11376)
* fix: only restart a flow on a version of its own path and workspace

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: pin cross-workspace restart version rejection

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 14:10:21 +00:00
Ruben FiszelandClaude Opus 5.5 90f9e59321 fix: only let a job's own token claim run lineage (#11367)
* fix: only let a job's own token claim its lineage on the run endpoints

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: drop an unclaimable run lineage instead of refusing the run

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: only let a job's own token run its workflow-as-code tasks

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 14:09:59 +00:00
Diego Imbert 47525b211a perf: shrink the module graph that gates first paint in dev (#11373)
* perf: shrink the module graph that gates first paint in dev

* docs: drop the stale synchronous-icons claim on the import card

* fix: replay search opened before its modal loads, guard lazy icons

* fix: only intercept search before load where the modal mounts

* perf: mount app-shell modals on first open and keep monaco off the shell

* perf: load the icon map on first read, not at module evaluation

* fix: report stale chunks with a reload toast, guard the home page against monaco
2026-09-27 10:09:45 +00:00
Ruben FiszelandClaude Opus 5.5 649c43e7c1 fix: run an AI agent tool on the worker its own tag selects (#11370)
* fix: run an AI agent tool on the worker its own tag selects

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: run a tagged agent tool inline when this worker serves its tag

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: give an inline agent tool a job token of its own

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: report a lost tool wait to the model and cancel tools on agent timeout

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-26 17:32:25 +00:00
Ruben FiszelandClaude Opus 5.5 d7a61de23f fix: never double-process a slow canceled flow in the zombie sweep (#11368)
* fix: leave a slow canceled flow to its live worker and bound its requeues

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF

* fix: give a canceled zombie flow a longer grace instead of guessing its worker

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF

* fix: retry a canceled zombie flow's forced completion until it lands

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF

* fix: claim a canceled zombie flow without waiting on its runtime row

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-26 16:18:42 +02:00
Ruben FiszelandClaude Opus 5.5 c2d8997549 fix: complete a canceled flow whose worker died between two steps (#11366)
* fix: complete a canceled flow whose worker died between two steps

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF

* fix: complete only the stranded canceled flow and let its parent process it

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF

* fix: requeue a stranded canceled flow for a worker to complete its cancel

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF

* fix: keep a requeued canceled flow's start time

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-26 14:40:01 +02:00
Ruben FiszelandClaude Opus 5.5 65cba2dbb7 perf: advance a flow step with one v2_job_status update (#11357)
* perf: advance a flow step with one v2_job_status update

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: state what advance_flow_status returning None means

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep the merged flow advance identical for rows without a status row or with a malformed status

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FbA4shbpCnRsUxir1GEfm

* docs: note the JSON null invariant behind the empty-path no-op

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FbA4shbpCnRsUxir1GEfm

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-26 00:18:59 +02:00
bebd762194 perf: complete a job in one statement on the common path (#11355)
* perf: complete a job in one statement on the common path

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: take completion locks in one order on every path

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep a losing zombie completion from touching its wac parent

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: leave a flow's ping alone when a step completes during its cancel

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: probe only this test's completion for the lock wait

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: stamp a wac child's kept duration when its completed row exists

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TJMJSJ2bDhYh7Shoh78Yyb

* perf: leave the parent ping out of completions with no flow to ping

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TJMJSJ2bDhYh7Shoh78Yyb

* docs: note that the two completion statements must stay in step

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TJMJSJ2bDhYh7Shoh78Yyb

* chore: update ee-repo-ref to 7a256cf353db7cf64a60a09fa0de7f3a8b27f626

This commit updates the EE repository reference after PR #830 was merged in windmill-ee-private.

Previous ee-repo-ref: 497137acb65e521568d46f3cbe1d66359f7f87ec

New ee-repo-ref: 7a256cf353db7cf64a60a09fa0de7f3a8b27f626

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-09-26 00:18:38 +02:00
Ruben FiszelandClaude Opus 5.5 e2be584ca5 fix: let custom workspace error handlers send email with the instance SMTP (#11365)
* fix: let custom workspace error handlers send email with the instance SMTP

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: state what the error handler email allowlist guarantees

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 22:18:40 +02:00
Ruben Fiszelandrubenfiszel 3974bbeac6 chore(main): release 1.818.0 (#11294)
* chore(main): release 1.818.0

* Apply automatic changes

---------

Co-authored-by: rubenfiszel <275584+rubenfiszel@users.noreply.github.com>
v1.818.0
2026-09-25 18:42:49 +00:00
hugocasaandClaude Opus 5.5 ca8a04a869 fix: allow results access inside nested functions in input transforms (#11358)
* fix: allow results access inside nested functions in input transforms

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: decode escaped bracket step ids and test deferred fetch errors

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: let quickjs decode bracket step ids and match quoted forms

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: prefetch results read through spread syntax

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: keep prefetched bracket literals on a single line

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: decode prefetch step literals as data and skip unparsable ones

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: run the results prefetch outside the expression scope

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: keep the transform expression a zero-arg iife after prefetch

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 18:41:52 +00:00
Ruben FiszelandClaude Opus 5.5 da866c5eff feat: alert on and optionally cancel jobs stuck on unserved tags (#11354)
* feat: alert on and optionally cancel jobs stuck on unserved tags

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: group stranded jobs in sql and recheck each job before canceling

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: guard stranded-job alerts and cancels against outages and pickups

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: count priority tags as served and retry lost stranded-job cancels

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: send one daily stranded-jobs alert that can be muted

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: cover native retries and finish lost stranded-job cancels unconditionally

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 17:48:21 +00:00
Ruben FiszelandClaude Opus 5.5 30bb62cd25 fix(cli): stub the API client over its real exports in tests (#11363)
* fix(cli): stub the API client over its real exports in tests

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cli): mock the API client once and dispatch to per-suite stubs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 17:43:08 +00:00
9a1c6e5081 feat: let a workspace withdraw operator schedule and trigger writes (#11226)
* feat: let a workspace withdraw operator schedule and trigger writes

Operators can create, edit and delete schedules and triggers today through the
API, CLI and MCP, while the operator_settings flags beside them only hide those
pages. An admin who wants operators to see what is scheduled without letting
them change it cannot express that. Add manage_schedules and manage_triggers as
enforced settings, gated at the schedule handlers and at the generic TriggerCrud
routes so every trigger kind is covered by one check.

They name capabilities operators already hold, so they are granted unless
withdrawn, and absence has to mean "never configured" rather than a value. The
read coalesces to true; the update endpoint merges into the stored jsonb with
the two fields as Option<bool>, so an omitted key keeps what is stored.
operator_settings is git-synced as a whole object, so a settings file written
before these keys existed reaches the endpoint on every pull, and a serde or SQL
default of either polarity would turn that pull into a silent withdrawal or
restoration.

The rights are read through a per-process cache, so withdrawing one publishes a
notify_event that drops the entry on every replica.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dsf6VC4MVLisiEoeQkgbr4

* feat: enforce operator write rights on the router and in the UI

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: close the capture gap and gate the trigger editors' write actions

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: gate acl writes and the native trigger drawer behind manage rights

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: refuse operator writes with 403 and gate sharing at the drawer

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf: resolve identity in the operator write gate only for writes

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: gate the suspended-jobs actions and stop the route check refusing reads

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: explain the empty-state create button when operator writes are withdrawn

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: audit operator settings changes and fold path writes into native rows

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: open locked editors read-only and group the operator settings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: skip email and azure lookups on editor open while triggers are locked

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: state each operator-rights rationale once in comments

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: address CI review findings on operator write rights

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep capture move gated and skip it in the builders while locked

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep admin and operator exclusive when setting a workspace role

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor: use the shared section component for operator settings groups

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
2026-09-25 17:34:58 +00:00
Ruben FiszelandClaude Opus 5.5 a1abb36d9f fix: relock importers on their own tag, not the bare dependency tag (#11359)
* fix: relock importers on their own tag, not the bare dependency tag

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: pin the tag of relocks triggered by a changed import

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 17:32:25 +00:00
Ruben FiszelandClaude Opus 5.5 53a5cfd17a perf: skip job-start pings and checkpoint read for short non-WAC jobs (#11356)
* perf: skip job-start pings and checkpoint read for short non-WAC jobs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: pin the wac language gate alongside is_wac_v2

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep the start memory sample for jobs shorter than one poll tick

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 18:39:08 +02:00
e14da5c6bc feat(bedrock): add OIDC role assumption as a fourth auth mode (#10936)
* feat(bedrock): add OIDC role assumption as a fourth auth mode

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn

* fix(bedrock): gate the OIDC cache correctly and assume the role once per job

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn

* refactor(bedrock): check the OIDC region before minting a token

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn

* fix(bedrock): keep OIDC session names collision-resistant, gate the copy on EE

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn

* fix(bedrock): check the OIDC region before reusing cached credentials

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn

* fix(bedrock): clear assumed-role sessions when AI settings change

invalidate_ai_request_cache_for_workspace cleared AI_REQUEST_CACHE only, so a
workspace's AI settings edit reset one cache and left the assumed-role sessions
keyed on the old config in place until STS expired them.

Also name the region requirement in the credentials-check hint, so following it
does not land on the OIDC path's region guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn

* chore: update ee-repo-ref to de73db2bacfdc3eaa2e63b1827178bc198d54e5c

This commit updates the EE repository reference after PR #770 was merged in windmill-ee-private.

Previous ee-repo-ref: c43dab1e69b1cb3f685e6df07bff634dc2a0b734

New ee-repo-ref: de73db2bacfdc3eaa2e63b1827178bc198d54e5c

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-09-25 18:19:38 +02:00
hugocasaandClaude Opus 5 94e4fb1c84 fix: carry labels when deploying variables, resources and folders (#11222)
* fix: carry labels when deploying variables, resources and folders across workspaces

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: clear a folder's labels in the target when the source has none

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: pin that the frontend deploy adapter carries variable labels

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-25 18:15:08 +02:00
hugocasaandClaude Opus 5 d3d5392917 feat: add an options field to the postgresql resource (#11223)
* feat: add an options field to the postgresql resource

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: keep a literal plus in postgres connection string parameters

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: pass postgres options to trigger connections

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: read DATABASE_URL options the way sqlx does

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: include postgres options in databaseUrlFromResource

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-25 17:40:13 +02:00
Alexander PetricandClaude Opus 5.5 f183bd43fb chore: retire the standalone lsp and multiplayer images from examples, drop lsp/Dockerfile (#11341)
* chore: retire the standalone lsp and multiplayer images from examples, drop lsp/Dockerfile

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* chore: drop the unbuilt DockerfileMultiplayer, document running the LSP from windmill-extra

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(examples): ecs terraform destroys cleanly and gives private instances no public ip

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(examples): give windmill-extra on ecs a WINDMILL_BASE_URL for multiplayer auth, address review

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(examples): make the ecs example upgrade cleanly from the standalone lsp/multiplayer stack

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(examples): name the extra target group by prefix so create_before_destroy can replace it

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(examples): note the brief editor-socket gap when upgrading the ecs example

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(examples): the debugger stays off after the ecs upgrade unless enabled

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 17:13:06 +02:00
hugocasaandClaude Opus 5 3bd89e92d8 feat: collapse the fork members setting and show its state in a badge (#11220)
* feat: collapse the fork members setting and show its state in a badge

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: show the fork members badge next to the title, only when on, like other section badges

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-25 17:12:51 +02:00
Alexander PetricandClaude Opus 5.5 8caf414301 fix(multiplayer): a malformed frame from one client no longer exits the server (#11353)
* fix(multiplayer): don't drop client messages during cold-start token verification

`wss.on('connection')` awaits `verifyToken()` before `setupWSConnection()`
attaches the 'message' listener. On a cold process that await includes the
first `/api/debug/jwks` fetch (~30ms on ECS). A y-websocket client sends sync
step 1 the instant the socket opens, and `ws` drops messages emitted with no
listener attached, so that step 1 was lost and never answered with step 2 —
the client's provider never became `synced`.

Buffer messages from the moment the connection is accepted and replay them, in
order, once `setupWSConnection()` has installed its handlers. Rejected
connections drop the buffer and close with the same 4401/4403 codes as before.

Also prefetch the public key at startup when WINDMILL_BASE_URL is set. That is
insurance, not the fix: a connection arriving before the prefetch resolves
still relies on the buffer.

Adds `npm test` in multiplayer/ (node:test, no docker or backend needed) with a
fake JWKS endpoint that answers with a delay, which holds the cold window open
and makes the race deterministic; wired into the existing test_extra CI job.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(multiplayer): cap what an unauthenticated peer can buffer pre-auth

Review follow-up.

The pre-auth buffer was unbounded: `ws` sets no `maxPayload` here and the JWKS
fetch has no timeout, so a peer that never authenticates could stream frames
into memory for as long as `verifyToken` was stalled. Cap it at 32 frames /
1 MiB — a real client only has sync step 1 and its first awareness update in
flight there — and close 1009 past that, dropping what was buffered.

A socket closed during verification (by the peer, or by that cap) is no longer
handed to setupWSConnection: it would be added to `doc.conns` with a 'close'
listener that can never fire.

The startup prefetch's .catch was dead code — getPublicKey() logs its own
failures and resolves to null rather than rejecting.

Test helper: pin REQUIRE_SIGNED_MULTIPLAYER_REQUESTS and BASE_INTERNAL_URL so an
ambient value cannot turn the rejection tests into false passes; bind the JWKS
server on port 0 instead of a released probe port, and retry the spawned server
on EADDRINUSE; destroy still-delayed JWKS responses on teardown, since
server.close() waits for in-flight requests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(multiplayer): gate the JWKS response instead of delaying it

Review follow-up.

The cold window was held open by a 1500 ms delay on the fake JWKS response, but
that timer started when the startup prefetch reached the fake server, not when
the client sent its first frame. A slow enough machine could load the key before
the client connected, and the race test would then pass without ever exercising
the buffer — a false pass.

The fake JWKS server now parks every response until the test calls release(), so
the server provably holds no key while the client is sending. The race test
releases only after both frames are written to the socket, and asserts the
server has not logged the key as loaded at that point; the flood test never
releases until after the cap has closed the connection.

What is left to wall-clock time is 250 ms for bytes already written to the socket
to cross loopback into an otherwise idle server, rather than a window that had to
cover process startup, connect and handshake.

Also drops the prefetch precondition from the forged-token and flood tests so
each test still maps to one behaviour. Suite runs in ~1.1s instead of ~5.3s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(multiplayer): survive a malformed frame instead of exiting the process

`setupWSConnection`'s message handler decoded whatever an authenticated peer
put on the wire with no guard: `decoding.readVarUint`,
`syncProtocol.readSyncMessage` and `awarenessProtocol.applyAwarenessUpdate` all
throw on input they cannot parse, `ws` re-emits a listener's exception on the
process, and server.mjs installs no `uncaughtException` handler. One bad frame
from one client therefore killed the whole multiplayer server, taking every
other document and every other client with it.

Catch decode/apply failures, log the document, the client address and the error
message (never the payload), and close only the offending connection with 1007
"invalid frame payload data". Frames that arrive once a connection is no longer
OPEN are ignored, so the replay of the pre-auth buffer stops at the first
refusal instead of applying the rest.

docker/entrypoint-extra.sh made that outage permanent: on a service exit it
logged a bare PID and then `wait`ed on the rest, so the container stayed up with
a dead service and the health checks in front of it — which probe the LSP — saw
nothing wrong. It now names the service that died, stops the others through the
same shutdown path SIGTERM uses, and exits non-zero so the orchestrator replaces
the container. The "no services enabled" branch still sleeps.

Tests: multiplayer/test/malformed_frame.test.mjs covers four malformed payloads
from an authenticated client and one replayed out of the pre-auth buffer,
asserting the 1007 close, a live server process, an undisturbed bystander and a
real edit still propagating. docker/test_entrypoint_extra.sh runs the real
entrypoint in a container with stub services.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(multiplayer): prove the replayed malformed frame really is buffered pre-auth

Assert the server has not yet logged the loaded key when the frame is written,
and give it the same in-flight margin as the cold-start tests before releasing
the JWKS response, so the frame provably goes through the replay path rather
than landing on an already-authenticated connection.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci(extra): run the entrypoint supervision tests in publish_extra

The multiplayer unit tests already run there; the entrypoint test needs only
docker and the checkout, so run it in the same job, before the image build.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(multiplayer): handle WebSocket protocol errors and bound the shutdown

Two crash paths of the same class as the malformed-frame one, from review.

`ws` fails a frame it cannot parse at the protocol level — an unmasked frame
from a client, a reserved opcode, a bad RSV bit — inside its Receiver, before
the application 'message' handler ever sees it, and `receiverOnError` ends with
`websocket.emit('error', err)`. With no 'error' listener that is an unhandled
EventEmitter error, so it exited the process just as a malformed payload did.
(A raw socket error such as ECONNRESET does not: ws 8.21.3's `socketOnError`
swallows those.) Add the listener on the accepted socket, before authentication
so the pre-auth window is covered too, and one on the server.

Log messages now go through `describeError`, which collapses whitespace and
truncates, so nothing that reaches an error message can forge or flood a log
line.

`stop_services` ended in a bare `wait`. On the `docker stop` path dockerd
provides the deadline; the "a service died" path signals itself, so a service
that is wedged or slow to honour SIGTERM would hold the container open
indefinitely — the state that path exists to prevent. Bound it: SIGTERM, wait
SHUTDOWN_GRACE_SECS (10 by default), then SIGKILL the stragglers by name.

Tests: multiplayer/test/socket_error.test.mjs (authenticated and pre-auth
illegal frames, asserting a live process and continued service), a
SIGTERM-ignoring stub scenario in docker/test_entrypoint_extra.sh, and that
harness is now bounded throughout — `timeout -k` on foreground runs, a watchdog
around the backgrounded ones, and an optional outer timeout on `docker run`.
`--entrypoint bash` so the documented windmill-extra:test override runs the
harness instead of the image's real entrypoint. The stubs publish a readiness
marker and the dying one waits for them, removing a startup race in the harness.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(multiplayer): never log peer bytes, and keep a fatal server error fatal

Two review findings on the previous commit, both mine to answer for.

`describeError` collapsed whitespace, which is not enough. An error message is
not always a fixed string: `applyAwarenessUpdate` runs `JSON.parse` on the
peer's bytes and V8 quotes ~30 bytes of the offending input back verbatim, ESC
included, so a peer could put terminal escapes and forged content into a log
line. Strip everything outside printable ASCII instead, and say so where the
comment previously claimed the messages were fixed strings.

`wss.on('error')` was worse than the crash it replaced for one case: `ws`
forwards the HTTP server's errors there, so a failed listen (EADDRINUSE) was
logged and the process then exited 0 — a clean shutdown as far as anything
upstream could tell. It now sets a non-zero exit code. Setting `process.exitCode`
rather than calling `process.exit()` keeps the log line from being truncated.

`openClient` in the test helpers now records the socket error it was already
swallowing, so a failed connection reports its cause instead of surfacing as a
bare `waitFor` timeout.

Tests: a malformed awareness frame whose state is `x\x1b[2J OWNED THE LOG` added
to the payload table, with every case now asserting exactly one refusal line and
no control characters in it (1 fail before, 0 after, 3 runs); and a server that
cannot listen must exit non-zero (1 fail before, 0 after, 3 runs). 14/14 on 5
consecutive runs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(multiplayer): make the exit-status helper robust to spawn and stdio races

Review follow-ups on runMultiplayerServerUntilExit.

Wait for 'close', not 'exit': 'exit' fires when the child terminates, which can
be before its stdio pipes are drained, and the caller reads the output. On the
EADDRINUSE path the child writes one line and exits immediately after, which is
exactly the shape that loses it.

Listen for 'error' too. A child that fails to spawn emits neither 'exit' nor
'close', so the promise would never settle and the SIGKILL guard could not help.

Report whether the guard fired, rather than leaving the caller to infer it from
the exit signal: `signal` is null for every child exit on Windows, so a server
that hung after the listen error would have looked like one that exited on its
own. The test asserts on that flag instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(multiplayer): assert the refusal line itself, and name hasSyncType for what it takes

Two review nits on the test helpers.

The `doc="..."` assertion searched the whole server log, where CONNECT and
DISCONNECT also name the document, so it would have passed even if the refusal
stopped naming anything. Every assertion about the refusal is now made against
the refusal line, which the test already isolates, and it also checks the peer
is named.

`hasKind` took a sync sub-type but was named as if it took any message kind, and
the two families overlap numerically (`syncStep1 === messageSync === 0`), so a
caller passing the wrong one got a silently wrong answer. No runtime check can
tell aliased numbers apart, so the fix is the name: `hasSyncType`, with the
overlap spelled out where the constants are declared.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(multiplayer): make the pre-auth tests prove which path they took

The replay test could not establish that its frame went through the pre-auth
buffer: a frame delivered after setup, into the live handler, produces the same
close code and the same refusal line, so if the in-flight margin were ever
missed the test would quietly become a duplicate of the main-loop cases rather
than fail. server.mjs now logs REPLAY when, and only when, it replays a buffered
pre-auth message — worth having on its own, since that path only runs when a
client beat the JWKS fetch on a slow-starting instance — and the test asserts on
it. Removing that log line turns the test red, which is the point.

The socket-error pre-auth test gated on `jwks.requests >= 1`, which the startup
warm-up already satisfies, so it proved nothing about the offender. What makes
it the pre-auth case is that the JWKS response stays parked for the whole test;
it now asserts the server never logged CONNECT, which is exact.

`killedByTimeout` was set before the kill, so a child that exited on its own just
before the timeout — with 'close' still pending on the stdio drain, the very
window this helper waits for — would have been reported as killed. It now claims
the rescue only when there was a live process to signal.

14/14 on eight consecutive runs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(multiplayer): document the frame-recording contract in openClient

ws hands every frame over as a Buffer under the default binaryType, text frames
included, so recording them as Uint8Array is lossless for both. Worth stating:
ws 7 delivered text frames as strings, where new Uint8Array(string) would have
been a silent zero-fill, and the difference is not visible at the call site.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(multiplayer): pin the close code for an illegal frame, and trim comments to the 4-line rule

Both socket-error tests waited for a close and never checked what it was, so an
abrupt 1006 teardown would have passed while the comment beside the payload
claimed 1002. `ws` sends 1002 for an unmasked frame in both the authenticated
and pre-auth cases, confirmed over repeated runs; that is now a named constant
asserted in each test, mirroring malformed_frame.test.mjs. Changing the expected
value turns both red.

The comments added by this branch also ran past the four lines AGENTS.md allows,
and several justified the change to a reader rather than stating the invariant.
Condensed to the invariant, at the site that would break it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 17:12:47 +02:00
hugocasaandClaude Opus 5 1d119e6b25 fix: correct the tool controls and name field of a nested AI agent (#11221)
* fix: hide tool controls a nested AI agent cannot use

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor: drop the tool navigation prop an agent tool now implies

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: ask for a tool name on every kind of agent tool

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: keep enabled_tools reachable on a linked nested agent tool

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: name the tool kinds the header's name field actually renders

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: state the tool-name rule without listing the kinds it covers

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-25 17:12:28 +02:00
Diego ImbertandClaude Opus 5.5 4fd7d62bd0 feat: rework the db manager: native grid, tabs, sql editor, joined columns (#11340)
* feat: replace ag-grid in the db manager table viewer with a native grid

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: add user-managed data, diagram and sql editor tabs to the db manager

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: add schema autocomplete to the db manager sql editor and polish its layout

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: add joined foreign key columns and draggable tabs to the db manager

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: move db manager tabs on drop instead of mid-drag

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: open followed foreign keys in a new db manager tab

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: scroll large tables, drop stale joins and return to the last tab on close

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: tolerate the db manager tabs going away while switching data table

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep saved joined columns while metadata loads, test joins in every dialect

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: put the db manager grid on the input surface in dark mode

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: address db manager review nits on joins, boolean keys and sql quoting

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: keep the dark border color on the left pinned column edge

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: tone down the db manager tree menu icons

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: write json and jsonb values from the db manager through a text cast

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: dim unrelated diagram tables less on hover

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: match json columns as text when deleting a db manager row

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: compare postgres json columns as text in exact db manager filters

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: pick the columns the db manager grid shows

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: toggle every db manager column from one checkbox

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: word joined columns as a view in the db manager

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: keep data table roles and access visible when unavailable, with the reason

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: keep data table and instance roles visible in settings when unavailable

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: hide a db manager column from its header menu

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: format db manager columns with a unit, significant digits and color rules

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: drop the before/after hints from the db manager unit picker

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: write the euro after the amount and keep units off non-numeric values

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: color rule presets, bold and italic, layered and reorderable rules in the db manager

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: decimals, thousands separator, compact notation and alignment in db manager column formats

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat: compact db manager format controls, a notation toggle group and a reset button

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: db manager format and filter nits

Keep the value's scale when decimals are auto, read boolean color-rule conditions as
booleans, let a rule's text color reach foreign-key links, close the formatter when the
columns picker opens, and filter BigQuery complex columns through TO_JSON_STRING.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 15:58:25 +02:00
Alexander PetricandClaude Opus 5.5 2e5f6d0e76 fix(frontend): keep the session when the persisted workspace is stale (#11344)
* fix(frontend): keep the session when the persisted workspace is stale

A single-use login link signs a different account in while `workspace` in
session/localStorage still names the previous account's workspace. `loadUser`
read that workspace, got no membership back, threw `Not logged in` and logged
the brand-new session out, landing on `/user/login?rd=...`.

A missing membership says nothing about the session, so forget the workspace
and continue down the no-workspace path, which logs out only when
`globalWhoami` shows the session itself is gone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(frontend): confirm the session before the workspace-picker redirect

`loadWithoutWorkspace` fired the `/user/workspaces?rd=…` navigation before
awaiting `globalWhoami`, so when the session turned out to be gone the logout
read whichever URL the race had left in `page.url` and carried the picker as
its `rd`. Ask first, then redirect.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(frontend): stop the loadUser comments overclaiming what they know

Neither comment can promise what it stated: `getUserExt` collapses every
failure into `undefined`, so the branch cannot tell a real non-membership from
a transient one, and `loadWithoutWorkspace` throws on any `globalWhoami`
rejection rather than only on a dead session. Say what each call actually
answers about.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 15:56:45 +02:00
Alexander PetricandClaude Opus 5.5 e8c3f9514e fix(multiplayer): don't drop client messages during cold-start token verification (#11343)
* fix(multiplayer): don't drop client messages during cold-start token verification

`wss.on('connection')` awaits `verifyToken()` before `setupWSConnection()`
attaches the 'message' listener. On a cold process that await includes the
first `/api/debug/jwks` fetch (~30ms on ECS). A y-websocket client sends sync
step 1 the instant the socket opens, and `ws` drops messages emitted with no
listener attached, so that step 1 was lost and never answered with step 2 —
the client's provider never became `synced`.

Buffer messages from the moment the connection is accepted and replay them, in
order, once `setupWSConnection()` has installed its handlers. Rejected
connections drop the buffer and close with the same 4401/4403 codes as before.

Also prefetch the public key at startup when WINDMILL_BASE_URL is set. That is
insurance, not the fix: a connection arriving before the prefetch resolves
still relies on the buffer.

Adds `npm test` in multiplayer/ (node:test, no docker or backend needed) with a
fake JWKS endpoint that answers with a delay, which holds the cold window open
and makes the race deterministic; wired into the existing test_extra CI job.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(multiplayer): cap what an unauthenticated peer can buffer pre-auth

Review follow-up.

The pre-auth buffer was unbounded: `ws` sets no `maxPayload` here and the JWKS
fetch has no timeout, so a peer that never authenticates could stream frames
into memory for as long as `verifyToken` was stalled. Cap it at 32 frames /
1 MiB — a real client only has sync step 1 and its first awareness update in
flight there — and close 1009 past that, dropping what was buffered.

A socket closed during verification (by the peer, or by that cap) is no longer
handed to setupWSConnection: it would be added to `doc.conns` with a 'close'
listener that can never fire.

The startup prefetch's .catch was dead code — getPublicKey() logs its own
failures and resolves to null rather than rejecting.

Test helper: pin REQUIRE_SIGNED_MULTIPLAYER_REQUESTS and BASE_INTERNAL_URL so an
ambient value cannot turn the rejection tests into false passes; bind the JWKS
server on port 0 instead of a released probe port, and retry the spawned server
on EADDRINUSE; destroy still-delayed JWKS responses on teardown, since
server.close() waits for in-flight requests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(multiplayer): gate the JWKS response instead of delaying it

Review follow-up.

The cold window was held open by a 1500 ms delay on the fake JWKS response, but
that timer started when the startup prefetch reached the fake server, not when
the client sent its first frame. A slow enough machine could load the key before
the client connected, and the race test would then pass without ever exercising
the buffer — a false pass.

The fake JWKS server now parks every response until the test calls release(), so
the server provably holds no key while the client is sending. The race test
releases only after both frames are written to the socket, and asserts the
server has not logged the key as loaded at that point; the flood test never
releases until after the cap has closed the connection.

What is left to wall-clock time is 250 ms for bytes already written to the socket
to cross loopback into an otherwise idle server, rather than a window that had to
cover process startup, connect and handshake.

Also drops the prefetch precondition from the forged-token and flood tests so
each test still maps to one behaviour. Suite runs in ~1.1s instead of ~5.3s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 15:54:11 +02:00