mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-10-04 08:02:23 +00:00
* fix pg_dump stuck on version 17 on nix * fix(datatables): refuse a malformed role annotation instead of ignoring it `-- Role operator`, `-- role operator;` and `-- role operator -- why` all failed the annotation parser's exact-match rule, so the query fell through to the data table's default role and ran, silently, under a login the author did not choose. Naming a role exists precisely to not do that. A leading comment whose first word is `role` is now an annotation attempt: the keyword matches case-insensitively, one trailing `;` is tolerated, and anything else is an error naming the line. Only callers that already know the target is a `datatable://` reference ever run this, so ordinary SQL keeps its comments. Also bumps the dev shell's postgres client to 18 — it trailed the server the dev database runs, which takes out every data table export, clone and fork-with-data. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): refuse a malformed role query string instead of ignoring it `?Role=analytics`, `?role=` and `?x=1&role=…` all fell through the reference parser's exact-match rule, so the connection resolved to the data table's default role and ran under a login the caller never asked for — the URI half of the same trap as a malformed `-- role` annotation. The key now matches case-insensitively, and anything else in the query string is an error naming it; `role` is the only parameter a reference takes. Callers that only need the entry keep a lenient `datatable_ref_name`, since they never act on the role. The DuckDB `ATTACH` parser propagates it rather than attaching under the default. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): carry the role annotation into the row_to_json retry The retry rebuilds its SQL from `pruneComments(code)`, so the leading comment block never reached the second attempt — and with it the `-- role <name>` line that decides which login the query runs as. The retry connected as the data table's default role instead, so a query the first attempt was denied could succeed on the second, reported as "recovered with the row_to_json fix". Carry the leading comment block over. The retry itself is unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * chore(datatables): don't mount the roles UI until the ACL editor lands Enforcement ships first. The permissions drawer is what turns roles on, and the catalog section is what creates them — both are only useful once there is a way to grant a role the privileges it needs, which arrives with the ACL editor. Left mounted they would offer a feature whose other half does not exist. The two components are complete and reviewed; only their call sites here are commented out, with a note pointing the follow-up PRs at them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): honour `-- role: x`, and fix the DuckDB attach test Two review findings, both real. `attach_datatable_parses_name_and_role` never compiled: `parse_attach_datatable` returns `Result<Option<_>>` now and one call site kept a single `unwrap`. Its `?Role=analytics` case also asserted a refusal, contradicting the parser in the same commit, which matches the key case-insensitively. Replaced with the cases that are genuinely malformed, and a positive one for the cased key. `-- role: analytics` fell through to the default role — the silent fallback the strict parser exists to remove, for the spelling most likely to be typed. The keyword now accepts an optional colon, attached or spaced, while a word that merely starts with it (`rolebased`) is still not an attempt. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): clone a fork's pointer instead of failing after the copy Forking a fork with cloning left an orphan database. The preflight resolves the pointer and sees the governing entry, so both endpoints ran and filled the new database; `apply_forked_datatable` then refused the inherited pointer and rolled the fork back, stranding a registered `wm_fork_*` that no entry names and whose name blocks the retry. Refusing earlier would have been the smaller change, but forking a fork and cloning worked before pointers existed, so it would trade an orphan for a regression. Resolve what the pointer names and write the terminal entry the clone needs: the whole `database` object rather than a patch of its `resource_path`, since a pointer has none, and `reference` removed with it. Also accepts `-- role=x` and `-- Role = x`, two more spellings that fell through to the default role. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): refuse to roll back the catalog while roles exist The down migration dropped the table and left every role behind: live Postgres logins whose passwords only that table carried, so after a revert Windmill could neither use, disable nor delete them, and re-applying could not recreate them because the names were taken. Cleaning up here is not possible either — dropping a role means reassigning what it owns in every instance database, and a migration runs in one — so it now refuses while the catalog is non-empty and says to delete the roles through instance settings, which does the cluster work. Also enforces the instance-only invariant the resolved-pointer clone relies on rather than only asserting it in a comment. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * refactor(datatables): settle clonability in one place, before anything is created A clone is three stages a workspace apart — `create_pg_database`, then `import_pg_database`, then `apply_forked_datatable` inside the fork transaction. Only the third can roll back, and `CREATE DATABASE` is not transactional, so any refusal that lives there strands a registered `wm_fork_*` that no entry names and whose name blocks the retry. That orphan has now been fixed three times, most recently reintroduced by a guard added one commit ago. Patching each new refusal into the first endpoint is not the fix; having two places that can refuse is. `ensure_datatable_is_clonable` now answers every reason a copy can be refused and returns what it resolved, and the stage that writes the entry only does the work. Also takes an ACCESS EXCLUSIVE lock before the rollback guard counts, so a role created concurrently cannot slip between the check and the drop. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): let a retried clone reclaim its own leftover database A clone creates its target database one request before it copies into it, and the fork that would name it is written a request after that. Any failure in between — a pg_dump error, a bad restore, a dropped connection, the source's roles changing mid-flow — left a registered `wm_fork_*` that no entry names, and every retry then failed on its name. This predates data table roles. `create_pg_database` now reclaims such a leftover before creating: only a `wm_fork_*` database Windmill registered as a data table database and that no data table or ducklake entry names, in any workspace, archived ones included. The drop never terminates connections, so a clone still copying into it makes the reclaim fail instead of being cut off. It is limited to callers who administer the source — reaching it is not enough, since on a data table without roles every member reaches it — and anyone else gets the refusal an existing database always got. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Revert "fix(datatables): let a retried clone reclaim its own leftover database" This reverts commit7dd3275a10. The reclaim tied the caller to the source they administer, but not to the database it dropped. Between another workspace's import and its final fork request, that workspace's target is full, registered, unnamed and has no open connection, so an admin of any instance data table could name it and have it dropped and recreated empty. The victim's fork would then commit pointing at the empty copy. Safe reclaim needs durable clone ownership and serialization with the request that names the database; until then the leftover stays, as it did before this PR. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(datatables): record the stale clone database as a known limitation A clone is three requests and `CREATE DATABASE` is not transactional, so a failure after the first leaves a registered `wm_fork_*` behind, as it did before data table roles. Accepted for this PR: it is harmless to data and goes away once the clone is a single server-side operation. The comment also records why the obvious fix is wrong: reclaiming the leftover on retry, without durable clone ownership, can drop another workspace's fully copied database between its import and its final fork request. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): bounce the streams reading a data table when it is deleted Deleting a governing data table, or the workspace that holds it, only collected the fork pointers it stranded, for the warning. A Postgres trigger or capture already streaming through one of those pointers kept the replication connection it opened while the pointer still resolved, so it went on dispatching the governing database's rows after the fork lost access — until its connection happened to restart. The governing workspace's own streams on a deleted entry did the same. Both deletion paths now bounce the affected listeners inside their own transaction, through the helper a permission change already uses, so a listener that reconnects re-resolves the entry and finds it gone. The helper is split so a caller can pass the (workspace, local name) pairs it already holds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): keep the fork schema baseline, and bounce streams on every removal Three fixes from review. `edit_datatable_config` took `forked_from` wholesale from the stored entry, so the fork schema diff's save of an advanced baseline was silently discarded and an applied change was offered again. Whether an entry carries a clone stamp is still carried from the store, since that is what marks its database droppable, but the baseline inside it is now taken from the request. The stranded-pointer warning and the stream bounce ran over the optional `deleted_datatables` hint, which the settings-sync CLI never sends, so removing a governing data table through `wmill` bounced nothing. Removals are now derived from the stored configuration against the saved one. `delete_workspace` read the pointers to bounce before its transaction, so a fork committing a pointer during the deletion was missed. The read now happens inside the transaction, after the workspace row is deleted: a fork's insert key-share locks that row through its parent foreign key, so it is either seen or fails on the missing parent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(datatables): keep Postgres triggers and data table roles apart A replication stream reads every row of every table whatever the data table's roles grant, and its listener checks access only when it connects. Rather than chase every way access can change and bounce the streams each one affects, a data table now carries one or the other: - a Postgres trigger or capture cannot be created on, or connect to, a data table under roles; - roles cannot be turned on while an enabled trigger or a live capture reads the data table, its own or a fork's through its pointer. The refusal names each one to disable. This removes the stream bounces on roles edits and on data table and workspace deletion, and the trigger gate that admitted admins. The fork schema baseline fix from the same review round is kept. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): refuse a Postgres trigger on a data table under roles when it is saved Creating or editing a trigger that points at a data table under roles was accepted, and its listener then retried the refused connection every 30 seconds forever. The save is now refused, and a trigger that reaches such a data table anyway (re-enabled, or cloned into a fork) is disabled by its listener with the reason, as a missing replication slot is. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): disable a data table role before deleting it Deleting a role reassigns and drops what it owns in each registered database on its own connection, and each of those passes commits as it goes. A database failing part-way left the role enabled in the catalog and able to log in, but already stripped in the databases reached before it. The role is now disabled in its own commit first, so a failed delete leaves a disabled role to retry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): serialize roles going on with a stream starting Turning roles on looked for enabled triggers and live captures once, without a lock anything starting a stream also took. A trigger enabled in that window could have its listener connect before roles committed, and a healthy listener never checks again. Both transitions now serialize on one advisory lock: roles going on hold it exclusive while they look, and trigger create, edit and enable, and capture setup and ping hold it shared while they commit. Either the look sees the stream, or the listener connects after roles are committed and refuses. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): wait out live listeners, and resolve stored names containing `?` Turning roles on counted a trigger as gone once disabled, and a capture once its client stopped pinging, but the listener keeps its replication connection until its next heartbeat notices. A trigger or capture whose listener pinged in the last 15 seconds, the window a server holds a listener for, now still counts as streaming. Data table names could contain `?` before they were restricted, and such entries are still stored. Splitting `?role=` off a reference misread them: `a?b` became `a` with an unknown parameter, and the clone checks looked at a different entry than the one copied. An entry stored under the whole reference is now looked up first, in the Postgres executor, DuckDB ATTACH and the clone checks. Agent workers cannot read the workspace and keep the strict parse, which refuses such a name rather than misreading it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): warn when a settings sync strands fork pointers A settings save reported the fork pointers left resolving to nothing only for the names in `deleted_datatables`, which `wmill sync push` never sends. The save now works out what it removed from the locked entries, and the CLI prints the stranded pointers it returns. Also correct the replication helper's contract: no role or admin check makes a replication connection safe, so a data table under roles is refused outright rather than gated as an admin operation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): refuse a save that drops a data table's roles through an undeclared rename A data table's roles follow its entry only through a declared rename. A settings sync sends the whole map and never declares one, so renaming a data table under roles there read as a delete and a new entry on the same database: the new entry carried no roles, and every caller connected as admin. Such a save is now refused, naming both entries. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): no entry without roles may newly reach a database under roles The previous guard only caught a new name replacing an entry under roles. A whole-map save could also repoint an existing entry without roles at that database, or another workspace could point one there, and every caller of that entry would connect as admin. The rule is now stated on the saved entries: one that carries no roles and newly points at an instance database any entry under roles uses, in this workspace or another, is refused. A declared rename carries its roles and passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * feat(datatables): move data table role catalog and resolution to the enterprise edition Roles are an Enterprise Edition feature. The catalog, the Postgres logins, CONNECT convergence, tenant evaluation and the role half of connection resolution move to windmill-ee-private. Every public function keeps its path and signature and forwards through datatable_roles_oss, which re-exports the enterprise implementation or, without it, refuses. Without the enterprise edition a data table under roles, or a caller naming a role, is refused a connection rather than resolved as admin, and the reach and admin-access checks refuse one under roles. A data table not under roles resolves as before in every edition, and an instance database keeps the CONNECT grants it was created with. The catalog lock, the stream lock, the tenant cascades and the permissions stripping stay in OSS: they only restrict. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * feat(datatables): move the data table permissions endpoints to the enterprise edition The permissions read, save and usable-roles handlers move to windmill-ee-private; the routes stay registered and, without the enterprise edition, answer that data table roles are an Enterprise Edition feature. ensure_governs_datatable and ensure_reaches_datatable keep their paths: the first refuses, the second passes a data table not under roles. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * feat(datatables): move the data table role catalog endpoints to the enterprise edition The superadmin list, create, update and delete handlers move to windmill-ee-private. The routes stay registered and, without the enterprise edition, refuse after authentication. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * test(datatables): run the roles tests on the enterprise edition, refusals without it Each test that exercises roles runs with private and enterprise. Two tests run without them: every roles route answers the Enterprise refusal, and a data table saved under roles, or a named role, is refused a connection while one not under roles resolves as before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * feat(datatables): gate the roles UI mount sites on an enterprise license Both mount sites are still commented out; the gate travels with them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * test(datatables): run the tenant matcher test on the enterprise edition The matcher it covers is enterprise code now, so without the enterprise edition the test hit the stub and failed the default windmill-common run. It runs with private and enterprise, and a counterpart without them asserts that no tenant list covers anyone, the wildcard and a workspace admin included. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * chore: update ee-repo-ref to a1873dbb67f2302b85ff5362f8387b48eccdb607 This commit updates the EE repository reference after PR #783 was merged in windmill-ee-private. Previous ee-repo-ref: 5c853e2c20eca6b748415fc0d6862a6ebfb5fec4 New ee-repo-ref: a1873dbb67f2302b85ff5362f8387b48eccdb607 Automated by sync-ee-ref workflow. * fix(datatables): refuse roles while a same-workspace alias reaches the database Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(datatables): add an ACL editor for data table roles Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(datatables): data table roles in the DB manager and raw apps Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: never add a role to the reference of a data table whose name contains '?' Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: read the roles of a data table whose name contains '?' The generated client leaves a '?' in a path param unencoded, so the lookup 404'd and the raw-app picker blocked Start on such a data table. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: take every pooled connection before the ACL apply locks Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: refresh grant options only after the ACL apply validates its plan Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix(datatables): refuse a reference naming both a legacy data table and a role When a workspace stores both `sales` and a legacy `sales?role=analytics`, the reference resolved to the legacy entry without a role, so browsing `sales` as `analytics` reached another data table. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: declare the default role in migrations written for a data table whose name contains '?' Such a data table connects as its default role without naming it, so the migrations the manager wrote for it declared no role and ran as admin. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): let CE migrations connect as an explicitly named admin Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: add only missing grant options before an ACL apply, never default privileges Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix(datatables): serialize roles going on with aliases saved from other workspaces Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(datatables): note that legacy names with ? cannot be migrated Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: run one data table ACL apply at a time per server before it connects Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * feat(datatables): set up an external instance cluster for data tables Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): send external cluster passwords as SCRAM verifiers Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): scope external cluster credential readers to the crate Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * [ee] feat(datatables): external_instance data tables on the external cluster Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: hold the ACL connection to the database that was authorized Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix(datatables): compare the external cluster settings under a row lock before storing setup Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): only drop external databases Windmill marked, and check use under the lock Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: build the ACL connection from the authorized data table entry Resolving the settings again could land on a resource with the same database name on another server, which the later entry checks never see. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: check ACL read reach against the entry it connects from Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * feat(datatables): put a data table's connection under Postgres roles A data table backed by the instance database resolved to exactly one Postgres connection, `custom_instance_user`, for everyone who could reach it at all. There was no way to say this job reads, that one writes, this one never sees the salaries table. A data table role is now a real Postgres login on the cluster, defined once for the instance by a superadmin and named exactly as they named it. A script that declares `-- role analytics` connects as `analytics`, and Postgres decides what it may touch — grants are ordinary SQL. Windmill answers only "may this caller ask for this role", from the tenant lists on the data table entry: `u/alice`, `g/analysts`, `f/finance` or `*`. A data table with no `permissions` block behaves exactly as before. Everything that opens a connection on someone's behalf goes through one chokepoint, `get_datatable_resource_from_db`, which takes the identity explicitly and fails closed when there is none. The role logs in as itself — never `SET ROLE`, which a script could `RESET ROLE` its way out of. A fork's data table entry becomes a pointer at the workspace that governs it rather than a copy of it. The settings clone used to hand a fork a byte-identical entry naming the parent's database, which a fork admin could edit to grant themselves `admin` there; a pointer has nothing local to edit, and its tenants are evaluated as a member of the governing workspace, by email. `permissions` is stripped from the workspace export and ignored on import: tenants name principals of one workspace, and a settings push is not where an access decision should be made. Operations that see the whole database whatever the roles grant stay with the governing workspace's admins: editing the roles, a migration that declares none, and opening a replication stream for a Postgres trigger or capture. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): gate the paths that reach a whole database as admin Auditing what still resolved through the unchecked resolver turned up three that act for a caller and hand back the admin connection: `resolve_pg_source_checked` (behind schema export, the full-schema read, database creation, import and the forked-database drop), the connection test, and the schema snapshot a fork clone takes of its parent. On a data table under roles each let any workspace member — or a fork admin who is nobody in the governing workspace — read or copy the whole database whatever its roles grant. All three now require admin reach on the governing workspace. A dump taken under a restricted role would be a silently truncated copy rather than an error, so refusing is the only right answer for the copy paths. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): confine roles to the instance database, and stop a fork reaching the parent's bookkeeping A data table role is a login on Windmill's own Postgres. Nothing stopped a workspace admin putting a *resource-backed* data table under roles, at which point the executor dialled the host that resource names — one the admin chose — with the role's real cluster password, and `CONNECT` is granted to every registered instance database. Both ends now refuse: the permissions endpoint rejects the save, and the chokepoint refuses to substitute credentials on a non-instance entry rather than trusting the record it read. Two more places reached the governing database without answering to it. The initial-migration generator returned a `pg_dump` of the whole schema to any member. And the migration rename/delete cascade followed a fork's pointer into the parent, so a fork admin renaming or removing their own local entry relabelled or wiped the parent's `_wm_migrations` — after which the parent re-runs every migration from zero. The remote half is now skipped when the entry resolves into another workspace, which is also just correct: a fork renaming what it calls a data table changes nothing about the data table. Also: revoking a tenant now bounces the replication streams of every workspace holding an entry that resolves here, not only the governing one, so a fork's trigger stops rather than living on inside its open connection; the instance role catalog and the governing workspace's tenant lists are no longer returned to someone who cannot edit them; and the tenant rename dedup collapses non-adjacent duplicates, per role rather than once any role changed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): fail loudly where a role or a pointer can be left half-recorded Three ways the feature could end up in a state nobody could see or undo. Creating a role writes the cluster first and the catalog second, but the catalog write was an `UPDATE` that matched nothing when the instance Postgres settings row was absent — leaving a live login with a password nobody recorded: invisible to the catalog, un-recreatable because the name is taken, and un-deletable because there is no entry to delete. It now errors, so the operation is retryable once the row is restored. Deleting a workspace only nulls the fork lineage; the data table entries pointing at it are left resolving to nothing. Sweeping them is not an option — turning a pointer back into a copy would hand each fork the database outright — so the delete now names the data tables it stranded, and resolving one says which workspace is missing rather than reporting a data table this workspace never had. `InstanceDatatableRole` derived `Debug` while holding a Postgres password; it is now hand-written so `{:?}` on the catalog cannot put a live credential in a log line. Adds the two branches the reviews found unpinned: a caller who is not a member of the governing workspace at all, and `NoIdentity` — the compatibility path for an agent worker that predates this and sends no job id. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): unbreak two operator messages and two comments that described other code The two strings this branch added for states an operator hits once — the catalog write that matched nothing, and the delete that stranded a pointer — were collapsed from their multi-line form with the indentation left in, so both rendered with a fourteen-space gap mid-sentence. `list_datatables` claimed to report a chain it cannot follow and then dropped it; it does drop it, and the comment now says why that is the right place to stay quiet. The non-superadmin check in `edit_datatable_config` was introduced as also covering references, which it does not and need not: `reference` is overwritten from the stored entry for every caller before the check runs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): serialize role catalog mutations, and state each helper's authorization contract The catalog is one JSON document, so create, rename, enable and delete are all read-modify-write. Two concurrent creates read the same snapshot, both succeed in the cluster, and the second write drops the first — leaving a live Postgres login with a password nobody recorded, which is the exact state the delete path exists to prevent. Every mutation now runs in one transaction holding an advisory lock across the read, the cluster DDL and the write, so a lost update cannot happen and a failure rolls the whole thing back. The DDL helpers take that transaction rather than the pool, which is what makes the lock cover them. Their statements moved off `sqlx::raw_sql`: the simple protocol is only needed for genuinely multi-statement SQL, and its future is not `Send`, which an axum handler holding the transaction requires. Each of these is one statement anyway. The new cross-crate surface now says what callers must do. `read_role_catalog` returns plaintext credentials; `create`/`rename`/`set_login`/`drop_instance_role` and `converge_connect_grants` mutate cluster-wide state; `read_datatable_entry` reads a workspace's raw config. All of them are superadmin-gated by their current handlers, but nothing said so at the definition, which is where the next caller looks. Also: the roles table reloads after a failed login toggle instead of leaving it claiming a flip that did not land; the rename affordance is the design-system `Button`, not a raw one; and `resolve_datatable_pg_as_caller` drops a `role` parameter no caller ever filled — browsing resolves as the data table's default until the database manager grows a picker. Why role passwords stay a plain `String` while the instance user's password beside them is a `StringOrSecretRef`, asked three times across reviews: that one is a secret ref because an operator supplies it and may want it from their own backend, while these are minted here and never entered by anyone, so there is nothing for a ref to point at. Encrypting generated secrets at rest is a separate change that would take the replication password with it. Now said at the field. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): give the role catalog its own row, out of reach of the config machinery Putting it inside `custom_instance_pg_databases` was the wrong call, and it cost two ways. The catalog serializes a generated Postgres password per role, and that row is the operator-facing instance config, so the passwords reached `get_instance_config` and its YAML editor — a live cluster credential in a response body, a UI field and any log of either. Worse in the other direction: `to_settings_map` strips the catalog, so a full-row upsert of that key writes the row back without it and the catalog is gone, while the cluster keeps every login it described. `custom_instance_replication_pwd` is the precedent and says exactly why — a generated secret, written only by the server, never operator-authored, hidden so the config machinery cannot read, rewrite or drop it. The catalog is the same thing, so it now has the same shape: `datatable_roles`, in `HIDDEN_SETTINGS`, `PROTECTED_SETTINGS` and the agent-worker denylist. No redaction to keep in step with three code paths, and no way for a neighbouring write to take it out. Two races on the same shared documents. `edit_datatable_config` read the stored data tables outside its transaction and then wrote the whole `datatable` document, so a permissions save committing in between was silently rolled back; it now reads under `FOR UPDATE`. And `set_datatable_permissions` validated role ids against the catalog before opening its transaction, so a deletion in between let it write a deleted role back — including as the default, which every later job then fails on; it now holds the catalog lock and the settings row across validation and write. Completes the authorization contracts the previous commit claimed but did not finish: `read_datatable_entry` (which it named and missed), `resolve_governing_datatable`, whose whole job is to answer for a workspace the caller may not belong to, and `converge_connect_grants_with`, which had not inherited its wrapper's. Also the generic Python SDK reference: `_format_py_params` learned the bare `*` last time, but `extract_py_functions` is a second formatter and still rendered `datatable(name, role)`, so code written from that page passed a keyword-only argument positionally. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): make the concurrency test pin the handlers, and the contracts describe what is enforced The concurrency test reimplemented the read-modify-write inline, so deleting the lock from all three handlers left it green — it pinned Postgres, not the code it was written for. It now drives `create_datatable_role` twice concurrently and asserts the catalog kept both names. Checked the way the last one should have been: removing the lock from the handler makes it fail with "wmtest_a_… is a live cluster login the catalog forgot". The contracts added last commit were stricter than this PR's own callers, which is worse than none — the next reader sees a rule already broken and learns to ignore it. `read_role_catalog` said superadmin-only while two of its four callers are open to any workspace member, and `converge_connect_grants` said superadmin while `set_datatable_permissions` reaches it as a workspace admin. Both were fine on substance: the rule that actually holds is about the credential never reaching a response, log, audit record or export, not about who may call. They now say that. `read_datatable_entry` gets the same treatment rather than the one the earlier message claimed for it: it is the primitive every resolution goes through, so it is deliberately open, and what must not escape is `permissions` — it names the governing workspace's users, groups and folders. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): close the last ways a role or a pointer can be left pointing at nothing The raw settings readers hand back whatever is in the row, so moving the catalog into its own `global_settings` key protected the config machinery and left `GET /settings/global/datatable_roles` and the settings listing returning every live password. Both now filter that one key. The neighbouring `custom_instance_replication_pwd` has the same shape and is not touched here: it predates this and widening the fix to it is a decision about an operator workflow, not a consequence of this change. Three ways a save could leave something resolving to nothing: A permissioned data table could be moved to a PostgreSQL resource. The block was carried across as a server-owned field, the runtime refuses roles on a resource-backed table, so the save succeeded and every job afterwards failed. Refused instead — turning roles off first is one step, and it keeps discarding an access decision something somebody chose. Renaming a governing data table left every fork pointing at the old name: the data table disappears from their pickers and their jobs stop, with nothing in the renaming workspace to suggest why. The rename now follows into the pointers in the same transaction. Deleting one cannot be followed the same way, so it is reported instead — the response names what it stranded, the way deleting a workspace does, and the fork's own error already says which workspace is gone. Also: `ensure_instance_db_grant_options_unchecked` claimed superadmin while the permissions handler reaches it as a workspace admin (the same class fixed last commit, one instance missed); the role entry kept an `instance_config_schema` derive it no longer needs; `write_role_catalog` was the one writer of that table not stamping `updated_at`; and the concurrency test dropped its roles only on success — a failing run is exactly the one that creates them without recording them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * refactor(datatables): put the role catalog in its own table, not in global_settings Five findings across three rounds were all the same choice. A set of live Postgres credentials was living in `global_settings`, which has generic read, list, write, config-export and CLI round-trip paths that know nothing about what they carry: the passwords reached the instance config and its YAML editor, a full-row upsert of a neighbouring key erased the catalog, `GET /settings/global/{key}` and the settings listing returned them raw, and this round the redaction that fixed the last two turned `wmill instance push` into something that wipes every password — a fix breaking the assumption the previous fix made. `POST /settings/global/datatable_roles` could also empty it outside the lock. The approved plan offered a table or `global_settings`, so this is the other option it already allowed rather than a new design. `datatable_role` is a table: no generic settings path can read it, list it, export it, write it or round-trip it, so none of the five needs a guard. The redaction, the hidden/protected/agent-denylist entries and the JSON document all go with it. One row per role also removes the read-modify-write the concurrency work was about: two concurrent creates are two inserts, and the unique index on `name` is what settles a collision. The advisory lock stays for the one window rows do not cover — `CREATE ROLE` is invisible to another transaction until commit, so without it both creates pass their `pg_roles` check. Also from this round: rename mappings are checked against the configuration they claim to describe, since fork pointers are rewritten from them — a caller could otherwise submit `main -> missing` against an unchanged config and repoint every fork of `main` at a name nothing has, and `A -> B` plus `B -> C` moved what pointed at `A` all the way to `C`. And the warning naming forks a delete stranded reached the response but not the screen: both the data table settings save and the workspace delete now show it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): validate a rename against the save it describes, and re-check under the locks Three from the round, all about deciding on state that could already have moved. A permission save resolved the data table and checked it was instance-backed before taking any lock, then wrote under one. A config save committing in between could move the table onto a PostgreSQL resource — recreating exactly what the transition guard refuses — or rename it, in which case the write targeted a key that no longer existed and reported success having changed nothing. It now re-resolves and re-checks on the locked state. Rename validation checked that the source existed before and the target existed after, which still accepts `main -> decoy` against a save that keeps both: every fork of `main` then follows onto a different data table, silently, because it keeps resolving. The rule is now the actual old-to-new key transition — a source may only survive if another rename took its name, and a target may only pre-exist if another rename freed it. That also stops two sources sharing one target, and it admits a swap, which the previous guard refused: `datatables` is keyed by name, so a swap cannot be done one save at a time, and refusing it was a regression against main. The pointer cascade now runs in two passes through a temporary name, the way the migration cascade one layer down already handles the same shape, so `A -> B` with `B -> C` moves each pointer once from what it named before the save. The tenant mutators say what they are for: they write an access decision for any workspace named, with an arbitrary mutation, and exist for the transaction that frees or renames a principal. Editing a decision on purpose belongs in the permissions endpoint. Carried in the same change: the stranded-fork list is a field rather than a phrase to grep out of a success string; the pointer cascade matches with `EXISTS` instead of a `LIKE` over the whole document, so a workspace whose pointers name something else is not rewritten to a byte-identical value under an exclusive lock; and `InstanceDatatableRole` drops the serde derives left over from the JSON document, one of which would emit `pwd`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): cascade on the leave route that is used, gate migrations before the admin connection, and drop a role atomically The tenant cascade on leaving went onto `/users/leave`. The UI and the generated client call `/workspaces/leave` — a different handler in a different crate with the same name — which deleted the membership and left `u/<username>` in the tenant lists. Leaving and rejoining therefore restored the access the leave was supposed to end, and a later account taking the username would have inherited it. The regression test drives the route the client actually calls; without the fix it fails with "leaving kept the tenant". The migration endpoints authorized too late. `run_datatable_migrations` opened the data table's admin connection, created `_wm_migrations` and read it before reaching the per-migration role check — so with nothing pending, nothing was checked at all. Rollback returned before its check when nothing was applied, and the status endpoint had none. All three now ask, before any connection is opened, whether the caller can reach the data table as any role at all; which role a given migration runs as is still decided per migration, and by the executor after that. Deleting a role committed the cluster drop and the catalog row, then swept the tenant lists in separate transactions. A sweep failing part-way left workspaces naming a role nothing can connect as, while the retry answered `NotFound` because the catalog entry was already gone. The sweep now runs in the same transaction, so the drop, the row and every tenant list commit together. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): refuse to copy a data table that is under roles pg_dump carries no roles and the import runs with --no-privileges, so a copied data table arrives owned by the admin connection with no GRANT for any role. The settings clone brings `permissions` across, so the fork's tenants pass Windmill's check, connect as the role they were given, and are denied by Postgres on everything: an entry that reads as configured and answers nothing. Refuse the copy — in the import endpoint before any data moves, and in the fork path the CLI takes. Replaying the source's owners and ACLs into the clone is what lifts this, and is a change of its own. Dropping `permissions` from the copy instead would be the unsafe half, since the copy holds the parent's rows. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): refuse the clone's database too, not only its data A clone is two endpoints: `create_pg_database` then `import_pg_database`. Only the second refused a data table under roles, so a fork asking to clone one created and registered an empty `wm_fork_…` instance database and then failed — and nothing collects it, since `drop_forked_datatable_databases` only drops entries carrying `forked_from` and no entry names this one. Refuse in both, so the clone stops before a database exists. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * nit worker error msg * fix pg_dump stuck on version 17 on nix * fix(datatables): refuse a malformed role annotation instead of ignoring it `-- Role operator`, `-- role operator;` and `-- role operator -- why` all failed the annotation parser's exact-match rule, so the query fell through to the data table's default role and ran, silently, under a login the author did not choose. Naming a role exists precisely to not do that. A leading comment whose first word is `role` is now an annotation attempt: the keyword matches case-insensitively, one trailing `;` is tolerated, and anything else is an error naming the line. Only callers that already know the target is a `datatable://` reference ever run this, so ordinary SQL keeps its comments. Also bumps the dev shell's postgres client to 18 — it trailed the server the dev database runs, which takes out every data table export, clone and fork-with-data. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): refuse a malformed role query string instead of ignoring it `?Role=analytics`, `?role=` and `?x=1&role=…` all fell through the reference parser's exact-match rule, so the connection resolved to the data table's default role and ran under a login the caller never asked for — the URI half of the same trap as a malformed `-- role` annotation. The key now matches case-insensitively, and anything else in the query string is an error naming it; `role` is the only parameter a reference takes. Callers that only need the entry keep a lenient `datatable_ref_name`, since they never act on the role. The DuckDB `ATTACH` parser propagates it rather than attaching under the default. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): carry the role annotation into the row_to_json retry The retry rebuilds its SQL from `pruneComments(code)`, so the leading comment block never reached the second attempt — and with it the `-- role <name>` line that decides which login the query runs as. The retry connected as the data table's default role instead, so a query the first attempt was denied could succeed on the second, reported as "recovered with the row_to_json fix". Carry the leading comment block over. The retry itself is unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * chore(datatables): don't mount the roles UI until the ACL editor lands Enforcement ships first. The permissions drawer is what turns roles on, and the catalog section is what creates them — both are only useful once there is a way to grant a role the privileges it needs, which arrives with the ACL editor. Left mounted they would offer a feature whose other half does not exist. The two components are complete and reviewed; only their call sites here are commented out, with a note pointing the follow-up PRs at them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): honour `-- role: x`, and fix the DuckDB attach test Two review findings, both real. `attach_datatable_parses_name_and_role` never compiled: `parse_attach_datatable` returns `Result<Option<_>>` now and one call site kept a single `unwrap`. Its `?Role=analytics` case also asserted a refusal, contradicting the parser in the same commit, which matches the key case-insensitively. Replaced with the cases that are genuinely malformed, and a positive one for the cased key. `-- role: analytics` fell through to the default role — the silent fallback the strict parser exists to remove, for the spelling most likely to be typed. The keyword now accepts an optional colon, attached or spaced, while a word that merely starts with it (`rolebased`) is still not an attempt. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): clone a fork's pointer instead of failing after the copy Forking a fork with cloning left an orphan database. The preflight resolves the pointer and sees the governing entry, so both endpoints ran and filled the new database; `apply_forked_datatable` then refused the inherited pointer and rolled the fork back, stranding a registered `wm_fork_*` that no entry names and whose name blocks the retry. Refusing earlier would have been the smaller change, but forking a fork and cloning worked before pointers existed, so it would trade an orphan for a regression. Resolve what the pointer names and write the terminal entry the clone needs: the whole `database` object rather than a patch of its `resource_path`, since a pointer has none, and `reference` removed with it. Also accepts `-- role=x` and `-- Role = x`, two more spellings that fell through to the default role. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): refuse to roll back the catalog while roles exist The down migration dropped the table and left every role behind: live Postgres logins whose passwords only that table carried, so after a revert Windmill could neither use, disable nor delete them, and re-applying could not recreate them because the names were taken. Cleaning up here is not possible either — dropping a role means reassigning what it owns in every instance database, and a migration runs in one — so it now refuses while the catalog is non-empty and says to delete the roles through instance settings, which does the cluster work. Also enforces the instance-only invariant the resolved-pointer clone relies on rather than only asserting it in a comment. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * refactor(datatables): settle clonability in one place, before anything is created A clone is three stages a workspace apart — `create_pg_database`, then `import_pg_database`, then `apply_forked_datatable` inside the fork transaction. Only the third can roll back, and `CREATE DATABASE` is not transactional, so any refusal that lives there strands a registered `wm_fork_*` that no entry names and whose name blocks the retry. That orphan has now been fixed three times, most recently reintroduced by a guard added one commit ago. Patching each new refusal into the first endpoint is not the fix; having two places that can refuse is. `ensure_datatable_is_clonable` now answers every reason a copy can be refused and returns what it resolved, and the stage that writes the entry only does the work. Also takes an ACCESS EXCLUSIVE lock before the rollback guard counts, so a role created concurrently cannot slip between the check and the drop. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): let a retried clone reclaim its own leftover database A clone creates its target database one request before it copies into it, and the fork that would name it is written a request after that. Any failure in between — a pg_dump error, a bad restore, a dropped connection, the source's roles changing mid-flow — left a registered `wm_fork_*` that no entry names, and every retry then failed on its name. This predates data table roles. `create_pg_database` now reclaims such a leftover before creating: only a `wm_fork_*` database Windmill registered as a data table database and that no data table or ducklake entry names, in any workspace, archived ones included. The drop never terminates connections, so a clone still copying into it makes the reclaim fail instead of being cut off. It is limited to callers who administer the source — reaching it is not enough, since on a data table without roles every member reaches it — and anyone else gets the refusal an existing database always got. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Revert "fix(datatables): let a retried clone reclaim its own leftover database" This reverts commit7dd3275a10. The reclaim tied the caller to the source they administer, but not to the database it dropped. Between another workspace's import and its final fork request, that workspace's target is full, registered, unnamed and has no open connection, so an admin of any instance data table could name it and have it dropped and recreated empty. The victim's fork would then commit pointing at the empty copy. Safe reclaim needs durable clone ownership and serialization with the request that names the database; until then the leftover stays, as it did before this PR. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(datatables): record the stale clone database as a known limitation A clone is three requests and `CREATE DATABASE` is not transactional, so a failure after the first leaves a registered `wm_fork_*` behind, as it did before data table roles. Accepted for this PR: it is harmless to data and goes away once the clone is a single server-side operation. The comment also records why the obvious fix is wrong: reclaiming the leftover on retry, without durable clone ownership, can drop another workspace's fully copied database between its import and its final fork request. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): bounce the streams reading a data table when it is deleted Deleting a governing data table, or the workspace that holds it, only collected the fork pointers it stranded, for the warning. A Postgres trigger or capture already streaming through one of those pointers kept the replication connection it opened while the pointer still resolved, so it went on dispatching the governing database's rows after the fork lost access — until its connection happened to restart. The governing workspace's own streams on a deleted entry did the same. Both deletion paths now bounce the affected listeners inside their own transaction, through the helper a permission change already uses, so a listener that reconnects re-resolves the entry and finds it gone. The helper is split so a caller can pass the (workspace, local name) pairs it already holds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): keep the fork schema baseline, and bounce streams on every removal Three fixes from review. `edit_datatable_config` took `forked_from` wholesale from the stored entry, so the fork schema diff's save of an advanced baseline was silently discarded and an applied change was offered again. Whether an entry carries a clone stamp is still carried from the store, since that is what marks its database droppable, but the baseline inside it is now taken from the request. The stranded-pointer warning and the stream bounce ran over the optional `deleted_datatables` hint, which the settings-sync CLI never sends, so removing a governing data table through `wmill` bounced nothing. Removals are now derived from the stored configuration against the saved one. `delete_workspace` read the pointers to bounce before its transaction, so a fork committing a pointer during the deletion was missed. The read now happens inside the transaction, after the workspace row is deleted: a fork's insert key-share locks that row through its parent foreign key, so it is either seen or fails on the missing parent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(datatables): keep Postgres triggers and data table roles apart A replication stream reads every row of every table whatever the data table's roles grant, and its listener checks access only when it connects. Rather than chase every way access can change and bounce the streams each one affects, a data table now carries one or the other: - a Postgres trigger or capture cannot be created on, or connect to, a data table under roles; - roles cannot be turned on while an enabled trigger or a live capture reads the data table, its own or a fork's through its pointer. The refusal names each one to disable. This removes the stream bounces on roles edits and on data table and workspace deletion, and the trigger gate that admitted admins. The fork schema baseline fix from the same review round is kept. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): refuse a Postgres trigger on a data table under roles when it is saved Creating or editing a trigger that points at a data table under roles was accepted, and its listener then retried the refused connection every 30 seconds forever. The save is now refused, and a trigger that reaches such a data table anyway (re-enabled, or cloned into a fork) is disabled by its listener with the reason, as a missing replication slot is. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): disable a data table role before deleting it Deleting a role reassigns and drops what it owns in each registered database on its own connection, and each of those passes commits as it goes. A database failing part-way left the role enabled in the catalog and able to log in, but already stripped in the databases reached before it. The role is now disabled in its own commit first, so a failed delete leaves a disabled role to retry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): serialize roles going on with a stream starting Turning roles on looked for enabled triggers and live captures once, without a lock anything starting a stream also took. A trigger enabled in that window could have its listener connect before roles committed, and a healthy listener never checks again. Both transitions now serialize on one advisory lock: roles going on hold it exclusive while they look, and trigger create, edit and enable, and capture setup and ping hold it shared while they commit. Either the look sees the stream, or the listener connects after roles are committed and refuses. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): wait out live listeners, and resolve stored names containing `?` Turning roles on counted a trigger as gone once disabled, and a capture once its client stopped pinging, but the listener keeps its replication connection until its next heartbeat notices. A trigger or capture whose listener pinged in the last 15 seconds, the window a server holds a listener for, now still counts as streaming. Data table names could contain `?` before they were restricted, and such entries are still stored. Splitting `?role=` off a reference misread them: `a?b` became `a` with an unknown parameter, and the clone checks looked at a different entry than the one copied. An entry stored under the whole reference is now looked up first, in the Postgres executor, DuckDB ATTACH and the clone checks. Agent workers cannot read the workspace and keep the strict parse, which refuses such a name rather than misreading it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): warn when a settings sync strands fork pointers A settings save reported the fork pointers left resolving to nothing only for the names in `deleted_datatables`, which `wmill sync push` never sends. The save now works out what it removed from the locked entries, and the CLI prints the stranded pointers it returns. Also correct the replication helper's contract: no role or admin check makes a replication connection safe, so a data table under roles is refused outright rather than gated as an admin operation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): refuse a save that drops a data table's roles through an undeclared rename A data table's roles follow its entry only through a declared rename. A settings sync sends the whole map and never declares one, so renaming a data table under roles there read as a delete and a new entry on the same database: the new entry carried no roles, and every caller connected as admin. Such a save is now refused, naming both entries. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): no entry without roles may newly reach a database under roles The previous guard only caught a new name replacing an entry under roles. A whole-map save could also repoint an existing entry without roles at that database, or another workspace could point one there, and every caller of that entry would connect as admin. The rule is now stated on the saved entries: one that carries no roles and newly points at an instance database any entry under roles uses, in this workspace or another, is refused. A declared rename carries its roles and passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * feat(datatables): move data table role catalog and resolution to the enterprise edition Roles are an Enterprise Edition feature. The catalog, the Postgres logins, CONNECT convergence, tenant evaluation and the role half of connection resolution move to windmill-ee-private. Every public function keeps its path and signature and forwards through datatable_roles_oss, which re-exports the enterprise implementation or, without it, refuses. Without the enterprise edition a data table under roles, or a caller naming a role, is refused a connection rather than resolved as admin, and the reach and admin-access checks refuse one under roles. A data table not under roles resolves as before in every edition, and an instance database keeps the CONNECT grants it was created with. The catalog lock, the stream lock, the tenant cascades and the permissions stripping stay in OSS: they only restrict. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * feat(datatables): move the data table permissions endpoints to the enterprise edition The permissions read, save and usable-roles handlers move to windmill-ee-private; the routes stay registered and, without the enterprise edition, answer that data table roles are an Enterprise Edition feature. ensure_governs_datatable and ensure_reaches_datatable keep their paths: the first refuses, the second passes a data table not under roles. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * feat(datatables): move the data table role catalog endpoints to the enterprise edition The superadmin list, create, update and delete handlers move to windmill-ee-private. The routes stay registered and, without the enterprise edition, refuse after authentication. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * test(datatables): run the roles tests on the enterprise edition, refusals without it Each test that exercises roles runs with private and enterprise. Two tests run without them: every roles route answers the Enterprise refusal, and a data table saved under roles, or a named role, is refused a connection while one not under roles resolves as before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * feat(datatables): gate the roles UI mount sites on an enterprise license Both mount sites are still commented out; the gate travels with them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * test(datatables): run the tenant matcher test on the enterprise edition The matcher it covers is enterprise code now, so without the enterprise edition the test hit the stub and failed the default windmill-common run. It runs with private and enterprise, and a counterpart without them asserts that no tenant list covers anyone, the wildcard and a workspace admin included. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * chore: update ee-repo-ref to a1873dbb67f2302b85ff5362f8387b48eccdb607 This commit updates the EE repository reference after PR #783 was merged in windmill-ee-private. Previous ee-repo-ref: 5c853e2c20eca6b748415fc0d6862a6ebfb5fec4 New ee-repo-ref: a1873dbb67f2302b85ff5362f8387b48eccdb607 Automated by sync-ee-ref workflow. * fix(datatables): refuse roles while a same-workspace alias reaches the database Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): let CE migrations connect as an explicitly named admin Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): serialize roles going on with aliases saved from other workspaces Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(datatables): note that legacy names with ? cannot be migrated Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(datatables): add an ACL editor for data table roles Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: take every pooled connection before the ACL apply locks Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: refresh grant options only after the ACL apply validates its plan Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: add only missing grant options before an ACL apply, never default privileges Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: run one data table ACL apply at a time per server before it connects Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: hold the ACL connection to the database that was authorized Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: build the ACL connection from the authorized data table entry Resolving the settings again could land on a resource with the same database name on another server, which the later entry checks never see. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: check ACL read reach against the entry it connects from Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * feat(datatables): Ducklake catalogs on the external instance cluster Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): never grant CREATEROLE to custom_instance_user on the external cluster Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(datatables): state the authorization contract of external database usage lookups Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): drop a DuckDB data table secret once its ATTACH has used it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * perf(datatables): resolve a workspace's data tables per pointer hop, not per entry Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): hold the parent's settings while a fork points at its data tables Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): refuse repointing the external cluster while it is in use, and keep verify-ca working for pg_dump Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): protect external databases pending fork cleanup, and describe Ducklake usage in the API Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(datatables): per-cluster data table role catalogs, with roles on the external instance cluster Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): write the external cluster setting under the lifecycle lock, and check fork targets are registered Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): register external fork catalogs under the lifecycle lock, and keep certificate verification in DuckDB attaches Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): keep certificate verification when DuckDB attaches an external data table Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): keep certificate verification when DuckDB attaches an external data table Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): refuse fork cleanup of an external database another workspace uses Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): refuse rolling back while external data tables are under roles, and type external_instance in the CLI Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): stop counting storage-only fork cleanup rows as uses of an external database Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): bind fork database copies to their workspace, and count every use before dropping one Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): authenticate instance database setup before writing its status, and keep a fork reservation across it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): create external databases only on a cluster setup succeeded on, and document the registry reader Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): keep only the most recently used DuckDB root certificate files Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): migrate fork reservations on workspace rename, and lock the parent's data tables for the whole fork Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): take the fork data table lock once, before the external cluster's Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(datatables): describe the external instance cluster and how to run one locally Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): keep fork reservations private, drop a cleaned-up entry with its database, and serialize cleanup with settings saves Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): hold the fork lock across a fork import, and carry the reservation inside the setup write Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): check the external cluster setting on its own transaction, and gate the registry probe Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(datatables): add an ACL editor for data table roles Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: take every pooled connection before the ACL apply locks Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: refresh grant options only after the ACL apply validates its plan Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: add only missing grant options before an ACL apply, never default privileges Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: run one data table ACL apply at a time per server before it connects Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: hold the ACL connection to the database that was authorized Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: build the ACL connection from the authorized data table entry Resolving the settings again could land on a resource with the same database name on another server, which the later entry checks never see. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * fix: check ACL read reach against the entry it connects from Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb * chore: update ee-repo-ref to 7e338e4dabf91689bfd7fb0333c6534040b17b59 This commit updates the EE repository reference after PR #787 was merged in windmill-ee-private. Previous ee-repo-ref: 0edd40979cf36bfba59323f3f6a0811ae1369cf5 New ee-repo-ref: 7e338e4dabf91689bfd7fb0333c6534040b17b59 Automated by sync-ee-ref workflow. * fix(datatables): keep DuckDB root certificate files in the job directory Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: drop the unused json import from the settings crate Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): drop the serde_json::json import left unused by the fork setup write Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): accept external_instance data tables in the settings form type Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): serialize fork database reservations, cleanup, setup and Ducklake saves on the database - A fork copy is created and registered under its workspace's fork lock, and refused once the workspace is archived; a rename re-migrates reservations under that lock after archiving. - Ducklake saves lock every instance database they newly name, as data table saves do. - Instance database setup holds the database's lock until its entry is written. - Fork import also holds the database's lock across the restore. - Cleanup re-reads the entry under its locks before dropping anything. - Non-superadmins no longer see fork copies reserved for workspaces they are not in. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): state the authorization contract of the external cluster status and write check Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): lock the fork copy at finalization, and order cleanup's registry write after its settings row - Fork finalization holds the copy's database lock from its availability check to the commit. - Cleanup removes the registry entry in its own transaction, in a task of its own, instead of on a second connection; a rename migrates reservations after its settings rewrites, so both take the settings rows before the registry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): hold the fork lock across a rename's settings copy Fork cleanup of the old id could otherwise drop a copy the renamed workspace goes on using. Also states the authorization contract of the instance database drop helpers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): refuse moving the external cluster to another host or port while it holds databases Every settings writer goes through the same check as removal: the login, TLS and maintenance database may still change, the cluster may not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(datatables): clone a data table under roles with its owners and grants (#11120) Claude-Session: https://claude.ai/code/session_01UbrtwiYNfayrmqouBJHwGV Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: open the raw app data table drawer when the workspace has none Selecting the first data table of an empty list passed undefined to the name check, which threw instead of opening the drawer on no data table. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): offer cloning a data table under roles where its grants can be replayed The server clones such a data table and replays the source's owners and grants, which only the Enterprise Edition does, so the fork wizard hid both clone options everywhere instead of on a build that cannot replay them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: name the placeholder the empty raw app data drawer renders Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): replace a job directory file at the DuckDB root certificate path instead of trusting it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: pin the enterprise refusal the role pickers read as 'not under roles' The server's sentence and the frontend's copy of it were coupled by nothing, so rewording either one turned every role picker on a community build into a failed lookup. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): type Ducklake catalogs on the external instance cluster in the settings form Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): read a roles answer only for the workspace it was asked in A fork and its parent each have their own roles on a data table of the same name, so an answer stamped with the name alone settled the role from the workspace the editor was acting on before. Also derive the AI table creation flag from the data replaced into the editor: data naming no data table left the flag on from before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): keep an instance database a settings save is waiting to name Cleanup for a database whose setup failed took the lock first, read no user, and dropped it while a save blocked on that same lock was about to commit a reference to it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): check the workspace stamp in the default database selector too Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): read what a save racing instance-database cleanup committed A transaction blocked on the lock may still roll back, so keeping the database for it stranded one whose name then blocks every retry: it is let through and its outcome read instead. The waiter query also matches this database's locks only, since pg_locks spans the cluster. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): tell a waiting request apart from the workspaces using a database Both callers render what cleanup returns as the workspaces that keep the database, so a waiting request's pid read as one of them. Each now words that case itself, and the give-up comment names where the kept name actually goes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): take the cleanup lock on a connection the pool cannot reclaim A session lock outlives the future holding it, so a cancellation between taking it and releasing it handed a locked session back to the pool, where every later settings save waits on it. Detached, the connection closes when it is dropped and the server releases the lock with the session. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): close the cleanup connection on drop instead of detaching it Detaching released the pool permit while the session stayed alive, so concurrent cleanups waiting on their locks could open as many connections as they liked. Closing on drop covers the same cancellation and keeps them counted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to 179d2454a063ee818a9387a1eafdd354a217a16c This commit updates the EE repository reference after PR #802 was merged in windmill-ee-private. Previous ee-repo-ref: c2e43f5b5ff753d339b70e232fd17b6ffecb054d New ee-repo-ref: 179d2454a063ee818a9387a1eafdd354a217a16c Automated by sync-ee-ref workflow. * chore: update ee-repo-ref to fd5b8af748f2c985b13e18d9ea30894f3bd7e9a3 This commit updates the EE repository reference after PR #798 was merged in windmill-ee-private. Previous ee-repo-ref: 3145e422d61d580f0a82804f075285c112879da0 New ee-repo-ref: fd5b8af748f2c985b13e18d9ea30894f3bd7e9a3 Automated by sync-ee-ref workflow. * fix(datatables): classify a fork import's target from the resolution that built its connection A second read could see the entry flipped to a resource and skip the reservation check while the connection already built still reached the instance cluster. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): run external database lifecycle on the caller's transaction, one lock order for fork cleanup Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): check a fork copy's reservation on the locked transaction's connection Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): make and drop fork copies of external data tables on the external cluster Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): refuse a fork whose external copy something already names Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): check an external fork copy's uses before removing its entry, so child fork pointers still count Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): migrate external cluster fork reservations on workspace rename Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(datatables): configure the external instance cluster and pick its databases from the UI Instance settings gets an External Postgres tab: the cluster's connection, the setup run with its report, and the databases Windmill created there. Data table and Ducklake settings offer the external_instance kind, with a picker over those databases. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): keep the external cluster form out of the instance settings the page bulk-saves The form seeded its key into the settings store after the page snapshotted them, so merely opening the tab made the page send a setting the validator refuses, failing an admin's unrelated save. The form is local state now, and only the setup writes the key. The database picker also treats a superadmin-only listing as authoritative: a workspace admin, who cannot list, no longer sees a saved database as missing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(tests): app paths are validated, so a guest app cannot be renamed to one with a space Path validation on apps landed after this test, which still expected a space to be accepted and created its fixture on a ':' path. The scopable-path guard it exists for is kept by planting that path on the row, which is now the only way an app can hold one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(datatables): one Managed Postgres tab, and a switch for Windmill's own database The instance settings tab covers both substrates Windmill administers: the external cluster, which can now be disabled once nothing sits on it, and Windmill's own database, which an operator can turn off so the cluster is the only one a workspace may newly name. A save that names it then refuses, as the create endpoint does. Workspace settings name them the way an admin meets them: Postgres Resource, and Managed instance, qualified as Internal or External only while both are on offer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(datatables): offer the external cluster in the add-data-table wizard, and give Ducklake its room back The wizard gets the external cluster as a fourth substrate: pick a database it already holds or name a new one, which the run creates before writing the entry, so Try again does not trip on a database the last attempt made. Ducklake's maintenance column moves into the settings popover next to the extra args, which the wider catalog select had squeezed the name box out of. A setting a component saves itself now also moves the page's baseline, or Discard would restore what it replaced and a later save would send that stale value back. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): only a database this wizard run made makes its retry idempotent A retry skipped the create whenever the name was registered, which let a run adopt a database another superadmin had made and share their data without the warning the existing-database branch shows. The run tracks what it created instead, and reports it so Discard can say the database is still there. Creating one now refreshes the registry the next run's default name and its validation read, and that default counts the external cluster's databases rather than the instance's. Row ownership carries the substrate: "instance" and "external_instance" can hold the same name, and a row repointed between them while a run probes must not read as that run's own. Also fixes the merge leftovers in the DuckDB executor's tests, which cargo check never compiles: the new PgDatabase field and the attach helper's job directory. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): shorten the picked managed kind, and say what an empty database list means The closed select has room for the qualifier, not the whole name, so a picked managed kind reads "Managed (Internal)". An external entry gets the tooltip its internal counterpart has, and a database picker with nothing in it says to type a name rather than "No items found". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): keep the wizard's external-database bookkeeping optional and parked The run deps gained a required field, which every existing caller of runSetup -- the sibling test suite included -- does not pass. It defaults to empty instead, the parked payload carries it across a Supabase redirect as it claims to, and a resumed run counts it as something left behind. Two regressions pin the create: made on a first attempt, skipped only for a database this run made. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): review an external-cluster database as one, not as a Postgres resource The review step fell through to the resource branch for the new provider, so it announced a connection already in the workspace for a database that does not exist yet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): show the server's refusal instead of an object in the toast Saving data table settings passed the whole error to the toast, which rendered the request and response as JSON, and the instance-database wizard replaced it with "check console". Both surface the message the API sent, through the helper that exists for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: point at the EE commit dropping the unused import Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): define a role on the cluster its data table sits on The permissions drawer created and listed roles without naming a cluster, so a role defined from an external data table became a login on Windmill's own, and the refresh then replaced the correct catalog with the internal one. The drawer passes the data table's cluster to both. The wizard also read a database it had just created as a name collision, which blocked returning to the failed attempt and finishing under the same name. And the merge had left the CONNECT-grant pass and its log in both the caller and the callee. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(datatables): manage the external cluster's role catalog, and name the cluster in the copy The roles drawer offered one catalog, so the external cluster's roles could only be reached through a data table sitting on it — and became unreachable once the last one was gone, while the cluster could not be unset until they were dropped. Opened from the page it now offers the clusters the instance has; a data table's own drawer still pins its cluster, and its copy names that cluster rather than "the instance". A DuckDB attach also decides whether to verify certificates the way every other Postgres connection does, so a resource carrying a root certificate is no longer downgraded to require here alone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(datatables): bind the roles list to the cluster it was read for Switching catalogs left the previous cluster's rows on screen and applied whichever response landed last, so a slower read could seat one cluster's logins under the other's heading while every control acted on the wrong id — invisible where a name exists on both. The switch clears the rows, a token discards a response the selection has moved past, and the controls stay inert while a catalog is being read. The drop confirmation also names the cluster whose databases it reaches. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to dfe7b8b9204d21e0264fbea1c6f6eedf9e738d56 This commit updates the EE repository reference after PR #810 was merged in windmill-ee-private. Previous ee-repo-ref: 30db1b33bda0446f5fc5dfe353fbb226d57a26d4 New ee-repo-ref: dfe7b8b9204d21e0264fbea1c6f6eedf9e738d56 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
2865 lines
108 KiB
Rust
2865 lines
108 KiB
Rust
/*
|
||
* Author: Ruben Fiszel
|
||
* Copyright: Windmill Labs, Inc 2023
|
||
* This file and its contents are licensed under the AGPLv3 License.
|
||
* Please see the included NOTICE for copyright information and
|
||
* LICENSE-AGPL for a copy of the license.
|
||
*/
|
||
|
||
use std::{
|
||
collections::{BTreeSet, HashMap},
|
||
time::Duration,
|
||
};
|
||
|
||
#[cfg(feature = "parquet")]
|
||
mod audit_logs_s3;
|
||
#[cfg(feature = "parquet")]
|
||
mod audit_logs_s3_backfill;
|
||
#[cfg(feature = "parquet")]
|
||
mod background_task;
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
mod datatable_roles_ee;
|
||
mod datatable_roles_oss;
|
||
#[cfg(feature = "private")]
|
||
mod ee;
|
||
pub mod ee_oss;
|
||
#[cfg(feature = "parquet")]
|
||
mod log_cleanup;
|
||
#[cfg(feature = "parquet")]
|
||
mod storage_usage;
|
||
|
||
use windmill_api_auth::{require_devops_role, require_super_admin, ApiAuthed};
|
||
use windmill_common::utils::HTTP_CLIENT_PERMISSIVE as HTTP_CLIENT;
|
||
use windmill_common::DB;
|
||
|
||
use ee_oss::validate_license_key;
|
||
use windmill_common::usernames::generate_instance_username_for_all_users;
|
||
|
||
use axum::{
|
||
body::Body,
|
||
extract::{Extension, Path, Query},
|
||
response::Response,
|
||
routing::{get, post},
|
||
Json, Router,
|
||
};
|
||
|
||
use serde::{Deserialize, Serialize};
|
||
use windmill_ai::ai_cache::bump_instance_ai_config_revision;
|
||
#[cfg(feature = "enterprise")]
|
||
use windmill_common::ee_oss::{send_critical_alert, CriticalAlertKind, CriticalErrorChannel};
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
use windmill_common::secret_backend::{
|
||
AwsSecretsManagerSettings, AzureKeyVaultSettings, SecretMigrationReport, VaultSettings,
|
||
};
|
||
use windmill_common::{
|
||
ee_oss::{get_license_plan, LicensePlan},
|
||
email_oss::{send_email_plain_text, SMTP_ENABLED},
|
||
error::{self, pg_error_message, JsonResult, Result},
|
||
get_database_url,
|
||
global_settings::{
|
||
ACCENT_COLOR_SETTING, AI_CONFIG_SETTING, APP_WORKSPACED_ROUTE_SETTING,
|
||
AUTOMATE_USERNAME_CREATION_SETTING, CRITICAL_ALERT_MUTE_UI_SETTING, CUSTOM_TAGS_SETTING,
|
||
DEFAULT_TAGS_WORKSPACES_SETTING, DISABLE_HUB_SETTING, EMAIL_DOMAIN_SETTING, ENV_SETTINGS,
|
||
EXTERNAL_INSTANCE_PG_SETTING,
|
||
GITHUB_APP_WEBHOOK_BASE_URL_SETTING, HTTP_ROUTE_DEFAULT_ALLOWED_ORIGINS_SETTING,
|
||
HTTP_ROUTE_WORKSPACED_ROUTE_SETTING, HUB_ACCESSIBLE_URL_SETTING, HUB_BASE_URL_SETTING,
|
||
INSTANCE_BANNER_SETTING, MAX_RETENTION_OVERRIDE_WORKSPACES,
|
||
MAX_TOKEN_EXPIRATION_DAYS_SETTING, MCP_DISABLE_TOKEN_QUERY_PARAM_SETTING,
|
||
RETENTION_PERIOD_SECS_OVERRIDES_SETTING, RUFF_CONFIG_SETTING, UNIQUE_ID_SETTING,
|
||
WORKSPACE_FAIRNESS_DURATION_SECS_SETTING, WORKSPACE_FAIRNESS_ENABLED_SETTING,
|
||
WORKSPACE_FAIRNESS_MAX_PERCENT_SETTING, WORKSPACE_FAIRNESS_MIN_TOTAL_SETTING,
|
||
WS_BASE_URL_SETTING,
|
||
},
|
||
instance_config::{self, ApplyMode, InstanceConfig},
|
||
server::Smtp,
|
||
};
|
||
use windmill_common::{
|
||
error::to_anyhow,
|
||
worker::{reload_custom_tags_setting, CLOUD_HOSTED},
|
||
PgDatabase,
|
||
};
|
||
|
||
/// Unauthenticated settings routes.
|
||
///
|
||
/// Used by the extra container (LSP service) to fetch non-sensitive instance
|
||
/// configuration like the shared ruff.toml content without needing to carry
|
||
/// a credential.
|
||
pub fn unauthed_service() -> Router {
|
||
Router::new().route("/ruff_config", get(get_ruff_config_unauthed))
|
||
}
|
||
|
||
/// Public endpoint that returns the instance-level ruff config as plain text
|
||
/// TOML. Returns an empty body when unset.
|
||
///
|
||
/// This is intentionally unauthenticated: ruff config is lint/format policy,
|
||
/// not a credential, and the extra container needs to pull it from any
|
||
/// deployment topology (docker-compose, k8s, local dev) without the extra
|
||
/// burden of shared secrets.
|
||
async fn get_ruff_config_unauthed(Extension(db): Extension<DB>) -> error::Result<Response> {
|
||
let value = sqlx::query_scalar!(
|
||
"SELECT value FROM global_settings WHERE name = $1",
|
||
RUFF_CONFIG_SETTING
|
||
)
|
||
.fetch_optional(&db)
|
||
.await?;
|
||
|
||
let body = value
|
||
.and_then(|v| v.as_str().map(|s| s.to_string()))
|
||
.unwrap_or_default();
|
||
|
||
Ok(Response::builder()
|
||
.status(200)
|
||
.header("content-type", "text/plain; charset=utf-8")
|
||
.header("cache-control", "no-store")
|
||
.body(Body::from(body))
|
||
.unwrap())
|
||
}
|
||
|
||
/// The announcement banner and accent color, which every signed-in session reads on each
|
||
/// full page load.
|
||
async fn get_instance_ui(
|
||
Extension(db): Extension<DB>,
|
||
_authed: ApiAuthed,
|
||
) -> JsonResult<windmill_common::global_settings::InstanceUi> {
|
||
Ok(Json(
|
||
windmill_common::global_settings::get_instance_ui(&db).await?,
|
||
))
|
||
}
|
||
|
||
pub fn global_service() -> Router {
|
||
#[warn(unused_mut)]
|
||
let r = Router::new()
|
||
// `/local` is the path in openapi.yaml, so every generated client (getLocal) calls it;
|
||
// `/envs` stays for callers that found the route in the code.
|
||
.route("/local", get(get_local_settings))
|
||
.route("/envs", get(get_local_settings))
|
||
.route(
|
||
"/global/{key}",
|
||
post(set_global_setting).get(get_global_setting),
|
||
)
|
||
.route("/instance_ui", get(get_instance_ui))
|
||
.route("/list_global", get(list_global_settings))
|
||
.route("/github_app_stale_webhooks", get(github_app_stale_webhooks))
|
||
.route(
|
||
"/instance_config",
|
||
get(get_instance_config).put(set_instance_config),
|
||
)
|
||
.route("/instance_config/yaml", get(get_instance_config_yaml))
|
||
.route("/test_smtp", post(test_email))
|
||
.route("/test_license_key", post(test_license_key))
|
||
.route("/send_stats", post(send_stats))
|
||
.route("/get_stats", get(get_stats))
|
||
.route(
|
||
"/latest_key_renewal_attempt",
|
||
get(get_latest_key_renewal_attempt),
|
||
)
|
||
.route("/renew_license_key", post(renew_license_key))
|
||
.route("/offline_license_status", get(get_offline_license_status))
|
||
.route("/instance_hash", get(get_instance_hash))
|
||
.route("/customer_portal", post(create_customer_portal_session))
|
||
.route("/test_critical_channels", post(test_critical_channels))
|
||
.route("/critical_alerts", get(get_critical_alerts))
|
||
.route(
|
||
"/critical_alerts/{id}/acknowledge",
|
||
post(acknowledge_critical_alert),
|
||
)
|
||
.route(
|
||
"/list_custom_instance_pg_databases",
|
||
post(list_custom_instance_pg_databases),
|
||
)
|
||
.route(
|
||
"/datatable_roles",
|
||
get(datatable_roles_oss::list_datatable_roles)
|
||
.post(datatable_roles_oss::create_datatable_role),
|
||
)
|
||
.route(
|
||
"/datatable_roles/{id}",
|
||
post(datatable_roles_oss::update_datatable_role)
|
||
.delete(datatable_roles_oss::delete_datatable_role),
|
||
)
|
||
.route(
|
||
"/refresh_custom_instance_user_pwd",
|
||
post(refresh_custom_instance_user_pwd),
|
||
)
|
||
.route(
|
||
"/external_instance_pg/status",
|
||
get(get_external_instance_pg_status),
|
||
)
|
||
.route(
|
||
"/external_instance_pg/setup",
|
||
post(setup_external_instance_pg),
|
||
)
|
||
.route(
|
||
"/external_instance_pg/databases",
|
||
get(list_external_instance_pg_databases),
|
||
)
|
||
.route(
|
||
"/external_instance_pg/databases/{name}",
|
||
post(create_external_instance_pg_database).delete(drop_external_instance_pg_database),
|
||
)
|
||
.route(
|
||
"/setup_custom_instance_pg_database/{name}",
|
||
post(setup_custom_instance_pg_database),
|
||
)
|
||
.route(
|
||
"/drop_custom_instance_pg_database/{name}",
|
||
post(drop_custom_instance_pg_database),
|
||
)
|
||
.route(
|
||
"/critical_alerts/acknowledge_all",
|
||
post(acknowledge_all_critical_alerts),
|
||
)
|
||
.route(
|
||
"/sync_cached_resource_types",
|
||
post(sync_cached_resource_types),
|
||
)
|
||
.route(
|
||
"/restart_worker_group/{worker_group}",
|
||
post(restart_worker_group),
|
||
);
|
||
|
||
// Vault/Azure KV integration routes (EE only - requires both private and enterprise features)
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
let r = r
|
||
.route("/test_secret_backend", post(test_secret_backend))
|
||
.route("/migrate_secrets_to_vault", post(migrate_secrets_to_vault))
|
||
.route(
|
||
"/migrate_secrets_to_database",
|
||
post(migrate_secrets_to_database),
|
||
)
|
||
.route("/test_azure_kv_backend", post(test_azure_kv_backend))
|
||
.route(
|
||
"/migrate_secrets_to_azure_kv",
|
||
post(migrate_secrets_to_azure_kv),
|
||
)
|
||
.route(
|
||
"/migrate_secrets_from_azure_kv",
|
||
post(migrate_secrets_from_azure_kv),
|
||
)
|
||
.route("/test_aws_sm_backend", post(test_aws_sm_backend))
|
||
.route(
|
||
"/migrate_secrets_to_aws_sm",
|
||
post(migrate_secrets_to_aws_sm),
|
||
)
|
||
.route(
|
||
"/migrate_secrets_from_aws_sm",
|
||
post(migrate_secrets_from_aws_sm),
|
||
);
|
||
|
||
#[cfg(feature = "parquet")]
|
||
{
|
||
return r
|
||
.route("/test_object_storage_config", post(test_s3_bucket))
|
||
.route(
|
||
"/object_storage_usage",
|
||
get(get_object_storage_usage).post(compute_object_storage_usage),
|
||
)
|
||
.route("/run_log_cleanup", post(run_log_cleanup))
|
||
.route("/log_cleanup_status", get(log_cleanup_status))
|
||
.route("/audit_logs_s3_status", get(audit_logs_s3_status))
|
||
.route("/audit_logs_s3_backfill", post(run_audit_logs_s3_backfill))
|
||
.route(
|
||
"/audit_logs_s3_backfill_status",
|
||
get(audit_logs_s3_backfill_status),
|
||
);
|
||
}
|
||
|
||
#[cfg(not(feature = "parquet"))]
|
||
{
|
||
return r;
|
||
}
|
||
}
|
||
|
||
#[derive(Deserialize)]
|
||
pub struct TestEmail {
|
||
pub to: String,
|
||
pub smtp: Smtp,
|
||
}
|
||
|
||
pub async fn test_email(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(test_email): Json<TestEmail>,
|
||
) -> error::Result<String> {
|
||
require_super_admin(&db, &authed).await?;
|
||
if !SMTP_ENABLED {
|
||
return Err(error::Error::Generic(
|
||
axum::http::StatusCode::NOT_IMPLEMENTED,
|
||
"This Windmill build was compiled without SMTP support, so no email can be sent."
|
||
.to_string(),
|
||
));
|
||
}
|
||
let smtp = test_email.smtp;
|
||
let to = test_email.to;
|
||
|
||
// A connection attempt covers TCP, the TLS handshake, EHLO and authentication against a remote
|
||
// provider; a tighter budget times out before the server ever states why it refused.
|
||
let client_timeout = Duration::from_secs(20);
|
||
send_email_plain_text(
|
||
"Test email from Windmill",
|
||
"Test email content",
|
||
vec![to],
|
||
smtp,
|
||
Some(client_timeout),
|
||
)
|
||
.await
|
||
// The SMTP layer already phrases its failures for an instance admin; the anyhow wrapper it
|
||
// comes back in would bury that behind "Internal: ... @<source location>".
|
||
.map_err(|e| match e {
|
||
error::Error::Anyhow { error, .. } => {
|
||
error::Error::Generic(axum::http::StatusCode::BAD_REQUEST, format!("{error:#}"))
|
||
}
|
||
e => e,
|
||
})?;
|
||
|
||
Ok("Sent test email".to_string())
|
||
}
|
||
|
||
#[cfg(feature = "parquet")]
|
||
use windmill_object_store::ObjectSettings;
|
||
|
||
#[cfg(feature = "parquet")]
|
||
use windmill_object_store::{
|
||
build_object_store_from_settings, build_public_object_store_from_settings,
|
||
};
|
||
|
||
#[cfg(feature = "parquet")]
|
||
pub async fn test_s3_bucket(
|
||
authed: ApiAuthed,
|
||
Extension(db): Extension<DB>,
|
||
Json(test_s3_bucket): Json<ObjectSettings>,
|
||
) -> error::Result<String> {
|
||
use bytes::Bytes;
|
||
use futures::StreamExt;
|
||
|
||
// The probe executes on the API server itself and reflects the upstream response into the
|
||
// error, so any authenticated caller could otherwise use it as an SSRF / port-scan primitive
|
||
// against the server's network, exfiltrate its ambient credentials, or write to its local
|
||
// disk (see validate_object_storage_test). That holds on self-hosted instances as much as on
|
||
// Cloud, so only super admins get the unrestricted path.
|
||
let is_super_admin = windmill_api_auth::is_super_admin_authed(&db, &authed).await?;
|
||
let restrict = !is_super_admin;
|
||
if restrict {
|
||
validate_object_storage_test(&test_s3_bucket)
|
||
.await
|
||
.map_err(|e| match e {
|
||
// A job token never counts as a super admin (it is capped at workspace admin), so
|
||
// a super admin calling this route from a script is told why rather than that
|
||
// they lack a privilege they hold.
|
||
error::Error::NotAuthorized(msg) if authed.job_id.is_some() => {
|
||
error::Error::NotAuthorized(format!(
|
||
"{msg} A job token ($WM_TOKEN) is never treated as a super admin; call \
|
||
this route with a user token instead."
|
||
))
|
||
}
|
||
e => e,
|
||
})?;
|
||
}
|
||
|
||
// The restricted client re-judges every address it connects to: the checks above resolve the
|
||
// endpoint separately from the connect, which a rebinding name answers differently.
|
||
let client = if restrict {
|
||
build_public_object_store_from_settings(test_s3_bucket).await?
|
||
} else {
|
||
build_object_store_from_settings(test_s3_bucket, Some(&db))
|
||
.await?
|
||
.store
|
||
};
|
||
|
||
let run = async {
|
||
let mut list = client.list(Some(
|
||
&windmill_object_store::object_store_reexports::Path::from("".to_string()),
|
||
));
|
||
match list.next().await {
|
||
Some(Err(e)) => {
|
||
tracing::error!("error listing bucket: {e:#}");
|
||
return Err(error::Error::internal_err(format!(
|
||
"Failed to list files in blob storage: {e:#}"
|
||
)));
|
||
}
|
||
Some(Ok(first_file)) => tracing::info!("Listed files: {:?}", first_file),
|
||
None => tracing::info!("No files in blob storage"),
|
||
}
|
||
|
||
let path = windmill_object_store::object_store_reexports::Path::from(format!(
|
||
"/test-s3-bucket-{uuid}",
|
||
uuid = uuid::Uuid::new_v4()
|
||
));
|
||
tracing::info!("Testing blob storage at path: {path}");
|
||
client
|
||
.put(
|
||
&path,
|
||
windmill_object_store::object_store_reexports::PutPayload::from_static(b"hello"),
|
||
)
|
||
.await
|
||
.map_err(|e| anyhow::anyhow!("error writing file to {path}: {e:#}"))?;
|
||
let content = client
|
||
.get(&path)
|
||
.await
|
||
.map_err(to_anyhow)?
|
||
.bytes()
|
||
.await
|
||
.map_err(to_anyhow)?;
|
||
if content != Bytes::from_static(b"hello") {
|
||
return Err(error::Error::internal_err(
|
||
"Failed to read back from blob storage".to_string(),
|
||
));
|
||
}
|
||
client.delete(&path).await.map_err(to_anyhow)?;
|
||
Ok::<String, error::Error>("Tested blob storage successfully".to_string())
|
||
};
|
||
|
||
if restrict {
|
||
// The object-store client is built with timeouts disabled, so a malicious endpoint could
|
||
// otherwise hold the API server connection open indefinitely.
|
||
tokio::time::timeout(Duration::from_secs(15), run)
|
||
.await
|
||
.map_err(|_| {
|
||
error::Error::internal_err("Object storage connectivity test timed out".to_string())
|
||
})?
|
||
} else {
|
||
run.await
|
||
}
|
||
}
|
||
|
||
// Hardening for the object-storage connectivity test by an untrusted (non-super-admin) caller.
|
||
// The probe runs on the API server, so without these constraints an authenticated
|
||
// user could coerce the server into connecting to arbitrary internal endpoints (SSRF), signing
|
||
// requests with the instance role (credential exfiltration), or reading/writing the server's local
|
||
// disk (filesystem object store).
|
||
#[cfg(feature = "parquet")]
|
||
async fn validate_object_storage_test(settings: &ObjectSettings) -> error::Result<()> {
|
||
fn non_empty(opt: &Option<String>) -> bool {
|
||
opt.as_ref().is_some_and(|s| !s.is_empty())
|
||
}
|
||
|
||
// Every refusal names the way out: the resource usually works in jobs (workers reach the
|
||
// endpoint directly), so without it the refusal reads as a broken resource.
|
||
const ALTERNATIVE: &str =
|
||
"Ask a super admin to run it, or test the resource from a script, which runs on a worker.";
|
||
|
||
// Reject backends that rely on the server's identity or local filesystem, require explicit
|
||
// credentials for the rest (so the server never falls back to its own ambient credentials), and
|
||
// resolve the host the client will actually connect to. We derive the *effective* endpoint here
|
||
// — mirroring build_*_from_settings: the region/account-derived default and the virtual-hosted
|
||
// bucket prefix — rather than only validating a caller-supplied `endpoint`, so caller-controlled
|
||
// `region`/`account_name`/`bucket` cannot smuggle an internal host past the check (e.g. an empty
|
||
// endpoint with region = "@169.254.169.254/" otherwise resolves to the cloud metadata service).
|
||
let effective_endpoint: Option<String> = match settings {
|
||
ObjectSettings::Filesystem(_) => {
|
||
return Err(error::Error::NotAuthorized(
|
||
"Testing a local filesystem object store requires a super admin: it runs on the \
|
||
Windmill server and reads and writes the server's local disk. Ask a super admin \
|
||
to run it."
|
||
.to_string(),
|
||
));
|
||
}
|
||
ObjectSettings::AwsOidc(_) => {
|
||
return Err(error::Error::NotAuthorized(format!(
|
||
"Testing OIDC-based object storage requires a super admin: it runs on the \
|
||
Windmill server with the server's own identity. {ALTERNATIVE}"
|
||
)));
|
||
}
|
||
ObjectSettings::S3(s3) => {
|
||
if !(non_empty(&s3.access_key) && non_empty(&s3.secret_key)) {
|
||
return Err(error::Error::NotAuthorized(format!(
|
||
"Testing S3 storage without an explicit access key and secret key requires a \
|
||
super admin: it runs on the Windmill server, which would use its own ambient \
|
||
credentials. {ALTERNATIVE}"
|
||
)));
|
||
}
|
||
let region = s3
|
||
.region
|
||
.clone()
|
||
.filter(|r| !r.is_empty())
|
||
.or_else(|| std::env::var("AWS_REGION").ok().filter(|r| !r.is_empty()))
|
||
.unwrap_or_else(|| "us-east-1".to_string());
|
||
let raw_endpoint = s3
|
||
.endpoint
|
||
.clone()
|
||
.filter(|e| !e.is_empty())
|
||
.or_else(|| std::env::var("S3_ENDPOINT").ok().filter(|e| !e.is_empty()))
|
||
.unwrap_or_else(|| format!("s3.{region}.amazonaws.com"));
|
||
Some(windmill_object_store::render_endpoint(
|
||
raw_endpoint,
|
||
!s3.allow_http.unwrap_or(true),
|
||
s3.port,
|
||
s3.path_style,
|
||
s3.bucket.clone().unwrap_or_default(),
|
||
))
|
||
}
|
||
ObjectSettings::Azure(azure) => {
|
||
if !non_empty(&azure.access_key) {
|
||
return Err(error::Error::NotAuthorized(format!(
|
||
"Testing Azure storage without an explicit access key requires a super admin: \
|
||
it runs on the Windmill server, which would use its own ambient credentials. \
|
||
{ALTERNATIVE}"
|
||
)));
|
||
}
|
||
Some(
|
||
azure
|
||
.endpoint
|
||
.clone()
|
||
.filter(|e| !e.is_empty())
|
||
.unwrap_or_else(|| format!("{}.blob.core.windows.net", azure.account_name)),
|
||
)
|
||
}
|
||
ObjectSettings::Gcs(gcs) => {
|
||
// Mirror `build_gcs_client`'s blank-key check (shared predicate): a blank/`{}` key falls
|
||
// back to the instance's ambient credentials there, so it must be rejected here too —
|
||
// otherwise an untrusted caller could probe with the server's identity (the very
|
||
// SSRF/credential-exfil this function guards against).
|
||
if windmill_object_store::gcs_service_account_key_is_blank(&gcs.service_account_key) {
|
||
return Err(error::Error::NotAuthorized(format!(
|
||
"Testing GCS storage without a service account key requires a super admin: \
|
||
it runs on the Windmill server, which would use its own ambient credentials. \
|
||
{ALTERNATIVE}"
|
||
)));
|
||
}
|
||
// The service-account-key JSON can override the data-plane URL (`gcs_base_url`) and the
|
||
// OAuth token endpoint (`token_uri`); the GCS client connects to whatever they point at.
|
||
// Validate every http(s) URL embedded in the key. When none override it, the host stays
|
||
// the public storage.googleapis.com, so no further check is needed.
|
||
if let Ok(serde_json::Value::Object(map)) =
|
||
serde_json::from_str::<serde_json::Value>(&gcs.service_account_key)
|
||
{
|
||
for value in map.values() {
|
||
if let Some(url) = value.as_str() {
|
||
// Match how the URL parser reads the value: leading whitespace/control is
|
||
// ignored and the scheme is case-insensitive.
|
||
let url =
|
||
url.trim_start_matches(|c: char| c.is_whitespace() || c.is_control());
|
||
if strip_http_scheme(url).is_some() {
|
||
validate_public_endpoint(url).await?;
|
||
}
|
||
}
|
||
}
|
||
}
|
||
None
|
||
}
|
||
};
|
||
|
||
// Block non-public network targets (internal services, cloud metadata, loopback, ...).
|
||
if let Some(endpoint) = effective_endpoint {
|
||
validate_public_endpoint(&endpoint).await?;
|
||
}
|
||
Ok(())
|
||
}
|
||
|
||
#[cfg(feature = "parquet")]
|
||
async fn validate_public_endpoint(endpoint: &str) -> error::Result<()> {
|
||
let host = extract_host(endpoint).ok_or_else(|| {
|
||
error::Error::BadRequest(format!("Invalid object storage endpoint: {endpoint}"))
|
||
})?;
|
||
|
||
let addrs: Vec<std::net::SocketAddr> = tokio::net::lookup_host((host.as_str(), 443u16))
|
||
.await
|
||
.map_err(|e| {
|
||
error::Error::BadRequest(format!(
|
||
"Could not resolve object storage endpoint '{host}': {e}"
|
||
))
|
||
})?
|
||
.collect();
|
||
|
||
if addrs.is_empty() {
|
||
return Err(error::Error::BadRequest(format!(
|
||
"Could not resolve object storage endpoint '{host}'"
|
||
)));
|
||
}
|
||
|
||
// Reject if any resolved address is non-public, which also defeats the simplest DNS-rebinding
|
||
// attempts (a name resolving to both a public and a private address).
|
||
for addr in addrs {
|
||
if windmill_common::ssrf::is_private_ip(&addr.ip()) {
|
||
// The resolved address stays out of the message: it is the server's resolver's
|
||
// answer, and this message is only ever shown to the caller being constrained.
|
||
return Err(error::Error::NotAuthorized(format!(
|
||
"Testing object storage at '{host}', which resolves to a private, loopback, or \
|
||
link-local address, requires a super admin: this test runs on the Windmill \
|
||
server, which is not allowed to probe internal addresses for non-super-admins. \
|
||
Ask a super admin to run it, or test the resource from a script, which runs on \
|
||
a worker."
|
||
)));
|
||
}
|
||
}
|
||
Ok(())
|
||
}
|
||
|
||
// Strip a leading `http://`/`https://` scheme case-insensitively (URL schemes are
|
||
// case-insensitive), returning the remainder when one was present.
|
||
#[cfg(feature = "parquet")]
|
||
fn strip_http_scheme(s: &str) -> Option<&str> {
|
||
for scheme in ["https://", "http://"] {
|
||
let b = scheme.as_bytes();
|
||
if s.len() >= b.len() && s.as_bytes()[..b.len()].eq_ignore_ascii_case(b) {
|
||
return Some(&s[b.len()..]);
|
||
}
|
||
}
|
||
None
|
||
}
|
||
|
||
#[cfg(feature = "parquet")]
|
||
fn extract_host(endpoint: &str) -> Option<String> {
|
||
let mut s = endpoint.trim();
|
||
if let Some(rest) = strip_http_scheme(s) {
|
||
s = rest;
|
||
}
|
||
s = s.split(['/', '?', '#', '\\']).next().unwrap_or(s);
|
||
if let Some((_, rest)) = s.rsplit_once('@') {
|
||
s = rest;
|
||
}
|
||
let host = if let Some(rest) = s.strip_prefix('[') {
|
||
// IPv6 literal, e.g. [::1]:9000
|
||
rest.split(']').next().unwrap_or(rest)
|
||
} else {
|
||
// host or host:port
|
||
s.split(':').next().unwrap_or(s)
|
||
}
|
||
.trim();
|
||
if host.is_empty() {
|
||
None
|
||
} else {
|
||
Some(host.to_string())
|
||
}
|
||
}
|
||
|
||
#[cfg(feature = "parquet")]
|
||
async fn get_object_storage_usage(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::JsonResult<Option<storage_usage::StorageUsageProgress>> {
|
||
require_super_admin(&db, &authed).await?;
|
||
Ok(Json(storage_usage::get_status(&db).await?))
|
||
}
|
||
|
||
#[cfg(feature = "parquet")]
|
||
async fn compute_object_storage_usage(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::Result<axum::http::StatusCode> {
|
||
require_super_admin(&db, &authed).await?;
|
||
storage_usage::try_start(&db).await?;
|
||
storage_usage::spawn_compute(db.clone());
|
||
Ok(axum::http::StatusCode::ACCEPTED)
|
||
}
|
||
|
||
#[cfg(feature = "parquet")]
|
||
async fn run_log_cleanup(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::Result<axum::http::StatusCode> {
|
||
require_super_admin(&db, &authed).await?;
|
||
log_cleanup::try_start(&db).await?;
|
||
log_cleanup::spawn_cleanup(db.clone());
|
||
Ok(axum::http::StatusCode::ACCEPTED)
|
||
}
|
||
|
||
#[cfg(feature = "parquet")]
|
||
async fn log_cleanup_status(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::JsonResult<Option<log_cleanup::LogCleanupProgress>> {
|
||
require_super_admin(&db, &authed).await?;
|
||
Ok(Json(log_cleanup::get_status(&db).await?))
|
||
}
|
||
|
||
#[cfg(feature = "parquet")]
|
||
async fn audit_logs_s3_status(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::JsonResult<Option<audit_logs_s3::AuditLogsS3ExportStatus>> {
|
||
require_super_admin(&db, &authed).await?;
|
||
Ok(Json(audit_logs_s3::get_status(&db).await?))
|
||
}
|
||
|
||
#[cfg(feature = "parquet")]
|
||
async fn run_audit_logs_s3_backfill(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(req): Json<audit_logs_s3_backfill::BackfillRequest>,
|
||
) -> error::Result<axum::http::StatusCode> {
|
||
require_super_admin(&db, &authed).await?;
|
||
if !matches!(get_license_plan().await, LicensePlan::Enterprise) {
|
||
return Err(error::Error::BadRequest(
|
||
"Audit log export to object storage is an Enterprise feature".to_string(),
|
||
));
|
||
}
|
||
audit_logs_s3_backfill::try_start(&db, req.from, req.to).await?;
|
||
audit_logs_s3_backfill::spawn_backfill(db.clone(), req.from, req.to);
|
||
Ok(axum::http::StatusCode::ACCEPTED)
|
||
}
|
||
|
||
#[cfg(feature = "parquet")]
|
||
async fn audit_logs_s3_backfill_status(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::JsonResult<Option<audit_logs_s3_backfill::AuditBackfillProgress>> {
|
||
require_super_admin(&db, &authed).await?;
|
||
Ok(Json(audit_logs_s3_backfill::get_status(&db).await?))
|
||
}
|
||
|
||
#[derive(Deserialize)]
|
||
pub struct TestKey {
|
||
pub license_key: String,
|
||
}
|
||
|
||
pub async fn test_license_key(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(TestKey { license_key }): Json<TestKey>,
|
||
) -> error::Result<String> {
|
||
require_super_admin(&db, &authed).await?;
|
||
let (_, expired, _offline_meta) = validate_license_key(license_key, Some(&db)).await?;
|
||
|
||
if expired {
|
||
Err(error::Error::BadRequest("Expired license key".to_string()))
|
||
} else {
|
||
Ok("Valid license key".to_string())
|
||
}
|
||
}
|
||
|
||
#[derive(serde::Serialize)]
|
||
pub struct InstanceHash {
|
||
pub instance_hash: Option<String>,
|
||
}
|
||
|
||
/// Returns the live cap status for an offline license, or `null` when no
|
||
/// offline license is loaded. Used by the superadmin settings panel.
|
||
pub async fn get_offline_license_status(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::JsonResult<Option<windmill_common::ee_oss::OfflineCapStatus>> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
let offline = (**windmill_common::ee_oss::LICENSE_OFFLINE_METADATA.load()).clone();
|
||
let is_offline = matches!(&offline, Some(m) if m.is_offline());
|
||
|
||
if !is_offline {
|
||
return Ok(Json(None));
|
||
}
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
let cap = windmill_common::ee_oss::enforce_offline_caps(&db)
|
||
.await
|
||
.map_err(|e| error::Error::internal_err(format!("enforce_offline_caps: {e:#}")))?;
|
||
#[cfg(not(feature = "enterprise"))]
|
||
let cap: Option<windmill_common::ee_oss::OfflineCapStatus> = None;
|
||
|
||
Ok(Json(cap))
|
||
}
|
||
|
||
/// Returns the per-instance binding hash that goes into offline license keys.
|
||
/// Admin invokes via `curl` with their personal token when requesting a key
|
||
/// from support.
|
||
pub async fn get_instance_hash(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::JsonResult<InstanceHash> {
|
||
require_super_admin(&db, &authed).await?;
|
||
#[cfg(feature = "enterprise")]
|
||
let hash = windmill_common::ee_oss::compute_instance_hash(&db)
|
||
.await
|
||
.map_err(|e| error::Error::internal_err(format!("compute_instance_hash: {e:#}")))?;
|
||
#[cfg(not(feature = "enterprise"))]
|
||
let hash: Option<String> = None;
|
||
Ok(Json(InstanceHash { instance_hash: hash }))
|
||
}
|
||
|
||
pub async fn get_local_settings(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::JsonResult<serde_json::Value> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
let mut settings = serde_json::Map::new();
|
||
for key in ENV_SETTINGS.iter() {
|
||
if let Some(value) = std::env::var(key).ok() {
|
||
settings.insert(key.to_string(), serde_json::Value::String(value));
|
||
}
|
||
}
|
||
Ok(Json(serde_json::Value::Object(settings)))
|
||
}
|
||
|
||
#[derive(serde::Deserialize)]
|
||
pub struct Value {
|
||
pub value: Option<serde_json::Value>,
|
||
}
|
||
|
||
pub async fn delete_global_setting(db: &DB, key: &str) -> error::Result<()> {
|
||
// ducklake_user_pg_pwd and ducklake_settings were old names stored as standalone global settings.
|
||
// Leave them for backward compatibility (CLI will try to delete them if not present in the yaml)
|
||
if key == "ducklake_user_pg_pwd"
|
||
|| key == "ducklake_settings"
|
||
|| key == "custom_instance_pg_databases"
|
||
{
|
||
tracing::error!("Tried to unset global setting {}, ignored", key);
|
||
return Ok(());
|
||
}
|
||
sqlx::query!("DELETE FROM global_settings WHERE name = $1", key,)
|
||
.execute(db)
|
||
.await?;
|
||
tracing::info!("Unset global setting {}", key);
|
||
Ok(())
|
||
}
|
||
/// Returns true when `key` is one of the workspace-fairness settings whose
|
||
/// writes must be gated to cloud only.
|
||
fn is_workspace_fairness_setting(key: &str) -> bool {
|
||
matches!(
|
||
key,
|
||
WORKSPACE_FAIRNESS_ENABLED_SETTING
|
||
| WORKSPACE_FAIRNESS_MAX_PERCENT_SETTING
|
||
| WORKSPACE_FAIRNESS_DURATION_SECS_SETTING
|
||
| WORKSPACE_FAIRNESS_MIN_TOTAL_SETTING
|
||
)
|
||
}
|
||
|
||
/// Enterprise gate for the workspace-fairness settings. Workspace fairness is
|
||
/// only useful on multi-tenant clusters where one workspace can starve other
|
||
/// workspaces sharing the same worker pool, and the feature is licensed as
|
||
/// part of Enterprise. Non-EE installs are rejected at write time; the runtime
|
||
/// dispatch additionally honours the `WORKSPACE_FAIRNESS_ENABLED` toggle.
|
||
async fn workspace_fairness_settings_allowed() -> bool {
|
||
matches!(get_license_plan().await, LicensePlan::Enterprise)
|
||
}
|
||
|
||
pub async fn set_global_setting(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Path(key): Path<String>,
|
||
Json(value): Json<Value>,
|
||
) -> error::Result<()> {
|
||
require_super_admin(&db, &authed).await?;
|
||
set_global_setting_internal(&db, key, value.value.unwrap_or(serde_json::Value::Null)).await
|
||
}
|
||
|
||
pub async fn set_global_setting_internal(
|
||
db: &DB,
|
||
key: String,
|
||
value: serde_json::Value,
|
||
) -> error::Result<()> {
|
||
let should_bump_instance_ai_revision = key == AI_CONFIG_SETTING;
|
||
let value = if key == "retention_period_secs" {
|
||
instance_config::clamp_retention_period(value)
|
||
} else {
|
||
value
|
||
};
|
||
|
||
// EE gate for workspace-fairness settings. Workspace fairness only matters
|
||
// on multi-tenant clusters; it is licensed as an Enterprise feature so the
|
||
// setter rejects writes from non-EE builds. Disabling/clearing writes are
|
||
// *always* allowed regardless of license, so an admin who downgrades from
|
||
// EE (or imports a row from a cloned EE DB) can always turn the cap off:
|
||
// - `Null` / empty-string → row delete
|
||
// - `Bool(false)` on `workspace_fairness_enabled` → explicit disable
|
||
// Without the `Bool(false)` carve-out, a stale `enabled=true` row from a
|
||
// downgrade would be impossible to flip off through the normal API/UI
|
||
// and the runtime path (which only checks the toggle) would keep
|
||
// throttling.
|
||
let is_clearing_value = matches!(&value, serde_json::Value::Null)
|
||
|| matches!(&value, serde_json::Value::String(s) if s.trim().is_empty())
|
||
|| (key == WORKSPACE_FAIRNESS_ENABLED_SETTING
|
||
&& matches!(&value, serde_json::Value::Bool(false)));
|
||
if is_workspace_fairness_setting(&key)
|
||
&& !is_clearing_value
|
||
&& !workspace_fairness_settings_allowed().await
|
||
{
|
||
return Err(error::Error::BadRequest(format!(
|
||
"{} requires an Enterprise license",
|
||
key
|
||
)));
|
||
}
|
||
|
||
if key == EXTERNAL_INSTANCE_PG_SETTING {
|
||
return windmill_common::external_instance_pg::write_external_instance_pg_setting(
|
||
db,
|
||
Some(&value),
|
||
)
|
||
.await;
|
||
}
|
||
|
||
run_setting_pre_write_hook(db, &key, &value).await?;
|
||
|
||
match value {
|
||
serde_json::Value::Null => {
|
||
if instance_config::PROTECTED_SETTINGS.contains(&key.as_str()) {
|
||
return Err(error::Error::BadRequest(format!(
|
||
"{key} is a protected setting and cannot be deleted"
|
||
)));
|
||
}
|
||
delete_global_setting(db, &key).await?;
|
||
}
|
||
serde_json::Value::String(ref x) if x.trim().is_empty() => {
|
||
if instance_config::PROTECTED_SETTINGS.contains(&key.as_str()) {
|
||
return Err(error::Error::BadRequest(format!(
|
||
"{key} is a protected setting and cannot be set to empty"
|
||
)));
|
||
}
|
||
delete_global_setting(db, &key).await?;
|
||
}
|
||
v => {
|
||
sqlx::query!(
|
||
"INSERT INTO global_settings (name, value) VALUES ($1, $2) ON CONFLICT (name) DO UPDATE SET value = EXCLUDED.value, updated_at = now()",
|
||
key,
|
||
v
|
||
)
|
||
.execute(db)
|
||
.await?;
|
||
tracing::info!(
|
||
"Set global setting {} to {}",
|
||
key,
|
||
instance_config::format_setting_value(&key, &v)
|
||
);
|
||
}
|
||
};
|
||
|
||
if should_bump_instance_ai_revision {
|
||
bump_instance_ai_config_revision();
|
||
}
|
||
|
||
// Tag reads are served from an in-memory cache that this process otherwise only
|
||
// refreshes on the next global-settings poll, so without this a refetch right after
|
||
// the write still returns the pre-write list. The setting is already persisted at
|
||
// this point, so a failed refresh must not be reported as a failed write — the
|
||
// poller retries it.
|
||
if key == CUSTOM_TAGS_SETTING {
|
||
if let Err(e) = reload_custom_tags_setting(db).await {
|
||
tracing::error!(error = %e, "Could not reload custom tags setting after write");
|
||
}
|
||
}
|
||
|
||
Ok(())
|
||
}
|
||
|
||
/// Run side-effect hooks for specific settings before writing to DB.
|
||
/// Extracted from `set_global_setting_internal` for reuse by the bulk endpoint.
|
||
async fn run_setting_pre_write_hook(
|
||
db: &DB,
|
||
key: &str,
|
||
value: &serde_json::Value,
|
||
) -> error::Result<()> {
|
||
match key {
|
||
// The instance AI config is written as an untyped blob through this generic
|
||
// endpoint, so it never passes the typed check the workspace handler applies.
|
||
// Rates that reach a cost total unbounded would make it negative or infinite.
|
||
AI_CONFIG_SETTING => {
|
||
windmill_ai::ai_types::validate_model_pricing_json(value)
|
||
.map_err(error::Error::BadRequest)?;
|
||
windmill_ai::ai_types::validate_token_maps_json(value)
|
||
.map_err(error::Error::BadRequest)?;
|
||
}
|
||
AUTOMATE_USERNAME_CREATION_SETTING => {
|
||
if value.as_bool().unwrap_or(false) {
|
||
generate_instance_username_for_all_users(db)
|
||
.await
|
||
.map_err(|err| {
|
||
error::Error::internal_err(format!(
|
||
"Failed to generate instance wide usernames: {}",
|
||
err
|
||
))
|
||
})?;
|
||
} else {
|
||
// Disabling is only allowed before any instance-wide username has been
|
||
// assigned. Once usernames exist they are globally unique and are baked
|
||
// into stored `u/<username>` identities (schedules, triggers, drafts,
|
||
// and non-member superadmin ownership). Disabling would drop back to
|
||
// workspace-local username uniqueness, letting a member reuse an
|
||
// existing instance username and silently take over those identities —
|
||
// so the setting is effectively one-way once derivation has taken
|
||
// effect. Re-saving `false` on an already-disabled instance is a no-op
|
||
// and stays allowed (guarded by the current-value check).
|
||
let currently_enabled = sqlx::query_scalar!(
|
||
"SELECT value FROM global_settings WHERE name = $1",
|
||
AUTOMATE_USERNAME_CREATION_SETTING
|
||
)
|
||
.fetch_optional(db)
|
||
.await?
|
||
.and_then(|v| v.as_bool())
|
||
.unwrap_or(true);
|
||
if currently_enabled {
|
||
let usernames_exist = sqlx::query_scalar!(
|
||
"SELECT EXISTS(SELECT 1 FROM password WHERE username IS NOT NULL)"
|
||
)
|
||
.fetch_one(db)
|
||
.await?
|
||
.unwrap_or(false);
|
||
if usernames_exist {
|
||
return Err(error::Error::BadRequest(
|
||
"automate_username_creation cannot be disabled once instance-wide usernames have been assigned: existing u/<username> identities (schedules, triggers, drafts, superadmin ownership) rely on those usernames staying stable and globally unique.".to_string(),
|
||
));
|
||
}
|
||
}
|
||
}
|
||
}
|
||
CRITICAL_ALERT_MUTE_UI_SETTING => {
|
||
if value.as_bool().unwrap_or(false) {
|
||
sqlx::query!("UPDATE alerts SET acknowledged = true")
|
||
.execute(db)
|
||
.await?;
|
||
}
|
||
}
|
||
APP_WORKSPACED_ROUTE_SETTING => {
|
||
let serde_json::Value::Bool(workspaced_route) = value else {
|
||
return Err(error::Error::BadRequest(format!(
|
||
"{} setting Expected to be boolean",
|
||
APP_WORKSPACED_ROUTE_SETTING
|
||
)));
|
||
};
|
||
|
||
// Cloud always scopes app custom paths by workspace_id (see
|
||
// `custom_path_exists` in apps.rs), so duplicates across workspaces
|
||
// are expected and this setting has no runtime effect on cloud.
|
||
if !*workspaced_route && !*CLOUD_HOSTED {
|
||
#[derive(Debug, Deserialize, Serialize)]
|
||
#[allow(unused)]
|
||
struct DuplicateApp {
|
||
custom_path: Option<String>,
|
||
path: String,
|
||
}
|
||
let duplicate_app = sqlx::query_as!(
|
||
DuplicateApp,
|
||
r#"
|
||
SELECT
|
||
path,
|
||
custom_path
|
||
FROM
|
||
app
|
||
WHERE
|
||
custom_path IN (
|
||
SELECT
|
||
custom_path
|
||
FROM
|
||
app
|
||
GROUP
|
||
BY custom_path
|
||
HAVING COUNT(*) > 1
|
||
)
|
||
ORDER BY custom_path
|
||
"#
|
||
)
|
||
.fetch_all(db)
|
||
.await?;
|
||
|
||
if !duplicate_app.is_empty() {
|
||
tracing::error!(
|
||
"Cannot disable {} setting as duplicate app with custom path were found: {:?}",
|
||
APP_WORKSPACED_ROUTE_SETTING,
|
||
&duplicate_app
|
||
);
|
||
|
||
#[derive(Serialize)]
|
||
struct ErrorResponse {
|
||
error: String,
|
||
details: Vec<DuplicateApp>,
|
||
}
|
||
|
||
let error_response = ErrorResponse {
|
||
error: "Duplicate custom paths detected".to_string(),
|
||
details: duplicate_app,
|
||
};
|
||
|
||
return Err(error::Error::JsonErr(
|
||
serde_json::to_value(error_response).unwrap(),
|
||
));
|
||
}
|
||
}
|
||
}
|
||
HTTP_ROUTE_DEFAULT_ALLOWED_ORIGINS_SETTING => {
|
||
// Rejected at write time rather than at boot: a mistyped origin
|
||
// matches no request, so it would silently block the very app it
|
||
// names with nothing but a log line to go on.
|
||
windmill_common::global_settings::parse_allowed_origins_setting(Some(value))?;
|
||
}
|
||
HTTP_ROUTE_WORKSPACED_ROUTE_SETTING => {
|
||
let serde_json::Value::Bool(workspaced_route) = value else {
|
||
return Err(error::Error::BadRequest(format!(
|
||
"{} setting expected to be boolean",
|
||
HTTP_ROUTE_WORKSPACED_ROUTE_SETTING
|
||
)));
|
||
};
|
||
|
||
// Cloud always scopes routes by workspace_id (see
|
||
// `route_path_key_exists` in windmill-trigger-http), so duplicates
|
||
// across workspaces are expected and this setting has no runtime
|
||
// effect on cloud.
|
||
if !*workspaced_route && !*CLOUD_HOSTED {
|
||
#[derive(Debug, Deserialize, Serialize)]
|
||
#[allow(unused)]
|
||
struct DuplicateRoute {
|
||
route_path: String,
|
||
workspace_id: String,
|
||
http_method: String,
|
||
}
|
||
let duplicate_routes = sqlx::query_as!(
|
||
DuplicateRoute,
|
||
r#"
|
||
SELECT
|
||
route_path,
|
||
workspace_id,
|
||
http_method::TEXT AS "http_method!"
|
||
FROM
|
||
http_trigger
|
||
WHERE
|
||
workspaced_route IS FALSE
|
||
AND route_path_key IN (
|
||
SELECT
|
||
route_path_key
|
||
FROM
|
||
http_trigger
|
||
WHERE
|
||
workspaced_route IS FALSE
|
||
GROUP BY
|
||
route_path_key, http_method
|
||
HAVING COUNT(*) > 1
|
||
)
|
||
ORDER BY route_path_key
|
||
"#
|
||
)
|
||
.fetch_all(db)
|
||
.await?;
|
||
|
||
if !duplicate_routes.is_empty() {
|
||
tracing::error!(
|
||
"Cannot disable {} setting as duplicate http routes were found: {:?}",
|
||
HTTP_ROUTE_WORKSPACED_ROUTE_SETTING,
|
||
&duplicate_routes
|
||
);
|
||
|
||
#[derive(Serialize)]
|
||
struct ErrorResponse {
|
||
error: String,
|
||
details: Vec<DuplicateRoute>,
|
||
}
|
||
|
||
let error_response = ErrorResponse {
|
||
error: "Duplicate HTTP route paths detected".to_string(),
|
||
details: duplicate_routes,
|
||
};
|
||
|
||
return Err(error::Error::JsonErr(
|
||
serde_json::to_value(error_response).unwrap(),
|
||
));
|
||
}
|
||
}
|
||
}
|
||
RETENTION_PERIOD_SECS_OVERRIDES_SETTING => {
|
||
// Reject a malformed map at write time so it can never be persisted. A persisted bad
|
||
// value (negative or non-integer) would fail to parse on the next server start and,
|
||
// because the loader fails closed (skips cleanup until a known-good value is read),
|
||
// silently disable ALL job-retention cleanup indefinitely. This shape check must stay in
|
||
// sync with `parse_retention_overrides` in backend/src/monitor.rs.
|
||
match value {
|
||
// Clearing (delete row) is handled by the caller; allow it through.
|
||
serde_json::Value::Null => {}
|
||
serde_json::Value::String(s) if s.trim().is_empty() => {}
|
||
serde_json::Value::Object(map) => {
|
||
if map.len() > MAX_RETENTION_OVERRIDE_WORKSPACES {
|
||
return Err(error::Error::BadRequest(format!(
|
||
"retention_period_secs_overrides: at most {MAX_RETENTION_OVERRIDE_WORKSPACES} per-workspace overrides are allowed, got {}",
|
||
map.len()
|
||
)));
|
||
}
|
||
for (ws, v) in map {
|
||
if !v.as_i64().is_some_and(|secs| secs >= 0) {
|
||
return Err(error::Error::BadRequest(format!(
|
||
"retention_period_secs_overrides: override for '{ws}' must be a non-negative integer number of seconds, got {v}"
|
||
)));
|
||
}
|
||
}
|
||
}
|
||
_ => {
|
||
return Err(error::Error::BadRequest(
|
||
"retention_period_secs_overrides must be a JSON object of {workspace_id: seconds}".to_string(),
|
||
));
|
||
}
|
||
}
|
||
}
|
||
GITHUB_APP_WEBHOOK_BASE_URL_SETTING => {
|
||
// A bad value here yields a webhook GitHub can never deliver to, and the
|
||
// failure only shows up much later as "falling back to polling" on a
|
||
// repository — so reject it at the boundary instead.
|
||
match value {
|
||
// Clearing (delete row) is handled by the caller; allow it through.
|
||
serde_json::Value::Null => {}
|
||
serde_json::Value::String(s) if s.trim().is_empty() => {}
|
||
serde_json::Value::String(s) => {
|
||
windmill_common::global_settings::validate_webhook_base_url(s).map_err(
|
||
|e| {
|
||
error::Error::BadRequest(format!(
|
||
"{GITHUB_APP_WEBHOOK_BASE_URL_SETTING}: {e}"
|
||
))
|
||
},
|
||
)?;
|
||
}
|
||
_ => {
|
||
return Err(error::Error::BadRequest(format!(
|
||
"{GITHUB_APP_WEBHOOK_BASE_URL_SETTING} must be a URL string"
|
||
)));
|
||
}
|
||
}
|
||
}
|
||
MAX_TOKEN_EXPIRATION_DAYS_SETTING => {
|
||
windmill_common::global_settings::parse_max_token_expiration_days(Some(value))
|
||
.map_err(|e| {
|
||
error::Error::BadRequest(format!("{MAX_TOKEN_EXPIRATION_DAYS_SETTING}: {e}"))
|
||
})?;
|
||
}
|
||
INSTANCE_BANNER_SETTING => {
|
||
match value {
|
||
// Clearing (delete row) is handled by the caller; allow it through.
|
||
serde_json::Value::Null => {}
|
||
serde_json::Value::String(s) if s.trim().is_empty() => {}
|
||
v => {
|
||
windmill_common::global_settings::validate_instance_banner(v).map_err(|e| {
|
||
error::Error::BadRequest(format!("{INSTANCE_BANNER_SETTING}: {e}"))
|
||
})?;
|
||
}
|
||
}
|
||
}
|
||
ACCENT_COLOR_SETTING => match value {
|
||
serde_json::Value::Null => {}
|
||
serde_json::Value::String(s) if s.trim().is_empty() => {}
|
||
v => windmill_common::global_settings::validate_accent_color(v)
|
||
.map_err(|e| error::Error::BadRequest(format!("{ACCENT_COLOR_SETTING}: {e}")))?,
|
||
},
|
||
_ => {}
|
||
}
|
||
Ok(())
|
||
}
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// Bulk instance config endpoints
|
||
// ---------------------------------------------------------------------------
|
||
|
||
async fn get_instance_config(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> JsonResult<InstanceConfig> {
|
||
require_super_admin(&db, &authed).await?;
|
||
let config = InstanceConfig::from_db(&db)
|
||
.await
|
||
.map_err(|e| error::Error::internal_err(e.to_string()))?;
|
||
Ok(Json(config))
|
||
}
|
||
|
||
async fn get_instance_config_yaml(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::Result<Response> {
|
||
require_super_admin(&db, &authed).await?;
|
||
let config = InstanceConfig::from_db(&db)
|
||
.await
|
||
.map_err(|e| error::Error::internal_err(e.to_string()))?;
|
||
let yaml = config
|
||
.to_sorted_yaml()
|
||
.map_err(|e| error::Error::internal_err(e))?;
|
||
Response::builder()
|
||
.header("content-type", "application/yaml")
|
||
.body(Body::from(yaml))
|
||
.map_err(|e| error::Error::internal_err(e.to_string()))
|
||
}
|
||
|
||
async fn set_instance_config(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(desired): Json<InstanceConfig>,
|
||
) -> error::Result<()> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
let current = InstanceConfig::from_db(&db)
|
||
.await
|
||
.map_err(|e| error::Error::internal_err(e.to_string()))?;
|
||
|
||
let desired_map = desired.global_settings.to_settings_map();
|
||
if !desired_map.is_empty() {
|
||
let current_map = current.global_settings.to_settings_map();
|
||
let mut settings_diff =
|
||
instance_config::diff_global_settings(¤t_map, &desired_map, ApplyMode::Merge);
|
||
let ai_config_changed = settings_diff
|
||
.upserts
|
||
.iter()
|
||
.any(|(key, _)| key == AI_CONFIG_SETTING);
|
||
|
||
// Mirror the per-key EE gate in `set_global_setting_internal`. Without
|
||
// this, the bulk endpoint would let a non-EE superadmin persist
|
||
// `workspace_fairness_*` rows even though the per-key API rejects them.
|
||
// Only block *non-disabling* upserts; deletes are allowed everywhere
|
||
// (already filtered into `settings_diff.removals`) and a
|
||
// `workspace_fairness_enabled=false` upsert is treated as a disable,
|
||
// so a downgraded instance can always turn the cap off via the bulk
|
||
// YAML endpoint too.
|
||
let upserts_touch_fairness_non_disable = settings_diff.upserts.iter().any(|(k, v)| {
|
||
if !is_workspace_fairness_setting(k) {
|
||
return false;
|
||
}
|
||
!(k == WORKSPACE_FAIRNESS_ENABLED_SETTING
|
||
&& matches!(v, serde_json::Value::Bool(false)))
|
||
});
|
||
if upserts_touch_fairness_non_disable && !workspace_fairness_settings_allowed().await {
|
||
return Err(error::Error::BadRequest(
|
||
"Workspace fairness settings require an Enterprise license".to_string(),
|
||
));
|
||
}
|
||
|
||
for (key, value) in &settings_diff.upserts {
|
||
if key != EXTERNAL_INSTANCE_PG_SETTING {
|
||
run_setting_pre_write_hook(&db, key, value).await?;
|
||
}
|
||
}
|
||
windmill_common::external_instance_pg::write_external_instance_pg_from_diff(
|
||
&db,
|
||
&mut settings_diff,
|
||
)
|
||
.await?;
|
||
|
||
instance_config::apply_settings_diff(&db, &settings_diff)
|
||
.await
|
||
.map_err(|e| error::Error::internal_err(e.to_string()))?;
|
||
|
||
if ai_config_changed {
|
||
bump_instance_ai_config_revision();
|
||
}
|
||
}
|
||
|
||
if !desired.worker_configs.is_empty() {
|
||
let current_wc: std::collections::BTreeMap<String, serde_json::Value> = current
|
||
.worker_configs
|
||
.iter()
|
||
.map(|(k, v)| {
|
||
(
|
||
k.clone(),
|
||
serde_json::to_value(v).expect("WorkerGroupConfig serialization cannot fail"),
|
||
)
|
||
})
|
||
.collect();
|
||
let desired_wc: std::collections::BTreeMap<String, serde_json::Value> = desired
|
||
.worker_configs
|
||
.iter()
|
||
.map(|(k, v)| {
|
||
(
|
||
k.clone(),
|
||
serde_json::to_value(v).expect("WorkerGroupConfig serialization cannot fail"),
|
||
)
|
||
})
|
||
.collect();
|
||
let configs_diff =
|
||
instance_config::diff_worker_configs(¤t_wc, &desired_wc, ApplyMode::Merge);
|
||
instance_config::apply_configs_diff(&db, &configs_diff)
|
||
.await
|
||
.map_err(|e| error::Error::internal_err(e.to_string()))?;
|
||
}
|
||
|
||
Ok(())
|
||
}
|
||
|
||
pub async fn get_global_setting(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Path(key): Path<String>,
|
||
) -> JsonResult<serde_json::Value> {
|
||
if !key.starts_with("default_error_handler_")
|
||
&& !key.starts_with("default_recovery_handler_")
|
||
&& !key.starts_with("default_success_handler_")
|
||
&& key != AUTOMATE_USERNAME_CREATION_SETTING
|
||
&& key != DEFAULT_TAGS_WORKSPACES_SETTING
|
||
&& key != HUB_BASE_URL_SETTING
|
||
// `wmill hub pull` reads it from a job, and no job token clears the gate. It binds an
|
||
// offline license only together with `license_key`, which stays gated.
|
||
&& key != UNIQUE_ID_SETTING
|
||
&& key != HUB_ACCESSIBLE_URL_SETTING
|
||
&& key != DISABLE_HUB_SETTING
|
||
&& key != EMAIL_DOMAIN_SETTING
|
||
&& key != APP_WORKSPACED_ROUTE_SETTING
|
||
&& key != HTTP_ROUTE_WORKSPACED_ROUTE_SETTING
|
||
// The route editor shows the inherited default to whoever is editing a
|
||
// trigger, who is usually not a superadmin. Not a secret either: any
|
||
// browser discovers the list by reading Access-Control-Allow-Origin off
|
||
// a response.
|
||
&& key != HTTP_ROUTE_DEFAULT_ALLOWED_ORIGINS_SETTING
|
||
&& key != WS_BASE_URL_SETTING
|
||
&& key != INSTANCE_BANNER_SETTING
|
||
// The token form reads it to stop offering expirations the server would shorten.
|
||
&& key != MAX_TOKEN_EXPIRATION_DAYS_SETTING
|
||
// Whoever is wiring up an MCP client reads it to know whether a URL-borne token
|
||
// would be refused, and they are usually not a superadmin. Not a secret: pointing
|
||
// any MCP client at the instance discovers the same answer.
|
||
&& key != MCP_DISABLE_TOKEN_QUERY_PARAM_SETTING
|
||
{
|
||
require_super_admin(&db, &authed).await?;
|
||
}
|
||
let value = sqlx::query!("SELECT value FROM global_settings WHERE name = $1", key)
|
||
.fetch_optional(&db)
|
||
.await?
|
||
.map(|x| x.value);
|
||
|
||
Ok(Json(value.unwrap_or_else(|| serde_json::Value::Null)))
|
||
}
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
#[derive(Deserialize, serde::Serialize)]
|
||
struct GlobalSetting {
|
||
name: String,
|
||
value: serde_json::Value,
|
||
}
|
||
|
||
/// Repositories whose registered webhook still points at a receiver the instance no
|
||
/// longer uses — what an admin has to re-save after changing the webhook base URL.
|
||
/// Read-only; changing the setting never moves a live hook on its own.
|
||
async fn github_app_stale_webhooks(
|
||
Extension(_db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> JsonResult<serde_json::Value> {
|
||
require_super_admin(&_db, &authed).await?;
|
||
#[cfg(all(feature = "enterprise", feature = "private"))]
|
||
{
|
||
let stale = windmill_common::git_sync_ee::stale_webhook_repos(&_db).await?;
|
||
return Ok(Json(serde_json::to_value(stale).map_err(|e| {
|
||
error::Error::internal_err(format!("Failed to serialize stale webhooks: {e}"))
|
||
})?));
|
||
}
|
||
#[cfg(not(all(feature = "enterprise", feature = "private")))]
|
||
Ok(Json(serde_json::json!([])))
|
||
}
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
async fn list_global_settings(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> JsonResult<Vec<GlobalSetting>> {
|
||
require_super_admin(&db, &authed).await?;
|
||
let settings = sqlx::query_as!(GlobalSetting, "SELECT name, value FROM global_settings")
|
||
.fetch_all(&db)
|
||
.await?;
|
||
|
||
Ok(Json(settings))
|
||
}
|
||
|
||
#[cfg(not(feature = "enterprise"))]
|
||
async fn list_global_settings() -> JsonResult<String> {
|
||
return Err(error::Error::BadRequest(
|
||
"Listing global settings not available on community edition".to_string(),
|
||
));
|
||
}
|
||
|
||
pub async fn send_stats(Extension(db): Extension<DB>, authed: ApiAuthed) -> Result<String> {
|
||
require_super_admin(&db, &authed).await?;
|
||
windmill_common::stats_oss::send_stats(
|
||
&HTTP_CLIENT,
|
||
&db,
|
||
windmill_common::stats_oss::SendStatsReason::Manual,
|
||
false,
|
||
)
|
||
.await?;
|
||
|
||
Ok("Sent stats".to_string())
|
||
}
|
||
|
||
async fn restart_worker_group(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Path(worker_group): Path<String>,
|
||
) -> error::Result<String> {
|
||
require_devops_role(&db, &authed).await?;
|
||
|
||
sqlx::query!(
|
||
"INSERT INTO notify_event (channel, payload) VALUES ('restart_worker_group', $1)",
|
||
worker_group
|
||
)
|
||
.execute(&db)
|
||
.await?;
|
||
|
||
Ok(format!(
|
||
"Restart signal sent to worker group '{worker_group}'"
|
||
))
|
||
}
|
||
|
||
#[derive(serde::Serialize)]
|
||
pub struct StatsDownload {
|
||
pub signature: String,
|
||
pub data: String,
|
||
}
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
pub async fn get_stats(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::JsonResult<StatsDownload> {
|
||
require_super_admin(&db, &authed).await?;
|
||
let stats = windmill_common::stats_oss::get_stats_payload(
|
||
&db,
|
||
&windmill_common::stats_oss::SendStatsReason::Manual,
|
||
false,
|
||
)
|
||
.await?;
|
||
let json =
|
||
serde_json::to_string(&stats).map_err(|e| error::Error::InternalErr(e.to_string()))?;
|
||
let signature = windmill_common::stats_oss::sign_stats(&json);
|
||
Ok(axum::Json(StatsDownload { signature, data: json }))
|
||
}
|
||
|
||
#[cfg(not(feature = "enterprise"))]
|
||
pub async fn get_stats() -> error::JsonResult<StatsDownload> {
|
||
Err(error::Error::BadRequest(
|
||
"Downloading telemetry is only available on enterprise edition".to_string(),
|
||
))
|
||
}
|
||
|
||
#[derive(serde::Serialize)]
|
||
pub struct KeyRenewalAttempt {
|
||
result: String,
|
||
attempted_at: chrono::DateTime<chrono::Utc>,
|
||
}
|
||
|
||
pub async fn get_latest_key_renewal_attempt(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> JsonResult<Option<KeyRenewalAttempt>> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
let last_attempt = sqlx::query!(
|
||
"SELECT value, created_at FROM metrics WHERE id = $1 ORDER BY created_at DESC LIMIT 1",
|
||
"license_key_renewal"
|
||
)
|
||
.fetch_optional(&db)
|
||
.await?;
|
||
|
||
match last_attempt {
|
||
Some(last_attempt) => {
|
||
let last_attempt_result = serde_json::from_value::<String>(last_attempt.value)
|
||
.map_err(|e| {
|
||
error::Error::internal_err(format!("Failed to parse last attempt: {}", e))
|
||
})?;
|
||
Ok(Json(Some(KeyRenewalAttempt {
|
||
result: last_attempt_result,
|
||
attempted_at: last_attempt.created_at,
|
||
})))
|
||
}
|
||
None => Ok(Json(None)),
|
||
}
|
||
}
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
#[derive(Deserialize)]
|
||
pub struct LicenseQuery {
|
||
license_key: Option<String>,
|
||
}
|
||
|
||
#[cfg(not(feature = "enterprise"))]
|
||
pub async fn renew_license_key() -> Result<String> {
|
||
return Err(error::Error::BadRequest(
|
||
"License key renewal not available on community edition".to_string(),
|
||
));
|
||
}
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
pub async fn renew_license_key(
|
||
Extension(db): Extension<DB>,
|
||
Query(LicenseQuery { license_key }): Query<LicenseQuery>,
|
||
authed: ApiAuthed,
|
||
) -> Result<String> {
|
||
require_super_admin(&db, &authed).await?;
|
||
let result = windmill_common::ee_oss::renew_license_key(
|
||
&HTTP_CLIENT,
|
||
&db,
|
||
license_key,
|
||
windmill_common::ee_oss::RenewReason::Manual,
|
||
)
|
||
.await;
|
||
|
||
if result != "success" {
|
||
return Err(error::Error::BadRequest(format!(
|
||
"Failed to renew license key: {}",
|
||
if result == "Unauthorized" {
|
||
"Invalid key".to_string()
|
||
} else {
|
||
result
|
||
}
|
||
)));
|
||
} else {
|
||
return Ok("Renewed license key".to_string());
|
||
}
|
||
}
|
||
|
||
#[cfg(not(feature = "enterprise"))]
|
||
pub async fn create_customer_portal_session() -> Result<String> {
|
||
return Err(error::Error::BadRequest(
|
||
"Customer portal is not available on community edition".to_string(),
|
||
));
|
||
}
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
pub async fn create_customer_portal_session(
|
||
Query(LicenseQuery { license_key }): Query<LicenseQuery>,
|
||
) -> Result<String> {
|
||
let url =
|
||
windmill_common::ee_oss::create_customer_portal_session(&HTTP_CLIENT, license_key).await?;
|
||
|
||
return Ok(url);
|
||
}
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
pub async fn test_critical_channels(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(test_critical_channels): Json<Vec<CriticalErrorChannel>>,
|
||
) -> Result<String> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
send_critical_alert(
|
||
"Test critical error".to_string(),
|
||
&db,
|
||
CriticalAlertKind::CriticalError,
|
||
Some(test_critical_channels),
|
||
)
|
||
.await;
|
||
Ok("Sent test critical error".to_string())
|
||
}
|
||
|
||
#[cfg(not(feature = "enterprise"))]
|
||
pub async fn test_critical_channels() -> Result<String> {
|
||
Ok("Critical channels require EE".to_string())
|
||
}
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
pub async fn get_critical_alerts(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Query(params): Query<windmill_alerting::AlertQueryParams>,
|
||
) -> JsonResult<serde_json::Value> {
|
||
require_devops_role(&db, &authed).await?;
|
||
|
||
windmill_alerting::get_critical_alerts(db, params, None).await
|
||
}
|
||
|
||
#[cfg(not(feature = "enterprise"))]
|
||
pub async fn get_critical_alerts() -> error::Error {
|
||
error::Error::NotFound("Critical Alerts require EE".to_string())
|
||
}
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
pub async fn acknowledge_critical_alert(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Path(id): Path<i32>,
|
||
) -> error::Result<String> {
|
||
require_devops_role(&db, &authed).await?;
|
||
windmill_alerting::acknowledge_critical_alert(db, None, id).await
|
||
}
|
||
|
||
#[cfg(not(feature = "enterprise"))]
|
||
pub async fn acknowledge_critical_alert() -> error::Error {
|
||
error::Error::NotFound("Critical Alerts require EE".to_string())
|
||
}
|
||
|
||
#[cfg(feature = "enterprise")]
|
||
pub async fn acknowledge_all_critical_alerts(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
) -> error::Result<String> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
windmill_alerting::acknowledge_all_critical_alerts(db, None).await
|
||
}
|
||
|
||
#[cfg(not(feature = "enterprise"))]
|
||
pub async fn acknowledge_all_critical_alerts() -> error::Error {
|
||
error::Error::NotFound("Critical Alerts require EE".to_string())
|
||
}
|
||
|
||
#[derive(Deserialize, Debug, Serialize)]
|
||
struct CustomInstanceDb {
|
||
logs: CustomInstanceDbLogs, // (Step, Message)[]
|
||
success: bool,
|
||
error: Option<String>,
|
||
tag: Option<String>,
|
||
#[serde(default, skip_serializing_if = "Vec::is_empty")]
|
||
used_by_workspaces: Vec<String>,
|
||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||
workspace_id: Option<String>,
|
||
}
|
||
|
||
#[derive(Deserialize, Debug, Serialize, Default)]
|
||
#[serde(default)]
|
||
struct CustomInstanceDbLogs {
|
||
super_admin: String,
|
||
#[serde(skip_serializing_if = "String::is_empty")]
|
||
database_credentials: String,
|
||
#[serde(skip_serializing_if = "String::is_empty")]
|
||
valid_dbname: String,
|
||
#[serde(skip_serializing_if = "String::is_empty")]
|
||
created_database: String,
|
||
#[serde(skip_serializing_if = "String::is_empty")]
|
||
db_connect: String,
|
||
#[serde(skip_serializing_if = "String::is_empty")]
|
||
grant_permissions: String,
|
||
#[serde(skip_serializing_if = "String::is_empty")]
|
||
replication_user: String,
|
||
#[serde(skip_serializing_if = "Option::is_none")]
|
||
replication_user_error: Option<String>,
|
||
#[serde(skip_serializing_if = "String::is_empty")]
|
||
user_connect: String,
|
||
}
|
||
|
||
async fn list_custom_instance_pg_databases(
|
||
authed: ApiAuthed,
|
||
Extension(db): Extension<DB>,
|
||
) -> JsonResult<HashMap<String, CustomInstanceDb>> {
|
||
let result = sqlx::query_scalar!(
|
||
r#"SELECT value->'databases' FROM global_settings WHERE name = 'custom_instance_pg_databases'"#,
|
||
)
|
||
.fetch_one(&db)
|
||
.await?
|
||
.ok_or_else(|| error::Error::ExecutionErr("Couldn't find custom_instance_pg_databases".to_string()))?;
|
||
let mut result: HashMap<String, CustomInstanceDb> =
|
||
serde_json::from_value(result).map_err(|e| {
|
||
error::Error::ExecutionErr(format!(
|
||
"couldn't parse custom_instance_pg_databases.databases : {}",
|
||
e.to_string()
|
||
))
|
||
})?;
|
||
|
||
if !windmill_api_auth::is_super_admin_authed(&db, &authed).await? {
|
||
// A fork copy's name gives away the workspace it was reserved for, so every pending fork on
|
||
// the instance would be listed. Kept for members of that workspace, and wherever the
|
||
// caller's workspaces use it, e.g. the fork it was finalized into.
|
||
let reserved_visible: BTreeSet<String> = sqlx::query_scalar(
|
||
r#"SELECT e.k FROM global_settings gs
|
||
CROSS JOIN LATERAL jsonb_each(gs.value->'databases') AS e(k, v)
|
||
WHERE gs.name = 'custom_instance_pg_databases' AND e.v->>'workspace_id' IS NOT NULL
|
||
AND (EXISTS (SELECT 1 FROM usr WHERE usr.email = $1
|
||
AND usr.workspace_id = e.v->>'workspace_id')
|
||
OR EXISTS (SELECT 1 FROM usr JOIN workspace_settings ws
|
||
ON ws.workspace_id = usr.workspace_id
|
||
CROSS JOIN LATERAL jsonb_each(
|
||
CASE WHEN jsonb_typeof(ws.datatable->'datatables') = 'object'
|
||
THEN ws.datatable->'datatables' ELSE '{}'::jsonb END) dt
|
||
WHERE usr.email = $1
|
||
AND dt.value->'database'->>'resource_type' = 'instance'
|
||
AND dt.value->'database'->>'resource_path' = e.k))"#,
|
||
)
|
||
.bind(&authed.email)
|
||
.fetch_all(&db)
|
||
.await?
|
||
.into_iter()
|
||
.collect();
|
||
result.retain(|dbname, entry| {
|
||
entry.workspace_id.is_none() || reserved_visible.contains(dbname)
|
||
});
|
||
// Which workspace reserved a copy is still only for superadmins.
|
||
for entry in result.values_mut() {
|
||
entry.workspace_id = None;
|
||
}
|
||
return Ok(Json(result));
|
||
}
|
||
{
|
||
// Enrich each database with the list of workspaces referencing it through
|
||
// either a ducklake catalog or a datatable database whose resource_type is
|
||
// 'instance'. Not stored in DB to avoid drift.
|
||
let usages = sqlx::query!(
|
||
r#"
|
||
SELECT ws.workspace_id AS "workspace_id!", entry->'catalog'->>'resource_path' AS dbname
|
||
FROM workspace_settings ws
|
||
CROSS JOIN LATERAL jsonb_each(
|
||
CASE WHEN jsonb_typeof(ws.ducklake->'ducklakes') = 'object'
|
||
THEN ws.ducklake->'ducklakes'
|
||
ELSE '{}'::jsonb END
|
||
) AS dl(k, entry)
|
||
WHERE entry->'catalog'->>'resource_type' = 'instance'
|
||
AND entry->'catalog'->>'resource_path' IS NOT NULL
|
||
UNION ALL
|
||
SELECT ws.workspace_id AS "workspace_id!", entry->'database'->>'resource_path' AS dbname
|
||
FROM workspace_settings ws
|
||
CROSS JOIN LATERAL jsonb_each(
|
||
CASE WHEN jsonb_typeof(ws.datatable->'datatables') = 'object'
|
||
THEN ws.datatable->'datatables'
|
||
ELSE '{}'::jsonb END
|
||
) AS dt(k, entry)
|
||
WHERE entry->'database'->>'resource_type' = 'instance'
|
||
AND entry->'database'->>'resource_path' IS NOT NULL
|
||
"#,
|
||
)
|
||
.fetch_all(&db)
|
||
.await?;
|
||
|
||
let mut by_db: HashMap<String, BTreeSet<String>> = HashMap::new();
|
||
for row in usages {
|
||
if let Some(dbname) = row.dbname {
|
||
by_db.entry(dbname).or_default().insert(row.workspace_id);
|
||
}
|
||
}
|
||
for (dbname, entry) in result.iter_mut() {
|
||
if let Some(workspaces) = by_db.remove(dbname) {
|
||
entry.used_by_workspaces = workspaces.into_iter().collect();
|
||
}
|
||
}
|
||
}
|
||
|
||
return Ok(Json(result));
|
||
}
|
||
|
||
async fn refresh_custom_instance_user_pwd(
|
||
authed: ApiAuthed,
|
||
Extension(db): Extension<DB>,
|
||
) -> JsonResult<()> {
|
||
require_super_admin(&db, &authed).await?;
|
||
windmill_common::utils::refresh_custom_instance_user_pwd(&db).await?;
|
||
windmill_common::utils::refresh_custom_instance_replication_user_pwd(&db).await?;
|
||
Ok(Json(()))
|
||
}
|
||
|
||
async fn get_external_instance_pg_status(
|
||
authed: ApiAuthed,
|
||
Extension(db): Extension<DB>,
|
||
) -> JsonResult<windmill_common::external_instance_pg::ExternalInstancePgStatus> {
|
||
require_super_admin(&db, &authed).await?;
|
||
Ok(Json(
|
||
windmill_common::external_instance_pg::external_instance_pg_status(&db).await?,
|
||
))
|
||
}
|
||
|
||
#[derive(Deserialize)]
|
||
struct SetupExternalInstancePgBody {
|
||
#[serde(default)]
|
||
rotate_passwords: bool,
|
||
}
|
||
|
||
async fn setup_external_instance_pg(
|
||
authed: ApiAuthed,
|
||
Extension(db): Extension<DB>,
|
||
Json(body): Json<SetupExternalInstancePgBody>,
|
||
) -> JsonResult<windmill_common::external_instance_pg::ExternalInstancePgSetupReport> {
|
||
require_super_admin(&db, &authed).await?;
|
||
let report = windmill_common::external_instance_pg::setup_external_instance_pg_unchecked(
|
||
&db,
|
||
body.rotate_passwords,
|
||
)
|
||
.await?;
|
||
let rotated = body.rotate_passwords.to_string();
|
||
let success = report.success.to_string();
|
||
windmill_audit::audit_oss::audit_log(
|
||
&db,
|
||
&authed,
|
||
"settings.setup_external_instance_pg",
|
||
windmill_audit::ActionKind::Update,
|
||
"global",
|
||
Some(&authed.email),
|
||
Some(
|
||
[
|
||
("rotate_passwords", rotated.as_str()),
|
||
("success", success.as_str()),
|
||
]
|
||
.into(),
|
||
),
|
||
)
|
||
.await?;
|
||
Ok(Json(report))
|
||
}
|
||
|
||
#[derive(Serialize)]
|
||
struct ExternalInstancePgDatabase {
|
||
#[serde(flatten)]
|
||
status: windmill_common::instance_config::CustomInstanceDb,
|
||
used_by_workspaces: Vec<String>,
|
||
}
|
||
|
||
async fn list_external_instance_pg_databases(
|
||
authed: ApiAuthed,
|
||
Extension(db): Extension<DB>,
|
||
) -> JsonResult<std::collections::BTreeMap<String, ExternalInstancePgDatabase>> {
|
||
require_super_admin(&db, &authed).await?;
|
||
let databases = windmill_common::external_instance_pg::external_instance_databases(&db).await?;
|
||
let mut usages =
|
||
windmill_common::external_instance_pg::external_instance_database_usages(&db).await?;
|
||
Ok(Json(
|
||
databases
|
||
.into_iter()
|
||
.map(|(name, status)| {
|
||
let used_by_workspaces = usages.remove(&name).unwrap_or_default();
|
||
(
|
||
name,
|
||
ExternalInstancePgDatabase {
|
||
status,
|
||
used_by_workspaces: used_by_workspaces.into_iter().collect(),
|
||
},
|
||
)
|
||
})
|
||
.collect(),
|
||
))
|
||
}
|
||
|
||
async fn create_external_instance_pg_database(
|
||
authed: ApiAuthed,
|
||
Extension(db): Extension<DB>,
|
||
Path(dbname): Path<String>,
|
||
Json(body): Json<SetupCustomInstanceDbBody>,
|
||
) -> JsonResult<()> {
|
||
require_super_admin(&db, &authed).await?;
|
||
let tag = body.tag.as_deref().unwrap_or("datatable");
|
||
let mut tx = db.begin().await?;
|
||
windmill_common::external_instance_pg::create_external_instance_database_unchecked(
|
||
&db, &mut tx, &dbname, tag, None,
|
||
)
|
||
.await?;
|
||
tx.commit().await?;
|
||
windmill_audit::audit_oss::audit_log(
|
||
&db,
|
||
&authed,
|
||
"settings.create_external_instance_pg_database",
|
||
windmill_audit::ActionKind::Create,
|
||
"global",
|
||
Some(&authed.email),
|
||
Some([("dbname", dbname.as_str()), ("tag", tag)].into()),
|
||
)
|
||
.await?;
|
||
Ok(Json(()))
|
||
}
|
||
|
||
async fn drop_external_instance_pg_database(
|
||
authed: ApiAuthed,
|
||
Extension(db): Extension<DB>,
|
||
Path(dbname): Path<String>,
|
||
) -> JsonResult<()> {
|
||
require_super_admin(&db, &authed).await?;
|
||
// A data table naming a dropped database fails on every job, far from the drop that caused it.
|
||
let mut tx = db.begin().await?;
|
||
windmill_common::external_instance_pg::drop_external_instance_database_unchecked(
|
||
&mut tx, &dbname, None,
|
||
)
|
||
.await?;
|
||
tx.commit().await?;
|
||
windmill_audit::audit_oss::audit_log(
|
||
&db,
|
||
&authed,
|
||
"settings.drop_external_instance_pg_database",
|
||
windmill_audit::ActionKind::Delete,
|
||
"global",
|
||
Some(&authed.email),
|
||
Some([("dbname", dbname.as_str())].into()),
|
||
)
|
||
.await?;
|
||
Ok(Json(()))
|
||
}
|
||
|
||
#[derive(Deserialize)]
|
||
struct SetupCustomInstanceDbBody {
|
||
tag: Option<String>,
|
||
}
|
||
|
||
async fn setup_custom_instance_pg_database(
|
||
authed: ApiAuthed,
|
||
Extension(db): Extension<DB>,
|
||
Path(dbname): Path<String>,
|
||
Json(body): Json<SetupCustomInstanceDbBody>,
|
||
) -> JsonResult<CustomInstanceDb> {
|
||
// Before anything is recorded: the status written below replaces the registry entry, and with it
|
||
// the workspace a fork copy is reserved for.
|
||
require_super_admin(&db, &authed).await?;
|
||
windmill_common::workspaces::ensure_instance_pg_available(&db).await?;
|
||
// Fork cleanup checks and drops the database and its entry under this lock. Held from before
|
||
// the setup creates the database to after its entry is written, neither lands on the other's
|
||
// half-done state: a dropped database with its entry written back, or the reverse.
|
||
let mut tx = db.begin().await?;
|
||
windmill_common::datatable_roles::lock_instance_databases_governance(&mut tx, [dbname.trim()])
|
||
.await?;
|
||
let mut logs = CustomInstanceDbLogs::default();
|
||
let result = setup_custom_instance_pg_database_inner(authed, &db, &dbname, &mut logs).await;
|
||
let success = result.is_ok();
|
||
let error = result.err().map(|e| e.to_string());
|
||
let status = CustomInstanceDb {
|
||
logs,
|
||
success,
|
||
error,
|
||
tag: body.tag,
|
||
used_by_workspaces: vec![],
|
||
workspace_id: None,
|
||
};
|
||
let status_json = serde_json::to_value(&status).map_err(to_anyhow)?;
|
||
// The fork reservation is carried over inside the write, from whatever the row holds then: a
|
||
// rename migrating it while the setup above ran would otherwise be overwritten with the value
|
||
// this request started from, stranding the copy under the archived workspace.
|
||
let saved = sqlx::query_scalar::<_, serde_json::Value>(
|
||
r#"UPDATE global_settings SET value = jsonb_set(value, '{databases}',
|
||
COALESCE(value->'databases', '{}'::jsonb)
|
||
|| jsonb_build_object($1::text, $2::jsonb || jsonb_build_object(
|
||
'workspace_id', value->'databases'->$1::text->'workspace_id')))
|
||
WHERE name = 'custom_instance_pg_databases'
|
||
RETURNING value->'databases'->$1::text"#,
|
||
)
|
||
.bind(&dbname)
|
||
.bind(&status_json)
|
||
.fetch_one(&mut *tx)
|
||
.await?;
|
||
tx.commit().await?;
|
||
let status: CustomInstanceDb = serde_json::from_value(saved).map_err(to_anyhow)?;
|
||
|
||
Ok(Json(status))
|
||
}
|
||
|
||
async fn setup_custom_instance_pg_database_inner(
|
||
authed: ApiAuthed,
|
||
db: &DB,
|
||
dbname: &str,
|
||
logs: &mut CustomInstanceDbLogs,
|
||
) -> Result<()> {
|
||
require_super_admin(db, &authed).await?;
|
||
logs.super_admin = "OK".to_string();
|
||
let wmill_pg_creds = PgDatabase::parse_uri(&get_database_url().await?.as_str().await)?;
|
||
logs.database_credentials = "OK".to_string();
|
||
|
||
// Validate name to ensure it only contains alphanumeric characters
|
||
// Prevents SQL injection on the instance database
|
||
lazy_static::lazy_static! {
|
||
// Must start with a letter, then alphanumeric/underscore/hyphen
|
||
static ref VALID_NAME: regex::Regex = regex::Regex::new(r"^[a-zA-Z][a-zA-Z0-9_-]*$").unwrap();
|
||
}
|
||
let dbname = dbname.trim();
|
||
if dbname.is_empty() {
|
||
return Err(error::Error::BadRequest(
|
||
"Database name cannot be empty".to_string(),
|
||
));
|
||
}
|
||
// PostgreSQL identifier limit is 63 bytes
|
||
if dbname.len() > 63 {
|
||
return Err(error::Error::BadRequest(
|
||
"Database name cannot exceed 63 characters".to_string(),
|
||
));
|
||
}
|
||
if !VALID_NAME.is_match(dbname) {
|
||
return Err(error::Error::BadRequest(
|
||
"Database name must start with a letter and contain only alphanumeric characters, underscores, or hyphens".to_string(),
|
||
));
|
||
}
|
||
// Additional check: block PostgreSQL reserved/special names
|
||
let lower = dbname.to_lowercase();
|
||
if lower == "template0" || lower == "template1" || lower == "postgres" {
|
||
return Err(error::Error::BadRequest(
|
||
"Cannot use reserved PostgreSQL database names".to_string(),
|
||
));
|
||
}
|
||
if wmill_pg_creds
|
||
.dbname
|
||
.trim()
|
||
.eq_ignore_ascii_case(dbname.trim())
|
||
{
|
||
return Err(error::Error::BadRequest(
|
||
"Database name cannot be the same as the main database".to_string(),
|
||
));
|
||
}
|
||
logs.valid_dbname = "OK".to_string();
|
||
|
||
let db_exists = sqlx::query_scalar!(
|
||
"SELECT EXISTS (SELECT 1 FROM pg_catalog.pg_database WHERE datname = $1)",
|
||
dbname
|
||
)
|
||
.fetch_one(db)
|
||
.await?
|
||
.unwrap_or(false);
|
||
|
||
let pg_creds = PgDatabase { dbname: dbname.to_string(), ..wmill_pg_creds };
|
||
|
||
logs.created_database = "SKIP".to_string();
|
||
if !db_exists {
|
||
// SAFETY: `dbname` has been validated by the VALID_NAME regex and length checks above (lines 1088–1120).
|
||
sqlx::query(&format!("CREATE DATABASE \"{dbname}\""))
|
||
.execute(db)
|
||
.await?;
|
||
logs.created_database = "OK".to_string();
|
||
}
|
||
|
||
// We have to connect to the newly created database as admin to grant permissions
|
||
let (client, connection) = pg_creds.connect(Some(db)).await?;
|
||
let join_handle = tokio::spawn(async move { connection.await });
|
||
|
||
logs.db_connect = "OK".to_string();
|
||
|
||
// SAFETY: `dbname` has been validated by the VALID_NAME regex and length checks above.
|
||
client
|
||
.batch_execute(&format!(
|
||
"GRANT CONNECT ON DATABASE \"{dbname}\" TO custom_instance_user;
|
||
GRANT USAGE ON SCHEMA public TO custom_instance_user;
|
||
GRANT CREATE ON SCHEMA public TO custom_instance_user;
|
||
GRANT CREATE ON DATABASE \"{dbname}\" TO custom_instance_user;
|
||
ALTER DEFAULT PRIVILEGES IN SCHEMA public
|
||
GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO custom_instance_user;
|
||
ALTER ROLE custom_instance_user CREATEROLE;"
|
||
))
|
||
.await
|
||
.map_err(|e| {
|
||
error::Error::ExecutionErr(format!(
|
||
"Failed to grant permissions to custom_instance_user: {}",
|
||
pg_error_message(&e),
|
||
))
|
||
})?;
|
||
|
||
logs.grant_permissions = "OK".to_string();
|
||
|
||
drop(client); // /!\ Drop before joining to avoid deadlock
|
||
windmill_common::shutdown_pg_connection(join_handle).await?;
|
||
|
||
// Roles are cluster-wide, so the dedicated role used by postgres trigger connections is
|
||
// provisioned on the main pool rather than on the new database. Reported as its own step
|
||
// rather than failing the setup: without the role the database still serves datatables, only
|
||
// postgres triggers on them break.
|
||
match windmill_common::utils::ensure_custom_instance_replication_user(db).await {
|
||
Ok(()) => logs.replication_user = "OK".to_string(),
|
||
Err(e) => {
|
||
tracing::error!("Failed to provision custom_instance_replication_user: {e:#}");
|
||
logs.replication_user = "FAIL".to_string();
|
||
logs.replication_user_error = Some(e.to_string());
|
||
}
|
||
}
|
||
|
||
// Everything above logged in as the DATABASE_URL user, but whatever uses the database logs
|
||
// in as custom_instance_user. A proxy that routes on the login name can accept one and
|
||
// refuse the other, which would otherwise surface only once a data table first connects.
|
||
let user_creds = PgDatabase {
|
||
user: Some(windmill_common::datatable_roles::CUSTOM_INSTANCE_USER.to_string()),
|
||
password: Some(windmill_common::utils::get_custom_pg_instance_password(db).await?),
|
||
..pg_creds
|
||
};
|
||
let (client, connection) = user_creds
|
||
.connect(Some(db))
|
||
.await
|
||
.map_err(|e| custom_instance_user_connect_error(&user_creds.host, dbname, e))?;
|
||
let join_handle = tokio::spawn(async move { connection.await });
|
||
logs.user_connect = "OK".to_string();
|
||
drop(client); // /!\ Drop before joining to avoid deadlock
|
||
windmill_common::shutdown_pg_connection(join_handle).await?;
|
||
|
||
Ok(())
|
||
}
|
||
|
||
fn custom_instance_user_connect_error(host: &str, dbname: &str, e: error::Error) -> error::Error {
|
||
let cause = match &e {
|
||
error::Error::Anyhow { error, .. } => format!("{error:#}"),
|
||
e => e.to_string(),
|
||
};
|
||
// Supavisor, Supabase's pooler, reads the tenant to route to from the login
|
||
// (`<user>.<project_ref>`), and only knows the logins configured for that tenant.
|
||
let lower = cause.to_lowercase();
|
||
let routing_refused = [
|
||
"enoidentifier",
|
||
"tenant identifier",
|
||
"tenant or user",
|
||
"tenant/user",
|
||
]
|
||
.iter()
|
||
.any(|signature| lower.contains(signature));
|
||
if routing_refused {
|
||
error::Error::BadConfig(format!(
|
||
"DATABASE_URL reaches Postgres through a connection pooler ({host}) that picks the \
|
||
server to route to from the login name, and it refused custom_instance_user, the \
|
||
role Windmill uses for instance databases ({cause}). Instance databases cannot be \
|
||
used through this pooler: use your own Postgres database instead, or point \
|
||
DATABASE_URL at the Postgres server directly rather than at the pooler."
|
||
))
|
||
} else {
|
||
error::Error::ExecutionErr(format!(
|
||
"Could not connect to {dbname} as custom_instance_user: {cause}"
|
||
))
|
||
}
|
||
}
|
||
|
||
async fn drop_custom_instance_pg_database(
|
||
authed: ApiAuthed,
|
||
Extension(db): Extension<DB>,
|
||
Path(dbname): Path<String>,
|
||
) -> Result<String> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
windmill_common::drop_custom_instance_database(&db, &dbname).await?;
|
||
|
||
Ok(format!("Database '{}' dropped successfully", dbname))
|
||
}
|
||
|
||
// ============================================================================
|
||
// Secret Backend Settings (HashiCorp Vault Integration) - Enterprise Edition
|
||
// ============================================================================
|
||
|
||
/// Test connection to a secret backend (HashiCorp Vault)
|
||
///
|
||
/// This endpoint validates that the Vault settings are correct and that
|
||
/// Windmill can successfully authenticate and communicate with Vault.
|
||
///
|
||
/// This is an Enterprise Edition feature.
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
pub async fn test_secret_backend(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(settings): Json<VaultSettings>,
|
||
) -> Result<String> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
windmill_common::secret_backend::test_vault_connection(&settings, Some(&db)).await?;
|
||
|
||
Ok("Successfully connected to HashiCorp Vault".to_string())
|
||
}
|
||
|
||
/// Migrate existing secrets from database to HashiCorp Vault
|
||
///
|
||
/// This endpoint reads all encrypted secrets from the database, decrypts them,
|
||
/// and stores them in HashiCorp Vault. The database values are NOT deleted
|
||
/// automatically to allow for rollback if needed.
|
||
///
|
||
/// This is an Enterprise Edition feature.
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
pub async fn migrate_secrets_to_vault(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(settings): Json<VaultSettings>,
|
||
) -> JsonResult<SecretMigrationReport> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
let report = windmill_common::secret_backend::migrate_secrets_to_vault(&db, &settings).await?;
|
||
|
||
Ok(Json(report))
|
||
}
|
||
|
||
/// Migrate secrets from HashiCorp Vault back to database
|
||
///
|
||
/// This endpoint reads all secrets from HashiCorp Vault, encrypts them using
|
||
/// the workspace encryption keys, and stores them in the database. The Vault
|
||
/// values are NOT deleted automatically to allow for rollback if needed.
|
||
///
|
||
/// This is an Enterprise Edition feature.
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
pub async fn migrate_secrets_to_database(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(settings): Json<VaultSettings>,
|
||
) -> JsonResult<SecretMigrationReport> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
let report =
|
||
windmill_common::secret_backend::migrate_secrets_to_database(&db, &settings).await?;
|
||
|
||
Ok(Json(report))
|
||
}
|
||
|
||
/// Test connection to Azure Key Vault
|
||
///
|
||
/// This is an Enterprise Edition feature.
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
pub async fn test_azure_kv_backend(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(settings): Json<AzureKeyVaultSettings>,
|
||
) -> Result<String> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
windmill_common::secret_backend::test_azure_kv_connection(&settings).await?;
|
||
|
||
Ok("Successfully connected to Azure Key Vault".to_string())
|
||
}
|
||
|
||
/// Migrate existing secrets from database to Azure Key Vault
|
||
///
|
||
/// This is an Enterprise Edition feature.
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
pub async fn migrate_secrets_to_azure_kv(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(settings): Json<AzureKeyVaultSettings>,
|
||
) -> JsonResult<SecretMigrationReport> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
let report =
|
||
windmill_common::secret_backend::migrate_secrets_to_azure_kv(&db, &settings).await?;
|
||
|
||
Ok(Json(report))
|
||
}
|
||
|
||
/// Migrate secrets from Azure Key Vault back to database
|
||
///
|
||
/// This is an Enterprise Edition feature.
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
pub async fn migrate_secrets_from_azure_kv(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(settings): Json<AzureKeyVaultSettings>,
|
||
) -> JsonResult<SecretMigrationReport> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
let report =
|
||
windmill_common::secret_backend::migrate_secrets_from_azure_kv(&db, &settings).await?;
|
||
|
||
Ok(Json(report))
|
||
}
|
||
|
||
/// Test connection to AWS Secrets Manager
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
pub async fn test_aws_sm_backend(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(settings): Json<AwsSecretsManagerSettings>,
|
||
) -> Result<String> {
|
||
require_super_admin(&db, &authed).await?;
|
||
windmill_common::secret_backend::test_aws_sm_connection(&settings).await?;
|
||
Ok("Successfully connected to AWS Secrets Manager".to_string())
|
||
}
|
||
|
||
/// Migrate existing secrets from database to AWS Secrets Manager
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
pub async fn migrate_secrets_to_aws_sm(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(settings): Json<AwsSecretsManagerSettings>,
|
||
) -> JsonResult<SecretMigrationReport> {
|
||
require_super_admin(&db, &authed).await?;
|
||
let report = windmill_common::secret_backend::migrate_secrets_to_aws_sm(&db, &settings).await?;
|
||
Ok(Json(report))
|
||
}
|
||
|
||
/// Migrate secrets from AWS Secrets Manager back to database
|
||
#[cfg(all(feature = "private", feature = "enterprise"))]
|
||
pub async fn migrate_secrets_from_aws_sm(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Json(settings): Json<AwsSecretsManagerSettings>,
|
||
) -> JsonResult<SecretMigrationReport> {
|
||
require_super_admin(&db, &authed).await?;
|
||
let report =
|
||
windmill_common::secret_backend::migrate_secrets_from_aws_sm(&db, &settings).await?;
|
||
Ok(Json(report))
|
||
}
|
||
|
||
// ============================================================================
|
||
// JWKS Endpoint for Vault JWT Authentication
|
||
// ============================================================================
|
||
|
||
/// JSON Web Key Set response structure
|
||
#[derive(Serialize)]
|
||
pub struct JwksResponse {
|
||
pub keys: Vec<serde_json::Value>,
|
||
}
|
||
|
||
/// Fallback JWKS endpoint used when OIDC support is not compiled in.
|
||
///
|
||
/// When built with `private` + `enterprise` + `openidconnect`, the route in
|
||
/// `windmill-api` dispatches to `oidc_oss::jwks` (re-exported from
|
||
/// `oidc_ee::jwks`) instead, which serves the actual public keys.
|
||
pub async fn get_jwks() -> JsonResult<JwksResponse> {
|
||
Ok(Json(JwksResponse { keys: vec![] }))
|
||
}
|
||
|
||
#[derive(serde::Deserialize, serde::Serialize)]
|
||
struct CachedResourceType {
|
||
#[allow(dead_code)]
|
||
id: i64,
|
||
name: String,
|
||
schema: Option<serde_json::Value>,
|
||
#[allow(dead_code)]
|
||
app: String,
|
||
description: Option<String>,
|
||
/// Doubly optional, and read through a wrapping deserializer: this struct also
|
||
/// decodes the on-disk cache, where an absent key means "written before the
|
||
/// column, leave the stored extension alone" and an explicit null means the hub
|
||
/// dropped it. Plain serde folds both into `None`.
|
||
#[serde(
|
||
default,
|
||
deserialize_with = "windmill_common::more_serde::double_option"
|
||
)]
|
||
format_extension: Option<Option<String>>,
|
||
/// Doubly optional like `format_extension`: no key leaves the stored name alone, an explicit
|
||
/// null (the hub naming nothing) clears it.
|
||
#[serde(
|
||
default,
|
||
deserialize_with = "windmill_common::more_serde::double_option"
|
||
)]
|
||
display_name: Option<Option<String>>,
|
||
}
|
||
|
||
#[derive(serde::Deserialize)]
|
||
struct HubResourceTypeRaw {
|
||
id: i64,
|
||
name: String,
|
||
schema: Option<String>,
|
||
app: String,
|
||
description: Option<String>,
|
||
#[serde(default)]
|
||
format_extension: Option<String>,
|
||
#[serde(
|
||
default,
|
||
deserialize_with = "windmill_common::more_serde::double_option"
|
||
)]
|
||
display_name: Option<Option<String>>,
|
||
}
|
||
|
||
async fn fetch_resource_types_from_hub() -> error::Result<Vec<CachedResourceType>> {
|
||
let response = HTTP_CLIENT
|
||
.get(format!(
|
||
"{}/resource_types/list",
|
||
windmill_common::DEFAULT_HUB_BASE_URL
|
||
))
|
||
.header("Accept", "application/json")
|
||
.send()
|
||
.await
|
||
.map_err(|e| error::Error::InternalErr(format!("Failed to fetch from hub: {}", e)))?;
|
||
|
||
if !response.status().is_success() {
|
||
return Err(error::Error::InternalErr(format!(
|
||
"Hub returned status {}",
|
||
response.status()
|
||
)));
|
||
}
|
||
|
||
let raw_types: Vec<HubResourceTypeRaw> = response
|
||
.json()
|
||
.await
|
||
.map_err(|e| error::Error::InternalErr(format!("Failed to parse hub response: {}", e)))?;
|
||
|
||
Ok(raw_types
|
||
.into_iter()
|
||
.filter_map(|rt| {
|
||
let schema = match rt.schema {
|
||
Some(s) => match serde_json::from_str(&s) {
|
||
Ok(v) => Some(v),
|
||
Err(_) => return None,
|
||
},
|
||
None => None,
|
||
};
|
||
Some(CachedResourceType {
|
||
id: rt.id,
|
||
name: rt.name,
|
||
schema,
|
||
app: rt.app,
|
||
description: rt.description,
|
||
format_extension: Some(rt.format_extension),
|
||
display_name: rt.display_name,
|
||
})
|
||
})
|
||
.collect())
|
||
}
|
||
|
||
#[derive(serde::Deserialize)]
|
||
struct SyncResourceTypesQuery {
|
||
name: Option<String>,
|
||
}
|
||
|
||
async fn sync_cached_resource_types(
|
||
Extension(db): Extension<DB>,
|
||
authed: ApiAuthed,
|
||
Query(SyncResourceTypesQuery { name }): Query<SyncResourceTypesQuery>,
|
||
) -> error::Result<String> {
|
||
require_super_admin(&db, &authed).await?;
|
||
|
||
use windmill_common::worker::HUB_RT_CACHE_DIR;
|
||
let cache_path = format!("{}/resource_types.json", *HUB_RT_CACHE_DIR);
|
||
|
||
// Manual sync is hub-first so it lands newly-published hub types on demand. The
|
||
// on-disk cache is only a fallback for when the hub is unreachable (airgapped
|
||
// installs / network error); refreshing it is left to the daily cache-rt cron and
|
||
// the startup sync in main.rs, which own the offline path.
|
||
let (resource_types, from_hub) = match fetch_resource_types_from_hub().await {
|
||
Ok(types) => {
|
||
tracing::info!("Fetched {} resource types live from the hub", types.len());
|
||
(types, true)
|
||
}
|
||
Err(hub_err) => {
|
||
tracing::warn!(
|
||
"Live hub fetch failed ({hub_err}), falling back to on-disk cache at {cache_path}"
|
||
);
|
||
match tokio::fs::read_to_string(&cache_path).await {
|
||
Ok(content) => {
|
||
let parsed = serde_json::from_str::<Vec<CachedResourceType>>(&content)
|
||
.map_err(|e| {
|
||
error::Error::InternalErr(format!(
|
||
"Failed to parse cached resource types: {}",
|
||
e
|
||
))
|
||
})?;
|
||
(parsed, false)
|
||
}
|
||
Err(_) => return Err(hub_err),
|
||
}
|
||
}
|
||
};
|
||
|
||
let mut synced_count = 0;
|
||
|
||
for rt in &resource_types {
|
||
// A name too long for the column counts as absent, leaving the stored one alone: one bad
|
||
// entry must not fail the upsert and end the rest of the sync.
|
||
let display_name = match &rt.display_name {
|
||
Some(Some(name)) if name.chars().count() > 100 => None,
|
||
other => other.clone(),
|
||
};
|
||
let exists: Option<bool> = sqlx::query_scalar!(
|
||
"SELECT EXISTS(SELECT 1 FROM resource_type WHERE workspace_id = 'admins' AND name = $1 AND schema IS NOT DISTINCT FROM $2 AND description IS NOT DISTINCT FROM $3 AND ($5 IS NOT TRUE OR format_extension IS NOT DISTINCT FROM $4) AND ($7 IS NOT TRUE OR display_name IS NOT DISTINCT FROM $6))",
|
||
&rt.name,
|
||
rt.schema.as_ref(),
|
||
rt.description.as_deref(),
|
||
rt.format_extension.clone().flatten(),
|
||
rt.format_extension.is_some(),
|
||
display_name.clone().flatten(),
|
||
display_name.is_some(),
|
||
)
|
||
.fetch_one(&db)
|
||
.await?;
|
||
|
||
if exists.unwrap_or(false) {
|
||
continue;
|
||
}
|
||
|
||
sqlx::query!(
|
||
// Whether the payload carried the key at all is what decides: present
|
||
// (even as null) is authoritative and may clear, absent means a cache
|
||
// written before the column and must leave the stored value alone.
|
||
"INSERT INTO resource_type (workspace_id, name, schema, description, format_extension, display_name, edited_at)
|
||
VALUES ('admins', $1, $2, $3, $4, $6, now())
|
||
ON CONFLICT (workspace_id, name) DO UPDATE
|
||
SET schema = EXCLUDED.schema, description = EXCLUDED.description,
|
||
-- A fileset is a set of files, so it cannot also be one file.
|
||
-- Create and update reject the pair; this writer bypasses both, so
|
||
-- it declines the extension rather than persisting the forbidden
|
||
-- combination onto a same-named local fileset.
|
||
format_extension = CASE
|
||
WHEN resource_type.is_fileset THEN NULL
|
||
WHEN $5 THEN EXCLUDED.format_extension
|
||
ELSE resource_type.format_extension END,
|
||
display_name = CASE WHEN $7 THEN EXCLUDED.display_name ELSE resource_type.display_name END,
|
||
edited_at = now()",
|
||
&rt.name,
|
||
rt.schema.as_ref(),
|
||
rt.description.as_deref(),
|
||
rt.format_extension.clone().flatten(),
|
||
rt.format_extension.is_some(),
|
||
display_name.clone().flatten(),
|
||
display_name.is_some(),
|
||
)
|
||
.execute(&db)
|
||
.await?;
|
||
|
||
synced_count += 1;
|
||
}
|
||
|
||
// If a specific type was requested and is still absent after syncing, surface an
|
||
// explicit not-found instead of a silent "Synced 0". Word it by source so the
|
||
// cache-fallback path does not claim it checked the hub.
|
||
if let Some(name) = name.as_deref() {
|
||
if !resource_types.iter().any(|rt| rt.name == name) {
|
||
let source = if from_hub {
|
||
"on the hub"
|
||
} else {
|
||
"in the cached resource types (hub unreachable)"
|
||
};
|
||
return Err(error::Error::NotFound(format!(
|
||
"resource type '{}' not found {}",
|
||
name, source
|
||
)));
|
||
}
|
||
}
|
||
|
||
Ok(format!(
|
||
"Synced {} resource types ({} unchanged)",
|
||
synced_count,
|
||
resource_types.len() - synced_count
|
||
))
|
||
}
|
||
|
||
#[cfg(test)]
|
||
mod tests {
|
||
use std::collections::BTreeMap;
|
||
use windmill_common::instance_config::{GlobalSettings, InstanceConfig, WorkerGroupConfig};
|
||
|
||
#[test]
|
||
fn supavisor_refusing_custom_instance_user_is_named() {
|
||
use windmill_common::error::{to_anyhow, Error};
|
||
let connect_error = |message: &str| {
|
||
let e = Error::from(to_anyhow(std::io::Error::other(message.to_string())));
|
||
super::custom_instance_user_connect_error(
|
||
"aws-0-eu-west-1.pooler.supabase.com",
|
||
"dt",
|
||
e,
|
||
)
|
||
};
|
||
for supavisor in [
|
||
"db error: FATAL: (ENOIDENTIFIER) no tenant identifier provided",
|
||
"db error: FATAL: Tenant or user not found",
|
||
] {
|
||
assert!(
|
||
matches!(connect_error(supavisor), Error::BadConfig(m) if m.contains("connection pooler")),
|
||
"{supavisor}"
|
||
);
|
||
}
|
||
assert!(matches!(
|
||
connect_error(
|
||
"db error: FATAL: password authentication failed for user \"custom_instance_user\""
|
||
),
|
||
Error::ExecutionErr(_)
|
||
));
|
||
}
|
||
|
||
#[test]
|
||
fn instance_config_yaml_round_trip() {
|
||
let config = InstanceConfig {
|
||
global_settings: GlobalSettings {
|
||
base_url: Some("https://windmill.example.com".to_string()),
|
||
retention_period_secs: Some(86400),
|
||
expose_metrics: Some(true),
|
||
..Default::default()
|
||
},
|
||
worker_configs: BTreeMap::from([(
|
||
"default".to_string(),
|
||
WorkerGroupConfig {
|
||
worker_tags: Some(vec!["deno".to_string(), "python3".to_string()]),
|
||
init_bash: Some("apt-get update".to_string()),
|
||
..Default::default()
|
||
},
|
||
)]),
|
||
};
|
||
|
||
let yaml = config.to_sorted_yaml().unwrap();
|
||
|
||
// Verify key fields appear in the YAML output
|
||
assert!(yaml.contains("base_url: https://windmill.example.com"));
|
||
assert!(yaml.contains("retention_period_secs: 86400"));
|
||
assert!(yaml.contains("expose_metrics: true"));
|
||
assert!(yaml.contains("default:"));
|
||
assert!(yaml.contains("- deno"));
|
||
assert!(yaml.contains("- python3"));
|
||
assert!(yaml.contains("init_bash: apt-get update"));
|
||
|
||
// Round-trip back to struct
|
||
let deserialized: InstanceConfig = serde_yml::from_str(&yaml).unwrap();
|
||
assert_eq!(
|
||
deserialized.global_settings.base_url.as_deref(),
|
||
Some("https://windmill.example.com")
|
||
);
|
||
assert_eq!(
|
||
deserialized.global_settings.retention_period_secs,
|
||
Some(86400)
|
||
);
|
||
assert_eq!(deserialized.global_settings.expose_metrics, Some(true));
|
||
let wc = &deserialized.worker_configs["default"];
|
||
assert_eq!(
|
||
wc.worker_tags.as_deref(),
|
||
Some(["deno".to_string(), "python3".to_string()].as_slice())
|
||
);
|
||
assert_eq!(wc.init_bash.as_deref(), Some("apt-get update"));
|
||
}
|
||
|
||
#[test]
|
||
fn sorted_yaml_global_settings_alphabetical() {
|
||
let config = InstanceConfig {
|
||
global_settings: GlobalSettings {
|
||
retention_period_secs: Some(3600),
|
||
base_url: Some("https://test.com".to_string()),
|
||
expose_metrics: Some(true),
|
||
email_domain: Some("example.com".to_string()),
|
||
..Default::default()
|
||
},
|
||
worker_configs: BTreeMap::new(),
|
||
};
|
||
|
||
let yaml = config.to_sorted_yaml().unwrap();
|
||
|
||
// Keys must appear in alphabetical order
|
||
let base_url_pos = yaml.find("base_url:").unwrap();
|
||
let email_pos = yaml.find("email_domain:").unwrap();
|
||
let expose_pos = yaml.find("expose_metrics:").unwrap();
|
||
let retention_pos = yaml.find("retention_period_secs:").unwrap();
|
||
|
||
assert!(
|
||
base_url_pos < email_pos && email_pos < expose_pos && expose_pos < retention_pos,
|
||
"global_settings keys should be alphabetically sorted, got yaml:\n{yaml}"
|
||
);
|
||
}
|
||
|
||
#[test]
|
||
fn sorted_yaml_worker_configs_default_and_native_first() {
|
||
let config = InstanceConfig {
|
||
global_settings: GlobalSettings::default(),
|
||
worker_configs: BTreeMap::from([
|
||
(
|
||
"gpu".to_string(),
|
||
WorkerGroupConfig {
|
||
init_bash: Some("echo gpu".to_string()),
|
||
..Default::default()
|
||
},
|
||
),
|
||
(
|
||
"native".to_string(),
|
||
WorkerGroupConfig {
|
||
init_bash: Some("echo native".to_string()),
|
||
..Default::default()
|
||
},
|
||
),
|
||
(
|
||
"default".to_string(),
|
||
WorkerGroupConfig {
|
||
init_bash: Some("echo default".to_string()),
|
||
..Default::default()
|
||
},
|
||
),
|
||
(
|
||
"alpha".to_string(),
|
||
WorkerGroupConfig {
|
||
init_bash: Some("echo alpha".to_string()),
|
||
..Default::default()
|
||
},
|
||
),
|
||
]),
|
||
};
|
||
|
||
let yaml = config.to_sorted_yaml().unwrap();
|
||
|
||
let default_pos = yaml.find("default:").unwrap();
|
||
let native_pos = yaml.find("native:").unwrap();
|
||
let alpha_pos = yaml.find("alpha:").unwrap();
|
||
let gpu_pos = yaml.find("gpu:").unwrap();
|
||
|
||
assert!(
|
||
default_pos < native_pos
|
||
&& native_pos < alpha_pos
|
||
&& alpha_pos < gpu_pos,
|
||
"worker_configs should have default, native first, then rest alphabetically, got yaml:\n{yaml}"
|
||
);
|
||
}
|
||
|
||
#[test]
|
||
fn sorted_yaml_roundtrips() {
|
||
let config = InstanceConfig {
|
||
global_settings: GlobalSettings {
|
||
base_url: Some("https://rt.test".to_string()),
|
||
retention_period_secs: Some(7200),
|
||
expose_metrics: Some(false),
|
||
..Default::default()
|
||
},
|
||
worker_configs: BTreeMap::from([
|
||
(
|
||
"default".to_string(),
|
||
WorkerGroupConfig {
|
||
worker_tags: Some(vec!["deno".to_string()]),
|
||
..Default::default()
|
||
},
|
||
),
|
||
(
|
||
"native".to_string(),
|
||
WorkerGroupConfig {
|
||
init_bash: Some("echo hi".to_string()),
|
||
..Default::default()
|
||
},
|
||
),
|
||
]),
|
||
};
|
||
|
||
let yaml = config.to_sorted_yaml().unwrap();
|
||
let deserialized: InstanceConfig = serde_yml::from_str(&yaml).unwrap();
|
||
|
||
assert_eq!(
|
||
deserialized.global_settings.base_url.as_deref(),
|
||
Some("https://rt.test")
|
||
);
|
||
assert_eq!(
|
||
deserialized.global_settings.retention_period_secs,
|
||
Some(7200)
|
||
);
|
||
assert_eq!(deserialized.global_settings.expose_metrics, Some(false));
|
||
assert_eq!(deserialized.worker_configs.len(), 2);
|
||
assert_eq!(
|
||
deserialized.worker_configs["default"]
|
||
.worker_tags
|
||
.as_deref(),
|
||
Some(["deno".to_string()].as_slice())
|
||
);
|
||
assert_eq!(
|
||
deserialized.worker_configs["native"].init_bash.as_deref(),
|
||
Some("echo hi")
|
||
);
|
||
}
|
||
}
|
||
|
||
#[cfg(all(test, feature = "parquet"))]
|
||
mod object_storage_test_hardening {
|
||
use super::{extract_host, validate_object_storage_test};
|
||
use windmill_object_store::ObjectSettings;
|
||
|
||
// IP literals (not hostnames) keep validate_public_endpoint deterministic — `lookup_host`
|
||
// parses them without any network round-trip.
|
||
fn gcs_settings(gcs_base_url: &str) -> ObjectSettings {
|
||
serde_json::from_value(serde_json::json!({
|
||
"type": "Gcs",
|
||
"bucket": "b",
|
||
"serviceAccountKey": { "gcs_base_url": gcs_base_url, "client_email": "x@y.z" }
|
||
}))
|
||
.unwrap()
|
||
}
|
||
|
||
#[tokio::test]
|
||
async fn rejects_gcs_internal_base_url() {
|
||
// gcs_base_url in the service-account key must not smuggle an internal host past the check,
|
||
// including via a mixed-case scheme (URL schemes are case-insensitive).
|
||
for url in [
|
||
"http://169.254.169.254",
|
||
"HTTP://169.254.169.254",
|
||
"Https://10.0.0.5",
|
||
] {
|
||
assert!(
|
||
validate_object_storage_test(&gcs_settings(url))
|
||
.await
|
||
.is_err(),
|
||
"{url} should be rejected"
|
||
);
|
||
}
|
||
}
|
||
|
||
#[tokio::test]
|
||
async fn allows_gcs_public_base_url() {
|
||
assert!(
|
||
validate_object_storage_test(&gcs_settings("https://8.8.8.8"))
|
||
.await
|
||
.is_ok()
|
||
);
|
||
}
|
||
|
||
#[tokio::test]
|
||
async fn rejects_gcs_blank_service_account_key() {
|
||
// A blank key makes build_gcs_client fall back to the instance's ambient credentials, so an
|
||
// untrusted caller must not be allowed to test with it. The `serviceAccountKey` field is
|
||
// serialized via serde's `as_string` (`to_string` of the JSON value), so the settings UI's
|
||
// "no key" empty object arrives as `"{}"` and a null as `"null"` — both must be rejected.
|
||
for key in [serde_json::json!({}), serde_json::json!(null)] {
|
||
let settings: ObjectSettings = serde_json::from_value(serde_json::json!({
|
||
"type": "Gcs",
|
||
"bucket": "b",
|
||
"serviceAccountKey": key
|
||
}))
|
||
.unwrap();
|
||
assert!(
|
||
validate_object_storage_test(&settings).await.is_err(),
|
||
"blank key {key:?} should be rejected"
|
||
);
|
||
}
|
||
}
|
||
|
||
#[test]
|
||
fn extracts_host_from_endpoint() {
|
||
let cases = [
|
||
("s3.amazonaws.com", Some("s3.amazonaws.com")),
|
||
("https://minio.internal:9000", Some("minio.internal")),
|
||
("http://10.0.0.5:9000/bucket", Some("10.0.0.5")),
|
||
("user:pass@host.example:443", Some("host.example")),
|
||
("[::1]:9000", Some("::1")),
|
||
("https://[fe80::1]/x", Some("fe80::1")),
|
||
("", None),
|
||
// Injection via region/bucket interpolation into the default endpoint string: the
|
||
// userinfo `@` and the path `/` must not hide the real authority from the host check.
|
||
(
|
||
"https://s3.@169.254.169.254/.amazonaws.com",
|
||
Some("169.254.169.254"),
|
||
),
|
||
(
|
||
"https://@169.254.169.254/mybucket.s3.amazonaws.com",
|
||
Some("169.254.169.254"),
|
||
),
|
||
("s3.#@169.254.169.254/x.amazonaws.com", Some("s3.")),
|
||
// Scheme is case-insensitive.
|
||
("HTTP://169.254.169.254", Some("169.254.169.254")),
|
||
];
|
||
for (input, expected) in cases {
|
||
assert_eq!(extract_host(input).as_deref(), expected, "input: {input}");
|
||
}
|
||
}
|
||
}
|