Files
windmill/backend/windmill-api-workspaces/src/datatable_acl.rs
T
e1e3692fbc feat: data table roles in the DB manager and raw apps (#11139)
* feat(datatables): put a data table's connection under Postgres roles

A data table backed by the instance database resolved to exactly one Postgres connection,
`custom_instance_user`, for everyone who could reach it at all. There was no way to say
this job reads, that one writes, this one never sees the salaries table.

A data table role is now a real Postgres login on the cluster, defined once for the
instance by a superadmin and named exactly as they named it. A script that declares
`-- role analytics` connects as `analytics`, and Postgres decides what it may touch —
grants are ordinary SQL. Windmill answers only "may this caller ask for this role", from
the tenant lists on the data table entry: `u/alice`, `g/analysts`, `f/finance` or `*`.
A data table with no `permissions` block behaves exactly as before.

Everything that opens a connection on someone's behalf goes through one chokepoint,
`get_datatable_resource_from_db`, which takes the identity explicitly and fails closed when
there is none. The role logs in as itself — never `SET ROLE`, which a script could
`RESET ROLE` its way out of.

A fork's data table entry becomes a pointer at the workspace that governs it rather than a
copy of it. The settings clone used to hand a fork a byte-identical entry naming the
parent's database, which a fork admin could edit to grant themselves `admin` there; a
pointer has nothing local to edit, and its tenants are evaluated as a member of the
governing workspace, by email. `permissions` is stripped from the workspace export and
ignored on import: tenants name principals of one workspace, and a settings push is not
where an access decision should be made.

Operations that see the whole database whatever the roles grant stay with the governing
workspace's admins: editing the roles, a migration that declares none, and opening a
replication stream for a Postgres trigger or capture.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): gate the paths that reach a whole database as admin

Auditing what still resolved through the unchecked resolver turned up three that act for a
caller and hand back the admin connection: `resolve_pg_source_checked` (behind schema
export, the full-schema read, database creation, import and the forked-database drop), the
connection test, and the schema snapshot a fork clone takes of its parent. On a data table
under roles each let any workspace member — or a fork admin who is nobody in the governing
workspace — read or copy the whole database whatever its roles grant.

All three now require admin reach on the governing workspace. A dump taken under a
restricted role would be a silently truncated copy rather than an error, so refusing is the
only right answer for the copy paths.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): confine roles to the instance database, and stop a fork reaching the parent's bookkeeping

A data table role is a login on Windmill's own Postgres. Nothing stopped a workspace admin
putting a *resource-backed* data table under roles, at which point the executor dialled the
host that resource names — one the admin chose — with the role's real cluster password, and
`CONNECT` is granted to every registered instance database. Both ends now refuse: the
permissions endpoint rejects the save, and the chokepoint refuses to substitute credentials
on a non-instance entry rather than trusting the record it read.

Two more places reached the governing database without answering to it. The initial-migration
generator returned a `pg_dump` of the whole schema to any member. And the migration
rename/delete cascade followed a fork's pointer into the parent, so a fork admin renaming or
removing their own local entry relabelled or wiped the parent's `_wm_migrations` — after
which the parent re-runs every migration from zero. The remote half is now skipped when the
entry resolves into another workspace, which is also just correct: a fork renaming what it
calls a data table changes nothing about the data table.

Also: revoking a tenant now bounces the replication streams of every workspace holding an
entry that resolves here, not only the governing one, so a fork's trigger stops rather than
living on inside its open connection; the instance role catalog and the governing workspace's
tenant lists are no longer returned to someone who cannot edit them; and the tenant rename
dedup collapses non-adjacent duplicates, per role rather than once any role changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): fail loudly where a role or a pointer can be left half-recorded

Three ways the feature could end up in a state nobody could see or undo.

Creating a role writes the cluster first and the catalog second, but the catalog write was an
`UPDATE` that matched nothing when the instance Postgres settings row was absent — leaving a
live login with a password nobody recorded: invisible to the catalog, un-recreatable because
the name is taken, and un-deletable because there is no entry to delete. It now errors, so
the operation is retryable once the row is restored.

Deleting a workspace only nulls the fork lineage; the data table entries pointing at it are
left resolving to nothing. Sweeping them is not an option — turning a pointer back into a copy
would hand each fork the database outright — so the delete now names the data tables it
stranded, and resolving one says which workspace is missing rather than reporting a data table
this workspace never had.

`InstanceDatatableRole` derived `Debug` while holding a Postgres password; it is now
hand-written so `{:?}` on the catalog cannot put a live credential in a log line.

Adds the two branches the reviews found unpinned: a caller who is not a member of the
governing workspace at all, and `NoIdentity` — the compatibility path for an agent worker that
predates this and sends no job id.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): unbreak two operator messages and two comments that described other code

The two strings this branch added for states an operator hits once — the catalog write that
matched nothing, and the delete that stranded a pointer — were collapsed from their multi-line
form with the indentation left in, so both rendered with a fourteen-space gap mid-sentence.

`list_datatables` claimed to report a chain it cannot follow and then dropped it; it does drop
it, and the comment now says why that is the right place to stay quiet. The non-superadmin
check in `edit_datatable_config` was introduced as also covering references, which it does not
and need not: `reference` is overwritten from the stored entry for every caller before the
check runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): serialize role catalog mutations, and state each helper's authorization contract

The catalog is one JSON document, so create, rename, enable and delete are all
read-modify-write. Two concurrent creates read the same snapshot, both succeed in the
cluster, and the second write drops the first — leaving a live Postgres login with a password
nobody recorded, which is the exact state the delete path exists to prevent. Every mutation
now runs in one transaction holding an advisory lock across the read, the cluster DDL and the
write, so a lost update cannot happen and a failure rolls the whole thing back. The DDL
helpers take that transaction rather than the pool, which is what makes the lock cover them.

Their statements moved off `sqlx::raw_sql`: the simple protocol is only needed for genuinely
multi-statement SQL, and its future is not `Send`, which an axum handler holding the
transaction requires. Each of these is one statement anyway.

The new cross-crate surface now says what callers must do. `read_role_catalog` returns
plaintext credentials; `create`/`rename`/`set_login`/`drop_instance_role` and
`converge_connect_grants` mutate cluster-wide state; `read_datatable_entry` reads a workspace's
raw config. All of them are superadmin-gated by their current handlers, but nothing said so at
the definition, which is where the next caller looks.

Also: the roles table reloads after a failed login toggle instead of leaving it claiming a flip
that did not land; the rename affordance is the design-system `Button`, not a raw one; and
`resolve_datatable_pg_as_caller` drops a `role` parameter no caller ever filled — browsing
resolves as the data table's default until the database manager grows a picker.

Why role passwords stay a plain `String` while the instance user's password beside them is a
`StringOrSecretRef`, asked three times across reviews: that one is a secret ref because an
operator supplies it and may want it from their own backend, while these are minted here and
never entered by anyone, so there is nothing for a ref to point at. Encrypting generated
secrets at rest is a separate change that would take the replication password with it. Now
said at the field.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): give the role catalog its own row, out of reach of the config machinery

Putting it inside `custom_instance_pg_databases` was the wrong call, and it cost two ways.
The catalog serializes a generated Postgres password per role, and that row is the
operator-facing instance config, so the passwords reached `get_instance_config` and its YAML
editor — a live cluster credential in a response body, a UI field and any log of either.
Worse in the other direction: `to_settings_map` strips the catalog, so a full-row upsert of
that key writes the row back without it and the catalog is gone, while the cluster keeps every
login it described.

`custom_instance_replication_pwd` is the precedent and says exactly why — a generated secret,
written only by the server, never operator-authored, hidden so the config machinery cannot
read, rewrite or drop it. The catalog is the same thing, so it now has the same shape:
`datatable_roles`, in `HIDDEN_SETTINGS`, `PROTECTED_SETTINGS` and the agent-worker denylist.
No redaction to keep in step with three code paths, and no way for a neighbouring write to
take it out.

Two races on the same shared documents. `edit_datatable_config` read the stored data tables
outside its transaction and then wrote the whole `datatable` document, so a permissions save
committing in between was silently rolled back; it now reads under `FOR UPDATE`. And
`set_datatable_permissions` validated role ids against the catalog before opening its
transaction, so a deletion in between let it write a deleted role back — including as the
default, which every later job then fails on; it now holds the catalog lock and the settings
row across validation and write.

Completes the authorization contracts the previous commit claimed but did not finish:
`read_datatable_entry` (which it named and missed), `resolve_governing_datatable`, whose whole
job is to answer for a workspace the caller may not belong to, and
`converge_connect_grants_with`, which had not inherited its wrapper's.

Also the generic Python SDK reference: `_format_py_params` learned the bare `*` last time, but
`extract_py_functions` is a second formatter and still rendered `datatable(name, role)`, so
code written from that page passed a keyword-only argument positionally.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): make the concurrency test pin the handlers, and the contracts describe what is enforced

The concurrency test reimplemented the read-modify-write inline, so deleting the lock from all
three handlers left it green — it pinned Postgres, not the code it was written for. It now
drives `create_datatable_role` twice concurrently and asserts the catalog kept both names.
Checked the way the last one should have been: removing the lock from the handler makes it
fail with "wmtest_a_… is a live cluster login the catalog forgot".

The contracts added last commit were stricter than this PR's own callers, which is worse than
none — the next reader sees a rule already broken and learns to ignore it.
`read_role_catalog` said superadmin-only while two of its four callers are open to any
workspace member, and `converge_connect_grants` said superadmin while
`set_datatable_permissions` reaches it as a workspace admin. Both were fine on substance: the
rule that actually holds is about the credential never reaching a response, log, audit record
or export, not about who may call. They now say that. `read_datatable_entry` gets the same
treatment rather than the one the earlier message claimed for it: it is the primitive every
resolution goes through, so it is deliberately open, and what must not escape is `permissions`
— it names the governing workspace's users, groups and folders.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): close the last ways a role or a pointer can be left pointing at nothing

The raw settings readers hand back whatever is in the row, so moving the catalog into its own
`global_settings` key protected the config machinery and left `GET /settings/global/datatable_roles`
and the settings listing returning every live password. Both now filter that one key. The
neighbouring `custom_instance_replication_pwd` has the same shape and is not touched here: it
predates this and widening the fix to it is a decision about an operator workflow, not a
consequence of this change.

Three ways a save could leave something resolving to nothing:

A permissioned data table could be moved to a PostgreSQL resource. The block was carried across
as a server-owned field, the runtime refuses roles on a resource-backed table, so the save
succeeded and every job afterwards failed. Refused instead — turning roles off first is one step,
and it keeps discarding an access decision something somebody chose.

Renaming a governing data table left every fork pointing at the old name: the data table
disappears from their pickers and their jobs stop, with nothing in the renaming workspace to
suggest why. The rename now follows into the pointers in the same transaction.

Deleting one cannot be followed the same way, so it is reported instead — the response names what
it stranded, the way deleting a workspace does, and the fork's own error already says which
workspace is gone.

Also: `ensure_instance_db_grant_options_unchecked` claimed superadmin while the permissions
handler reaches it as a workspace admin (the same class fixed last commit, one instance missed);
the role entry kept an `instance_config_schema` derive it no longer needs; `write_role_catalog`
was the one writer of that table not stamping `updated_at`; and the concurrency test dropped its
roles only on success — a failing run is exactly the one that creates them without recording them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* refactor(datatables): put the role catalog in its own table, not in global_settings

Five findings across three rounds were all the same choice. A set of live Postgres credentials
was living in `global_settings`, which has generic read, list, write, config-export and CLI
round-trip paths that know nothing about what they carry: the passwords reached the instance
config and its YAML editor, a full-row upsert of a neighbouring key erased the catalog,
`GET /settings/global/{key}` and the settings listing returned them raw, and this round the
redaction that fixed the last two turned `wmill instance push` into something that wipes every
password — a fix breaking the assumption the previous fix made. `POST /settings/global/datatable_roles`
could also empty it outside the lock.

The approved plan offered a table or `global_settings`, so this is the other option it already
allowed rather than a new design. `datatable_role` is a table: no generic settings path can read
it, list it, export it, write it or round-trip it, so none of the five needs a guard. The
redaction, the hidden/protected/agent-denylist entries and the JSON document all go with it.

One row per role also removes the read-modify-write the concurrency work was about: two
concurrent creates are two inserts, and the unique index on `name` is what settles a collision.
The advisory lock stays for the one window rows do not cover — `CREATE ROLE` is invisible to
another transaction until commit, so without it both creates pass their `pg_roles` check.

Also from this round: rename mappings are checked against the configuration they claim to
describe, since fork pointers are rewritten from them — a caller could otherwise submit
`main -> missing` against an unchanged config and repoint every fork of `main` at a name nothing
has, and `A -> B` plus `B -> C` moved what pointed at `A` all the way to `C`. And the warning
naming forks a delete stranded reached the response but not the screen: both the data table
settings save and the workspace delete now show it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): validate a rename against the save it describes, and re-check under the locks

Three from the round, all about deciding on state that could already have moved.

A permission save resolved the data table and checked it was instance-backed before taking any
lock, then wrote under one. A config save committing in between could move the table onto a
PostgreSQL resource — recreating exactly what the transition guard refuses — or rename it, in
which case the write targeted a key that no longer existed and reported success having changed
nothing. It now re-resolves and re-checks on the locked state.

Rename validation checked that the source existed before and the target existed after, which
still accepts `main -> decoy` against a save that keeps both: every fork of `main` then follows
onto a different data table, silently, because it keeps resolving. The rule is now the actual
old-to-new key transition — a source may only survive if another rename took its name, and a
target may only pre-exist if another rename freed it. That also stops two sources sharing one
target, and it admits a swap, which the previous guard refused: `datatables` is keyed by name, so
a swap cannot be done one save at a time, and refusing it was a regression against main. The
pointer cascade now runs in two passes through a temporary name, the way the migration cascade
one layer down already handles the same shape, so `A -> B` with `B -> C` moves each pointer once
from what it named before the save.

The tenant mutators say what they are for: they write an access decision for any workspace named,
with an arbitrary mutation, and exist for the transaction that frees or renames a principal.
Editing a decision on purpose belongs in the permissions endpoint.

Carried in the same change: the stranded-fork list is a field rather than a phrase to grep out of
a success string; the pointer cascade matches with `EXISTS` instead of a `LIKE` over the whole
document, so a workspace whose pointers name something else is not rewritten to a byte-identical
value under an exclusive lock; and `InstanceDatatableRole` drops the serde derives left over from
the JSON document, one of which would emit `pwd`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): cascade on the leave route that is used, gate migrations before the admin connection, and drop a role atomically

The tenant cascade on leaving went onto `/users/leave`. The UI and the generated client call
`/workspaces/leave` — a different handler in a different crate with the same name — which
deleted the membership and left `u/<username>` in the tenant lists. Leaving and rejoining
therefore restored the access the leave was supposed to end, and a later account taking the
username would have inherited it. The regression test drives the route the client actually
calls; without the fix it fails with "leaving kept the tenant".

The migration endpoints authorized too late. `run_datatable_migrations` opened the data table's
admin connection, created `_wm_migrations` and read it before reaching the per-migration role
check — so with nothing pending, nothing was checked at all. Rollback returned before its check
when nothing was applied, and the status endpoint had none. All three now ask, before any
connection is opened, whether the caller can reach the data table as any role at all; which role
a given migration runs as is still decided per migration, and by the executor after that.

Deleting a role committed the cluster drop and the catalog row, then swept the tenant lists in
separate transactions. A sweep failing part-way left workspaces naming a role nothing can connect
as, while the retry answered `NotFound` because the catalog entry was already gone. The sweep now
runs in the same transaction, so the drop, the row and every tenant list commit together.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): refuse to copy a data table that is under roles

pg_dump carries no roles and the import runs with --no-privileges, so a copied
data table arrives owned by the admin connection with no GRANT for any role.
The settings clone brings `permissions` across, so the fork's tenants pass
Windmill's check, connect as the role they were given, and are denied by
Postgres on everything: an entry that reads as configured and answers nothing.

Refuse the copy — in the import endpoint before any data moves, and in the fork
path the CLI takes. Replaying the source's owners and ACLs into the clone is
what lifts this, and is a change of its own. Dropping `permissions` from the
copy instead would be the unsafe half, since the copy holds the parent's rows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): refuse the clone's database too, not only its data

A clone is two endpoints: `create_pg_database` then `import_pg_database`. Only
the second refused a data table under roles, so a fork asking to clone one
created and registered an empty `wm_fork_…` instance database and then failed —
and nothing collects it, since `drop_forked_datatable_databases` only drops
entries carrying `forked_from` and no entry names this one.

Refuse in both, so the clone stops before a database exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* nit worker error msg

* fix pg_dump stuck on version 17 on nix

* fix(datatables): refuse a malformed role annotation instead of ignoring it

`-- Role operator`, `-- role operator;` and `-- role operator -- why` all failed
the annotation parser's exact-match rule, so the query fell through to the data
table's default role and ran, silently, under a login the author did not choose.
Naming a role exists precisely to not do that.

A leading comment whose first word is `role` is now an annotation attempt: the
keyword matches case-insensitively, one trailing `;` is tolerated, and anything
else is an error naming the line. Only callers that already know the target is a
`datatable://` reference ever run this, so ordinary SQL keeps its comments.

Also bumps the dev shell's postgres client to 18 — it trailed the server the dev
database runs, which takes out every data table export, clone and fork-with-data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): refuse a malformed role query string instead of ignoring it

`?Role=analytics`, `?role=` and `?x=1&role=…` all fell through the reference
parser's exact-match rule, so the connection resolved to the data table's default
role and ran under a login the caller never asked for — the URI half of the same
trap as a malformed `-- role` annotation.

The key now matches case-insensitively, and anything else in the query string is
an error naming it; `role` is the only parameter a reference takes. Callers that
only need the entry keep a lenient `datatable_ref_name`, since they never act on
the role. The DuckDB `ATTACH` parser propagates it rather than attaching under
the default.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): carry the role annotation into the row_to_json retry

The retry rebuilds its SQL from `pruneComments(code)`, so the leading comment
block never reached the second attempt — and with it the `-- role <name>` line
that decides which login the query runs as. The retry connected as the data
table's default role instead, so a query the first attempt was denied could
succeed on the second, reported as "recovered with the row_to_json fix".

Carry the leading comment block over. The retry itself is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* chore(datatables): don't mount the roles UI until the ACL editor lands

Enforcement ships first. The permissions drawer is what turns roles on, and the
catalog section is what creates them — both are only useful once there is a way
to grant a role the privileges it needs, which arrives with the ACL editor. Left
mounted they would offer a feature whose other half does not exist.

The two components are complete and reviewed; only their call sites here are
commented out, with a note pointing the follow-up PRs at them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): honour `-- role: x`, and fix the DuckDB attach test

Two review findings, both real.

`attach_datatable_parses_name_and_role` never compiled: `parse_attach_datatable`
returns `Result<Option<_>>` now and one call site kept a single `unwrap`. Its
`?Role=analytics` case also asserted a refusal, contradicting the parser in the
same commit, which matches the key case-insensitively. Replaced with the cases
that are genuinely malformed, and a positive one for the cased key.

`-- role: analytics` fell through to the default role — the silent fallback the
strict parser exists to remove, for the spelling most likely to be typed. The
keyword now accepts an optional colon, attached or spaced, while a word that
merely starts with it (`rolebased`) is still not an attempt.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): clone a fork's pointer instead of failing after the copy

Forking a fork with cloning left an orphan database. The preflight resolves the
pointer and sees the governing entry, so both endpoints ran and filled the new
database; `apply_forked_datatable` then refused the inherited pointer and rolled
the fork back, stranding a registered `wm_fork_*` that no entry names and whose
name blocks the retry.

Refusing earlier would have been the smaller change, but forking a fork and
cloning worked before pointers existed, so it would trade an orphan for a
regression. Resolve what the pointer names and write the terminal entry the
clone needs: the whole `database` object rather than a patch of its
`resource_path`, since a pointer has none, and `reference` removed with it.

Also accepts `-- role=x` and `-- Role = x`, two more spellings that fell through
to the default role.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): refuse to roll back the catalog while roles exist

The down migration dropped the table and left every role behind: live Postgres
logins whose passwords only that table carried, so after a revert Windmill could
neither use, disable nor delete them, and re-applying could not recreate them
because the names were taken. Cleaning up here is not possible either — dropping
a role means reassigning what it owns in every instance database, and a
migration runs in one — so it now refuses while the catalog is non-empty and
says to delete the roles through instance settings, which does the cluster work.

Also enforces the instance-only invariant the resolved-pointer clone relies on
rather than only asserting it in a comment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* refactor(datatables): settle clonability in one place, before anything is created

A clone is three stages a workspace apart — `create_pg_database`, then
`import_pg_database`, then `apply_forked_datatable` inside the fork transaction.
Only the third can roll back, and `CREATE DATABASE` is not transactional, so any
refusal that lives there strands a registered `wm_fork_*` that no entry names
and whose name blocks the retry.

That orphan has now been fixed three times, most recently reintroduced by a
guard added one commit ago. Patching each new refusal into the first endpoint is
not the fix; having two places that can refuse is. `ensure_datatable_is_clonable`
now answers every reason a copy can be refused and returns what it resolved, and
the stage that writes the entry only does the work.

Also takes an ACCESS EXCLUSIVE lock before the rollback guard counts, so a role
created concurrently cannot slip between the check and the drop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): let a retried clone reclaim its own leftover database

A clone creates its target database one request before it copies into it, and
the fork that would name it is written a request after that. Any failure in
between — a pg_dump error, a bad restore, a dropped connection, the source's
roles changing mid-flow — left a registered `wm_fork_*` that no entry names,
and every retry then failed on its name. This predates data table roles.

`create_pg_database` now reclaims such a leftover before creating: only a
`wm_fork_*` database Windmill registered as a data table database and that no
data table or ducklake entry names, in any workspace, archived ones included.
The drop never terminates connections, so a clone still copying into it makes
the reclaim fail instead of being cut off. It is limited to callers who
administer the source — reaching it is not enough, since on a data table
without roles every member reaches it — and anyone else gets the refusal an
existing database always got.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Revert "fix(datatables): let a retried clone reclaim its own leftover database"

This reverts commit 7dd3275a10.

The reclaim tied the caller to the source they administer, but not to the
database it dropped. Between another workspace's import and its final fork
request, that workspace's target is full, registered, unnamed and has no open
connection, so an admin of any instance data table could name it and have it
dropped and recreated empty. The victim's fork would then commit pointing at
the empty copy. Safe reclaim needs durable clone ownership and serialization
with the request that names the database; until then the leftover stays, as it
did before this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(datatables): record the stale clone database as a known limitation

A clone is three requests and `CREATE DATABASE` is not transactional, so a
failure after the first leaves a registered `wm_fork_*` behind, as it did
before data table roles. Accepted for this PR: it is harmless to data and goes
away once the clone is a single server-side operation.

The comment also records why the obvious fix is wrong: reclaiming the leftover
on retry, without durable clone ownership, can drop another workspace's fully
copied database between its import and its final fork request.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): bounce the streams reading a data table when it is deleted

Deleting a governing data table, or the workspace that holds it, only collected
the fork pointers it stranded, for the warning. A Postgres trigger or capture
already streaming through one of those pointers kept the replication connection
it opened while the pointer still resolved, so it went on dispatching the
governing database's rows after the fork lost access — until its connection
happened to restart. The governing workspace's own streams on a deleted entry
did the same.

Both deletion paths now bounce the affected listeners inside their own
transaction, through the helper a permission change already uses, so a
listener that reconnects re-resolves the entry and finds it gone. The helper is
split so a caller can pass the (workspace, local name) pairs it already holds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep the fork schema baseline, and bounce streams on every removal

Three fixes from review.

`edit_datatable_config` took `forked_from` wholesale from the stored entry, so
the fork schema diff's save of an advanced baseline was silently discarded and
an applied change was offered again. Whether an entry carries a clone stamp is
still carried from the store, since that is what marks its database droppable,
but the baseline inside it is now taken from the request.

The stranded-pointer warning and the stream bounce ran over the optional
`deleted_datatables` hint, which the settings-sync CLI never sends, so removing
a governing data table through `wmill` bounced nothing. Removals are now derived
from the stored configuration against the saved one.

`delete_workspace` read the pointers to bounce before its transaction, so a fork
committing a pointer during the deletion was missed. The read now happens inside
the transaction, after the workspace row is deleted: a fork's insert key-share
locks that row through its parent foreign key, so it is either seen or fails on
the missing parent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(datatables): keep Postgres triggers and data table roles apart

A replication stream reads every row of every table whatever the data table's
roles grant, and its listener checks access only when it connects. Rather than
chase every way access can change and bounce the streams each one affects, a
data table now carries one or the other:

- a Postgres trigger or capture cannot be created on, or connect to, a data
  table under roles;
- roles cannot be turned on while an enabled trigger or a live capture reads
  the data table, its own or a fork's through its pointer. The refusal names
  each one to disable.

This removes the stream bounces on roles edits and on data table and workspace
deletion, and the trigger gate that admitted admins. The fork schema baseline
fix from the same review round is kept.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): refuse a Postgres trigger on a data table under roles when it is saved

Creating or editing a trigger that points at a data table under roles was
accepted, and its listener then retried the refused connection every 30
seconds forever. The save is now refused, and a trigger that reaches such a
data table anyway (re-enabled, or cloned into a fork) is disabled by its
listener with the reason, as a missing replication slot is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): disable a data table role before deleting it

Deleting a role reassigns and drops what it owns in each registered
database on its own connection, and each of those passes commits as it
goes. A database failing part-way left the role enabled in the catalog and
able to log in, but already stripped in the databases reached before it.
The role is now disabled in its own commit first, so a failed delete
leaves a disabled role to retry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): serialize roles going on with a stream starting

Turning roles on looked for enabled triggers and live captures once,
without a lock anything starting a stream also took. A trigger enabled in
that window could have its listener connect before roles committed, and a
healthy listener never checks again. Both transitions now serialize on one
advisory lock: roles going on hold it exclusive while they look, and
trigger create, edit and enable, and capture setup and ping hold it shared
while they commit. Either the look sees the stream, or the listener
connects after roles are committed and refuses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): wait out live listeners, and resolve stored names containing `?`

Turning roles on counted a trigger as gone once disabled, and a capture
once its client stopped pinging, but the listener keeps its replication
connection until its next heartbeat notices. A trigger or capture whose
listener pinged in the last 15 seconds, the window a server holds a
listener for, now still counts as streaming.

Data table names could contain `?` before they were restricted, and such
entries are still stored. Splitting `?role=` off a reference misread them:
`a?b` became `a` with an unknown parameter, and the clone checks looked at
a different entry than the one copied. An entry stored under the whole
reference is now looked up first, in the Postgres executor, DuckDB ATTACH
and the clone checks. Agent workers cannot read the workspace and keep
the strict parse, which refuses such a name rather than misreading it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): warn when a settings sync strands fork pointers

A settings save reported the fork pointers left resolving to nothing only
for the names in `deleted_datatables`, which `wmill sync push` never sends.
The save now works out what it removed from the locked entries, and the
CLI prints the stranded pointers it returns.

Also correct the replication helper's contract: no role or admin check
makes a replication connection safe, so a data table under roles is
refused outright rather than gated as an admin operation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): refuse a save that drops a data table's roles through an undeclared rename

A data table's roles follow its entry only through a declared rename. A
settings sync sends the whole map and never declares one, so renaming a
data table under roles there read as a delete and a new entry on the same
database: the new entry carried no roles, and every caller connected as
admin. Such a save is now refused, naming both entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): no entry without roles may newly reach a database under roles

The previous guard only caught a new name replacing an entry under roles.
A whole-map save could also repoint an existing entry without roles at
that database, or another workspace could point one there, and every
caller of that entry would connect as admin. The rule is now stated on
the saved entries: one that carries no roles and newly points at an
instance database any entry under roles uses, in this workspace or
another, is refused. A declared rename carries its roles and passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move data table role catalog and resolution to the enterprise edition

Roles are an Enterprise Edition feature. The catalog, the Postgres logins,
CONNECT convergence, tenant evaluation and the role half of connection
resolution move to windmill-ee-private. Every public function keeps its path
and signature and forwards through datatable_roles_oss, which re-exports the
enterprise implementation or, without it, refuses.

Without the enterprise edition a data table under roles, or a caller naming a
role, is refused a connection rather than resolved as admin, and the reach and
admin-access checks refuse one under roles. A data table not under roles
resolves as before in every edition, and an instance database keeps the
CONNECT grants it was created with. The catalog lock, the stream lock, the
tenant cascades and the permissions stripping stay in OSS: they only restrict.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move the data table permissions endpoints to the enterprise edition

The permissions read, save and usable-roles handlers move to
windmill-ee-private; the routes stay registered and, without the enterprise
edition, answer that data table roles are an Enterprise Edition feature.
ensure_governs_datatable and ensure_reaches_datatable keep their paths: the
first refuses, the second passes a data table not under roles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move the data table role catalog endpoints to the enterprise edition

The superadmin list, create, update and delete handlers move to
windmill-ee-private. The routes stay registered and, without the enterprise
edition, refuse after authentication.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* test(datatables): run the roles tests on the enterprise edition, refusals without it

Each test that exercises roles runs with private and enterprise. Two tests run
without them: every roles route answers the Enterprise refusal, and a data
table saved under roles, or a named role, is refused a connection while one
not under roles resolves as before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): gate the roles UI mount sites on an enterprise license

Both mount sites are still commented out; the gate travels with them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* test(datatables): run the tenant matcher test on the enterprise edition

The matcher it covers is enterprise code now, so without the enterprise
edition the test hit the stub and failed the default windmill-common run. It
runs with private and enterprise, and a counterpart without them asserts that
no tenant list covers anyone, the wildcard and a workspace admin included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* chore: update ee-repo-ref to a1873dbb67f2302b85ff5362f8387b48eccdb607

This commit updates the EE repository reference after PR #783 was merged in windmill-ee-private.

Previous ee-repo-ref: 5c853e2c20eca6b748415fc0d6862a6ebfb5fec4

New ee-repo-ref: a1873dbb67f2302b85ff5362f8387b48eccdb607

Automated by sync-ee-ref workflow.

* fix(datatables): refuse roles while a same-workspace alias reaches the database

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): add an ACL editor for data table roles

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): data table roles in the DB manager and raw apps

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: never add a role to the reference of a data table whose name contains '?'

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: read the roles of a data table whose name contains '?'

The generated client leaves a '?' in a path param unencoded, so the lookup
404'd and the raw-app picker blocked Start on such a data table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: take every pooled connection before the ACL apply locks

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: refresh grant options only after the ACL apply validates its plan

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix(datatables): refuse a reference naming both a legacy data table and a role

When a workspace stores both `sales` and a legacy `sales?role=analytics`, the
reference resolved to the legacy entry without a role, so browsing `sales` as
`analytics` reached another data table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: declare the default role in migrations written for a data table whose name contains '?'

Such a data table connects as its default role without naming it, so the
migrations the manager wrote for it declared no role and ran as admin.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): let CE migrations connect as an explicitly named admin

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: add only missing grant options before an ACL apply, never default privileges

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix(datatables): serialize roles going on with aliases saved from other workspaces

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(datatables): note that legacy names with ? cannot be migrated

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: run one data table ACL apply at a time per server before it connects

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: hold the ACL connection to the database that was authorized

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: build the ACL connection from the authorized data table entry

Resolving the settings again could land on a resource with the same
database name on another server, which the later entry checks never see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: check ACL read reach against the entry it connects from

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* feat(datatables): put a data table's connection under Postgres roles

A data table backed by the instance database resolved to exactly one Postgres connection,
`custom_instance_user`, for everyone who could reach it at all. There was no way to say
this job reads, that one writes, this one never sees the salaries table.

A data table role is now a real Postgres login on the cluster, defined once for the
instance by a superadmin and named exactly as they named it. A script that declares
`-- role analytics` connects as `analytics`, and Postgres decides what it may touch —
grants are ordinary SQL. Windmill answers only "may this caller ask for this role", from
the tenant lists on the data table entry: `u/alice`, `g/analysts`, `f/finance` or `*`.
A data table with no `permissions` block behaves exactly as before.

Everything that opens a connection on someone's behalf goes through one chokepoint,
`get_datatable_resource_from_db`, which takes the identity explicitly and fails closed when
there is none. The role logs in as itself — never `SET ROLE`, which a script could
`RESET ROLE` its way out of.

A fork's data table entry becomes a pointer at the workspace that governs it rather than a
copy of it. The settings clone used to hand a fork a byte-identical entry naming the
parent's database, which a fork admin could edit to grant themselves `admin` there; a
pointer has nothing local to edit, and its tenants are evaluated as a member of the
governing workspace, by email. `permissions` is stripped from the workspace export and
ignored on import: tenants name principals of one workspace, and a settings push is not
where an access decision should be made.

Operations that see the whole database whatever the roles grant stay with the governing
workspace's admins: editing the roles, a migration that declares none, and opening a
replication stream for a Postgres trigger or capture.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): gate the paths that reach a whole database as admin

Auditing what still resolved through the unchecked resolver turned up three that act for a
caller and hand back the admin connection: `resolve_pg_source_checked` (behind schema
export, the full-schema read, database creation, import and the forked-database drop), the
connection test, and the schema snapshot a fork clone takes of its parent. On a data table
under roles each let any workspace member — or a fork admin who is nobody in the governing
workspace — read or copy the whole database whatever its roles grant.

All three now require admin reach on the governing workspace. A dump taken under a
restricted role would be a silently truncated copy rather than an error, so refusing is the
only right answer for the copy paths.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): confine roles to the instance database, and stop a fork reaching the parent's bookkeeping

A data table role is a login on Windmill's own Postgres. Nothing stopped a workspace admin
putting a *resource-backed* data table under roles, at which point the executor dialled the
host that resource names — one the admin chose — with the role's real cluster password, and
`CONNECT` is granted to every registered instance database. Both ends now refuse: the
permissions endpoint rejects the save, and the chokepoint refuses to substitute credentials
on a non-instance entry rather than trusting the record it read.

Two more places reached the governing database without answering to it. The initial-migration
generator returned a `pg_dump` of the whole schema to any member. And the migration
rename/delete cascade followed a fork's pointer into the parent, so a fork admin renaming or
removing their own local entry relabelled or wiped the parent's `_wm_migrations` — after
which the parent re-runs every migration from zero. The remote half is now skipped when the
entry resolves into another workspace, which is also just correct: a fork renaming what it
calls a data table changes nothing about the data table.

Also: revoking a tenant now bounces the replication streams of every workspace holding an
entry that resolves here, not only the governing one, so a fork's trigger stops rather than
living on inside its open connection; the instance role catalog and the governing workspace's
tenant lists are no longer returned to someone who cannot edit them; and the tenant rename
dedup collapses non-adjacent duplicates, per role rather than once any role changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): fail loudly where a role or a pointer can be left half-recorded

Three ways the feature could end up in a state nobody could see or undo.

Creating a role writes the cluster first and the catalog second, but the catalog write was an
`UPDATE` that matched nothing when the instance Postgres settings row was absent — leaving a
live login with a password nobody recorded: invisible to the catalog, un-recreatable because
the name is taken, and un-deletable because there is no entry to delete. It now errors, so
the operation is retryable once the row is restored.

Deleting a workspace only nulls the fork lineage; the data table entries pointing at it are
left resolving to nothing. Sweeping them is not an option — turning a pointer back into a copy
would hand each fork the database outright — so the delete now names the data tables it
stranded, and resolving one says which workspace is missing rather than reporting a data table
this workspace never had.

`InstanceDatatableRole` derived `Debug` while holding a Postgres password; it is now
hand-written so `{:?}` on the catalog cannot put a live credential in a log line.

Adds the two branches the reviews found unpinned: a caller who is not a member of the
governing workspace at all, and `NoIdentity` — the compatibility path for an agent worker that
predates this and sends no job id.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): unbreak two operator messages and two comments that described other code

The two strings this branch added for states an operator hits once — the catalog write that
matched nothing, and the delete that stranded a pointer — were collapsed from their multi-line
form with the indentation left in, so both rendered with a fourteen-space gap mid-sentence.

`list_datatables` claimed to report a chain it cannot follow and then dropped it; it does drop
it, and the comment now says why that is the right place to stay quiet. The non-superadmin
check in `edit_datatable_config` was introduced as also covering references, which it does not
and need not: `reference` is overwritten from the stored entry for every caller before the
check runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): serialize role catalog mutations, and state each helper's authorization contract

The catalog is one JSON document, so create, rename, enable and delete are all
read-modify-write. Two concurrent creates read the same snapshot, both succeed in the
cluster, and the second write drops the first — leaving a live Postgres login with a password
nobody recorded, which is the exact state the delete path exists to prevent. Every mutation
now runs in one transaction holding an advisory lock across the read, the cluster DDL and the
write, so a lost update cannot happen and a failure rolls the whole thing back. The DDL
helpers take that transaction rather than the pool, which is what makes the lock cover them.

Their statements moved off `sqlx::raw_sql`: the simple protocol is only needed for genuinely
multi-statement SQL, and its future is not `Send`, which an axum handler holding the
transaction requires. Each of these is one statement anyway.

The new cross-crate surface now says what callers must do. `read_role_catalog` returns
plaintext credentials; `create`/`rename`/`set_login`/`drop_instance_role` and
`converge_connect_grants` mutate cluster-wide state; `read_datatable_entry` reads a workspace's
raw config. All of them are superadmin-gated by their current handlers, but nothing said so at
the definition, which is where the next caller looks.

Also: the roles table reloads after a failed login toggle instead of leaving it claiming a flip
that did not land; the rename affordance is the design-system `Button`, not a raw one; and
`resolve_datatable_pg_as_caller` drops a `role` parameter no caller ever filled — browsing
resolves as the data table's default until the database manager grows a picker.

Why role passwords stay a plain `String` while the instance user's password beside them is a
`StringOrSecretRef`, asked three times across reviews: that one is a secret ref because an
operator supplies it and may want it from their own backend, while these are minted here and
never entered by anyone, so there is nothing for a ref to point at. Encrypting generated
secrets at rest is a separate change that would take the replication password with it. Now
said at the field.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): give the role catalog its own row, out of reach of the config machinery

Putting it inside `custom_instance_pg_databases` was the wrong call, and it cost two ways.
The catalog serializes a generated Postgres password per role, and that row is the
operator-facing instance config, so the passwords reached `get_instance_config` and its YAML
editor — a live cluster credential in a response body, a UI field and any log of either.
Worse in the other direction: `to_settings_map` strips the catalog, so a full-row upsert of
that key writes the row back without it and the catalog is gone, while the cluster keeps every
login it described.

`custom_instance_replication_pwd` is the precedent and says exactly why — a generated secret,
written only by the server, never operator-authored, hidden so the config machinery cannot
read, rewrite or drop it. The catalog is the same thing, so it now has the same shape:
`datatable_roles`, in `HIDDEN_SETTINGS`, `PROTECTED_SETTINGS` and the agent-worker denylist.
No redaction to keep in step with three code paths, and no way for a neighbouring write to
take it out.

Two races on the same shared documents. `edit_datatable_config` read the stored data tables
outside its transaction and then wrote the whole `datatable` document, so a permissions save
committing in between was silently rolled back; it now reads under `FOR UPDATE`. And
`set_datatable_permissions` validated role ids against the catalog before opening its
transaction, so a deletion in between let it write a deleted role back — including as the
default, which every later job then fails on; it now holds the catalog lock and the settings
row across validation and write.

Completes the authorization contracts the previous commit claimed but did not finish:
`read_datatable_entry` (which it named and missed), `resolve_governing_datatable`, whose whole
job is to answer for a workspace the caller may not belong to, and
`converge_connect_grants_with`, which had not inherited its wrapper's.

Also the generic Python SDK reference: `_format_py_params` learned the bare `*` last time, but
`extract_py_functions` is a second formatter and still rendered `datatable(name, role)`, so
code written from that page passed a keyword-only argument positionally.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): make the concurrency test pin the handlers, and the contracts describe what is enforced

The concurrency test reimplemented the read-modify-write inline, so deleting the lock from all
three handlers left it green — it pinned Postgres, not the code it was written for. It now
drives `create_datatable_role` twice concurrently and asserts the catalog kept both names.
Checked the way the last one should have been: removing the lock from the handler makes it
fail with "wmtest_a_… is a live cluster login the catalog forgot".

The contracts added last commit were stricter than this PR's own callers, which is worse than
none — the next reader sees a rule already broken and learns to ignore it.
`read_role_catalog` said superadmin-only while two of its four callers are open to any
workspace member, and `converge_connect_grants` said superadmin while
`set_datatable_permissions` reaches it as a workspace admin. Both were fine on substance: the
rule that actually holds is about the credential never reaching a response, log, audit record
or export, not about who may call. They now say that. `read_datatable_entry` gets the same
treatment rather than the one the earlier message claimed for it: it is the primitive every
resolution goes through, so it is deliberately open, and what must not escape is `permissions`
— it names the governing workspace's users, groups and folders.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): close the last ways a role or a pointer can be left pointing at nothing

The raw settings readers hand back whatever is in the row, so moving the catalog into its own
`global_settings` key protected the config machinery and left `GET /settings/global/datatable_roles`
and the settings listing returning every live password. Both now filter that one key. The
neighbouring `custom_instance_replication_pwd` has the same shape and is not touched here: it
predates this and widening the fix to it is a decision about an operator workflow, not a
consequence of this change.

Three ways a save could leave something resolving to nothing:

A permissioned data table could be moved to a PostgreSQL resource. The block was carried across
as a server-owned field, the runtime refuses roles on a resource-backed table, so the save
succeeded and every job afterwards failed. Refused instead — turning roles off first is one step,
and it keeps discarding an access decision something somebody chose.

Renaming a governing data table left every fork pointing at the old name: the data table
disappears from their pickers and their jobs stop, with nothing in the renaming workspace to
suggest why. The rename now follows into the pointers in the same transaction.

Deleting one cannot be followed the same way, so it is reported instead — the response names what
it stranded, the way deleting a workspace does, and the fork's own error already says which
workspace is gone.

Also: `ensure_instance_db_grant_options_unchecked` claimed superadmin while the permissions
handler reaches it as a workspace admin (the same class fixed last commit, one instance missed);
the role entry kept an `instance_config_schema` derive it no longer needs; `write_role_catalog`
was the one writer of that table not stamping `updated_at`; and the concurrency test dropped its
roles only on success — a failing run is exactly the one that creates them without recording them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* refactor(datatables): put the role catalog in its own table, not in global_settings

Five findings across three rounds were all the same choice. A set of live Postgres credentials
was living in `global_settings`, which has generic read, list, write, config-export and CLI
round-trip paths that know nothing about what they carry: the passwords reached the instance
config and its YAML editor, a full-row upsert of a neighbouring key erased the catalog,
`GET /settings/global/{key}` and the settings listing returned them raw, and this round the
redaction that fixed the last two turned `wmill instance push` into something that wipes every
password — a fix breaking the assumption the previous fix made. `POST /settings/global/datatable_roles`
could also empty it outside the lock.

The approved plan offered a table or `global_settings`, so this is the other option it already
allowed rather than a new design. `datatable_role` is a table: no generic settings path can read
it, list it, export it, write it or round-trip it, so none of the five needs a guard. The
redaction, the hidden/protected/agent-denylist entries and the JSON document all go with it.

One row per role also removes the read-modify-write the concurrency work was about: two
concurrent creates are two inserts, and the unique index on `name` is what settles a collision.
The advisory lock stays for the one window rows do not cover — `CREATE ROLE` is invisible to
another transaction until commit, so without it both creates pass their `pg_roles` check.

Also from this round: rename mappings are checked against the configuration they claim to
describe, since fork pointers are rewritten from them — a caller could otherwise submit
`main -> missing` against an unchanged config and repoint every fork of `main` at a name nothing
has, and `A -> B` plus `B -> C` moved what pointed at `A` all the way to `C`. And the warning
naming forks a delete stranded reached the response but not the screen: both the data table
settings save and the workspace delete now show it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): validate a rename against the save it describes, and re-check under the locks

Three from the round, all about deciding on state that could already have moved.

A permission save resolved the data table and checked it was instance-backed before taking any
lock, then wrote under one. A config save committing in between could move the table onto a
PostgreSQL resource — recreating exactly what the transition guard refuses — or rename it, in
which case the write targeted a key that no longer existed and reported success having changed
nothing. It now re-resolves and re-checks on the locked state.

Rename validation checked that the source existed before and the target existed after, which
still accepts `main -> decoy` against a save that keeps both: every fork of `main` then follows
onto a different data table, silently, because it keeps resolving. The rule is now the actual
old-to-new key transition — a source may only survive if another rename took its name, and a
target may only pre-exist if another rename freed it. That also stops two sources sharing one
target, and it admits a swap, which the previous guard refused: `datatables` is keyed by name, so
a swap cannot be done one save at a time, and refusing it was a regression against main. The
pointer cascade now runs in two passes through a temporary name, the way the migration cascade
one layer down already handles the same shape, so `A -> B` with `B -> C` moves each pointer once
from what it named before the save.

The tenant mutators say what they are for: they write an access decision for any workspace named,
with an arbitrary mutation, and exist for the transaction that frees or renames a principal.
Editing a decision on purpose belongs in the permissions endpoint.

Carried in the same change: the stranded-fork list is a field rather than a phrase to grep out of
a success string; the pointer cascade matches with `EXISTS` instead of a `LIKE` over the whole
document, so a workspace whose pointers name something else is not rewritten to a byte-identical
value under an exclusive lock; and `InstanceDatatableRole` drops the serde derives left over from
the JSON document, one of which would emit `pwd`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): cascade on the leave route that is used, gate migrations before the admin connection, and drop a role atomically

The tenant cascade on leaving went onto `/users/leave`. The UI and the generated client call
`/workspaces/leave` — a different handler in a different crate with the same name — which
deleted the membership and left `u/<username>` in the tenant lists. Leaving and rejoining
therefore restored the access the leave was supposed to end, and a later account taking the
username would have inherited it. The regression test drives the route the client actually
calls; without the fix it fails with "leaving kept the tenant".

The migration endpoints authorized too late. `run_datatable_migrations` opened the data table's
admin connection, created `_wm_migrations` and read it before reaching the per-migration role
check — so with nothing pending, nothing was checked at all. Rollback returned before its check
when nothing was applied, and the status endpoint had none. All three now ask, before any
connection is opened, whether the caller can reach the data table as any role at all; which role
a given migration runs as is still decided per migration, and by the executor after that.

Deleting a role committed the cluster drop and the catalog row, then swept the tenant lists in
separate transactions. A sweep failing part-way left workspaces naming a role nothing can connect
as, while the retry answered `NotFound` because the catalog entry was already gone. The sweep now
runs in the same transaction, so the drop, the row and every tenant list commit together.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): refuse to copy a data table that is under roles

pg_dump carries no roles and the import runs with --no-privileges, so a copied
data table arrives owned by the admin connection with no GRANT for any role.
The settings clone brings `permissions` across, so the fork's tenants pass
Windmill's check, connect as the role they were given, and are denied by
Postgres on everything: an entry that reads as configured and answers nothing.

Refuse the copy — in the import endpoint before any data moves, and in the fork
path the CLI takes. Replaying the source's owners and ACLs into the clone is
what lifts this, and is a change of its own. Dropping `permissions` from the
copy instead would be the unsafe half, since the copy holds the parent's rows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): refuse the clone's database too, not only its data

A clone is two endpoints: `create_pg_database` then `import_pg_database`. Only
the second refused a data table under roles, so a fork asking to clone one
created and registered an empty `wm_fork_…` instance database and then failed —
and nothing collects it, since `drop_forked_datatable_databases` only drops
entries carrying `forked_from` and no entry names this one.

Refuse in both, so the clone stops before a database exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* nit worker error msg

* fix pg_dump stuck on version 17 on nix

* fix(datatables): refuse a malformed role annotation instead of ignoring it

`-- Role operator`, `-- role operator;` and `-- role operator -- why` all failed
the annotation parser's exact-match rule, so the query fell through to the data
table's default role and ran, silently, under a login the author did not choose.
Naming a role exists precisely to not do that.

A leading comment whose first word is `role` is now an annotation attempt: the
keyword matches case-insensitively, one trailing `;` is tolerated, and anything
else is an error naming the line. Only callers that already know the target is a
`datatable://` reference ever run this, so ordinary SQL keeps its comments.

Also bumps the dev shell's postgres client to 18 — it trailed the server the dev
database runs, which takes out every data table export, clone and fork-with-data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): refuse a malformed role query string instead of ignoring it

`?Role=analytics`, `?role=` and `?x=1&role=…` all fell through the reference
parser's exact-match rule, so the connection resolved to the data table's default
role and ran under a login the caller never asked for — the URI half of the same
trap as a malformed `-- role` annotation.

The key now matches case-insensitively, and anything else in the query string is
an error naming it; `role` is the only parameter a reference takes. Callers that
only need the entry keep a lenient `datatable_ref_name`, since they never act on
the role. The DuckDB `ATTACH` parser propagates it rather than attaching under
the default.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR

* fix(datatables): carry the role annotation into the row_to_json retry

The retry rebuilds its SQL from `pruneComments(code)`, so the leading comment
block never reached the second attempt — and with it the `-- role <name>` line
that decides which login the query runs as. The retry connected as the data
table's default role instead, so a query the first attempt was denied could
succeed on the second, reported as "recovered with the row_to_json fix".

Carry the leading comment block over. The retry itself is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* chore(datatables): don't mount the roles UI until the ACL editor lands

Enforcement ships first. The permissions drawer is what turns roles on, and the
catalog section is what creates them — both are only useful once there is a way
to grant a role the privileges it needs, which arrives with the ACL editor. Left
mounted they would offer a feature whose other half does not exist.

The two components are complete and reviewed; only their call sites here are
commented out, with a note pointing the follow-up PRs at them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): honour `-- role: x`, and fix the DuckDB attach test

Two review findings, both real.

`attach_datatable_parses_name_and_role` never compiled: `parse_attach_datatable`
returns `Result<Option<_>>` now and one call site kept a single `unwrap`. Its
`?Role=analytics` case also asserted a refusal, contradicting the parser in the
same commit, which matches the key case-insensitively. Replaced with the cases
that are genuinely malformed, and a positive one for the cased key.

`-- role: analytics` fell through to the default role — the silent fallback the
strict parser exists to remove, for the spelling most likely to be typed. The
keyword now accepts an optional colon, attached or spaced, while a word that
merely starts with it (`rolebased`) is still not an attempt.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): clone a fork's pointer instead of failing after the copy

Forking a fork with cloning left an orphan database. The preflight resolves the
pointer and sees the governing entry, so both endpoints ran and filled the new
database; `apply_forked_datatable` then refused the inherited pointer and rolled
the fork back, stranding a registered `wm_fork_*` that no entry names and whose
name blocks the retry.

Refusing earlier would have been the smaller change, but forking a fork and
cloning worked before pointers existed, so it would trade an orphan for a
regression. Resolve what the pointer names and write the terminal entry the
clone needs: the whole `database` object rather than a patch of its
`resource_path`, since a pointer has none, and `reference` removed with it.

Also accepts `-- role=x` and `-- Role = x`, two more spellings that fell through
to the default role.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): refuse to roll back the catalog while roles exist

The down migration dropped the table and left every role behind: live Postgres
logins whose passwords only that table carried, so after a revert Windmill could
neither use, disable nor delete them, and re-applying could not recreate them
because the names were taken. Cleaning up here is not possible either — dropping
a role means reassigning what it owns in every instance database, and a
migration runs in one — so it now refuses while the catalog is non-empty and
says to delete the roles through instance settings, which does the cluster work.

Also enforces the instance-only invariant the resolved-pointer clone relies on
rather than only asserting it in a comment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* refactor(datatables): settle clonability in one place, before anything is created

A clone is three stages a workspace apart — `create_pg_database`, then
`import_pg_database`, then `apply_forked_datatable` inside the fork transaction.
Only the third can roll back, and `CREATE DATABASE` is not transactional, so any
refusal that lives there strands a registered `wm_fork_*` that no entry names
and whose name blocks the retry.

That orphan has now been fixed three times, most recently reintroduced by a
guard added one commit ago. Patching each new refusal into the first endpoint is
not the fix; having two places that can refuse is. `ensure_datatable_is_clonable`
now answers every reason a copy can be refused and returns what it resolved, and
the stage that writes the entry only does the work.

Also takes an ACCESS EXCLUSIVE lock before the rollback guard counts, so a role
created concurrently cannot slip between the check and the drop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): let a retried clone reclaim its own leftover database

A clone creates its target database one request before it copies into it, and
the fork that would name it is written a request after that. Any failure in
between — a pg_dump error, a bad restore, a dropped connection, the source's
roles changing mid-flow — left a registered `wm_fork_*` that no entry names,
and every retry then failed on its name. This predates data table roles.

`create_pg_database` now reclaims such a leftover before creating: only a
`wm_fork_*` database Windmill registered as a data table database and that no
data table or ducklake entry names, in any workspace, archived ones included.
The drop never terminates connections, so a clone still copying into it makes
the reclaim fail instead of being cut off. It is limited to callers who
administer the source — reaching it is not enough, since on a data table
without roles every member reaches it — and anyone else gets the refusal an
existing database always got.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Revert "fix(datatables): let a retried clone reclaim its own leftover database"

This reverts commit 7dd3275a10.

The reclaim tied the caller to the source they administer, but not to the
database it dropped. Between another workspace's import and its final fork
request, that workspace's target is full, registered, unnamed and has no open
connection, so an admin of any instance data table could name it and have it
dropped and recreated empty. The victim's fork would then commit pointing at
the empty copy. Safe reclaim needs durable clone ownership and serialization
with the request that names the database; until then the leftover stays, as it
did before this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(datatables): record the stale clone database as a known limitation

A clone is three requests and `CREATE DATABASE` is not transactional, so a
failure after the first leaves a registered `wm_fork_*` behind, as it did
before data table roles. Accepted for this PR: it is harmless to data and goes
away once the clone is a single server-side operation.

The comment also records why the obvious fix is wrong: reclaiming the leftover
on retry, without durable clone ownership, can drop another workspace's fully
copied database between its import and its final fork request.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): bounce the streams reading a data table when it is deleted

Deleting a governing data table, or the workspace that holds it, only collected
the fork pointers it stranded, for the warning. A Postgres trigger or capture
already streaming through one of those pointers kept the replication connection
it opened while the pointer still resolved, so it went on dispatching the
governing database's rows after the fork lost access — until its connection
happened to restart. The governing workspace's own streams on a deleted entry
did the same.

Both deletion paths now bounce the affected listeners inside their own
transaction, through the helper a permission change already uses, so a
listener that reconnects re-resolves the entry and finds it gone. The helper is
split so a caller can pass the (workspace, local name) pairs it already holds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep the fork schema baseline, and bounce streams on every removal

Three fixes from review.

`edit_datatable_config` took `forked_from` wholesale from the stored entry, so
the fork schema diff's save of an advanced baseline was silently discarded and
an applied change was offered again. Whether an entry carries a clone stamp is
still carried from the store, since that is what marks its database droppable,
but the baseline inside it is now taken from the request.

The stranded-pointer warning and the stream bounce ran over the optional
`deleted_datatables` hint, which the settings-sync CLI never sends, so removing
a governing data table through `wmill` bounced nothing. Removals are now derived
from the stored configuration against the saved one.

`delete_workspace` read the pointers to bounce before its transaction, so a fork
committing a pointer during the deletion was missed. The read now happens inside
the transaction, after the workspace row is deleted: a fork's insert key-share
locks that row through its parent foreign key, so it is either seen or fails on
the missing parent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(datatables): keep Postgres triggers and data table roles apart

A replication stream reads every row of every table whatever the data table's
roles grant, and its listener checks access only when it connects. Rather than
chase every way access can change and bounce the streams each one affects, a
data table now carries one or the other:

- a Postgres trigger or capture cannot be created on, or connect to, a data
  table under roles;
- roles cannot be turned on while an enabled trigger or a live capture reads
  the data table, its own or a fork's through its pointer. The refusal names
  each one to disable.

This removes the stream bounces on roles edits and on data table and workspace
deletion, and the trigger gate that admitted admins. The fork schema baseline
fix from the same review round is kept.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): refuse a Postgres trigger on a data table under roles when it is saved

Creating or editing a trigger that points at a data table under roles was
accepted, and its listener then retried the refused connection every 30
seconds forever. The save is now refused, and a trigger that reaches such a
data table anyway (re-enabled, or cloned into a fork) is disabled by its
listener with the reason, as a missing replication slot is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): disable a data table role before deleting it

Deleting a role reassigns and drops what it owns in each registered
database on its own connection, and each of those passes commits as it
goes. A database failing part-way left the role enabled in the catalog and
able to log in, but already stripped in the databases reached before it.
The role is now disabled in its own commit first, so a failed delete
leaves a disabled role to retry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): serialize roles going on with a stream starting

Turning roles on looked for enabled triggers and live captures once,
without a lock anything starting a stream also took. A trigger enabled in
that window could have its listener connect before roles committed, and a
healthy listener never checks again. Both transitions now serialize on one
advisory lock: roles going on hold it exclusive while they look, and
trigger create, edit and enable, and capture setup and ping hold it shared
while they commit. Either the look sees the stream, or the listener
connects after roles are committed and refuses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): wait out live listeners, and resolve stored names containing `?`

Turning roles on counted a trigger as gone once disabled, and a capture
once its client stopped pinging, but the listener keeps its replication
connection until its next heartbeat notices. A trigger or capture whose
listener pinged in the last 15 seconds, the window a server holds a
listener for, now still counts as streaming.

Data table names could contain `?` before they were restricted, and such
entries are still stored. Splitting `?role=` off a reference misread them:
`a?b` became `a` with an unknown parameter, and the clone checks looked at
a different entry than the one copied. An entry stored under the whole
reference is now looked up first, in the Postgres executor, DuckDB ATTACH
and the clone checks. Agent workers cannot read the workspace and keep
the strict parse, which refuses such a name rather than misreading it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): warn when a settings sync strands fork pointers

A settings save reported the fork pointers left resolving to nothing only
for the names in `deleted_datatables`, which `wmill sync push` never sends.
The save now works out what it removed from the locked entries, and the
CLI prints the stranded pointers it returns.

Also correct the replication helper's contract: no role or admin check
makes a replication connection safe, so a data table under roles is
refused outright rather than gated as an admin operation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): refuse a save that drops a data table's roles through an undeclared rename

A data table's roles follow its entry only through a declared rename. A
settings sync sends the whole map and never declares one, so renaming a
data table under roles there read as a delete and a new entry on the same
database: the new entry carried no roles, and every caller connected as
admin. Such a save is now refused, naming both entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* fix(datatables): no entry without roles may newly reach a database under roles

The previous guard only caught a new name replacing an entry under roles.
A whole-map save could also repoint an existing entry without roles at
that database, or another workspace could point one there, and every
caller of that entry would connect as admin. The rule is now stated on
the saved entries: one that carries no roles and newly points at an
instance database any entry under roles uses, in this workspace or
another, is refused. A declared rename carries its roles and passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move data table role catalog and resolution to the enterprise edition

Roles are an Enterprise Edition feature. The catalog, the Postgres logins,
CONNECT convergence, tenant evaluation and the role half of connection
resolution move to windmill-ee-private. Every public function keeps its path
and signature and forwards through datatable_roles_oss, which re-exports the
enterprise implementation or, without it, refuses.

Without the enterprise edition a data table under roles, or a caller naming a
role, is refused a connection rather than resolved as admin, and the reach and
admin-access checks refuse one under roles. A data table not under roles
resolves as before in every edition, and an instance database keeps the
CONNECT grants it was created with. The catalog lock, the stream lock, the
tenant cascades and the permissions stripping stay in OSS: they only restrict.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move the data table permissions endpoints to the enterprise edition

The permissions read, save and usable-roles handlers move to
windmill-ee-private; the routes stay registered and, without the enterprise
edition, answer that data table roles are an Enterprise Edition feature.
ensure_governs_datatable and ensure_reaches_datatable keep their paths: the
first refuses, the second passes a data table not under roles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): move the data table role catalog endpoints to the enterprise edition

The superadmin list, create, update and delete handlers move to
windmill-ee-private. The routes stay registered and, without the enterprise
edition, refuse after authentication.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* test(datatables): run the roles tests on the enterprise edition, refusals without it

Each test that exercises roles runs with private and enterprise. Two tests run
without them: every roles route answers the Enterprise refusal, and a data
table saved under roles, or a named role, is refused a connection while one
not under roles resolves as before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* feat(datatables): gate the roles UI mount sites on an enterprise license

Both mount sites are still commented out; the gate travels with them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* test(datatables): run the tenant matcher test on the enterprise edition

The matcher it covers is enterprise code now, so without the enterprise
edition the test hit the stub and failed the default windmill-common run. It
runs with private and enterprise, and a counterpart without them asserts that
no tenant list covers anyone, the wildcard and a workspace admin included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb

* chore: update ee-repo-ref to a1873dbb67f2302b85ff5362f8387b48eccdb607

This commit updates the EE repository reference after PR #783 was merged in windmill-ee-private.

Previous ee-repo-ref: 5c853e2c20eca6b748415fc0d6862a6ebfb5fec4

New ee-repo-ref: a1873dbb67f2302b85ff5362f8387b48eccdb607

Automated by sync-ee-ref workflow.

* fix(datatables): refuse roles while a same-workspace alias reaches the database

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): let CE migrations connect as an explicitly named admin

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): serialize roles going on with aliases saved from other workspaces

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(datatables): note that legacy names with ? cannot be migrated

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): add an ACL editor for data table roles

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: take every pooled connection before the ACL apply locks

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: refresh grant options only after the ACL apply validates its plan

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: add only missing grant options before an ACL apply, never default privileges

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: run one data table ACL apply at a time per server before it connects

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: hold the ACL connection to the database that was authorized

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: build the ACL connection from the authorized data table entry

Resolving the settings again could land on a resource with the same
database name on another server, which the later entry checks never see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: check ACL read reach against the entry it connects from

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix(datatables): drop a DuckDB data table secret once its ATTACH has used it

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf(datatables): resolve a workspace's data tables per pointer hop, not per entry

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): hold the parent's settings while a fork points at its data tables

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(datatables): add an ACL editor for data table roles

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: take every pooled connection before the ACL apply locks

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: refresh grant options only after the ACL apply validates its plan

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: add only missing grant options before an ACL apply, never default privileges

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: run one data table ACL apply at a time per server before it connects

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: hold the ACL connection to the database that was authorized

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: build the ACL connection from the authorized data table entry

Resolving the settings again could land on a resource with the same
database name on another server, which the later entry checks never see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* fix: check ACL read reach against the entry it connects from

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BRoYE5ZeAVvrDYdfhDAYXb

* chore: update ee-repo-ref to 7e338e4dabf91689bfd7fb0333c6534040b17b59

This commit updates the EE repository reference after PR #787 was merged in windmill-ee-private.

Previous ee-repo-ref: 0edd40979cf36bfba59323f3f6a0811ae1369cf5

New ee-repo-ref: 7e338e4dabf91689bfd7fb0333c6534040b17b59

Automated by sync-ee-ref workflow.

* feat(datatables): clone a data table under roles with its owners and grants (#11120)

Claude-Session: https://claude.ai/code/session_01UbrtwiYNfayrmqouBJHwGV

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: open the raw app data table drawer when the workspace has none

Selecting the first data table of an empty list passed undefined to the name
check, which threw instead of opening the drawer on no data table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): offer cloning a data table under roles where its grants can be replayed

The server clones such a data table and replays the source's owners and grants,
which only the Enterprise Edition does, so the fork wizard hid both clone
options everywhere instead of on a build that cannot replay them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: name the placeholder the empty raw app data drawer renders

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: pin the enterprise refusal the role pickers read as 'not under roles'

The server's sentence and the frontend's copy of it were coupled by nothing,
so rewording either one turned every role picker on a community build into a
failed lookup.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): read a roles answer only for the workspace it was asked in

A fork and its parent each have their own roles on a data table of the same
name, so an answer stamped with the name alone settled the role from the
workspace the editor was acting on before.

Also derive the AI table creation flag from the data replaced into the editor:
data naming no data table left the flag on from before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): keep an instance database a settings save is waiting to name

Cleanup for a database whose setup failed took the lock first, read no user,
and dropped it while a save blocked on that same lock was about to commit a
reference to it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): check the workspace stamp in the default database selector too

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): read what a save racing instance-database cleanup committed

A transaction blocked on the lock may still roll back, so keeping the database
for it stranded one whose name then blocks every retry: it is let through and
its outcome read instead. The waiter query also matches this database's locks
only, since pg_locks spans the cluster.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): tell a waiting request apart from the workspaces using a database

Both callers render what cleanup returns as the workspaces that keep the
database, so a waiting request's pid read as one of them. Each now words that
case itself, and the give-up comment names where the kept name actually goes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): take the cleanup lock on a connection the pool cannot reclaim

A session lock outlives the future holding it, so a cancellation between
taking it and releasing it handed a locked session back to the pool, where
every later settings save waits on it. Detached, the connection closes when it
is dropped and the server releases the lock with the session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(datatables): close the cleanup connection on drop instead of detaching it

Detaching released the pool permit while the session stayed alive, so
concurrent cleanups waiting on their locks could open as many connections as
they liked. Closing on drop covers the same cancellation and keeps them
counted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref to fd5b8af748f2c985b13e18d9ea30894f3bd7e9a3

This commit updates the EE repository reference after PR #798 was merged in windmill-ee-private.

Previous ee-repo-ref: 3145e422d61d580f0a82804f075285c112879da0

New ee-repo-ref: fd5b8af748f2c985b13e18d9ea30894f3bd7e9a3

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-09-21 12:03:01 +00:00

2035 lines
83 KiB
Rust

/*
* Author: Ruben Fiszel
* Copyright: Windmill Labs, Inc 2022
* This file and its contents are licensed under the AGPLv3 License.
* Please see the included NOTICE for copyright information and
* LICENSE-AGPL for a copy of the license.
*/
//! Ownership and grants on the objects of an instance data table.
//!
//! [`datatable_permissions`](crate::datatable_permissions) decides who may connect as which role;
//! this decides what each role may then touch. Every change is a real `GRANT`, `REVOKE`,
//! `ALTER ... OWNER TO` or `ALTER DEFAULT PRIVILEGES`, so Postgres is what enforces it.
//!
//! Reading is open to anyone who reaches the data table. Planning and applying are for those who
//! administer it — admins of the workspace that governs it, and superadmins. All of it is
//! Enterprise Edition ([`crate::datatable_acl_oss`]).
use std::collections::BTreeMap;
use axum::{
extract::{Extension, Path, Query},
routing::{get, post},
Json, Router,
};
use serde::{Deserialize, Serialize};
use tokio::sync::mpsc;
use tokio_postgres::error::{DbError, SqlState};
use tokio_postgres::AsyncMessage;
use windmill_api_auth::ApiAuthed;
use windmill_audit::audit_oss::audit_log;
use windmill_audit::ActionKind;
use windmill_common::datatable_roles::{
lock_role_catalog, quote_ident, read_role_catalog, read_role_catalog_tx, DatatableRoleCatalog,
ADMIN_DATATABLE_ROLE, CUSTOM_INSTANCE_USER,
};
use windmill_common::error::{pg_error_message, Error, JsonResult, Result};
use windmill_common::workspaces::{resolve_governing_datatable, DataTable, GoverningDatatable};
use windmill_common::{PgDatabase, DB};
use crate::datatable_permissions::{ensure_governs_datatable, ensure_reaches_governing_datatable};
pub(crate) fn routes() -> Router {
Router::new()
.route("/datatable_acl/{datatable_name}", get(get_datatable_acl))
.route(
"/datatable_acl/{datatable_name}/plan",
post(plan_datatable_acl),
)
.route(
"/datatable_acl/{datatable_name}/apply",
post(apply_datatable_acl),
)
}
/// What a read or a change is about.
#[derive(Deserialize, Serialize, Debug, Clone, PartialEq)]
#[serde(tag = "kind", rename_all = "snake_case")]
pub enum AclTarget {
/// The data table's own database — where the privilege to create schemas lives.
Database,
Schema {
schema: String,
},
Table {
schema: String,
table: String,
},
}
impl AclTarget {
/// The schema the target is in, absent for the database itself.
pub(crate) fn schema(&self) -> Option<&str> {
match self {
AclTarget::Database => None,
AclTarget::Schema { schema } => Some(schema),
AclTarget::Table { schema, .. } => Some(schema),
}
}
/// What it is called in a message.
pub(crate) fn label(&self, dbname: &str) -> String {
match self {
AclTarget::Database => dbname.to_string(),
AclTarget::Schema { schema } => schema.clone(),
AclTarget::Table { schema, table } => format!("{schema}.{table}"),
}
}
}
#[derive(Deserialize, Debug)]
pub struct AclTargetQuery {
kind: String,
schema: Option<String>,
table: Option<String>,
}
impl TryFrom<AclTargetQuery> for AclTarget {
type Error = Error;
fn try_from(q: AclTargetQuery) -> Result<Self> {
match (q.kind.as_str(), q.schema, q.table) {
("database", _, _) => Ok(AclTarget::Database),
("schema", Some(schema), _) => Ok(AclTarget::Schema { schema }),
("table", Some(schema), Some(table)) => Ok(AclTarget::Table { schema, table }),
("schema" | "table", None, _) => {
Err(Error::BadRequest("This target needs a schema".to_string()))
}
("table", _, None) => Err(Error::BadRequest(
"A table target needs a table".to_string(),
)),
(kind, _, _) => Err(Error::BadRequest(format!("Unknown ACL target '{kind}'"))),
}
}
}
/// Where a set of privileges applies, relative to the target.
///
/// `Future*` covers what does not exist yet: those become `ALTER DEFAULT PRIVILEGES`, which only
/// binds objects created by the roles it names.
#[derive(Deserialize, Serialize, Debug, Clone, Copy, PartialEq)]
#[serde(rename_all = "snake_case")]
pub enum GrantScope {
/// The target itself — the database, the schema, or the table.
Target,
AllTables,
AllSequences,
AllFunctions,
FutureTables,
FutureSequences,
FutureFunctions,
}
impl GrantScope {
pub(crate) fn is_future(&self) -> bool {
matches!(
self,
GrantScope::FutureTables | GrantScope::FutureSequences | GrantScope::FutureFunctions
)
}
}
/// A change to plan. One at a time: each is confirmed against its own SQL.
#[derive(Deserialize, Serialize, Debug, Clone)]
#[serde(tag = "type", rename_all = "snake_case")]
pub enum AclChange {
/// Hand the target — and, for a schema, everything already in it but an extension's members,
/// which stay with the extension — to another role.
SetOwner {
role: String,
},
Grant {
role: String,
privileges: Vec<String>,
scope: GrantScope,
},
Revoke {
role: String,
privileges: Vec<String>,
scope: GrantScope,
/// Objects inside the target, empty for the target itself. `ON ALL TABLES` grants read
/// back per object, so they are revoked per object — and the same privileges on several
/// of them are revoked together.
#[serde(default)]
objects: Vec<AclObject>,
},
}
impl AclChange {
/// The role the change is about, as the editor names it.
fn role(&self) -> &str {
match self {
AclChange::SetOwner { role }
| AclChange::Grant { role, .. }
| AclChange::Revoke { role, .. } => role,
}
}
}
#[derive(Deserialize, Debug)]
pub struct AclChangeRequest {
pub target: AclTarget,
pub change: AclChange,
/// The statements the plan showed. An apply runs only those: it plans again and refuses if the
/// result differs.
#[serde(default)]
pub statements: Option<Vec<String>>,
}
/// An object inside a schema, named the way `REVOKE ... ON <keyword>` needs it.
#[derive(Deserialize, Serialize, Debug, Clone, PartialEq)]
pub struct AclObject {
pub name: String,
/// `TABLE`, `SEQUENCE`, `FUNCTION`, `PROCEDURE` or `TYPE`: what the object is.
/// [`object_keyword`] turns it into the keyword a revoke takes; a type has none.
pub kind: String,
/// A routine is identified by its argument types, not by its name: two `f` in one schema are
/// two objects. Absent for everything else.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub args: Option<String>,
}
/// A grant as the database has it, under the role names the editor uses.
#[derive(Serialize, Debug, PartialEq)]
pub struct AclGrant {
/// A data table role's name, `admin` for `custom_instance_user`, else the raw Postgres role
/// (`PUBLIC` included).
pub grantee: String,
pub privileges: Vec<String>,
/// `None` for the target itself, else the object inside it.
#[serde(skip_serializing_if = "Option::is_none")]
pub object: Option<AclObject>,
/// `TABLES` / `SEQUENCES` / `FUNCTIONS` / `TYPES` (or, on the database, `SCHEMAS`) when this
/// is a default privilege, which applies to objects that do not exist yet.
#[serde(skip_serializing_if = "Option::is_none")]
pub future: Option<String>,
/// Where the grant comes from, each role once: who granted it, or for a default privilege the
/// role whose future objects it covers. A revoke of some of the privileges takes them back from
/// every source that gave them.
pub sources: Vec<AclSource>,
}
/// One role a grant comes from.
#[derive(Serialize, Debug, PartialEq)]
pub struct AclSource {
/// Under the same names as [`AclGrant::grantee`].
pub role: String,
/// What `role` gave of the grant's privileges. A revoke is held back only by a source out of
/// reach that gave some of what it takes back.
pub privileges: Vec<String>,
/// Whether this data table's connection can take back what `role` gave: on an object only the
/// owner's grants when it acts for the owner, or else its own (`grant_source!`); for a default
/// privilege, a creating role it acts for. What a source out of reach gave is not revocable
/// from here; privileges only other sources gave still are.
pub reachable: bool,
}
#[derive(Serialize, Debug)]
pub struct DatatableAclInfo {
/// Under the same names as [`AclGrant::grantee`].
pub owner: String,
/// The roles a change may name: `admin`, then every role of the instance catalog. Only for a
/// caller who may change anything, as the catalog is in the permissions drawer.
pub roles: Vec<String>,
/// Whether this caller may plan and apply changes: they administer the data table, on an
/// edition that has the planner.
pub editable: bool,
/// Whether the data table is a clone, whose grants stay as they were copied from its source.
pub clone: bool,
/// Whether the server is Postgres 17 or later, which added the `MAINTAIN` table privilege.
pub supports_maintain: bool,
/// The database the target lives in, which no target carries itself.
pub dbname: String,
pub grants: Vec<AclGrant>,
/// What the target holds that is a target of its own: a database's schemas, a schema's tables.
pub children: Vec<String>,
}
#[derive(Serialize, Debug)]
pub struct AclPlan {
pub statements: Vec<String>,
pub warnings: Vec<String>,
}
/// The Postgres role a role name stands for. A data table role is a login named exactly like the
/// role, so this is the identity — except `admin`, which is `custom_instance_user`.
///
/// Anything else is refused, never resolved to some default: every statement a plan writes names
/// the role it is about.
pub(crate) fn pg_role_of(name: &str, catalog: &DatatableRoleCatalog) -> Result<String> {
if name == ADMIN_DATATABLE_ROLE {
return Ok(CUSTOM_INSTANCE_USER.to_string());
}
if catalog.values().any(|r| r.name == name) {
return Ok(name.to_string());
}
Err(Error::BadRequest(format!(
"'{name}' is not a data table role of this instance"
)))
}
/// The reverse of [`pg_role_of`], for display. A role that is not a data table role reads back as
/// itself.
pub(crate) fn role_name_of(pg_role: &str) -> String {
if pg_role == CUSTOM_INSTANCE_USER {
ADMIN_DATATABLE_ROLE.to_string()
} else {
pg_role.to_string()
}
}
/// Every role a change may name: `admin` first, then the catalog.
fn role_names(catalog: &DatatableRoleCatalog) -> Vec<String> {
let mut names: Vec<String> = catalog.values().map(|r| r.name.clone()).collect();
names.sort();
names.insert(0, ADMIN_DATATABLE_ROLE.to_string());
names
}
fn ensure_instance(governing: &GoverningDatatable) -> Result<()> {
if governing.is_instance() {
return Ok(());
}
Err(Error::BadRequest(format!(
"Data table '{}' is backed by a Postgres resource, so its access is managed on that \
server directly. Only a data table on the Windmill instance's own database has data \
table roles to grant to.",
governing.name
)))
}
/// The data table's `admin` connection, and the notices Postgres sends on it.
///
/// Authorization: connects as `custom_instance_user` with the instance's own credentials and checks
/// nothing. Callers MUST have authorized the request first — a request about to be refused must
/// not get as far as this connection.
async fn connect_as_admin_unchecked(
db: &DB,
governing: &GoverningDatatable,
) -> Result<(
tokio_postgres::Client,
mpsc::UnboundedReceiver<DbError>,
String,
)> {
ensure_instance(governing)?;
// Built from the authorized entry, never by resolving the settings again: a save in between
// could point the entry at a resource on another server and back, and this connection would
// then alter a database the later checks of the entry never see.
let mut pg = PgDatabase::parse_uri(&windmill_common::get_database_url().await?.as_str().await)?;
pg.dbname = governing
.datatable
.database
.as_ref()
.expect("a governing entry owns a database")
.resource_path
.clone();
pg.user = Some(CUSTOM_INSTANCE_USER.to_string());
pg.password = Some(windmill_common::utils::get_custom_pg_instance_password(db).await?);
let dbname = pg.dbname.clone();
let (client, notices) = connect_with_notices(db, &pg).await?;
Ok((client, notices, dbname))
}
/// A connection to `pg`, and the notices Postgres sends on it — which is where a grant or revoke
/// that changed nothing is reported ([`execute_acl_statements`]).
pub(crate) async fn connect_with_notices(
db: &DB,
pg: &PgDatabase,
) -> Result<(tokio_postgres::Client, mpsc::UnboundedReceiver<DbError>)> {
let (client, connection) = pg.connect(Some(db)).await?;
Ok((
client,
drive_with_notices(connection, windmill_common::TokioPgConnection::poll_message),
))
}
type PollMessage<C> =
fn(
&mut C,
&mut std::task::Context<'_>,
) -> std::task::Poll<Option<std::result::Result<AsyncMessage, tokio_postgres::Error>>>;
/// Drive `connection` in the background with `poll`, forwarding its notices.
pub(crate) fn drive_with_notices<C: Send + 'static>(
mut connection: C,
poll: PollMessage<C>,
) -> mpsc::UnboundedReceiver<DbError> {
// Unbounded: the driver must never wait on the receiver, which only drains once the statement
// the driver is carrying has completed.
let (notices_tx, notices) = mpsc::unbounded_channel();
tokio::spawn(async move {
loop {
match std::future::poll_fn(|cx| poll(&mut connection, cx)).await {
Some(Ok(AsyncMessage::Notice(notice))) => {
let _ = notices_tx.send(notice);
}
Some(Ok(_)) => {}
Some(Err(e)) => {
tracing::error!("Datatable ACL connection error: {e}");
break;
}
None => break,
}
}
});
notices
}
/// Run `statements` in order on `tx`, failing on the first that errors or that Postgres only warns
/// about. A privilege the connection cannot pass on is a warning to Postgres (`01007` / `01006`),
/// which then carries on having changed nothing; returning drops the transaction, rolling back
/// everything before it.
///
/// Authorization: none. Runs `statements` as the connection `tx` is on; callers MUST have authorized
/// changing that database's access, and built the statements themselves.
pub(crate) async fn execute_acl_statements(
tx: &tokio_postgres::Transaction<'_>,
notices: &mut mpsc::UnboundedReceiver<DbError>,
statements: &[String],
) -> Result<()> {
while notices.try_recv().is_ok() {}
for statement in statements {
tx.batch_execute(statement).await.map_err(|e| {
Error::ExecutionErr(format!(
"Failed to run `{statement}`: {}",
pg_error_message(&e)
))
})?;
while let Ok(notice) = notices.try_recv() {
if *notice.code() == SqlState::WARNING_PRIVILEGE_NOT_GRANTED
|| *notice.code() == SqlState::WARNING_PRIVILEGE_NOT_REVOKED
{
return Err(Error::ExecutionErr(format!(
"`{statement}` did not take effect ({}), so nothing was applied",
notice.message()
)));
}
}
}
Ok(())
}
/// An object whose ownership follows the schema's.
#[derive(Debug, PartialEq)]
pub(crate) struct OwnedObject {
/// The keyword `ALTER ... OWNER TO` takes for this kind of object.
pub(crate) keyword: &'static str,
/// How Postgres names the object (`pg_identify_object`): schema-qualified, quoted where
/// needed, with a routine's arguments or an operator class's access method. It goes into the
/// statement as it is.
pub(crate) identity: String,
}
/// Everything a schema's change of owner takes along, in schema `$1`, as (kind, identity, owner
/// oid), the first two as `pg_identify_object` gives them. The plan, the check of what this
/// connection may move and the check after the move all read this one list, so what is moved and
/// what is checked cannot differ.
///
/// Read from what depends on the schema rather than catalog by catalog, so no kind of object is
/// left out by omission; array types, row types and indexes depend on another object instead.
/// Left out: an object's internal parts (a range type's constructors and multirange, an identity
/// column's sequence), a sequence tied to a column (it follows its table, and refuses an owner of
/// its own), an extension's members, and what has no `ALTER ... OWNER` at all (an extension, a text
/// search parser or template). Postgres records no owner for the bootstrap superuser, oid 10.
macro_rules! schema_owned_objects {
() => {
"SELECT o.type AS kind, o.identity, COALESCE(s.refobjid, 10::oid) AS owner
FROM pg_depend d
CROSS JOIN LATERAL pg_identify_object(d.classid, d.objid, 0) o
LEFT JOIN pg_shdepend s
ON s.dbid = (SELECT oid FROM pg_database WHERE datname = current_database())
AND s.classid = d.classid AND s.objid = d.objid AND s.deptype = 'o'
WHERE d.refclassid = 'pg_namespace'::regclass AND d.deptype = 'n'
AND d.refobjid = (SELECT oid FROM pg_namespace WHERE nspname = $1)
AND d.classid <> ALL(ARRAY['pg_extension'::regclass, 'pg_ts_parser'::regclass,
'pg_ts_template'::regclass]::oid[])
AND NOT EXISTS (
SELECT 1 FROM pg_depend x
WHERE x.classid = d.classid AND x.objid = d.objid AND x.objsubid = 0
AND (x.deptype IN ('i', 'e')
OR (x.deptype = 'a' AND d.classid = 'pg_class'::regclass
AND x.refobjsubid <> 0)))"
};
}
#[allow(unused_imports)]
pub(crate) use schema_owned_objects;
/// The keyword `ALTER ... OWNER TO` takes for a kind of object, as `pg_identify_object` names the
/// kind. A kind missing here is refused rather than skipped, which would leave it behind.
pub(crate) fn owned_keyword(kind: &str) -> Option<&'static str> {
Some(match kind {
"table" => "TABLE",
"view" => "VIEW",
"materialized view" => "MATERIALIZED VIEW",
"sequence" => "SEQUENCE",
"foreign table" => "FOREIGN TABLE",
"type" => "TYPE",
"function" | "procedure" | "aggregate" => "ROUTINE",
"collation" => "COLLATION",
"conversion" => "CONVERSION",
"operator" => "OPERATOR",
"operator class" => "OPERATOR CLASS",
"operator family" => "OPERATOR FAMILY",
"statistics object" => "STATISTICS",
"text search dictionary" => "TEXT SEARCH DICTIONARY",
"text search configuration" => "TEXT SEARCH CONFIGURATION",
_ => return None,
})
}
async fn read_owned_objects(
client: &tokio_postgres::Client,
schema: &str,
) -> Result<Vec<OwnedObject>> {
let rows = client
.query(
concat!(
"SELECT kind, identity FROM (",
schema_owned_objects!(),
") o ORDER BY kind, identity"
),
&[&schema],
)
.await
.map_err(|e| {
Error::internal_err(format!(
"Failed to list the objects of schema '{schema}': {}",
pg_error_message(&e)
))
})?;
rows.into_iter()
.map(|row| {
let kind: &str = row.get(0);
let identity: String = row.get(1);
match owned_keyword(kind) {
Some(keyword) => Ok(OwnedObject { keyword, identity }),
None => Err(Error::BadRequest(format!(
"{identity} is a {kind}, whose owner cannot be changed from here, and schema \
{schema} would change hands without it. Move it to another schema first."
))),
}
})
.collect()
}
/// Default privileges set database-wide (no `IN SCHEMA`), as `(defaclrole, defaclobjtype, grantee,
/// privilege_type)`. Postgres stores such an entry as a whole acl, the creator's own privileges
/// included, so only what goes beyond the built-in default is a grant. It applies in every schema
/// on top of the schema's own defaults, which cannot take it back.
macro_rules! database_wide_defaults {
() => {
"SELECT d.defaclrole, d.defaclobjtype, a.grantee, a.privilege_type
FROM pg_default_acl d, aclexplode(d.defaclacl) a
WHERE d.defaclnamespace = 0
AND NOT EXISTS (
SELECT 1 FROM aclexplode(acldefault(
CASE d.defaclobjtype WHEN 'S' THEN 's' ELSE d.defaclobjtype END, d.defaclrole)) x
WHERE x.grantee = a.grantee AND x.privilege_type = a.privilege_type)"
};
}
/// `TABLES`, `SEQUENCES`, `FUNCTIONS` or `TYPES`, as a `defaclobjtype` names them.
fn default_objects_keyword(objtype: &str) -> &'static str {
match objtype {
"r" => "TABLES",
"S" => "SEQUENCES",
"T" => "TYPES",
_ => "FUNCTIONS",
}
}
/// What a schema's owner holds on what gets created there later. A change of owner hands the new
/// owner the same and takes these back: otherwise every former owner keeps reaching whatever the
/// other roles create there.
#[derive(Debug, PartialEq)]
pub(crate) struct FormerOwnerDefaults {
pub(crate) pg_role: String,
/// (creating role, `TABLES` / `SEQUENCES` / `FUNCTIONS` / `TYPES`), one per default privilege
/// it holds.
pub(crate) defaults: Vec<(String, &'static str)>,
}
/// The default privileges schema `schema`'s owner holds there — `None` when it holds none, or is
/// `new_owner` already. Refused when one was set by a role this connection cannot act for: only
/// a member of the creating role may change its defaults, so the revoke would fail at apply. Also
/// refused when the owner holds defaults set database-wide, which no change to this schema takes
/// back.
async fn read_former_owner_defaults(
client: &tokio_postgres::Client,
schema: &str,
new_owner: &str,
) -> Result<Option<FormerOwnerDefaults>> {
let database_wide = client
.query_opt(
concat!(
"SELECT pg_get_userbyid(n.nspowner), pg_get_userbyid(g.defaclrole),
g.defaclobjtype::text
FROM pg_namespace n, (",
database_wide_defaults!(),
") g
WHERE n.nspname = $1 AND g.grantee = n.nspowner
AND g.defaclobjtype IN ('r', 'S', 'f', 'T')
AND n.nspowner <> (SELECT oid FROM pg_roles WHERE rolname = $2)
ORDER BY 2, 3
LIMIT 1"
),
&[&schema, &new_owner],
)
.await
.map_err(|e| {
Error::internal_err(format!(
"Failed to read the database-wide default privileges: {}",
pg_error_message(&e)
))
})?;
if let Some(row) = database_wide {
let pg_role: String = row.get(0);
let creator: String = row.get(1);
return Err(Error::BadRequest(format!(
"{} holds default privileges set database-wide by {}, which no default of schema \
{schema} takes back, so it would keep reaching what is created here after the change \
of owner. Revoke them database-wide first: ALTER DEFAULT PRIVILEGES FOR ROLE {} \
REVOKE ALL PRIVILEGES ON {} FROM {}",
role_name_of(&pg_role),
role_name_of(&creator),
quote_ident(&creator),
default_objects_keyword(row.get(2)),
quote_ident(&pg_role)
)));
}
let rows = client
.query(
"SELECT DISTINCT pg_get_userbyid(n.nspowner), pg_get_userbyid(d.defaclrole),
d.defaclobjtype::text, pg_has_role(d.defaclrole, 'USAGE')
FROM pg_namespace n
JOIN pg_default_acl d ON d.defaclnamespace = n.oid
CROSS JOIN LATERAL aclexplode(d.defaclacl) a
WHERE n.nspname = $1 AND a.grantee = n.nspowner
-- What the owner gives itself is about the objects it creates, not about owning
-- the schema, so a move leaves it alone.
AND a.grantee <> d.defaclrole
AND d.defaclobjtype IN ('r', 'S', 'f', 'T')
AND n.nspowner <> (SELECT oid FROM pg_roles WHERE rolname = $2)
ORDER BY 2, 3",
&[&schema, &new_owner],
)
.await
.map_err(|e| {
Error::internal_err(format!(
"Failed to read the default privileges of schema '{schema}': {}",
pg_error_message(&e)
))
})?;
let Some(first) = rows.first() else {
return Ok(None);
};
let pg_role: String = first.get(0);
let plural = |row: &tokio_postgres::Row| default_objects_keyword(row.get(2));
if let Some(row) = rows.iter().find(|row| !row.get::<_, bool>(3)) {
let creator: String = row.get(1);
return Err(Error::BadRequest(format!(
"{} holds default privileges in schema {schema} from {creator}, which this data \
table's connection cannot act for, so they would outlive the change of owner. \
Revoke them as {creator} first: ALTER DEFAULT PRIVILEGES FOR ROLE {} IN SCHEMA {} \
REVOKE ALL PRIVILEGES ON {} FROM {}",
role_name_of(&pg_role),
quote_ident(&creator),
quote_ident(schema),
plural(row),
quote_ident(&pg_role)
)));
}
Ok(Some(FormerOwnerDefaults {
defaults: rows
.iter()
.map(|row| (row.get::<_, String>(1), plural(row)))
.collect(),
pg_role,
}))
}
/// What the catalog holds that a plan depends on, read before planning so the planner stays pure.
#[derive(Debug, Default, PartialEq)]
pub(crate) struct CatalogFacts {
/// Every role that may create objects here, but the one the change is about.
pub(crate) other_pg_roles: Vec<String>,
/// For a schema's change of owner: what moves along with it.
pub(crate) existing_objects: Vec<OwnedObject>,
/// For a schema's change of owner: the defaults it takes back from the owner it replaces.
pub(crate) former_owner: Option<FormerOwnerDefaults>,
/// For a revoke: each grant it takes back.
pub(crate) revoked_grants: Vec<RevokedGrant>,
}
/// A grant a revoke takes back, as the catalog records it: what `source` gave on `object` (the
/// target itself when `None`), or for a default privilege on what `source` creates later.
#[derive(Debug, PartialEq)]
pub(crate) struct RevokedGrant {
pub(crate) object: Option<AclObject>,
/// The role that made the grant: its grantor, or the creating role of a default privilege.
pub(crate) source: String,
pub(crate) privileges: Vec<String>,
}
/// An `aclexplode` row's source and whether this connection can take back what it gave, on an
/// object owned by `$owner`. A REVOKE speaks for the grantor Postgres picks itself — the owner,
/// when the connection acts for the owner, and otherwise the connection — and `GRANTED BY` names
/// nobody else, so a grant any other role made stays whatever the connection runs.
macro_rules! grant_source {
($owner:literal) => {
concat!(
"pg_get_userbyid(a.grantor), a.grantor = CASE WHEN pg_has_role(",
$owner,
", 'USAGE') THEN ",
$owner,
" ELSE (SELECT oid FROM pg_roles WHERE rolname = current_user) END"
)
};
}
/// What relation `$2` of schema `$1` grants role `$3`, as (source, whether this connection can
/// take it back, privilege).
const RELATION_GRANTS: &str = concat!(
"SELECT ",
grant_source!("c.relowner"),
", a.privilege_type
FROM pg_class c
JOIN pg_namespace n ON n.oid = c.relnamespace,
aclexplode(COALESCE(c.relacl, acldefault(
CASE c.relkind WHEN 'S' THEN 's' ELSE 'r' END::\"char\", c.relowner))) a
WHERE n.nspname = $1 AND c.relname = $2
AND c.relkind = ANY(ARRAY['r','p','v','m','S','f']::\"char\"[])
AND a.grantee = (SELECT oid FROM pg_roles WHERE rolname = $3)"
);
/// The grants a revoke takes back, read from the catalog rather than from the request: one per
/// object and source that gave `pg_role` any of `privileges`. Refused when a source's grant is out
/// of this connection's reach (`grant_source!`, or for a default privilege a creating role it does
/// not act for): the revoke would leave that grant in place.
async fn read_revoked_grants(
client: &tokio_postgres::Client,
dbname: &str,
target: &AclTarget,
scope: GrantScope,
objects: &[AclObject],
privileges: &[String],
pg_role: &str,
) -> Result<Vec<RevokedGrant>> {
let read_error = |e: tokio_postgres::Error| {
Error::internal_err(format!(
"Failed to read what the revoke takes back: {}",
pg_error_message(&e)
))
};
let wanted: Vec<String> = privileges.iter().map(|p| p.to_uppercase()).collect();
// (object, how it reads in a refusal, one row per source and privilege)
let mut read = Vec::new();
match (scope, target) {
(scope, AclTarget::Schema { schema }) if scope.is_future() => {
let (objtype, plural) = match scope {
GrantScope::FutureTables => ("r", "tables"),
GrantScope::FutureSequences => ("S", "sequences"),
_ => ("f", "functions"),
};
let database_wide = client
.query(
concat!(
"SELECT pg_get_userbyid(g.defaclrole), g.privilege_type FROM (",
database_wide_defaults!(),
") g
WHERE g.defaclobjtype::text = $1
AND g.grantee = (SELECT oid FROM pg_roles WHERE rolname = $2)
ORDER BY 1, 2"
),
&[&objtype, &pg_role],
)
.await
.map_err(read_error)?;
let mut still_granted: BTreeMap<String, Vec<String>> = BTreeMap::new();
for row in database_wide {
let privilege: String = row.get(1);
if wanted.contains(&privilege) {
still_granted.entry(row.get(0)).or_default().push(privilege);
}
}
if let Some((creator, taken)) = still_granted.into_iter().next() {
return Err(Error::BadRequest(format!(
"{} also receives {} on {plural} {} creates through a default privilege set \
database-wide, which no default of schema {schema} takes back. Revoke it \
database-wide first: ALTER DEFAULT PRIVILEGES FOR ROLE {} REVOKE {} ON {} \
FROM {}",
role_name_of(pg_role),
taken.join(", "),
role_name_of(&creator),
quote_ident(&creator),
taken.join(", "),
default_objects_keyword(objtype),
quote_ident(pg_role)
)));
}
let rows = client
.query(
"SELECT pg_get_userbyid(d.defaclrole), pg_has_role(d.defaclrole, 'USAGE'),
a.privilege_type
FROM pg_default_acl d
JOIN pg_namespace n ON n.oid = d.defaclnamespace,
aclexplode(d.defaclacl) a
WHERE n.nspname = $1 AND d.defaclobjtype::text = $2
AND a.grantee = (SELECT oid FROM pg_roles WHERE rolname = $3)",
&[schema, &objtype, &pg_role],
)
.await
.map_err(read_error)?;
read.push((
None,
format!("{plural} created later in schema {schema}"),
rows,
));
}
(GrantScope::Target, _) if objects.is_empty() => {
let rows = match target {
AclTarget::Database => {
client
.query(
concat!(
"SELECT ",
grant_source!("d.datdba"),
", a.privilege_type
FROM pg_database d,
aclexplode(COALESCE(d.datacl, acldefault('d', d.datdba))) a
WHERE d.datname = current_database()
AND a.grantee = (SELECT oid FROM pg_roles WHERE rolname = $1)"
),
&[&pg_role],
)
.await
}
AclTarget::Schema { schema } => {
client
.query(
concat!(
"SELECT ",
grant_source!("n.nspowner"),
", a.privilege_type
FROM pg_namespace n,
aclexplode(COALESCE(n.nspacl, acldefault('n', n.nspowner))) a
WHERE n.nspname = $1
AND a.grantee = (SELECT oid FROM pg_roles WHERE rolname = $2)"
),
&[schema, &pg_role],
)
.await
}
AclTarget::Table { schema, table } => {
client
.query(RELATION_GRANTS, &[schema, table, &pg_role])
.await
}
}
.map_err(read_error)?;
read.push((None, target.label(dbname), rows));
}
(GrantScope::Target, AclTarget::Schema { schema }) => {
for object in objects {
let rows = match object_keyword(&object.kind)? {
"ROUTINE" => {
client
.query(
concat!(
"SELECT ",
grant_source!("p.proowner"),
", a.privilege_type
FROM pg_proc p
JOIN pg_namespace n ON n.oid = p.pronamespace,
aclexplode(COALESCE(p.proacl, acldefault('f', p.proowner))) a
WHERE n.nspname = $1 AND p.proname = $2
AND pg_get_function_identity_arguments(p.oid) = $3
AND a.grantee = (SELECT oid FROM pg_roles WHERE rolname = $4)"
),
&[
schema,
&object.name,
&object.args.as_deref().unwrap_or(""),
&pg_role,
],
)
.await
}
_ => {
client
.query(RELATION_GRANTS, &[schema, &object.name, &pg_role])
.await
}
}
.map_err(read_error)?;
read.push((
Some(object.clone()),
format!("{} {schema}.{}", object.kind.to_lowercase(), object.name),
rows,
));
}
}
// Every other scope and target is the planner's to refuse.
_ => {}
}
let mut revoked = Vec::new();
for (object, label, rows) in read {
let mut by_source: BTreeMap<String, (bool, Vec<String>)> = BTreeMap::new();
for row in rows {
let privilege: String = row.get(2);
if wanted.contains(&privilege) {
by_source
.entry(row.get(0))
.or_insert_with(|| (row.get(1), vec![]))
.1
.push(privilege);
}
}
for (source, (reachable, mut privileges)) in by_source {
if !reachable {
return Err(Error::BadRequest(format!(
"{} on {label} was granted to {} by {source}, and Postgres takes a grant \
back only through the role that made it, which this data table's \
connection cannot speak for here. Revoke it as {source}.",
privileges.join(", "),
role_name_of(pg_role),
)));
}
privileges.sort();
privileges.dedup();
revoked.push(RevokedGrant { object: object.clone(), source, privileges });
}
}
Ok(revoked)
}
/// The keyword a `REVOKE ... ON` takes for one object, checked rather than interpolated: it lands
/// in SQL unquoted.
pub(crate) fn object_keyword(kind: &str) -> Result<&'static str> {
match kind.to_uppercase().as_str() {
"TABLE" | "VIEW" | "MATERIALIZED VIEW" | "FOREIGN TABLE" => Ok("TABLE"),
"SEQUENCE" => Ok("SEQUENCE"),
// `FUNCTION` names no procedure; `ROUTINE` names either.
"FUNCTION" | "PROCEDURE" | "ROUTINE" => Ok("ROUTINE"),
other => Err(Error::BadRequest(format!("Unknown object kind '{other}'"))),
}
}
/// Every object of a schema, named the way the catalog names it.
async fn read_schema_objects(
client: &tokio_postgres::Client,
schema: &str,
) -> Result<Vec<AclObject>> {
let rows = client
.query(
"SELECT CASE c.relkind WHEN 'S' THEN 'SEQUENCE' ELSE 'TABLE' END, c.relname, NULL::text
FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE n.nspname = $1 AND c.relkind = ANY(ARRAY['r','p','v','m','S','f']::\"char\"[])
UNION ALL
SELECT CASE p.prokind WHEN 'p' THEN 'PROCEDURE' ELSE 'FUNCTION' END, p.proname,
pg_get_function_identity_arguments(p.oid)
FROM pg_proc p JOIN pg_namespace n ON n.oid = p.pronamespace
WHERE n.nspname = $1",
&[&schema],
)
.await
.map_err(|e| {
Error::internal_err(format!(
"Failed to list the objects of schema '{schema}': {}",
pg_error_message(&e)
))
})?;
Ok(rows
.into_iter()
.map(|row| AclObject { kind: row.get(0), name: row.get(1), args: row.get(2) })
.collect())
}
/// Replace the objects a revoke names with the catalog's own entry for each.
///
/// A routine is identified by its argument types, and those go into the statement as written —
/// there is no quoting for them — so the request may name an object but never spell one: what
/// reaches the SQL is read back from Postgres. An object that resolves to nothing is refused rather
/// than dropped, since a revoke that silently covers less than it says is worse than an error.
async fn resolve_acl_objects(
client: &tokio_postgres::Client,
target: &AclTarget,
objects: &[AclObject],
) -> Result<Vec<AclObject>> {
if objects.is_empty() {
return Ok(vec![]);
}
let Some(schema) = target.schema() else {
return Err(Error::BadRequest(
"A database has no objects of its own to revoke on".to_string(),
));
};
let known = read_schema_objects(client, schema).await?;
objects
.iter()
.map(|requested| {
let keyword = object_keyword(&requested.kind)?;
known
.iter()
.find(|k| {
k.name == requested.name
&& k.args == requested.args
&& object_keyword(&k.kind).is_ok_and(|k| k == keyword)
})
.cloned()
.ok_or_else(|| {
Error::NotFound(format!(
"'{}' is not an object of schema '{schema}'",
requested.name
))
})
})
.collect()
}
async fn read_owner(client: &tokio_postgres::Client, target: &AclTarget) -> Result<Option<String>> {
let row = match target {
AclTarget::Database => {
client
.query_opt(
"SELECT pg_get_userbyid(datdba) FROM pg_database WHERE datname = current_database()",
&[],
)
.await
}
AclTarget::Schema { schema } => {
client
.query_opt(
// `public` is owned by `pg_database_owner`, a placeholder role whose membership
// is whoever owns the database — naming it back would say nothing, so resolve
// it to that owner.
"SELECT pg_get_userbyid(owner) FROM (
SELECT CASE WHEN n.nspowner = (SELECT oid FROM pg_roles WHERE rolname = 'pg_database_owner')
THEN (SELECT d.datdba FROM pg_database d WHERE d.datname = current_database())
ELSE n.nspowner END AS owner
FROM pg_namespace n WHERE n.nspname = $1
) o",
&[schema],
)
.await
}
AclTarget::Table { schema, table } => {
client
.query_opt(
"SELECT pg_get_userbyid(c.relowner)
FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE n.nspname = $1 AND c.relname = $2",
&[schema, table],
)
.await
}
}
.map_err(|e| Error::internal_err(format!("Failed to read the owner: {}", pg_error_message(&e))))?;
Ok(row.map(|row| row.get(0)))
}
async fn read_children(client: &tokio_postgres::Client, target: &AclTarget) -> Result<Vec<String>> {
let rows = match target {
AclTarget::Database => {
client
.query(
"SELECT nspname::text FROM pg_namespace
WHERE nspname <> 'information_schema' AND nspname NOT LIKE 'pg\\_%'
ORDER BY nspname",
&[],
)
.await
}
AclTarget::Schema { schema } => {
client
.query(
"SELECT c.relname::text
FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE n.nspname = $1 AND c.relkind = ANY(ARRAY['r','p']::\"char\"[])
ORDER BY c.relname",
&[schema],
)
.await
}
AclTarget::Table { .. } => return Ok(vec![]),
}
.map_err(|e| {
Error::internal_err(format!(
"Failed to list what the target holds: {}",
pg_error_message(&e)
))
})?;
Ok(rows.into_iter().map(|row| row.get(0)).collect())
}
async fn get_datatable_acl(
authed: ApiAuthed,
Extension(db): Extension<DB>,
Path((w_id, datatable_name)): Path<(String, String)>,
Query(query): Query<AclTargetQuery>,
) -> JsonResult<DatatableAclInfo> {
crate::datatable_acl_oss::ensure_datatable_acl_available()?;
let target: AclTarget = query.try_into()?;
let governing = resolve_governing_datatable(&db, &w_id, &datatable_name).await?;
ensure_reaches_governing_datatable(&db, &w_id, &datatable_name, &governing, &authed).await?;
ensure_instance(&governing)?;
let editable = ensure_governs_datatable(&db, &authed, &w_id, &governing)
.await
.is_ok();
let roles = if editable {
role_names(&read_role_catalog(&db).await?)
} else {
vec![]
};
let (client, _notices, dbname) = connect_as_admin_unchecked(&db, &governing).await?;
let owner = read_owner(&client, &target)
.await?
.ok_or_else(|| Error::NotFound(format!("{} not found", target.label(&dbname))))?;
let grants = read_grants(&client, &target).await?;
let supports_maintain: bool = client
.query_one(
"SELECT current_setting('server_version_num')::int >= 170000",
&[],
)
.await
.map_err(|e| {
Error::internal_err(format!(
"Failed to read the server version: {}",
pg_error_message(&e)
))
})?
.get(0);
let children = read_children(&client, &target).await?;
Ok(Json(DatatableAclInfo {
owner: role_name_of(&owner),
roles,
editable,
clone: governing.governor.is_some(),
supports_maintain,
dbname,
grants,
children,
}))
}
async fn read_grants(client: &tokio_postgres::Client, target: &AclTarget) -> Result<Vec<AclGrant>> {
// `aclexplode` turns an acl array into one row per (grantee, privilege); grantee 0 is PUBLIC,
// which has no name to resolve. A NULL acl is not "no access" but Postgres's built-in default —
// the owner holds everything and, on a routine, PUBLIC may EXECUTE — hence `acldefault`. The
// owner's own entries are left out: what it holds comes with ownership, which the owner shows,
// not with a grant a revoke here could take back. Each row ends with its source — the grantor,
// or a default privilege's creating role — and whether this connection can take back what it
// gave.
// Column-level grants (`pg_attribute.attacl`) are not supported yet: they are neither read here
// nor revocable from the editor.
let mut rows = match target {
AclTarget::Database => {
let mut out = client
.query(
concat!(
"SELECT CASE WHEN a.grantee = 0 THEN 'PUBLIC' ELSE pg_get_userbyid(a.grantee) END,
a.privilege_type, NULL::text, NULL::text, NULL::text, NULL::text, ",
grant_source!("d.datdba"),
" FROM pg_database d, aclexplode(COALESCE(d.datacl, acldefault('d', d.datdba))) a
WHERE d.datname = current_database() AND a.grantee <> d.datdba"
),
&[],
)
.await
.map_err(grant_read_error)?;
out.extend(
client
.query(
// Default privileges set database-wide reach what is created in every
// schema, so they are the database's to show rather than any schema's.
concat!(
"SELECT CASE WHEN g.grantee = 0 THEN 'PUBLIC' ELSE pg_get_userbyid(g.grantee) END,
g.privilege_type, NULL::text,
CASE g.defaclobjtype
WHEN 'r' THEN 'TABLES' WHEN 'S' THEN 'SEQUENCES'
WHEN 'f' THEN 'FUNCTIONS' WHEN 'n' THEN 'SCHEMAS'
ELSE 'TYPES' END, NULL::text, NULL::text,
pg_get_userbyid(g.defaclrole), pg_has_role(g.defaclrole, 'USAGE')
FROM (",
database_wide_defaults!(),
") g"
),
&[],
)
.await
.map_err(grant_read_error)?,
);
out
}
AclTarget::Schema { schema } => {
let mut out = client
.query(
concat!(
"SELECT CASE WHEN a.grantee = 0 THEN 'PUBLIC' ELSE pg_get_userbyid(a.grantee) END,
a.privilege_type, NULL::text, NULL::text, NULL::text, NULL::text, ",
grant_source!("n.nspowner"),
" FROM pg_namespace n, aclexplode(COALESCE(n.nspacl, acldefault('n', n.nspowner))) a
WHERE n.nspname = $1 AND a.grantee <> n.nspowner"
),
&[schema],
)
.await
.map_err(grant_read_error)?;
out.extend(
client
.query(
concat!(
"SELECT CASE WHEN a.grantee = 0 THEN 'PUBLIC' ELSE pg_get_userbyid(a.grantee) END,
a.privilege_type, c.relname, NULL::text,
CASE c.relkind WHEN 'S' THEN 'SEQUENCE' ELSE 'TABLE' END,
NULL::text, ",
grant_source!("c.relowner"),
" FROM pg_class c
JOIN pg_namespace n ON n.oid = c.relnamespace,
aclexplode(COALESCE(c.relacl, acldefault(
CASE c.relkind WHEN 'S' THEN 's' ELSE 'r' END::\"char\", c.relowner))) a
WHERE n.nspname = $1
AND c.relkind = ANY(ARRAY['r','p','v','m','S','f']::\"char\"[])
AND a.grantee <> c.relowner"
),
&[schema],
)
.await
.map_err(grant_read_error)?,
);
out.extend(
client
.query(
// Routines carry their own acl in `pg_proc`; without this a grant made here
// would vanish on the next read and could never be revoked back.
concat!(
"SELECT CASE WHEN a.grantee = 0 THEN 'PUBLIC' ELSE pg_get_userbyid(a.grantee) END,
a.privilege_type, p.proname, NULL::text,
CASE p.prokind WHEN 'p' THEN 'PROCEDURE' ELSE 'FUNCTION' END,
pg_get_function_identity_arguments(p.oid), ",
grant_source!("p.proowner"),
" FROM pg_proc p
JOIN pg_namespace n ON n.oid = p.pronamespace,
aclexplode(COALESCE(p.proacl, acldefault('f', p.proowner))) a
WHERE n.nspname = $1 AND a.grantee <> p.proowner"
),
&[schema],
)
.await
.map_err(grant_read_error)?,
);
out.extend(
client
.query(
// `USAGE` on a type is what lets a role use it in a column. Only a type the
// schema holds in its own right has an acl: an array, a row type and a
// multirange answer to their element, table or range.
concat!(
"SELECT CASE WHEN a.grantee = 0 THEN 'PUBLIC' ELSE pg_get_userbyid(a.grantee) END,
a.privilege_type, t.typname, NULL::text, 'TYPE', NULL::text, ",
grant_source!("t.typowner"),
" FROM pg_type t
JOIN pg_depend d ON d.classid = 'pg_type'::regclass AND d.objid = t.oid
AND d.refclassid = 'pg_namespace'::regclass AND d.deptype = 'n',
aclexplode(COALESCE(t.typacl, acldefault('T', t.typowner))) a
WHERE d.refobjid = (SELECT oid FROM pg_namespace WHERE nspname = $1)
AND a.grantee <> t.typowner
AND NOT EXISTS (
SELECT 1 FROM pg_depend x
WHERE x.classid = 'pg_type'::regclass AND x.objid = t.oid
AND x.objsubid = 0 AND x.deptype = 'i')"
),
&[schema],
)
.await
.map_err(grant_read_error)?,
);
out.extend(
client
.query(
// What a creating role set is taken back `FOR ROLE` that role, which only
// a role acting for it may do.
"SELECT CASE WHEN a.grantee = 0 THEN 'PUBLIC' ELSE pg_get_userbyid(a.grantee) END,
a.privilege_type, NULL::text,
CASE d.defaclobjtype
WHEN 'r' THEN 'TABLES' WHEN 'S' THEN 'SEQUENCES'
WHEN 'f' THEN 'FUNCTIONS' ELSE 'TYPES' END, NULL::text, NULL::text,
pg_get_userbyid(d.defaclrole), pg_has_role(d.defaclrole, 'USAGE')
FROM pg_default_acl d
JOIN pg_namespace n ON n.oid = d.defaclnamespace,
aclexplode(d.defaclacl) a
WHERE n.nspname = $1",
&[schema],
)
.await
.map_err(grant_read_error)?,
);
out
}
AclTarget::Table { schema, table } => client
.query(
concat!(
"SELECT CASE WHEN a.grantee = 0 THEN 'PUBLIC' ELSE pg_get_userbyid(a.grantee) END,
a.privilege_type, NULL::text, NULL::text, NULL::text, NULL::text, ",
grant_source!("c.relowner"),
" FROM pg_class c
JOIN pg_namespace n ON n.oid = c.relnamespace,
aclexplode(COALESCE(c.relacl, acldefault(
CASE c.relkind WHEN 'S' THEN 's' ELSE 'r' END::\"char\", c.relowner))) a
WHERE n.nspname = $1 AND c.relname = $2 AND a.grantee <> c.relowner"
),
&[schema, table],
)
.await
.map_err(grant_read_error)?,
};
// One row per privilege and source — and, for default privileges, per creating role. Fold them
// back into one entry per grantee and object that keeps every source and what each gave: a
// revoke takes a privilege back from every source that gave it.
let mut folded: BTreeMap<
(
String,
Option<(String, String, Option<String>)>,
Option<String>,
),
(Vec<String>, BTreeMap<String, (bool, Vec<String>)>),
> = BTreeMap::new();
for row in rows.drain(..) {
let grantee: String = row.get(0);
let privilege: String = row.get(1);
let object: Option<String> = row.get(2);
let future: Option<String> = row.get(3);
let object_kind: Option<String> = row.get(4);
let object_args: Option<String> = row.get(5);
let source: String = row.get(6);
let reachable: bool = row.get(7);
let (privileges, sources) = folded
.entry((
role_name_of(&grantee),
object.map(|name| {
(
name,
object_kind.unwrap_or_else(|| "TABLE".to_string()),
object_args,
)
}),
future,
))
.or_default();
sources
.entry(role_name_of(&source))
.or_insert_with(|| (reachable, vec![]))
.1
.push(privilege.clone());
privileges.push(privilege);
}
Ok(folded
.into_iter()
.map(|((grantee, object, future), (mut privileges, sources))| {
privileges.sort();
privileges.dedup();
AclGrant {
grantee,
privileges,
object: object.map(|(name, kind, args)| AclObject { name, kind, args }),
future,
sources: sources
.into_iter()
.map(|(role, (reachable, mut privileges))| {
privileges.sort();
privileges.dedup();
AclSource { role, privileges, reachable }
})
.collect(),
}
})
.collect())
}
fn grant_read_error(e: tokio_postgres::Error) -> Error {
Error::internal_err(format!("Failed to read grants: {}", pg_error_message(&e)))
}
/// Changing a data table's access is administering it. Checked in full before anything connects
/// with the instance's credentials.
async fn authorize_acl_change(
db: &DB,
authed: &ApiAuthed,
w_id: &str,
datatable_name: &str,
) -> Result<GoverningDatatable> {
let governing = resolve_governing_datatable(db, w_id, datatable_name).await?;
ensure_governs_datatable(db, authed, w_id, &governing).await?;
ensure_instance(&governing)?;
Ok(governing)
}
static APPLY_SLOT: tokio::sync::Semaphore = tokio::sync::Semaphore::const_new(1);
/// A role passes on only privileges it holds with grant option, and an instance database
/// provisioned before data table roles gave `custom_instance_user` none. Adds that option to its
/// database and `public` privileges, and nothing else: default privileges are left alone, since a
/// schema's change of owner is planned against them. Best-effort, as a grant it fails to enable is
/// refused when it runs.
async fn ensure_grant_options(client: &tokio_postgres::Client, db: &DB, dbname: &str) {
let held = client
.query_one(
"SELECT has_database_privilege(current_database(), 'CONNECT WITH GRANT OPTION')
AND has_database_privilege(current_database(), 'CREATE WITH GRANT OPTION')
AND (to_regnamespace('public') IS NULL
OR (has_schema_privilege('public', 'USAGE WITH GRANT OPTION')
AND has_schema_privilege('public', 'CREATE WITH GRANT OPTION')))",
&[],
)
.await
.is_ok_and(|row| row.get::<_, bool>(0));
if held {
return;
}
if let Err(e) = grant_options_as_server(db, dbname).await {
tracing::warn!("Could not enable grant options on '{dbname}': {e}");
}
}
/// Only the database's owner, the server's own Postgres user, can hand out an option it holds.
async fn grant_options_as_server(db: &DB, dbname: &str) -> Result<()> {
let server = PgDatabase::parse_uri(&windmill_common::get_database_url().await?.as_str().await)?;
let creds = PgDatabase { dbname: dbname.to_string(), ..server };
let (client, connection) = creds.connect(Some(db)).await?;
let join_handle = tokio::spawn(async move { connection.await });
let role = quote_ident(CUSTOM_INSTANCE_USER);
let result = client
.batch_execute(&format!(
"GRANT CONNECT, CREATE ON DATABASE {} TO {role} WITH GRANT OPTION;
DO $$ BEGIN
IF to_regnamespace('public') IS NOT NULL THEN
GRANT USAGE, CREATE ON SCHEMA public TO {role} WITH GRANT OPTION;
END IF;
END $$;",
quote_ident(dbname)
))
.await;
drop(client);
windmill_common::shutdown_pg_connection(join_handle).await?;
result.map_err(|e| {
Error::internal_err(format!(
"Failed to grant options on '{dbname}': {}",
pg_error_message(&e)
))
})
}
/// Whether the governing entry an apply was authorized on is still the one in the settings, read
/// under the lock: a save in between could have pointed it at another database or changed its roles.
fn entry_unchanged(governing: &GoverningDatatable, entry_now: Option<serde_json::Value>) -> bool {
let Some(Ok(now)) = entry_now.map(serde_json::from_value::<DataTable>) else {
return false;
};
match (
serde_json::to_value(&now),
serde_json::to_value(&governing.datatable),
) {
(Ok(now), Ok(authorized)) => now == authorized,
_ => false,
}
}
/// Plan one change against the catalog and the database as they are now.
async fn build_plan(
client: &tokio_postgres::Client,
dbname: &str,
catalog: &DatatableRoleCatalog,
target: &AclTarget,
change: &AclChange,
) -> Result<AclPlan> {
let role = change.role();
// `admin` is the login the data table itself reaches Postgres through, and the one every
// change here runs as: a revoke that lands leaves nothing able to grant it back.
if matches!(change, AclChange::Revoke { .. }) && role == ADMIN_DATATABLE_ROLE {
return Err(Error::BadRequest(format!(
"'{ADMIN_DATATABLE_ROLE}' is how this data table reaches its database; \
its own access is not revocable from here"
)));
}
let pg_role = pg_role_of(role, catalog)?;
let change = match change {
AclChange::Revoke { role, privileges, scope, objects } => AclChange::Revoke {
role: role.clone(),
privileges: privileges.clone(),
scope: *scope,
objects: resolve_acl_objects(client, target, objects).await?,
},
change => change.clone(),
};
// Default privileges are recorded per creating role, and a schema's new owner is kept in reach
// of what the others create there, so both are written for every role there is.
let other_pg_roles = role_names(catalog)
.iter()
.map(|name| pg_role_of(name, catalog))
.filter(|r| r.as_ref().map_or(true, |r| *r != pg_role))
.collect::<Result<Vec<_>>>()?;
let (existing_objects, former_owner) = match (&change, target) {
(AclChange::SetOwner { .. }, AclTarget::Schema { schema }) => (
read_owned_objects(client, schema).await?,
read_former_owner_defaults(client, schema, &pg_role).await?,
),
_ => (vec![], None),
};
let revoked_grants = match &change {
AclChange::Revoke { privileges, scope, objects, .. } => {
read_revoked_grants(
client, dbname, target, *scope, objects, privileges, &pg_role,
)
.await?
}
_ => vec![],
};
if matches!(change, AclChange::SetOwner { .. }) {
if let Some((object, owner)) = unmanaged_owner(client, target).await? {
return Err(Error::BadRequest(format!(
"{object} is owned by {owner}, which this data table's connection cannot act \
for, so its owner cannot be changed from here"
)));
}
}
let facts = CatalogFacts { other_pg_roles, existing_objects, former_owner, revoked_grants };
let mut plan =
crate::datatable_acl_oss::plan_statements(target, &change, dbname, &pg_role, &facts)?;
if matches!(change, AclChange::SetOwner { .. }) {
if let Some(missing) = missing_owner_privilege(client, target, &pg_role).await? {
plan.warnings.push(format!(
"{role} does not have {missing}, which Postgres requires of a new owner, so this \
will be refused. Grant it first."
));
}
}
Ok(plan)
}
/// Postgres only hands an object to a role that could have created it: a table to one with
/// `CREATE` on its schema, a schema to one with `CREATE` on the database.
async fn missing_owner_privilege(
client: &tokio_postgres::Client,
target: &AclTarget,
pg_role: &str,
) -> Result<Option<String>> {
let (row, missing) = match target {
AclTarget::Table { schema, .. } => (
client
.query_one(
"SELECT has_schema_privilege($1::name, $2::text, 'CREATE')",
&[&pg_role, schema],
)
.await,
format!("CREATE on schema {schema}"),
),
AclTarget::Schema { .. } => (
client
.query_one(
"SELECT has_database_privilege($1::name, current_database(), 'CREATE')",
&[&pg_role],
)
.await,
"CREATE on the database".to_string(),
),
AclTarget::Database => return Ok(None),
};
let has: bool = row
.map_err(|e| {
Error::internal_err(format!(
"Failed to read the new owner's privileges: {}",
pg_error_message(&e)
))
})?
.get(0);
Ok((!has).then_some(missing))
}
/// The first thing a change of owner would move that this connection cannot act for, with its
/// owner. Postgres lets only a member of the current owner move an object, and every change runs as
/// `custom_instance_user`, so an object it does not hold the owner of — `public`, owned by the
/// database's owner, above all — is refused here rather than at apply. Running as the instance's
/// own user instead would reach objects Windmill never created.
async fn unmanaged_owner(
client: &tokio_postgres::Client,
target: &AclTarget,
) -> Result<Option<(String, String)>> {
let read_error = |e: tokio_postgres::Error| {
Error::internal_err(format!(
"Failed to read who owns what the change moves: {}",
pg_error_message(&e)
))
};
match target {
AclTarget::Database => Ok(None),
AclTarget::Table { schema, table } => {
let row = client
.query_opt(
"SELECT pg_get_userbyid(c.relowner)
FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE n.nspname = $1 AND c.relname = $2 AND NOT pg_has_role(c.relowner, 'USAGE')",
&[schema, table],
)
.await
.map_err(read_error)?;
Ok(row.map(|row| (format!("{schema}.{table}"), row.get(0))))
}
AclTarget::Schema { schema } => {
let row = client
.query_opt(
concat!(
"SELECT ord, identity, pg_get_userbyid(owner) FROM (
SELECT 0 AS ord, NULL::text AS identity, n.nspowner AS owner
FROM pg_namespace n WHERE n.nspname = $1
UNION ALL
SELECT 1, identity, owner FROM (",
schema_owned_objects!(),
") m
) o WHERE NOT pg_has_role(owner, 'USAGE') ORDER BY ord, identity LIMIT 1"
),
&[schema],
)
.await
.map_err(read_error)?;
Ok(row.map(|row| {
let label = row
.get::<_, Option<String>>(1)
.unwrap_or_else(|| format!("schema {schema}"));
(label, row.get(2))
}))
}
}
}
async fn plan_datatable_acl(
authed: ApiAuthed,
Extension(db): Extension<DB>,
Path((w_id, datatable_name)): Path<(String, String)>,
Json(req): Json<AclChangeRequest>,
) -> JsonResult<AclPlan> {
crate::datatable_acl_oss::ensure_datatable_acl_available()?;
let governing = authorize_acl_change(&db, &authed, &w_id, &datatable_name).await?;
let catalog = read_role_catalog(&db).await?;
let (client, _notices, dbname) = connect_as_admin_unchecked(&db, &governing).await?;
Ok(Json(
build_plan(&client, &dbname, &catalog, &req.target, &req.change).await?,
))
}
async fn apply_datatable_acl(
authed: ApiAuthed,
Extension(db): Extension<DB>,
Path((w_id, datatable_name)): Path<(String, String)>,
Json(req): Json<AclChangeRequest>,
) -> Result<String> {
crate::datatable_acl_oss::ensure_datatable_acl_available()?;
let confirmed = req.statements.as_ref().ok_or_else(|| {
Error::BadRequest(
"An apply runs exactly the statements its plan showed; plan the change first"
.to_string(),
)
})?;
// Everything that needs the pool happens before the locks: once `tx` holds them, a second pool
// connection could wait forever on a pool that concurrent applies, queued on the same locks,
// have exhausted.
let governing = authorize_acl_change(&db, &authed, &w_id, &datatable_name).await?;
// Applies queue on an instance-wide lock while each holds a direct connection to the instance's
// Postgres; unbounded, the queue alone could exhaust its connection limit. One at a time per
// server, and the ones waiting hold no connection at all.
let _slot = APPLY_SLOT
.acquire()
.await
.map_err(|e| Error::internal_err(format!("ACL apply slot closed: {e}")))?;
let (mut client, mut notices, dbname) = connect_as_admin_unchecked(&db, &governing).await?;
ensure_grant_options(&client, &db, &dbname).await;
// Held until the change is committed: a role renamed or dropped meanwhile would change what
// the plan names, and a settings save could move the entry onto another database. Taken in the
// same order as the permissions save, so the two cannot deadlock.
let mut tx = db.begin().await?;
lock_role_catalog(&mut tx).await?;
let entry_now = sqlx::query_scalar::<_, Option<serde_json::Value>>(
"SELECT datatable->'datatables'->$2 FROM workspace_settings WHERE workspace_id = $1 FOR UPDATE",
)
.bind(&governing.workspace_id)
.bind(&governing.name)
.fetch_optional(&mut *tx)
.await?
.flatten();
let catalog = read_role_catalog_tx(&mut tx).await?;
let plan = build_plan(&client, &dbname, &catalog, &req.target, &req.change).await?;
if !entry_unchanged(&governing, entry_now) || &plan.statements != confirmed {
return Err(Error::BadRequest(
"The data table or its roles changed since this was planned, so it would no longer \
run what was confirmed. Plan it again."
.to_string(),
));
}
// One transaction: a half-applied ownership transfer leaves one schema's objects owned by two
// different roles.
let pg_tx = client.transaction().await.map_err(|e| {
Error::internal_err(format!(
"Failed to open a transaction on the data table: {}",
pg_error_message(&e)
))
})?;
execute_acl_statements(&pg_tx, &mut notices, &plan.statements).await?;
// A schema's objects were listed before the transaction opened; one committed since would stay
// with its old owner. One created while this transaction is still open can still slip past, as
// Postgres has no lock that holds creation in a schema back. That is benign: it stays with its
// creator, like anything created there later, and moving that table fixes it.
if let (AclChange::SetOwner { role }, AclTarget::Schema { schema }) = (&req.change, &req.target)
{
let new_owner = pg_role_of(role, &catalog)?;
let straggler = pg_tx
.query_opt(
concat!(
"SELECT identity FROM (",
schema_owned_objects!(),
") m WHERE owner <> (SELECT oid FROM pg_roles WHERE rolname = $2)
ORDER BY 1 LIMIT 1"
),
&[schema, &new_owner],
)
.await
.map_err(|e| {
Error::internal_err(format!(
"Failed to check what the schema holds: {}",
pg_error_message(&e)
))
})?;
if let Some(row) = straggler {
return Err(Error::BadRequest(format!(
"{} appeared while this ran and would keep its old owner, so nothing was \
applied. Plan it again.",
row.get::<_, String>(0)
)));
}
}
let target_label = req.target.label(&dbname);
audit_log(
&mut *tx,
&authed,
"workspaces.datatable_acl",
ActionKind::Update,
&governing.workspace_id,
Some(&governing.name),
Some(
[
("target", target_label.as_str()),
("change", change_kind(&req.change)),
("role", req.change.role()),
]
.into(),
),
)
.await?;
pg_tx.commit().await.map_err(|e| {
Error::internal_err(format!(
"Failed to commit the changes: {}",
pg_error_message(&e)
))
})?;
tx.commit().await?;
windmill_common::feature_usage::log_feature_usage(
"datatable",
"acl_applied",
change_kind(&req.change),
);
Ok(format!("Updated access on {target_label}"))
}
/// What kind of change, never what it named: the telemetry key and the audit's summary.
fn change_kind(change: &AclChange) -> &'static str {
match change {
AclChange::SetOwner { .. } => "owner",
AclChange::Grant { scope, .. } | AclChange::Revoke { scope, .. } if scope.is_future() => {
"default_privileges"
}
AclChange::Grant { .. } => "grant",
AclChange::Revoke { .. } => "revoke",
}
}
#[cfg(test)]
mod tests {
use super::*;
use windmill_common::datatable_roles::InstanceDatatableRole;
#[test]
fn a_role_is_its_own_postgres_role_except_admin() {
let catalog: DatatableRoleCatalog = BTreeMap::from([(
"role1".to_string(),
InstanceDatatableRole { name: "analytics".to_string(), enabled: true, pwd: None },
)]);
assert_eq!(
pg_role_of("admin", &catalog).unwrap(),
"custom_instance_user"
);
assert_eq!(pg_role_of("analytics", &catalog).unwrap(), "analytics");
assert_eq!(role_name_of("custom_instance_user"), "admin");
assert_eq!(role_name_of("analytics"), "analytics");
// A catalog id, a name the catalog lacks, or the admin login spelled out never stands for
// some other role.
for unknown in ["role1", "operator", "custom_instance_user", "PUBLIC", ""] {
assert!(
matches!(pg_role_of(unknown, &catalog), Err(Error::BadRequest(_))),
"{unknown}"
);
}
}
/// A connection to the test's own database, the way the handlers reach a data table's.
async fn catalog_client(pool: &sqlx::PgPool) -> tokio_postgres::Client {
let mut config: tokio_postgres::Config =
std::env::var("DATABASE_URL").unwrap().parse().unwrap();
config.dbname(pool.connect_options().get_database().unwrap());
let (client, connection) = config.connect(tokio_postgres::NoTls).await.unwrap();
tokio::spawn(connection);
client
}
/// What a revoke takes back is read from the catalog, per object and source, and only the
/// privileges it asks for: the planner renders exactly this, and refuses an empty read.
#[sqlx::test(migrations = false)]
async fn a_revoke_reads_back_only_what_it_asks_for(pool: sqlx::PgPool) {
let client = catalog_client(&pool).await;
// A predefined role, so that nothing is granted outside the test's own database.
client
.batch_execute(
"CREATE SCHEMA granted;
CREATE TABLE granted.g (id int);
GRANT SELECT, INSERT ON granted.g TO pg_read_all_data;",
)
.await
.unwrap();
let target = AclTarget::Table { schema: "granted".to_string(), table: "g".to_string() };
let revoked = read_revoked_grants(
&client,
"db",
&target,
GrantScope::Target,
&[],
&["select".to_string()],
"pg_read_all_data",
)
.await
.unwrap();
assert_eq!(revoked.len(), 1, "{revoked:?}");
assert_eq!(revoked[0].object, None);
assert_eq!(revoked[0].privileges, ["SELECT"]);
let held_none = read_revoked_grants(
&client,
"db",
&target,
GrantScope::Target,
&[],
&["update".to_string()],
"pg_read_all_data",
)
.await
.unwrap();
assert!(held_none.is_empty(), "{held_none:?}");
// Each source says what it gave: that, and not the whole row, is what a revoke of some of
// its privileges is held back by.
let grants = read_grants(
&client,
&AclTarget::Schema { schema: "granted".to_string() },
)
.await
.unwrap();
let on_g = grants
.iter()
.find(|g| {
g.grantee == "pg_read_all_data" && g.object.as_ref().is_some_and(|o| o.name == "g")
})
.unwrap();
assert_eq!(on_g.sources.len(), 1, "{:?}", on_g.sources);
assert_eq!(on_g.sources[0].privileges, ["INSERT", "SELECT"]);
assert!(on_g.sources[0].reachable);
}
/// A default privilege set database-wide applies in every schema on top of the schema's own,
/// which cannot take it back: the database shows it, and neither a schema's revoke of it nor a
/// change of owner away from its grantee may go ahead as if it were gone.
#[sqlx::test(migrations = false)]
async fn a_schema_cannot_take_back_a_database_wide_default(pool: sqlx::PgPool) {
let client = catalog_client(&pool).await;
client
.batch_execute(
"CREATE SCHEMA owned AUTHORIZATION pg_read_all_data;
ALTER DEFAULT PRIVILEGES GRANT SELECT ON TABLES TO pg_read_all_data;
ALTER DEFAULT PRIVILEGES IN SCHEMA owned
GRANT SELECT, INSERT ON TABLES TO pg_read_all_data;
CREATE SCHEMA creators AUTHORIZATION pg_monitor;
ALTER DEFAULT PRIVILEGES FOR ROLE pg_monitor
GRANT SELECT ON TABLES TO pg_read_all_stats;",
)
.await
.unwrap();
let creator: String = client
.query_one("SELECT current_user::text", &[])
.await
.unwrap()
.get(0);
let grants = read_grants(&client, &AclTarget::Database).await.unwrap();
let database_wide = grants
.iter()
.find(|g| g.grantee == "pg_read_all_data" && g.future.as_deref() == Some("TABLES"))
.unwrap_or_else(|| panic!("{grants:?}"));
assert_eq!(database_wide.privileges, ["SELECT"]);
// A database-wide entry also holds its creator's own privileges, which come with creating
// and are no grant: neither a row nor a refusal may stem from them.
assert!(
!grants
.iter()
.any(|g| g.future.is_some() && (g.grantee == creator || g.grantee == "pg_monitor")),
"{grants:?}"
);
assert_eq!(
database_wide.sources.len(),
1,
"{:?}",
database_wide.sources
);
assert_eq!(database_wide.sources[0].role, creator);
let schema = AclTarget::Schema { schema: "owned".to_string() };
let still_granted = read_revoked_grants(
&client,
"db",
&schema,
GrantScope::FutureTables,
&[],
&["select".to_string()],
"pg_read_all_data",
)
.await;
assert!(
matches!(&still_granted, Err(Error::BadRequest(m)) if m.contains("database-wide")),
"{still_granted:?}"
);
let schema_only = read_revoked_grants(
&client,
"db",
&schema,
GrantScope::FutureTables,
&[],
&["insert".to_string()],
"pg_read_all_data",
)
.await
.unwrap();
assert_eq!(schema_only.len(), 1, "{schema_only:?}");
assert_eq!(schema_only[0].privileges, ["INSERT"]);
let moved = read_former_owner_defaults(&client, "owned", "pg_write_all_data").await;
assert!(
matches!(&moved, Err(Error::BadRequest(m)) if m.contains("database-wide")),
"{moved:?}"
);
let moved_from_creator =
read_former_owner_defaults(&client, "creators", "pg_write_all_data").await;
assert!(
matches!(moved_from_creator, Ok(None)),
"{moved_from_creator:?}"
);
}
/// Defaults on types go with a schema's owner like the other kinds: once a database-wide
/// default takes PUBLIC's USAGE on types away, they are all that reaches a new type.
#[sqlx::test(migrations = false)]
async fn a_schemas_former_owner_defaults_include_types(pool: sqlx::PgPool) {
let client = catalog_client(&pool).await;
client
.batch_execute(
"CREATE SCHEMA typed AUTHORIZATION pg_read_all_data;
ALTER DEFAULT PRIVILEGES IN SCHEMA typed GRANT USAGE ON TYPES TO pg_read_all_data;
ALTER DEFAULT PRIVILEGES FOR ROLE pg_read_all_data IN SCHEMA typed
GRANT SELECT ON TABLES TO pg_read_all_data;",
)
.await
.unwrap();
let creator: String = client
.query_one("SELECT current_user::text", &[])
.await
.unwrap()
.get(0);
let former = read_former_owner_defaults(&client, "typed", "pg_write_all_data")
.await
.unwrap();
assert_eq!(
former,
Some(FormerOwnerDefaults {
pg_role: "pg_read_all_data".to_string(),
defaults: vec![(creator, "TYPES")],
})
);
}
/// A kind of object the list misses stays with its old owner while the schema changes hands,
/// which only a real catalog shows.
#[sqlx::test(migrations = false)]
async fn a_schemas_owner_change_takes_every_object_in_it(pool: sqlx::PgPool) {
let client = catalog_client(&pool).await;
client
.batch_execute(
"CREATE SCHEMA moved;
CREATE TABLE moved.t (id serial PRIMARY KEY, a int, b int);
CREATE STATISTICS moved.st ON a, b FROM moved.t;
CREATE TABLE moved.pt (id int) PARTITION BY RANGE (id);
CREATE TABLE moved.pt1 PARTITION OF moved.pt FOR VALUES FROM (0) TO (10);
CREATE TYPE moved.r AS RANGE (subtype = float8);
CREATE TYPE moved.pair AS (a int, b int);
CREATE TYPE moved.mood AS ENUM ('ok');
CREATE DOMAIN moved.tags AS text[];
CREATE COLLATION moved.coll (provider = libc, locale = 'C');",
)
.await
.unwrap();
let owned = read_owned_objects(&client, "moved").await.unwrap();
// Not the serial's sequence, the range's constructors and multirange, or any array or row
// type: each follows the object it belongs to.
assert_eq!(
owned
.iter()
.map(|o| (o.keyword, o.identity.as_str()))
.collect::<Vec<_>>(),
[
("COLLATION", "moved.coll"),
("STATISTICS", "moved.st"),
("TABLE", "moved.pt"),
("TABLE", "moved.pt1"),
("TABLE", "moved.t"),
("TYPE", "moved.mood"),
("TYPE", "moved.pair"),
("TYPE", "moved.r"),
("TYPE", "moved.tags"),
]
);
}
}