mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-10-03 16:02:12 +00:00
9f40cdca6258a9218770ebe1e7b89179e7d6667f
14935
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9f40cdca62 |
feat: add pull_batch to claim jobs for many waiting workers at once (#11350)
* feat: add pull_batch to claim jobs for many waiting workers at once Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep jobs admitted by earlier batch passes when a re-pull fails Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep suspended flows first on every batch re-pull pass Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
ede4103b00 |
feat: share AI guidance with the CLI skills and lint flow groups (#11397)
* feat: lint AI agent tool names and flow groups in wmill lint * refactor: assemble chat and CLI AI guidance from one topic table * feat: share flow groups, reuse and pipeline guidance with the CLI skills * feat: share raw app, data table and secret guidance between chat and CLI * docs: document the shared AI guidance source for contributors * fix: keep wmill lint running on flows with malformed collections * fix: reject skill descriptions that are not plain YAML text * fix: tighten fence typos, script base scope and app prompt order * fix: skip tool name checks on agent steps linked to a saved agent * fix: catch any misspelled prompt fence and soften the tool name claim * fix: align cli eval harness with the files and steps wmill init adds * fix: drop cli eval checks that expect unrequested deploy commands * fix: list ansible as mainless and c# Main in script base guidance |
||
|
|
d44c901647 |
replace the AI sessions beta banner with a feedback link (#11401)
* feat: make the AI sessions beta banner dismissible Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: move AI sessions feedback link into assistant settings Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor: reduce the sessions beta gate setter to opt-in only Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: size the sessions activate buttons with unifiedSize Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep the sessions page activate button at its previous height Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
c34abf7330 | fix: refuse flow preview restarts from runs the caller cannot read (#11407) | ||
|
|
8741843d2e |
feat: external instance cluster for data tables and Ducklake catalogs (#11197)
* fix pg_dump stuck on version 17 on nix * fix(datatables): refuse a malformed role annotation instead of ignoring it `-- Role operator`, `-- role operator;` and `-- role operator -- why` all failed the annotation parser's exact-match rule, so the query fell through to the data table's default role and ran, silently, under a login the author did not choose. Naming a role exists precisely to not do that. A leading comment whose first word is `role` is now an annotation attempt: the keyword matches case-insensitively, one trailing `;` is tolerated, and anything else is an error naming the line. Only callers that already know the target is a `datatable://` reference ever run this, so ordinary SQL keeps its comments. Also bumps the dev shell's postgres client to 18 — it trailed the server the dev database runs, which takes out every data table export, clone and fork-with-data. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): refuse a malformed role query string instead of ignoring it `?Role=analytics`, `?role=` and `?x=1&role=…` all fell through the reference parser's exact-match rule, so the connection resolved to the data table's default role and ran under a login the caller never asked for — the URI half of the same trap as a malformed `-- role` annotation. The key now matches case-insensitively, and anything else in the query string is an error naming it; `role` is the only parameter a reference takes. Callers that only need the entry keep a lenient `datatable_ref_name`, since they never act on the role. The DuckDB `ATTACH` parser propagates it rather than attaching under the default. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ti5HyeTikPMYyW8YSdiHR * fix(datatables): carry the role annotation into the row_to_json retry The retry rebuilds its SQL from `pruneComments(code)`, so the leading comment block never reached the second attempt — and with it the `-- role <name>` line that decides which login the query runs as. The retry connected as the data table's default role instead, so a query the first attempt was denied could succeed on the second, reported as "recovered with the row_to_json fix". Carry the leading comment block over. The retry itself is unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * chore(datatables): don't mount the roles UI until the ACL editor lands Enforcement ships first. The permissions drawer is what turns roles on, and the catalog section is what creates them — both are only useful once there is a way to grant a role the privileges it needs, which arrives with the ACL editor. Left mounted they would offer a feature whose other half does not exist. The two components are complete and reviewed; only their call sites here are commented out, with a note pointing the follow-up PRs at them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): honour `-- role: x`, and fix the DuckDB attach test Two review findings, both real. `attach_datatable_parses_name_and_role` never compiled: `parse_attach_datatable` returns `Result<Option<_>>` now and one call site kept a single `unwrap`. Its `?Role=analytics` case also asserted a refusal, contradicting the parser in the same commit, which matches the key case-insensitively. Replaced with the cases that are genuinely malformed, and a positive one for the cased key. `-- role: analytics` fell through to the default role — the silent fallback the strict parser exists to remove, for the spelling most likely to be typed. The keyword now accepts an optional colon, attached or spaced, while a word that merely starts with it (`rolebased`) is still not an attempt. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): clone a fork's pointer instead of failing after the copy Forking a fork with cloning left an orphan database. The preflight resolves the pointer and sees the governing entry, so both endpoints ran and filled the new database; `apply_forked_datatable` then refused the inherited pointer and rolled the fork back, stranding a registered `wm_fork_*` that no entry names and whose name blocks the retry. Refusing earlier would have been the smaller change, but forking a fork and cloning worked before pointers existed, so it would trade an orphan for a regression. Resolve what the pointer names and write the terminal entry the clone needs: the whole `database` object rather than a patch of its `resource_path`, since a pointer has none, and `reference` removed with it. Also accepts `-- role=x` and `-- Role = x`, two more spellings that fell through to the default role. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): refuse to roll back the catalog while roles exist The down migration dropped the table and left every role behind: live Postgres logins whose passwords only that table carried, so after a revert Windmill could neither use, disable nor delete them, and re-applying could not recreate them because the names were taken. Cleaning up here is not possible either — dropping a role means reassigning what it owns in every instance database, and a migration runs in one — so it now refuses while the catalog is non-empty and says to delete the roles through instance settings, which does the cluster work. Also enforces the instance-only invariant the resolved-pointer clone relies on rather than only asserting it in a comment. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * refactor(datatables): settle clonability in one place, before anything is created A clone is three stages a workspace apart — `create_pg_database`, then `import_pg_database`, then `apply_forked_datatable` inside the fork transaction. Only the third can roll back, and `CREATE DATABASE` is not transactional, so any refusal that lives there strands a registered `wm_fork_*` that no entry names and whose name blocks the retry. That orphan has now been fixed three times, most recently reintroduced by a guard added one commit ago. Patching each new refusal into the first endpoint is not the fix; having two places that can refuse is. `ensure_datatable_is_clonable` now answers every reason a copy can be refused and returns what it resolved, and the stage that writes the entry only does the work. Also takes an ACCESS EXCLUSIVE lock before the rollback guard counts, so a role created concurrently cannot slip between the check and the drop. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjfMkJyKzodxkobqGZ6Lqb * fix(datatables): let a retried clone reclaim its own leftover database A clone creates its target database one request before it copies into it, and the fork that would name it is written a request after that. Any failure in between — a pg_dump error, a bad restore, a dropped connection, the source's roles changing mid-flow — left a registered `wm_fork_*` that no entry names, and every retry then failed on its name. This predates data table roles. `create_pg_database` now reclaims such a leftover before creating: only a `wm_fork_*` database Windmill registered as a data table database and that no data table or ducklake entry names, in any workspace, archived ones included. The drop never terminates connections, so a clone still copying into it makes the reclaim fail instead of being cut off. It is limited to callers who administer the source — reaching it is not enough, since on a data table without roles every member reaches it — and anyone else gets the refusal an existing database always got. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Revert "fix(datatables): let a retried clone reclaim its own leftover database" This reverts commit |
||
|
|
797147ea7a |
fix: report a worker's last job when it ran under one poll interval (#11399)
* fix: report a worker's last job when it ran under one poll interval Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: share the unreported job slot with the interactive worker shell Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: pin that a main-loop ping without a job keeps the last one Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
97fa55b719 |
feat: detect and alert when a schedule skips occurrences (#10917)
* docs: plan for detecting skipped schedule occurrences Design plan only, no implementation. Records the scheduler's re-anchoring behaviour, the measurements behind it, and the three-piece design that came out of reviewing the alternatives. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * docs: state the user-facing outcome in the schedule plan The plan described the mechanism but never what a user would see, which made it hard to judge what the work is worth. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * docs: state which cause the schedule plan catches, and correct its scope Records which of the two causes each piece covers, and corrects the overrun scope: a script schedule carrying retry or dynamic_skip is pushed as a SingleStepFlow, so it re-arms at step 0 entry and its occurrences overlap like a flow's. Resolves the no_flow_overlap question, splits the read-time work into bounded detection and editor-only counting behind measured croner costs, and fixes the delivery order. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: count the occurrences a schedule skipped A schedule that overruns its interval, or waits for a worker, silently loses the occurrences in between: the scheduler keeps one queued occurrence and re-anchors on the clock, so nothing records that a run was due and never happened. Recovers the sequence from rows that already exist rather than writing per occurrence. `push_scheduled_job` anchors on `now_from_db` inside the transaction that inserts the job, and `v2_job.created_at` defaults to that same transaction timestamp, so `scheduled_for = find_next(created_at)` holds exactly and the whole occurrence history is derivable. The schedules list reports how many of the recent runs were followed by a lost occurrence, and a new occurrences endpoint carries the per-run wait and duration behind it. Detection is one `find_next` per gap, which stays bounded on a full page; counting walks the gap and runs only for a single schedule. The one write is `occurrence_baseline_at`, advanced at create, edit, re-enable and re-arm. Gaps older than it span a pause, a cron change, a re-enable or a reconciler re-arm, none of which mean runs were lost. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: show the wait and run time behind a schedule's skipped occurrences The list badge says a schedule is losing runs; this says which of the two causes did it. A large wait means not enough workers, a long run means the job outgrew its interval, and the pair is what tells them apart. Sits under the existing upcoming-events panel, so due and overdue read together. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: flag a schedule that is running late right now Reconstruction is retrospective: a gap only appears once the next occurrence has a row, which needs the current one to finish. A schedule wedged mid-run shows nothing until it moves, which is the case an operator most wants to see. An occurrence still in flight past the time its own successor was due will cost that successor, so `now > find_next(scheduled_for)` is the signal, needing no threshold and self-calibrating across a daily and a per-minute schedule. It applies only where occurrences serialize; an overlapping schedule starts its successor on time and would flag constantly while healthy. The queue is read in one aggregating pass keyed on (trigger, runnable_path) rather than a subquery per schedule, and an overlapping schedule holds more than one root row, hence the aggregate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: run the schedule overrun alert from the monitor pass Wires `schedule_overrun_alerts` in next to `jobs_waiting_alerts`, every 30 iterations (~5 min). Its Enterprise implementation lives in windmill-labs/windmill-ee-private#772; only the wiring and the OSS stub are here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * chore: refresh the sqlx offline cache Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: record and alert when a schedule skips occurrences push_scheduled_job compares each chained occurrence with the slot after the previous one. A gap is written to schedule.skipped_occurrences off the push transaction, alerts once when a clean schedule starts skipping, and recovers on the next clean chain. The schedules list shows a badge. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * feat: alert only on a streak of skipping runs, keep a recent skip visible The skip state now describes the current streak and is written in the push transaction, so it commits or rolls back with the push. The alert fires once when 3 runs in a row skipped, and the list keeps a muted badge for 7 days after the latest skip. Editing or toggling a schedule resets it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * fix: name the missed-occurrence state after what it counts, alert only once committed Renames the columns to late_run_streak, missed_occurrences and last_missed_at, keeps the missed count after a streak ends so the muted badge can show it, and rewords both badges. The alert task now reads the streak FOR SHARE, which waits for the push transaction, so a push that rolls back and retries alerts once. A failed slot count leaves the streak untouched, and a schedule deleted mid-push no longer fails it. Adds an integration test for the streak and its reset. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * fix: recover the late run alert, store the missed slot, name it missed throughout The alert now recovers (and so acknowledges itself) when a streak that alerted ends on a run on time, under the schedule:{path} resource used by the other trigger alerts. last_missed_at records the last missed cron slot rather than when the late run chained, and the counting helpers say missed, since skipped already names occurrences queued and not run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * fix: scope the late run alert to its workspace, acknowledge it on edit, toggle and delete Recovery acknowledges alerts by resource alone, so the resource now carries the workspace. Editing, toggling or deleting a schedule clears its streak and a disabled or deleted one never chains a run on time, so those handlers acknowledge its open alert after committing. Past the 1000-slot cap, last_missed_at falls back to the detection time. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 * refactor: raise the late run alert like the other critical alerts Drops the recovery, the workspace-scoped resource and the acknowledgement on edit, toggle and delete: the alert now fires once per streak with no resource and is acknowledged from the alerts feed, as the trigger and job failure alerts are. The FOR SHARE read stays, so a push that rolls back across the flow path's retries still alerts once. Notes in openapi that past 1000 misses in one late run the count is a floor and the time approximate. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LJ9wpjWp2YgLUSqt1Ai5d6 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
17448c97d3 |
count trigger suspend, resume and discard (#11405)
* feat(telemetry): count trigger suspend, resume and discard Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor: shorten the telemetry disclosure to one line per category Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: correct the resume branch comments and note the pre-commit fire count Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: name feature adoption in the telemetry disclosure Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore: update ee-repo-ref to 9855e1b7a43a0a33e04f8accf1c497af3fd9b139 This commit updates the EE repository reference after PR #834 was merged in windmill-ee-private. Previous ee-repo-ref: 1d5b128ec956c156fe549cf099ba0dbc1b6467bf New ee-repo-ref: 9855e1b7a43a0a33e04f8accf1c497af3fd9b139 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
f367eaf6d0 |
feat: run turns in several flow chat conversations at once (#11202)
* feat: run turns in several flow chat conversations at once Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep finished turns finished and cached chats current in the flow chat pool Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: attribute a turn's rows by job id as well as sequence Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: count only real stream updates and retry the job-id read Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep a chat that holds an unsent draft, and take one back when its first turn is withdrawn Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: ignore a stale running-turn snapshot, keep a withdrawn chat's draft, poll after clean stream ends Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: follow the turn running now when the listing named one already over Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep replacement turns and SSE fallback moving * fix: keep replacement turn handoffs active * fix: preserve unread badge line height * fix: settle local fallback handoffs * fix: settle refused turn handoffs * fix: scope turn handoffs to conversation * fix: drop stale turn handoffs * refactor: move the queued message and 409 handling into per-conversation turns Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address cubic's review of the parallel flow chat turns Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: clear a stale failure on refresh, and tighten the docs and test waits Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: recover running rows past the first page, and drop the failure a re-read disproves Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: check the running-turn query at compile time, and narrow what a refresh clears Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: settle a failed turn only from an answer that turn wrote Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: settle a failed turn from its own answer, and only while it is still the failure shown Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: drop a failure whose answer arrived even when a newer turn owns the error Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: free an answered failure whatever the turn that started meanwhile is doing Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: drop a rows read that a turn outran, rather than merging it under newer messages Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: drop a rows read whose conversation was left and opened again Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: hand over a file still being read when its composer goes Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: count a drop's routing as work in flight, so its file is handed over too Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: hold the send until every file a conversation is owed has landed Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: keep a panel mounted per conversation instead of handing its draft over Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep the withdrawn chat whose composer was written in, not the empty one Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep the chat in front of the reader when both withdrawn composers were written in Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep a retry's own run arguments when a turn elsewhere refuses it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: name panels apart across pools, and read a flow's inputs when its chat is built Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e68ff0d969 |
feat: add tree view to the schedules page (#11360)
* feat: add tree view to the schedules page Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep schedule job previews visible inside tree folders Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep job preview loading while hovered and close it on scroll Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep schedules outside u/ and f/ in the tree view Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: add tree view to the trigger list pages (#11400) * feat: add tree view to the trigger list pages Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor: let a TreeViewState own the tree view setting Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: drop the doubled bottom border at the end of a tree Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
651a6b01f5 |
chore(main): release 1.819.0 (#11364)
* chore(main): release 1.819.0 * Apply automatic changes --------- Co-authored-by: rubenfiszel <275584+rubenfiszel@users.noreply.github.com>v1.819.0 |
||
|
|
cedd6dc901 |
fix: hold interpolated references and captures to the token path scopes (#11391)
* fix: hold interpolated references and captures to the token path scopes Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: let a resource read cover its own linked secret variable Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: resolve policy-granted app upload resources on the viewer's rls Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: cover multi-secret linked variables and keep capture paths out of refusals Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
163a4ffa4e |
fix: scope flow resume to its workspace and minting to the job's run (#11392)
* fix: scope flow resume to its workspace and minting to the job's run Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: name the lineage columns resume minting checks Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
14a2619ad2 |
fix: gate batch rerun on job read access, scope started_at to workspace (#11387)
* fix: gate batch rerun on job read access and scope started_at lookup to the workspace Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: assert batch rerun denial comes from the read gate Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
3eaf2888c0 |
fix: scope workspace dependencies create to the path workspace (#11385)
* fix: scope workspace dependencies create to the path workspace Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: state the workspace_id must-match contract in the spec and struct Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
5e59cefd1f |
fix(frontend): apply operator write locks from the session's operating workspace (#11395)
Claude-Session: https://claude.ai/code/session_01VAj4mmm2YThZLVkivrgsbb Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
f4dcaf3e45 |
fix: list only the paths the caller can read in path autocomplete (#11388)
* fix: list only the paths the caller can read in path autocomplete Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: bound the path autocomplete cache by total path count Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
ec6ec1b06f |
fix: judge IPv4 embedded in IPv6 and pin the object storage test connect (#11389)
* fix: judge IPv4 embedded in IPv6 and pin the object storage test connect Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: let the public-only object store client reach the egress proxy Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: refuse private IP literals and the proxy host in the public-only store client Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: refuse the egress proxy as a target whether named or an IP literal Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
22b5a1cd62 | fix: keep test panel controls off the args form in debug mode (#11382) | ||
|
|
d76a962331 |
feat: add hub sync button to the resource types tab (#11375)
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
d9c7d71f07 |
fix: redesign run not found page and fix switching to the right workspace (#11374)
* fix: redesign run not found page and clear stale not-found on workspace switch Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor: use design-system Button for workspace rows on run not found page Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
60ef82196e |
fix: keep smtp_clicktracking_off when syncing instance config (#11372)
* fix: keep smtp_clicktracking_off when syncing instance config Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: name smtp regression test after what it guards Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
990a726409 |
feat: add provenance claims to job OIDC tokens (#11369)
* feat: add provenance claims to job OIDC tokens and mark preview sub Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: require a flow or script job's version to belong to its path for deployed Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: derive app script paths server-side and test job provenance in CE Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: count an app script as deployed only when a deployed app run stamped it Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: state the deployed condition for the preview sub prefix Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: keep the plain OIDC sub for previews by users who can write the path Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: refuse OIDC tokens to previews by users who cannot write the path Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor: keep OIDC token issuance unchanged, leaving provenance to the claims Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor: name the root job's trigger claim root_trigger_kind Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to 421cf2a8b4f98b421e93c0fc7c1c378314a66e50 This commit updates the EE repository reference after PR #831 was merged in windmill-ee-private. Previous ee-repo-ref: 7acd384875deba4b01a502e628a153b11c82eecb New ee-repo-ref: 421cf2a8b4f98b421e93c0fc7c1c378314a66e50 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
3838cd6ee0 |
fix: stop the schedule enabled toggle from showing unsaved changes (#11390)
* fix: keep schedule enabled toggle from reading as unsaved changes Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: only fold the enabled toggle into the baseline when it is deployed Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: skip the enabled revert once the drawer moved to another schedule Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
893e64f630 |
fix: only restart a flow on a version of its own path and workspace (#11376)
* fix: only restart a flow on a version of its own path and workspace Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: pin cross-workspace restart version rejection Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
90f9e59321 |
fix: only let a job's own token claim run lineage (#11367)
* fix: only let a job's own token claim its lineage on the run endpoints Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: drop an unclaimable run lineage instead of refusing the run Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: only let a job's own token run its workflow-as-code tasks Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
47525b211a |
perf: shrink the module graph that gates first paint in dev (#11373)
* perf: shrink the module graph that gates first paint in dev * docs: drop the stale synchronous-icons claim on the import card * fix: replay search opened before its modal loads, guard lazy icons * fix: only intercept search before load where the modal mounts * perf: mount app-shell modals on first open and keep monaco off the shell * perf: load the icon map on first read, not at module evaluation * fix: report stale chunks with a reload toast, guard the home page against monaco |
||
|
|
649c43e7c1 |
fix: run an AI agent tool on the worker its own tag selects (#11370)
* fix: run an AI agent tool on the worker its own tag selects Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: run a tagged agent tool inline when this worker serves its tag Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: give an inline agent tool a job token of its own Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: report a lost tool wait to the model and cancel tools on agent timeout Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
d7a61de23f |
fix: never double-process a slow canceled flow in the zombie sweep (#11368)
* fix: leave a slow canceled flow to its live worker and bound its requeues Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF * fix: give a canceled zombie flow a longer grace instead of guessing its worker Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF * fix: retry a canceled zombie flow's forced completion until it lands Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF * fix: claim a canceled zombie flow without waiting on its runtime row Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
c2d8997549 |
fix: complete a canceled flow whose worker died between two steps (#11366)
* fix: complete a canceled flow whose worker died between two steps Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF * fix: complete only the stranded canceled flow and let its parent process it Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF * fix: requeue a stranded canceled flow for a worker to complete its cancel Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF * fix: keep a requeued canceled flow's start time Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzGsxKey5g3kycKwGmpKNF --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
65cba2dbb7 |
perf: advance a flow step with one v2_job_status update (#11357)
* perf: advance a flow step with one v2_job_status update Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: state what advance_flow_status returning None means Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep the merged flow advance identical for rows without a status row or with a malformed status Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FbA4shbpCnRsUxir1GEfm * docs: note the JSON null invariant behind the empty-path no-op Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FbA4shbpCnRsUxir1GEfm --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
bebd762194 |
perf: complete a job in one statement on the common path (#11355)
* perf: complete a job in one statement on the common path Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: take completion locks in one order on every path Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep a losing zombie completion from touching its wac parent Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: leave a flow's ping alone when a step completes during its cancel Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: probe only this test's completion for the lock wait Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: stamp a wac child's kept duration when its completed row exists Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TJMJSJ2bDhYh7Shoh78Yyb * perf: leave the parent ping out of completions with no flow to ping Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TJMJSJ2bDhYh7Shoh78Yyb * docs: note that the two completion statements must stay in step Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TJMJSJ2bDhYh7Shoh78Yyb * chore: update ee-repo-ref to 7a256cf353db7cf64a60a09fa0de7f3a8b27f626 This commit updates the EE repository reference after PR #830 was merged in windmill-ee-private. Previous ee-repo-ref: 497137acb65e521568d46f3cbe1d66359f7f87ec New ee-repo-ref: 7a256cf353db7cf64a60a09fa0de7f3a8b27f626 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
e2be584ca5 |
fix: let custom workspace error handlers send email with the instance SMTP (#11365)
* fix: let custom workspace error handlers send email with the instance SMTP Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: state what the error handler email allowlist guarantees Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
3974bbeac6 |
chore(main): release 1.818.0 (#11294)
* chore(main): release 1.818.0 * Apply automatic changes --------- Co-authored-by: rubenfiszel <275584+rubenfiszel@users.noreply.github.com>v1.818.0 |
||
|
|
ca8a04a869 |
fix: allow results access inside nested functions in input transforms (#11358)
* fix: allow results access inside nested functions in input transforms Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: decode escaped bracket step ids and test deferred fetch errors Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: let quickjs decode bracket step ids and match quoted forms Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: prefetch results read through spread syntax Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: keep prefetched bracket literals on a single line Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: decode prefetch step literals as data and skip unparsable ones Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: run the results prefetch outside the expression scope Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: keep the transform expression a zero-arg iife after prefetch Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
da866c5eff |
feat: alert on and optionally cancel jobs stuck on unserved tags (#11354)
* feat: alert on and optionally cancel jobs stuck on unserved tags Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: group stranded jobs in sql and recheck each job before canceling Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: guard stranded-job alerts and cancels against outages and pickups Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: count priority tags as served and retry lost stranded-job cancels Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: send one daily stranded-jobs alert that can be muted Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: cover native retries and finish lost stranded-job cancels unconditionally Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
30bb62cd25 |
fix(cli): stub the API client over its real exports in tests (#11363)
* fix(cli): stub the API client over its real exports in tests Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(cli): mock the API client once and dispatch to per-suite stubs Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
9a1c6e5081 |
feat: let a workspace withdraw operator schedule and trigger writes (#11226)
* feat: let a workspace withdraw operator schedule and trigger writes Operators can create, edit and delete schedules and triggers today through the API, CLI and MCP, while the operator_settings flags beside them only hide those pages. An admin who wants operators to see what is scheduled without letting them change it cannot express that. Add manage_schedules and manage_triggers as enforced settings, gated at the schedule handlers and at the generic TriggerCrud routes so every trigger kind is covered by one check. They name capabilities operators already hold, so they are granted unless withdrawn, and absence has to mean "never configured" rather than a value. The read coalesces to true; the update endpoint merges into the stored jsonb with the two fields as Option<bool>, so an omitted key keeps what is stored. operator_settings is git-synced as a whole object, so a settings file written before these keys existed reaches the endpoint on every pull, and a serde or SQL default of either polarity would turn that pull into a silent withdrawal or restoration. The rights are read through a per-process cache, so withdrawing one publishes a notify_event that drops the entry on every replica. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Dsf6VC4MVLisiEoeQkgbr4 * feat: enforce operator write rights on the router and in the UI Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: close the capture gap and gate the trigger editors' write actions Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: gate acl writes and the native trigger drawer behind manage rights Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: refuse operator writes with 403 and gate sharing at the drawer Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * perf: resolve identity in the operator write gate only for writes Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: gate the suspended-jobs actions and stop the route check refusing reads Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: explain the empty-state create button when operator writes are withdrawn Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: audit operator settings changes and fold path writes into native rows Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: open locked editors read-only and group the operator settings Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: skip email and azure lookups on editor open while triggers are locked Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: state each operator-rights rationale once in comments Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: address CI review findings on operator write rights Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep capture move gated and skip it in the builders while locked Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep admin and operator exclusive when setting a workspace role Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor: use the shared section component for operator settings groups Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Ruben Fiszel <ruben@windmill.dev> |
||
|
|
a1abb36d9f |
fix: relock importers on their own tag, not the bare dependency tag (#11359)
* fix: relock importers on their own tag, not the bare dependency tag Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: pin the tag of relocks triggered by a changed import Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
53a5cfd17a |
perf: skip job-start pings and checkpoint read for short non-WAC jobs (#11356)
* perf: skip job-start pings and checkpoint read for short non-WAC jobs Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: pin the wac language gate alongside is_wac_v2 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep the start memory sample for jobs shorter than one poll tick Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
e14da5c6bc |
feat(bedrock): add OIDC role assumption as a fourth auth mode (#10936)
* feat(bedrock): add OIDC role assumption as a fourth auth mode Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn * fix(bedrock): gate the OIDC cache correctly and assume the role once per job Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn * refactor(bedrock): check the OIDC region before minting a token Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn * fix(bedrock): keep OIDC session names collision-resistant, gate the copy on EE Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn * fix(bedrock): check the OIDC region before reusing cached credentials Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn * fix(bedrock): clear assumed-role sessions when AI settings change invalidate_ai_request_cache_for_workspace cleared AI_REQUEST_CACHE only, so a workspace's AI settings edit reset one cache and left the assumed-role sessions keyed on the old config in place until STS expired them. Also name the region requirement in the credentials-check hint, so following it does not land on the OIDC path's region guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YCVb91fZp3dqM14KRPzoEn * chore: update ee-repo-ref to de73db2bacfdc3eaa2e63b1827178bc198d54e5c This commit updates the EE repository reference after PR #770 was merged in windmill-ee-private. Previous ee-repo-ref: c43dab1e69b1cb3f685e6df07bff634dc2a0b734 New ee-repo-ref: de73db2bacfdc3eaa2e63b1827178bc198d54e5c Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
94e4fb1c84 |
fix: carry labels when deploying variables, resources and folders (#11222)
* fix: carry labels when deploying variables, resources and folders across workspaces Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: clear a folder's labels in the target when the source has none Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: pin that the frontend deploy adapter carries variable labels Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d3d5392917 |
feat: add an options field to the postgresql resource (#11223)
* feat: add an options field to the postgresql resource Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: keep a literal plus in postgres connection string parameters Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: pass postgres options to trigger connections Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: read DATABASE_URL options the way sqlx does Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: include postgres options in databaseUrlFromResource Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f183bd43fb |
chore: retire the standalone lsp and multiplayer images from examples, drop lsp/Dockerfile (#11341)
* chore: retire the standalone lsp and multiplayer images from examples, drop lsp/Dockerfile Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * chore: drop the unbuilt DockerfileMultiplayer, document running the LSP from windmill-extra Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(examples): ecs terraform destroys cleanly and gives private instances no public ip Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(examples): give windmill-extra on ecs a WINDMILL_BASE_URL for multiplayer auth, address review Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(examples): make the ecs example upgrade cleanly from the standalone lsp/multiplayer stack Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(examples): name the extra target group by prefix so create_before_destroy can replace it Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(examples): note the brief editor-socket gap when upgrading the ecs example Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(examples): the debugger stays off after the ecs upgrade unless enabled Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
3bd89e92d8 |
feat: collapse the fork members setting and show its state in a badge (#11220)
* feat: collapse the fork members setting and show its state in a badge Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: show the fork members badge next to the title, only when on, like other section badges Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8caf414301 |
fix(multiplayer): a malformed frame from one client no longer exits the server (#11353)
* fix(multiplayer): don't drop client messages during cold-start token verification
`wss.on('connection')` awaits `verifyToken()` before `setupWSConnection()`
attaches the 'message' listener. On a cold process that await includes the
first `/api/debug/jwks` fetch (~30ms on ECS). A y-websocket client sends sync
step 1 the instant the socket opens, and `ws` drops messages emitted with no
listener attached, so that step 1 was lost and never answered with step 2 —
the client's provider never became `synced`.
Buffer messages from the moment the connection is accepted and replay them, in
order, once `setupWSConnection()` has installed its handlers. Rejected
connections drop the buffer and close with the same 4401/4403 codes as before.
Also prefetch the public key at startup when WINDMILL_BASE_URL is set. That is
insurance, not the fix: a connection arriving before the prefetch resolves
still relies on the buffer.
Adds `npm test` in multiplayer/ (node:test, no docker or backend needed) with a
fake JWKS endpoint that answers with a delay, which holds the cold window open
and makes the race deterministic; wired into the existing test_extra CI job.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(multiplayer): cap what an unauthenticated peer can buffer pre-auth
Review follow-up.
The pre-auth buffer was unbounded: `ws` sets no `maxPayload` here and the JWKS
fetch has no timeout, so a peer that never authenticates could stream frames
into memory for as long as `verifyToken` was stalled. Cap it at 32 frames /
1 MiB — a real client only has sync step 1 and its first awareness update in
flight there — and close 1009 past that, dropping what was buffered.
A socket closed during verification (by the peer, or by that cap) is no longer
handed to setupWSConnection: it would be added to `doc.conns` with a 'close'
listener that can never fire.
The startup prefetch's .catch was dead code — getPublicKey() logs its own
failures and resolves to null rather than rejecting.
Test helper: pin REQUIRE_SIGNED_MULTIPLAYER_REQUESTS and BASE_INTERNAL_URL so an
ambient value cannot turn the rejection tests into false passes; bind the JWKS
server on port 0 instead of a released probe port, and retry the spawned server
on EADDRINUSE; destroy still-delayed JWKS responses on teardown, since
server.close() waits for in-flight requests.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(multiplayer): gate the JWKS response instead of delaying it
Review follow-up.
The cold window was held open by a 1500 ms delay on the fake JWKS response, but
that timer started when the startup prefetch reached the fake server, not when
the client sent its first frame. A slow enough machine could load the key before
the client connected, and the race test would then pass without ever exercising
the buffer — a false pass.
The fake JWKS server now parks every response until the test calls release(), so
the server provably holds no key while the client is sending. The race test
releases only after both frames are written to the socket, and asserts the
server has not logged the key as loaded at that point; the flood test never
releases until after the cap has closed the connection.
What is left to wall-clock time is 250 ms for bytes already written to the socket
to cross loopback into an otherwise idle server, rather than a window that had to
cover process startup, connect and handshake.
Also drops the prefetch precondition from the forged-token and flood tests so
each test still maps to one behaviour. Suite runs in ~1.1s instead of ~5.3s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(multiplayer): survive a malformed frame instead of exiting the process
`setupWSConnection`'s message handler decoded whatever an authenticated peer
put on the wire with no guard: `decoding.readVarUint`,
`syncProtocol.readSyncMessage` and `awarenessProtocol.applyAwarenessUpdate` all
throw on input they cannot parse, `ws` re-emits a listener's exception on the
process, and server.mjs installs no `uncaughtException` handler. One bad frame
from one client therefore killed the whole multiplayer server, taking every
other document and every other client with it.
Catch decode/apply failures, log the document, the client address and the error
message (never the payload), and close only the offending connection with 1007
"invalid frame payload data". Frames that arrive once a connection is no longer
OPEN are ignored, so the replay of the pre-auth buffer stops at the first
refusal instead of applying the rest.
docker/entrypoint-extra.sh made that outage permanent: on a service exit it
logged a bare PID and then `wait`ed on the rest, so the container stayed up with
a dead service and the health checks in front of it — which probe the LSP — saw
nothing wrong. It now names the service that died, stops the others through the
same shutdown path SIGTERM uses, and exits non-zero so the orchestrator replaces
the container. The "no services enabled" branch still sleeps.
Tests: multiplayer/test/malformed_frame.test.mjs covers four malformed payloads
from an authenticated client and one replayed out of the pre-auth buffer,
asserting the 1007 close, a live server process, an undisturbed bystander and a
real edit still propagating. docker/test_entrypoint_extra.sh runs the real
entrypoint in a container with stub services.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(multiplayer): prove the replayed malformed frame really is buffered pre-auth
Assert the server has not yet logged the loaded key when the frame is written,
and give it the same in-flight margin as the cold-start tests before releasing
the JWKS response, so the frame provably goes through the replay path rather
than landing on an already-authenticated connection.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* ci(extra): run the entrypoint supervision tests in publish_extra
The multiplayer unit tests already run there; the entrypoint test needs only
docker and the checkout, so run it in the same job, before the image build.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(multiplayer): handle WebSocket protocol errors and bound the shutdown
Two crash paths of the same class as the malformed-frame one, from review.
`ws` fails a frame it cannot parse at the protocol level — an unmasked frame
from a client, a reserved opcode, a bad RSV bit — inside its Receiver, before
the application 'message' handler ever sees it, and `receiverOnError` ends with
`websocket.emit('error', err)`. With no 'error' listener that is an unhandled
EventEmitter error, so it exited the process just as a malformed payload did.
(A raw socket error such as ECONNRESET does not: ws 8.21.3's `socketOnError`
swallows those.) Add the listener on the accepted socket, before authentication
so the pre-auth window is covered too, and one on the server.
Log messages now go through `describeError`, which collapses whitespace and
truncates, so nothing that reaches an error message can forge or flood a log
line.
`stop_services` ended in a bare `wait`. On the `docker stop` path dockerd
provides the deadline; the "a service died" path signals itself, so a service
that is wedged or slow to honour SIGTERM would hold the container open
indefinitely — the state that path exists to prevent. Bound it: SIGTERM, wait
SHUTDOWN_GRACE_SECS (10 by default), then SIGKILL the stragglers by name.
Tests: multiplayer/test/socket_error.test.mjs (authenticated and pre-auth
illegal frames, asserting a live process and continued service), a
SIGTERM-ignoring stub scenario in docker/test_entrypoint_extra.sh, and that
harness is now bounded throughout — `timeout -k` on foreground runs, a watchdog
around the backgrounded ones, and an optional outer timeout on `docker run`.
`--entrypoint bash` so the documented windmill-extra:test override runs the
harness instead of the image's real entrypoint. The stubs publish a readiness
marker and the dying one waits for them, removing a startup race in the harness.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(multiplayer): never log peer bytes, and keep a fatal server error fatal
Two review findings on the previous commit, both mine to answer for.
`describeError` collapsed whitespace, which is not enough. An error message is
not always a fixed string: `applyAwarenessUpdate` runs `JSON.parse` on the
peer's bytes and V8 quotes ~30 bytes of the offending input back verbatim, ESC
included, so a peer could put terminal escapes and forged content into a log
line. Strip everything outside printable ASCII instead, and say so where the
comment previously claimed the messages were fixed strings.
`wss.on('error')` was worse than the crash it replaced for one case: `ws`
forwards the HTTP server's errors there, so a failed listen (EADDRINUSE) was
logged and the process then exited 0 — a clean shutdown as far as anything
upstream could tell. It now sets a non-zero exit code. Setting `process.exitCode`
rather than calling `process.exit()` keeps the log line from being truncated.
`openClient` in the test helpers now records the socket error it was already
swallowing, so a failed connection reports its cause instead of surfacing as a
bare `waitFor` timeout.
Tests: a malformed awareness frame whose state is `x\x1b[2J OWNED THE LOG` added
to the payload table, with every case now asserting exactly one refusal line and
no control characters in it (1 fail before, 0 after, 3 runs); and a server that
cannot listen must exit non-zero (1 fail before, 0 after, 3 runs). 14/14 on 5
consecutive runs.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(multiplayer): make the exit-status helper robust to spawn and stdio races
Review follow-ups on runMultiplayerServerUntilExit.
Wait for 'close', not 'exit': 'exit' fires when the child terminates, which can
be before its stdio pipes are drained, and the caller reads the output. On the
EADDRINUSE path the child writes one line and exits immediately after, which is
exactly the shape that loses it.
Listen for 'error' too. A child that fails to spawn emits neither 'exit' nor
'close', so the promise would never settle and the SIGKILL guard could not help.
Report whether the guard fired, rather than leaving the caller to infer it from
the exit signal: `signal` is null for every child exit on Windows, so a server
that hung after the listen error would have looked like one that exited on its
own. The test asserts on that flag instead.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(multiplayer): assert the refusal line itself, and name hasSyncType for what it takes
Two review nits on the test helpers.
The `doc="..."` assertion searched the whole server log, where CONNECT and
DISCONNECT also name the document, so it would have passed even if the refusal
stopped naming anything. Every assertion about the refusal is now made against
the refusal line, which the test already isolates, and it also checks the peer
is named.
`hasKind` took a sync sub-type but was named as if it took any message kind, and
the two families overlap numerically (`syncStep1 === messageSync === 0`), so a
caller passing the wrong one got a silently wrong answer. No runtime check can
tell aliased numbers apart, so the fix is the name: `hasSyncType`, with the
overlap spelled out where the constants are declared.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(multiplayer): make the pre-auth tests prove which path they took
The replay test could not establish that its frame went through the pre-auth
buffer: a frame delivered after setup, into the live handler, produces the same
close code and the same refusal line, so if the in-flight margin were ever
missed the test would quietly become a duplicate of the main-loop cases rather
than fail. server.mjs now logs REPLAY when, and only when, it replays a buffered
pre-auth message — worth having on its own, since that path only runs when a
client beat the JWKS fetch on a slow-starting instance — and the test asserts on
it. Removing that log line turns the test red, which is the point.
The socket-error pre-auth test gated on `jwks.requests >= 1`, which the startup
warm-up already satisfies, so it proved nothing about the offender. What makes
it the pre-auth case is that the JWKS response stays parked for the whole test;
it now asserts the server never logged CONNECT, which is exact.
`killedByTimeout` was set before the kill, so a child that exited on its own just
before the timeout — with 'close' still pending on the stdio drain, the very
window this helper waits for — would have been reported as killed. It now claims
the rescue only when there was a live process to signal.
14/14 on eight consecutive runs.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(multiplayer): document the frame-recording contract in openClient
ws hands every frame over as a Buffer under the default binaryType, text frames
included, so recording them as Uint8Array is lossless for both. Worth stating:
ws 7 delivered text frames as strings, where new Uint8Array(string) would have
been a silent zero-fill, and the difference is not visible at the call site.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(multiplayer): pin the close code for an illegal frame, and trim comments to the 4-line rule
Both socket-error tests waited for a close and never checked what it was, so an
abrupt 1006 teardown would have passed while the comment beside the payload
claimed 1002. `ws` sends 1002 for an unmasked frame in both the authenticated
and pre-auth cases, confirmed over repeated runs; that is now a named constant
asserted in each test, mirroring malformed_frame.test.mjs. Changing the expected
value turns both red.
The comments added by this branch also ran past the four lines AGENTS.md allows,
and several justified the change to a reader rather than stating the invariant.
Condensed to the invariant, at the site that would break it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
||
|
|
1d119e6b25 |
fix: correct the tool controls and name field of a nested AI agent (#11221)
* fix: hide tool controls a nested AI agent cannot use Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor: drop the tool navigation prop an agent tool now implies Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: ask for a tool name on every kind of agent tool Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: keep enabled_tools reachable on a linked nested agent tool Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: name the tool kinds the header's name field actually renders Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: state the tool-name rule without listing the kinds it covers Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4fd7d62bd0 |
feat: rework the db manager: native grid, tabs, sql editor, joined columns (#11340)
* feat: replace ag-grid in the db manager table viewer with a native grid Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: add user-managed data, diagram and sql editor tabs to the db manager Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: add schema autocomplete to the db manager sql editor and polish its layout Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: add joined foreign key columns and draggable tabs to the db manager Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: move db manager tabs on drop instead of mid-drag Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: open followed foreign keys in a new db manager tab Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: scroll large tables, drop stale joins and return to the last tab on close Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: tolerate the db manager tabs going away while switching data table Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep saved joined columns while metadata loads, test joins in every dialect Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: put the db manager grid on the input surface in dark mode Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: address db manager review nits on joins, boolean keys and sql quoting Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: keep the dark border color on the left pinned column edge Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: tone down the db manager tree menu icons Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: write json and jsonb values from the db manager through a text cast Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: dim unrelated diagram tables less on hover Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: match json columns as text when deleting a db manager row Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: compare postgres json columns as text in exact db manager filters Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: pick the columns the db manager grid shows Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: toggle every db manager column from one checkbox Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: word joined columns as a view in the db manager Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: keep data table roles and access visible when unavailable, with the reason Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: keep data table and instance roles visible in settings when unavailable Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: hide a db manager column from its header menu Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: format db manager columns with a unit, significant digits and color rules Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: drop the before/after hints from the db manager unit picker Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: write the euro after the amount and keep units off non-numeric values Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: color rule presets, bold and italic, layered and reorderable rules in the db manager Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: decimals, thousands separator, compact notation and alignment in db manager column formats Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat: compact db manager format controls, a notation toggle group and a reset button Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: db manager format and filter nits Keep the value's scale when decimals are auto, read boolean color-rule conditions as booleans, let a rule's text color reach foreign-key links, close the formatter when the columns picker opens, and filter BigQuery complex columns through TO_JSON_STRING. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
2e5f6d0e76 |
fix(frontend): keep the session when the persisted workspace is stale (#11344)
* fix(frontend): keep the session when the persisted workspace is stale A single-use login link signs a different account in while `workspace` in session/localStorage still names the previous account's workspace. `loadUser` read that workspace, got no membership back, threw `Not logged in` and logged the brand-new session out, landing on `/user/login?rd=...`. A missing membership says nothing about the session, so forget the workspace and continue down the no-workspace path, which logs out only when `globalWhoami` shows the session itself is gone. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(frontend): confirm the session before the workspace-picker redirect `loadWithoutWorkspace` fired the `/user/workspaces?rd=…` navigation before awaiting `globalWhoami`, so when the session turned out to be gone the logout read whichever URL the race had left in `page.url` and carried the picker as its `rd`. Ask first, then redirect. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(frontend): stop the loadUser comments overclaiming what they know Neither comment can promise what it stated: `getUserExt` collapses every failure into `undefined`, so the branch cannot tell a real non-membership from a transient one, and `loadWithoutWorkspace` throws on any `globalWhoami` rejection rather than only on a dead session. Say what each call actually answers about. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
e8c3f9514e |
fix(multiplayer): don't drop client messages during cold-start token verification (#11343)
* fix(multiplayer): don't drop client messages during cold-start token verification
`wss.on('connection')` awaits `verifyToken()` before `setupWSConnection()`
attaches the 'message' listener. On a cold process that await includes the
first `/api/debug/jwks` fetch (~30ms on ECS). A y-websocket client sends sync
step 1 the instant the socket opens, and `ws` drops messages emitted with no
listener attached, so that step 1 was lost and never answered with step 2 —
the client's provider never became `synced`.
Buffer messages from the moment the connection is accepted and replay them, in
order, once `setupWSConnection()` has installed its handlers. Rejected
connections drop the buffer and close with the same 4401/4403 codes as before.
Also prefetch the public key at startup when WINDMILL_BASE_URL is set. That is
insurance, not the fix: a connection arriving before the prefetch resolves
still relies on the buffer.
Adds `npm test` in multiplayer/ (node:test, no docker or backend needed) with a
fake JWKS endpoint that answers with a delay, which holds the cold window open
and makes the race deterministic; wired into the existing test_extra CI job.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(multiplayer): cap what an unauthenticated peer can buffer pre-auth
Review follow-up.
The pre-auth buffer was unbounded: `ws` sets no `maxPayload` here and the JWKS
fetch has no timeout, so a peer that never authenticates could stream frames
into memory for as long as `verifyToken` was stalled. Cap it at 32 frames /
1 MiB — a real client only has sync step 1 and its first awareness update in
flight there — and close 1009 past that, dropping what was buffered.
A socket closed during verification (by the peer, or by that cap) is no longer
handed to setupWSConnection: it would be added to `doc.conns` with a 'close'
listener that can never fire.
The startup prefetch's .catch was dead code — getPublicKey() logs its own
failures and resolves to null rather than rejecting.
Test helper: pin REQUIRE_SIGNED_MULTIPLAYER_REQUESTS and BASE_INTERNAL_URL so an
ambient value cannot turn the rejection tests into false passes; bind the JWKS
server on port 0 instead of a released probe port, and retry the spawned server
on EADDRINUSE; destroy still-delayed JWKS responses on teardown, since
server.close() waits for in-flight requests.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(multiplayer): gate the JWKS response instead of delaying it
Review follow-up.
The cold window was held open by a 1500 ms delay on the fake JWKS response, but
that timer started when the startup prefetch reached the fake server, not when
the client sent its first frame. A slow enough machine could load the key before
the client connected, and the race test would then pass without ever exercising
the buffer — a false pass.
The fake JWKS server now parks every response until the test calls release(), so
the server provably holds no key while the client is sending. The race test
releases only after both frames are written to the socket, and asserts the
server has not logged the key as loaded at that point; the flood test never
releases until after the cap has closed the connection.
What is left to wall-clock time is 250 ms for bytes already written to the socket
to cross loopback into an otherwise idle server, rather than a window that had to
cover process startup, connect and handshake.
Also drops the prefetch precondition from the forged-token and flood tests so
each test still maps to one behaviour. Suite runs in ~1.1s instead of ~5.3s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|