Commit Graph
839 Commits
Author SHA1 Message Date
hugocasaandClaude Fable 5.1 03a8a0edf2 fix: read the provider for url-token repos, await the origin before defaulting, and visit unchecked repos last
The maintenance pass sorted repositories with no recorded check first on
the premise that they cost nothing, but a token-in-URL remote on a host
that is not GitLab is probed every pass and never records a check, so it
held the head of the list ahead of the tokens that expire. Such
repositories now sort last.

The card decided its delivery defaults before the origin lookup landed,
so a freshly picked GitLab repository never got webhook delivery; the two
lookups are awaited together. The resource editor offers to replace a
token only where it is held, not in a fork that borrows it, and the
replace flow refuses a URL it cannot parse instead of keying the token to
it. Attaching a stored credential to a commit-hash probe now requires
admin, matching the installation credential beside it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C1xHmkxuxYb1GYvth1BS75
2026-09-07 18:19:17 +02:00
hugocasaandClaude Opus 5 b683b14da7 fix: gate the credential pass budget on the features that use it
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C1xHmkxuxYb1GYvth1BS75
2026-09-07 11:50:36 +02:00
hugocasa 03c7069f03 Merge remote-tracking branch 'origin/main' into gitlab-app-integration-exploration
# Conflicts:
#	backend/ee-repo-ref.txt
#	backend/windmill-api-workspaces/src/workspaces_extra.rs
2026-09-07 11:18:08 +02:00
fce635d3c4 feat: guest app execution mode, a role that takes no seat (#10929)
* feat: guest app execution mode, a fourth role that takes no seat

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: make the guest grant a server-minted label, not a declarable scope

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* chore: pin ee-repo-ref to the guest session companion branch

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: close the relabel hole, guest embed tokens, read-path switch, custom-path entry

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: guest tokens are not rescopable and guest embed tokens keep the sentinel

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: guest-derived tokens share one constraint set; gate sign-in on guest discovery

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: the label alone governs a guest; refuse guests with accounts; unserialize discovery

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: guest discovery fails closed; SAML aborts if the guest cookie write fails

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* refactor: enforce the guest switch once at the auth door; sign-in for a guest of another app

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: guest app-mode decided once at the on-behalf resolver; clear a stale guest session before offering another app's sign-in

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: a guest may use anonymous apps; await the stale-session logout; trim comments

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: a guest's path confinement waits for the app's mode, so anonymous apps stay open to it

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: guest target survives http (Lax cookie), rides SAML RelayState; tell account holders on arrival

* fix: a guest uses an anonymous app as itself; S3 uploads confined by app mode

* fix: a guest upload needs an app policy; a missing app does not skip the confinement

* fix: guests are gated on the Enterprise plan server-side; pin ee-repo-ref

* fix: the guest plan gate fails closed on non-enterprise builds; settings report the effective switch

* fix: guest controls read the plan, not the key; gate the guest tests on the features they need

* docs: tighten the guest session invariant comments

* feat: 100 free guests per 30 days, then a quarter seat each on Enterprise and a hard cap elsewhere; superadmin guest list; refusals reach the page

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: the cap is exact, an account ends a guest session at the door, popups close, and guest mode survives the CLI round trip

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* feat: a superadmin switch over guests for the whole instance; the pre-existing-user flag keeps its meaning

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: drop the dead guest-access helper, name the instance setting once, guests tab states, CE save order

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: a guest app path is refused at the mint if it could widen the scope; the instance toggle waits for its reload

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: guests stop at the launched-by-me job grant; canonical app paths at the mint and discovery; the toggle ends on the stored value

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: only the scope grammar's own characters bar an app path from guests, refused at deploy as well as at the mint

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: the deploy-time guest path guard checks the destination of a rename and refuses a leading slash

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: a workspace rename keeps the guest switch; the rename guard reads the deployed mode under the row lock

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* fix: guest_activity follows a workspace rename and goes with a workspace delete

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* chore: pin ee-repo-ref to the state-bound guest target

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* chore: pin ee-repo-ref; the guest cookie is never cleared by a callback

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* docs: the workspace-scoped guest_activity delete moves an instance-wide count; assert the mint records the guest

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* test: the seeded allowance is a day old, so only the mint can write today's guest_activity row

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BayTppRCstWX6qTf3LMco5

* chore: update ee-repo-ref to 1a10132e4f3cb442c7d0c2cf6e5d92d150bf6e07

This commit updates the EE repository reference after PR #769 was merged in windmill-ee-private.

Previous ee-repo-ref: 32841072aa396bff91d30bd91854fa348cb3c439

New ee-repo-ref: 1a10132e4f3cb442c7d0c2cf6e5d92d150bf6e07

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-09-04 22:47:28 +02:00
hugocasa b94f9ac7c6 Merge remote-tracking branch 'origin/main' into gitlab-app-integration-exploration
# Conflicts:
#	backend/ee-repo-ref.txt
#	frontend/src/lib/components/ResourceEditor.svelte
2026-09-04 18:05:23 +02:00
hugocasaandClaude Opus 5 09280aefd6 fix: keep credential status out of exports and clear stale webhook warnings
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C1xHmkxuxYb1GYvth1BS75
2026-09-03 17:35:35 +02:00
hugocasaandClaude Opus 5 bf528ef072 fix: recreate a missing webhook from credential maintenance
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C1xHmkxuxYb1GYvth1BS75
2026-09-03 13:57:53 +02:00
hugocasaandClaude Opus 5 5eecf0e8b1 fix: refuse to finish a check whose repository has been repointed
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C1xHmkxuxYb1GYvth1BS75
2026-09-03 13:35:20 +02:00
hugocasaandClaude Opus 5 ed3cb9acb5 fix: resolve the check marker's repository from its path, not a stored url
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C1xHmkxuxYb1GYvth1BS75
2026-09-03 13:22:41 +02:00
hugocasaandClaude Opus 5 1b21708261 fix: bound the credential maintenance pass and gate the gitlab picker on a license
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C1xHmkxuxYb1GYvth1BS75
2026-09-03 12:29:15 +02:00
hugocasaandClaude Opus 5 d472193e5b feat: add retention cleanup for the otel_traces table (#10949)
* feat: add retention cleanup for the otel_traces table

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLhUaCPpLRAa29rSZDjS28

* fix: vacuum otel_traces and badge its retention setting EE

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLhUaCPpLRAa29rSZDjS28

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-02 19:41:46 +02:00
hugocasa bb071efd7c feat: receive gitlab push webhooks for instant git sync pull 2026-09-02 13:24:16 +02:00
hugocasa b237aed906 fix: gate credential maintenance on enterprise and alert on stalled renewal 2026-09-02 12:31:58 +02:00
hugocasa 9750bbb3d8 fix: strip server-owned credential status and correct expiry copy 2026-09-02 12:17:29 +02:00
hugocasa b5431e2e6e feat: track and rotate gitlab git-sync repository tokens 2026-09-02 11:58:36 +02:00
Ruben Fiszelandwindmill-internal-app[bot] cfcfe298dd feat(ai-chat): make reusable skills ai_skill resources you select per workspace (#10914)
* feat(ai-chat): make reusable skills ai_skill resources you select per workspace

* chore: pin the ee ref to the skill telemetry counters

* fix: address review findings on skill authoring, import and migration

* fix: enforce skill selection in read_skill and stop imports clobbering resources

* feat: carry format_extension from the hub into synced resource types

* fix: let an edit set or clear a resource type's format_extension

* fix: regenerate the sqlx cache and close the review round findings

* fix: close the round-2 findings on folder ACLs, cached sync and truncation

* refactor: make the skills migration non-destructive and use design-system inputs

* fix: close the round-4 findings on folder owners, startup sync and truncation

* fix: clear obsolete extensions, guard folder owners, and report skipped skills

* fix: honor explicit-null extensions and report same-type migration conflicts

* fix: scope skill actions to the committed workspace and paginate the listing

* fix: keep the drawer scoped to the live workspace and surface truncation

* fix: discard a skills refresh for a workspace the chat has left

* chore: update ee-repo-ref to 6efe7a73c745c2e1377a34498523c00d89010a3d

This commit updates the EE repository reference after PR #764 was merged in windmill-ee-private.

Previous ee-repo-ref: 55998c142bc72edd08532748af1974b16035658d

New ee-repo-ref: 6efe7a73c745c2e1377a34498523c00d89010a3d

Automated by sync-ee-ref workflow.

---------

Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-09-01 12:51:27 +00:00
aa4a6ffd66 fix: track outstanding service log files on the rows themselves (#10894)
* fix: track outstanding service log files on the rows themselves

Adds `log_file.indexed_at` so the service log ingest can read outstanding rows
instead of walking a cursor over `log_ts`. A row registered after the pass had
gone by its minute was skipped for good, and no ordering fixes that — an arrival
sequence fails the same way, since a row can take a lower value and commit after
a higher one has moved the cursor past it.

The migration marks existing rows with a sentinel; the first pass returns the
ones the old cursor had not reached to the queue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EPAP96jJNYpPQ8bpxZcU1C

* [ee] refactor: drop the claim/confirm phase from the service log ingest queue

Two states are enough: a row is outstanding or it is marked. The migration no
longer creates the index for the claim sentinel, and the sqlx cache loses the
two queries the event-time cursor used.

* [ee] fix: make re-indexing a service log file idempotent

Corrects the `init_last_log_file_sent` note: a rewritten row keeps the
`indexed_at` it had, so one the indexers already took is not offered again.

* [ee] fix: let a rebuild take the rows it covered out of the ingest queue

Adds the query that releases them; the index layout stays v4.

* [ee] fix: index the lookup a rebuild releases rows by

A rebuild takes rows out of the queue by the file it read out of the store, which
is the one lookup that arrives without a `log_ts`. The primary key is
`(hostname, log_ts)`, so nothing covered it and each batch scanned every
outstanding row — worst in exactly the state a rebuild follows. Verified at 50k
outstanding rows: sequential scan becomes an index scan.

Also records `log_file.indexed_at` in the schema reference.

* [ee] fix: treat a state handed back without its line count as behind

* [ee] fix: give the converted state a line count

* [ee] fix: keep the converted cursor from being rewound by the rebuild

* [ee] fix: inherit the legacy cursor from one source, not field by field

* [ee] fix: count a file's lines against the buffer before reading it

* [ee] fix: bound the row buffer on what it holds, not on reported counts

* [ee] fix: settle the upgrade from the store rather than from event time

* [ee] docs: describe the conversion's second half as it now works

* [ee] refactor: settle the upgrade with one rebuild instead of reconciling

The migration records existing rows as done rather than marking them with a
sentinel: the indexer puts back what the old cursor had not reached on its first
pass, which is the only place that cursor's position is known.

* [ee] fix: repair the rows the old cursor skipped instead of recording them as done

The migration marks pre-existing rows with a sentinel again, so the indexer can
tell them from rows registered since and put the window's worth back on the queue.

* [ee] fix: keep a source file whole in one partition

* [ee] revert the file-atomic partition change

* [ee] fix: dedupe the public reads, and repair an index without a cursor

* [ee] fix: repair an index whose cursor is gone, and keep what the repair found

* [ee] fix: seed a pass from both axes of what a rebuild recovered

* [ee] fix: settle the cursor on what the store holds, not on what was read

* [ee] fix: an empty rebuild must not claim ground it has not covered

* [ee] test: pin the cursor a rebuild settles on

* chore: update ee-repo-ref to bc0c7051585194474078b6c1941a3fb73893d9e5

This commit updates the EE repository reference after PR #755 was merged in windmill-ee-private.

Previous ee-repo-ref: 328f5a90afeae9c683bf3294f0d9eb293a3e1a92

New ee-repo-ref: bc0c7051585194474078b6c1941a3fb73893d9e5

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-08-31 14:06:29 +02:00
d91ee4614a feat: day-partition the service log index and expire whole chunks (#10893)
* feat: day-partition the service log index and expire whole chunks

The service log index becomes one tantivy index per UTC day. The substance is
in windmill-ee-private#753; this side carries the EE ref and moves the log
indexer writer instead of cloning it, because sealing a chunk takes sole
ownership of its tantivy writer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EeUmYWCJeaaHutiHLfZBQX

* fix: do not adopt the superseded watermark after an explicit index clear

A clear asks for the retention window to be read again, and a watermark says
it already has been — and the v3 copy in object storage is kept for rollback,
so it outlives the local one the clear removes. Both copies of that watermark
are now read and the newer wins, for the same reason the v4 one is taken from
the store when it is ahead: a replica that lost the lock keeps a local file
frozen where it stopped while the store went on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EeUmYWCJeaaHutiHLfZBQX

* fix: delete a day's raw files at its checkpoint, and rebuild whole days

Bumps the EE ref for windmill-ee-private#753.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EeUmYWCJeaaHutiHLfZBQX

* fix: make an interrupted rebuild detectable, and pin the rebuild floor

Bumps the EE ref for windmill-ee-private#753.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EeUmYWCJeaaHutiHLfZBQX

* fix: keep the rebuild marker in the object store, not on local disk

Bumps the EE ref for windmill-ee-private#753.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EeUmYWCJeaaHutiHLfZBQX

* fix: two more routes to a partial index being accepted as complete

Bumps the EE ref for windmill-ee-private#753.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EeUmYWCJeaaHutiHLfZBQX

* fix: trust a local chunk only when the tracker vouches for it

Bumps the EE ref for windmill-ee-private#753.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EeUmYWCJeaaHutiHLfZBQX

* chore: condense the stale-chunk guard's doc to the four-line limit

Bumps the EE ref for windmill-ee-private#753.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EeUmYWCJeaaHutiHLfZBQX

* chore: update ee-repo-ref to 17ef439b087b400889ff19109be9d2c810142278

This commit updates the EE repository reference after PR #753 was merged in windmill-ee-private.

Previous ee-repo-ref: 3e79901b4742906d2285dd943e24fac0f735f199

New ee-repo-ref: 17ef439b087b400889ff19109be9d2c810142278

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-08-30 07:16:09 +02:00
815de49e23 feat: make the service log retention period an instance setting (#10889)
* feat: make the service log retention period an instance setting

Service log retention was a hardcoded 14 days with no override, unlike job retention. It
becomes the `service_log_retention_secs` global setting (env `SERVICE_LOG_RETENTION_SECS`,
default unchanged at 14 days), reloaded on change like the other retention settings.

The constant becomes `DEFAULT_SERVICE_LOG_RETENTION_SECS` and every reader goes through
`service_log_retention_secs()`, so the `log_file` sweep, the object-storage orphan scan, the
columnar store's compaction and pruning, the retrieval clamp and the search index's trim
window all follow the configured value.

Loaded outside `initial_load`'s `server_mode` guard: a dedicated indexer trims the search
index to a window derived from this value and is not a server.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN

* fix: never let a non-positive service log retention expire every log

Every service log cutoff is `now - retention`, so a `0` or negative window puts the cutoff
at or after `now` and the next sweep reads the whole history as expired — deleting the
`log_file` rows and their object-storage files irreversibly.

`0` is reachable two ways now that the window is configurable: it is what an operator types
by analogy with the job retention period sitting directly above it, where `0` does mean keep
forever; and `SecondsInput` writes a `0` into a field that was merely focused, so saving the
Jobs panel is enough. Service logs always have a window, so clamp an unusable value back to
the default in the accessor every reader already goes through. The upper bound is where
`chrono::Duration::seconds` panics, which would abort the sweep that reads it.

The settings field rejects a non-positive value rather than silently correcting it, and its
description now names the database rows too — they are swept on every instance, including
one with no object storage configured.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN

* fix: address review findings on the service log retention setting

- Bound the monitor's `log_file` sweep. Every process rotates a log file a minute, so lowering
  the retention can make one ordinary setting change expire millions of rows; the unbounded
  `DELETE ... RETURNING` materialized all of them, and their deletion futures, in a single
  tick. Batched like the settings-page cleanup on the same table.
- Make the retention atomic private and give it one writer, so a value that would expire every
  service log cannot reach a cutoff by any path, and say so in the log when one is rejected
  rather than falling back silently.
- Cap the retention at a century. The previous ceiling only bounded `TimeDelta` construction,
  while consumers compute `now - retention`, which panics past year 262143, and build a
  Postgres interval that overflows well before the old cap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN

* fix: cap an oversized service log retention instead of shortening it

The two unusable directions were landing on the same fallback, so configuring a retention
above the ceiling silently produced 14 days — deleting logs the operator had asked to keep
for longer. Too large now caps at the maximum, which preserves that intent; only a
non-positive value, which would expire everything and has no upward reading, falls back to
the default.

Also bound the `log_file` drain to ten batches per pass: `monitor_db` runs under a 600s
timeout that cancels every maintenance future in the same `join!` and reports a critical
error, so a backlog large enough to need batching has to drain across ticks, the way the
neighbouring sweeps already do. The settings field carries the upper bound too, and the
superseded query's offline entry is dropped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN

* fix: route the new log-file registration cutoff through the retention accessor

`send_log_files_to_object_store` arrived on main while this branch was open and reads the
retention directly. The atomic behind it is private now, so it goes through the accessor like
every other consumer — which also means the cutoff it uses to skip registering already-expired
files follows the configured retention rather than a fixed two weeks.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN

* fix: say why every mode loads the service log retention setting

A worker registers its rotated log files against the retention cutoff, so the comment naming
only the indexer no longer covers why the setting sits outside the `server_mode` guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN

* fix: file service log retention under Monitoring, not Jobs

Service logs are the Windmill processes' own logs — every process rotates and registers its
own, no job involved — so the Jobs panel was grouping by the shape of the widget rather than
by the subject. It sits under Monitoring now, beside the Indexer panel that holds the other
service-log window.

Its own section rather than inside that panel: the panel is badged EE, while this governs the
database sweep that runs on every instance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN

* chore: update ee-repo-ref to a6e3533b26195918a17fea58646f71d2bbcde288

This commit updates the EE repository reference after PR #752 was merged in windmill-ee-private.

Previous ee-repo-ref: 1d93da24bd166b9a5a5cc204034a1d35ffc88474

New ee-repo-ref: a6e3533b26195918a17fea58646f71d2bbcde288

Automated by sync-ee-ref workflow.

* feat: say on the service logs page where the logs actually are

The retention number alone does not tell an operator what it governs, and the answer differs
by instance. Two states are worth calling out because they are the ones where retention does
not mean what it looks like:

Without instance object storage, each process keeps its files on its own disk. The page lists
what every host wrote, since the rows are in the shared database, but can only open the files
of the replica serving the request, and a host's files go with it when it is replaced.

With object storage but "Delete logs from s3 periodically" off — the backend default, since
uploads are gated on a store existing while deletions are gated on that toggle — expiring a
log removes the row and the local file and leaves the uploaded copy behind for good.

The retention field itself now names every copy it covers and says that full-text search
reaches back at most that far, and less when the indexer's own window is shorter.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN

* fix: describe raw log files as the transient copy they became

Retiring the raw files landed while this was being written: the indexer now deletes each one
as soon as it is ingested, and the log viewer rebuilds a file from the columnar store once the
raw copy is gone. So the durable copy is the store, and warning that an uploaded file is kept
forever when periodic s3 deletion is off only holds where no indexer runs to ingest it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN

* chore: point ee-repo-ref at the EE compile fix

EE main does not build on its own: extracting the index-window expression and adding a fourth
copy of it landed in separate PRs that never conflicted textually. windmill-ee-private#756 is
the one-line fix; this pins it so CI has a tree that compiles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-08-29 19:39:14 +02:00
Ruben FiszelandClaude Opus 5 c8172480b0 fix: register every rotated service log file exactly once (#10891)
* fix: register every rotated service log file exactly once

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QTxWg4Sx57UodA9RpFMJm

* chore: refresh sqlx cache for the log_file watermark query

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QTxWg4Sx57UodA9RpFMJm

* fix: skip service log files past the retention cutoff on catch-up

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QTxWg4Sx57UodA9RpFMJm

* refactor: name the shutdown flush for what it does and scope its doc claims

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QTxWg4Sx57UodA9RpFMJm

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 11:53:21 +02:00
7c1a785f75 feat: serve service log retrieval from a columnar parquet store (#10886)
* feat: always write service log files as json so they index structured

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016ijGGPCFkYhVzisFexAHYx

* feat: serve service log retrieval from a columnar parquet store

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016ijGGPCFkYhVzisFexAHYx

* feat: shrink the service log index to the per-host count it still serves

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016ijGGPCFkYhVzisFexAHYx

* fix: reclaim the superseded service log index on upgrade

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016ijGGPCFkYhVzisFexAHYx

* fix: address review findings in the service log store

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016ijGGPCFkYhVzisFexAHYx

* chore: update ee-repo-ref to ad9e899dfd2ee4e3d18ecf06d016f821968c5a83

This commit updates the EE repository reference after PR #751 was merged in windmill-ee-private.

Previous ee-repo-ref: 6ad4064f9d58d83612b42b4ec870384994d64bcb

New ee-repo-ref: ad9e899dfd2ee4e3d18ecf06d016f821968c5a83

Automated by sync-ee-ref workflow.

* fix: address review nits on the service log store

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016ijGGPCFkYhVzisFexAHYx

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-08-29 09:51:59 +02:00
b72ccc3593 fix: key build artifact caches on a runnable's inline modules (#10819)
* fix: key build artifact caches on a runnable's inline modules

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: seal the cache-key base and skip prebundling multi-file bun scripts

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: tighten cache-key invariant comments and name the retained-artifact residual

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: version the build artifact keyspace so pre-fix artifacts are abandoned

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: namespace the artifact cache by keyspace version instead of the hash preimage

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: namespace module-bearing artifacts instead of versioning the whole keyspace

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: pin the cache-name base seal and name the retained-artifact residual

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: bump ee ref for agent-worker module resolution fix

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: align agent-worker module resolution with the worker for previews by hash

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: drop calculate_hash imports left unused by artifact_cache_name

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: update ee-repo-ref to 2d6c66b32f20d9605c6a677727473ab66fcc8a87

This commit updates the EE repository reference after PR #743 was merged in windmill-ee-private.

Previous ee-repo-ref: efce983cae3d53175bbb286a10205a2a360c2a9e

New ee-repo-ref: 2d6c66b32f20d9605c6a677727473ab66fcc8a87

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
2026-08-28 16:40:22 +02:00
hugocasaandClaude Opus 5 ffdf17ef8d fix: force HTTP router rebuild on trigger-change notification (#10849)
* fix: force HTTP router rebuild on trigger-change notification

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: coalesce http trigger change events into one forced rebuild

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: retry the coalesced http router rebuild when it fails

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: mark http routers stale when a forced rebuild fails

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: keep the router invalidation across an in-flight rebuild

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 08:23:25 +02:00
hugocasaandClaude Opus 5 2906504125 feat: add instance setting to mute zombie job restart alerts (#10813)
* feat: add instance setting to opt out of zombie job restart alerts

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: preserve explicit false for default-on boolean instance settings

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor: invert zombie restart alert setting to a mute flag

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 22:36:21 +02:00
Ruben FiszelandClaude Opus 5 4b406e37c0 fix: size the ephemeral job token to the job timeout it must serve (#10804)
* fix: size job token to the premium cloud job timeout

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: give the job token setup headroom and drop dead MAX_TIMEOUT_DURATION

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: cap job token setup slack so self-hosted tokens stay at 7d

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 09:30:08 +00:00
Ruben FiszelandClaude Opus 5 3c8e4b43fd fix: resolve a script path to its new version as soon as the lock lands (#10794)
* fix: resolve a script path to its new version as soon as the lock lands

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011H5ygpzQHkPeYsjiP9GzBy

* fix: tell MCP script deploy callers to stop polling on a lock error

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011H5ygpzQHkPeYsjiP9GzBy

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 11:54:51 +02:00
1f59841a67 feat: add WM_ROOT_WORKSPACE, the closest dev or prod workspace of a job (#10776)
* feat: add WM_ROOT_WORKSPACE, the closest dev or prod workspace of a job

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HBMJwogo6jJ1P55uvB3YpF

* fix: do not cache a failed root-workspace lookup, and sweep on fork create

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HBMJwogo6jJ1P55uvB3YpF

* fix: shorten the agent-worker root-workspace TTL and pin the sweep wiring

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HBMJwogo6jJ1P55uvB3YpF

* chore: update ee-repo-ref to a2fa58e5301d3865dd06ad73519e20ba7a5af0f0

This commit updates the EE repository reference after PR #736 was merged in windmill-ee-private.

Previous ee-repo-ref: 07a9d26a79a403ae27c48abd508a6699f2c87c49

New ee-repo-ref: a2fa58e5301d3865dd06ad73519e20ba7a5af0f0

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-08-20 15:40:22 +02:00
Ruben FiszelandClaude Opus 5 de90e44650 docs: document DATABASE_URL_FILE in the env var help (#10773)
Claude-Session: https://claude.ai/code/session_0125f1Vwj7pR9oY8NCxLXwtW

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 12:35:35 +02:00
Ruben FiszelandClaude Opus 5 0258f3f81b perf: unblock workers before the API router is built (#10711)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 16:43:51 +02:00
Ruben FiszelandClaude Opus 5 c9ddddda1b log the settings a failed read left unapplied (#10709)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 21:15:33 +02:00
53eb94659b feat(telemetry): extend feature-usage tracking beyond AI features (#10681)
* feat(telemetry): extend feature-usage tracking to long-tail features

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: describe telemetry as product feature usage rather than AI usage

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(telemetry): trim disclosure copy and drop unused pick origin

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(telemetry): count trigger fires per run and key hub picks from hub data

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(telemetry): slugify hub keys and order both writers' upserts

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(telemetry): key native trigger adoption by service so it matches fires

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref for native trigger adoption fix

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(telemetry): move feature-usage collection into the ee crate

* docs: point feature-telemetry at the moved registry and rust writer

* docs: correct the trigger-fire gate comment to match measured step counts

* docs: put the private-build caveat on the verification step

* chore: update ee-repo-ref to f079db9e7962a413b349c4ff8036080894f30771

This commit updates the EE repository reference after PR #725 was merged in windmill-ee-private.

Previous ee-repo-ref: 055adb80416f9339c9a28ae7fbaeadad30d74959

New ee-repo-ref: f079db9e7962a413b349c4ff8036080894f30771

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-08-14 18:50:38 +02:00
Ruben FiszelandClaude Opus 5 30f5d2e766 perf: declare a settings pass instead of reading one setting at a time (#10698)
* perf: read global_settings once per settings-load pass

`initial_load` reads several dozen settings back to back, one
`SELECT value FROM global_settings WHERE name = $1` each: 50 serialized round
trips before a worker is ready, 32 before a server is. On localhost that is
~20ms and invisible; against a real database it is 50x the RTT per process
start, which `EXIT_AFTER_N_JOBS` turns into a per-job cost.

`with_global_settings_snapshot` reads the whole table (12 rows on a typical
instance) into a tokio task-local, and `load_value_from_global_settings`
serves from it. Scoping it to the task is what keeps the single-setting
reload paths correct: a `notify_global_setting_change` event for one key runs
outside any scope and still reads the database, so a live settings change
reaches a running worker as before. Agent workers hold an HTTP connection
with no snapshot to take and are unchanged.

`load_smtp_config` and `reload_custom_tags_setting` had their own inline
copies of the same query; they go through the shared loader so they land in
the snapshot too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: state the snapshot contract on the reader and the query

`load_value_from_global_settings` is called from ~10 crates and one of them
writes a setting then immediately re-reads it through
`reload_custom_tags_setting`; say on the function itself that a scope, when
one is installed, serves the read and leaves `db` unused.

The query comment claimed the table is a handful of rows. It is not bounded
that way: `workspace_dependencies_map_rebuilt:<workspace_id>` adds a row per
workspace and never removes it. Those dynamically named rows are also why the
snapshot fetches the whole table instead of the wanted names, so state that
as the reason.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: bound the settings snapshot and keep it out of two reads

Three review findings, all real:

The snapshot fetched the whole table, which is not bounded by the settings
that exist: `workspace_dependencies_map_rebuilt:<workspace_id>` adds a row per
workspace with no cleanup path, and no settings pass reads one. It now fetches
only statically named rows, and reads of a `<prefix>:<id>` name skip the
snapshot and go to the database. Correctness does not rest on that naming
convention — a colon-free dynamic name would simply be in the snapshot and
still answered correctly — only the bound does.

A snapshot query that failed inside an enclosing snapshot awaited the body
bare, so its reads were served by the outer snapshot rather than falling
through as documented. The task-local carries an explicit bypass state and the
failure path scopes it.

`reload_jwt_secret_setting` decided whether to generate-and-upsert the JWT
secret from a snapshot-served read, so a replica booting alongside another
could overwrite the secret it had just generated and invalidate its tokens.
That read goes through the new `load_value_from_global_settings_fresh`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep the snapshot query on the primary-key index

`name NOT LIKE '%:%'` bounded the rows returned but not the work: a leading
wildcard cannot use the index, so Postgres read every row anyway. Against
50k dynamically named rows it plans as a seq scan of 516 buffers whether or
not seqscans are enabled — and worker connections disable them, so the plan
was one the query shape forbade rather than one the planner chose.

`name = ANY($1)` over an explicit list plans as a bitmap index scan, 7
buffers, bounded by the listed names rather than by table size. That list is
also exactly the set the snapshot may answer from, so a name outside it falls
through to the database instead of reading as unset: listing a setting is a
performance choice, never a correctness one, which is what keeps the list
safe to maintain by hand.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: declare a settings pass instead of reading one setting at a time

Replaces the prefetch-list snapshot with a pass the call sites build
themselves. `SettingsPass` collects the reads `initial_load` will make as
`(name, applier)` pairs, fetches them together, then replays the appliers in
declaration order.

Declaring is what makes the batch exact. The same `if server_mode` /
`if *CLOUD_HOSTED` / `cfg` branches that used to guard a read now guard a
declaration, so the fetch asks for what this process needs and nothing else,
and there is no list of setting names to keep in sync with anything.

Ordering is preserved end to end: appliers run in the order they were
declared, and non-setting work in the middle of the sequence keeps its place
as a step, so nothing moves and nothing runs twice. Steps that need several
settings at once take them together.

The batch distinguishes three states where a per-setting read only ever
produced two at a given call site:

- a value,
- genuinely unset, which several settings must see in order to restore a
  default when the setting is cleared,
- could not be read, which must leave the in-memory value alone. Collapsing
  this into "unset" would let one failed query reset workspace fairness and
  the queue caps across a cluster.

Over HTTP the reads go out together rather than sequentially, so an agent
worker's settings load costs one round instead of ~36, with no new endpoint.
A setting an agent may not request still resolves to unset, as the
per-setting call returned for it.

`reload_*` keeps working per setting for the notify path, sharing its apply
half with the pass. The wrappers no caller was left using are dropped.

worker startup: 50 queries -> 2 (the batch, and jwt_secret which stays its
own read so the pass cannot sit between reading it absent and upserting a
replacement over another replica's).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: run the pass's non-setting steps in declaration order too

Review round found the settings pass had a gap: the reads were declared but
the work interleaved between them still awaited inline, so it all ran before
`pass.run` applied anything.

`manage_audit_partitions` therefore saw `AUDIT_LOG_RETENTION_DAYS` at its
compile-time default rather than the configured value, and dropped every
partition past that default. An instance keeping 30 days on CE lost the
14-to-30-day band on startup and on every full-reload tick. The
`STORE_AUDIT_LOGS_S3` export anchor had the same cause: the gate read `false`
before the setting applied, so an env-var-enabled export never anchored and
its first tick skipped the rows committed before it.

`action` exists so a step keeps its place in the sequence; every remaining
inline await is now one, which fixes both and leaves no phase where a read
can observe a value the pass has not applied yet.

Two more from the same round:

A batch that fails as a whole now falls back to per-setting reads. Skipping
every applier preserves known-good state on a reload tick, but a starting
process has none, and would have run on compile-time defaults until the next
full reload twelve hours later.

`FORCE_RUBY_REPOS` is honored again: the batched url-list path parsed without
the `FORCE_` check its per-setting counterpart applied, so the override was
silently dropped. `load_setting_value` never had one, so the third helper was
never affected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: declare the object-store and worker-config steps in the pass too

Two awaits were left running ahead of `pass.run`, so the settings they read
were still at their compile-time defaults.

The object-store reload is the one that matters: an AWS OIDC store mints its
first token against an issuer built from `BASE_URL` (`oidc_ee.rs`), and with
`OTEL_ENVIRONMENT` set nothing loads that before this pass does, so the store
signed with the unset default, left `OBJECT_STORE_SETTINGS` empty and fell
back to the ten-second retry while startup carried on.

`reload_worker_config` calls `store_pull_query`, which reads the workspace
fairness knobs. It happened to converge because the enabled flag re-stores the
query when it changes, but it was reading defaults on the way there.

Both are steps now, which is also what the earlier fix should have covered:
the only await left outside a step is `pass.run` itself.

Also from the same round: `fetch_settings_batch`'s doc comment had been
stranded on the helper inserted above it, and the batch-failure fallback
re-ran the same reads on an agent worker, where the batch already is the
per-setting read.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: point the setting-loader docs at functions that still exist

`reload_setting` went with the other wrappers no caller was left using, but
two doc links still referenced it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: decide the jwt secret in sql so the read can be batched

`reload_jwt_secret_setting` generated a secret whenever its read came back
absent or unparseable, and upserted it unconditionally. Two replicas booting
against an empty row therefore each installed their own and rejected each
other's tokens, and the same happened on a running cluster whenever the row
was deleted or set to a non-string. Keeping the read next to the write kept
the window narrow but never closed it, and it was the reason this one setting
could not go through the settings pass.

`get_or_create_jwt_secret` puts the decision in the statement instead:

    INSERT ... ON CONFLICT (name) DO UPDATE SET value = EXCLUDED.value
    WHERE jsonb_typeof(global_settings.value) <> 'string'
    RETURNING value

First writer wins, a usable secret is never overwritten, and an empty
RETURNING is how a caller learns another process's secret stands. The `WHERE`
also keeps a normal startup from writing at all, which matters because
`notify_global_setting_change` fires on every write to this table and an
unconditional upsert would have made each start trigger a cluster-wide reload.

Because the statement decides rather than the caller's read, a stale value is
harmless and `jwt_secret` is now an ordinary declaration. Worker startup is a
single batch round.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep a failed read from dropping a FORCE_ override or clearing a setting

Two ways a read that did not succeed was being treated as an answer.

A `FORCE_` override used to be checked before the read, so a failed read
could not affect it. Moving that check into the parser put it behind a value
arriving, and a failed read skips its applier, so a forced private registry
fell back to the public index and a forced `settings.xml` was deleted from
disk by the Maven step that follows it. Forced settings are declared as steps
with no read now: the override outranks the database, so there is nothing to
fetch and nothing to lose when a fetch fails.

The setting loaders were passing `v.ok().flatten()` to their appliers, which
turns a database error into "unset". Most appliers ignore `None`, but
`apply_tag_per_workspace_workspaces` clears the workspace whitelist with it,
making every workspace eligible for per-workspace tags, and
`apply_fork_workspace_tag_append_fork_suffix` stores `false`. Both are also
reached from the notify handlers, so a blip during a reload changed routing
for the cluster. They take `?` now, as the code they replaced did by leaving
the error arm empty, and the other five are converted with them so an applier
that later grows a `None` branch cannot inherit the problem.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: route hub_api_secret through the FORCE-aware declaration

`HUB_API_SECRET` lives in an `ArcSwap` rather than an `Arc<RwLock<_>>`, so it
could not use `option_setting` and was declared by hand with a bare `setting`
plus `parse_option_setting_value` — which is exactly the path that skips the
`FORCE_` handling, so a failed read still dropped `FORCE_HUB_API_SECRET`.

The rule now lives in `option_setting_with`, which takes the store closure and
leaves `option_setting` a wrapper over it, so a setting held in something other
than an `RwLock` reaches it too rather than having to reimplement it.

The three remaining hand-written parses are `parse_setting_value`, which has no
`FORCE_` handling to miss: `load_setting_value` never had the check either.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 18:03:33 +02:00
Ruben FiszelandClaude Opus 5 6d03784d4b fix: keep non traffic-serving processes out of coordinated restarts (#10694)
* fix: key server_heartbeat row on hostname so restarts reuse one row

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: trim announce_server_started doc to the durable constraints

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: only traffic-serving processes take part in coordinated restarts

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: name every non traffic-serving mode in the restart-gate comments

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: narrow the restart-gate comments to claims that hold

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 15:10:30 +02:00
Ruben FiszelandClaude Opus 5 22eadab67d perf: resolve the worker external IP in the background (#10697)
* perf: resolve the worker external IP in the background

`run_workers` awaited `external_ip::get_ip()` — an HTTPS GET to
hub.windmill.dev — before spawning any worker, so every worker process paid
that round trip before its first job pull. Measured on a CE debug build it was
120-450 ms of a ~200-500 ms startup, and behind a firewall the call does not
fail fast: it burns its whole 5 s connect timeout, on every process start. That
cost is per-job under EXIT_AFTER_N_JOBS.

The value is informational (it is only written to `worker_ping.ip`, which the
workers list displays so users can whitelist the address), so nothing needs to
wait on it. It now resolves into a process-wide cache off the startup path, and
`WORKER_EXTERNAL_IP` supplies it explicitly for deployments that know their
egress address or have no egress at all.

Until it resolves the ping carries no IP, which `insert_ping_query` now
COALESCEs so a reclaimed row keeps the address the previous process wrote
instead of being blanked. The main loop reports the IP as soon as it lands
rather than on the next periodic tick, so a short-lived process still records
it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep unknown worker IPs out of the whitelist alert

Review follow-ups:

- `WhitelistIp` filtered only the `'unretrievable IP'` sentinel, so the `'NO IP'`
  one a pending or failed lookup now leaves in the row would be offered as an
  address to whitelist. It filters both.
- Register `WORKER_EXTERNAL_IP` in `ENV_SETTINGS` so operators can confirm from
  the instance settings view that it took effect.
- The worker tracked whether it had reported the IP by re-reading the cache
  after each ping rather than remembering what the ping carried, so a lookup
  landing mid-ping marked it reported without it reaching the row. The value is
  read once and threaded through `insert_ping` / `update_worker_ping_full`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: report a sentinel IP once the lookup has definitively failed

Keeping the previous process's address on a reclaimed `worker_ping` row is right
while the lookup is still in flight, but not once it has failed: the row would
advertise an address nothing has confirmed, and the whitelist alert would offer
it. A failed lookup now reports `UNKNOWN_IP`, leaving NULL to mean "in flight".

`WORKER_EXTERNAL_IP` is rejected when longer than the `varchar(50)` column
rather than panicking the worker on its initial ping, which is a hard failure.

Adds the regression guard for the `ON CONFLICT` semantics: reverting to
`ip = EXCLUDED.ip` would compile and blank every reclaimed row.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep the agent initial ping acceptable to older servers

An agent worker routinely runs against a server of a different version, and one
predating the background lookup rejects an initial ping carrying no IP — which
`run_worker` turns into a panic, so a newly upgraded agent would crash-loop
against it. The not-resolved-yet case goes over the wire as the sentinel
instead, and the server maps it back so a reclaimed row still keeps its address
while resolution is pending.

Also documents `ip` as the one conditional exception to `insert_ping_query`'s
"only `started_at` and `jobs_executed` survive a restart", and adds
`WORKER_EXTERNAL_IP` to the README env-var table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: deliver the resolved IP to servers that only take it at registration

A server predating the background lookup applies `ip` from the initial ping
only, and ignores it on the periodic ones. An agent registering before its
lookup resolves would therefore keep the sentinel forever on such a server,
where it used to report its real address. It registers a second time once the
address is known, skipping that when the address is still unknown, when the
server is reached over SQL and needs no second registration, or once a job has
run, since registering clears the row's current job.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: re-register the resolved IP even after a job has run

Gating the second registration on "this process has not run a job yet" meant an
agent that pulled queued work before its lookup resolved never delivered the
address to a server that only takes one at registration. No job of the worker is
in flight where that runs, so the gate bought nothing beyond the last job's id,
which the next job refills.

Documents the two cases where WORKER_EXTERNAL_IP stops being an optimisation and
becomes the only way to report an address: an agent against such a server, and a
process shorter-lived than the lookup.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* revert: drop the WORKER_EXTERNAL_IP escape hatch

Supplying the address by hand skips the hub lookup, which is not something to
make easy. Resolving it in the background is what keeps it off the startup path;
opting out of it is a separate decision this does not need to take.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: distinguish an IP never established from one that could not be retrieved

`NO IP` was doing double duty: the column default for a row whose lookup has not
resolved, and the marker for one that failed. An operator reading the workers
list could not tell "not resolved yet" from "this instance cannot reach the
hub", and the latter is the actionable one. A failed lookup now reports
`unretrievable IP`, which is also what it reported before the lookup moved off
the startup path.

That leaves `NO IP` meaning only "no address established", which is what an
agent sends while its lookup is in flight and what the server maps back to
"unresolved" — so the wire sentinel no longer collides with the failure marker,
and an agent delivers the failure to a server that only reads an IP at
registration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 14:27:21 +02:00
Ruben FiszelandClaude Opus 5 2fcce4526a feat: add EXIT_AFTER_N_JOBS worker mode for environment cleanup (#10671)
* feat: add EXIT_AFTER_N_JOBS worker mode for environment cleanup

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address review findings on the EXIT_AFTER_N_JOBS worker mode

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address round-2 review findings on EXIT_AFTER_N_JOBS

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address round-3 review findings on EXIT_AFTER_N_JOBS

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: bound WORKER_SUFFIX length and document the same-worker drain

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: validate the assembled worker name length

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 06:14:55 +02:00
hugocasaandClaude Opus 5 c09de594b6 feat: version resource values with history, diff and restore (#10596)
* feat: version resource values with history, diff and restore

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: record resource versions in a trigger so direct writes are covered

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: show the selected version's value and tighten history write access

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf: gate resource version recording in trigger WHEN clauses

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: clear a resource's past versions, and address review nits

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: restore the displayed version and keep author attribution on pooled writes

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: scope history to the selected workspace and gate clearing on ownership

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: gate restore on write access and clearing on the signed-in workspace

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(frontend): share the version-history row between script and resource drawers

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf: trim resource version history in the monitor sweep, not on write

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(frontend): match the script versions drawer shell for resource history

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf(frontend): highlight version values instead of mounting monaco

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(frontend): match the script drawer's code preview presentation

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: rank version trim in one windowed pass instead of a correlated delete

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(frontend): treat the newest version as current by position

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf: gate the resource version trim to an hourly sweep

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: unnest the version row action and correct the trim cadence docs

* perf: cap the history listing and use sets for reference lookup

* feat: warn when a resource is written more than 60 times a minute

* fix: lower the resource write advisory to 20 per minute

* fix: discard stale history loads and never diff against an unread value

* fix: correct the write advisory boundary and document the eviction lock

* fix: read history and the live value from one snapshot

* refactor: read the drawer's diff baseline from versions, not the live resource

* fix: open the history drawer with no version selected

* fix: disarm the clear confirmation and clear the pane when the selection moves

* fix: explain the missing diff and drop a guard that can no longer fire

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 21:32:18 +02:00
Ruben FiszelandClaude Opus 5 c59b60c729 fix: keep the same_worker pin when a suspend ends without approval (#10552)
* fix: keep the same_worker pin when a suspend ends without approval

A disapproved or timed-out approval gate hands the flow back through the
UpdateFlow channel with unrecoverable = true. That flag means "the previous
step's worker died", and it is read by six sites. Five of them happen to want
what it does here, but continue_on_same_worker and continue_with_runners do
not: the worker that ran the approval step is alive, so unpinning the error
handler and routing it by tag breaks the ./shared contract of a same_worker
flow and can land it on a worker group that cannot run it — the same defect
#10551 fixed for the three producers that hand back a live flow.

Replace the boolean with StepFailureKind so the suspend producer can say
"worker alive, but this failure is not the module's to handle" instead of
overstating a worker death. The failed module's error policy is deliberately
still bypassed: the failure is recorded against the step the gate was holding
back, which never ran, so its retry would re-open the gate and its
continue_on_error would skip it outright (verified: the gated step is marked
Failure with a nil job id and the flow jumps past it). suspend.
continue_on_disapprove_timeout remains the way to continue past a gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(flow-editor): flag that continue on error does not cover the approval gate

A resolved approval is recorded against the step the gate holds back, not
the step carrying the suspend, so continue_on_error never sees it: the flow
still stops on a disapproval or timeout. Point users at
suspend.continue_on_disapprove_timeout, which is what actually continues past
a gate, whenever both settings are on and that one is not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 23:59:42 +02:00
Ruben FiszelandClaude Opus 5 25084170d7 feat: make job subprocess oom_score_adj configurable (#10443)
* feat: make job subprocess oom_score_adj configurable via JOB_OOM_SCORE_ADJ

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: warn when JOB_OOM_SCORE_ADJ leaves no gap over the worker's own score

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* style: drop em dash from JOB_OOM_SCORE_ADJ doc comment

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: warn on any oom_score_adj gap too small to steer the OOM killer

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 11:50:47 +00:00
Alexander Petric e0d6dc1a19 fix: harden flow-orchestration token refresh (mint from job_perms) (#10419)
* fix: harden flow-orchestration token refresh (mint from job_perms)

* refactor: address review nits on flow token refresh
2026-07-31 00:05:15 +02:00
81b23a2ba0 feat: make the fork lineage the only deploy relationship (#10410)
* feat: make the fork lineage the only deploy relationship

`workspace_settings.deploy_to` (2023) and `workspace.parent_workspace_id` (2025)
both expressed "which workspace does this one deploy into". Fork creation and
dev-workspace attach seeded both, but nothing kept them in agreement, so every
reader picked one and they disagreed.

Drop `deploy_to`. A migration folds surviving pairs into the lineage: a sole
claimant on a target with no dev workspace becomes that target's dev workspace
and keeps its own job tags, while many-to-one pairs become plain forks. Pairs
that the lineage cannot express -- dangling target, self-reference, chain,
mutual -- are reported and left unlinked.

Job tags were never lineage-aware: `per_workspace_tag` mapped any parented
workspace to its parent while `$workspace` interpolated the raw id, so a fork
running a script tagged `<tag>-$workspace` produced a tag no worker serves and
the job queued forever. Both paths now resolve to the nearest ancestor whose id
an admin would provision workers for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: preserve unconvertible deploy links and sweep tag caches on reparent

Review findings on the deploy_to unification:

- convert chains instead of discarding them, and keep whatever the lineage
  cannot express in workspace_deploy_to_unmigrated so the down migration can
  restore it
- ignore soft-deleted workspaces when choosing between a dev workspace and a
  plain fork; an archived claimant was demoting live pairs
- mirror attach_dev_workspace's git-sync strip, which the migration skipped
- sweep the tag cache over whole subtrees on rename and delete: tag resolution
  now walks ancestors, so a nested fork kept a tag nothing serves
- call a dev workspace a dev workspace in the settings copy
- redirect a root away from ?tab=deploy_to instead of rendering an empty target

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: detect lineage cycles and record archived links in the deploy_to migration

Second review round on the unification:

- detect cycles over the lineage as it would exist after conversion, not over
  the deploy_to graph alone: a root whose target was one of its own forks
  closed a loop that no deploy_to edge revealed
- record an archived source's link instead of filtering it out entirely, which
  dropped it with the column
- treat a fork whose deploy_to merely repeats its parent as redundant rather
  than reporting every pre-existing fork as unmigrated
- read the row count from the lineage update rather than the git-sync one
- sweep the tag cache when archiving a dev workspace, the last site that
  mutates is_dev_workspace without one

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: resolve $workspace on preprocessed flow tags regardless of $args

Third review round on the unification:

- a flow tag containing only `$workspace` skipped interpolation entirely on the
  preprocessed path, because the branch that ran it keys on `$args`. The raw
  tag was written back and named a queue no worker serves. Resolve `$workspace`
  before the branch and leave `$args` to it.
- record the new table's foreign key in the schema summary
- describe what the archive tag sweep actually does: the dev flag is cleared for
  any archived workspace, which is why it is unconditional

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep the deploy_to leftovers table only when it holds something

* fix: sweep tag caches on archive only where the dev flag actually changes

* feat: broadcast lineage changes and walk ws_specific ancestors only

- propagate tag-cache invalidation across processes over notify_events: the
  cache is per-process, so replicas kept resolving stale lineage for the TTL.
  The listener clears the whole cache rather than tracking ids, since a single
  mutation invalidates an unbounded set of descendants and lineage changes are
  rare admin actions.
- narrow list_ws_specific_versions to ancestors: walking down as well made a
  root fan out over its entire live fork subtree, and each member costs an
  identity lookup plus an RLS switch and probe. Ancestors are bounded by the
  fork depth limit.
- probe the leftovers table unqualified so rollback restores on a PG_SCHEMA
  install, where search_path is not public
- drop the nativets client method for the removed edit_deploy_to endpoint

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: let a prod see its dev workspace in ws_specific, and stop the walk oscillating

Descending into plain forks made a root fan out over its whole live fork
subtree, but a dev workspace is the paired editable environment rather than a
throwaway copy, so a prod should still see it. There is at most one per parent
and attach rejects nested dev chains, so that edge stays bounded.

The edges run both ways, so the recursion never converged: it bounced
parent<->dev until the depth cap on every call, 33 rows for a two-member set.
A visited-path guard ends the walk when nothing new is reachable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep dev pairings unnested, gate the delete broadcast, cover the ws_specific walk

Fifth review round:

- a root that already owns a dev workspace no longer converts: linking it under
  its deploy target would leave that dev nested beneath a fork, the shape
  attach_dev_workspace refuses to create. The link is preserved instead.
- broadcast a lineage change on delete only when descendants are orphaned.
  Deleting a leaf, which ephemeral fork churn does constantly, changes nobody
  else's resolution and was making every replica drop its whole tag cache.
- call list_ws_specific_versions in a test. plpgsql defers everything past a raw
  parse to the first call, so replaying the migration only proved it parses.
- use unwrap_or_default for the descendant sweeps, which run after the
  transaction has committed; a transient failure must not fail the request
- trim the traversal comment to the four-line limit

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: cache the renamed tally query and clear instance alerts on conversion

The integration test's query was never cached: `cargo sqlx prepare` without
--all-targets skips test targets entirely, and renaming its fixture workspace
changed the query text. Regenerated with --all-targets --features
all_sqlx_features,private, which is what lets the EE-gated otel test compile.

Also from review:
- clear error_handler_fallback_to_instance_alerts on converted workspaces.
  Dispatch ignores it once a parent exists, but the settings page keeps
  submitting the stored true, which the API rejects on a fork.
- restore the schema summary row to the file's name: columns format and put it
  back in alphabetical order

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: never cache an unresolvable tag workspace, and unadvertise the removed endpoint

- lookup_tag_workspace cached a "no row" result as self-resolution. A rename
  resolves the new id before its row lands, so a fork could be pinned to its own
  wm-fork-* id -- which nothing serves -- for the whole TTL, and its schedules
  kept re-pushing onto that dead tag. Fall back for the call without caching,
  matching how the error path already behaved.
- change_workspace_id swept its children but never itself. Sweep the new and old
  ids and broadcast unconditionally, since a rename always changes lineage.
- openapi-deref.{json,yaml} are served to clients via include_str!, so they were
  advertising edit_deploy_to after it started 404ing. The audit-action enum
  keeps the entry: historical rows still carry it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: align the served YAML spec with the JSON one and correct two comments

- the YAML deref lost the removed path but kept deploy_to on get_settings,
  so the two served specs disagreed. Both are now identical.
- the rename-sweep comment blamed cached-unresolvable lookups, which the same
  commit stopped caching. The real reason is that workspace ids are
  reclaimable, so a new id can carry a previous occupant's resolution.
- the instance-alert comment claimed the settings page submits the stored true
  and gets a 400. It hides the option on a fork and sends false; the hazard is
  the value outliving the pairing and re-enabling alerts after a detach.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref to 82da6cb2bafeda18acd6b70c599013a12117ecb0

This commit updates the EE repository reference after PR #694 was merged in windmill-ee-private.

Previous ee-repo-ref: f9ddf6a75aa13d1c13a3d7216a361a96f75ca435

New ee-repo-ref: 82da6cb2bafeda18acd6b70c599013a12117ecb0

Automated by sync-ee-ref workflow.

* fix: grant the deploy_to preservation table to the windmill roles

* test: drop the one-shot migration tests, keep the ws_specific execution guard

The two conversion tests replayed the migration against the fully-migrated
schema, which is not how it runs -- in production it runs mid-sequence against
the schema as of that point. A later migration touching workspace or
workspace_settings would break them without breaking anything real, and sqlx
checksums already freeze a released migration. They earned their keep finding
the archived-claimant and nested-dev cases during development; there is nothing
left for them to guard.

list_ws_specific_versions is different: it is live, no caller exercises it, and
plpgsql only parses a function body until first call.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: invalidate a reclaimed fork id cluster-wide without flushing every entry

Gating the delete broadcast on orphaned descendants stopped leaf churn flushing
every replica, but fork ids are reclaimable: the deleting process invalidated
locally while every other replica kept the old parent for the TTL, so a job
pushed in a recreated fork routed to the previous parent's tag.

The broadcast payload now carries meaning. A workspace id drops that one entry,
used for leaf deletion where exactly one id changed what it denotes. The `*`
sentinel drops everything, used for attach, detach, archive, rename and
deletions that orphan descendants -- reshaping a subtree no single id names.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: name the right broadcast for each invalidation case

* docs: attach does invalidate the tag cache; the resolver walks the whole chain

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-07-30 14:20:27 +00:00
Ruben FiszelandClaude Opus 5 4d3ff0299f feat: mark failed jobs as resolved so handled failures stop showing red (#10319)
* feat: mark failed jobs as resolved so handled failures stop showing red

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: constrain auto-resolve to the proven retry chain and honor resolved filter everywhere

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: apply resolved filter to queue-union, concurrency and delete paths, bound note

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: sweep resolutions on workspace delete, verify helper args, enforce UI limits

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: count resolution note in characters on both sides of the API

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: skip the queue lookup for cancel-all under the resolved-only filter

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: converge retry auto-resolution from either commit order, keep notes on re-resolve

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: correct the idempotency claim on the retry auto-resolve sweep

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: gate resolution notes and attribution behind enterprise, add note popover

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: hide resolution from operators, exclude flow steps, enforce EE licence at runtime

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: add job_resolution.automatic to the summarized schema

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: preserve stored attribution when re-resolving without a valid licence

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: condense the attribution-preservation comment to four lines

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: validate resolution notes by code point instead of a UTF-16 maxlength

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep the resolution popover open when a note is rejected

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: offer to resolve the original failure after a successful re-run

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: verify supersession server-side and stop re-runs overwriting notes

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: apply tag scope to the superseding run

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: exclude obscured cross-workspace runs from resolution actions

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 09:36:27 +02:00
30d8104edc feat: alert on expired online license key (#10295)
* [ee] feat: alert on expired online license key

Wire alert_on_online_license_expired into the periodic monitor loop
(verify_license_key_f), server-mode gated so only servers report to the
alerts table and critical channels. Add the ee_oss stub so the
enterprise-without-private build still compiles.

Closes the gap where an expired online (renewable) license force-set
externally (env var / CI / k8s) halted all jobs with only a stdout
tracing::error! and no critical alert.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: point ee-repo-ref at license-expired-alert EE branch

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref to 9e63619688588fda434dfa06ed585f050e75ee67

This commit updates the EE repository reference after PR #682 was merged in windmill-ee-private.

Previous ee-repo-ref: 7847711567a9c66da6a5c85d176b8cacc5aa7a7f

New ee-repo-ref: 9e63619688588fda434dfa06ed585f050e75ee67

Automated by sync-ee-ref workflow.

* chore: bump ee-repo-ref to license-expiry alert review fixes

Point at the EE follow-up (windmill-ee-private#683): cross-replica dedup via
acquire_lock and no false recovery on malformed key replacement.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref to eb8733a36010a7509b438903e95edfa293da079b

This commit updates the EE repository reference after PR #683 was merged in windmill-ee-private.

Previous ee-repo-ref: 562a306d9b66624b28ff90958f8b45d8db5481b9

New ee-repo-ref: eb8733a36010a7509b438903e95edfa293da079b

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-07-24 10:53:24 +02:00
Ruben FiszelandClaude Opus 4.8 f02df7fc45 feat(monitor): make between-steps zombie flows hand-recoverable (#10287)
* feat(monitor): make between-steps zombie flows hand-recoverable

When a worker is OOM-killed mid state-transition, the flow is reaped as a
between-steps zombie (children all success, module still InProgress). We do
not auto-recover (a re-driven transition can OOM again), so instead:

- Append actionable recovery guidance to the cancellation reason when the
  reaped step's state is derivable (every child a success completion): which
  step, iterations completed, raise memory then restart-from-step (UI + API).
- Restart-from-step now reuses a zombie step verbatim (InProgress with all
  children successful) and restarts from the next step, so no completed child
  re-runs; downstream steps re-derive its result from flow_jobs on demand.
- Cast flow_status ::text in the reaper query: reading the jsonb column as
  Box<str> included the binary version byte and silently failed FlowStatus
  parsing (disabling the restart-not-yet-started branch since the v2 migration).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(monitor): only reuse a between-steps zombie step that provably finished

Address review findings on the zombie-restart reuse path:

- Require structural completeness (FlowStatusModule::is_between_steps_complete):
  a serial for-loop / branch-all reaped mid-fan-out has an all-success prefix but
  unrun remaining iterations, so the cursor must sit on the last element; while-loops
  are never derivable (continuation is a post-iteration condition). Parallel
  containers preallocate all children, so success alone is conclusive. Shared by the
  monitor guidance and the restart resolution.
- Decline reuse when the step carries stop_after_if / stop_after_all_iters_if: those
  predicates decide whether downstream steps run, and reuse would bypass them; such a
  step re-runs instead.
- Decline reuse when the zombie step is the last module (advancing past it lands on
  the failure step); it falls back to the existing re-run path.
- Unit tests for is_between_steps_complete and an integration test asserting a
  mid-iteration serial-loop zombie is re-run, not reused.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(monitor): align zombie recovery guidance with restart eligibility

Address CI review findings:

- Exclude skip_if / suspend / sleep (not just stop predicates) from reuse via
  FlowModule::allows_zombie_reuse, so a skipped/suspend-armed step is never
  synthesized as Success (which would strand a restart waiting on an approval it
  never armed).
- The reaper does not load the flow definition, so it cannot know whether restart
  will reuse or re-run a given step; reword the guidance to state both outcomes
  (reuse where derivable, re-run for the flow's last step or one carrying a
  stop/skip condition, approval, or sleep) instead of promising "no re-run".
- Make the mid-iteration regression test exercise the cursor-completeness guard:
  a downstream step makes the loop non-final, so reuse is prevented only by the
  guard; a truncated loop result would then fail the assertion.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(monitor): never let zombie reuse swallow a nested restart request

A nested restart (RestartedFrom.nested) descends into the restart step's child to
re-run an inner step. For an eligible zombie BranchOne/Subflow the outer
branch_or_iteration_n is None, so reuse fired, skipped the container, and the
explicitly requested inner step never re-ran. Thread the presence of a nested
chain into restarted_flows_resolution and decline reuse when set. Regression test
added (RED without the guard: the nested target is reused instead of re-run).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(monitor): don't auto-requeue preprocessor zombies as unstarted flows

The ::text parse fix re-activated the "hasn't started yet, restart it" branch,
but its `modules[0] == WaitingForPriorSteps` check also matches a flow whose
preprocessor is still InProgress (step == -1, first module waiting). Requeuing
such a flow re-runs the preprocessor, duplicating side effects / repeating the
OOM. Gate the branch on FlowStatus::is_not_yet_started, which also requires the
preprocessor (if any) to be WaitingForPriorSteps. Unit-tested.

Also drop the numbered procedural narration from the happy-path test comments.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(monitor): only emit restart guidance for restartable (deployed, top-level) flows

The recovery guidance points operators at the run page's "Re-start from" button
and the restart API, but both require a top-level deployed flow: a preview has no
flow path (the button is hidden, the API 400s) and a subflow child restarts via
its root, not itself. Gate the guidance on runnable_path IS NOT NULL AND
parent_job IS NULL so previews/subflows keep the existing wording instead of
being told to use a button/endpoint that isn't there. Verified end-to-end: a
reaped preview gets no RECOVERY block, a reaped deployed flow does.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(monitor): gate recovery guidance on kind='flow' to match the restart surface

Addresses review nit: a pathful editor preview (kind='flowpreview' with a
runnable_path) satisfied the previous runnable_path check but the run page only
renders the "Re-start from" button for kind='flow'. Match that condition exactly
so previews/singlestepflow keep the plain wording. Verified end-to-end: a reaped
pathful preview now gets no RECOVERY block.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(monitor): disable zombie reuse for raw-flow (editor preview) restarts

A JobPayload::RawFlow restart queues the request's current, possibly EDITED,
definition, but restarted_flows_resolution validates reuse against the completed
job's STORED definition. For an eligible preview zombie, editing the restart step
and restarting from it would synthesize Success from the old children and skip the
edit. Thread allow_zombie_reuse into the resolver (true only for
JobPayload::RestartedFlow, which queues the stored definition) and decline reuse
for raw-flow restarts. Regression test added (RED without the guard: the edited
step is skipped and the old result is reused).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(sqlx): add offline cache for zombie_flow_recovery test queries

The integration test's UPDATE v2_job_completed queries had no .sqlx entry, so the
CI SQLX_OFFLINE build of the test failed to compile. Regenerated with
--all-targets --features deno_core,quickjs to capture the test-target queries.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(monitor): drop procedural narration from the raw-flow zombie test

Per AGENTS.md (comments record constraints, not narration): remove the two
step-describing comments the reviewer flagged; the test doc comment already
carries the durable rationale.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(monitor): restrict zombie reuse to monitor-reaped flows

The reuse predicate matched the InProgress/all-children-success shape without
checking provenance, so an ordinary force-cancel at the same boundary (a child
succeeded before its parent transition landed) would also be reused, dropping
the usual restart-from-step re-run. Gate reuse on canceled_by = 'monitor' (the
username the zombie reaper cancels with). Regression test added (RED without the
guard: a user-cancelled flow reuses the child instead of re-running it).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(monitor): reuse zombie step on Some(0) too, so the run-page button works

The run page's "Re-start from" button always sends branch_or_iteration_n = 0
(never omits it), but reuse only fired for None, so the exact UI path the
recovery message points to would re-run the children instead of reusing them.
Treat a whole-step restart (None or Some(0)) as reuse-eligible; Some(n>=1) keeps
the explicit partial-container restart. Verified against the live EE restart API
with branch_or_iteration_n=0: all loop-iteration child UUIDs are reused. Happy-
path test now sends Some(0) to match the button.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 19:09:28 +02:00
Ruben FiszelandClaude Opus 4.8 fa3644281f fix(monitor): diagnose zombie-flow OOM on the transition worker, not q.worker (#10286)
handle_zombie_flows() joined worker_ping on the flow's queue-row worker
(q.worker) only. In nested/subflow/forloop cases that worker is frequently
NOT the one that performed the flow's final state transition, so all OOM
detection ran against a healthy bystander and the diagnostic wrongly told
on-call to chase a deadlock / SIGQUIT.

Change 1: when q.worker itself looks healthy, additionally search worker_ping
for a plausible culprit — a different worker on the same pod (worker_instance)
or worker group whose last ping clusters around the flow's last_ping and has
since gone silent (the signature of a worker OOM-killed mid state-transition),
picking the one closest in time to the transition. If found, the message names
that worker and routes to the OOM explanation. The extra query runs only on the
rare zombie path and fails soft (warn + fall back) so it never blocks the cancel.

Change 2: reorder the healthy-q.worker hint to lead with "check whether a
different worker on the same pod/group was OOM-killed around {last_ping}",
keeping deadlock/SIGQUIT as the secondary possibility.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 11:40:23 +00:00
Ruben Fiszel ddec2abbb3 feat(jobs): cap total queued jobs per workspace on cloud (#10218)
* feat(jobs): cap total queued jobs per workspace on cloud

A workspace could flood the queue with an unbounded number of jobs across
many concurrency keys and scripts (or keyless jobs), which the per-key
cap from #10197 does not bound. Add a companion instance-wide ceiling on
a workspace's total queued jobs.

check_workspace_queue_cap rejects a push once the workspace has
WORKSPACE_MAX_QUEUED_JOBS (default 20000, superadmin-configurable, 0 to
disable) jobs queued, cloud-only and runtime-gated on CLOUD_HOSTED like
the per-key cap. It runs on every push, so it applies even to premium
workspaces and catches parallel for-loop floods. Jobs already queued
still drain; only new pushes past the ceiling are rejected, so an
in-flight flow only fails to push further work while at the ceiling.

The setting loader self-gates on CLOUD_HOSTED so it is never loaded off
cloud, from initial load or a settings-change reload. The depth count is
bounded by the cap via LIMIT so a runaway backlog never costs an
unbounded scan on the push path.

* docs(jobs): note the workspace cap is a soft ceiling and the depth helper is count-only

Records the two review points as constraints: the cap does not serialize
admission (a soft ceiling by design, like the per-key cap), and
workspace_queue_depth is pub only for the test, returns a count not job
data, and leaves authorization to the caller.
2026-07-20 22:09:59 +02:00
11fda89b52 feat(telemetry): generic feature-usage telemetry with AI session metrics (#10200)
* feat(telemetry): add generic feature_usage table and batched logging endpoint

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(telemetry): log AI session usage events and document them in telemetry settings

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(telemetry): use escape sequence instead of literal NUL bytes in buffer key

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(telemetry): validate dimensions, decouple retention, keepalive flush

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(telemetry): allowlist feature-usage dimensions and index retention scans

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(telemetry): pin tool-name allowlist and deploy session attribution

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(telemetry): route AI chat usage through feature_usage and drop ai_chat_usage

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(telemetry): slim dimension validation to registered kinds plus key shape

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(telemetry): backfill ai_chat_usage into feature_usage before dropping it

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(telemetry): disclose provider and model identifiers in telemetry settings text

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(telemetry): issue all flush chunks before awaiting so pagehide keeps them

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: update ee-repo-ref to 6306c072a50937ea9af44a5bcf42345543207486

This commit updates the EE repository reference after PR #672 was merged in windmill-ee-private.

Previous ee-repo-ref: 964f242a0eb44db7f7d26636cc8d76aeabea2b73

New ee-repo-ref: 6306c072a50937ea9af44a5bcf42345543207486

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
2026-07-20 20:56:57 +02:00
87be041c09 fix(git-sync): avoid percent-encoded colon in git-sync hub script path (#10213)
* fix(git-sync): avoid percent-encoded colon in git-sync hub script path

The git-sync init/detection hub script slug contained a colon stored as
`%3A` in the run-by-path URL. The generated API client re-encodes path
params with encodeURI, turning `%3A` into `%253A` (double-encoding). Some
hardened reverse proxies / WAFs reject double URL-encoding and return a
bare 400 before the request reaches Windmill, breaking git-sync repository
detection on those instances.

The hub resolves scripts by numeric id and ignores the slug, so dropping
the colon from the slug is behavior-neutral (same script, same id-keyed
worker cache) while producing a colon-free run URL.

Also force-cache GIT_SYNC_PULL_SCRIPT_PATH at build alongside
LATEST_GIT_SYNC_SCRIPT_PATH so the backend-driven pull script is always
baked into the image for airgapped workers, instead of relying on an
incidental hubPaths.json overlap.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: bump ee-repo-ref for git-init slug match fix

Pulls in windmill-ee-private#676 so the EE is_git_init_script check matches
the colon-free git-init hub slug (GitHub App token grant for git-sync jobs).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: bump ee-repo-ref for git-init slug helper + test

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref to a3adea1ffb406e709cc480871df58fab6c51aca1

This commit updates the EE repository reference after PR #676 was merged in windmill-ee-private.

Previous ee-repo-ref: cef4e008ef62dec434aa9bb3ec783db8aff6a1c1

New ee-repo-ref: a3adea1ffb406e709cc480871df58fab6c51aca1

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-07-20 20:36:34 +02:00
Ruben Fiszel 71f2d47cb4 feat: cap queued jobs per concurrency key on cloud (#10197)
* feat: cap queued jobs per concurrency key on cloud

* fix: close preprocessed-flow bypass and bound concurrency cap scan

* fix: only cap concurrency keys with an active concurrent_limit

* chore: only load concurrency key cap setting when cloud hosted

* fix: reject queued-job import on cloud
2026-07-20 12:33:40 +02:00
Ruben FiszelandClaude Opus 4.8 f6e36f862e chore(schedules): lower reconciler re-arm back-off cap to 8 passes (#10182)
Follow-up to #10179. The exponential back-off between reconciler re-arm
retries of a persistently-failing schedule capped at 32 passes (~2.7h at
the default 5-min reconcile cadence). Lower the cap to 8 (~40min) so a
schedule fixed out of band (a lapsed license renewed, a bad cron corrected
directly in the DB) auto-recovers within a few passes, while still cutting
the retry rate sharply versus retrying every pass. Fixes via the UI/API
re-arm immediately and are unaffected.

Fixes WIN-2198

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-18 08:49:41 +02:00
Ruben FiszelandClaude Opus 4.8 c82056cfde fix(schedules): stop disabling schedules on transient push errors (#10179)
* fix(schedules): stop disabling schedules on transient push errors

A scheduled flow whose next-occurrence push failed after retry exhaustion
used to be disabled, killing a healthy schedule over a transient DB blip
(pool contention, statement timeout). Now that the unarmed-schedule
reconciler exists (#10174), transient failures no longer disable: the
current occurrence runs to completion and the reconciler re-arms the next
occurrence once this run leaves the queue.

In the flow schedule-push path after retry exhaustion we now branch on the
error: QuotaExceeded/NotFound still disable (the schedule's own fault, and
rearm_schedule would otherwise leave them enabled-yet-unarmed forever),
while transient errors are only reported and the flow continues.

The previous iteration returned a SchedulePushZombieError to force a zombie
restart; that is removed, because zombie detection cancels (does not
restart) same-worker flows, so it would have lost the current run of a
same-worker scheduled flow. The now-obsolete SchedulePushZombieError type
and its catch in worker.rs are deleted.

Fixes WIN-2198

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(schedules): back off and surface repeated reconciler re-arm failures

The unarmed-schedule reconciler retried a schedule that could not be
re-armed on every pass, forever, logging only to the server. With the flow
schedule-push path no longer disabling on non-transient errors, a
persistently-broken push (bad stored cron/timezone/args, lapsed license
key) now stays enabled and would spin in that loop silently.

The reconciler now tracks consecutive re-arm failures per schedule:
exponential back-off (2, 4, 8, … passes, capped) between retries so a
broken schedule is not hammered, and after 3 consecutive failures it
surfaces the cause once (records schedule.error + raises a critical alert)
without disabling. Both reset the moment the schedule re-arms, which also
clears the recorded error.

Verified end-to-end on a running server: a flow schedule with a corrupted
cron stays enabled, retries back off, the error is surfaced after the
third failure, and it re-arms and clears the error once the cron is fixed.

Fixes WIN-2198

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 22:48:15 +02:00