mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-09-06 00:02:13 +00:00
815de49e2322f85ca92b1e41a2bcd22591ebe93f
129
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
815de49e23 |
feat: make the service log retention period an instance setting (#10889)
* feat: make the service log retention period an instance setting Service log retention was a hardcoded 14 days with no override, unlike job retention. It becomes the `service_log_retention_secs` global setting (env `SERVICE_LOG_RETENTION_SECS`, default unchanged at 14 days), reloaded on change like the other retention settings. The constant becomes `DEFAULT_SERVICE_LOG_RETENTION_SECS` and every reader goes through `service_log_retention_secs()`, so the `log_file` sweep, the object-storage orphan scan, the columnar store's compaction and pruning, the retrieval clamp and the search index's trim window all follow the configured value. Loaded outside `initial_load`'s `server_mode` guard: a dedicated indexer trims the search index to a window derived from this value and is not a server. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN * fix: never let a non-positive service log retention expire every log Every service log cutoff is `now - retention`, so a `0` or negative window puts the cutoff at or after `now` and the next sweep reads the whole history as expired — deleting the `log_file` rows and their object-storage files irreversibly. `0` is reachable two ways now that the window is configurable: it is what an operator types by analogy with the job retention period sitting directly above it, where `0` does mean keep forever; and `SecondsInput` writes a `0` into a field that was merely focused, so saving the Jobs panel is enough. Service logs always have a window, so clamp an unusable value back to the default in the accessor every reader already goes through. The upper bound is where `chrono::Duration::seconds` panics, which would abort the sweep that reads it. The settings field rejects a non-positive value rather than silently correcting it, and its description now names the database rows too — they are swept on every instance, including one with no object storage configured. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN * fix: address review findings on the service log retention setting - Bound the monitor's `log_file` sweep. Every process rotates a log file a minute, so lowering the retention can make one ordinary setting change expire millions of rows; the unbounded `DELETE ... RETURNING` materialized all of them, and their deletion futures, in a single tick. Batched like the settings-page cleanup on the same table. - Make the retention atomic private and give it one writer, so a value that would expire every service log cannot reach a cutoff by any path, and say so in the log when one is rejected rather than falling back silently. - Cap the retention at a century. The previous ceiling only bounded `TimeDelta` construction, while consumers compute `now - retention`, which panics past year 262143, and build a Postgres interval that overflows well before the old cap. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN * fix: cap an oversized service log retention instead of shortening it The two unusable directions were landing on the same fallback, so configuring a retention above the ceiling silently produced 14 days — deleting logs the operator had asked to keep for longer. Too large now caps at the maximum, which preserves that intent; only a non-positive value, which would expire everything and has no upward reading, falls back to the default. Also bound the `log_file` drain to ten batches per pass: `monitor_db` runs under a 600s timeout that cancels every maintenance future in the same `join!` and reports a critical error, so a backlog large enough to need batching has to drain across ticks, the way the neighbouring sweeps already do. The settings field carries the upper bound too, and the superseded query's offline entry is dropped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN * fix: route the new log-file registration cutoff through the retention accessor `send_log_files_to_object_store` arrived on main while this branch was open and reads the retention directly. The atomic behind it is private now, so it goes through the accessor like every other consumer — which also means the cutoff it uses to skip registering already-expired files follows the configured retention rather than a fixed two weeks. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN * fix: say why every mode loads the service log retention setting A worker registers its rotated log files against the retention cutoff, so the comment naming only the indexer no longer covers why the setting sits outside the `server_mode` guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN * fix: file service log retention under Monitoring, not Jobs Service logs are the Windmill processes' own logs — every process rotates and registers its own, no job involved — so the Jobs panel was grouping by the shape of the widget rather than by the subject. It sits under Monitoring now, beside the Indexer panel that holds the other service-log window. Its own section rather than inside that panel: the panel is badged EE, while this governs the database sweep that runs on every instance. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN * chore: update ee-repo-ref to a6e3533b26195918a17fea58646f71d2bbcde288 This commit updates the EE repository reference after PR #752 was merged in windmill-ee-private. Previous ee-repo-ref: 1d93da24bd166b9a5a5cc204034a1d35ffc88474 New ee-repo-ref: a6e3533b26195918a17fea58646f71d2bbcde288 Automated by sync-ee-ref workflow. * feat: say on the service logs page where the logs actually are The retention number alone does not tell an operator what it governs, and the answer differs by instance. Two states are worth calling out because they are the ones where retention does not mean what it looks like: Without instance object storage, each process keeps its files on its own disk. The page lists what every host wrote, since the rows are in the shared database, but can only open the files of the replica serving the request, and a host's files go with it when it is replaced. With object storage but "Delete logs from s3 periodically" off — the backend default, since uploads are gated on a store existing while deletions are gated on that toggle — expiring a log removes the row and the local file and leaves the uploaded copy behind for good. The retention field itself now names every copy it covers and says that full-text search reaches back at most that far, and less when the indexer's own window is shorter. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN * fix: describe raw log files as the transient copy they became Retiring the raw files landed while this was being written: the indexer now deletes each one as soon as it is ingested, and the log viewer rebuilds a file from the columnar store once the raw copy is gone. So the durable copy is the store, and warning that an uploaded file is kept forever when periodic s3 deletion is off only holds where no indexer runs to ingest it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN * chore: point ee-repo-ref at the EE compile fix EE main does not build on its own: extracting the index-window expression and adding a fourth copy of it landed in separate PRs that never conflicted textually. windmill-ee-private#756 is the one-line fix; this pins it so CI has a tree that compiles. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsnpNSM6K3oyjwntRwJtVN --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
2906504125 |
feat: add instance setting to mute zombie job restart alerts (#10813)
* feat: add instance setting to opt out of zombie job restart alerts Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: preserve explicit false for default-on boolean instance settings Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor: invert zombie restart alert setting to a mute flag Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ef99a739dd |
fix(github-app): complete the self-managed setup instructions, render the page header (#10683)
* docs(github-app): state the pull-direction permissions and the App owner field The in-product "How to create a GitHub App" panel only listed Contents and Metadata, which covers the push direction of git sync. Webhooks, pull requests and checks are what the git to Windmill direction needs, and a GHE Cloud (*.ghe.com) app also needs App owner, whose field hint was the only place saying so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(instance-settings): render the GitHub App page header The branch tested the pre-rename category name, so the page rendered with no header at all. Naming the header after the category duplicates the card below it, so the card that holds the app credentials is now labelled for what it is, next to the webhook base url card. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
71b9989daa |
feat: auto-build binaries to object storage on deployment (#10673)
* feat: auto-build binaries to object storage on deployment Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: queue the auto-build from pre-locked deploys and off the lock slot Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: materialize companion modules before a deploy-time build Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep a build job from stamping lock_error_logs on a healthy script Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: de-flake test_flow_lock_all and surface the lock error it hides Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: trim drafting history from the flow-lock fixture comments Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: stop a binary build from restarting dedicated workers Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep the build-job marker off the agent wire and out of user args Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2d2cdb7a99 |
allowlist resource_type and escape fuzzy-search highlights (#10509)
* fix: allowlist resource_type and escape search highlights * fix: bound path length and keep marked-label offsets entity-aware * fix: match postgres word-char semantics and drop double-escaping * fix: sanitize db constraint and rls errors instead of relying on the regex |
||
|
|
318c9f0073 |
feat(git-sync): dedicated base url for GitHub webhook delivery (#10411)
* feat(git-sync): let GitHub webhooks register a dedicated base url Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): validate the webhook base url and apply it on change Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: pin ee ref for the git-sync webhook base url change Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): validate and reconcile the webhook base url on every write path Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): route every declarative settings writer through the same rules Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): let the reconciler own the webhook field write-back Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): make the webhook base url validators agree across UI and server Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): lock the workspace row across git_sync read-modify-writes Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: pin the webhook base url validator to its server counterpart Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): retry a failed webhook move on every re-apply of the setting Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): retry pending webhook moves on every declarative re-apply Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): reject non-string webhook base urls and bound the sweep Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): reject credential-bearing webhook base urls Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): keep credentials out of webhook base url validation errors Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): redact through the last authority @ when reporting a bad url Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): stop echoing unparsed webhook base urls instead of scrubbing them Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): never echo a submitted webhook base url in validation errors Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): keep the submitted scheme out of validation errors Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(git-sync): drop the webhook sweep, surface stale receivers in settings Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): refresh the stale webhook list when settings are saved Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): mark registered_url nullable and drop the duplicated field error Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: pin ee ref after dropping the reconcile lock and CAS Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(git-sync): refresh the stale webhook list on category saves too Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to aa05ca8e97fc8265cd724753a80db37f83243254 This commit updates the EE repository reference after PR #695 was merged in windmill-ee-private. Previous ee-repo-ref: 3e6cd9226b68707233ae2434511fe5131dce808b New ee-repo-ref: aa05ca8e97fc8265cd724753a80db37f83243254 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
621718b32f |
fix(ai): resolve deployment-pinned Azure base URLs to the v1 surface (#10362)
An Azure OpenAI base URL naming a deployment, such as the `https://<res>.openai.azure.com/openai/deployments/<id>` format that `openai_azure_base_path` documents, was appended to as-is. That names the legacy surface, which serves only with an `api-version` query and answers 404 without one, so both the proxy and the AI agent step reached a route that does not exist. Such a base now resolves to the resource root and the v1 surface, like every other Azure shape. The deployment in the URL is redundant there: the v1 surface takes it from the request body. `azure_foundry_root` recovers the root the same way, so a Foundry resource on such a base builds its Claude URL from the root too. The instance-settings help text promised the URL pins the model for every workspace, which that surface never delivered; it now says where the model comes from. Verified against a live Azure OpenAI resource: the previous URLs 404 and the ones built now return 200. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
434c4ac7c8 |
chore: pin ruff to 0.16.0 and keep the python editor rule set stable (#10331)
* chore: pin ruff to 0.16.0 and keep the python editor rule set stable * chore: keep the ruff config path rationale at a single site |
||
|
|
ddec2abbb3 |
feat(jobs): cap total queued jobs per workspace on cloud (#10218)
* feat(jobs): cap total queued jobs per workspace on cloud A workspace could flood the queue with an unbounded number of jobs across many concurrency keys and scripts (or keyless jobs), which the per-key cap from #10197 does not bound. Add a companion instance-wide ceiling on a workspace's total queued jobs. check_workspace_queue_cap rejects a push once the workspace has WORKSPACE_MAX_QUEUED_JOBS (default 20000, superadmin-configurable, 0 to disable) jobs queued, cloud-only and runtime-gated on CLOUD_HOSTED like the per-key cap. It runs on every push, so it applies even to premium workspaces and catches parallel for-loop floods. Jobs already queued still drain; only new pushes past the ceiling are rejected, so an in-flight flow only fails to push further work while at the ceiling. The setting loader self-gates on CLOUD_HOSTED so it is never loaded off cloud, from initial load or a settings-change reload. The depth count is bounded by the cap via LIMIT so a runaway backlog never costs an unbounded scan on the push path. * docs(jobs): note the workspace cap is a soft ceiling and the depth helper is count-only Records the two review points as constraints: the cap does not serialize admission (a soft ceiling by design, like the per-key cap), and workspace_queue_depth is pub only for the test, returns a count not job data, and leaves authorization to the caller. |
||
|
|
71f2d47cb4 |
feat: cap queued jobs per concurrency key on cloud (#10197)
* feat: cap queued jobs per concurrency key on cloud * fix: close preprocessed-flow bypass and bound concurrency cap scan * fix: only cap concurrency keys with an active concurrent_limit * chore: only load concurrency key cap setting when cloud hosted * fix: reject queued-job import on cloud |
||
|
|
ff774c46bf |
feat: add per-workspace job-retention override (#10050)
* feat: add per-workspace job-retention override (EE) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to 2ba6a2a75b6fc97858b306b2c98ada481e363c10 This commit updates the EE repository reference after PR #658 was merged in windmill-ee-private. Previous ee-repo-ref: e7fb36acd813cd717bcf05f5aafbf81de271d618 New ee-repo-ref: 2ba6a2a75b6fc97858b306b2c98ada481e363c10 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
577ceeee86 |
perf(audit): re-anchor S3 audit export on enable + opt-in backfill (#9818)
* [ee] perf(audit): re-anchor S3 audit export on enable + opt-in backfill
The S3/GCS audit-log export's steady-state query filters by `age(xmin)`
(unindexable), so the only scan bound is the timestamp floor. On a fresh
enable the floor was epoch, and on a re-enable the cursor resumed from its
pre-disable position — either way the first run scanned the whole
`audit_partitioned` table. Under a `statement_timeout` (e.g. Aiven) that scan
never completes: the cursor never advances, nothing is exported, and the
repeated full scans saturate the database.
Re-anchor on enable (EE companion, windmill-ee-private#634):
- New trigger migration records a recent timestamp floor instead of the epoch
sentinel and `DO UPDATE`s the cursor to the current snapshot xmin on
re-enable, so the export always resumes from ~now and never rescans history.
Includes a one-time fixup for legacy epoch-sentinel checkpoints on upgrade.
Opt-in historical backfill (new `audit_logs_s3_backfill` module + endpoints):
- Exports a chosen `[from, to)` window on demand, scanning strictly by
`timestamp` (the partition key) in bounded keyset pages — each query is an
index scan capped at one page (verified via EXPLAIN: later partitions
`never executed`, ~11ms/page), so it stays well under any statement timeout
regardless of window size. Writes alongside the steady-state objects under
logs/audit/, without touching the xmin cursor.
- POST /settings/audit_logs_s3_backfill {from,to} (super-admin + Enterprise),
GET /settings/audit_logs_s3_backfill_status.
Also repurposes the status endpoint's `bootstrapping` flag to mean "draining a
backlog" (the cursor is capped and catching up), and updates the setting
description to point operators at the backfill.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audit): heartbeat backfill lease per object; bump EE ref
Address review (cubic): persist progress (refreshing the lease heartbeat) after
every object PUT in the backfill page loop, not only once per page, so the gap
between heartbeats stays well under STALE_HEARTBEAT_SECS even on slow uploads
and another replica can't re-claim mid-page and run a concurrent backfill.
Bumps ee-repo-ref.txt to pull in the EE test-race fix (folding the backlog-drain
regression into the single audit e2e test).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audit): reject unstable backfill windows; bump EE ref
Address review (P1): the backfill keyset-pages over rows visible at scan time
and declares completion when the scan runs dry, but a row's `timestamp` is its
inserting transaction's `xact_start`. A window whose upper bound is recent or in
the future could silently omit a transaction that started inside `[from, to)`
but commits after the scan passed that timestamp. `try_start` now rejects any
`to` newer than the oldest in-flight `xact_start` (everything strictly older
than the oldest running transaction is committed and stable), using the same
trustworthy stats gating as the exporter's floor (restricted role / 2PC → a
7-day-old cutoff).
Bumps ee-repo-ref.txt for the EE monotonic-checkpoint fix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audit): re-anchor legacy epoch checkpoints instead of synthetic floor
Address review (P1): the legacy-checkpoint fixup stamped last_oldest_inflight_ts
to now()-7d while leaving the old last_xmin in place. On an instance that
enabled export on the old code >7 days ago and got stuck before the first
successful batch, the next run would filter post-enable rows older than 7 days
out via `timestamp >= ts_floor` while still advancing last_xmin over the
interval — silently dropping them (the same floor-vs-cursor loss class fixed
elsewhere in this PR), and contradicting the "nothing committed after enabling
is skipped" guarantee.
A stuck epoch-sentinel checkpoint cannot be safely resumed (its backlog can be
arbitrarily old, so any recent floor prunes rows the cursor then skips, and an
epoch floor reintroduces the full scan). Re-anchor it to the migration's current
snapshot xmin instead — exactly like a fresh enable — so the export resumes
cleanly from ~now and the never-exported pre-upgrade window is recovered via the
opt-in backfill rather than silently dropped. Reword the setting description so
it no longer implies the disabled/legacy window is covered by the cursor.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(audit): end-to-end integration tests for the object-store backfill
The backfill previously had only SQL-level/EXPLAIN validation. Add real
integration tests (in-memory object store, sqlx::test) exercising the public
path:
- backfill_exports_window_in_pages: with the page size forced to 2 rows, a
settled 3-day window is exported across multiple keyset pages; asserts every
in-window row lands exactly once, rows outside [from,to) are excluded, a day
that straddles a page boundary yields more than one object, progress counts
match, and a re-run is idempotent (deterministic keys overwritten, no dupes).
- backfill_rejects_unstable_window: a future/live `to` is rejected as unstable,
a window safely in the past is accepted.
Adds a test-only PAGE_ROWS override so multi-page behaviour is exercised with a
handful of rows.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(audit): note backfill scope is audit_partitioned only
Make explicit that, like the steady-state export, the backfill reads only
audit_partitioned; the pre-partitioning `audit` table is intentionally out of
scope (not a missed case).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audit): reject backfill windows before the partitioned boundary
Address review (Codex P1): the backfill reads only audit_partitioned, but
pre-partitioning history lives in the legacy `audit` table (still read by audit
list/get via UNION ALL, and retained for the configured period — 365 days by
default on EE). Since the setting text points operators at this API for
"pre-existing history", a window overlapping legacy rows would report completion
while silently omitting them.
Per the decision to not export the legacy table, reject instead of silently
omit: try_start now rejects a `from` earlier than the oldest audit_partitioned
timestamp (every legacy row predates the partition cutover, so a `from` at/after
that boundary can never overlap them). Reworded the setting text to scope the
backfill to the partitioned era. Added a regression test, plus an RAII guard
(cubic P2) so the test-only globals are restored even if an assertion panics.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audit): backfill object keys per-window; require trustworthy settled cutoff
Address review (two P1s):
- Object-key overwrite loss: keys were `dt=<day>/audit_backfill_<min_id>.ndjson`.
A narrower, overlapping backfill can start a day's page at the same first row
(same min_id) but hold fewer rows, and `put` would overwrite a broader run's
object — silently dropping the rows only that object held. Include the
requested window in the key so different ranges write disjoint objects (same
window re-runs stay idempotent; consumers dedupe overlapping rows by id). New
regression test (verified red→green).
- Untrustworthy settled cutoff: when min(xact_start) isn't trustworthy (role
lacks pg_read_all_stats/superuser, or a prepared 2PC txn exists), the old
now()-7d fallback could still let an old transaction commit rows inside an
accepted window after the scan, so a "complete" backfill silently missed them.
Since a backfill asserts completeness, reject in those cases instead of
falling back. (The continuous exporter keeps its 7-day fallback — it only
claims bounded lag.)
Also makes the tests robust under the parallel runner: run_backfill takes the
store as a param, so tests pass a local in-memory store (no global
OBJECT_STORE_SETTINGS race) and serialize on the PAGE_ROWS override.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audit): reject backfill overlapping legacy table; regen deref openapi; trim migration comment
Address review (1 P1 + 2 P2):
- Empty-partition backfill (P1): the min(audit_partitioned) guard no-ops when
audit_partitioned is empty, so an upgraded instance with legacy `audit` rows
but no partitioned rows yet would accept a window and complete with zero rows,
silently omitting the legacy rows. Check the legacy `audit` table directly:
reject any window that overlaps a legacy row (subsumes the boundary check and
covers the empty-partitioned case). Test updated accordingly.
- openapi-deref (P2): regenerate openapi-deref.yaml/json (served via include_str!)
so /openapi.{yaml,json} expose the new backfill endpoints.
- Migration comment (P2): trim the PR-history narration to the durable
constraints (why a recent floor and a monotonic cursor are required), per
AGENTS.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: update ee-repo-ref to b821fecccbcba2efed544890576bf2b84321d70d
This commit updates the EE repository reference after PR #634 was merged in windmill-ee-private.
Previous ee-repo-ref: 6b191b77aabcf77658ad4f9031576e0d7b66bf89
New ee-repo-ref: b821fecccbcba2efed544890576bf2b84321d70d
Automated by sync-ee-ref workflow.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
|
||
|
|
796230d90a |
fix(workspaces): add instance setting to disable workspace invite/add emails (#9643)
* feat(workspaces): add skip_email option to invite_user and add_user endpoints The workspace invite_user and add_user API endpoints unconditionally sent notification emails when SMTP was configured, with no way to suppress them per-request. This is noise for automated workflows that programmatically add users to workspaces. Add an optional `skip_email: Option<bool>` field to `NewWorkspaceInvite` and `NewWorkspaceUser`, following the existing pattern on `NewUser` used by POST /api/users/create, and guard the `send_email_if_possible` calls with `if !nu.skip_email.unwrap_or(false)`. The field is optional, so existing clients are unaffected. The auto-add code paths in workspaces_ee.rs (domain-based and instance-group auto-add) are auto-triggered and take no API parameter, so they are left as-is. Fixes WIN-2068 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(workspaces): make workspace invite/add emails toggleable via instance setting Replace the per-request skip_email approach with an instance-level setting `disable_workspace_invite_emails`. When enabled, the email notifications sent by the workspace invite_user and add_user endpoints are suppressed. Useful for instances where users are added programmatically (e.g. CI pipelines that fork workspaces and add users) and the invite emails are noise. Backend: - Add `DISABLE_WORKSPACE_INVITE_EMAILS_SETTING` global setting constant. - Guard the `send_email_if_possible` calls in invite_user and add_user with a read of that setting (via the existing `load_value_from_global_settings` helper). Defaults to false, so existing behavior is unchanged. - Revert the per-request `skip_email` field on NewWorkspaceInvite / NewWorkspaceUser and the corresponding openapi additions. Frontend: - Expose the setting as a boolean toggle in the SMTP tab of the instance settings (superadmin). The auto-add paths in workspaces_ee.rs are unaffected. Fixes WIN-2068 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(frontend): gate disable_workspace_invite_emails toggle behind EE Email delivery (send_email_if_possible) is a no-op outside the EE/private build, so the toggle has no effect on a pure-OSS instance. Add `ee_only: ''` to match the sibling SMTP settings: the toggle is grayed out (with an EE badge) on non-EE instances instead of rendering as an active no-op control. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(frontend): don't EE-gate disable_workspace_invite_emails toggle The earlier ee_only addition was based on the false premise that the workspace invite/add emails are license-gated. They are not: SMTP configuration (SmtpSettings) and email sending (send_email_if_possible) have no enterpriseLicense check — they only require the closed-source build with SMTP configured. The sibling smtp_settings carries ee_only: '' but its smtp_connect field renders no SettingCard label, so that flag is inert (no badge, no disable). On a plain boolean field ee_only is fully active, which incorrectly grayed out the toggle and showed an EE badge. Drop ee_only so the control matches the actual non-license-gated behavior. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
afddfe8445 |
feat(worker): #ssh directive to run a bash script on a remote SSH host (#9479)
* feat(worker): #ssh directive to run a bash script on a remote SSH host Add a first-class `#ssh <resource_path>` bash directive that reroutes a normal bash script to run on a remote host reached over SSH (a jump/utility node) instead of on the worker, with full parity: typed positional args in, structured result out, live streamed logs, cancellation, and remote exit-code propagation. It mirrors the existing `# sandbox <image>` precedent: the directive is parsed in handle_bash_job and reroutes to a specialized handler that reuses handle_child for all execution plumbing. - windmill-common: BashAnnotations::ssh_target() parser (+ unit test) and the ssh_execution_enabled instance setting (off by default) - windmill-worker: reroute hook in bash_executor + ssh_executor_oss shim. OSS returns a clear "enterprise feature" error; the real handler lives in ssh_executor_ee.rs (private feature) and is gated by a valid enterprise license + the instance setting. - examples/usecase/ssh-execution-wrapper: the ssh_target resource type, a userland wrapper (no-license fallback), and a README documenting both paths and the trade-offs vs agent workers. EE companion: windmill-labs/windmill-ee-private (ee-repo-ref.txt bumped). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(worker): ssh host-key opt-in, 0600 key write, instance setting UI * chore: update ee-repo-ref * feat(worker): #ssh $arg form to take the ssh target from a job argument * fix(worker): #ssh token must look like a target; $arg restricted to path strings * fix(worker): tighten #ssh parser to exact directive; add -- ssh destination guard * chore: update ee-repo-ref to d45b9a6cbe40f7fe5d322c850c50f64a6980e4f0 This commit updates the EE repository reference after PR #609 was merged in windmill-ee-private. Previous ee-repo-ref: 2804f1aa8e74b3a7733aeb6f5044d5085193872a New ee-repo-ref: d45b9a6cbe40f7fe5d322c850c50f64a6980e4f0 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
7590b28108 |
feat(sandbox): pull/extract images with crane instead of podman (#9455)
* feat(sandbox): pull/extract images with crane instead of podman (+ add to image)
The sandboxed container runtime (`# sandbox <image>`) only ever pulls + flattens an
image (nsjail does the run), so a full container engine is overkill — and podman was
never actually in any Dockerfile, so the merged feature couldn't run in the shipped
image. Switch to crane (google/go-containerregistry): a single ~25MB static binary,
no daemon/store/root/privileged.
- docker_v2.rs: crane export -> flattened rootfs tar, crane config -> OCI config,
crane digest -> content-addressed rootfs+config cache (cross-job dedup + automatic
freshness), crane manifest -> pre-download size guard. DOCKER_CONFIG authfile dir.
Cache eviction prunes the rootfs-tar cache by mtime (LRU). Pull policy honored via a
ref->digest cache (missing/never reuse without a registry hit).
- Dockerfile + docker/DockerfileSlim{,Ee}: install the crane binary (Full/FullEe and
the EE image inherit it via FROM the base image).
- docs + UI text + instance-setting descriptions updated (download size is compressed;
cache is the rootfs-tar cache).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(sandbox): address CI review — digest-pinned fetch, size cap on every job, eviction race
Codex P1s:
- Fetch by the resolved digest (name@digest), not the mutable tag, so content can't
diverge from the digest the cache is keyed under if a tag moves mid-fetch.
- Enforce the size cap on EVERY job via a cached {digest}.size sidecar (no registry call
on cache reuse), so lowering the limit rejects already-cached oversized images.
- Eviction race: hardlink the cache tar into the job dir before tar -xf (pins the inode
against concurrent eviction) and re-fetch if it was evicted first.
Claude P2s: atomic config sidecar (tmp+rename) + tolerate torn parse; soften the LRU
comment (mtime = creation order); sweep orphaned *.tmp.* and .size on eviction.
+digest_key/ref_key unit tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(sandbox): P1 cross-fs cache staging (EXDEV), Dockerfile arch fail-fast
CI re-review (Claude + Codex P1): the eviction-race hardlink crosses filesystems in the
shipped deployments — the cache is its own volume (/tmp/windmill/cache) while the job dir
is on the container fs — so hard_link returns EXDEV (not NotFound) and every sandbox job
fails. Fall back to tokio::fs::copy on a non-NotFound link error; copy reads through the
source inode so it still survives a concurrent eviction.
Also: Dockerfiles fail fast with a clear error on an unsupported arch instead of building
a 404 crane URL; ref->digest file written via tmp+rename (no torn read under missing/never).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(sandbox): say 'oldest by creation time' not 'LRU' for cache eviction
Codex P2: the code evicts by tar creation time (cache hits don't touch mtime), so the
user-facing docs + instance-setting text shouldn't claim true LRU.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
1727271e19 |
feat: sandboxed daemonless container runtime via '# sandbox <image>' (#9453)
* feat: add sandboxed docker v2 runtime via '# docker <image>' Run a container image as a subprogram of the job's own nsjail sandbox: extract the image rootfs with podman (rootless) and run it chrooted inside the job's nsjail, so the container inherits the job's confinement and is safe under nsjail / for untrusted code. Selected by '# docker <image>'; a bare '# docker' keeps the v1 (dind) path untouched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: default to daemonless docker (drop dind from compose, allow docker on cloud) docker-compose no longer ships the dind sidecar (v2 is daemonless: podman + nsjail in the worker); removed the dind service, DOCKER_HOST env, depends_on and volume. Removed the language-picker guard that blocked Docker scripts on the multi-tenant platform, now that v2 makes docker safe to run sandboxed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: select sandboxed container via # sandbox <image>; add pull policy + size guards - Surface moved from '# docker <image>' to '# sandbox <image>' (groups under the sandbox annotation; '# docker' stays v1-only, '# sandbox' stays nsjail-bash). - SANDBOX_IMAGE_PULL_POLICY (default 'newer') so moving tags don't go stale. - SANDBOX_IMAGE_MAX_SIZE_MB rejects oversized images before extraction. - SANDBOX_IMAGE_CACHE_MAX_MB best-effort LRU eviction of podman's image store. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(sandbox): support # volume, honor nsjail tmp instance settings, v2 docker template - Thread shared_mount into the sandbox container nsjail config so '# volume' mounts (and the same-worker /tmp/shared folder) apply inside the container. - Use resolve_nsjail_tmp_mount_block for the container's /tmp so it honors the same nsjail_tmp_backing / nsjail_tmpfs_size_mb instance settings as other nsjail jobs. - docker-compose comment + the editor's Docker template now use '# sandbox <image>'. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(sandbox): make image size/cache/pull-policy UI instance settings Convert SANDBOX_IMAGE_* from worker env vars to DB-backed instance settings (sandbox_image_max_size_mb, sandbox_image_cache_max_mb, sandbox_image_pull_policy), hot-reloaded via the same mechanism as nsjail_tmpfs_size_mb and configurable in #superadmin-settings. No worker restart needed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(sandbox): windmill-managed registry — default registry + private auth Two new instance settings: - sandbox_image_default_registry: prepended to unqualified image refs (alpine -> <registry>/alpine); fully-qualified refs untouched. - sandbox_registry_auth: docker/podman auth.json blob written to a per-job authfile (0600, removed with the job) and passed to podman --authfile for private registries. Both hot-reloaded and configurable in #superadmin-settings. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(sandbox): protobuf-safe proto_str escaper, atomic 0600 authfile, registry tests Addresses local-review P2s: proto_str now emits valid protobuf octal escapes for control/non-ASCII bytes (not Rust \u{..} that nsjail would reject); the registry authfile is created 0600 atomically (no world-readable window); add a registry_qualified table test + a non-ASCII proto_str case. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(sandbox): P0 — deliver image env via nsjail envar:, never the launcher process env CI review (P0): the image's OCI Env (attacker-controlled keys+values) was applied to the nsjail launcher process via .envs(), so a hostile image could set LD_PRELOAD/ LD_LIBRARY_PATH/LD_AUDIT on nsjail itself and execute code as the worker outside the jail. Now the image env is rendered as proto-escaped 'envar:' directives (child-only) and nsjail's process env carries only windmill-trusted keys (reserved vars + proxy). Also: warn instead of silently bypassing the size guard on inspect failure; reset the eviction guard via a Drop guard (no stuck flag on panic/early-return). +render_envars test. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(sandbox): P0 symlink-write escape via rootfs script; P1 redact registry-auth logging CI review: - P0 (Codex): the body was written into the image-controlled rootfs as .windmill_docker_main.sh via write_file (follows symlinks) — a hostile image could plant that path as a symlink to a host file and capture the worker's write before nsjail starts. Now the body is passed straight to 'sh -c <body> sh <args>'; no file is written into the rootfs at all. - P1 (Codex): sandbox_registry_auth flowed through the generic setting loader which logs the value (raw auth.json credentials). Replaced with a secret-aware reload that loads directly and logs only a redacted 'configured=' message. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(sandbox): redact sandbox_registry_auth in instance-settings write log too The settings API also logs 'Set global setting <key> to <value>' via format_setting_value; add sandbox_registry_auth to SENSITIVE_SETTINGS so the credential is redacted there as well as on reload. * fix(sandbox): don't silently disable cache eviction on podman images parse error Re-review (cubic/Claude P2): serde_json::from_slice(...).unwrap_or_default() meant any parse hiccup (e.g. podman omitting Size/Created via omitempty for a zero value, or schema drift) silently degraded to an empty Vec and disabled eviction with no log. Now Size/Created are #[serde(default)] (a missing omitempty key -> 0, not a whole-array parse failure) and a real parse error warns + breaks instead of being swallowed. --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8bf7fd2c92 | feat(queue): stochastic admission + EE availability of workspace fairness algorithm (#9321) | ||
|
|
de2e243313 |
feat(queue): per-workspace fairness cap on the shared cloud worker pool (#9303)
* feat(queue): cloud-only per-workspace fairness cap on the shared worker pool
On `app.windmill.dev` the cluster runs a single default worker group, so a
single workspace flooding the queue can degrade quality of service for
everyone else. This adds an opt-in mechanism that caps any single workspace
at a configurable share of the shared worker pool when it has been
dominating cluster activity for more than a configurable window.
Detection signal counts both currently-running jobs and jobs completed in
the rolling window, so it catches workspaces hogging slots with long jobs
**and** workspaces spamming many tiny jobs (where no individual job's
started_at is old, but throughput share dominates).
Refresh is coordinated cluster-wide via a single UPDATE on
`background_task_state`: the `WHERE updated_at < now() - interval` predicate
combined with row-level locking means only one process per refresh cycle
actually runs the aggregation, regardless of fleet size. Every other
process gets the freshly written value in the same round trip via
`UNION ALL ... LIMIT 1`. Heavy aggregation rate stays at ~0.2-0.5 qps for
the whole cluster.
Pull queries are split: the existing query string and its bind shape stay
bit-identical to today, so the planner keeps using the same indexes when
fairness is off or no workspace is currently capped. A separate
`WORKER_PULL_QUERIES_FAIRNESS` adds `AND workspace_id <> ALL($2::text[])`
and is only materialized while the feature is enabled.
Hard-gated to `CLOUD_HOSTED=true` + BASE_URL host == app.windmill.dev at
three layers: frontend `cloudonly: true`, API setter rejection in
`set_global_setting_internal`, runtime check in `fairness_active`. Settings
are exposed under Jobs in the instance-settings UI; defaults are off so
the change is a no-op for self-hosted.
Two-pass pull guarantees no worker idling: if every queued job belongs to
a capped workspace, the second pass uses the unmodified pull queries.
Cap re-asserts on the next refresh.
Fixes WIN-1982
* fix(queue): address CI review findings on workspace fairness
Six fixes from the four-reviewer cross-check on #9303:
1. **Aggregation evaluation (Codex P1).** The previous `INSERT ... ON CONFLICT
DO UPDATE WHERE updated_at < ...` had the heavy `v2_job_queue ∪
v2_job_completed` aggregation inlined into `VALUES`, which Postgres
evaluates for every contender to build the proposed row — losing the
"one heavy aggregation per cycle cluster-wide" property the design
advertises. Split into three small statements: (a) cheap claim with
constant `VALUES`, (b) winner-only `UPDATE ... SET value = jsonb_build_object('overloaded', <agg>)`
(Postgres only evaluates `SET` per row matching `WHERE`, so losers never
compute the aggregation), (c) read for everyone. Heavy query now truly
runs ~0.2-0.5 qps cluster-wide regardless of fleet size.
2. **Numeric setting wraparound (cubic P1).** `u64 as u32` and downstream
`u32 as i32` could silently flip sign and feed `make_interval(secs => -N)`,
making `now() - interval` a future timestamp and disabling the
completed-jobs half of the activity signal. Clamp `duration_secs` to
[1, 86400] and `min_total_jobs` to [0, u32::MAX] before storing.
3. **`/instance_config` bypass (cubic/Claude/Codex P2).** Bulk config endpoint
sidestepped `set_global_setting_internal`'s gate; a self-hosted superadmin
could persist `workspace_fairness_*` rows via the bulk path. Mirror the
per-key check in `set_instance_config` upsert flow.
4. **DB error coerced to false (Claude P2).** `load_workspace_fairness_enabled`
collapsed `Err(_)` to `false` and unconditionally swapped the atomic — a
transient DB blip during notify-event propagation toggled the feature off
cluster-wide (and triggered a `store_pull_query` rebuild precisely when load
is highest). Now propagates the error so the atomic stays at its prior value.
5. **Refresh failure cooldown (Claude P2).** Storing `0` removed the rate
limit entirely; every subsequent pull spawned a new refresh task. Leave
`LAST_REFRESH_MICROS` at `now_us` (already written by the CAS) so the
natural interval acts as the cooldown.
6. **Visibility + duplication (Pi P2).** Mark `make_pull_query_fairness` as
`pub(crate)`. Move the duplicated `BASE_URL host == app.windmill.dev`
parser into `windmill-common::worker::is_cloud_production_host` and share
it between the API setter and the runtime path.
Verified locally:
- `POST /api/settings/global/workspace_fairness_enabled` → 400 (per-key gate)
- `PUT /api/settings/instance_config` with fairness key → 400 (bulk gate)
- `cargo check --workspace --features=private,enterprise,quickjs` — clean
Refs WIN-1982.
* fix(queue): second round of CI review nits on workspace fairness
Three issues raised by the Codex/Claude re-review of commit
|
||
|
|
b656dc6cdc |
feat(nsjail): optional disk-backed /tmp via instance setting (#9272)
* feat(nsjail): optional disk-backed /tmp via instance setting * test(nsjail): unit-test tmp mount resolver and narrow visibility * refactor(nsjail): switch tmp backing to select + conditional UI * ui(nsjail): make tmpfs the visible default in /tmp backing select * fix(nsjail): refuse preexisting jail_tmp to block symlink escape * fix(nsjail): allow jail_tmp reuse on sequential nsjail calls Codex flagged that python/ruby/rust executors invoke nsjail twice per job_dir (install then run). The previous resolver treated any preexisting jail_tmp as hostile and silently fell back to tmpfs on the second call, so disk-backed mode never reached the main script run for those langs. Use symlink_metadata().is_dir() to distinguish a real directory left by an earlier call in the same job_dir (safe to reuse) from a symlink or other entity (still refused, as the codebase-tar escape requires). Also loosen the frontend visibility predicate: only hide nsjail settings when job_isolation is explicitly 'none' or 'unshare', so deployments that enable nsjail via DISABLE_NSJAIL=false with no DB setting can still see the controls. |
||
|
|
1169371d48 |
feat: add UV_PYTHON_INSTALL_MIRROR env and instance setting (#9271)
* feat: add UV_PYTHON_INSTALL_MIRROR env and instance setting Allows operators to point `uv python install` at a private mirror of the python-build-standalone releases. Configurable via the `UV_PYTHON_INSTALL_MIRROR` env var or the `uv_python_install_mirror` instance setting, with the env var as the boot fallback and the instance setting taking precedence at reload. Fixes WIN-1966 Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: hoist uv_python_install_mirror binding above sandboxing branch The non-sandboxed uv pip install branch referenced a binding that was only declared inside the sandboxed branch. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: neutral placeholder for uv_python_install_mirror The previous placeholder was the default public URL the setting is meant to redirect away from. A neutral example mirror URL is clearer. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
9111f8908d |
feat(nsjail): make tmpfs size configurable via instance setting (#9261)
* feat(nsjail): make tmpfs size configurable via instance setting Adds a new `nsjail_tmpfs_size_mb` instance setting that overrides the size of the `/tmp` tmpfs mount inside the nsjail sandbox across all languages. When unset, the existing per-language defaults (500MB or 800MB) continue to apply, so no behavior change for existing deployments. The setting is exposed under Settings → Jobs and is read at job execution time, so changes take effect on the next job without a restart. Fixes WIN-1963 Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(nsjail): unify default tmpfs size to 800MB Previously each executor passed its own per-language default (500MB or 800MB) to resolve_nsjail_tmpfs_size. Unify on a single DEFAULT_NSJAIL_TMPFS_SIZE_BYTES constant (800MB) so the placeholder behavior is consistent across languages. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(nsjail): resolve tmpfs size outside ruby download closure The download.ruby config render runs inside a sync closure passed to par_install_language_dependencies_seq, so `.await` on resolve_nsjail_tmpfs_size() was a compile error under the `ruby` feature. Resolve the size once before the closure and capture the string instead. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(nsjail): rename resolver to *_bytes and clarify fallback Addresses CI review feedback: - Rename `resolve_nsjail_tmpfs_size` to `resolve_nsjail_tmpfs_size_bytes` so the returned unit is unambiguous at the call site (cubic P2). - Fix the `NSJAIL_TMPFS_SIZE_MB` doc comment that still said "per-language default" — there is no per-language fallback anymore, all unset values resolve to the unified 800MB `DEFAULT_NSJAIL_TMPFS_SIZE_BYTES` (codex/pi P2). - Expand the resolver doc to call out that `Some(0)` and negative values also fall back, since the match arm is `Some(mb) if mb > 0`. No behavior change. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ba6fb7021b |
feat: export audit logs to a dedicated object store folder (#9207)
* feat: export audit logs to dedicated object store folder * fix: gap-free audit export via snapshot-xmin gate and stable object keys * test: add integration test for audit log object store exporter * fix: cursor audit export on snapshot xmin to prevent id-leapfrog loss * fix: protect audit s3 checkpoint from config sync and bound export interval * fix: anchor audit s3 checkpoint at enable time to not skip first-window rows * fix: anchor first audit export at the enable transaction's xid * fix: use epoch timestamp floor on first audit export run to not drop old backlog * fix: anchor audit export at startup for env-var enable path * fix: anchor audit export via enabling-txn snapshot xmin trigger * fix: bound the bootstrap audit export to MAX_XID_INTERVAL per tick * refactor: store audit export cursor in background_task_state, add status endpoint * docs: align store_audit_logs_s3 setting text with the actual enable-boundary contract * [ee] refactor: move audit s3 export core logic to EE, gate on Enterprise license * chore: update ee-repo-ref to ec3cd353245e1cdf6a290528dbd7f2ac2498386c This commit updates the EE repository reference after PR #579 was merged in windmill-ee-private. Previous ee-repo-ref: 4ffc6d5f874e64d7dc4a147b4e73baa6c44867a5 New ee-repo-ref: ec3cd353245e1cdf6a290528dbd7f2ac2498386c Automated by sync-ee-ref workflow. --------- Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
1d279e7a1e |
feat: add min release age instance settings for bun and uv (#8956)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> |
||
|
|
4cf53a44bb |
feat: add auto-login SSO provider instance setting (#8929)
* [ee] feat: add auto-login SSO provider instance setting Adds an instance-level `auto_login_provider` setting that, when set to an OAuth provider key (e.g. "okta") or "saml", causes the login page to auto-redirect users to the configured SSO flow on mount. Useful for orgs with a single SSO where the provider button grid adds a pointless extra click. - Backend: new global setting constant, read from DB in list_logins handler and returned as the `auto_login` field in the response - Frontend: Login.svelte auto-redirects in loadLogins() when the configured provider is actually present in the response - Escape hatch: `?no_sso=1` skips the auto-redirect and shows the normal login form (admin fallback when SSO is broken) - No redirect loop: if the `error` prop is set (SSO callback failed), the redirect is skipped - Admin UI: new text field under Auth/OAuth/SAML in instance settings Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: skip auto-redirect on /user/login page Auto-redirect should only fire on embeds where the user did not explicitly navigate to a login screen (public app popup, approval pages). Visiting /user/login is an explicit sign-in action — often by an admin who needs password fallback — so we must never hijack it. Gate the logic on a new `autoRedirect` prop (default true). The main login page passes `autoRedirect={false}`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to b7157d55fb9f8d8f7aeb7b1fb69bc935af895a2f This commit updates the EE repository reference after PR #547 was merged in windmill-ee-private. Previous ee-repo-ref: e32a48499d206a24e0c12817b465775321b0ee41 New ee-repo-ref: b7157d55fb9f8d8f7aeb7b1fb69bc935af895a2f Automated by sync-ee-ref workflow. * fix: handle popup-blocked auto-redirect in popup mode When Login is embedded with popup=true (public app), auto-redirect funnels through window.open() without a user gesture — browsers block it by default, leaving the user stuck on "Signing you in…". Detect window.open returning null, clean up listeners, reset autoRedirecting so the provider button grid re-renders, and surface a toast pointing users at the manual button. The grid click retains its user gesture and passes the popup blocker. Also extracts a redirectSaml() helper so the SAML auto-redirect path and the SSO button click share the same logic. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
0cfa131254 |
feat: add disable_password_login global setting (#8873)
Adds an instance-level toggle that hides the email/password form on the login page and rejects password login, password reset request, and password reset endpoints server-side. Useful for OAuth/SAML-only deployments. - New `disable_password_login` global setting + lazy_static AtomicBool - `load_disable_password_login` loader wired into monitor initial_load and notify_global_setting_change listener - Unauthenticated `GET /auth/is_password_login_disabled` endpoint so the login page can hide the password form when enabled - Toggle in Instance Settings → Auth/OAuth/SAML - Login.svelte hides the password form and the "Log in without third-party" toggle when the setting is on Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
4dc54ca3aa |
fix: persist indexer max_index_time_window_secs setting (#8821)
* fix: persist indexer max_index_time_window_secs setting Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: toggle UX for indexer time window cap Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
3f5841f84d |
feat: instance-level ruff config auto-pulled by LSP container (#8803)
* feat: add instance-level ruff config auto-pulled by LSP container Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: move ruff config to new LSP tab in instance settings Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
506b7f55e1 |
fix: zero-downtime coordinated restarts for OTEL and other setting changes (#8768)
* fix: zero-downtime coordinated restarts for OTEL and other setting changes Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: use background_task_state for server heartbeats and fix stale heartbeat detection Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: show restart propagation toast when saving settings that trigger server restarts Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
c09a4311fd |
fix: remove stale KMS openapi/description, restore stripped doc comments
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
09bbc18bb7 |
feat: add AWS Secrets Manager as secret storage backend (Beta) (#8734)
* feat: add AWS KMS as secret backend (EE) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: switch from AWS KMS to AWS Secrets Manager as secret backend Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test: add AWS Secrets Manager integration tests (requires LocalStack) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: mark AWS Secrets Manager as beta Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: remove leftover KMS handler functions from api-settings Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to include AWS Secrets Manager EE impl Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: use full commit hash in ee-repo-ref.txt Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * sqlx --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
02d0ee9198 | feat: add object storage usage view and manual log cleanup (#8724) | ||
|
|
dcd615fdc3 |
feat: add Azure Key Vault as secret storage backend (#8704)
* feat: add --main flag to write_latest_ee_ref.sh to point to latest EE main Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: add Azure Key Vault as secret storage backend (EE) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref.txt to azure-key-vault-support branch Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: add token auth, insecure TLS for emulator, and integration tests Adds optional `token` field to AzureKeyVaultSettings for direct Bearer auth (bypasses OAuth2), enables self-signed cert acceptance in token mode, and includes 4 integration tests against the Azure KV emulator. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref.txt Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: handle Azure KV soft-delete and emulator quirks - Purge soft-deleted secrets after delete to allow name reuse - Retry set_secret on 409 Conflict (purge stale soft-deleted secret) - Accept self-signed certs when using static token (emulator mode) - Work around emulator version-ordering bug in CRUD test Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref.txt Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to 47b0d9d5d163efdab1e145ee012bdb2eb1373b78 This commit updates the EE repository reference after PR #511 was merged in windmill-ee-private. Previous ee-repo-ref: d432d78bda151d611d8065162de7c1b7edce92e9 New ee-repo-ref: 47b0d9d5d163efdab1e145ee012bdb2eb1373b78 Automated by sync-ee-ref workflow. * fix: accept token OR client_secret in Azure KV validation, add token UI field - isAzureKvConfigValid() now accepts either client_secret or token - Added token input field to the Azure KV config form for emulator/dev use Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
9ceab730d7 |
feat: add DB health diagnostic dashboard for superadmins (#8574)
* feat: add DB health diagnostic dashboard for superadmins Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * Update SQLx metadata * fix: improve db health query performance Bound large_results scan to last N jobs (configurable via scan_limit query param, default 10K) instead of full-table pg_column_size sort. Replace N+1 datatable size queries with single batched pg_class lookup. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Update SQLx metadata Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * sqlx --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> Co-authored-by: Ruben Fiszel <ruben@windmill.dev> |
||
|
|
935fb44c84 |
fix: GitHub Enterprise Server support for self-managed GitHub Apps (#8507)
* fix: GitHub Enterprise Server (GHE) support for self-managed GitHub Apps - Fix GHE installation URL: use /github-apps/ path instead of /apps/ for non-github.com hosts - Fix double decodeURIComponent on OAuth state param (URLSearchParams already decodes) - Add client_id to self-managed GitHub App validation - Bump hub scripts to GHE-compatible versions (sync, test, init, clone) - Bump LATEST_GIT_SYNC_SCRIPT_PATH to hub/28176 - Rename "GitHub Enterprise App" → "GitHub App" in UI labels (it works for both) - Formatting fixes in GhesAppSettings.svelte and gh_success page EE ref: windmill-labs/windmill-ee-private@09c9ed1 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * Update SQLx metadata * fix: handle GHE Cloud (*.ghe.com) app installation URL path GHE Cloud uses /apps/ like github.com, not /github-apps/ like self-hosted GHES. Docs: https://docs.github.com/en/enterprise-cloud@latest/apps/using-github-apps/installing-a-github-app-from-a-third-party Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: handle GHE Cloud (*.ghe.com) installation URL and update ee-repo-ref GHE Cloud uses /apps/ like github.com, not /github-apps/ like self-hosted GHES. Docs: https://docs.github.com/en/enterprise-cloud@latest/apps/using-github-apps/installing-a-github-app-from-a-third-party Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: update hubPaths to deprecate 28176 and use 28180 as latest sync script Aligns with main's LATEST_GIT_SYNC_SCRIPT_PATH bump in PR #8532. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * chore: update ee-repo-ref to 6bb0ff0 (includes GHE fixes) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
9b3e558d84 |
feat: add instance setting to enforce workspace prefix for HTTP routes (#8528)
* feat: add instance-level setting to enforce workspace prefix for HTTP routes
Add `http_route_workspaced_route` instance setting that forces all HTTP routes
to use workspace prefix (`/api/r/{workspace_id}/{route}`), mirroring the existing
`app_workspaced_route` setting for apps.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: bump http trigger version on setting change to invalidate route cache
The route cache is version-based, not TTL-based. Without bumping the
version sequence when the instance setting changes, cached routes would
continue serving with the old prefix behavior until a route is
created/updated/deleted or the server restarts.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: immediately refresh HTTP routers on setting change
The route cache polls every 60 seconds, but bumping the version sequence
only makes the next poll pick up changes. Explicitly call refresh_routers
after the setting reload so routes are rebuilt immediately.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
||
|
|
db5e03610d |
feat: add instance-level AI settings (#8453)
* feat: add instance-level AI settings with workspace fallback Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add AI step to onboarding setup wizard Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: thread workspace prop through resource editor and disable chat offset Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * Revert "fix: thread workspace prop through resource editor and disable chat offset" This reverts commit 9fea9cc0c239f6432d1fef1487c45e74ab752e21. * fix: set workspace store and disable chat offset during AI setup step Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: thread workspace and disableChatOffset props through resource editors Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: populate workspace and user stores for AI step path component Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: initialize AI clients for test key during onboarding Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor: extract AI config state into InstanceAISettings component Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor: move AI config state ownership into AISettings component Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * Persist instance AI settings before navigation * Reload effective workspace AI state after save * Scope AI key tests to the rendered workspace * Add post-create AI onboarding for new workspaces * Unify instance AI settings header * Fix instance AI drawer offset on workspace selection * Add instance AI fallback settings behavior * Update sqlx metadata * Update sqlx metadata * Clarify active instance AI in workspace settings * Refresh workspace AI state after instance AI save * Declare instance AI summary in API schema * Normalize empty instance AI config handling * Clean up workspace AI settings UI * Unify AI config provider checks * Split AI settings metadata from effective config * Propagate instance AI cache invalidation across servers * Fix AI settings dirty state tracking * Update sqlx metadata --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
7d9fb57368 |
feat: DB-backed instance events webhook with superadmin UI (#8402)
* feat: make instance events webhook URL configurable via superadmin UI The instance events webhook was previously only configurable via the INSTANCE_EVENTS_WEBHOOK env var, requiring a restart to change. This adds a DB-backed global setting with a UI in superadmin settings under Monitoring > Webhooks, while keeping the env var as an override. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review - prometheus timer bug and cleaner cache init - Bind prometheus timer to `let timer` and call `stop_and_record()` after the POST (was silently discarded before) - Use `Option<Instant>` with `map_or` instead of `checked_sub` trick for clearer "not yet read" semantics Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: remove env var mention from webhook setting description Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: list all instance events explicitly in webhook description Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: restore send_instance_event guard with AtomicBool for DB setting Use a shared Arc<AtomicBool> between send_instance_event and the event loop so we skip channel sends when no webhook is configured (env or DB). Starts optimistic (true) so the first event triggers a DB read, then the loop updates it after each cache refresh. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: use static AtomicBool + notify handler for webhook guard Replace the Arc<AtomicBool> instance field with a global static INSTANCE_EVENTS_WEBHOOK_DB_ENABLED, updated by the notify_global_setting_change handler in main.rs. This follows the established pattern (like REQUIRE_PREEXISTING_USER_FOR_OAUTH) and avoids the deadlock where the bool could never flip back to true. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: single Arc<RwLock<Option<String>>> for instance webhook URL Replace the separate INSTANCE_EVENTS_WEBHOOK env var lazy_static and INSTANCE_EVENTS_WEBHOOK_DB_ENABLED AtomicBool with a single shared variable. Initialized from env var, then the reload function overwrites from DB (falls back to env var when DB has no value). Follows the same pattern as SCIM_TOKEN and other settings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Ruben Fiszel <ruben@windmill.dev> |
||
|
|
73fe45b6cb |
feat: workspace-specific registry overrides (#8406)
* feat: add workspace-specific registry overrides Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: move workspace registries to end of registries tab Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: workspace overrides use field selector instead of showing all fields Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: polish workspace registries UI to match design guidelines Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: show field selector directly and fix addField initialization logic Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: namespace pip_resolution_cache by workspace when registry overrides exist Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: namespace binary/bundle caches by workspace when registry overrides exist Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * perf: zero-cost cache suffix when no workspace overrides exist Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: reload workspace_registries via notify events on setting change Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review findings - Fix discardCategory not reverting workspace_registries changes - Fix get_no_default: convert to async fn with owned Uuid param - Fix append_logs: use windmill_queue import already available - Fix ruby URL parsing: support both comma and whitespace delimiters - Add WorkspaceRegistryMap type alias to reduce inline type noise Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * all --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
372023e995 |
feat: add ws_base_url instance setting for WebSocket URL override (#8405)
* feat: add ws_base_url instance setting to override WebSocket base URL Allow deployments behind reverse proxies to route WebSocket traffic (LSP, debugger, multiplayer) to a different host/port than the main frontend via a new instance setting. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: move ws_base_url to Advanced section with toggle and connectivity test - Move setting from Core to Advanced > WebSocket section - Render as toggle "Custom websocket base url from frontend to multiplayer/lsp/debugger" with conditional URL text field - Add Test connectivity button (always visible) that checks HTTP health and WebSocket ping for all three services (LSP, Multiplayer, Debugger) - Add /ws/ping and /ws/health endpoints to LSP service - Add /ws_mp/health HTTP and __ping__ WS handlers to multiplayer service - Add /ping WS handler to debugger service - Add CORS headers to health endpoints for cross-origin testing Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: toggle enabled check and testWs promise resolution - Fix enabled derived to check only for null (not empty string), otherwise the toggle never turns on since toggleEnabled sets '' - Fix testWs onclose handler to resolve(false) so the promise doesn't hang if the server closes without sending a message Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: make connectivity test work with existing services - HTTP test: accept plain text "ok"/"okay" (old services) in addition to JSON {"status": "ok"} (new services), reject HTML (SPA fallback) - WS test: resolve on onopen (connection established) instead of waiting for a specific pong message, so the test works even with services that don't have the new /ping handler yet Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
2e430c4c0b |
feat: add GitHub Enterprise Server (GHES) support for GitHub App git sync (#8344)
* feat: add GitHub Enterprise Server (GHES) support for GitHub App git sync Add a self-managed GitHub App mode alongside the existing managed (stats.windmill.dev) mode, enabling git sync for GitHub Enterprise Server and custom GitHub App installations. Backend: - Parameterize GitHub API URLs (no more hardcoded github.com) - Add GITHUB_ENTERPRISE_APP_SETTING global setting - Add OpenAPI specs for ghes_installation_callback and ghes_config endpoints Frontend: - Add instance settings UI for configuring self-managed GitHub Apps with setup instructions and validation - GHES installation flow in gh_success page - Dynamic installation URL based on GHES config - Increase git sync test connection timeout to 10s - Block "Review changes" save when settings are invalid EE companion PR: windmill-labs/windmill-ee-private#<PR_NUMBER> Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * chore: update ee-repo-ref to c74c86b78a66b976fd9968b21f77903723e668ec This commit updates the EE repository reference after PR #459 was merged in windmill-ee-private. Previous ee-repo-ref: 45e4550110799525b5502cf072c8af8132492638 New ee-repo-ref: c74c86b78a66b976fd9968b21f77903723e668ec Automated by sync-ee-ref workflow. * sqlx --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> Co-authored-by: Ruben Fiszel <ruben@windmill.dev> |
||
|
|
fe1519f128 |
feat: support minimal telemetry mode (#8243)
* feat: support minimal telemetry mode for EE When EE customers disable telemetry, send a reduced payload with only license-compliance data instead of ignoring the setting. Job usage data is excluded in minimal mode. The telemetry settings UI now shows in EE with context-appropriate descriptions for both CE and EE. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * chore: update ee-repo-ref for telemetry-minimal Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: make telemetry toggle label and description license-aware Show "Minimal telemetry" with EE-specific description on EE, and "Disable telemetry" with CE-specific description on CE. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * chore: update ee-repo-ref Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * Update commit hash in ee-repo-ref.txt * Update reference hash in ee-repo-ref.txt * chore: update ee-repo-ref to 2f52c015bc6c81391234fa87b27ee1d4cd3a48a3 This commit updates the EE repository reference after PR #440 was merged in windmill-ee-private. Previous ee-repo-ref: 3628ed51426d8d29b3d5c62864ba256b7f9eab17 New ee-repo-ref: 2f52c015bc6c81391234fa87b27ee1d4cd3a48a3 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
2aef01d18c |
feat: partition audit log table by day with configurable retention (#8292)
* feat: partition audit log table by day with configurable retention Introduce daily range partitioning for audit logs to replace expensive DELETE-based retention with instant DROP TABLE per partition. - Create `audit_partitioned` table alongside existing `audit` table - New inserts go to `audit_partitioned`, reads UNION ALL both tables - Monitor creates future partitions and drops expired ones - Add `audit_log_retention_days` instance setting (default 365 days) - Old `audit` table empties naturally via existing DELETE cleanup Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add audit log retention setting to Core instance settings UI Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: bump audit partitioning migration timestamp to avoid collision Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * chore: update ee-repo-ref.txt for audit partitioning Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add RLS/grants to audit_partitioned, run partition mgmt hourly, CE default 14d - Add grants for windmill_user/windmill_admin and all 5 RLS policies - Move manage_audit_partitions to hourly via should_run(120) - Default retention: 14 days CE, 365 days EE - Download JSON button is now icon-only Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address code review — quote SQL identifiers, add workspace index, deduplicate retention logic - Quote partition names in dynamic SQL for defense in depth - Add idx_audit_partitioned_workspace(workspace_id, timestamp DESC) index - Extract audit_log_retention_days() helper to deduplicate retention logic Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * chore: update ee-repo-ref for audit insert error handling Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * chore: update ee-repo-ref to cef4dfc45e6d6344c5d8d107bd2b4d1bf9bbdd64 This commit updates the EE repository reference after PR #450 was merged in windmill-ee-private. Previous ee-repo-ref: f09284bb257d461bcbe3c50fe31eb6f1e7eafee5 New ee-repo-ref: cef4dfc45e6d6344c5d8d107bd2b4d1bf9bbdd64 Automated by sync-ee-ref workflow. * fix: create audit partitions on startup in initial_load Ensures partitions exist before any requests arrive, closing the gap between server start and the first hourly monitor run. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
0c4d72cfe3 |
feat: add indexer time window setting (default 7 days) (#8290)
* feat: add indexer time window setting (default 7 days) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add time window note to search UIs Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: fetch indexer time window from API in search UIs Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * chore: update ee-repo-ref to 9df755c57fbfc88f4a724e1ea51b1d5f5af4fe52 This commit updates the EE repository reference after PR #447 was merged in windmill-ee-private. Previous ee-repo-ref: c17f16bf45091272974e3aa8009cdf5cc15669bf New ee-repo-ref: 9df755c57fbfc88f4a724e1ea51b1d5f5af4fe52 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
f67b8159ad |
warn about missing <clear /> in nuget config and make description optional (#8281)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
e56ccd200b |
feat: token expiration notifications (#8190)
* feat: add token expiration notifications via email, critical alerts, and webhooks - Monitor loop checks for tokens expiring within 7 days and sends email notifications to token owners. Tracks notification state via new `expiry_notified` column on the token table to avoid duplicates. - When tokens expire and are deleted, owners are also notified. - Critical alerts (in-app UI) are gated behind a new instance setting `critical_alerts_on_token_expiry` (off by default); emails are always sent regardless of the setting. - Add TokenExpiringSoon and TokenExpired webhook message variants for workspace webhook integrations. - Frontend: show expiration badges and a warning banner on the tokens table for tokens expiring within 30 days. - Exclude session and ephemeral tokens from all notifications. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor: use separate token_expiry_notification table for dedup - Replace `expiry_notified` column on token table with a dedicated `token_expiry_notification` table (token, expiration) - Insert notification row on token creation via shared `register_token_expiry_notification()` helper - Delete notification row atomically when sending the notification - Clean up orphaned rows in `delete_expired_items()` - No FK constraint to avoid cascade overhead on token deletions - Add index on expiration column for efficient range queries Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: calendar-based expiration badge and move notification cleanup - Fix daysUntilExpiration to compare calendar dates instead of time diff - Move notification row cleanup from delete_expired_items to check_expiring_tokens to keep it off the hot path - Use simple expiration <= now() index scan instead of NOT EXISTS join Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
63ebae8829 |
feat: replace hub error toasts with warning alerts and add disable hub setting (#8225)
* feat: replace hub error toasts with warning alerts and add disable hub setting Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: guard hub script cache refresh when hub is disabled Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
c70307d3f2 |
fix: show sync endpoint timeout setting on all instances (#8170)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> |
||
|
|
13daebf88a |
fix: restore email domain (MX) setting in instance settings UI (#8152)
The email_domain setting was accidentally removed from the frontend instance settings in a recent onboarding cleanup. The backend still fully supports it. This restores the setting in the Core section. Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
9eb15312f6 |
feat: add .npmrc support for private npm registries (#8039)
* feat: add .npmrc support for private npm registries Add a new `npmrc` instance setting that accepts full .npmrc file content for configuring private npm registries. Works with bun (native .npmrc support since 1.1.18), deno (native .npmrc support in 2.x), and the npm proxy (parses default registry + auth token from .npmrc). Legacy `npm_config_registry` and `bunfig_install_scopes` fields are now hidden when empty, so new users only see the .npmrc field. Also fixes a pre-existing race condition where gen_bunfig was called after start_child_process. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * all --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
57ca7dbca0 |
improve instance settings drawer UX (#8002)
* fix(frontend): prevent false dirty state in instance settings on load Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(frontend): handle undefined python version in select binding Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor(frontend): extract SaveButton component and improve drawer header UX Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor(frontend): replace inline diff with diff drawer and simplify save flow Save now saves immediately instead of requiring a two-step confirm flow. Diff view opens in a separate drawer with split/unified toggle instead of replacing the form content inline. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(frontend): preserve dirty state when toggling YAML mode in instance settings syncFormToYaml() was setting yamlCodeInitial to the current modified YAML, causing hasUnsavedChanges to become false when entering YAML mode with pending form changes. Build yamlCodeInitial from initialValues instead. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(frontend): clear dirty state after saving in YAML mode Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * reduce save button timeout * feat(frontend): add review changes button to unsaved changes confirmation modal Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(frontend): address code review issues from PR #8002 Remove unnecessary IIFE wrappers in handleSave/handleSaveAndCloseDiff, fix stale on:close reference on diff drawer, clip SaveButton overlay with overflow-hidden, make DiffEditor respond reactively to inlineDiff prop instead of using {#key} destroy/recreate, and revert normalizeValue object check to original simpler behavior. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(frontend): remove tab-switch confirmation modal in full settings mode In full mode, the save button saves all settings across all categories, so switching tabs cannot lose unsaved changes. Remove the per-category dirty check, confirmation modal, and unused ConfirmationModal import. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(frontend): prevent SMTP toggles from creating false dirty state Use getter/setter bind:checked so Toggle reads undefined as false without writing it back to the store. This prevents visiting the SMTP tab from mutating smtp_settings and triggering a false unsaved diff. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(frontend): prevent OTEL toggles from creating false dirty state Same fix as SMTP toggles: use getter/setter bind:checked so Toggle reads undefined as false without writing it back to the store. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor(frontend): use recursive normalizeValue for dirty state instead of per-component fixes Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor(frontend): replace save button with always-visible review changes button Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix(frontend): address PR review comments on DiffEditor and SaveButton Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |