WindmillFinder's ModuleSpec lacked origin, so __file__ was never set on
loaded modules. inspect.getfile() then raised "is a built-in module",
breaking typeguard's @typechecked and anything else that introspects
module source. Use spec_from_file_location() which sets origin correctly.
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: hugocasa <hugo@casademont.ch>
* perf: cap resource content sent to the search modal
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: address review — fence the LATERAL, flag partial search, add cap test
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: pluralize the truncation notice and link the cap to its openapi doc
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: drop sampling params on Claude models that reject them
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: scope the sampling-param claim to what was probed and split the bedrock test
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: build the disable body through the resolver instead of asserting a rejected shape
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: use the Gemini 3.1 Pro id that actually resolves
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: Bedrock Sonnet 5 cannot disable thinking, unlike the native API
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: confine path-scoped jobs:run tokens to their runnable's jobs
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: project singlestepflow onto its runnable and confine kind-only run scopes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep every by-id job read reachable by a jobs:run token
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: whitelist the dbt and wac-approval by-id job reads for run tokens
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: let an apps:run scope satisfy job-read confinement for that app's runs
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: apply run-scope confinement on top of the approval-token read bypass
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: confine the resume-secret job reads to the run scope as well
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: let the global AI chat call connected MCP servers as the user
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: address review findings on the chat MCP tools
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: connect MCP servers from a predefined list in chat and agent steps
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: show the OAuth redirect URL in the instance connect settings
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: clarify the OAuth redirect URL copy in instance settings
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: match the instance settings warning style and drop the redirect tooltip
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: use the standard warning alert for the redirect url mismatch
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: correct the GitHub token guidance in the MCP registry
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: warn when an OAuth connect lacks the scopes an MCP server needs
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: request the connect's scopes when the oauth popup is opened directly
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: connect an oauth-app MCP server without leaving the panel
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: seed connect scopes from the instance config only
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: make the chat use only the MCP servers you turn on
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor: align the MCP connect UI with the design system
* feat: make a pasted url the default way to connect an mcp server
* feat: show provider icons on the suggested mcp servers
* fix: make both mcp sign-in paths behave the same and stop reloading on toggle
* fix: clarify the mcp tool step's server field and drop its info alert
* fix: name the mcp resource in the tool step and move the transport note into the connect box
* fix: drop the redundant description on the mcp resource field
* fix: make the mcp connections trigger icon-only
* fix: scope enabled mcp servers to the account and address review nits
* fix: wait for connect scopes and create session connections in the operating workspace
* feat: move mcp connections into the chat's plus menu and fix review findings
* fix: show mcp servers as checkboxes so off reads as a state
* feat: give menu rows an on/off switch and use it for mcp servers
* fix: lead the mcp menu rows with the switch
* feat: keep the menu open while toggling and simplify the connect card
* fix: ask for the server before the credential in the connect card
* fix: show one credential path at a time in the connect card
* fix: label the path field and move token guidance into its tooltip
* fix: open straight into connect and keep the server menu scannable
* feat: warn when an mcp connection lands outside your own space
* refactor: require the workspace on the mcp connect components and rename the oauth child
* fix: replace the oauth variable on reconnect and bound every mcp result
* feat: show a connected server's provider icon in the connections list
* feat: resolve mcp provider icons from the url and clarify the path field
* style: align the mcp connect card with the design system surfaces
* style: drop the redundant oauth support line and name the scopes oauth scopes
* feat: keep the mcp connect card open in the connections drawer
* feat: preopen the mcp connect card under the agent step resource picker
* feat: resolve a typed mcp url to its registry entry and describe the token field
* style: name both mcp connect actions connect
* style: name the mcp oauth actions connect with the provider
* style: say in the path description what the connect action will save
* style: name the resource type in the mcp connect path description
* feat: cache mcp provider icons and confirm disconnect in a modal
* fix: keep the mcp menu switches live and the disconnect modal above the drawer
* style: fall back to the plug icon in the mcp menu rows
* fix: never destroy a foreign variable or resource when connecting an mcp server
* fix: prove a token variable is ours before writing it and bound mcp search failures
* fix: pin an mcp oauth popup to the target it was opened for
* fix: bind an mcp credential to the server and popup it was requested for
* fix: bound mcp tool calls with a deadline and drop stale server listings
* fix: keep the disconnect confirmation handler returning void
* fix: tie the mcp tool cache to the resource revision and the grant to its scopes
* fix: verify mcp read-only server-side, keep oauth connector mounted
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(telemetry): extend feature-usage tracking to long-tail features
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: describe telemetry as product feature usage rather than AI usage
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor(telemetry): trim disclosure copy and drop unused pick origin
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(telemetry): count trigger fires per run and key hub picks from hub data
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(telemetry): slugify hub keys and order both writers' upserts
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(telemetry): key native trigger adoption by service so it matches fires
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: update ee-repo-ref for native trigger adoption fix
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor(telemetry): move feature-usage collection into the ee crate
* docs: point feature-telemetry at the moved registry and rust writer
* docs: correct the trigger-fire gate comment to match measured step counts
* docs: put the private-build caveat on the verification step
* chore: update ee-repo-ref to f079db9e7962a413b349c4ff8036080894f30771
This commit updates the EE repository reference after PR #725 was merged in windmill-ee-private.
Previous ee-repo-ref: 055adb80416f9339c9a28ae7fbaeadad30d74959
New ee-repo-ref: f079db9e7962a413b349c4ff8036080894f30771
Automated by sync-ee-ref workflow.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
* refactor: combine the per-minute counters onto one shared helper
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep dashmap in windmill-store for the azure devops token cache
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor: name the sweep counter for what it counts
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: refresh AI provider model defaults and capability metadata
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: send explicit thinking disable for Claude and cap Opus 4.1 output
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: resolve mistral-medium-latest window and OpenRouter Claude 5 off
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: cover au. bedrock geo and Fable 5 caching, revert unverified mistral ladder
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: scope the Anthropic explicit disable to models that think by default
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: translate the reasoning off sentinel on the backend Anthropic path
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: translate the reasoning off sentinel on the Bedrock Converse path
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor: share the reasoning off sentinel and make its translation testable
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* perf: read global_settings once per settings-load pass
`initial_load` reads several dozen settings back to back, one
`SELECT value FROM global_settings WHERE name = $1` each: 50 serialized round
trips before a worker is ready, 32 before a server is. On localhost that is
~20ms and invisible; against a real database it is 50x the RTT per process
start, which `EXIT_AFTER_N_JOBS` turns into a per-job cost.
`with_global_settings_snapshot` reads the whole table (12 rows on a typical
instance) into a tokio task-local, and `load_value_from_global_settings`
serves from it. Scoping it to the task is what keeps the single-setting
reload paths correct: a `notify_global_setting_change` event for one key runs
outside any scope and still reads the database, so a live settings change
reaches a running worker as before. Agent workers hold an HTTP connection
with no snapshot to take and are unchanged.
`load_smtp_config` and `reload_custom_tags_setting` had their own inline
copies of the same query; they go through the shared loader so they land in
the snapshot too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: state the snapshot contract on the reader and the query
`load_value_from_global_settings` is called from ~10 crates and one of them
writes a setting then immediately re-reads it through
`reload_custom_tags_setting`; say on the function itself that a scope, when
one is installed, serves the read and leaves `db` unused.
The query comment claimed the table is a handful of rows. It is not bounded
that way: `workspace_dependencies_map_rebuilt:<workspace_id>` adds a row per
workspace and never removes it. Those dynamically named rows are also why the
snapshot fetches the whole table instead of the wanted names, so state that
as the reason.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: bound the settings snapshot and keep it out of two reads
Three review findings, all real:
The snapshot fetched the whole table, which is not bounded by the settings
that exist: `workspace_dependencies_map_rebuilt:<workspace_id>` adds a row per
workspace with no cleanup path, and no settings pass reads one. It now fetches
only statically named rows, and reads of a `<prefix>:<id>` name skip the
snapshot and go to the database. Correctness does not rest on that naming
convention — a colon-free dynamic name would simply be in the snapshot and
still answered correctly — only the bound does.
A snapshot query that failed inside an enclosing snapshot awaited the body
bare, so its reads were served by the outer snapshot rather than falling
through as documented. The task-local carries an explicit bypass state and the
failure path scopes it.
`reload_jwt_secret_setting` decided whether to generate-and-upsert the JWT
secret from a snapshot-served read, so a replica booting alongside another
could overwrite the secret it had just generated and invalidate its tokens.
That read goes through the new `load_value_from_global_settings_fresh`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the snapshot query on the primary-key index
`name NOT LIKE '%:%'` bounded the rows returned but not the work: a leading
wildcard cannot use the index, so Postgres read every row anyway. Against
50k dynamically named rows it plans as a seq scan of 516 buffers whether or
not seqscans are enabled — and worker connections disable them, so the plan
was one the query shape forbade rather than one the planner chose.
`name = ANY($1)` over an explicit list plans as a bitmap index scan, 7
buffers, bounded by the listed names rather than by table size. That list is
also exactly the set the snapshot may answer from, so a name outside it falls
through to the database instead of reading as unset: listing a setting is a
performance choice, never a correctness one, which is what keeps the list
safe to maintain by hand.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor: declare a settings pass instead of reading one setting at a time
Replaces the prefetch-list snapshot with a pass the call sites build
themselves. `SettingsPass` collects the reads `initial_load` will make as
`(name, applier)` pairs, fetches them together, then replays the appliers in
declaration order.
Declaring is what makes the batch exact. The same `if server_mode` /
`if *CLOUD_HOSTED` / `cfg` branches that used to guard a read now guard a
declaration, so the fetch asks for what this process needs and nothing else,
and there is no list of setting names to keep in sync with anything.
Ordering is preserved end to end: appliers run in the order they were
declared, and non-setting work in the middle of the sequence keeps its place
as a step, so nothing moves and nothing runs twice. Steps that need several
settings at once take them together.
The batch distinguishes three states where a per-setting read only ever
produced two at a given call site:
- a value,
- genuinely unset, which several settings must see in order to restore a
default when the setting is cleared,
- could not be read, which must leave the in-memory value alone. Collapsing
this into "unset" would let one failed query reset workspace fairness and
the queue caps across a cluster.
Over HTTP the reads go out together rather than sequentially, so an agent
worker's settings load costs one round instead of ~36, with no new endpoint.
A setting an agent may not request still resolves to unset, as the
per-setting call returned for it.
`reload_*` keeps working per setting for the notify path, sharing its apply
half with the pass. The wrappers no caller was left using are dropped.
worker startup: 50 queries -> 2 (the batch, and jwt_secret which stays its
own read so the pass cannot sit between reading it absent and upserting a
replacement over another replica's).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: run the pass's non-setting steps in declaration order too
Review round found the settings pass had a gap: the reads were declared but
the work interleaved between them still awaited inline, so it all ran before
`pass.run` applied anything.
`manage_audit_partitions` therefore saw `AUDIT_LOG_RETENTION_DAYS` at its
compile-time default rather than the configured value, and dropped every
partition past that default. An instance keeping 30 days on CE lost the
14-to-30-day band on startup and on every full-reload tick. The
`STORE_AUDIT_LOGS_S3` export anchor had the same cause: the gate read `false`
before the setting applied, so an env-var-enabled export never anchored and
its first tick skipped the rows committed before it.
`action` exists so a step keeps its place in the sequence; every remaining
inline await is now one, which fixes both and leaves no phase where a read
can observe a value the pass has not applied yet.
Two more from the same round:
A batch that fails as a whole now falls back to per-setting reads. Skipping
every applier preserves known-good state on a reload tick, but a starting
process has none, and would have run on compile-time defaults until the next
full reload twelve hours later.
`FORCE_RUBY_REPOS` is honored again: the batched url-list path parsed without
the `FORCE_` check its per-setting counterpart applied, so the override was
silently dropped. `load_setting_value` never had one, so the third helper was
never affected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: declare the object-store and worker-config steps in the pass too
Two awaits were left running ahead of `pass.run`, so the settings they read
were still at their compile-time defaults.
The object-store reload is the one that matters: an AWS OIDC store mints its
first token against an issuer built from `BASE_URL` (`oidc_ee.rs`), and with
`OTEL_ENVIRONMENT` set nothing loads that before this pass does, so the store
signed with the unset default, left `OBJECT_STORE_SETTINGS` empty and fell
back to the ten-second retry while startup carried on.
`reload_worker_config` calls `store_pull_query`, which reads the workspace
fairness knobs. It happened to converge because the enabled flag re-stores the
query when it changes, but it was reading defaults on the way there.
Both are steps now, which is also what the earlier fix should have covered:
the only await left outside a step is `pass.run` itself.
Also from the same round: `fetch_settings_batch`'s doc comment had been
stranded on the helper inserted above it, and the batch-failure fallback
re-ran the same reads on an agent worker, where the batch already is the
per-setting read.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: point the setting-loader docs at functions that still exist
`reload_setting` went with the other wrappers no caller was left using, but
two doc links still referenced it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: decide the jwt secret in sql so the read can be batched
`reload_jwt_secret_setting` generated a secret whenever its read came back
absent or unparseable, and upserted it unconditionally. Two replicas booting
against an empty row therefore each installed their own and rejected each
other's tokens, and the same happened on a running cluster whenever the row
was deleted or set to a non-string. Keeping the read next to the write kept
the window narrow but never closed it, and it was the reason this one setting
could not go through the settings pass.
`get_or_create_jwt_secret` puts the decision in the statement instead:
INSERT ... ON CONFLICT (name) DO UPDATE SET value = EXCLUDED.value
WHERE jsonb_typeof(global_settings.value) <> 'string'
RETURNING value
First writer wins, a usable secret is never overwritten, and an empty
RETURNING is how a caller learns another process's secret stands. The `WHERE`
also keeps a normal startup from writing at all, which matters because
`notify_global_setting_change` fires on every write to this table and an
unconditional upsert would have made each start trigger a cluster-wide reload.
Because the statement decides rather than the caller's read, a stale value is
harmless and `jwt_secret` is now an ordinary declaration. Worker startup is a
single batch round.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep a failed read from dropping a FORCE_ override or clearing a setting
Two ways a read that did not succeed was being treated as an answer.
A `FORCE_` override used to be checked before the read, so a failed read
could not affect it. Moving that check into the parser put it behind a value
arriving, and a failed read skips its applier, so a forced private registry
fell back to the public index and a forced `settings.xml` was deleted from
disk by the Maven step that follows it. Forced settings are declared as steps
with no read now: the override outranks the database, so there is nothing to
fetch and nothing to lose when a fetch fails.
The setting loaders were passing `v.ok().flatten()` to their appliers, which
turns a database error into "unset". Most appliers ignore `None`, but
`apply_tag_per_workspace_workspaces` clears the workspace whitelist with it,
making every workspace eligible for per-workspace tags, and
`apply_fork_workspace_tag_append_fork_suffix` stores `false`. Both are also
reached from the notify handlers, so a blip during a reload changed routing
for the cluster. They take `?` now, as the code they replaced did by leaving
the error arm empty, and the other five are converted with them so an applier
that later grows a `None` branch cannot inherit the problem.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: route hub_api_secret through the FORCE-aware declaration
`HUB_API_SECRET` lives in an `ArcSwap` rather than an `Arc<RwLock<_>>`, so it
could not use `option_setting` and was declared by hand with a bare `setting`
plus `parse_option_setting_value` — which is exactly the path that skips the
`FORCE_` handling, so a failed read still dropped `FORCE_HUB_API_SECRET`.
The rule now lives in `option_setting_with`, which takes the store closure and
leaves `option_setting` a wrapper over it, so a setting held in something other
than an `RwLock` reaches it too rather than having to reimplement it.
The three remaining hand-written parses are `parse_setting_value`, which has no
`FORCE_` handling to miss: `load_setting_value` never had the check either.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: add trigger_history table with source tracking
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: gate trigger history reads on scopes and harden its writers
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: filter trigger history scopes in SQL and match the cleared-handler diff
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: record a trigger restore from the trashbin in its history
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: record bulk http trigger creates and document the recording boundary
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: lock the trigger row when capturing its history preimage
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: only record an auto-disable that actually flipped the schedule
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: state the auto-disable invariant once instead of at four call sites
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: render trigger history changes as a structured field diff
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: make a server-initiated disable atomic with its history row
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: note that the auto-disable savepoint takes no pool connection
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: note the flow fallback is the last chance to disable
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: never leave a trigger enabled because its history row failed
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: retry the disable history row instead of dropping it on first failure
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: use the design-system Button for the change-value expander
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: hold the trigger row lock across its disable history row
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the history-loss alert out of the listener cancellation race
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: read the history workspace through the trigger-workspace seam
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: key server_heartbeat row on hostname so restarts reuse one row
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: trim announce_server_started doc to the durable constraints
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: only traffic-serving processes take part in coordinated restarts
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: name every non traffic-serving mode in the restart-gate comments
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: narrow the restart-gate comments to claims that hold
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* perf: cache resolved python interpreter path across worker restarts
Every worker process start spawned two `uv python find` subprocesses to
re-discover an interpreter path that had not changed, and every python job
spawned one more. The resolved paths are now memoized in a small JSON file next
to PY_INSTALL_DIR, which outlives the process, so a restarted worker (notably
under EXIT_AFTER_N_JOBS) reuses what the previous one resolved.
An entry is only served when the uv binary is the same one that produced it and
the interpreter is still on disk; otherwise it falls through to a real
`uv python find`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: address review findings on the python path cache
- resolve uv through PATH on windows, where `metadata("uv")` looked in the
worker's current directory and silently disabled the cache
- stat uv with tokio::fs instead of blocking the runtime, and compute the
identity once per resolution instead of once per read and twice per write
- store one file per version instead of a shared map, so workers resolving
different versions concurrently cannot drop each other's entry
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the windows uv PATH probe off the async runtime
The lazy static resolving uv through PATH stats candidate entries synchronously,
so its first use is moved onto a blocking thread.
Also records why an entry keyed on a minor-only version does not pin a patch:
uv answers such a request with its minor-version link and re-points it on a patch
install, so the memoized path follows the upgrade.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* perf: back off the interactive worker shell under EXIT_AFTER_N_JOBS
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: address review nits on the shell backoff docs and periodic warning
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* perf: only give the worker shell its sub-second cadence during a live session
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* perf: resolve the worker external IP in the background
`run_workers` awaited `external_ip::get_ip()` — an HTTPS GET to
hub.windmill.dev — before spawning any worker, so every worker process paid
that round trip before its first job pull. Measured on a CE debug build it was
120-450 ms of a ~200-500 ms startup, and behind a firewall the call does not
fail fast: it burns its whole 5 s connect timeout, on every process start. That
cost is per-job under EXIT_AFTER_N_JOBS.
The value is informational (it is only written to `worker_ping.ip`, which the
workers list displays so users can whitelist the address), so nothing needs to
wait on it. It now resolves into a process-wide cache off the startup path, and
`WORKER_EXTERNAL_IP` supplies it explicitly for deployments that know their
egress address or have no egress at all.
Until it resolves the ping carries no IP, which `insert_ping_query` now
COALESCEs so a reclaimed row keeps the address the previous process wrote
instead of being blanked. The main loop reports the IP as soon as it lands
rather than on the next periodic tick, so a short-lived process still records
it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep unknown worker IPs out of the whitelist alert
Review follow-ups:
- `WhitelistIp` filtered only the `'unretrievable IP'` sentinel, so the `'NO IP'`
one a pending or failed lookup now leaves in the row would be offered as an
address to whitelist. It filters both.
- Register `WORKER_EXTERNAL_IP` in `ENV_SETTINGS` so operators can confirm from
the instance settings view that it took effect.
- The worker tracked whether it had reported the IP by re-reading the cache
after each ping rather than remembering what the ping carried, so a lookup
landing mid-ping marked it reported without it reaching the row. The value is
read once and threaded through `insert_ping` / `update_worker_ping_full`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: report a sentinel IP once the lookup has definitively failed
Keeping the previous process's address on a reclaimed `worker_ping` row is right
while the lookup is still in flight, but not once it has failed: the row would
advertise an address nothing has confirmed, and the whitelist alert would offer
it. A failed lookup now reports `UNKNOWN_IP`, leaving NULL to mean "in flight".
`WORKER_EXTERNAL_IP` is rejected when longer than the `varchar(50)` column
rather than panicking the worker on its initial ping, which is a hard failure.
Adds the regression guard for the `ON CONFLICT` semantics: reverting to
`ip = EXCLUDED.ip` would compile and blank every reclaimed row.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the agent initial ping acceptable to older servers
An agent worker routinely runs against a server of a different version, and one
predating the background lookup rejects an initial ping carrying no IP — which
`run_worker` turns into a panic, so a newly upgraded agent would crash-loop
against it. The not-resolved-yet case goes over the wire as the sentinel
instead, and the server maps it back so a reclaimed row still keeps its address
while resolution is pending.
Also documents `ip` as the one conditional exception to `insert_ping_query`'s
"only `started_at` and `jobs_executed` survive a restart", and adds
`WORKER_EXTERNAL_IP` to the README env-var table.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: deliver the resolved IP to servers that only take it at registration
A server predating the background lookup applies `ip` from the initial ping
only, and ignores it on the periodic ones. An agent registering before its
lookup resolves would therefore keep the sentinel forever on such a server,
where it used to report its real address. It registers a second time once the
address is known, skipping that when the address is still unknown, when the
server is reached over SQL and needs no second registration, or once a job has
run, since registering clears the row's current job.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: re-register the resolved IP even after a job has run
Gating the second registration on "this process has not run a job yet" meant an
agent that pulled queued work before its lookup resolved never delivered the
address to a server that only takes one at registration. No job of the worker is
in flight where that runs, so the gate bought nothing beyond the last job's id,
which the next job refills.
Documents the two cases where WORKER_EXTERNAL_IP stops being an optimisation and
becomes the only way to report an address: an agent against such a server, and a
process shorter-lived than the lookup.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* revert: drop the WORKER_EXTERNAL_IP escape hatch
Supplying the address by hand skips the hub lookup, which is not something to
make easy. Resolving it in the background is what keeps it off the startup path;
opting out of it is a separate decision this does not need to take.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: distinguish an IP never established from one that could not be retrieved
`NO IP` was doing double duty: the column default for a row whose lookup has not
resolved, and the marker for one that failed. An operator reading the workers
list could not tell "not resolved yet" from "this instance cannot reach the
hub", and the latter is the actionable one. A failed lookup now reports
`unretrievable IP`, which is also what it reported before the lookup moved off
the startup path.
That leaves `NO IP` meaning only "no address established", which is what an
agent sends while its lookup is in flight and what the server maps back to
"unresolved" — so the wire sentinel no longer collides with the failure marker,
and an agent delivers the failure to a server that only reads an IP at
registration.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: stream audit logs in batches when a page is slow to load
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: bound streamed page size and clear stale rows on stop
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the runs batch cap and drop rows of a replaced query on failure
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: ignore stop once a load has settled and reset paging when one fails
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: update ee-repo-ref to aab7da6e1f8b1fadacc2208913a5d6596f06f922
This commit updates the EE repository reference after PR #727 was merged in windmill-ee-private.
Previous ee-repo-ref: 59ba8d7ce9ce1de0814b159b3813c2ac2a49239a
New ee-repo-ref: aab7da6e1f8b1fadacc2208913a5d6596f06f922
Automated by sync-ee-ref workflow.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
* docs(agents): scope agent guidance to where it loads
AGENTS.md loads in every session. Three of its sections only ever applied to
one directory, and docs/autonomous-mode.md was unreferenced by anything in the
repo, so none of its content was in effect.
- Move "Verifying Backend Changes" to backend/CLAUDE.md, "Verifying Frontend
Changes" and "Banned Patterns" to frontend/CLAUDE.md. They now load when
working under those directories, which is when they apply.
- Update the two cross-references that pointed at the moved sections (pr and
svelte-frontend skills).
- Delete docs/autonomous-mode.md. Its "don't stop early" half is already in
.webmux.yaml's oneshot system prompt, which actually loads; its trigger was
bypassPermissions, which does not imply an absent user; and it restated
AGENTS.md and the pr skill with copies that had drifted (hardcoded ports,
relative screenshot paths). Salvaged the UI traps it uniquely documented
into frontend/CLAUDE.md and dropped the three stale profile references.
AGENTS.md drops ~3.6k characters with no guidance lost.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(agents): guidance for building a feature — reuse, telemetry, live verification
Three recurring gaps, all cases where a pointer existed but nothing triggered
on it.
Component reuse. The svelte-frontend skill documented three components with
props, which reads as the whole catalog; the barrel exports 23 and common/ has
34 subdirectories against those 23. So "never use raw HTML elements" was an
instruction agents could not follow. Added a mandatory discovery step: read
the barrel, grep the tree, and treat the documented three as examples.
Brand guidelines. frontend/brand-guidelines.md is 34k characters referenced by
bare path, which nothing opens speculatively. Added a table mapping what you
are building to the section that governs it, entered with grep rather than a
full read.
Product telemetry. feature_usage has 14 registered actions across three
features, and an unregistered (feature, kind) pair is dropped by
valid_feature_usage_event with a bare continue — no error, still a 204 — so
frontend-only instrumentation silently records nothing. New
docs/feature-telemetry.md carries the criteria for when to instrument, the
four-step recipe including the allowlist and the InstanceSettings disclosure,
and the privacy rules. Raised in the plan for user-facing work, not as a
separate question, and not at all for bugfixes or refactors.
Also: validation now ends at exercising the change on the running instance,
with standing permission to spin up whatever that takes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(dev): correct the worktree dev-environment guidance
Several things agents were told to do did not match what the machine does.
- Env discovery pointed at .env / .env.local / backend/.env. In a webmux
worktree the real values are in $(git rev-parse --git-dir)/webmux/runtime.env
(BACKEND_PORT, FRONTEND_PORT, DATABASE_URL, CARGO_FEATURES, WM_DB_NAME),
sourced by every pane and undocumented. Reading it is also not blocked by the
Read(**/.env) deny rules, which the old instruction walked straight into.
- The database name rule said branch-with-underscores. worktree-common.sh uses
the worktree directory basename, and Postgres truncates at 63 characters, so
branch hugo/win-2340-… resolves to windmill_win_2340_…_and_eval with no hugo_
prefix and the tail chopped. A wrong DATABASE_URL guts the sqlx cache.
- The restart procedure said "tmux pane 1" and sent keys to an undefined
<pane1>. Pane 1 is the backend under the full profile and the frontend under
frontendOnly. Replaced with finding the pane by pane_current_command,
recovering the live feature set from the running process (CARGO_FEATURES in
runtime.env only records what the pane started with), and restarting in place.
- Added recovery for an orphaned backend holding the port: it reparents to
systemd when its shell dies, so it survives anything that looks like cleanup.
Three checks before killing a single pid, because pkill -f windmill takes out
every sibling worktree.
- Agents spawned their own servers because AGENTS.md opened by telling them to.
Now it checks for the existing panes first; the spawn commands are scoped to
a plain checkout.
- New EE worktrees branched from the EE repo's local main, which nothing
fast-forwards, so they started behind the commit pinned in
backend/ee-repo-ref.txt — the one CI builds against. They now base on the pin,
falling back to main only when it is unreadable.
- Enabled webmux autoPull so local main stays current; new worktrees are
branched from it. Documented what WM_CLONE_DB does, including that it
terminates every connection to the base windmill database.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(skills): vendor grilling/architecture skills; tighten PR ready and review rounds
Vendors five skills from https://github.com/mattpocock/skills (MIT, pinned at
84fdeffd12f2ee307994d1eb6feb48173b6e0502). They are one dependency closure:
grill-me is a stub that runs grilling, and improve-codebase-architecture draws
its vocabulary from codebase-design and its CONTEXT.md upkeep from
domain-modeling. .agents/skills/UPSTREAM.md records the license, the pin, and
the four local deltas so a refresh stays a diff:
- flattened the upstream engineering/ and productivity/ split
- rewrote bundled-file links to repo-root paths, since relative links break
when read through the .claude/skills symlink
- dropped the upstream agents/openai.yaml packaging metadata
- removed every ADR path. This repo has not adopted ADRs, and a skill that
offers to create them is how the practice arrives by side effect rather than
by decision.
PR workflow changes, all in the pr skill:
- A round that never starts is usually a conflict with main, not a CI outage.
Resolve by merging, not rebasing — a rebase rewrites the head SHA that round
verdicts and the clean-round marker are keyed to. If the merge advances
backend/ee-repo-ref.txt, the EE worktree has to follow or
cargo check --features private compiles a tree neither the author nor CI
intends.
- A clean round no longer means an automatic flip to ready. Wide blast radius
(*_ee.rs, migrations, OpenAPI or the generated client, auth paths, shared
worker infrastructure, a new public surface) asks first; self-contained
changes flip. Unattended, the judgement holds and the action degrades: flip
the small ones, leave the rest at a clean draft with the reason in the PR
body.
- Rounds that never converge are usually structural. After three without
convergence, stop, name the module the findings cluster around, and suggest
improve-codebase-architecture rather than burning more CI.
AGENTS.local.md (gitignored, with CLAUDE.local.md importing it) holds the
ready/ask calibration, recorded as dated observations rather than a rule.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(dev): state that each worktree gets its own fresh database
The per-worktree section warned which DATABASE_URL to use but never said where
the database comes from: the post-create hook creates and migrates a new one
per worktree, so it starts with none of the main instance's workspaces, scripts
or flows. WM_CLONE_DB was documented only as a comment in .webmux.yaml, which
reads as how things work rather than as a per-project opt-in.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(sqlx): script the cache backup/restore instead of documenting it
The update-sqlx skill spelled out a cp/comm/rm dance around `cargo sqlx
prepare`, which empties backend/.sqlx before regenerating — a failed run leaves
the cache gutted (observed: 2350 -> 142 entries), and a --all-targets run in a
CE checkout fails that way every time. Three problems with documenting it:
- The backup path was the literal /tmp/sqlx_backup, shared by every worktree.
Two concurrent runs overwrite each other's backup, which is the only thing
standing between a failed prepare and a gutted cache.
- The restore was a copy-pasted `rm -rf .sqlx && cp -r ... && cp ...` chain.
- Skipping the backup is what turns a routine failure into a lost cache, and a
convention is easier to skip than a command.
sqlx-cache.sh has backup / newq / restore, keeps state in a per-worktree
directory, and leaves the judgement call where it belongs: `newq` prints each
added entry's query field for review, and only `restore` writes them in.
Also adds the general rule that scratch files belong outside the checkout —
anything written into the tree has to be deleted again, and rm prompts each
time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(agents): state why a routine cleanup prompts, and where scratch goes
The guard hook already auto-allows a plain rm whose operands are under /tmp or
inside a git checkout in $HOME, so deleting a temp dir or a stale .sqlx entry
costs nothing. What prompts is the command shape: the hook's tokenizer defers on
&&, ;, redirects, quotes and $VAR, so a chained cleanup falls through to the
Bash(rm:*) ask rule.
That was recorded only inside a paragraph about screenshot file paths in
frontend/CLAUDE.md, where nobody looking for it would find it. Stated in Core
Principles instead, alongside the rule that scratch belongs outside the tree —
for the reason that actually applies, which is not committing junk rather than
avoiding prompts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(security): deny agent edits to the permission hooks and project settings
.claude/hooks/guard-rm-outside-tmp.sh and guard-main-branch.sh are the
enforcement points for everything the permission rules are meant to catch, and
nothing stopped an agent editing them. One sed -i disables the guard for every
later command, silently, and the deny list in .claude/settings.json has the same
exposure.
Defence in depth rather than a boundary: an agent with arbitrary bash can still
delete, and this may only close the Edit-tool path if Bash writes are not
covered by Edit deny rules. It costs nothing and removes the cheapest way to
turn the guards off. Changing them now means editing the files by hand, which is
the intent.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: address review round findings on head 3f47dc1
- backend/ and frontend/ guidance was Claude-only. Codex and Pi read AGENTS.md,
not CLAUDE.md, so moving "Verifying Backend/Frontend Changes" and the
$bindable ban out of the root AGENTS.md made them invisible to two of the
three CLIs this repo supports. Renamed both to AGENTS.md with a one-line
@AGENTS.md CLAUDE.md beside them, matching what the repo already does at the
root and in ai_evals/, and retargeted the four references.
- sqlx-cache.sh aborted with exit 2 and no output when .sqlx was empty:
list_entries ran `ls -1 ./*.json`, and an unmatched glob under
`set -euo pipefail` killed the script. An empty cache is precisely what a
failed prepare leaves behind, so it broke in the one case it exists for.
Replaced with a glob loop; reproduced the failure and verified the fix.
- The oneshot prompt ("never leave the PR sitting in draft") contradicted the
"Flip, or ask first" rule added in the same PR, which tells unattended runs to
leave wide-blast-radius changes as clean drafts. The prompt now defers to the
skill for the flip decision and keeps only "never stop at an unreviewed
draft".
- Bundled-resource references in the vendored skills were markdown links to
`.agents/skills/...`, which resolve relative to the file, not the repo root.
Replaced with inline paths stating they are repo-root relative.
- The PR-ready calibration file was write-only: the skill said to record
answers there but never to read it. It is now consulted before deciding.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Revert "chore(security): deny agent edits to the permission hooks and project settings"
This reverts commit 3f47dc1692.
* fix: address round 2 nits
- backend/AGENTS.md told agents to persist CARGO_FEATURES in runtime.env, but
webmux regenerates that file from metadata and .env.local every time the
worktree is opened, so the setting is lost on the next reopen. The persistent
source is .env.local, which scripts/post-create.sh already writes.
- UPSTREAM.md still described the vendoring delta as rewriting bundled-file
*links* to repo-root paths. 555f063 replaced them with plain paths in prose,
because a markdown target resolves relative to the file — a repo-root link is
just as broken as a sibling-relative one through the symlink. Replaying the
old wording on a refresh would reintroduce the bug UPSTREAM.md exists to
prevent.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(skills): correct the UPSTREAM.md link-rewrite delta
The delta note still described rewriting bundled-file *links* to repo-root
paths. 555f063 replaced them with plain paths in prose, because a markdown
target resolves relative to the file containing it — a repo-root link is as
broken as a sibling-relative one read through the symlink. Replaying the old
wording on a refresh would reintroduce exactly the bug UPSTREAM.md exists to
prevent.
The preceding commit's message claimed this fix; the edit had failed on a
stale anchor and only the backend/AGENTS.md half landed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(dev): describe what a fresh worktree database actually contains
Exercising a real worktree creation showed the previous wording ("none of your
workspaces, scripts or flows") reads as an empty database. It is a bootstrap
instance: the admins workspace, the admin@windmill.dev superadmin, the license
key copied from the base database, and the migration seeds — observed as
u/admin/hub_sync and the default app theme resource.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: expand AZURE_DEVOPS_TOKEN placeholder in backend git probes
* fix: require azure token placeholder to be http userinfo
* fix: scrub probe credentials from git stderr and harden token mint
* fix: confine azure token placeholder to azure devops hosts
* fix: require https and authorize azure reference at write time
* fix: require workspace admin to configure an azure token reference
* fix: name the azure reference in the admin-required error
* fix: git sync missed metadata-only deploys, deploy check missed job link
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: skip the deploy hook when the mute toggle matched no row
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: pin ee ref forward of main so the bump only adds this change
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: update ee-repo-ref to a65162b22b127b54c0686095ee1b16b04e3111f7
This commit updates the EE repository reference after PR #724 was merged in windmill-ee-private.
Previous ee-repo-ref: ac5f646c3ace7e5841200c6b83b34fb4371340d9
New ee-repo-ref: a65162b22b127b54c0686095ee1b16b04e3111f7
Automated by sync-ee-ref workflow.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
* feat: auto-build binaries to object storage on deployment
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: queue the auto-build from pre-locked deploys and off the lock slot
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: materialize companion modules before a deploy-time build
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep a build job from stamping lock_error_logs on a healthy script
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: de-flake test_flow_lock_all and surface the lock error it hides
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: trim drafting history from the flow-lock fixture comments
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: stop a binary build from restarting dedicated workers
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the build-job marker off the agent wire and out of user args
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: close SSRF bypasses in git URL validation (DNS + redirects)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: name the remedy when a git probe stops at a redirect
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: retry the .git form when a probe stops at a same-host redirect
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the .git retry on the validated host for pathless URLs
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: add EXIT_AFTER_N_JOBS worker mode for environment cleanup
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: address review findings on the EXIT_AFTER_N_JOBS worker mode
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: address round-2 review findings on EXIT_AFTER_N_JOBS
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: address round-3 review findings on EXIT_AFTER_N_JOBS
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: bound WORKER_SUFFIX length and document the same-worker drain
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: validate the assembled worker name length
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: bound go compilation memory with GOMEMLIMIT
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: bound the whole go build tree, not each toolchain process
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the go build memlimit and parallelism atomic
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: log the go limits actually installed and stop serializing small workers
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: make go build parallelism authoritative over persisted GOFLAGS
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: canonicalize the go build -p value and floor the module-step budget
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: parse GOMAXPROCS for -p the way the go runtime does
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: read GOMAXPROCS with go's own grammar and report limits neutrally
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: derive go build parallelism from the cgroup quota over its own period
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep go's minimum build parallelism under sub-CPU quotas
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the windows 1CU cap out of go's two-compiler floor
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: record that a worker runs one job at a time
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: scope the one-job-at-a-time rule away from native workers
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: home search matches each term instead of the whole query verbatim
* docs: state the search term cap and drop unreachable test cases
* fix: treat a term-less search as no filter and trim the comment
* fix: a term-less search matches nothing instead of the whole page
* feat: match the homepage fuzzy search exactly in the runnables endpoint
* docs: say apostrophes stay in terms; test summary-less and draft rows
* docs: separate an empty search from one holding no terms
* docs: state that terms split on ASCII alphanumerics only
* fix: escape and validate custom env var names in the nativets prologue
Custom workspace environment variable names were spliced verbatim into the
generated NativeTS/Bun JS prologue (both the `const {name}` binding and the
`process.env['{name}']` assignment), while only the value was escaped. A
non-identifier name could therefore alter the generated program.
- Add `escape_js_single_quoted` / `is_valid_js_identifier` helpers.
- worker.rs and bun_executor.rs: escape the name as a string literal, and only
emit the `const {name}` binding for valid identifiers.
- set_environment_variable: reject non-identifier names on write (deletion stays
unrestricted so existing rows remain removable).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: address review — reserved-word const gate, grandfathered-name editability
- Gate the `const {name}` prologue binding on `can_bind_as_prologue_const`, which
additionally excludes JS reserved words and the prologue's own bindings
(`process`, `BASE_URL`, `BASE_INTERNAL_URL`); such names would otherwise emit a
SyntaxError that breaks every NativeTS run. They are still exposed via
`process.env['{name}']`.
- set_environment_variable: only enforce the identifier check for names that don't
already exist, so editing the value of a pre-existing non-identifier name (the
edit UI resubmits the name) isn't rejected with no in-product fix.
- Document the name constraint on the endpoint in openapi.yaml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: exclude eval/arguments from const gate; skip existence query on valid names
- Strict-mode ES modules forbid `eval` and `arguments` as binding names, so add
them to the non-bindable set — otherwise an env var named `eval`/`arguments`
emits `const eval = ...`, a SyntaxError that breaks every NativeTS run.
- set_environment_variable: run the existence check only when the name isn't a
valid identifier, so the common (valid-name) path skips the extra query; trim
the rationale comment.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: allow `async` as a prologue const binding; note reserved-bindings coupling
`async` is a contextual keyword, not a reserved word — `const async = ...` is
valid, so it needn't be excluded from the const binding. Also cross-reference the
prologue head from PROLOGUE_RESERVED_BINDINGS so the two stay in sync.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix: authorize GET /concurrency_groups/{job_id}/key per job
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: answer 404 for an inaccessible and an unknown job alike
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: clear orphaned usr_to_group rows on service account creation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: pin service account creation over orphaned usr_to_group rows
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: update ee-repo-ref to 60c20e686cead73ff075512b15c6e2d6232beca6
This commit updates the EE repository reference after PR #723 was merged in windmill-ee-private.
Previous ee-repo-ref: 5c2c553f960abcd7988fdac8830dd36c066160ad
New ee-repo-ref: 60c20e686cead73ff075512b15c6e2d6232beca6
Automated by sync-ee-ref workflow.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
create_schedule opened the RLS transaction (user_db.begin) first, then ran
reads that deliberately use the non-RLS `db` pool — fork-ness and
permissioned_as/email resolution — while holding it. Acquiring a second pooled
connection while the tx holds one self-deadlocks on a single-connection pool
(embedded Postgres, PgBouncer statement mode, any max_connections=1 setup): the
read blocks on the sqlx acquire timeout, then errors.
Move those reads (and the ScheduleType::from_str validation) above
user_db.begin(). They don't depend on the tx and bypass RLS by design, so the
result is semantically identical; the RLS transaction is simply opened later and
held for less time. Same class of fix as #9970 (migration bootstrap on the
migrator's held connection).
Note: sibling paths keep the same latent pattern on branches this change does
not touch (push_scheduled_job reads the pool under the tx for flow schedules;
edit_schedule/set_enabled for cross-user permissioned_as) — a possible follow-up.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix: bound postgres result collection so it cannot OOM the worker
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: render the sql result limit exactly so the error can be set verbatim
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: point the fraction rationale at the renderer that still emits them
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* perf: stop re-parsing every collected row to rebuild it as a RawValue
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* style: drop a dangling doc line and an unrelated rustfmt reflow
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: bound duckdb result collection so an oversized result cannot OOM the worker
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the duckdb cap a worker-survival limit rather than a cloud product one
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: refuse an oversized blob before it expands to one json value per byte
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: share one expansion budget across a row's values, nested ones included
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: bound the row's own serialization so escaping cannot outgrow the budget
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: charge a json column before parsing it into a value tree
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: trim the json budget rationale and name what the budget does not cover
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: gate the sql result size limit on the duckdb feature
Its only consumer is the duckdb executor, so the minimal build compiled it
as dead code and failed under -D warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: bound how much disk a single duckdb job can spill
* fix: name the env var and correct duckdb's unreachable spill-cap advice
* fix: do not blame an unset env var for duckdb's default spill cap
* style: keep the duckdb spill-cap invariant comments within four lines
* docs: size the duckdb spill cap against the disk cloud pods actually use
* fix: tell MCP clients which tool parameters may be omitted
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor: make the mcp property-key rename testable and shorten the hint
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the mcp omission hint from calling flow inputs optional
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: skip the mcp omission hint on a parameterless tool
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: design for where the npm proxy keeps cached registry content
* feat(npm-proxy): keep package files on disk and in the object store
* fix(npm-proxy): degrade when the cache is unwritable, stream and bound it
* fix(npm-proxy): keep the happy path off the heap and isolate pull scratch
* fix(npm-proxy): bound the upload, verify pulled trees, keep oversized manifests
* fix(npm-proxy): protect live scratch, bound uploads by parts, refuse traversals
* fix: let the blocking unpack own the scratch it writes into
* fix: replace a cache directory that is not a package instead of deferring to it
* fix: evict by moving a package off the live path, not by deleting it in place
* fix: leave a package the sweep cannot move rather than deleting it in place
* fix: take one registry snapshot through a cache miss
* fix: stamp a pulled package as used so the sweep does not evict it first
* fix: explain duckdb failures caused by job isolation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: apply the isolation policy to the schema-sync pre-pass
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: bump ee ref for the out-of-memory hint wording
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: bump the bundled DuckDB engine to 1.5.5
The 1.5.5 duckdb crate no longer hands back a 96-bit `rust_decimal`, so a
DECIMAL wider than that renders instead of panicking inside an `extern "C"`
frame — which, being unable to unwind, aborted the whole worker process and
left the job running as a zombie. `SELECT
'1234567890123456789012345678.9012345678'::DECIMAL(38, 10)` was enough.
Adapting to the crate's API: `Value` is now `#[non_exhaustive]` and gained
`UHugeInt` and `Geometry`, and `rust_decimal` became an optional feature that
the `decimal`/`numeric` argument path still needs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: address review findings on the duckdb bump
Run the FFI crate's own tests in CI: it is excluded from the workspace, so the
`cargo test --all` in backend-test never reached them and the new guard against
the worker-aborting DECIMAL would not have run. build_dev.sh now honors a
caller-pinned CARGO_TARGET_DIR so the test build reuses that compile instead of
building the bundled engine a second time.
Also pin UHUGEINT rendering, and correct the rust_decimal rationale —
`Decimal::new` is public without the feature, so the reason is that the feature
reproduces the exact binding the crate used to derive, not that nothing else can.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: address review nits on the duckdb bump
Name the unsupported DuckDB type rather than dumping the value, which may be
arbitrarily large or hold data that does not belong in an error message, and
say which column it came from.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: pin the ee ref to the narrowed duckdb extension allowlist
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: keep duckdb spilling behind the local-filesystem fence
* chore: repin the duckdb fork after adding the reset-test exclusion
* docs: stop claiming the duckdb patch has been filed upstream
* docs: point the backend duckdb bullet at the fork's rationale
* fix: place lock_temp_directory so no existing struct member moves
* fix: skip the extension-load guard when the repo is unreachable
* refactor: trim the fork comments and fail the extension guard in CI
* chore: repin the duckdb fork onto upstream duckdb-rs main
* fix: keep the engine patch applying on a CRLF checkout
* docs: link the upstream issue tracking the underlying problem
* chore: repin the duckdb fork onto the patch as filed upstream
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: update ee-repo-ref to 88568d11162ffa11723e7955e613224bab4f0568
This commit updates the EE repository reference after PR #720 was merged in windmill-ee-private.
Previous ee-repo-ref: 22f075c1164d9dd5a3ba92d682905aabd071d273
New ee-repo-ref: 88568d11162ffa11723e7955e613224bab4f0568
Automated by sync-ee-ref workflow.
* chore: repin the duckdb fork onto the cmake/fmt build fix
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: pin the immutability half of lock_temp_directory
The spill test proves the exemption works; nothing proved the lock that makes
it sound. A rebase could drop the refusals and leave every other tripwire green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>