Experimental and env-gated (defaults off — stock behavior unchanged).
Read path is made fully non-blocking on both sides so neither the worker
nor the server waits on the stage behind it:
Postgres -> [server-side buffer] -> HTTP -> [agent-side buffer] -> worker
- batch pull/complete: N jobs per pull, N completions per flush, so the
per-job HTTP round-trips and per-job dequeue queries are amortized
- server-side prefetch: a single background puller per job-tag-set runs
the SKIP LOCKED dequeue ahead of demand and coalesces concurrent agents,
moving the PG query off the request critical path
- agent-side prefetch: a background refiller keeps the local job buffer
deep (high low-watermark + several refills in flight) so the worker
loop pops without a network call
- empty-pull backoff fix: don't apply the idle long-sleep to a transient
empty pull from a prefetching consumer
- AGENT_PROF: per-phase profiler + auto-rendered per-job time-budget SVG
Single agent goes from ~2.5k to ~125k+ j/s. Multi-agent collapse is
reproduced but its root cause is still open (not established to be PG).
See benchmarks/agent_worker_scaling.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Earlier README told users to run a second `kubectl port-forward` and
hard-coded port 33405. Both wrong: wm_sim up calls portForwardApi() in
helm_deploy.ts:213, picks a free local port, prints `[helm] API reachable
at http://127.0.0.1:<port>`, and keeps the forward alive as a child of
the wm_sim process.
Updated workflow: read the URL out of wm_sim's output, pass it to main.ts
as --host. Kept a note about restarting the forward manually when the
kubelet drops it under heavy bench load (real failure mode we hit).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
sim/README.md was last touched May 27 and described the pre-minikube
docker/podman-based "three node modes" design with k8s mode listed as
"planned for M5". The reality is now k8s-only via minikube + helm, and
the workflow / config-files / outputs all changed.
New sim/README.md covers:
- The actual `wm_sim up` CLI form against the external
windmill-helm-charts chart and smoke.yaml / local.yaml overlays.
- Two-shell workflow: wm_sim up in one, port-forward + main.ts in the
other, report at reports/<timestamp>/dashboard.svg.
- Measurement subsystem table: per-pod CPU/mem sampler DS, JSONL pollers
(pod_timeline, oom, pg_latency, pg_conn, node_load), throughput
capture, failed-jobs fetch, pgBadger.
- The three reliability fixes we leaned on this session and why they
matter: sampler host-log file (kubelet log rotation), readiness
rollout-complete check (cgroup_mutex starvation), procs_running for
oversaturation (not load1 / D-state overcount).
- Dashboard panel index + the shared bench_start_ms x-axis origin.
- Workload catalog (io_4phase + flood variants + ops_day + cpu).
- Agent worker setup with the Secret-based JWT.
- Test entry points (util_metrics_test, util_panel_snapshot_test).
- Operational notes (port-forward fragility, minikube stop safety).
- File-by-file code layout under sim/.
benchmarks/README.md gains a "Cluster benchmarks — sim mode" section
linking to sim/README.md plus a quick-start snippet so the path from
the legacy benchmark suite docs into the cluster workflow is obvious.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Stand up a minikube-backed simulation subsystem for benching Windmill under
realistic multi-node load, with a per-bench measurement pipeline and a
dashboard renderer that consolidates throughput, queue depth, per-node CPU,
PG latency/conns, OOM events, and per-node CPU-util-vs-oversaturation into
one SVG report.
Sim infrastructure (sim/):
- k8s_provisioner: minikube up + heterogeneous node sizing from topology JSON
- helm_deploy: helm install Windmill with smoke.yaml + local.yaml overlays
- image_cache: pre-load required images so bench bringup is offline-safe
- toxiproxy_k8s: per-node toxiproxy DaemonSet for cross-node latency injection
- cpu_sampler_k8s: privileged DS reading per-cgroup cpu.stat at 10Hz, dual-
writes to stdout AND a host-mounted log file (/var/log/wm-sim-cpu-sampler/
sampler.tsv) so heavy benches no longer lose early samples to kubelet log
rotation
- pg_logging: ALTER SYSTEM + SIGHUP to enable verbose PG logging without restart
- pgbadger: post-bench PG log analysis HTML report
- readiness: pre-bench cluster health check (samplers stable ≥30s, workers
ready, PG responsive, queue empty, **deploy.status rollout-complete**) —
the rollout-complete check catches mid-rolling-update fires that previously
starved m04's sampler under cgroup_mutex contention
Per-bench JSONL pollers, started/finalized alongside the bench loop:
- pod_timeline: 1Hz workers-per-node Ready counts (used for the workers panel)
- oom_poller: live OOM event capture (kernel + kubelet evictions + cgroup)
- pg_latency_poller: 4Hz psql \\timing on SELECT 1 vs kubectl-exec roundtrip
- pg_conn_poller: 1Hz pg_stat_activity by state (active/idle/idle_in_xact)
- node_load_poller: 2Hz /proc/loadavg + /proc/stat procs_running per node
Dashboard renderer (sim/render_report.ts + graph.ts):
- Util group: one panel per node with translucent orange oversaturation area
BEHIND solid blue CPU-util area, 100% reference line, phase-boundary verticals.
cols:2 grid wraps after 2 panels per row.
- PG node tinted with [PG] flag in legend across the dashboard.
- Phase-boundary verticals + push-window shaded zones layered consistently.
- All x-axes switched from wall-clock HH:MM to relative seconds-from-bench-
start. Shared origin sourced from meta.json's bench_start_ms so 0s on every
panel = the same wall-clock moment (previously each chart picked its own
earliest sample as origin, causing drift between panels).
Oversaturation metric, with explicit fallback:
- Primary: (procs_running - ncpu) / ncpu × 100 — true CPU run-queue pressure.
- Fallback to load1 when procs_running is missing (older reports).
- load1 overcounted previously because it includes uninterruptible D-state
procs (PG backends in disk I/O, cgroup_mutex waits), inflating "saturation"
by 5-10x under load.
- Pure helper extracted to sim/util_metrics.ts; 8 unit tests cover the
procs_running > load1 preference, the clamp-at-zero, invalid-ncpu cases.
Sampler reliability:
- HostPath log file in addition to stdout so the bench's scp-based collector
bypasses kubelet log rotation entirely.
- main.ts truncates the host log file on every node before pushers start
(parallel ssh, best-effort) so it doesn't grow unbounded across runs.
- Collector falls back to kubectl-logs when scp fails for any node.
Workloads (workloads/):
- io_4phase: four-phase IO step (idle → 2.5s → 500ms → 150ms jobs)
- io_150ms_flood / io_300ms_flood / io_1s_flood / io_2s_flood: single-phase
flood configs to isolate the worker-host CFS context-switch storm vs PG
contention regime
- burst, ops_day, cpu_*, etc. for other scenarios
Tests:
- sim/util_metrics_test.ts — 8 cases for computeOversatPct
- sim/util_panel_snapshot_test.ts — 5 assertions guarding util-panel SVG
invariants (orange behind blue, 100% ref line, relative-time ticks NOT
wall-clock, phase-boundary verticals, shared-origin override)
Helm values:
- sim/values/smoke.yaml — bench-tuned: workers w/ no CPU limit & low mem
request, PG w/ 3-core request + wm-critical priorityClass + oomImmune +
maxConnections, app w/ wm-critical + oomImmune + no resource limits.
- sim/values/local.example.yaml — template for the gitignored local.yaml
that carries the EE license key.
- Depends on the wm-critical PriorityClass + oomImmune + maxConnections
knobs landing in windmill-helm-charts (separate PR).
graph.ts additions:
- areaFills param: ordered list of per-kind translucent area fills drawn
before lines, used by the util panel for orange-behind-blue layering
- lineColorOverrides: pin per-kind line colors so oversaturation reliably
renders orange regardless of d3 ordinal-color insertion order
- highlightKindToken: substring-match flag for the PG-node tint in Node CPU
- xRelativeOriginMs: shared bench-start origin for the relative-time x-axis
- DataPointMulti is now exported for downstream tests
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Operators are read-only and cannot create native triggers, yet the
google/github/nextcloud integration picker routes had no authorization
gate, letting any workspace member drive the admin-configured
integration's upstream API and enumerate its data (Drive files, repo
names, calendars, events).
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: authenticate slack callback payload with per-workspace hmac
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: regression tests for unauthenticated slack callback decryption
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: verify slack submission signature before resume + close workspace oracle
Addresses review: verify private_metadata HMAC before handle_resume_action so a
tampered/unsigned submission is rejected up front, and map get_workspace_key
failure to the generic 401 so the status code is not a workspace-existence oracle.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: domain-separate slack payload hmac from resume-secret hmac
Both MAC families key Hmac<Sha256> on the same per-workspace key; resume secrets
are distributed to approvers in resume URLs, so add a fixed domain tag
(slack_payload_v1) to the slack payload MAC to make the two non-interchangeable
by construction rather than by byte-layout coincidence.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A flow inline step whose id is a Python keyword (e.g. `in`) crashed with a
`SyntaxError`: the wrapper emits `from {pkg} import {step_id} as inner_script`,
and `from x import in as y` is invalid Python.
The codegen already prefixes `_` to path segments that start with a digit
(`1234` → `_1234`); this extends that guard to Python hard keywords (`in` →
`_in`) in `compute_python_module_dir` and on the leaf in `compute_py_codegen`
and `prepare_wrapper`. The relative-imports write path inherits it for free.
Fixes#8893
* fix: prevent token label collision bypassing job read access control
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: regression tests for token label collision job read access
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: bind job-read override fast-path to permissioned_as_email
Replaces the reserved-label / label-* exclusion approach: webhook-/http-/email-
labels are created through the public token API by the trigger panels, so they
cannot be reserved, and blocking label-* regressed legitimate re-reads. Instead
the username_override fast-path now requires the job's permissioned_as_email
(non-forgeable, never derived from the label) to equal the caller's email.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(otel): propagate inbound W3C traceparent to job spans
Capture the inbound traceparent header at the run endpoints
(WebhookArgs::to_args_from_format) into a reserved _wm_traceparent arg key
(gated on OTEL_TRACING_ENABLED), riding the args jsonb like
_ENTRYPOINT_OVERRIDE. At pickup, create_span_with_name attaches a span link
from the job's worker span to the originating distributed trace, so a job
triggered by an instrumented service is connected to the caller's trace
while keeping its UUID-derived trace id (trace-by-job-id unaffected).
The link/parse logic lives in the EE otel modules; this OSS side only
captures the header and calls the (no-op outside EE) hook. Companion EE PR
required.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: bump ee-repo-ref to inbound-trace-propagation EE branch
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(agents): don't attribute work to specific customers in repo content
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(otel): relocate job + script spans into the inbound trace
Builds on the captured _wm_traceparent: the worker job span is re-parented on
the inbound caller context, the script subprocess's TRACEPARENT env is the
inbound context (so its spans join the caller's trace), and the context is
propagated to flow steps so the whole flow relocates. Carried to the worker via
a new LogContext.inbound_traceparent field. Non-inbound jobs are unchanged.
Adds a relocation integration test. Companion EE PR required.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: bump ee-repo-ref to inbound-trace-propagation relocate commit
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(otel): harden inbound traceparent capture
Address review feedback:
- strip any caller-supplied _wm_traceparent from args/extra before stashing the
header-captured value, so the reserved key is Windmill-controlled only
- valid_w3c_traceparent: reject version ff and require lowercase hex, so we don't
forward an inbound header that downstream OTel parsers would reject
- clarify that the capture helper does not validate the W3C format (done at use)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: update ee-repo-ref to 2c7964460327fab5e3a27c0f74b8d6f26ab7f79a
This commit updates the EE repository reference after PR #604 was merged in windmill-ee-private.
Previous ee-repo-ref: 8fc04fb105dc49769205f7174d551a0d134d1bec
New ee-repo-ref: 2c7964460327fab5e3a27c0f74b8d6f26ab7f79a
Automated by sync-ee-ref workflow.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
* fix: distinguish canceled jobs in runs
* fix: order status=failure|canceled by completed_at to use partial index
The new `status` query param replaced the legacy `success=false` filter on
the Runs page, but the ORDER BY switch in list_completed_jobs_query only
flipped to v2_job_completed.completed_at for success==Some(false). With
status=failure|canceled (and success=None), the query fell back to ordering
by v2_job.created_at, which the partial index
ix_v2_job_completed_failure_workspace (workspace_id, completed_at DESC WHERE
status IN ('failure','canceled')) cannot serve.
EXPLAIN ANALYZE on 500k rows (1% failure/canceled): ordering by completed_at
uses the partial index (~150 buffers, 0.3ms); ordering by created_at scans
the v2_job created_at index and probes/discards 99% of rows via the join
(~49k buffers, 31ms). Switch the ordering to completed_at for
failure/canceled so the partial index serves both filtering and ordering.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: trim order-by regression test to the failure/canceled case
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: only treat canceled as a terminal status icon for completed jobs
Guard the canceled branch in JobStatusIcon and getJobStatusKind with
`'success' in job` so a job that is still running while being canceled keeps
its running icon/favicon until it completes, instead of immediately showing
the gray Canceled state. Also clarify the openapi `status` param is an exact
match (status=success excludes skipped, unlike success=true).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(sandbox): pull/extract images with crane instead of podman (+ add to image)
The sandboxed container runtime (`# sandbox <image>`) only ever pulls + flattens an
image (nsjail does the run), so a full container engine is overkill — and podman was
never actually in any Dockerfile, so the merged feature couldn't run in the shipped
image. Switch to crane (google/go-containerregistry): a single ~25MB static binary,
no daemon/store/root/privileged.
- docker_v2.rs: crane export -> flattened rootfs tar, crane config -> OCI config,
crane digest -> content-addressed rootfs+config cache (cross-job dedup + automatic
freshness), crane manifest -> pre-download size guard. DOCKER_CONFIG authfile dir.
Cache eviction prunes the rootfs-tar cache by mtime (LRU). Pull policy honored via a
ref->digest cache (missing/never reuse without a registry hit).
- Dockerfile + docker/DockerfileSlim{,Ee}: install the crane binary (Full/FullEe and
the EE image inherit it via FROM the base image).
- docs + UI text + instance-setting descriptions updated (download size is compressed;
cache is the rootfs-tar cache).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(sandbox): address CI review — digest-pinned fetch, size cap on every job, eviction race
Codex P1s:
- Fetch by the resolved digest (name@digest), not the mutable tag, so content can't
diverge from the digest the cache is keyed under if a tag moves mid-fetch.
- Enforce the size cap on EVERY job via a cached {digest}.size sidecar (no registry call
on cache reuse), so lowering the limit rejects already-cached oversized images.
- Eviction race: hardlink the cache tar into the job dir before tar -xf (pins the inode
against concurrent eviction) and re-fetch if it was evicted first.
Claude P2s: atomic config sidecar (tmp+rename) + tolerate torn parse; soften the LRU
comment (mtime = creation order); sweep orphaned *.tmp.* and .size on eviction.
+digest_key/ref_key unit tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(sandbox): P1 cross-fs cache staging (EXDEV), Dockerfile arch fail-fast
CI re-review (Claude + Codex P1): the eviction-race hardlink crosses filesystems in the
shipped deployments — the cache is its own volume (/tmp/windmill/cache) while the job dir
is on the container fs — so hard_link returns EXDEV (not NotFound) and every sandbox job
fails. Fall back to tokio::fs::copy on a non-NotFound link error; copy reads through the
source inode so it still survives a concurrent eviction.
Also: Dockerfiles fail fast with a clear error on an unsupported arch instead of building
a 404 crane URL; ref->digest file written via tmp+rename (no torn read under missing/never).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(sandbox): say 'oldest by creation time' not 'LRU' for cache eviction
Codex P2: the code evicts by tar creation time (cache hits don't touch mtime), so the
user-facing docs + instance-setting text shouldn't claim true LRU.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: make C# dotnet target framework configurable via DOTNET_TARGET_FRAMEWORK
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: include dotnet target framework in C# binary cache key
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: add sandboxed docker v2 runtime via '# docker <image>'
Run a container image as a subprogram of the job's own nsjail sandbox:
extract the image rootfs with podman (rootless) and run it chrooted inside the
job's nsjail, so the container inherits the job's confinement and is safe under
nsjail / for untrusted code. Selected by '# docker <image>'; a bare '# docker'
keeps the v1 (dind) path untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: default to daemonless docker (drop dind from compose, allow docker on cloud)
docker-compose no longer ships the dind sidecar (v2 is daemonless: podman + nsjail
in the worker); removed the dind service, DOCKER_HOST env, depends_on and volume.
Removed the language-picker guard that blocked Docker scripts on the multi-tenant
platform, now that v2 makes docker safe to run sandboxed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: select sandboxed container via # sandbox <image>; add pull policy + size guards
- Surface moved from '# docker <image>' to '# sandbox <image>' (groups under the
sandbox annotation; '# docker' stays v1-only, '# sandbox' stays nsjail-bash).
- SANDBOX_IMAGE_PULL_POLICY (default 'newer') so moving tags don't go stale.
- SANDBOX_IMAGE_MAX_SIZE_MB rejects oversized images before extraction.
- SANDBOX_IMAGE_CACHE_MAX_MB best-effort LRU eviction of podman's image store.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(sandbox): support # volume, honor nsjail tmp instance settings, v2 docker template
- Thread shared_mount into the sandbox container nsjail config so '# volume' mounts
(and the same-worker /tmp/shared folder) apply inside the container.
- Use resolve_nsjail_tmp_mount_block for the container's /tmp so it honors the same
nsjail_tmp_backing / nsjail_tmpfs_size_mb instance settings as other nsjail jobs.
- docker-compose comment + the editor's Docker template now use '# sandbox <image>'.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(sandbox): make image size/cache/pull-policy UI instance settings
Convert SANDBOX_IMAGE_* from worker env vars to DB-backed instance settings
(sandbox_image_max_size_mb, sandbox_image_cache_max_mb, sandbox_image_pull_policy),
hot-reloaded via the same mechanism as nsjail_tmpfs_size_mb and configurable in
#superadmin-settings. No worker restart needed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(sandbox): windmill-managed registry — default registry + private auth
Two new instance settings:
- sandbox_image_default_registry: prepended to unqualified image refs (alpine ->
<registry>/alpine); fully-qualified refs untouched.
- sandbox_registry_auth: docker/podman auth.json blob written to a per-job authfile
(0600, removed with the job) and passed to podman --authfile for private registries.
Both hot-reloaded and configurable in #superadmin-settings.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(sandbox): protobuf-safe proto_str escaper, atomic 0600 authfile, registry tests
Addresses local-review P2s: proto_str now emits valid protobuf octal escapes for
control/non-ASCII bytes (not Rust \u{..} that nsjail would reject); the registry
authfile is created 0600 atomically (no world-readable window); add a
registry_qualified table test + a non-ASCII proto_str case.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(sandbox): P0 — deliver image env via nsjail envar:, never the launcher process env
CI review (P0): the image's OCI Env (attacker-controlled keys+values) was applied to
the nsjail launcher process via .envs(), so a hostile image could set LD_PRELOAD/
LD_LIBRARY_PATH/LD_AUDIT on nsjail itself and execute code as the worker outside the
jail. Now the image env is rendered as proto-escaped 'envar:' directives (child-only)
and nsjail's process env carries only windmill-trusted keys (reserved vars + proxy).
Also: warn instead of silently bypassing the size guard on inspect failure; reset the
eviction guard via a Drop guard (no stuck flag on panic/early-return). +render_envars test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(sandbox): P0 symlink-write escape via rootfs script; P1 redact registry-auth logging
CI review:
- P0 (Codex): the body was written into the image-controlled rootfs as
.windmill_docker_main.sh via write_file (follows symlinks) — a hostile image could
plant that path as a symlink to a host file and capture the worker's write before
nsjail starts. Now the body is passed straight to 'sh -c <body> sh <args>'; no file
is written into the rootfs at all.
- P1 (Codex): sandbox_registry_auth flowed through the generic setting loader which
logs the value (raw auth.json credentials). Replaced with a secret-aware reload that
loads directly and logs only a redacted 'configured=' message.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(sandbox): redact sandbox_registry_auth in instance-settings write log too
The settings API also logs 'Set global setting <key> to <value>' via format_setting_value;
add sandbox_registry_auth to SENSITIVE_SETTINGS so the credential is redacted there as
well as on reload.
* fix(sandbox): don't silently disable cache eviction on podman images parse error
Re-review (cubic/Claude P2): serde_json::from_slice(...).unwrap_or_default() meant any
parse hiccup (e.g. podman omitting Size/Created via omitempty for a zero value, or
schema drift) silently degraded to an empty Vec and disabled eviction with no log.
Now Size/Created are #[serde(default)] (a missing omitempty key -> 0, not a whole-array
parse failure) and a real parse error warns + breaks instead of being swallowed.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ee repo ref
* fix(ee-ref): pin to EE commit that includes read_only create_session_token fix
The previous pin (f7a83d9) carried only the connect_config_template change and
dropped Ruben's read_only=false fix (EE 3742e06). CE #9371 made
create_session_token require 6 args, so the EE overlay fails check_ee_full with
an arity error without it. Bump the pin to 9be38de, which includes both fixes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: update ee-repo-ref to fb106b89cdf4088b004dac6062adb029f3923887
This commit updates the EE repository reference after PR #603 was merged in windmill-ee-private.
Previous ee-repo-ref: 9be38def879f702cd0b134d9e71bbb17fbb9cfa4
New ee-repo-ref: fb106b89cdf4088b004dac6062adb029f3923887
Automated by sync-ee-ref workflow.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
Adds DatabricksIcon.svelte (brand mark, #FF3621) and registers it under
`databricks` in the shared APP_TO_ICON_COMPONENT map, so both the app and
hub frontends pick it up for the new Databricks hub integration.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
Adds AdobeAcrobatSignIcon.svelte and registers `adobe_acrobat_sign` in
APP_TO_ICON_COMPONENT, for the Adobe Acrobat Sign hub integration
(windmill-labs/windmill-integrations#143).
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* oauth: add ServiceNow provider; make per-instance OAuth registry-driven
ServiceNow's OAuth endpoints are per-instance
(https://<instance>.service-now.com/oauth_auth.do + /oauth_token.do), like
Snowflake's. Rather than add another bespoke special-case, generalize:
a registry entry may carry a `connect_config_template` (label/placeholder/
help_url + {instance}-templated auth_url/token_url + req_body_auth +
optional extra_params_key/strip_suffix). The instance-settings UI renders
one generic instance-name input for any such provider and substitutes
{instance} to build the per-client connect_config — a new per-instance
provider needs only a JSON entry, no frontend code.
- oauth_connect.json: servicenow + snowflake_oauth now carry a
connect_config_template (snowflake keeps its account_identifier
extra_params key for backward compatibility).
- windmill-oauth: add the ConnectConfigTemplate struct (frontend-only
metadata; the backend's existing connect_config override resolves the
concrete URLs generically — no other backend change).
- AuthSettings/InstanceSettings: replace the Snowflake + ServiceNow
special-cases with one registry-driven path (instanceInputs map,
setupTemplatedOauthUrls, loadInstanceInputs); per-instance providers are
derived from the registry for the builtins list + dropdown.
Pairs with windmill-integrations#139 (ServiceNow hub integration).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci: point ee-repo-ref at servicenow-oauth EE branch (revert at merge)
Temporary CI pointer so check_ee_full / cargo_test build against the EE
slack-literal fix (windmill-ee-private#602). Revert to a pinned SHA once
that EE PR is merged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Wiz star logomark (brand blue #0254EC) for the shared icon map
(APP_TO_ICON_COMPONENT), for windmill-integrations#144.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(flows): early stop can include the stopping step's result in the raised error
When a step uses Early Stop with "Raise an error message if stopped", the
flow result was entirely replaced with a static error object
({"error": {"name": "EarlyStopError", "message": "..."}}), discarding the
stopping step's own output. This made it impossible to stop+fail a flow
while preserving the data the step produced (e.g. an API that returns
HTTP 200 with a userErrors payload).
Add an opt-in `error_include_result` flag on StopAfterIf. When enabled on
the raise-error path, the raised payload becomes
{"error": {...}, "result": <step result>} instead of dropping the result.
Default is false, so existing behavior is unchanged. The option is threaded
through the worker's stop-after-if handling (including stop_after_all_iters_if
for loops/branchall) and exposed in the flow editor's Early Stop panel.
Fixes WIN-2012
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(flows): cover early-stop error_include_result payload shaping
Add a regression test asserting that a step using Early Stop with a raised
error message and error_include_result=true fails the flow while preserving
the step output as {"error": {..}, "result": <step result>}, and that with
the flag off the result is the bare {"error": {..}} object.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(flows): nest early-stop step result inside the error object
Embed the stopping step's result under `error.result` rather than as a
top-level sibling of `error`. This keeps the flow result shape as
`{ "error": { .. } }` — identical to a normal error — so consumers that
key off the top-level shape (single `error` key) keep working, while the
data is still preserved for those that look inside the error object.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(flows): always include the stopping step's result in early-stop errors
Drop the opt-in `error_include_result` gate. Since the step result is nested
inside the error object (`error.result`), the top-level result shape stays
`{ "error": .. }` — identical to a normal error — so consumers that detect or
parse failures by the top-level shape are unaffected. Gating it added schema
surface, plumbing, and a UI toggle for no real compatibility benefit.
Now, whenever a step early-stops with a raised error message, the flow fails
and the raised error embeds the stopping step's own result under
`error.result` (aggregated iteration results for loops/branchall). This
reverts the `StopAfterIf.error_include_result` field, its threading, the
OpenAPI/generated-client surface, and the editor toggle; the "Raise an error
message" tooltip now notes that the step result is included.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(flows): gate early-stop result inclusion behind opt-in flag
Re-introduce the per-step `error_include_result` flag (default off) instead
of always embedding the step result. Although nesting the result under
`error.result` keeps the result *shape* backward-compatible, it does not
address data exposure: a failed flow's result is propagated to synchronous
webhook callers, the flow's failure module, and the workspace/global error
handler (commonly a Slack/email/outbound-webhook notifier). Always including
the step output would surface previously-redacted intermediate data to all of
those sinks for every existing error-stop flow.
Gating keeps the existing behavior (bare `{ "error": .. }`) as the default and
only embeds `error.result` when the flow author explicitly opts in, matching
the original issue's intent.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(flows): omit error_include_result when false; refresh generated prompts
- Add `skip_serializing_if = "is_false"` to `StopAfterIf.error_include_result`
so serialized flows are byte-identical when the flag is off. Fixes the
`flowmodule_serde` round-trip test (cargo_test) and avoids churn on existing
flows.
- Regenerate `system_prompts/auto-generated/` and `cli/src/guidance/skills.gen.ts`
for the new OpenFlow `error_include_result` property. Fixes check-freshness.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(flows): cover error_include_result for the loop "stop after all iters" path
Add a regression test for the stop_after_all_iters_if branch, where `nresult`
already holds the aggregated iteration results — confirming `error.result`
carries each iteration's output (distinct from the per-step fallback path).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cli): push whole raw app instead of treating frontend files as scripts
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(cli): shorten raw-app handleFile comment
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor: resolve workspace imports via /f/,/u/ not $f/,$u/ aliases
Keep the CLI managed tsconfig.wmill.json / `refresh tsconfig` / Deno
import-map QoL from #9378, but re-key it on the existing /f/,/u/ workspace
paths instead of the new $f/,$u/ specifiers. Verified /f/,/u/ resolves in
tsc, Bun, Deno, the in-app ATA editor, and the worker, so the $-prefixed
alias added no value. Drop the $f/,$u/ handling from the parser, dep-map,
deno_executor, bun loaders, ATA, relative_imports and monaco paths; revert
the windmill-parser-wasm-ts bump (1.714.0 -> 1.695.0). Also fold in the
cli/package-lock.json sync for the already-committed pg-gateway dependency.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: drop duplicate relative-path check and restore rustfmt formatting
Follow-up cleanups to the previous commit's full-file reverts, which
restored pre-#9378 state that main had since improved:
- relative_imports.ts: remove the redundant duplicate d.startsWith('/')
(pre-#9378 had it; #9378 had repurposed that line, so main has no dup).
- windmill-parser-ts/src/lib.rs: restore the multi-line new_source_file(...)
formatting required by backend/rustfmt.toml (the single-line revert would
fail `cargo fmt --check`). Now differs from main only by the $f//$u/ removal.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: let flow AI chat create and edit sticky notes
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: strengthen flow AI guidance to prefer groups for organizing flows
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: harden flow note validation (validate position/size, document color default and group acceptance)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: make AI-created free notes draggable by seeding default position and size
Free notes need explicit geometry to be draggable/resizable in the editor; UI-created notes always set position+size but agent-created notes omitted both, so they couldn't be moved until resized. Seed defaults in validateFlowNotes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`cli/src/utils/utils.ts` imported `VERSION` from `cli/src/main.ts`, while
`main.ts` transitively imports `utils.ts` (via `workspace.ts`). When a module
load order entered the graph through `workspace.ts -> utils.ts -> main.ts`,
`main.ts`'s top-level command tree ran while `workspace.ts` was still
mid-initialization, so the `workspace` binding was still in its temporal dead
zone at `.command("workspace", workspace)`:
ReferenceError: Cannot access 'workspace' before initialization
This surfaced as 56 failing CLI tests on Windows CI (the Windows runner's test
module-load order triggers the bad path; it reproduces on any platform via
`bun -e 'await import("./src/commands/workspace/workspace.ts")'`).
Move `VERSION` to `cli/src/core/constants.ts` (already the "minimal imports"
module), re-export it from `main.ts` for backwards compatibility, and have
`utils.ts` read it from `constants.ts` — eliminating the cycle. Release tooling
(`.github/change-versions*.sh`) is updated to rewrite the `VERSION` line in its
new location.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When a dependency job for an app is triggered by a relative/workspace
import (e.g. an imported script was updated), handle_app_dependency_job
re-appended the version captured at job-creation time to the versions
array. On a git-sync/CLI push that deploys both the imported script and
the importing app in the same batch, the script's dependency job
snapshots the app's old version; the app push then creates a newer
version (uploading its bundle against that new version); finally the
relock runs and re-appends the old version, making it latest again.
For raw apps this is fatal: bundle_secret is computed from the latest
version, so the served HTML requests /apps_u/get_data/v/<secret>.{js,css}
for a version that has no stored bundle -> 404 and a white screen.
Manually redeploying fixes it until the next merge re-triggers the revert.
Two changes:
- Re-query the current latest version to relock (mirrors the flow
dependency handler, #8673), so we don't lock a stale snapshot.
- Guard the re-publish append with `versions[array_upper(...)] = $1` so
it is a single atomic, never-demoting statement: it can only re-append
the version that is already latest, never revert to an older one. A
relock never creates a new app_version, so there is never a version to
legitimately promote here.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: support $f/ and $u/ import path aliases for scripts
$f/ and $u/ are local-friendly aliases for the absolute workspace
import paths /f/ and /u/. Unlike the /-prefixed form (which local tools
treat as a filesystem-root path), the $-prefixed form is a bare specifier
that can be remapped via tsconfig paths / Deno import maps, so the same
import resolves on the Windmill worker and in a local editor.
- worker: recognize $f//$u/ in the Deno import map and both Bun loaders
- dep-map/parser: normalize $f/->f/, $u/->u/ for lockgen + dep tracking
- cli: emit $f/$u path aliases in generated tsconfig.json / deno.json
- frontend: ATA + Monaco paths resolve $f//$u/ type hints in the editor
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cli): split generated tsconfig into managed + user file with refresh command
Mirror the AGENTS.cli.md/AGENTS.md prompts model for the IDE tsconfig so the
recommended settings can evolve without ever clobbering user customizations:
- tsconfig.wmill.json: wmill-managed, always refreshed, holds recommended
compilerOptions incl. the $f/$u path aliases (Deno: import_map.wmill.json)
- tsconfig.json: user-owned, created once, just extends the managed file;
warn (never auto-edit) when an existing one doesn't reference it
- add 'wmill refresh tsconfig'; init generates it unconditionally (no longer
gated behind resource-type namespace / a bound workspace)
- regenerate CLI guidance docs for the new subcommand
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cli): address PR review on $f/ tsconfig generation
- handle existing deno.jsonc so we don't shadow it with a new deno.json
(P1 identified by cubic)
- fix the bun-types hint that pointed users at the managed do-not-edit
tsconfig.wmill.json; tell them to install + re-run 'wmill refresh tsconfig'
- document the .ts-extension-only local-resolution limitation (cross-flavor
.bun.ts/.deno.ts/.fetch.ts scripts won't resolve in a local editor)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cli): warn when a project's tsconfig isn't wired to tsconfig.wmill.json
Mirror the prompts freshness check for the managed tsconfig so users with an
existing setup actually discover they're missing $f//$u/ resolution:
- embed a version hash in tsconfig.wmill.json (excludes the env-dependent
bun-types 'types' entry so it doesn't false-positive)
- add warnIfTsconfigStale to the main.ts freshness hook, gated identically to
the prompts check (skips init/refresh/help/version). When a tsconfig.json
exists it warns one line (stderr) if the managed file is missing, not
referenced via extends, or out of date; silent for non-TS projects.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(cli): make tsconfig setup equivalent to prompts (auto-wire + stale-only)
Unify the two managed-file systems so they behave identically:
- auto-wire an existing unlinked tsconfig.json/deno.json on init/refresh
(add extends / importMap; merge into an array extends), instead of only
warning. Parses JSON and falls back to a warning when it can't round-trip
(JSONC comments, or a conflicting deno imports/importMap) — never corrupts.
- narrow warnIfTsconfigStale to stale-only, gated on the managed file
existing, exactly like warnIfPromptsStale: it no longer nags about a
missing or unlinked tsconfig.json, so a deliberately-custom/unlinked setup
stays silent and a not-yet-initialized project isn't bothered.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cli): place tsconfig.wmill.json first in extends to preserve user base config
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cli): migrate legacy tsconfig and require consent for custom configs
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(cli): align prompts wiring to the same consent model
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(cli): bump windmill-parser-wasm-ts to 1.714.0 for $f/ $u/ aliases
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(worker): resolve $f/ and $u/ in deno lock generation
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: narrow relative-imports lock-gen guard to deno import-map failure
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(cli): sync bun.lock with windmill-parser-wasm-ts 1.714.0
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cli): warn when a custom tsconfig's paths would shadow $f/ $u/ aliases
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: add datatable tool coverage to global ai_evals (stage 0+1)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: add seeded datatable difficulty-ladder global ai_evals (stage 2)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: skipJudge datatable evals and make stringIncludesAnyOf existential
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: make ai_evals datatable mock reflect SQL writes
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: gate workspace fork creation in sessions behind enterprise license
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: gate session fork creation on CE workspace cap, not EE license
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The variable and resource value caches (backing
`GET /api/w/{w}/variables/get_value/{path}?allow_cache=true` and
`.../resources/get_value_interpolated/{path}?allow_cache=true`) are consulted
before the per-folder RLS query and store the already-decrypted value. The
resource cache was keyed only by `workspace:path` with no caller identity, so a
cache entry warmed by a privileged peer using `allow_cache=true` could be
returned to a caller with no access to the resource's folder on a cache hit
within the 30s TTL — leaking another folder's decrypted secrets.
Scope both caches to the caller's full authorization identity. The key is now
`auth_identity(authed):workspace:path`, where `auth_identity` is a SHA-256 of the
caller's effective authorization context (email, username, is_admin, is_operator,
sorted groups, sorted folders, sorted scopes) — mirroring
`job_read_access_cache_key`. Email alone is insufficient: the same email can
resolve to different effective permissions via job/owner-scoped tokens, so a
lower-privilege context must not reuse a higher-privilege context's entry.
Job-context resource interpolation is handled correctly: only `$WM_*` contextual
variables are resolved (and only when a `job_id` is present). The interpolation
reports whether the value contains a `$WM_*` placeholder
(`transform_json_value_tracked` + an `AtomicBool`). A value containing one is
job-dependent — even on a no-job read where it's left unresolved — and is never
cached (so a later job read never gets a stale placeholder or another job's
context). Any value without a `$WM_*` placeholder is job-independent and cached
under the identity key, shared across job contexts, so reads carrying a `job_id`
still hit the cache.
BEHAVIOR CHANGE: custom workspace environment variables are no longer interpolated
into resource values via `$NAME` (this was undocumented and prevented caching of
any `$`-prefixed value). Custom envs remain available to scripts/workers as before.
Built-in `$WM_*` contextual variables in resource values are unchanged.
The variable cache previously wrote with an identity-scoped key but read with the
unscoped key, so it never hit (a latent functional bug that happened to be safe).
Aligning the read path enables the cache and makes it identity-scoped by
construction. Secret variables are cached too, but the entry carries the
`is_secret` flag so a cache hit re-runs the per-read side effects a secret read
performs — the EE `variables.decrypt_secret` audit and running-job secret
registration (factored into `audit_decrypt_secret`, shared by both paths).
The unused `invalidate_{variable,resource}_cache` helpers can no longer target
identity-scoped entries; documented the constraint and refreshed the stale
key-format docs on the cache statics.
Tests:
- integration regression for both caches: a folder-scoped user warms the cache via
allow_cache=true, then a user without folder access is denied (401) and never
receives the cached value.
- integration regression that variables (secret included) are served from cache.
- integration regression for job context: plain and non-`$WM_` `$`-string resources
stay cached and are served under a job_id, while a `$WM_*` resource (warmed without
a job_id) is not cached.
- unit tests for `auth_identity`.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): resolve MCP resource token via caller RLS + SSRF-guard url
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): clone user_db for oauth2 refresh and drop advisory ids from comments
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): disable redirects on MCP client to prevent SSRF bypass
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>