c1b59f70dd feat(ai-agent): add compaction memory that summarizes older context (#10928)
* feat(ai-agent): add autocompacted memory that summarizes older context

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): compact on final-answer turns and count what a turn appended

Address the pre-push review findings on the compaction path:

- A turn the model answers without a tool call left the agent loop on its first
  iteration, so a chat-shaped step never compacted and reloaded the whole
  conversation on every later turn. Compaction now also runs after the loop.
- The trigger measured only the last request, so a single large tool result
  could carry the next one past the window without ever crossing 80%.
- The summarization call re-sent the usage-tracking request shape on endpoints
  the loop had already learned to drop it for.
- The flat 8000-token summary reserve swallowed the whole target on a small
  context window, leaving one message in the tail and summarizing the rest.
- A response cut off inside the <analysis> scratchpad was accepted as a summary.
- The chat-mode memory default was a shared object the step form edited in place.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): keep Anthropic prompt counts and compact once per response

Address the first CI review round on the compaction path:

- Anthropic's streaming parser dropped `message_start`, the only event carrying
  the prompt-side counts, so a native Anthropic run reported no input tokens at
  all and compaction fell back to a character estimate.
- A loop that exits without issuing another request — a structured-output turn
  does — reached the post-loop pass still holding the previous measurement and
  compacted a second time, or retried a failure with nothing changed.
- The summarization call inherited the step's `max_completion_tokens`; a low one
  truncates the summary inside its scratchpad, which counts as a failure and
  disables compaction after three of them.
- A fired trigger that found nothing to summarize said nothing.
- Memory already over the window — a lowered `context_window`, or a step moved
  over from `auto` — had no way back, since compaction only ran after an
  accepted request. It now also runs once before the first one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): state the summary's own completion cap and drop the pre-flight pass

- The summarization call asked for no completion cap at all, which is "uncapped"
  only on the OpenAI-shaped providers: Anthropic substitutes 64000, over several
  Claude models' output ceiling, and Bedrock leaves the model's own small default,
  short enough to cut the response off inside its scratchpad. It now asks for the
  reserve the split already set aside, raised to the step's cap when that is larger.
- Compaction no longer runs before the first request. The fallbacks the loop learns
  from a rejection are not known that early, so on exactly the endpoints that need
  them the summarization was malformed by construction: it failed, spent a strike,
  and the first agent request still carried the oversized conversation. A memory
  already past the window is repaired on the turn after a request the endpoint
  accepts, rather than by a pass that cannot succeed there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): ask the summary for exactly the room the split reserved

The split scales its reserve down on a small window while the request asked for
a flat 8000, so the two diverged below an 80k window: on a 4k/8k model the cap
alone exceeded the window and every summarization was refused, and on a 20k one
a full-length summary could land the conversation back over the trigger and
compact its own previous summary on the next response. Both now read one
`summary_reserve_tokens`.

The call also no longer inherits the step's reasoning effort. Every provider
counts thinking against that same budget, so a high-effort model could spend the
whole reserve before writing anything and return a summary cut off inside its
scratchpad; the compaction prompt asks for an `<analysis>` block, which is the
reasoning this call needs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): charge the compaction budget for tools and the system prompt

The tail budget was the whole target, but a request also carries the system
prompt compaction keeps and the tool definitions, which are not in the message
list at all. On a small window those are most of it: a tail sized to the full
target left the next request back over the trigger, compacting again every
response, and the no-usage estimate missed the tool schemas entirely so it could
fail to trigger at all. Both now account for them.

The reserve also gains a floor. It is the summary's output cap as well as the
room the split leaves, and scaled down without one a small window gave a
structured nine-section summary a few hundred tokens — truncated inside its
scratchpad every time, which is discarded, which switches the mode off after
three.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): count Gemini's tool-use prompt tokens in an agent step's usage

Gemini splits a tool-using turn's input across `promptTokenCount` and a disjoint
`toolUsePromptTokenCount`, and its thinking apart from `candidatesTokenCount`.
The agent step's parser read only the headline fields, so every tool-using turn
under-reported both — and the compaction trigger, which runs off the reported
prompt, could not see the tool results that grew it. It now goes through the same
helpers the proxy path already used.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): calibrate the compaction estimate against the measured prompt

Two rounds running, the finding was "the character estimate cannot see input X"
— tool schemas, then S3 attachments, which are short paths in the message list
and whole images by the time a provider counts them. Enumerating those is a list
that only grows, so the estimate is now scaled to the one number that is ground
truth: what the provider charged for the last request. Attachments, tokenizer
drift and whatever comes next fall out of that, because the estimate is only
ever used relative to itself.

Also stop the Gemini helpers turning an absent count into `Some(0)`. Downstream,
absent means "fall back to estimating the conversation" while zero reads as an
empty prompt and would hold the trigger below its threshold for the whole run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): charge attachments what they cost and let a heavy short prefix compact

The calibration conserved the conversation's total cost but spread it by
character count, so an attachment — a short S3 path in the message list, a whole
image or PDF once a provider expands it — was charged to the text messages around
it and stayed nearly free in the split. It now carries a nominal cost of its own,
which the calibration corrects a residual on rather than the whole gap.

The four-message minimum also refused exactly the case that fix is for: an
attachment arriving on the first or second turn can pass the trigger before four
removable messages exist, and summarizing even one of them saves most of the
prompt. A prefix worth a quarter of the window is now enough on its own.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): never summarize a prefix holding only a previous summary

The message-count floor was carrying a second job: a fresh summary sits in a one
or two message prefix, so requiring four declined it. The share threshold added
last commit admits it, and a summary is reserve-sized by construction — so the
post-compaction shape could spend one summarization per response swapping a
summary for another the same size, shrinking nothing and losing fidelity each
time. A previous summary no longer counts towards that threshold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): take the context window from the model and drop the estimate calibration

Brings compaction in line with how the AI session does the same job, which had
already answered these three questions.

- The window is looked up from the model. `MODEL_CONTEXT_WINDOWS` in
  `windmill-ai/src/model_context.rs` mirrors the session's table in
  `copilot/modelConfig.ts`, entry for entry and with the same matching rules;
  each side points at the other, since a model added to one and not the other
  compacts at two different sizes. A step's `context_window` becomes the
  override for what the lookup cannot serve, and chat mode writes none.
- Provider usage is normalized where the provider's quirk is, not at the
  consumer. `TokenUsage::with_cache_beside_input` raises `input_tokens` to the
  whole prompt for Anthropic and Bedrock, which report their cached prefix
  beside it; the OpenAI shape already counts it inside. `prompt_tokens()` is
  then just `input_tokens`, rather than inferring the shape from whether a
  write count is present.
- The estimator is no longer calibrated against the measured prompt. The
  session uses the provider's count when it has one and a chars/4 estimate
  otherwise, with nothing in between, and a tail sized a little wrong only
  compacts again a turn later.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* feat(ai-agent): summarize memory down to what the database can store

Without an instance object store, memory is a 100KB database row cut
from its oldest message, the summary included, so compaction on a
mainstream model never got to keep anything across runs. A step that
persists there now runs its post-loop compaction pass against the
smaller of the model's window and the cap at chars/4, about 25k tokens:
the loop keeps the whole window, and what is written is a summary plus a
tail that fits. The run logs when that pass summarizes, and how many
messages the write dropped when one still overshoots.

The editor's storage warning on the option is removed: nothing exposes
the instance storage to it, so it keyed on the workspace S3 setting,
which is unrelated to where memory goes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): get a complete, billed summary out of every provider

Compaction against the real providers turned up four things the stub
could not: Gemini and OpenAI's reasoning models think by default and
bill it against the same cap the summary must fit in, so the
summarization request now asks them for their least (none, low); an
OpenAI Responses call that hits max_output_tokens ends in
response.incomplete, whose usage the parser dropped, so that
summarization went unbilled; a summary that quotes </summary> when it
describes its own instruction was cut off at the quote, on the agent
step and the AI session alike; and the prefix could end on an unanswered
user message, after which the instruction reads as part of that turn
(Anthropic merges the two outright). The tail now starts on a user
message, and both prompts tell the model the instruction is not part of
the conversation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): compact down to half the window, on the agent step and the AI session

The gap between the 80% trigger and the target is what one compaction
buys, and every summarization request carries most of the window. At a
70% target a 128k model summarized about 13k tokens of prefix for a
summary of up to 8k, so each ~100k-token request bought a few turns of
room before the next one re-summarized the previous summary. At 50% the
same request frees about 30k.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): drop the workspace-S3 memory hint and state the database bound in the tooltip

The memory field warned that memory is kept in the database whenever the
workspace had no S3 storage. That setting has no bearing on where memory
goes: the instance object store decides, and nothing exposes it to the
editor. The field's tooltip now describes both memory kinds and states
the database bound unconditionally; the run log says what happened.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): send the summarizer its tool history as text

The summarization request carries no tool definitions, and Bedrock's
Converse API rejects toolUse/toolResult blocks that arrive without them,
so on Bedrock every summarization of a prefix holding a tool call failed
silently until the breaker tripped. The prefix's tool calls and results
now reach the summarizer rendered as text, on the agent step and in the
AI session's compaction, which goes through the same proxy.

Also drops the TokenUsage::prompt_tokens accessor, which had become a
plain read of the normalized input_tokens, and shortens the context
window field's description.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): price attachments from the provider count, bound storage in bytes, effort per pro model

Addresses two Codex rounds and a leftovers audit.

- Attachments were priced at a flat 1500 tokens in the split, so a
  multi-page PDF (tens of thousands of tokens to the provider, a short
  S3 path in the message list) could be kept in the tail or leave no
  prefix worth summarizing. They are now priced from the provider's
  count for the request that carried them, less that request's text,
  with the 1500 floor where nothing was counted.
- The database storage bound measured the provider's token count, but
  the 100KB cap is bytes and repetitive text packs several characters
  per token. The persist pass now measures the serialized conversation.
- The summarizer forced `low` on every reasoning model, which the pro
  variants reject (gpt-5-pro takes only high, gpt-5.2-pro starts at
  medium); they now get no effort.
- Dropped the unused prompt_tokens accessor and its orphaned assert, an
  unused PartialEq, a needlessly public lookup, and fully-qualified
  Gemini calls; refreshed stale comments and the memory_id schema doc;
  regenerated the flow schema artifacts.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): evict a heavy attachment into the summarized prefix, not the tail

Pricing attachments from the provider count was not enough on its own: a
leading attachment is a user message, and the boundary rule pulled the
last unanswered user turn back into the kept tail to keep it with its
answer. For a heavy attachment that dragged it into the tail — or, at
the front, emptied the prefix — so it was never summarized and rode
every request. The boundary now moves forward instead, keeping that
user turn and its answer in the summarized prefix. Verified on the
running instance: a 25k-token PDF on a 30k window is summarized out on
the turn it overflows, and later turns drop from 26k to ~1.5k tokens.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): keep the forward boundary move off tool results and the prefix start

The forward move that keeps an unanswered user turn out of the tail had
two edges the third Codex round found: advancing past the user could
land the boundary on a tool result (its tool_calls then summarized away,
orphaning it), and with no system prompt the summarizable prefix starts
at 0, so a trigger firing while the tail estimate fit everything indexed
below the start and panicked the task. The forward scan now skips
tool-opening boundaries, and the move is guarded above the prefix start.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): drop the step temperature from the summary request

OpenAI's reasoning models (gpt-5-mini, gpt-5.1, gpt-5.2) reject
`temperature` alongside any reasoning effort but their own default, so a
step configured with a temperature made every summarization fail once
the summarizer forced a low effort — history then grew unchecked. The
internal summary call now omits the step's temperature: a structured
extraction does not need a set one, and omitting it sidesteps each
provider's temperature-versus-reasoning rules. Confirmed against the API
that low + temperature is refused on those models.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): compact an oversized loaded memory before the first request

Compaction was reactive, taken only after a request the endpoint
accepted, so the fallbacks the loop learns from a rejection are known
first. But a memory loaded from an earlier run can already exceed this
run's window — the step was switched to a smaller model, or a run under
a wider one persisted more than fits — and that first request then
overflows and fails the run, with every retry reloading the same
history and failing again. A pass is now taken up front, off the
character estimate, before the first request. It uses the default
request shape; an endpoint needing a fallback may reject this one
summary, which is non-fatal, and mainstream providers need none.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): under the storage bound, trigger on the max of bytes and model tokens

The storage-bound pass measured only the serialized row size, so an
attachment — a few bytes as an S3 path but nearly the whole model
context — read as tiny and the pass skipped a compaction the model
needed. It now takes the larger of the byte measure and the model's
token count, since repetitive text is few tokens but many bytes and an
attachment is the reverse; either being over must fire a pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): drop oldest turns when a summary cannot fit the window, as the AI session does

An oversized loaded memory (a step switched to a smaller model, or an
object-store run that persisted more than a later model's window holds)
left a prefix larger than the summarizer's own window, so the summary
request overflowed and failed, the memory was untouched, and every
retry failed the same way. The AI session handles this by falling back
from summarization to dropping the oldest turns down to the target;
compaction here now does the same. When a summary cannot run — it
failed, the breaker is tripped, or nothing is worth folding — the oldest
turns are dropped until the conversation fits and opens on a user
message, keeping the newest turn. The next request then always fits.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): drop whole turns only, keep the storage pass to bytes, refresh the count after a rewrite

Three edges the seventh Codex round found, all in the drop-oldest
fallback and the storage-bound measure:

- drop_oldest_to_fit dropped to any point that freed enough, which could
  strand a tool result whose tool_calls went with the messages before
  it. It now drops whole turns only, always landing the boundary on a
  user message and never splitting the newest turn; a lone turn too big
  for the window is left whole rather than broken.
- The storage-bound pass measured the whole model prompt against the
  shrunk 25k window, so a large tool roster and the system prompt —
  neither written to the row — tripped it on a conversation the row
  easily held. It measures the serialized bytes alone now; the model's
  own window is enforced by the in-loop passes and the pre-first-request
  pass, so the persisted size is all this pass is for.
- A compaction rewrites the message list, so the provider's count for
  the request that produced it no longer lines up. The count is now
  cleared after any pass that rewrites the conversation, so a later pass
  measures the estimate over the actual messages instead of a stale,
  larger prompt (which could decline a summary that already fit and then
  drop it). The step temperature, no longer sent to the summarizer on
  any path, is dropped from the request struct.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): measure only the persisted messages against the storage cap

Persistence strips the system prompt before writing the memory row, but
the storage pass was serializing every message including it, so a large
system prompt with a tiny conversation reported far over the storage
trigger, and the fallback dropped the one real turn, run after run. The
storage measure now serializes only the non-system messages, matching
what the row actually holds.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): run the model-window pass before the storage-bytes pass post-loop

A turn the model answered without a tool call broke before the in-loop
compaction check, so on database-backed memory its only pass was the
storage one, which measures bytes. An attachment fills the model context
but is a few bytes in the row, so that turn never compacted and a
follow-up could overflow the model. The post-loop now runs a
model-window pass first, off the provider's count, then the
storage-bytes pass when the row is smaller than the model — both limits
enforced for a chat-shaped step, not just the one that happens to bind.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix: simplify agent compaction and preserve execution history

* fix: remove unused compaction history setting

* fix: preserve answers and recover rejected agent context

* refactor: make agent compaction transactional

* fix: skip agent summaries that cannot fit retained context

* fix: explain skipped agent context compaction

* fix: retain recent agent memory when storage compaction cannot fit

* fix: start retained agent memory at a user turn

* fix: reject unsafe agent memory truncation on storage fallback

* docs: clarify agent context window override scope

* fix: keep recent turns verbatim when compaction memory outgrows storage

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix: keep the compaction summary out of the agent's answers

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix: shorten the agent context window help text

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* chore: update ee-repo-ref to 942d4013f36edac1fc9a9addbdb02198db1c7a05

This commit updates the EE repository reference after PR #812 was merged in windmill-ee-private.

Previous ee-repo-ref: 8ca1682ce6106ba6ea96894fbe606dac64102eb6

New ee-repo-ref: 942d4013f36edac1fc9a9addbdb02198db1c7a05

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
2026-09-29 19:01:53 +02:00
2024-02-08 16:09:11 +01:00
2026-02-03 17:50:56 +00:00
2026-09-29 08:33:39 +02:00
2023-08-26 09:17:07 +02:00
2026-08-19 14:18:37 +00:00
2025-04-21 16:53:23 +02:00
2026-04-16 06:26:25 -07:00
2025-12-08 10:34:16 +00:00
2025-03-07 15:08:19 +01:00
2022-05-05 04:25:58 +02:00
2024-02-08 16:09:11 +01:00

windmill.dev

Open-source developer platform for internal code: APIs, background jobs, workflows and UIs. Self-hostable alternative to Retool, Pipedream, Superblocks and a simplified Temporal with autogenerated UIs and custom UIs to trigger workflows and scripts as internal apps.

Scripts are turned into sharable UIs automatically, and can be composed together into flows or used into richer apps built with low-code. Supported languages: Python, TypeScript, Go, Bash, SQL, GraphQL, PowerShell, Rust, and more.

Package version Docker Image CI Package version

Commit activity Discord Shield

Try it - Website - Docs - Discord - Hub - Contributing

Windmill - Developer platform for APIs, background jobs, workflows and UIs

Windmill is fully open-sourced (AGPLv3) and Windmill Labs offers dedicated instances and commercial support and licenses.

Windmill Diagram

https://github.com/user-attachments/assets/d80de1d9-64de-4d89-aacd-6df23fa81fc4

Main Concepts

  1. Define a minimal and generic script in Python, TypeScript, Go or Bash that solves a specific task. The code can be defined in the provided Web IDE or synchronized with your own GitHub repo (e.g. through VS Code extension): provided Web IDE or synchronized with your own GitHub repo (e.g. through VS Code extension):

Step 1

  1. Your scripts parameters are automatically parsed and generate a frontend.

Step 2

Step 3

  1. Make it flow! You can chain your scripts or scripts made by the community shared on WindmillHub.

Step 3

  1. Build complex UIs on top of your scripts and flows.

Step 4

Scripts and flows can be triggered by schedules, webhooks, HTTP routes, Kafka, WebSockets, emails, and more.

Build your entire infra on top of Windmill!

Show me some actual script code

//import any dependency  from npm
import * as wmill from "windmill-client";
import * as cowsay from "cowsay@1.5.0";

// fill the type, or use the +Resource type to get a type-safe reference to a resource
type Postgresql = {
  host: string;
  port: number;
  user: string;
  dbname: string;
  sslmode: string;
  password: string;
};

export async function main(
  a: number,
  b: "my" | "enum",
  c: Postgresql,
  d = "inferred type string from default arg",
  e = { nested: "object" }
  //f: wmill.Base64
) {
  const email = process.env["WM_EMAIL"];
  // variables are permissioned and by path
  let variable = await wmill.getVariable("f/company-folder/my_secret");
  const lastTimeRun = await wmill.getState();
  // logs are printed and always inspectable
  console.log(cowsay.say({ text: "hello " + email + " " + lastTimeRun }));
  await wmill.setState(Date.now());

  // return is serialized as JSON
  return { foo: d, variable };
}

Local Development

Windmill supports multiple ways to develop locally and sync with your instance:

Tool Description
CLI Sync scripts from local files or GitHub, run scripts/flows from the command line
VS Code Extension Edit and test scripts & flows directly from VS Code / Cursor with full IDE support
Git Sync Two-way sync between Windmill and your Git repository
Claude Code AI-assisted development with Claude for scripts, flows, and apps

https://github.com/user-attachments/assets/c541c326-e9ae-4602-a09a-1989aaded1e9

You can run scripts locally by passing the right environment variables for the wmill client library to fetch resources and variables from your instance. See local development docs.

Stack

  • Database: Postgres (compatible with Aurora, Cloud SQL, Neon, Azure PostgreSQL)
  • Backend: Rust - stateless API servers and workers pulling jobs from a Postgres queue
  • Frontend: Svelte 5
  • Sandboxing: nsjail and PID namespace isolation
  • Runtimes:
    • TypeScript/JavaScript: Bun (default) and Deno
    • Python: python3 with uv for dependency management
    • Go, Bash, PowerShell, PHP, Rust, C#, Java, Ansible

Fastest Self-Hostable Workflow Engine

We have compared Windmill to other self-hostable workflow engines (Airflow, Prefect & Temporal) and Windmill is the most performant solution for both benchmarks: one flow composed of 40 lightweight tasks & one flow composed of 10 long-running tasks.

All methodology & results on our Benchmarks page.

Fastest workflow engine

Security

  • Sandboxing: nsjail for filesystem/resource isolation, and PID namespace isolation (enabled by default) to prevent jobs from accessing worker process memory
  • Secrets: One encryption key per workspace for credentials stored in Windmill's K/V store. We recommend encrypting the Postgres database as well.

See Security documentation for details.

Performance

Once a job started, there is no overhead compared to running the same script on the node with its corresponding runner (Deno/Go/Python/Bash). The added latency from a job being pulled from the queue, started, and then having its result sent back to the database is ~50ms. A typical lightweight deno job will take around 100ms total.

Architecture

How to self-host

For detailed setup options, see Self-Host documentation.

Docker compose

Deploy Windmill with 3 files (docker-compose.yml, Caddyfile, .env):

curl https://raw.githubusercontent.com/windmill-labs/windmill/main/docker-compose.yml -o docker-compose.yml
curl https://raw.githubusercontent.com/windmill-labs/windmill/main/Caddyfile -o Caddyfile
curl https://raw.githubusercontent.com/windmill-labs/windmill/main/.env -o .env

docker compose up -d

Go to http://localhost - default credentials: admin@windmill.dev / changeme

Using an external database: Set DATABASE_URL in .env to point to your managed Postgres (AWS RDS, GCP Cloud SQL, Azure, Neon, etc.) and set db replicas to 0.

Kubernetes (Helm charts)

helm repo add windmill https://windmill-labs.github.io/windmill-helm-charts/
helm install windmill-chart windmill/windmill --namespace=windmill --create-namespace

See windmill-helm-charts for configuration options.

Cloud providers

Windmill works on AWS (EKS/ECS), GCP, Azure, Ubicloud, Fly.io, Render.com, Hetzner, Digital Ocean, and others. Rule of thumb: 1 worker per 1vCPU and 1-2 GB RAM.

OAuth, SSO & SMTP

Configure OAuth and SSO (Google Workspace, Microsoft/Azure, Okta) directly from the superadmin UI. See documentation.

License

The Community Edition is free to use internally. For commercial redistribution or managed services, contact sales@windmill.dev. See LICENSE and Pricing for details.

The "Community Edition" of Windmill available in the docker images hosted under ghcr.io/windmill-labs/windmill and the github binary releases contains the files under the AGPLv3 and Apache 2 sources but also includes proprietary and non-public code and features which are not open source and under the following terms: Windmill Labs, Inc. grants a right to use all the features of the "Community Edition" for free without restrictions other than the limits and quotas set in the software and a right to distribute the community edition as is but not to sell, resell, serve Windmill as a managed service, modify or wrap under any form without an explicit agreement.

The binary compilable from source code in this repository without the "enterprise" feature flag is open-source under the LICENSE-AGPLv3 License terms and conditions.

To re-expose directly any Windmill parts to your users as a feature of your product, with the exception of iframed public Windmill "apps", or to build a feature on top of "Windmill Community Edition" that you sell commercially or embed in a distributable product or binary, you must get a commercial license. Contact us at sales@windmill.dev if you have any questions. To do the same from the binary compiled from the source code in this repository without the "enterprise" feature flag, you must comply with the AGPLv3 license terms and conditions or get a commercial license from Windmill Labs, Inc.

To use Windmill "Community Edition" as is internally in your organization, or to use its APIs as is, you do NOT need a commercial license.

Integrations

In Windmill, integrations are referred to as resources and resource types. Each Resource has a Resource Type that defines the schema that the resource needs to implement.

On self-hosted instances, you might want to import all the approved resource types from WindmillHub. A setup script will prompt you to have it being synced automatically everyday.

Environment Variables

Environment Variable name Default Description Api Server/Worker/All
DATABASE_URL The Postgres database url. All
WORKER_GROUP default The worker group the worker belongs to and get its configuration pulled from Worker
MODE standalone The mode if the binary. Possible values: standalone, worker, server, agent All
METRICS_ADDR None (ee only) The socket addr at which to expose Prometheus metrics at the /metrics path. Set to "true" to expose it on port 8001 All
JSON_FMT false Output the logs in json format instead of logfmt All
BASE_URL http://localhost:8000 The base url that is exposed publicly to access your instance. Is overriden by the instance settings if any. Server
ZOMBIE_JOB_TIMEOUT 30 The timeout after which a job is considered to be zombie if the worker did not send pings about processing the job (every server check for zombie jobs every 30s) Server
RESTART_ZOMBIE_JOBS true If true then a zombie job is restarted (in-place with the same uuid and some logs), if false the zombie job is failed Server
NATIVE_MODE false Enable native mode: sets NUM_WORKERS=8, rejects non-native jobs (nativets, postgresql, mysql, etc.) Worker
SLEEP_QUEUE 50 The number of ms to sleep in between the last check for new jobs in the DB. It is multiplied by NUM_WORKERS such that in average, for one worker instance, there is one pull every SLEEP_QUEUE ms. Worker
KEEP_JOB_DIR false Keep the job directory after the job is done. Useful for debugging. Worker
EXIT_AFTER_N_JOBS None Exit the worker process after it has executed that many jobs, so that a supervisor restarts it and no process runs more than that many, bar the steps of a same-worker flow it has started, which it always finishes (set it to 1 for a process per job; jobs handed to a dedicated worker, and the worker's own init and periodic scripts, do not count). Not counting the init and periodic scripts means they run again on every restart: an init script's runtime is added to the latency of every batch of that many jobs, and a periodic script fires once per process start whatever its interval says. The worker's shell in the workers page also starts backed off rather than after the two minutes it otherwise takes, since a process due to be recycled cannot count on living that long: the first command of a session can wait up to 15s, later ones are immediate. For deployments that isolate executions by process lifetime rather than with nsjail; note that a container restart resets the process, not the container filesystem, so caches and /tmp survive it. The worker name is then derived from the hostname instead of being random, so the restarted worker keeps its row in the workers list (an agent worker keeps the row but restarts its job count). Use one worker per process: workers of one process share its environment, so the first to reach the limit shuts the others down too. Worker
WORKER_SUFFIX None Pins the last part of the worker name, which is otherwise random, so that a restarted worker keeps its row in the workers list. Only needed when several worker processes of the same worker group run on one host, since the name is derived from the hostname: give each of them a distinct value, as two processes sharing one must never happen. At most 64 letters, digits and underscores; anything else is refused at startup. Worker
LICENSE_KEY (EE only) None License key checked at startup for the Enterprise Edition of Windmill Worker
SLACK_SIGNING_SECRET None The signing secret of your Slack app. See Slack documentation Server
COOKIE_DOMAIN None The domain of the cookie. If not set, the cookie will be set by the browser based on the full origin Server
DENO_PATH /usr/bin/deno The path to the deno binary. Worker
PYTHON_PATH The path to the python binary if wanting to not have it managed by uv. Worker
GO_PATH /usr/bin/go The path to the go binary. Worker
GOPRIVATE The GOPRIVATE env variable to use private go modules Worker
GOPROXY The GOPROXY env variable to use Worker
NETRC The netrc content to use a private go registry Worker
PY_CONCURRENT_DOWNLOADS 20 Sets the maximum number of in-flight concurrent python downloads that windmill will perform at any given time. Worker
PATH None The path environment variable, usually inherited Worker
HOME None The home directory to use for Go and Bash , usually inherited Worker
DATABASE_CONNECTIONS 50 (Server)/3 (Worker) The max number of connections in the database connection pool All
SUPERADMIN_SECRET None A token that would let the caller act as a virtual superadmin superadmin@windmill.dev Server
TIMEOUT_WAIT_RESULT 20 The number of seconds to wait before timeout on the 'run_wait_result' endpoint Worker
QUEUE_LIMIT_WAIT_RESULT None The number of max jobs in the queue before rejecting immediately the request in 'run_wait_result' endpoint. Takes precedence on the query arg. If none is specified, there are no limit. Worker
DENO_AUTH_TOKENS None Custom DENO_AUTH_TOKENS to pass to worker to allow the use of private modules Worker
DISABLE_RESPONSE_LOGS false Disable response logs Server
CREATE_WORKSPACE_REQUIRE_SUPERADMIN true If true, only superadmins can create new workspaces Server
MIN_FREE_DISK_SPACE_MB 15000 Minimum amount of free space on worker. Sends critical alert if worker has less free space. Worker
RUN_UPDATE_CA_CERTIFICATE_AT_START false If true, runs CA certificate update command at startup before other initialization All
RUN_UPDATE_CA_CERTIFICATE_PATH /usr/sbin/update-ca-certificates Path to the CA certificate update command/script to run when RUN_UPDATE_CA_CERTIFICATE_AT_START is true All
GOOGLE_APPLICATION_CREDENTIALS None (ee only) Credentials file for GCP Pub/Sub triggers that authenticate as the instance rather than through a gcloud resource (workspace admins only). Application default credentials also resolve the gcloud well-known file and the GCE metadata server. Workload Identity Federation files work with the file, url and aws credential sources; the executable source is not supported. Server

Run a local dev setup

We recommend using Nix. See ./frontend/README_DEV.md for all options.

Frontend only

Uses the backend of https://app.windmill.dev with local frontend (hot-reload):

cd frontend
npm install
npm run generate-backend-client  # or generate-backend-client-mac on Mac
npm run dev

Windmill available at http://localhost/

Backend + Frontend

See the ./frontend/README_DEV.md file for all running options.

  1. Start a local Postgres database using for instance the start-dev-db.sh script which will make a database available at postgres://postgres:changeme@localhost:5432/windmill Then run the migrations using the following command:
    cargo install sqlx-cli
    env DATABASE_URL=<YOUR_DATABASE_URL> sqlx migrate run
    
    This will also avoid compile time issue with sqlx's query! macro.
  2. (optional, linux only) Install nsjail and have it accessible in your PATH
  3. Install bun, deno and python3 (+ any languages you want to use), have the bins at /usr/bin/bun,/usr/bin/deno, and /usr/local/bin/python3 or set the corresponding environment variables.
  4. (optional) Install the lld linker
  5. Go to frontend/:
    1. npm install, npm run generate-backend-client then REMOTE=http://localhost:8000 npm run dev
    2. You might need to set some extra heap space for the node runtime export NODE_OPTIONS="--max-old-space-size=4096"
    3. Create an empty frontend/build folder using mkdir frontend/build
  6. Go to backend/:
    1. env DATABASE_URL=<YOUR_DATABASE_URL> RUST_LOG=info cargo run
    2. You can specify any feature flag you want to enable, for example cargo run --features python to enable the python executor.
  7. Windmill should be available at http://localhost:3000

Contributing

At this time, we are not seeking outside contribution. Bug reports and feature requests remain very welcome, and small, trivially-verified PRs that fix a problem are still accepted. See CONTRIBUTING.md for the full policy.

Contributors

© 2023-2026 Windmill Labs, Inc.

S
Description
Open-source developer platform to power your entire infra and turn scripts into webhooks, workflows and UIs. Fastest workflow engine (13x vs Airflow). Open-source alternative to Retool and Temporal.
Readme
890 MiB
Languages
Rust 34.5%
TypeScript 25.6%
Svelte 23.9%
HTML 12.2%
JavaScript 1.5%
Other 2.1%