mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-10-03 00:02:08 +00:00
* feat(ai-agent): add autocompacted memory that summarizes older context Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): compact on final-answer turns and count what a turn appended Address the pre-push review findings on the compaction path: - A turn the model answers without a tool call left the agent loop on its first iteration, so a chat-shaped step never compacted and reloaded the whole conversation on every later turn. Compaction now also runs after the loop. - The trigger measured only the last request, so a single large tool result could carry the next one past the window without ever crossing 80%. - The summarization call re-sent the usage-tracking request shape on endpoints the loop had already learned to drop it for. - The flat 8000-token summary reserve swallowed the whole target on a small context window, leaving one message in the tail and summarizing the rest. - A response cut off inside the <analysis> scratchpad was accepted as a summary. - The chat-mode memory default was a shared object the step form edited in place. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): keep Anthropic prompt counts and compact once per response Address the first CI review round on the compaction path: - Anthropic's streaming parser dropped `message_start`, the only event carrying the prompt-side counts, so a native Anthropic run reported no input tokens at all and compaction fell back to a character estimate. - A loop that exits without issuing another request — a structured-output turn does — reached the post-loop pass still holding the previous measurement and compacted a second time, or retried a failure with nothing changed. - The summarization call inherited the step's `max_completion_tokens`; a low one truncates the summary inside its scratchpad, which counts as a failure and disables compaction after three of them. - A fired trigger that found nothing to summarize said nothing. - Memory already over the window — a lowered `context_window`, or a step moved over from `auto` — had no way back, since compaction only ran after an accepted request. It now also runs once before the first one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): state the summary's own completion cap and drop the pre-flight pass - The summarization call asked for no completion cap at all, which is "uncapped" only on the OpenAI-shaped providers: Anthropic substitutes 64000, over several Claude models' output ceiling, and Bedrock leaves the model's own small default, short enough to cut the response off inside its scratchpad. It now asks for the reserve the split already set aside, raised to the step's cap when that is larger. - Compaction no longer runs before the first request. The fallbacks the loop learns from a rejection are not known that early, so on exactly the endpoints that need them the summarization was malformed by construction: it failed, spent a strike, and the first agent request still carried the oversized conversation. A memory already past the window is repaired on the turn after a request the endpoint accepts, rather than by a pass that cannot succeed there. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): ask the summary for exactly the room the split reserved The split scales its reserve down on a small window while the request asked for a flat 8000, so the two diverged below an 80k window: on a 4k/8k model the cap alone exceeded the window and every summarization was refused, and on a 20k one a full-length summary could land the conversation back over the trigger and compact its own previous summary on the next response. Both now read one `summary_reserve_tokens`. The call also no longer inherits the step's reasoning effort. Every provider counts thinking against that same budget, so a high-effort model could spend the whole reserve before writing anything and return a summary cut off inside its scratchpad; the compaction prompt asks for an `<analysis>` block, which is the reasoning this call needs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): charge the compaction budget for tools and the system prompt The tail budget was the whole target, but a request also carries the system prompt compaction keeps and the tool definitions, which are not in the message list at all. On a small window those are most of it: a tail sized to the full target left the next request back over the trigger, compacting again every response, and the no-usage estimate missed the tool schemas entirely so it could fail to trigger at all. Both now account for them. The reserve also gains a floor. It is the summary's output cap as well as the room the split leaves, and scaled down without one a small window gave a structured nine-section summary a few hundred tokens — truncated inside its scratchpad every time, which is discarded, which switches the mode off after three. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): count Gemini's tool-use prompt tokens in an agent step's usage Gemini splits a tool-using turn's input across `promptTokenCount` and a disjoint `toolUsePromptTokenCount`, and its thinking apart from `candidatesTokenCount`. The agent step's parser read only the headline fields, so every tool-using turn under-reported both — and the compaction trigger, which runs off the reported prompt, could not see the tool results that grew it. It now goes through the same helpers the proxy path already used. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): calibrate the compaction estimate against the measured prompt Two rounds running, the finding was "the character estimate cannot see input X" — tool schemas, then S3 attachments, which are short paths in the message list and whole images by the time a provider counts them. Enumerating those is a list that only grows, so the estimate is now scaled to the one number that is ground truth: what the provider charged for the last request. Attachments, tokenizer drift and whatever comes next fall out of that, because the estimate is only ever used relative to itself. Also stop the Gemini helpers turning an absent count into `Some(0)`. Downstream, absent means "fall back to estimating the conversation" while zero reads as an empty prompt and would hold the trigger below its threshold for the whole run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): charge attachments what they cost and let a heavy short prefix compact The calibration conserved the conversation's total cost but spread it by character count, so an attachment — a short S3 path in the message list, a whole image or PDF once a provider expands it — was charged to the text messages around it and stayed nearly free in the split. It now carries a nominal cost of its own, which the calibration corrects a residual on rather than the whole gap. The four-message minimum also refused exactly the case that fix is for: an attachment arriving on the first or second turn can pass the trigger before four removable messages exist, and summarizing even one of them saves most of the prompt. A prefix worth a quarter of the window is now enough on its own. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): never summarize a prefix holding only a previous summary The message-count floor was carrying a second job: a fresh summary sits in a one or two message prefix, so requiring four declined it. The share threshold added last commit admits it, and a summary is reserve-sized by construction — so the post-compaction shape could spend one summarization per response swapping a summary for another the same size, shrinking nothing and losing fidelity each time. A previous summary no longer counts towards that threshold. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): take the context window from the model and drop the estimate calibration Brings compaction in line with how the AI session does the same job, which had already answered these three questions. - The window is looked up from the model. `MODEL_CONTEXT_WINDOWS` in `windmill-ai/src/model_context.rs` mirrors the session's table in `copilot/modelConfig.ts`, entry for entry and with the same matching rules; each side points at the other, since a model added to one and not the other compacts at two different sizes. A step's `context_window` becomes the override for what the lookup cannot serve, and chat mode writes none. - Provider usage is normalized where the provider's quirk is, not at the consumer. `TokenUsage::with_cache_beside_input` raises `input_tokens` to the whole prompt for Anthropic and Bedrock, which report their cached prefix beside it; the OpenAI shape already counts it inside. `prompt_tokens()` is then just `input_tokens`, rather than inferring the shape from whether a write count is present. - The estimator is no longer calibrated against the measured prompt. The session uses the provider's count when it has one and a chars/4 estimate otherwise, with nothing in between, and a tail sized a little wrong only compacts again a turn later. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * feat(ai-agent): summarize memory down to what the database can store Without an instance object store, memory is a 100KB database row cut from its oldest message, the summary included, so compaction on a mainstream model never got to keep anything across runs. A step that persists there now runs its post-loop compaction pass against the smaller of the model's window and the cap at chars/4, about 25k tokens: the loop keeps the whole window, and what is written is a summary plus a tail that fits. The run logs when that pass summarizes, and how many messages the write dropped when one still overshoots. The editor's storage warning on the option is removed: nothing exposes the instance storage to it, so it keyed on the workspace S3 setting, which is unrelated to where memory goes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): get a complete, billed summary out of every provider Compaction against the real providers turned up four things the stub could not: Gemini and OpenAI's reasoning models think by default and bill it against the same cap the summary must fit in, so the summarization request now asks them for their least (none, low); an OpenAI Responses call that hits max_output_tokens ends in response.incomplete, whose usage the parser dropped, so that summarization went unbilled; a summary that quotes </summary> when it describes its own instruction was cut off at the quote, on the agent step and the AI session alike; and the prefix could end on an unanswered user message, after which the instruction reads as part of that turn (Anthropic merges the two outright). The tail now starts on a user message, and both prompts tell the model the instruction is not part of the conversation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): compact down to half the window, on the agent step and the AI session The gap between the 80% trigger and the target is what one compaction buys, and every summarization request carries most of the window. At a 70% target a 128k model summarized about 13k tokens of prefix for a summary of up to 8k, so each ~100k-token request bought a few turns of room before the next one re-summarized the previous summary. At 50% the same request frees about 30k. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): drop the workspace-S3 memory hint and state the database bound in the tooltip The memory field warned that memory is kept in the database whenever the workspace had no S3 storage. That setting has no bearing on where memory goes: the instance object store decides, and nothing exposes it to the editor. The field's tooltip now describes both memory kinds and states the database bound unconditionally; the run log says what happened. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): send the summarizer its tool history as text The summarization request carries no tool definitions, and Bedrock's Converse API rejects toolUse/toolResult blocks that arrive without them, so on Bedrock every summarization of a prefix holding a tool call failed silently until the breaker tripped. The prefix's tool calls and results now reach the summarizer rendered as text, on the agent step and in the AI session's compaction, which goes through the same proxy. Also drops the TokenUsage::prompt_tokens accessor, which had become a plain read of the normalized input_tokens, and shortens the context window field's description. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): price attachments from the provider count, bound storage in bytes, effort per pro model Addresses two Codex rounds and a leftovers audit. - Attachments were priced at a flat 1500 tokens in the split, so a multi-page PDF (tens of thousands of tokens to the provider, a short S3 path in the message list) could be kept in the tail or leave no prefix worth summarizing. They are now priced from the provider's count for the request that carried them, less that request's text, with the 1500 floor where nothing was counted. - The database storage bound measured the provider's token count, but the 100KB cap is bytes and repetitive text packs several characters per token. The persist pass now measures the serialized conversation. - The summarizer forced `low` on every reasoning model, which the pro variants reject (gpt-5-pro takes only high, gpt-5.2-pro starts at medium); they now get no effort. - Dropped the unused prompt_tokens accessor and its orphaned assert, an unused PartialEq, a needlessly public lookup, and fully-qualified Gemini calls; refreshed stale comments and the memory_id schema doc; regenerated the flow schema artifacts. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): evict a heavy attachment into the summarized prefix, not the tail Pricing attachments from the provider count was not enough on its own: a leading attachment is a user message, and the boundary rule pulled the last unanswered user turn back into the kept tail to keep it with its answer. For a heavy attachment that dragged it into the tail — or, at the front, emptied the prefix — so it was never summarized and rode every request. The boundary now moves forward instead, keeping that user turn and its answer in the summarized prefix. Verified on the running instance: a 25k-token PDF on a 30k window is summarized out on the turn it overflows, and later turns drop from 26k to ~1.5k tokens. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): keep the forward boundary move off tool results and the prefix start The forward move that keeps an unanswered user turn out of the tail had two edges the third Codex round found: advancing past the user could land the boundary on a tool result (its tool_calls then summarized away, orphaning it), and with no system prompt the summarizable prefix starts at 0, so a trigger firing while the tail estimate fit everything indexed below the start and panicked the task. The forward scan now skips tool-opening boundaries, and the move is guarded above the prefix start. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): drop the step temperature from the summary request OpenAI's reasoning models (gpt-5-mini, gpt-5.1, gpt-5.2) reject `temperature` alongside any reasoning effort but their own default, so a step configured with a temperature made every summarization fail once the summarizer forced a low effort — history then grew unchecked. The internal summary call now omits the step's temperature: a structured extraction does not need a set one, and omitting it sidesteps each provider's temperature-versus-reasoning rules. Confirmed against the API that low + temperature is refused on those models. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): compact an oversized loaded memory before the first request Compaction was reactive, taken only after a request the endpoint accepted, so the fallbacks the loop learns from a rejection are known first. But a memory loaded from an earlier run can already exceed this run's window — the step was switched to a smaller model, or a run under a wider one persisted more than fits — and that first request then overflows and fails the run, with every retry reloading the same history and failing again. A pass is now taken up front, off the character estimate, before the first request. It uses the default request shape; an endpoint needing a fallback may reject this one summary, which is non-fatal, and mainstream providers need none. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): under the storage bound, trigger on the max of bytes and model tokens The storage-bound pass measured only the serialized row size, so an attachment — a few bytes as an S3 path but nearly the whole model context — read as tiny and the pass skipped a compaction the model needed. It now takes the larger of the byte measure and the model's token count, since repetitive text is few tokens but many bytes and an attachment is the reverse; either being over must fire a pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): drop oldest turns when a summary cannot fit the window, as the AI session does An oversized loaded memory (a step switched to a smaller model, or an object-store run that persisted more than a later model's window holds) left a prefix larger than the summarizer's own window, so the summary request overflowed and failed, the memory was untouched, and every retry failed the same way. The AI session handles this by falling back from summarization to dropping the oldest turns down to the target; compaction here now does the same. When a summary cannot run — it failed, the breaker is tripped, or nothing is worth folding — the oldest turns are dropped until the conversation fits and opens on a user message, keeping the newest turn. The next request then always fits. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): drop whole turns only, keep the storage pass to bytes, refresh the count after a rewrite Three edges the seventh Codex round found, all in the drop-oldest fallback and the storage-bound measure: - drop_oldest_to_fit dropped to any point that freed enough, which could strand a tool result whose tool_calls went with the messages before it. It now drops whole turns only, always landing the boundary on a user message and never splitting the newest turn; a lone turn too big for the window is left whole rather than broken. - The storage-bound pass measured the whole model prompt against the shrunk 25k window, so a large tool roster and the system prompt — neither written to the row — tripped it on a conversation the row easily held. It measures the serialized bytes alone now; the model's own window is enforced by the in-loop passes and the pre-first-request pass, so the persisted size is all this pass is for. - A compaction rewrites the message list, so the provider's count for the request that produced it no longer lines up. The count is now cleared after any pass that rewrites the conversation, so a later pass measures the estimate over the actual messages instead of a stale, larger prompt (which could decline a summary that already fit and then drop it). The step temperature, no longer sent to the summarizer on any path, is dropped from the request struct. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): measure only the persisted messages against the storage cap Persistence strips the system prompt before writing the memory row, but the storage pass was serializing every message including it, so a large system prompt with a tiny conversation reported far over the storage trigger, and the fallback dropped the one real turn, run after run. The storage measure now serializes only the non-system messages, matching what the row actually holds. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix(ai-agent): run the model-window pass before the storage-bytes pass post-loop A turn the model answered without a tool call broke before the in-loop compaction check, so on database-backed memory its only pass was the storage one, which measures bytes. An attachment fills the model context but is a few bytes in the row, so that turn never compacted and a follow-up could overflow the model. The post-loop now runs a model-window pass first, off the provider's count, then the storage-bytes pass when the row is smaller than the model — both limits enforced for a chat-shaped step, not just the one that happens to bind. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix: simplify agent compaction and preserve execution history * fix: remove unused compaction history setting * fix: preserve answers and recover rejected agent context * refactor: make agent compaction transactional * fix: skip agent summaries that cannot fit retained context * fix: explain skipped agent context compaction * fix: retain recent agent memory when storage compaction cannot fit * fix: start retained agent memory at a user turn * fix: reject unsafe agent memory truncation on storage fallback * docs: clarify agent context window override scope * fix: keep recent turns verbatim when compaction memory outgrows storage Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix: keep the compaction summary out of the agent's answers Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * fix: shorten the agent context window help text Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH * chore: update ee-repo-ref to 942d4013f36edac1fc9a9addbdb02198db1c7a05 This commit updates the EE repository reference after PR #812 was merged in windmill-ee-private. Previous ee-repo-ref: 8ca1682ce6106ba6ea96894fbe606dac64102eb6 New ee-repo-ref: 942d4013f36edac1fc9a9addbdb02198db1c7a05 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> Co-authored-by: Ruben Fiszel <ruben@windmill.dev>