feat(ai-agent): add compaction memory that summarizes older context (#10928)

* feat(ai-agent): add autocompacted memory that summarizes older context

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): compact on final-answer turns and count what a turn appended

Address the pre-push review findings on the compaction path:

- A turn the model answers without a tool call left the agent loop on its first
  iteration, so a chat-shaped step never compacted and reloaded the whole
  conversation on every later turn. Compaction now also runs after the loop.
- The trigger measured only the last request, so a single large tool result
  could carry the next one past the window without ever crossing 80%.
- The summarization call re-sent the usage-tracking request shape on endpoints
  the loop had already learned to drop it for.
- The flat 8000-token summary reserve swallowed the whole target on a small
  context window, leaving one message in the tail and summarizing the rest.
- A response cut off inside the <analysis> scratchpad was accepted as a summary.
- The chat-mode memory default was a shared object the step form edited in place.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): keep Anthropic prompt counts and compact once per response

Address the first CI review round on the compaction path:

- Anthropic's streaming parser dropped `message_start`, the only event carrying
  the prompt-side counts, so a native Anthropic run reported no input tokens at
  all and compaction fell back to a character estimate.
- A loop that exits without issuing another request — a structured-output turn
  does — reached the post-loop pass still holding the previous measurement and
  compacted a second time, or retried a failure with nothing changed.
- The summarization call inherited the step's `max_completion_tokens`; a low one
  truncates the summary inside its scratchpad, which counts as a failure and
  disables compaction after three of them.
- A fired trigger that found nothing to summarize said nothing.
- Memory already over the window — a lowered `context_window`, or a step moved
  over from `auto` — had no way back, since compaction only ran after an
  accepted request. It now also runs once before the first one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): state the summary's own completion cap and drop the pre-flight pass

- The summarization call asked for no completion cap at all, which is "uncapped"
  only on the OpenAI-shaped providers: Anthropic substitutes 64000, over several
  Claude models' output ceiling, and Bedrock leaves the model's own small default,
  short enough to cut the response off inside its scratchpad. It now asks for the
  reserve the split already set aside, raised to the step's cap when that is larger.
- Compaction no longer runs before the first request. The fallbacks the loop learns
  from a rejection are not known that early, so on exactly the endpoints that need
  them the summarization was malformed by construction: it failed, spent a strike,
  and the first agent request still carried the oversized conversation. A memory
  already past the window is repaired on the turn after a request the endpoint
  accepts, rather than by a pass that cannot succeed there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): ask the summary for exactly the room the split reserved

The split scales its reserve down on a small window while the request asked for
a flat 8000, so the two diverged below an 80k window: on a 4k/8k model the cap
alone exceeded the window and every summarization was refused, and on a 20k one
a full-length summary could land the conversation back over the trigger and
compact its own previous summary on the next response. Both now read one
`summary_reserve_tokens`.

The call also no longer inherits the step's reasoning effort. Every provider
counts thinking against that same budget, so a high-effort model could spend the
whole reserve before writing anything and return a summary cut off inside its
scratchpad; the compaction prompt asks for an `<analysis>` block, which is the
reasoning this call needs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): charge the compaction budget for tools and the system prompt

The tail budget was the whole target, but a request also carries the system
prompt compaction keeps and the tool definitions, which are not in the message
list at all. On a small window those are most of it: a tail sized to the full
target left the next request back over the trigger, compacting again every
response, and the no-usage estimate missed the tool schemas entirely so it could
fail to trigger at all. Both now account for them.

The reserve also gains a floor. It is the summary's output cap as well as the
room the split leaves, and scaled down without one a small window gave a
structured nine-section summary a few hundred tokens — truncated inside its
scratchpad every time, which is discarded, which switches the mode off after
three.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): count Gemini's tool-use prompt tokens in an agent step's usage

Gemini splits a tool-using turn's input across `promptTokenCount` and a disjoint
`toolUsePromptTokenCount`, and its thinking apart from `candidatesTokenCount`.
The agent step's parser read only the headline fields, so every tool-using turn
under-reported both — and the compaction trigger, which runs off the reported
prompt, could not see the tool results that grew it. It now goes through the same
helpers the proxy path already used.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): calibrate the compaction estimate against the measured prompt

Two rounds running, the finding was "the character estimate cannot see input X"
— tool schemas, then S3 attachments, which are short paths in the message list
and whole images by the time a provider counts them. Enumerating those is a list
that only grows, so the estimate is now scaled to the one number that is ground
truth: what the provider charged for the last request. Attachments, tokenizer
drift and whatever comes next fall out of that, because the estimate is only
ever used relative to itself.

Also stop the Gemini helpers turning an absent count into `Some(0)`. Downstream,
absent means "fall back to estimating the conversation" while zero reads as an
empty prompt and would hold the trigger below its threshold for the whole run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): charge attachments what they cost and let a heavy short prefix compact

The calibration conserved the conversation's total cost but spread it by
character count, so an attachment — a short S3 path in the message list, a whole
image or PDF once a provider expands it — was charged to the text messages around
it and stayed nearly free in the split. It now carries a nominal cost of its own,
which the calibration corrects a residual on rather than the whole gap.

The four-message minimum also refused exactly the case that fix is for: an
attachment arriving on the first or second turn can pass the trigger before four
removable messages exist, and summarizing even one of them saves most of the
prompt. A prefix worth a quarter of the window is now enough on its own.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): never summarize a prefix holding only a previous summary

The message-count floor was carrying a second job: a fresh summary sits in a one
or two message prefix, so requiring four declined it. The share threshold added
last commit admits it, and a summary is reserve-sized by construction — so the
post-compaction shape could spend one summarization per response swapping a
summary for another the same size, shrinking nothing and losing fidelity each
time. A previous summary no longer counts towards that threshold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): take the context window from the model and drop the estimate calibration

Brings compaction in line with how the AI session does the same job, which had
already answered these three questions.

- The window is looked up from the model. `MODEL_CONTEXT_WINDOWS` in
  `windmill-ai/src/model_context.rs` mirrors the session's table in
  `copilot/modelConfig.ts`, entry for entry and with the same matching rules;
  each side points at the other, since a model added to one and not the other
  compacts at two different sizes. A step's `context_window` becomes the
  override for what the lookup cannot serve, and chat mode writes none.
- Provider usage is normalized where the provider's quirk is, not at the
  consumer. `TokenUsage::with_cache_beside_input` raises `input_tokens` to the
  whole prompt for Anthropic and Bedrock, which report their cached prefix
  beside it; the OpenAI shape already counts it inside. `prompt_tokens()` is
  then just `input_tokens`, rather than inferring the shape from whether a
  write count is present.
- The estimator is no longer calibrated against the measured prompt. The
  session uses the provider's count when it has one and a chars/4 estimate
  otherwise, with nothing in between, and a tail sized a little wrong only
  compacts again a turn later.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* feat(ai-agent): summarize memory down to what the database can store

Without an instance object store, memory is a 100KB database row cut
from its oldest message, the summary included, so compaction on a
mainstream model never got to keep anything across runs. A step that
persists there now runs its post-loop compaction pass against the
smaller of the model's window and the cap at chars/4, about 25k tokens:
the loop keeps the whole window, and what is written is a summary plus a
tail that fits. The run logs when that pass summarizes, and how many
messages the write dropped when one still overshoots.

The editor's storage warning on the option is removed: nothing exposes
the instance storage to it, so it keyed on the workspace S3 setting,
which is unrelated to where memory goes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): get a complete, billed summary out of every provider

Compaction against the real providers turned up four things the stub
could not: Gemini and OpenAI's reasoning models think by default and
bill it against the same cap the summary must fit in, so the
summarization request now asks them for their least (none, low); an
OpenAI Responses call that hits max_output_tokens ends in
response.incomplete, whose usage the parser dropped, so that
summarization went unbilled; a summary that quotes </summary> when it
describes its own instruction was cut off at the quote, on the agent
step and the AI session alike; and the prefix could end on an unanswered
user message, after which the instruction reads as part of that turn
(Anthropic merges the two outright). The tail now starts on a user
message, and both prompts tell the model the instruction is not part of
the conversation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): compact down to half the window, on the agent step and the AI session

The gap between the 80% trigger and the target is what one compaction
buys, and every summarization request carries most of the window. At a
70% target a 128k model summarized about 13k tokens of prefix for a
summary of up to 8k, so each ~100k-token request bought a few turns of
room before the next one re-summarized the previous summary. At 50% the
same request frees about 30k.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): drop the workspace-S3 memory hint and state the database bound in the tooltip

The memory field warned that memory is kept in the database whenever the
workspace had no S3 storage. That setting has no bearing on where memory
goes: the instance object store decides, and nothing exposes it to the
editor. The field's tooltip now describes both memory kinds and states
the database bound unconditionally; the run log says what happened.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): send the summarizer its tool history as text

The summarization request carries no tool definitions, and Bedrock's
Converse API rejects toolUse/toolResult blocks that arrive without them,
so on Bedrock every summarization of a prefix holding a tool call failed
silently until the breaker tripped. The prefix's tool calls and results
now reach the summarizer rendered as text, on the agent step and in the
AI session's compaction, which goes through the same proxy.

Also drops the TokenUsage::prompt_tokens accessor, which had become a
plain read of the normalized input_tokens, and shortens the context
window field's description.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): price attachments from the provider count, bound storage in bytes, effort per pro model

Addresses two Codex rounds and a leftovers audit.

- Attachments were priced at a flat 1500 tokens in the split, so a
  multi-page PDF (tens of thousands of tokens to the provider, a short
  S3 path in the message list) could be kept in the tail or leave no
  prefix worth summarizing. They are now priced from the provider's
  count for the request that carried them, less that request's text,
  with the 1500 floor where nothing was counted.
- The database storage bound measured the provider's token count, but
  the 100KB cap is bytes and repetitive text packs several characters
  per token. The persist pass now measures the serialized conversation.
- The summarizer forced `low` on every reasoning model, which the pro
  variants reject (gpt-5-pro takes only high, gpt-5.2-pro starts at
  medium); they now get no effort.
- Dropped the unused prompt_tokens accessor and its orphaned assert, an
  unused PartialEq, a needlessly public lookup, and fully-qualified
  Gemini calls; refreshed stale comments and the memory_id schema doc;
  regenerated the flow schema artifacts.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): evict a heavy attachment into the summarized prefix, not the tail

Pricing attachments from the provider count was not enough on its own: a
leading attachment is a user message, and the boundary rule pulled the
last unanswered user turn back into the kept tail to keep it with its
answer. For a heavy attachment that dragged it into the tail — or, at
the front, emptied the prefix — so it was never summarized and rode
every request. The boundary now moves forward instead, keeping that
user turn and its answer in the summarized prefix. Verified on the
running instance: a 25k-token PDF on a 30k window is summarized out on
the turn it overflows, and later turns drop from 26k to ~1.5k tokens.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): keep the forward boundary move off tool results and the prefix start

The forward move that keeps an unanswered user turn out of the tail had
two edges the third Codex round found: advancing past the user could
land the boundary on a tool result (its tool_calls then summarized away,
orphaning it), and with no system prompt the summarizable prefix starts
at 0, so a trigger firing while the tail estimate fit everything indexed
below the start and panicked the task. The forward scan now skips
tool-opening boundaries, and the move is guarded above the prefix start.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): drop the step temperature from the summary request

OpenAI's reasoning models (gpt-5-mini, gpt-5.1, gpt-5.2) reject
`temperature` alongside any reasoning effort but their own default, so a
step configured with a temperature made every summarization fail once
the summarizer forced a low effort — history then grew unchecked. The
internal summary call now omits the step's temperature: a structured
extraction does not need a set one, and omitting it sidesteps each
provider's temperature-versus-reasoning rules. Confirmed against the API
that low + temperature is refused on those models.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): compact an oversized loaded memory before the first request

Compaction was reactive, taken only after a request the endpoint
accepted, so the fallbacks the loop learns from a rejection are known
first. But a memory loaded from an earlier run can already exceed this
run's window — the step was switched to a smaller model, or a run under
a wider one persisted more than fits — and that first request then
overflows and fails the run, with every retry reloading the same
history and failing again. A pass is now taken up front, off the
character estimate, before the first request. It uses the default
request shape; an endpoint needing a fallback may reject this one
summary, which is non-fatal, and mainstream providers need none.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): under the storage bound, trigger on the max of bytes and model tokens

The storage-bound pass measured only the serialized row size, so an
attachment — a few bytes as an S3 path but nearly the whole model
context — read as tiny and the pass skipped a compaction the model
needed. It now takes the larger of the byte measure and the model's
token count, since repetitive text is few tokens but many bytes and an
attachment is the reverse; either being over must fire a pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): drop oldest turns when a summary cannot fit the window, as the AI session does

An oversized loaded memory (a step switched to a smaller model, or an
object-store run that persisted more than a later model's window holds)
left a prefix larger than the summarizer's own window, so the summary
request overflowed and failed, the memory was untouched, and every
retry failed the same way. The AI session handles this by falling back
from summarization to dropping the oldest turns down to the target;
compaction here now does the same. When a summary cannot run — it
failed, the breaker is tripped, or nothing is worth folding — the oldest
turns are dropped until the conversation fits and opens on a user
message, keeping the newest turn. The next request then always fits.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): drop whole turns only, keep the storage pass to bytes, refresh the count after a rewrite

Three edges the seventh Codex round found, all in the drop-oldest
fallback and the storage-bound measure:

- drop_oldest_to_fit dropped to any point that freed enough, which could
  strand a tool result whose tool_calls went with the messages before
  it. It now drops whole turns only, always landing the boundary on a
  user message and never splitting the newest turn; a lone turn too big
  for the window is left whole rather than broken.
- The storage-bound pass measured the whole model prompt against the
  shrunk 25k window, so a large tool roster and the system prompt —
  neither written to the row — tripped it on a conversation the row
  easily held. It measures the serialized bytes alone now; the model's
  own window is enforced by the in-loop passes and the pre-first-request
  pass, so the persisted size is all this pass is for.
- A compaction rewrites the message list, so the provider's count for
  the request that produced it no longer lines up. The count is now
  cleared after any pass that rewrites the conversation, so a later pass
  measures the estimate over the actual messages instead of a stale,
  larger prompt (which could decline a summary that already fit and then
  drop it). The step temperature, no longer sent to the summarizer on
  any path, is dropped from the request struct.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): measure only the persisted messages against the storage cap

Persistence strips the system prompt before writing the memory row, but
the storage pass was serializing every message including it, so a large
system prompt with a tiny conversation reported far over the storage
trigger, and the fallback dropped the one real turn, run after run. The
storage measure now serializes only the non-system messages, matching
what the row actually holds.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix(ai-agent): run the model-window pass before the storage-bytes pass post-loop

A turn the model answered without a tool call broke before the in-loop
compaction check, so on database-backed memory its only pass was the
storage one, which measures bytes. An attachment fills the model context
but is a few bytes in the row, so that turn never compacted and a
follow-up could overflow the model. The post-loop now runs a
model-window pass first, off the provider's count, then the
storage-bytes pass when the row is smaller than the model — both limits
enforced for a chat-shaped step, not just the one that happens to bind.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix: simplify agent compaction and preserve execution history

* fix: remove unused compaction history setting

* fix: preserve answers and recover rejected agent context

* refactor: make agent compaction transactional

* fix: skip agent summaries that cannot fit retained context

* fix: explain skipped agent context compaction

* fix: retain recent agent memory when storage compaction cannot fit

* fix: start retained agent memory at a user turn

* fix: reject unsafe agent memory truncation on storage fallback

* docs: clarify agent context window override scope

* fix: keep recent turns verbatim when compaction memory outgrows storage

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix: keep the compaction summary out of the agent's answers

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* fix: shorten the agent context window help text

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViJyjUmidDYV2m6ifQdLeH

* chore: update ee-repo-ref to 942d4013f36edac1fc9a9addbdb02198db1c7a05

This commit updates the EE repository reference after PR #812 was merged in windmill-ee-private.

Previous ee-repo-ref: 8ca1682ce6106ba6ea96894fbe606dac64102eb6

New ee-repo-ref: 942d4013f36edac1fc9a9addbdb02198db1c7a05

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
This commit is contained in:
hugocasa
2026-09-29 19:01:53 +02:00
committed by GitHub
co-authored by Claude Opus 4.8 windmill-internal-app[bot] Ruben Fiszel
parent cf5c49c3dc
commit c1b59f70dd
32 changed files with 2394 additions and 395 deletions
+6 -3
View File
@@ -303,7 +303,7 @@ pub struct GeminiUsageMetadata {
/// 60 tool-use + 17 candidates + 52 thoughts against a `totalTokenCount` of 146, so
/// leaving them out under-reports the input of every tool-using turn. Cached tokens
/// are not added here, being already part of `promptTokenCount`.
fn gemini_prompt_tokens(usage: &GeminiUsageMetadata) -> i32 {
pub(crate) fn gemini_prompt_tokens(usage: &GeminiUsageMetadata) -> i32 {
usage
.prompt_token_count
.unwrap_or(0)
@@ -313,7 +313,7 @@ fn gemini_prompt_tokens(usage: &GeminiUsageMetadata) -> i32 {
/// Output tokens as billed: Gemini counts thinking apart from `candidatesTokenCount`
/// but charges it at the output rate, so a reply that thought would otherwise be
/// reported as far cheaper than it was.
fn gemini_completion_tokens(usage: &GeminiUsageMetadata) -> i32 {
pub(crate) fn gemini_completion_tokens(usage: &GeminiUsageMetadata) -> i32 {
usage
.candidates_token_count
.unwrap_or(0)
@@ -1041,7 +1041,10 @@ mod tests {
// through unchanged and the cached share is reported alongside it; thinking
// is billed as output but counted apart from the candidates.
assert_eq!(usage_chunk["usage"]["prompt_tokens"], 1000);
assert_eq!(usage_chunk["usage"]["prompt_tokens_details"]["cached_tokens"], 900);
assert_eq!(
usage_chunk["usage"]["prompt_tokens_details"]["cached_tokens"],
900
);
assert_eq!(usage_chunk["usage"]["completion_tokens"], 120);
}
+1
View File
@@ -6,6 +6,7 @@ pub mod ai_providers;
pub mod ai_types;
pub mod credentials;
pub mod image_handler;
pub mod model_context;
pub mod providers;
pub mod proxy;
pub mod query_builder;
+165
View File
@@ -0,0 +1,165 @@
//! What a model's context window holds, for the code that has to keep a conversation
//! inside it.
/// Assumed window for a model that is not in the table. Conservative on purpose: a
/// guess that is too small only compacts earlier, which is recoverable, while one that
/// is too large overflows and the provider raises the context error itself.
pub const DEFAULT_CONTEXT_WINDOW: usize = 128_000;
/// Context windows of the models we know, most specific entry first — the first name
/// found in the bare model id wins, so vendor-namespaced and date-suffixed ids
/// (`anthropic.claude-sonnet-4-6-…-v1:0`, `gpt-5.2-2026-01-01`) still resolve.
/// Conservative family fallbacks sit below the explicit entries.
///
/// **Mirrors `MODEL_CONTEXT_WINDOWS` in
/// `frontend/src/lib/components/copilot/modelConfig.ts`**, which the AI session's own
/// compaction reads. Add a model to one and it must go in the other, or the same model
/// compacts at one size in a chat and another in an agent step.
const MODEL_CONTEXT_WINDOWS: &[(&str, usize)] = &[
// Anthropic — Sonnet/Opus 4.6+ ship a 1M window at standard pricing (GA);
// Haiku, older Claude models (3.x, 4.0, 4.1, 4.5) and date-suffixed Claude 4
// base ids (claude-sonnet-4-20250514) fall through to 200K
("claude-fable-5", 1_000_000),
("claude-mythos-5", 1_000_000),
("claude-opus-5", 1_000_000),
("claude-sonnet-5", 1_000_000),
("claude-opus-4-8", 1_000_000),
("claude-opus-4-7", 1_000_000),
("claude-opus-4-6", 1_000_000),
("claude-sonnet-4-6", 1_000_000),
("claude", 200_000),
// OpenAI — gpt-5 covers the base family (-mini / -nano) and the 5.1/5.2
// revisions, all 400K; 5.4/5.5 moved to 1M and 5.6 to 1.05M
("gpt-5-6", 1_050_000),
("gpt-5-5", 1_000_000),
("gpt-5-4", 1_000_000),
("gpt-5", 400_000),
("gpt-4-1", 1_000_000),
("gpt-4o", 128_000),
("o4-mini", 200_000),
("o3", 200_000),
// Google — the 2.5 / 3 / 3.1 Gemini families are all 1M
("gemini-3-1", 1_000_000),
("gemini-3", 1_000_000),
("gemini-2-5", 1_000_000),
// DeepSeek — the V4 family (pro / flash) is 1M. The deepseek-chat /
// deepseek-reasoner aliases were retired 2026-07-24 but can still sit in a
// saved selection, so they keep resolving to the window they had.
("deepseek-v4", 1_000_000),
("deepseek-chat", 1_000_000),
("deepseek-reasoner", 1_000_000),
("deepseek", 128_000),
// Alibaba — Qwen3-Max is 256K. No qwen family fallback: variant windows range
// from 8K (character models) to 1M, too wide for even a conservative guess
("qwen3-max", 256_000),
// Others — Mistral Medium 3.5 is 256K, reachable under both its version and
// the `-latest` alias. There is deliberately no `mistral-medium` family row:
// pinned older snapshots are 128K, and over-claiming a window overflows it.
("mistral-medium-3-5", 256_000),
("mistral-medium-latest", 256_000),
("llama", 128_000),
("codestral", 32_000),
];
/// The bare model id an entry is matched against: lowercased, the leading `~` and any
/// vendor path segment dropped, the `:`-suffixed route removed, and `.` collapsed to
/// `-`. Version separators differ by route to the same model — Anthropic writes
/// `claude-opus-4-8`, OpenRouter `anthropic/claude-opus-4.8` — so collapsing both sides
/// keeps one entry covering every route.
fn bare_model_id(model: &str) -> String {
let normalized = model.trim().to_lowercase();
let normalized = normalized.strip_prefix('~').unwrap_or(&normalized);
let last = normalized.rsplit('/').next().unwrap_or(normalized);
let base = match last.find(':') {
Some(0) | None => last,
Some(colon) => &last[..colon],
};
base.replace('.', "-")
}
/// Whether this is one of OpenAI's reasoning models (gpt-5 and later, the o-series),
/// which think by default. Mirrors `requiresMaxCompletionTokens` in the frontend's
/// `modelConfig.ts`: the o-series match wants a digit after the `o` so that ids such
/// as `open-mistral-*` stay out.
pub fn is_openai_reasoning_model(model: &str) -> bool {
let id = bare_model_id(model);
id.starts_with("gpt-5")
|| (id.starts_with('o') && id[1..].starts_with(|c: char| c.is_ascii_digit()))
}
/// The window this model is known to hold, `None` for one not in the table.
fn known_model_context_window(model: &str) -> Option<usize> {
let id = bare_model_id(model);
MODEL_CONTEXT_WINDOWS
.iter()
.find(|(name, _)| matches_entry(&id, name))
.map(|(_, window)| *window)
}
/// The window to plan against: the model's where it is known, the conservative
/// assumption otherwise.
pub fn model_context_window(model: &str) -> usize {
known_model_context_window(model).unwrap_or(DEFAULT_CONTEXT_WINDOW)
}
/// An entry matches anywhere in the id, so a vendor prefix or a date suffix does not
/// hide it. One ending on a version digit must not run into a longer version:
/// `gpt-4-1` would otherwise claim `gpt-4-1106-preview`. A family fallback ending on a
/// letter gets no such guard — a version welded straight onto the name (`llama3-1`) is
/// what it exists to catch.
fn matches_entry(id: &str, name: &str) -> bool {
let guard_digits = name.ends_with(|c: char| c.is_ascii_digit());
id.match_indices(name).any(|(at, _)| {
!guard_digits
|| !id[at + name.len()..]
.chars()
.next()
.is_some_and(|c| c.is_ascii_digit())
})
}
#[cfg(test)]
mod tests {
use super::*;
/// One entry has to cover every route to the same model: the vendor path and dot
/// versions OpenRouter uses, Bedrock's prefixed and `-v1:0`-suffixed ids, and the
/// date suffixes the providers' own ids carry.
#[test]
fn every_route_to_a_model_resolves_to_one_window() {
for id in [
"claude-opus-4-8",
"anthropic/claude-opus-4.8",
"~anthropic/claude-opus-4.8",
"anthropic.claude-opus-4-8-20260101-v1:0",
] {
assert_eq!(known_model_context_window(id), Some(1_000_000), "{id}");
}
// Falls through the explicit rows to the family fallback.
assert_eq!(
known_model_context_window("claude-3-5-haiku"),
Some(200_000)
);
}
/// The guard that keeps an entry ending on a version digit from claiming a longer
/// version of the same family.
#[test]
fn a_version_entry_does_not_claim_a_longer_version() {
assert_eq!(known_model_context_window("gpt-4.1"), Some(1_000_000));
assert_eq!(known_model_context_window("gpt-4-1106-preview"), None);
assert_eq!(known_model_context_window("gpt-5-mini"), Some(400_000));
assert_eq!(known_model_context_window("gpt-5.6"), Some(1_050_000));
}
/// An unknown model still gets a number, so nothing downstream has to carry a
/// "no limit" case that would let a conversation grow unbounded.
#[test]
fn an_unknown_model_falls_back_to_the_assumed_window() {
assert_eq!(known_model_context_window("some-local-llm"), None);
assert_eq!(
model_context_window("some-local-llm"),
DEFAULT_CONTEXT_WINDOW
);
}
}
@@ -782,7 +782,7 @@ impl QueryBuilder for AnthropicQueryBuilder {
// Convert Anthropic usage to TokenUsage
let usage = anthropic_usage.map(|u| {
TokenUsage::from_input_output(u.input_tokens, u.output_tokens)
.with_cache(u.cache_read_input_tokens, u.cache_creation_input_tokens)
.with_cache_beside_input(u.cache_read_input_tokens, u.cache_creation_input_tokens)
});
Ok(ParsedResponse::Text {
+1 -1
View File
@@ -1148,7 +1148,7 @@ impl BedrockQueryBuilder {
Some(token_usage.output_tokens()),
Some(token_usage.total_tokens()),
)
.with_cache(
.with_cache_beside_input(
token_usage
.cache_read_input_tokens()
.map(|v| i32::try_from(v).unwrap_or(i32::MAX)),
+50 -11
View File
@@ -1,10 +1,11 @@
use crate::{
ai_google::{
gemini_event_to_openai_sse_chunks, gemini_response_to_openai, openai_messages_to_gemini,
openai_tools_to_gemini, parse_gemini_response, parse_gemini_sse_event,
sanitize_schema_for_google, GeminiFunctionDeclaration, GeminiGenerationConfig,
GeminiImageContent, GeminiImageRequest, GeminiImageResponse, GeminiInlineData, GeminiPart,
GeminiPredictContent, GeminiTextRequest, GeminiThinkingConfig, GeminiTool,
gemini_completion_tokens, gemini_event_to_openai_sse_chunks, gemini_prompt_tokens,
gemini_response_to_openai, openai_messages_to_gemini, openai_tools_to_gemini,
parse_gemini_response, parse_gemini_sse_event, sanitize_schema_for_google,
GeminiFunctionDeclaration, GeminiGenerationConfig, GeminiImageContent, GeminiImageRequest,
GeminiImageResponse, GeminiInlineData, GeminiPart, GeminiPredictContent, GeminiTextRequest,
GeminiThinkingConfig, GeminiTool,
},
image_handler::{download_and_encode_s3_image, prepare_messages_for_api},
proxy::{ProxyBuildArgs, ProxyRequest},
@@ -685,12 +686,32 @@ impl QueryBuilder for GoogleAIQueryBuilder {
stream_event_processor.send(event, &mut events_str).await?;
}
let usage = gemini_usage.map(|u| {
TokenUsage::new(
u.prompt_token_count,
u.candidates_token_count,
u.total_token_count,
)
// Through the shared helpers, which fold in the two counts Gemini reports beside
// the headline ones: tool-use prompt tokens, disjoint from `promptTokenCount` and
// present on every tool-using turn, and thinking tokens. Reading the headline
// fields alone under-reports the prompt of exactly the conversations compaction
// has to notice.
let usage = gemini_usage.and_then(|u| {
let has_prompt =
u.prompt_token_count.is_some() || u.tool_use_prompt_token_count.is_some();
let has_completion =
u.candidates_token_count.is_some() || u.thoughts_token_count.is_some();
if !has_prompt && !has_completion && u.total_token_count.is_none() {
// An endpoint that reported nothing has to stay `None`. A zeroed count
// reads downstream as an empty prompt, where absent means "fall back to
// estimating the conversation".
return None;
}
let prompt = has_prompt.then(|| gemini_prompt_tokens(&u));
let completion = has_completion.then(|| gemini_completion_tokens(&u));
Some(TokenUsage::new(
prompt,
completion,
u.total_token_count.or_else(|| match (prompt, completion) {
(Some(p), Some(c)) => Some(p.saturating_add(c)),
_ => None,
}),
))
});
Ok(ParsedResponse::Text {
@@ -1014,4 +1035,22 @@ mod tests {
.headers
.contains(&("x-goog-api-key".to_string(), "api-key".to_string())));
}
/// Gemini reports a tool-using turn's input across two disjoint fields. Reading only
/// the headline one hides the tool results from the agent step's usage — and from the
/// compaction trigger, which is what would notice them.
#[test]
fn gemini_agent_usage_counts_tool_use_and_thinking_tokens() {
let usage = crate::ai_google::GeminiUsageMetadata {
prompt_token_count: Some(17),
candidates_token_count: Some(17),
total_token_count: Some(146),
tool_use_prompt_token_count: Some(60),
thoughts_token_count: Some(52),
..Default::default()
};
assert_eq!(gemini_prompt_tokens(&usage), 77);
assert_eq!(gemini_completion_tokens(&usage), 69);
}
}
+43
View File
@@ -6,6 +6,49 @@ use crate::ai_types::OpenAIToolCall;
use crate::proxy::{ProxyBuildArgs, ProxyRequest};
use crate::types::*;
/// Only explicit input-context rejections warrant dropping conversation history.
/// Rate limits, output-token limits and generic validation errors must propagate.
pub fn is_context_length_error(message: &str) -> bool {
let message = message.to_ascii_lowercase();
message.contains("context_length_exceeded")
|| message.contains("maximum context length")
|| message.contains("prompt is too long")
|| message.contains("exceed context limit")
|| message.contains("input is too long for requested model")
|| (message.contains("input token count") && message.contains("exceeds the maximum"))
|| message.contains("too many input tokens")
}
#[cfg(test)]
mod context_error_tests {
use super::is_context_length_error;
#[test]
fn recognizes_context_rejections_without_retrying_other_provider_errors() {
for message in [
r#"{"error":{"code":"context_length_exceeded"}}"#,
"This model's maximum context length is 128000 tokens",
"prompt is too long: 210000 tokens > 200000 maximum",
"input length and `max_tokens` exceed context limit: 199000 + 8192 > 200000",
"The input token count (10000) exceeds the maximum number of tokens allowed (8192)",
"ValidationException: Input is too long for requested model.",
"ValidationException: Too many input tokens",
] {
assert!(is_context_length_error(message), "{message}");
}
for message in [
"Rate limit exceeded: tokens per minute",
"max_tokens exceeds the maximum output tokens",
"Invalid tool schema",
"Additional properties are not allowed: stream_options",
"Request body too large",
"Internal server error",
] {
assert!(!is_context_length_error(message), "{message}");
}
}
}
/// Arguments for building an AI request
pub struct BuildRequestArgs<'a> {
pub messages: &'a [OpenAIMessage],
+84 -8
View File
@@ -304,7 +304,16 @@ pub enum AnthropicDelta {
Unknown,
}
/// Anthropic usage information from message_delta event
/// The `message` envelope of a `message_start` event. Only its usage is read: the
/// prompt-side counts appear here and nowhere else in the stream.
#[derive(Deserialize, Debug)]
pub struct AnthropicStreamMessage {
#[serde(default)]
pub usage: Option<AnthropicUsage>,
}
/// Anthropic usage information, reported across `message_start` (prompt side) and
/// `message_delta` (completion side)
#[derive(Deserialize, Debug, Clone)]
pub struct AnthropicUsage {
#[serde(default)]
@@ -322,7 +331,10 @@ pub struct AnthropicUsage {
#[serde(tag = "type")]
pub enum AnthropicSSEEvent {
#[serde(rename = "message_start")]
MessageStart {},
MessageStart {
#[serde(default)]
message: Option<AnthropicStreamMessage>,
},
#[serde(rename = "content_block_start")]
ContentBlockStart { index: usize, content_block: AnthropicContentBlockStart },
#[serde(rename = "content_block_delta")]
@@ -371,7 +383,8 @@ pub struct AnthropicSSEParser {
pub annotations: Vec<UrlCitation>,
/// Whether web search was used in this response
pub used_websearch: bool,
/// Token usage from message_delta event
/// Token usage, merged from the `message_start` (prompt side) and `message_delta`
/// (completion side) events
pub usage: Option<AnthropicUsage>,
/// Claude thinking block accumulated from `thinking`/`signature` deltas
/// (or a redacted block). Attached to the first tool call of the turn so it
@@ -574,14 +587,35 @@ impl SSEParser for AnthropicSSEParser {
let error_msg = message.unwrap_or_else(|| "Unknown error".to_string());
tracing::error!("Anthropic streaming error: {}", error_msg);
}
AnthropicSSEEvent::MessageStart { message } => {
// The only event carrying the prompt-side counts. `message_delta`
// reports the completion, so dropping this one leaves the request
// with no input token count at all.
if let Some(usage) = message.and_then(|message| message.usage) {
self.usage = Some(usage);
}
}
AnthropicSSEEvent::MessageDelta { usage } => {
if let Some(u) = usage {
self.usage = Some(u);
match &mut self.usage {
// Field by field, so the prompt counts from `message_start`
// survive a delta that only reports the completion.
Some(existing) => {
existing.input_tokens = u.input_tokens.or(existing.input_tokens);
existing.output_tokens = u.output_tokens.or(existing.output_tokens);
existing.cache_read_input_tokens = u
.cache_read_input_tokens
.or(existing.cache_read_input_tokens);
existing.cache_creation_input_tokens = u
.cache_creation_input_tokens
.or(existing.cache_creation_input_tokens);
}
None => self.usage = Some(u),
}
}
}
// Ignore other events
AnthropicSSEEvent::MessageStart {}
| AnthropicSSEEvent::MessageStop {}
AnthropicSSEEvent::MessageStop {}
| AnthropicSSEEvent::Ping {}
| AnthropicSSEEvent::Unknown => {}
}
@@ -794,6 +828,11 @@ pub enum OpenAIResponsesSSEEvent {
#[serde(rename = "response.completed")]
Completed { response: OpenAIResponsesResponse },
/// Response ended by `max_output_tokens` or a filter: the same object, with the
/// tokens spent so far, so it is billed the same
#[serde(rename = "response.incomplete")]
Incomplete { response: OpenAIResponsesResponse },
/// Response created
#[serde(rename = "response.created")]
Created {},
@@ -969,8 +1008,8 @@ impl SSEParser for OpenAIResponsesSSEParser {
});
}
OpenAIResponsesSSEEvent::Completed { response } => {
// Extract usage from response.completed event
OpenAIResponsesSSEEvent::Completed { response }
| OpenAIResponsesSSEEvent::Incomplete { response } => {
if let Some(usage) = response.usage {
self.usage = Some(usage);
}
@@ -1041,6 +1080,43 @@ mod tests {
assert_eq!(token_usage.total_tokens, Some(4820));
}
struct NoopSink;
#[async_trait::async_trait]
impl StreamEventSink for NoopSink {
async fn send(
&self,
_event: StreamingEvent,
_events_str: &mut String,
) -> Result<(), Error> {
Ok(())
}
}
/// The prompt-side counts arrive only on `message_start` and the completion total
/// only on `message_delta`; a parser that keeps just the last one reports a request
/// with no input tokens at all.
#[tokio::test]
async fn anthropic_usage_merges_message_start_and_message_delta() {
let mut parser = AnthropicSSEParser::new(Box::new(NoopSink));
parser
.parse_event_data(
r#"{"type":"message_start","message":{"id":"msg_1","role":"assistant","usage":{"input_tokens":4821,"output_tokens":1,"cache_read_input_tokens":4096,"cache_creation_input_tokens":128}}}"#,
)
.await
.unwrap();
parser
.parse_event_data(r#"{"type":"message_delta","usage":{"output_tokens":312}}"#)
.await
.unwrap();
let usage = parser.usage.expect("usage");
assert_eq!(usage.input_tokens, Some(4821));
assert_eq!(usage.output_tokens, Some(312));
assert_eq!(usage.cache_read_input_tokens, Some(4096));
assert_eq!(usage.cache_creation_input_tokens, Some(128));
}
#[test]
fn openai_responses_usage_maps_cached_input_tokens() {
// Payload shape returned by the OpenAI Responses API.
+44 -1
View File
@@ -82,6 +82,14 @@ pub enum Memory {
#[serde(default, deserialize_with = "deserialize_null_as_zero")]
context_length: usize,
},
Compaction {
/// Overrides the window looked up from the model. Only a step the lookup cannot
/// serve — a Custom AI deployment, a model id not in the table — needs one, and
/// a value larger than the model actually serves lets the conversation overflow
/// before the summary is ever taken.
#[serde(default, deserialize_with = "deserialize_null_as_zero")]
context_window: usize,
},
/// Written before `window`. Its `memory_id` stays a fallback behind the run's memory id.
Auto {
#[serde(default, deserialize_with = "deserialize_null_as_zero")]
@@ -388,13 +396,28 @@ impl TokenUsage {
Self::new(input, output, total)
}
/// Add cache token information
/// Records a cached prefix the provider counted *inside* `input_tokens`, which the
/// OpenAI-shaped ones do. The counts are kept only for the cost split.
pub fn with_cache(mut self, read: Option<i32>, write: Option<i32>) -> Self {
self.cache_read_input_tokens = read;
self.cache_write_input_tokens = write;
self
}
/// Records a cached prefix the provider reported *beside* `input_tokens` rather than
/// inside it, which Anthropic and Bedrock do. `input_tokens` is raised to the whole
/// prompt, so `input_tokens` means the same thing whatever served the request; the
/// cache counts stay as the subsets they have become, and the total follows.
pub fn with_cache_beside_input(mut self, read: Option<i32>, write: Option<i32>) -> Self {
let beside = read.unwrap_or(0).saturating_add(write.unwrap_or(0));
self.input_tokens = self.input_tokens.map(|input| input.saturating_add(beside));
self.total_tokens = match (self.input_tokens, self.output_tokens) {
(Some(input), Some(output)) => Some(input.saturating_add(output)),
_ => self.total_tokens.map(|total| total.saturating_add(beside)),
};
self.with_cache(read, write)
}
pub fn is_empty(&self) -> bool {
self.input_tokens.is_none()
&& self.output_tokens.is_none()
@@ -939,6 +962,26 @@ mod tests {
use super::*;
use std::collections::HashMap;
/// Whichever provider served the request, `input_tokens` has to end up meaning the
/// whole prompt. Leaving the Anthropic shape as reported under-states it by the
/// entire cached prefix, which is exactly what compaction has to notice.
#[test]
fn both_provider_cache_shapes_report_the_whole_prompt() {
// OpenAI-shaped: cached_tokens is already inside input_tokens.
let openai = TokenUsage::new(Some(1000), Some(10), Some(1010)).with_cache(Some(800), None);
assert_eq!(openai.input_tokens, Some(1000));
assert_eq!(openai.total_tokens, Some(1010));
// Anthropic-shaped: the cached prefix is reported beside input_tokens, and the
// total follows the prompt it is folded into.
let anthropic = TokenUsage::from_input_output(Some(200), Some(10))
.with_cache_beside_input(Some(5000), Some(300));
assert_eq!(anthropic.input_tokens, Some(5500));
assert_eq!(anthropic.total_tokens, Some(5510));
// Kept as the subsets they have become, for the cost split.
assert_eq!(anthropic.cache_read_input_tokens, Some(5000));
}
/// Helper to create a simple string type schema
fn string_schema() -> OpenAPISchema {
OpenAPISchema {
File diff suppressed because it is too large Load Diff
+1
View File
@@ -1,6 +1,7 @@
// AI executor module structure
// This module will contain all AI-related execution logic
pub mod compaction;
pub mod stream_event_processor;
pub mod tools;
pub mod utils;
File diff suppressed because it is too large Load Diff
+36 -4
View File
@@ -30,16 +30,21 @@ pub async fn read_from_db(
}
}
/// Write AI agent memory to database with size checking and truncation
/// Write AI agent memory to database with size checking and truncation.
///
/// Returns how many of the oldest messages were dropped to fit, so the step can say so
/// on the run: from the flow's side a truncation is invisible, the agent simply having
/// forgotten the start of its conversation by the next one.
pub async fn write_to_db(
db: &DB,
workspace_id: &str,
conversation_id: Uuid,
step_id: &str,
messages: &[OpenAIMessage],
) -> Result<(), Error> {
allow_truncation: bool,
) -> Result<usize, Error> {
if messages.is_empty() {
return Ok(());
return Ok(0);
}
// Serialize messages and check size
@@ -49,6 +54,13 @@ pub async fn write_to_db(
// Truncate if necessary
if size_bytes > MAX_MEMORY_SIZE_BYTES {
// Compaction owns its exchange boundaries, including after an object-store failure.
if !allow_truncation {
return Err(Error::ExecutionErr(format!(
"Memory exceeds database capacity ({} > {} bytes); existing memory was not updated",
size_bytes, MAX_MEMORY_SIZE_BYTES
)));
}
tracing::warn!(
"Memory size ({} bytes) exceeds limit ({} bytes) for workspace={} conversation={} step={}. Truncating messages. Use S3 storage in workspace settings to store full conversation history.",
size_bytes,
@@ -78,7 +90,7 @@ pub async fn write_to_db(
.execute(db)
.await?;
Ok(())
Ok(messages.len() - messages_to_store.len())
}
/// Delete all memory for a conversation from database
@@ -121,3 +133,23 @@ fn truncate_messages(
Ok(result)
}
#[cfg(test)]
mod tests {
use super::*;
#[tokio::test]
async fn oversized_compaction_memory_is_rejected_before_touching_the_database() {
let db = sqlx::postgres::PgPoolOptions::new()
.connect_lazy_with(sqlx::postgres::PgConnectOptions::new());
db.close().await;
let message: OpenAIMessage = serde_json::from_value(serde_json::json!({
"role": "user", "content": "x".repeat(MAX_MEMORY_SIZE_BYTES)
}))
.unwrap();
let error = write_to_db(&db, "test", Uuid::nil(), "a", &[message], false)
.await
.unwrap_err();
assert!(error.to_string().contains("Memory exceeds database capacity"));
}
}
+20 -5
View File
@@ -30,14 +30,29 @@ pub async fn write_to_memory(
conversation_id: Uuid,
step_id: &str,
messages: &[OpenAIMessage],
) -> anyhow::Result<()> {
allow_truncation: bool,
) -> anyhow::Result<usize> {
if messages.is_empty() {
return Ok(());
return Ok(0);
}
memory_common::write_to_db(db, workspace_id, conversation_id, step_id, messages)
.await
.map_err(|e| anyhow::anyhow!("Database write failed: {e:?}"))
memory_common::write_to_db(
db,
workspace_id,
conversation_id,
step_id,
messages,
allow_truncation,
)
.await
.map_err(|e| anyhow::anyhow!("Database write failed: {e:?}"))
}
/// How many bytes a persisted memory can hold, `None` where nothing bounds it.
/// In OSS: always the database limit
#[cfg(not(all(feature = "private", feature = "enterprise")))]
pub async fn memory_storage_capacity_bytes() -> Option<usize> {
Some(memory_common::MAX_MEMORY_SIZE_BYTES)
}
/// Delete all memory for a conversation from storage
+1 -1
View File
File diff suppressed because one or more lines are too long
+78 -11
View File
@@ -44,8 +44,10 @@ which memory it is:
- **Agent: managed memory.** `memory` is a brain key, so it moves with a saved agent.
`{ kind: window, context_length }` has Windmill store the conversation and replay its last N
messages; `{ kind: off }` keeps none. An absent `memory` means off, the default: the editor turns
it on when chat input is enabled. `auto` and `manual` are the older spellings and are still read.
messages; `{ kind: compaction }` replays all of it and summarizes the older part as it fills the
model's context window (see below); `{ kind: off }` keeps none. An absent `memory` means off, the
default: the editor turns it on when chat input is enabled, as compaction, since a chat has no
end. `auto` and `manual` are the older spellings and are still read.
- **Run: memory id.** `flow_status.memory_id`, set when the run is queued: the chat conversation
id, an app chat session id, or the `memory_id` run parameter. Any string is accepted, and one
that is not a uuid is hashed to a v5 uuid scoped to the workspace and the flow the run started
@@ -69,25 +71,90 @@ The worker reconciles them once per agent invocation, nested agent tools include
1. A legacy `auto` or `manual` memory: read as the editor that wrote it ran it. `manual` replays
its list; `auto` uses the run's memory id, else the id baked into it, else runs stateless.
Neither history input is read. An `auto` without a count, or with 0, is off and read as such.
2. Managed memory: the memory id is the step's, else the run's. With no memory id the agent runs
stateless, and a step `previous_messages` is ignored.
2. Managed memory, `window` or `compaction`: the memory id is the step's, else the run's. With no
memory id the agent runs stateless, and a step `previous_messages` is ignored. Compaction still
bounds a stateless run's own loop, which is where a long tool sequence overflows.
3. Memory off: the history is `previous_messages`, else nothing. Memory is neither read nor
written, and a step `memory_id` is ignored.
Each ignored input and each stateless fallback is written to the job log.
### Compaction
`{ kind: compaction }` keeps the whole conversation and lets a summary, rather than a message count,
decide what leaves the prompt.
The window it plans against comes from the model, through `MODEL_CONTEXT_WINDOWS` in
`windmill-ai/src/model_context.rs`, falling back to 128000 for an id the table does not list. That
table mirrors the one the AI session's own compaction reads
(`frontend/src/lib/components/copilot/modelConfig.ts`) and the two have to be updated together. The
step's `context_window` overrides it, for a Custom AI deployment or a model the table cannot name;
setting it too large never trips the trigger and the provider raises the context error itself.
Workspace AI chat `context_window_per_model` overrides are separate and are not inherited by
agent steps; configure a custom deployment's size on the step itself.
`windmill-worker/src/ai/compaction.rs` keeps two histories: model context, which can be compacted,
and the execution record, which retains the loaded history and every message produced by the run.
Returned results and max-iteration partial results use the execution record. Compaction cannot remove or
reorder the action messages the flow viewer indexes, or the MCP results it finds by call ID.
Before each provider request, including the first, the worker checks the projected context size.
It reserves the larger of 20% of the model window and the configured maximum output tokens.
At the remaining input budget it summarizes an older prefix, retaining up to 20K estimated tokens
of recent complete exchanges (at most half the input budget on smaller models).
The newest exchange is always retained. A user prompt stays with its
first response; later tool rounds can be compacted within a single turn, but a call and its results
are never split. A prefix containing only an earlier summary is not summarized again.
The summarization request carries no tools; its tool exchanges are rendered as text because
Bedrock rejects tool blocks without definitions. A dedicated system instruction asks for a factual
handoff from a labelled transcript, keeping the compaction instruction outside that transcript;
media parts remain available. Its output cap matches the reserved summary budget,
and its temperature and reasoning settings are independent of the step's answer settings.
A replacement is built separately and installed only if it reduces projected context and fits
the estimated input budget. An empty, oversized or failed summary leaves the context and its
usage measurement untouched; compaction never evicts messages. Three consecutive failed attempts
disable summary requests for the run. Estimates schedule compaction but do not reject model calls.
The projection uses normalized `TokenUsage::input_tokens` from the last request plus a `bytes/4`
estimate of appended messages. Without usage, or after rewriting context, it estimates the whole
prompt including tools. S3 descriptors get a nominal attachment allowance; actual attachment costs
and tokenizer differences remain approximate. Provider parsing owns usage normalization: Anthropic
and Bedrock report cached input separately, while OpenAI-shaped providers include it in input tokens.
If a provider explicitly rejects the context size, the worker summarizes all older complete
exchanges and retries that request once if a usable replacement was produced. This also recovers
from undercounted attachments loaded from a previous run, without provider-specific tokenizers.
Unrelated errors are not retried this way. The newest exchange is still retained, so a request
that cannot fit even after recovery fails; estimates do not guarantee every first request fits.
If recovery fails or the retry is rejected, the provider error is returned without saving changed
memory. A successful checkpoint uses the same execution record as before compaction.
Memory is stored per (memory id, step id), in `ai_agent_memory` or S3 at
`memory/{workspace}/{memory id}/{step}.json`. The chat transcript (`flow_conversation_message`)
always follows the run's id, even when a step sets its own. Nothing expires stored memory: deleting
a chat conversation deletes its memory, and a memory named by a string id stays until it is
overwritten.
always follows the run's id, even when a step sets its own. Persistence independently checks
serialized memory against the database's 100KB limit (`MAX_MEMORY_SIZE_BYTES`),
when `memory_storage_capacity_bytes` reports one. System messages and tool definitions are not
stored, so they do not count towards this byte limit. If oversized, it requests one checkpoint
of all older exchanges, then measures bytes again. If memory still cannot fit, persistence saves
the largest suffix of complete exchanges that fits and starts with a user message, dropping
older exchanges and logging that loss. This selection changes neither model context nor the
returned execution record.
If no user-starting suffix fits, the write is skipped and the flow log explains that the
next run will load the previous saved memory. The completed answer succeeds in either case.
If object storage fails and falls back to the database, an oversized compaction write is
rejected without changing existing memory; the flow log reports the failed save.
There is no model-window compaction after the final answer unless storage needs
it. Persistence reads model context independently of the execution record returned by the step.
Nothing expires stored memory: deleting a chat conversation deletes its memory, and a memory named
by a string id stays until it is overwritten.
Compatibility runs one way. New workers read every older shape. The editor rewrites a legacy step
only when the author changes it, so a flow nobody edits keeps running on older workers, while a
step saved with `window` or a history input needs a worker that knows them. An id an older editor
baked into `memory` stays a fallback behind the run's id until the author chooses *Keep as memory
id* or *Use the run's memory id*. In a chat flow it is dropped on save, since the conversation id
always took precedence there.
step saved with `window`, `compaction` or a history input needs a worker that knows them. An id an
older editor baked into `memory` stays a fallback behind the run's id until the author chooses
*Keep as memory id* or *Use the run's memory id*. In a chat flow it is dropped on save, since the
conversation id always took precedence there.
## Drafts
@@ -59,7 +59,8 @@ import { getEffectiveModelContextWindow } from '../modelConfig'
import {
getCompactionSummaryPrompt,
formatCompactSummary,
buildSummaryMessageContent
buildSummaryMessageContent,
toolExchangesAsText
} from './compactionPrompt'
import { dfs } from '$lib/components/flows/previousResults'
import { redactFileArgs, redactSecretArgs } from '$lib/components/job_args'
@@ -171,7 +172,10 @@ import { PlanModeController, type PlanModeHost } from './planModeController.svel
// cannot see — the upcoming completion and tool results, system-prompt/tool-
// schema changes from mode switches, and the estimate's chars/4 error.
const COMPACTION_TRIGGER_RATIO = 0.8
const COMPACTION_TARGET_RATIO = 0.7
// The gap below the trigger is what one compaction buys: each summarization request
// carries most of the window, and a target close to the trigger spends that on a few
// turns of room and summarizes its own previous summary again soon after.
const COMPACTION_TARGET_RATIO = 0.5
// How often a running turn is offered to the mid-turn checkpoint (see
// sendRequest). The whole transcript is rewritten on each accepted checkpoint,
// so this bounds the write rate; it also bounds how much of a turn a tab that
@@ -1552,7 +1556,7 @@ export class AIChatManager implements ChatViewHost {
[
// Strip image blobs from the summarizer input — the summary text stands in
// for them, so re-sending base64 to the summarizer only wastes tokens.
...stripImagePartsFromMessages(sanitizeToolCallArguments(prefix)),
...toolExchangesAsText(stripImagePartsFromMessages(sanitizeToolCallArguments(prefix))),
{ role: 'user', content: getCompactionSummaryPrompt() }
],
abortController,
@@ -3061,14 +3061,14 @@ describe('AIChatManager context compaction', () => {
it('compacts the stored history before sending once reported usage projects over the trigger', async () => {
const manager = new AIChatManager()
manager.messages = [
{ role: 'user', content: 'a'.repeat(400_000) }, // ~100k estimated tokens
{ role: 'assistant', content: 'b'.repeat(400_000) }, // ~100k
{ role: 'user', content: 'a'.repeat(1_200_000) }, // ~300k estimated tokens
{ role: 'assistant', content: 'b'.repeat(1_200_000) }, // ~300k
{ role: 'user', content: 'c'.repeat(400) },
{ role: 'assistant', content: 'd'.repeat(400) }
]
// Provider fact: 850k used. Projected past the 800k trigger, so ~150k
// must be freed to come back to the 700k target — the first user +
// assistant pair (~200k estimated).
// Provider fact: 850k used. Projected past the 800k trigger, so ~350k
// must be freed to come back to the 500k target — the first user +
// assistant pair (~600k estimated).
manager.contextUsage = 850_000
manager.instructions = 'next question'
const saveChat = vi.spyOn(manager.historyManager, 'saveChat')
@@ -3086,7 +3086,7 @@ describe('AIChatManager context compaction', () => {
// compaction-time save) so a rolled-back turn keeps a consistent value
// 4th arg: the modified-items mask rides on every save (undefined here —
// this bare manager never initialised tracking).
expect(saveChat).toHaveBeenCalledWith(expect.anything(), expect.anything(), 650_000, undefined)
expect(saveChat).toHaveBeenCalledWith(expect.anything(), expect.anything(), 250_000, undefined)
// At commit, the no-report turn clears the stored value; the readable
// number falls back to estimating the now-tiny compacted history
expect(manager.contextUsage).toBeUndefined()
@@ -2,9 +2,38 @@ import { describe, expect, it } from 'vitest'
import {
buildSummaryMessageContent,
formatCompactSummary,
getCompactionSummaryPrompt
getCompactionSummaryPrompt,
toolExchangesAsText
} from './compactionPrompt'
describe('toolExchangesAsText', () => {
it('renders tool calls and results as plain text messages', () => {
const sent = toolExchangesAsText([
{ role: 'user', content: 'weather?' },
{
role: 'assistant',
content: null,
tool_calls: [
{
id: 'call_1',
type: 'function',
function: { name: 'get_weather', arguments: '{"city":"Paris"}' }
}
]
},
{ role: 'tool', tool_call_id: 'call_1', content: 'sunny' },
{ role: 'assistant', content: 'It is sunny.' }
])
expect(sent.every((m) => m.role !== 'tool' && !('tool_calls' in m))).toBe(true)
expect(sent[1]).toEqual({
role: 'assistant',
content: '[Called tool `get_weather` with arguments: {"city":"Paris"}]'
})
expect(sent[2]).toEqual({ role: 'user', content: '[Result of tool `get_weather`: sunny]' })
expect(sent[3]).toEqual({ role: 'assistant', content: 'It is sunny.' })
})
})
describe('formatCompactSummary', () => {
it('strips the analysis scratchpad and unwraps the summary block', () => {
const raw = `<analysis>
@@ -65,6 +94,15 @@ chronological thinking the model should not keep
expect(formatted).not.toContain('before output')
})
it('keeps the whole summary when it quotes its own tags', () => {
const raw = `<analysis>notes</analysis>
<summary>1. The user asked for an <analysis> block then a <summary></summary> block.
2. Work continued.</summary>`
expect(formatCompactSummary(raw)).toBe(
'1. The user asked for an block then a block.\n2. Work continued.'
)
})
it('strips every analysis block, not just the first, when the summary is untagged', () => {
const raw = '<analysis>first</analysis>\nkept one\n<analysis>second</analysis>\nkept two'
const formatted = formatCompactSummary(raw)
@@ -1,3 +1,5 @@
import type { ChatCompletionMessageParam } from 'openai/resources/chat/completions.mjs'
// Summary-based compaction: when a conversation approaches the model's context
// window, the older prefix is replaced by an LLM-generated structured summary
// while the recent tail is kept verbatim. The summary precedes the kept tail,
@@ -24,6 +26,8 @@ const NO_TOOLS_TRAILER =
// strips before the summary reaches context.
const SUMMARY_PROMPT = `Your task is to create a detailed summary of the conversation so far. This summary will be placed at the start of a continuing session; newer messages that build on this context will follow after it (you do not see them here). Summarize thoroughly so that someone reading only your summary and then the newer messages can fully understand what happened and continue the work without losing context.
The message you are reading now is an instruction, not part of the conversation. Summarize only the messages above it: do not describe this instruction, do not list it among the user's messages, pending tasks or current work, and do not mention the <analysis> or <summary> tags inside your summary.
This is a conversation with Windmill's global workspace assistant. It inspects workspace items and authors them as per-user drafts — scripts, flows, apps, resources, variables, triggers, and schedules — then deploys those drafts and test-runs scripts and flows. It works with items by their workspace path (e.g. \`u/alice/sync_orders\`, \`f/team/my_flow\`); it does NOT edit files on a filesystem. Frame the summary in those terms.
Before providing your final summary, wrap your analysis in <analysis> tags to organize your thoughts. In your analysis:
@@ -105,7 +109,9 @@ export function formatCompactSummary(raw: string): string {
// real summary boundary.
let formatted = raw.replace(/<analysis>[\s\S]*?<\/analysis>/gi, '')
const summaryMatch = formatted.match(/<summary>([\s\S]*?)<\/summary>/i)
// Greedy to the last closer: the summary describes the instruction that asked for
// it, tags included, and stopping at a quoted </summary> cuts it off mid-sentence.
const summaryMatch = formatted.match(/<summary>([\s\S]*)<\/summary>/i)
if (summaryMatch) {
formatted = (summaryMatch[1] ?? '').trim()
} else {
@@ -125,6 +131,34 @@ export function formatCompactSummary(raw: string): string {
return formatted.replace(/\n{3,}/g, '\n\n').trim()
}
/**
* The prefix with every tool exchange rendered as text. The summarization request
* carries no tool definitions, and Bedrock rejects tool-use and tool-result blocks
* that arrive without them; the summary only needs what was called and what came back.
*/
export function toolExchangesAsText(
messages: ChatCompletionMessageParam[]
): ChatCompletionMessageParam[] {
const toolNames = new Map<string, string>()
return messages.map((m) => {
if (m.role === 'assistant' && m.tool_calls?.length) {
const lines = m.tool_calls.map((t) => {
if (t.type !== 'function') return `[Called tool ${t.type}]`
toolNames.set(t.id, t.function.name)
return `[Called tool \`${t.function.name}\` with arguments: ${t.function.arguments}]`
})
const text = typeof m.content === 'string' && m.content ? m.content + '\n\n' : ''
return { role: 'assistant', content: text + lines.join('\n') }
}
if (m.role === 'tool') {
const name = toolNames.get(m.tool_call_id) ?? 'unknown'
const result = typeof m.content === 'string' ? m.content : JSON.stringify(m.content)
return { role: 'user', content: `[Result of tool \`${name}\`: ${result}]` }
}
return m
})
}
/**
* Wraps a formatted summary as the content of the user message that replaces
* the summarized prefix in the conversation.
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -69,6 +69,9 @@ export function requiresMaxCompletionTokens(model: string) {
// (trim/compaction, the usage indicator) go through
// getEffectiveModelContextWindow, whose conservative 128K fallback keeps a
// limit enforced and is surfaced to the user as an assumed window.
//
// Keep entries aligned with MODEL_CONTEXT_WINDOWS in
// `backend/windmill-ai/src/model_context.rs`, used by AI agent compaction.
const MODEL_CONTEXT_WINDOWS: [name: string, contextWindow: number][] = [
// Anthropic — Sonnet/Opus 4.6+ ship a 1M window at standard pricing (GA);
// Haiku, older Claude models (3.x, 4.0, 4.1, 4.5) and date-suffixed Claude 4
@@ -110,8 +110,9 @@ describe('memoryOptionLabel', () => {
// The ignored-input note names the setting by the same label its own button carries.
it('names each memory option the way the field renders it', () => {
expect(memoryOptionLabel({ kind: 'manual', messages: [] })).toBe('Previous messages (legacy)')
expect(memoryOptionLabel({ kind: 'auto', context_length: 4 })).toBe('On (legacy)')
expect(memoryOptionLabel({ kind: 'window', context_length: 10 })).toBe('On')
expect(memoryOptionLabel({ kind: 'auto', context_length: 4 })).toBe('Last messages (legacy)')
expect(memoryOptionLabel({ kind: 'window', context_length: 10 })).toBe('Last messages')
expect(memoryOptionLabel({ kind: 'compaction', context_window: 128000 })).toBe('Compaction')
// Keeping no messages runs as off, whichever kind says so.
expect(memoryOptionLabel({ kind: 'window', context_length: 0 })).toBe('Off')
expect(memoryOptionLabel({ kind: 'auto' })).toBe('Off')
@@ -130,9 +131,15 @@ describe('memoryPropertyFor', () => {
expect(kinds({ kind: 'auto', context_length: 4, memory_id: 'x' })).toEqual([
'off',
'window',
'compaction',
'auto'
])
expect(kinds({ kind: 'manual', messages: [] })).toEqual(['off', 'window', 'manual'])
expect(kinds({ kind: 'manual', messages: [] })).toEqual([
'off',
'window',
'compaction',
'manual'
])
const autoVariant = (value: unknown) => memoryPropertyFor(property, value).oneOf.at(-1)
expect(autoVariant({ kind: 'auto', context_length: 4 }).properties.memory_id).toBeUndefined()
expect(
@@ -52,16 +52,20 @@ export function agentTestInputTransforms(
}
}
/** What turning managed memory on writes. */
export const DEFAULT_AGENT_MEMORY: MemoryConfig = { kind: 'window', context_length: 10 }
/** What turning managed memory on writes. Both call sites are chat mode, and a chat conversation
* is open-ended: it keeps everything and summarizes the older part rather than dropping messages
* off the front. No context window, so the run reads the one known for the model it ends up on. */
export const DEFAULT_AGENT_MEMORY: MemoryConfig = { kind: 'compaction' }
/** The docs section on how an agent's memory is named and kept. */
export const AGENT_MEMORY_DOCS_URL =
'https://www.windmill.dev/docs/core_concepts/ai_agents#memory-auto--manual'
/** Whether Windmill stores and replays the agent's conversation, mirroring the worker: `window`, or
* its older spelling `auto`, with a message count above 0. A legacy `manual` list is not managed. */
/** Whether Windmill stores and replays the agent's conversation, mirroring the worker: `compaction`,
* which keeps all of it, or `window` and its older spelling `auto` with a message count above 0. A
* legacy `manual` list is not managed. */
export function keepsManagedMemory(memory: any): boolean {
if (memory?.kind === 'compaction') return true
return (memory?.kind === 'window' || memory?.kind === 'auto') && Boolean(memory.context_length)
}
@@ -90,6 +94,11 @@ export function historyInputApplies(
/** A memory setting in words, for a linked agent's summary. */
export function describeMemoryPolicy(memory: any): string {
if (memory?.kind === 'compaction') {
return memory.context_window
? `Whole conversation, summarized near ${memory.context_window} tokens`
: "Whole conversation, summarized as it fills the model's context window"
}
if (keepsManagedMemory(memory)) return `Last ${memory.context_length} messages`
if (memory?.kind === 'manual') return 'Off, sends previous messages saved with the agent'
return 'Off'
@@ -165,7 +174,8 @@ export const AGENT_FIELDS: AgentFieldSpec[] = [
key: 'memory',
group: 'messages',
label: 'Managed memory',
tooltip: 'Windmill stores the conversation and sends its last messages with each request.',
tooltip:
'Windmill stores the conversation and sends it with each request: its last messages, or a summary of the older ones with the recent ones verbatim. Without instance object storage, saved memory is limited to 100KB. Compaction tries a summary, then keeps the newest complete conversation that fits, starting with a user message. If none fits, memory is not updated.',
implicit: { kind: 'off' },
defaultHint: 'Default: off',
textOnly: true
@@ -1,6 +1,7 @@
<script lang="ts">
import { Alert, Button } from '$lib/components/common'
import { keepsManagedMemory } from '../agentFormFields'
import { MEMORY_OPTION_LABELS } from '../flowInfers'
interface Props {
/** The agent's input transforms. Converting a legacy setting writes `memory` and the step input
@@ -10,15 +11,9 @@
/** Whether the step's own memory id and previous messages are on this form. A saved agent has
* neither: they belong to each step linking it. */
historyOnStep?: boolean
s3StorageConfigured?: boolean
}
let {
args = $bindable(),
chatInputEnabled = false,
historyOnStep = false,
s3StorageConfigured = true
}: Props = $props()
let { args = $bindable(), chatInputEnabled = false, historyOnStep = false }: Props = $props()
let memory = $derived(
args?.memory?.type === 'static'
@@ -41,6 +36,7 @@
// An `auto` setting whose saved id is never read runs exactly like the current setting for its
// state, so switching to that setting is the only choice.
let legacyEquivalent = $derived(memory?.kind === 'auto' && !legacyMemoryId)
let equivalentLabel = $derived(on ? MEMORY_OPTION_LABELS.window : MEMORY_OPTION_LABELS.off)
// The older setting never read the step's own memory id, so a conversion that promises the same
// behaviour, or the run's id, drops it rather than bringing it to life. Off keeps ignoring it.
@@ -70,11 +66,6 @@
}
</script>
{#if on && !s3StorageConfigured}
<p class="mt-1 text-2xs text-hint">
Without S3 storage on the workspace, memory is kept in the database, up to 100KB per memory.
</p>
{/if}
{#if legacyMessages}
<Alert type="info" title="Older memory setting" class="mt-2">
<div class="flex flex-col gap-2">
@@ -133,9 +124,7 @@
<Alert type="info" title="Older memory setting" class="mt-2">
<div class="flex flex-col gap-2">
<span>
An earlier version of the editor saved this setting. It works the same as {on
? 'On'
: 'Off'}.
An earlier version of the editor saved this setting. It works the same as {equivalentLabel}.
</span>
<div class="flex">
<Button
@@ -144,7 +133,7 @@
btnClasses="bg-surface"
onclick={switchToEquivalent}
>
Switch to {on ? 'On' : 'Off'}
Switch to {equivalentLabel}
</Button>
</div>
</div>
@@ -580,7 +580,6 @@
bind:args
{chatInputEnabled}
historyOnStep={scopedFields.some((f) => f.key === 'previous_messages')}
s3StorageConfigured={s3Storage.current}
/>
{:else if spec.key === 'memory_id' && memoryIdOffered}
{@render transformField(
@@ -10,8 +10,9 @@ import { AGENT_HISTORY_KEYS } from './agentFormFields'
* picked cannot drift from the button they see. */
export const MEMORY_OPTION_LABELS: Record<string, string> = {
off: 'Off',
window: 'On',
auto: 'On (legacy)',
window: 'Last messages',
compaction: 'Compaction',
auto: 'Last messages (legacy)',
manual: 'Previous messages (legacy)'
}
@@ -59,7 +60,7 @@ export const AI_AGENT_SCHEMA: Schema = {
memory: {
type: 'object',
description:
'Windmill stores the conversation and sends its last messages with each request.',
'Windmill stores the conversation and sends it with each request, keeping either its last messages or a summary of the older ones.',
enumLabels: MEMORY_OPTION_LABELS,
// Chat mode keys memory on the conversation, so a chat whose agent has memory off
// forgets every turn. Enabling chat mode turns it on; this keeps it there. A step
@@ -88,6 +89,20 @@ export const AI_AGENT_SCHEMA: Schema = {
}
},
required: ['kind', 'context_length']
},
{
type: 'object',
title: 'compaction',
properties: {
kind: { type: 'string', enum: ['compaction'] },
context_window: {
type: 'number',
title: 'Context window',
description:
"Leave empty to use the model's context window. Set it for custom models."
}
},
required: ['kind']
}
],
showExpr: "fields.output_type !== 'image'"
+23 -2
View File
@@ -571,6 +571,25 @@ components:
required:
- kind
MemoryCompaction:
type: object
description: |
Keeps the whole memory named by the run's memory id (or the step's `memory_id`), replacing
its older part with a summary as the conversation approaches the model's context window.
Without a memory id the agent runs without memory, and compaction bounds the run's own loop.
properties:
kind:
type: string
enum:
- compaction
context_window:
type: integer
description: |
Overrides the context window looked up from the model, in tokens. Only a model
Windmill does not know needs one; those fall back to 128000.
required:
- kind
MemoryMessage:
type: object
description: A single message in conversation history
@@ -609,6 +628,7 @@ components:
oneOf:
- $ref: '#/components/schemas/MemoryOff'
- $ref: '#/components/schemas/MemoryWindow'
- $ref: '#/components/schemas/MemoryCompaction'
- $ref: '#/components/schemas/MemoryAuto'
- $ref: '#/components/schemas/MemoryManual'
discriminator:
@@ -616,6 +636,7 @@ components:
mapping:
'off': '#/components/schemas/MemoryOff'
window: '#/components/schemas/MemoryWindow'
compaction: '#/components/schemas/MemoryCompaction'
auto: '#/components/schemas/MemoryAuto'
manual: '#/components/schemas/MemoryManual'
@@ -1090,8 +1111,8 @@ components:
parameter). Leave unset to use the run's memory id. A fixed value shares one memory
across every run; an expression such as `flow_input.customer_id` keeps one memory per
key. When it evaluates to an empty value the agent runs without memory. Read only
while `memory` is `window`: it is ignored when memory is off, and an older `auto` or
`manual` memory reads neither history input.
while `memory` is `window` or `compaction`: it is ignored when memory is off, and an
older `auto` or `manual` memory reads neither history input.
previous_messages:
allOf:
- $ref: '#/components/schemas/InputTransform'
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long