mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-09-07 08:02:40 +00:00
bebd8bef8ba86f0cc016db3827ed1d8ee8c577fc
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9c557859c5 |
feat: AI agent evals: datasets, scored runs and comparison (#10633)
* feat: eval datasets and standalone runs for reusable AI agents Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: agent eval drawer with case editor, runs and capture entry points Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: document AI agent eval datasets and standalone runs Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: say how many eval cases the list is not showing Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address review findings on eval datasets - keep an edited case's conversation and tool inputs: serde(flatten) silently drops Box<RawValue> fields, so the update payload is spelled out - remount the case editor per case so one case's turns cannot leak into another - require jobs:read / flow_conversations:read on the capture endpoints, which UserDB does not gate by token scope - take the dataset lock in create and update so a delete cannot be undone by a concurrent metadata write, and delete cases before metadata - load more cases beyond the first page, and stop capping the agent picker - record that the version stamp is taken at enqueue, not at resolution Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-2 review findings on eval datasets - block operators from dataset and case writes - pass the editor's operating workspace through the drawer and the capture request, instead of assuming the navigation workspace - discard superseded case-list responses so switching datasets cannot land the previous dataset's cases - reject a dataset without a case_id (or vice versa) rather than running an inline case under a dangling association - run unsaved edits inline instead of silently running the stored case - surface the API error body on a failed run - fetch dataset metadata concurrently when listing - $bindable() without a default on the optional open prop - correct the permission and enqueue-time-version wording in the docs Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: run an untouched saved case by reference again The editor writes back keys the stored case omits, so comparing the raw objects reported every unedited case as edited: the run went inline and lost the dataset/case stamp its history depends on. Compare a normalized form, and pin it with a test. Also scope the history query to the drawer's workspace and drop superseded responses. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: show a dataset's cases as a table, and fix round-4 review findings The case list showed one case at a time with no overview. It is now a table with the case, where it was captured from, and its last run — the last-run column is a single jobs query on the path stamp rather than a request per row. Review fixes in the same file: - keep the edit baseline on the selected case rather than looking it up in the loaded page, so a case beyond page 1 is not treated as unedited and run stale - release the loading state when a superseded case load returns early - reload every loaded page after a write instead of collapsing to page 1 - last remaining 'resolved to' wording in the version tooltip Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: run a dataset as an experiment, with scorers as runnables An experiment runs every case of a dataset against one subject and records the exact case set it executed, so a result set stays reproducible while the dataset keeps changing. Each case runs as its own small flow — the agent, then a step per scorer — so a case keeps the run stamp, history query and trajectory view a single run already has, and scorers need no orchestration of their own. Results are read back per step by node id rather than by walking a nested loop's status. A scorer is any runnable taking (input, output, expected): a script, a flow, or a reusable agent used as a judge. A judge is prompted with the case and the answer as one JSON message; a script or flow receives them as named arguments. Scores accept a bare number, a boolean or {score}. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: results table for an experiment, with scorer columns One row per case: status, the agent's answer, and a column per scorer, with the mean per scorer above the table and a link into each case's run for its trajectory. Averages skip cases a scorer produced no number for — counting a missing score as zero would read as a regression. The drawer's left pane becomes Cases / Results, and Results carries the scorer picker and Run dataset. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: compare an experiment against a baseline Per-scorer deltas on each row and on the mean, and a filter down to the rows that regressed. Rows join by case id, so a case added after the baseline ran has no delta instead of counting as a change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-5 review findings on experiments - match scorers by label when diffing two experiments; joining by array position subtracted one scorer from another whenever the scorer sets differed - report a row's status from the case job, not the agent step, so a case whose scorer failed no longer reads as a success - delete a dataset's experiments with it: they hold copies of its cases, and a recreated dataset of the same path would have exposed them - select the experiment that Run dataset just started instead of leaving the table on the previous one - expected is scored now, so stop describing it as having no consumer Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-6 review findings on experiments - hold the dataset lock across an experiment launch, so a delete landing between reading the cases and writing the experiment cannot recreate the deleted dataset's inputs - match scorers between experiments on kind and path, not on label: labels default to a path's last segment, so f/a/quality and f/b/quality compared against each other - average mean deltas over the cases both runs scored; comparing each run's own average reported a regression from a case the baseline never ran, with no regressed row to point at - openapi: the row status is the job's, which is also canceled/skipped; runEval takes scorers; the update-case body no longer advertises source, which the handler deliberately ignores - record why the experiment prefix cannot reach a sibling dataset Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-7 review findings on experiments - release the dataset lock for the push loop and retake it for the write, re-checking the dataset still exists: holding it across the whole launch made every capture and case edit on that dataset 409 until the last job queued - assemble experiment results with bounded concurrency; a 100-case, 3-scorer experiment was 400 sequential lookups, each itself several queries - clear the baseline when it becomes the selected experiment, which was comparing a run against itself and reporting zero deltas - take the header mean over the same cases as its delta while comparing, so the two numbers beside each other describe the same set - a canceled or skipped case is no longer the same grey dot as a running one Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-8 review findings on experiments - verify the dataset's identity, not just its existence, before recording an experiment: the path can be deleted and recreated during the push loop, and the experiment holds copies of the old dataset's cases - give the recording lock a longer budget than a case edit, since its jobs are already queued and giving up strands them, and say so when it fails - keep score lookups sequential within a case: nesting two bounded streams multiplied into 32 in-flight queries against a 50-connection pool - clear a baseline that no longer belongs to the loaded experiments, so switching datasets does not leave comparison mode on with nothing to compare - keep a scorer's own mean when the baseline never ran it, instead of blanking a column full of numbers - EvalCaseDraft.expected no longer claims nothing scores it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: do not trust an experiment's job ids, and require write to record one Experiment objects live in workspace object storage, which a script can write directly, and results are read on the unrestricted pool — so a forged experiment naming another flow job returned output the jobs API would have refused. Only jobs this server stamped with that experiment's id are read now. Also from round 9: - recording an experiment requires write on the dataset, not read: it persists into the dataset's namespace and its shared list - clear the results table when the selection changes and surface a failed load, instead of labelling the previous experiment's numbers as the new one's - a storage fault is no longer reported as a deleted dataset - the lock-timeout message at the recording site no longer says to retry, which would run the whole dataset again on top of the jobs already queued - ExperimentRow.status documents canceled and skipped Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: bind the experiment trust check to the requested dataset The previous check matched jobs on the experiment id alone, which the stored object supplies — so copying another dataset's experiment JSON under a readable key carried its jobs' output along with it. A job is now only read if it was stamped for this experiment *and* for the dataset the caller's read access was checked against, and an experiment that names a different dataset is not served from this key at all. Also from round 10: - add the .sqlx entry for that query; without it every SQLX_OFFLINE build failed - serve results over GET: as POST the route-scope middleware classified a read as ai_evals:write, locking read-only tokens out of their own results - clear the selected and baseline experiments synchronously when the dataset changes, so the previous dataset's id is not requested under the new one Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-11 review findings on experiments and scorers - give scorers the whole case input, not just the message: an answer that came from attachments or a replayed conversation could not be judged on it - accept a judge's boolean and structured {score} answers, including stringified ones, and pin every documented scorer shape with a test - record an experiment for the cases that did launch when a later push fails, instead of leaving those jobs running with nothing to attribute them to - do not capture a preview parent's synthetic runnable_path as a host flow; the saved case could not be rerun - clear the case table before loading a dataset and surface a failed load, so a failure cannot leave the previous dataset's cases under the new name - keep the results table through a refresh of the same experiment - exclude flow-step jobs from the per-case last-run lookup - drop case sets from the experiment list, which is only used to pick a run - report a database failure at the recording lock as itself, not as contention Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-12 review findings on capture and run history - load flow_node.flow for flownode parents: an agent inside a deployed branch or loop captured without its agent, host flow or tool bindings - decide host_flow_path by whether the path resolves to a flow, not by job kind: excluding previews wholesale also dropped the flow editor's step test, whose path is real - page the per-case last-run lookup by created_before until the loaded cases are covered; one page of 200 reported older cases as never run - do not record an experiment when nothing launched - only attach the case input to a job when a scorer will read it - keep the case table through a save; only a different dataset clears it - drop the superseded duplicate comment on the score parser Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: stop refetching run history on every case write Reading the case list before the first await made the whole job-history query a dependency of it, so every save, delete and Load more refetched up to 1000 job rows and blanked the column. Read untracked instead. - an empty Last run cell now distinguishes never-ran from not-found-within the page bound, which the comment already claimed and the cell did not - reloading a dataset no longer replaces a populated table with a skeleton - keep the score-parser comment that describes every shape it handles Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: keep eval datasets in Postgres instead of object storage Datasets, cases and experiments become rows (`eval_dataset`, `eval_case`, `eval_experiment`, `eval_experiment_case`) rather than objects under a `wmill_eval_datasets/` prefix. What a run produced is still the job's: only case inputs and an experiment's case snapshot are stored. This removes the machinery the object store needed: - The advisory lock and the read-modify-write of a per-dataset JSONL. A case is a row, so there is nothing to serialize. - The launch-time identity check on the dataset. The foreign key makes a concurrent delete fail the transaction instead. - The trust guard on an experiment's job ids, which existed because a script can write workspace object storage directly and could forge an experiment naming somebody else's job. An experiment now chooses every job id and records itself before pushing anything, so a launch that dies partway leaves a recorded case whose job is missing rather than a running job nothing accounts for; cases that never reached the queue are removed again. Row-level security on `eval_dataset` is the authority on who may read or write a dataset, so `extra_perms` grants work and the rule is not mirrored in Rust. Cases and experiments carry a read policy derived from their dataset and no write policy: they are written on the unrestricted pool after the dataset row itself has been asked, with `SELECT ... FOR UPDATE`, whether the caller may write it. Cases are capped at 256 KiB each and 10 000 per dataset, refused rather than truncated. Attachments are S3 references, not inline bytes, so a case that approaches either cap is a mistake rather than a use case. Evals no longer need the `parquet` feature or a configured workspace object storage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * style: align the eval drawer with the design system - Scorer chips are `Badge`s rather than a hand-rolled bordered span, and the section header is a `Label` with its tooltip, as are the case editor's fields (which also gets the label colour right). - The results table showed status as a coloured bullet, which says nothing to a colour-blind reader. It now carries the same icons the runs table uses, with the status as its accessible name. - Feedback colours move to the `-500` shades the brand guidelines name. - The conversation JSON error uses `TextInput`'s `error` prop for the border and the caption style for the message, as elsewhere. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: author an expected answer, tags and attachments on a case Every scorer is handed `(input, output, expected)`, but nothing could produce an `expected` except a conversation capture: the case editor had no field for it and a captured run left it empty. So: - The editor gains Expected, Tags and a read-only list of the attachments a captured case carries. Expected is plain text, or JSON when the answer has structure. - Capturing from an AI agent run keeps what that run answered, which is the only moment a reference answer exists for free. The results table also laid itself out by content, so a long answer pushed the scores — the numbers the table exists for — off the edge of the pane. It is fixed-layout now, with the text columns bounded. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: expected is captured from a run and can be authored Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: link a saved agent when inserting an ai agent step "AI Agent" in the step picker was a leaf that always created a blank step, so reusing a saved agent meant inserting a blank one, opening its step input and linking it there. It is a category now, like Flow and AI Sandbox, listing the workspace's `ai_agent` resources next to a blank option, filtered by the picker's own search. A picked agent produces a step that is already linked rather than one linked afterwards: `agent` set, no tools, and only the flow-local `user_message`/`user_attachments` transforms. Seeding the brain keys there would leave transforms a linked step never reads and that `AgentResourceBar` strips on its next link change. Each `on:new` forwarder rebuilds the insert detail field by field instead of spreading it, so a new field is dropped unless the forwarder names it. `agentPath` is typed on both `GraphEventHandlers.insert` and `FlowGraphV2`'s `onInsert` so the next one to forget it fails the check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: restore the link on cancel and simplify the agent bar Cancel on an agent edit forked the step into a standalone copy, which is the opposite of what the word means and needed a paragraph under the card to explain. It discards the edits and re-links the step now, leaving the agent untouched; diverging from an agent is Unlink's job, on the linked card. This flow's `tool_inputs` survive the round trip as overrides, so Cancel no longer folds them into the tools the way Unlink does. Linking a step to a saved agent happens in the step picker at insert time, so the bar's own resource picker is gone and "Save as agent" is the one action left. Its `+` button was a trap besides: it opened the generic resource form, where an agent would have to be written as raw JSON. The card itself was `surface-secondary`, the sections token, so in dark mode it was darker than the pane and read as a sunken well rather than an elevated card. It uses `surface-tertiary` as the brand table prescribes, its tool chips are `Badge`s, and the editing card no longer overflows the pane and clips its own buttons. The remaining tooltip follows the inline `Label` convention rather than sitting in a flex row whose gap stacked on the trigger's own margin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: rework the AI agent evals surface into one table Evals become a single pane: a dataset of cases, one column per scorer, one row per case, with the run being looked at chosen from the toolbar. Runs are permanent. Running the whole dataset opens one; running a single case records nothing at all — it is a job, and looking at what it did is not a claim that it belongs in the history. Its result and its scores sit over the row until they are saved as a run, which carries the cases that were not rerun and the scoring jobs themselves, so the number that is saved is the number that was looked at. A scorer is a runnable: a judge agent or a script, created in one click and edited in place. Scores carry a reason and per-assertion checks, shown on hover with a rescore button. What ran is always named. A run records the agent version, or — for a configuration that is not deployed — a hash of it, so a table can say that its numbers describe an agent that no longer exists: those rows dim and the table offers to rerun. An agent's draft can be run directly instead of the deployed value, and once those edits are deployed the runs that made them are recognised as that version. A step with no agent of its own is evaluable too, and saving it as an agent moves its history onto it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: keep an agent's in-progress edits on the agent Editing a linked agent forks it into the step, which is what makes the edits runnable there — but the agent is what is being edited, so that is where the unsaved state belongs. The edit is mirrored into the agent's own resource draft as it is made. It then survives leaving the flow, shows the agent as drafted wherever it appears, and is what evals run when asked to run the draft rather than what is deployed. Deploying or cancelling clears it; opening Edit without changing anything does not mark the agent as drafted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: shape the evals surface around a saved agent Evals hang off an `ai_agent` resource, so the surface is now only ever about one: the `draft` subject kind, the standalone-step subject and the move that carried a step's history onto a newly saved agent are gone. - A run is permanent and numbered per agent. Running a single case is a trial: it answers in the panel and never touches the table. - "Run scorers only" opens a run of its own that reuses the answers of the run you are looking at, so a scorer added later measures what already ran without calling the agent again. - A draft run whose configuration is later deployed is stamped, once, to the version it became, so its label stops reading `v23 + edits` forever. - A scorer can carry a pass threshold, read off the scores already recorded. - The table is the case, its answer and one number per scorer; datasets are created and edited in a drawer; a run that executed an earlier state of the current draft says so above the table, in one line. - Which agent a step is, whether it is being edited, and which version it is on is a strip above the step's tabs, because it is true of every tab. - Capturing a case from a step test or a conversation is dropped, and with it the `memory` override on a linked step that nothing set. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: run past versions of an agent, and number versions per resource The evals home becomes one table of every run of the agent, whichever dataset each is of, with one badge per scorer. A list spanning datasets cannot hold every dataset's scorers to look a name up, so a score carries its name and kind with its number, and thresholds are joined in per run and column. Run now asks what to run: the latest agent, resolved when the run executes as a flow step does, any past version, or the unsaved edits. Pinning is a subject kind of its own, since a linked step resolves the resource live and inlining is the only way to run a version that is no longer current. Scorers move into the edit-dataset drawer. The column header over a run reports and nothing else: a run is permanent, and a control there that changed the columns would edit the past from the one place that must not. Adding one offers four ways rather than two, writing and reusing being different jobs, and both new kinds open with a summary filled in. Versions are numbered per resource. `resource_version.id` is one identity sequence for the whole table, so an agent saved nine times read v4 ... v24, and the gaps counted writes in workspaces the reader cannot see. The id stays how a version is addressed; the new number is what it is called, in the resource history drawer as well as here. It is assigned on write rather than counted on read because trimming past the cap and clearing a history both take the oldest rows, and counting the survivors would renumber a version a run already names. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: read the dataset a remembered selection names Reopening the evals modal restored the last dataset from storage as a bare path, without reading the row it names. Every "is this already the one?" test compared against that selection, so all of them short-circuited and the dataset was never loaded: editing it opened a drawer with no summary, no scorers and no cases. The remembered path is now brought into context the same way any other choice is, and the tests compare against the dataset that is loaded rather than the one that is selected, so a selection can no longer stand for a read that did not happen. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: give dialogs a trail in their header A dialog deep enough to navigate had nowhere to say where you were: the header held a fixed title, and the way back was a control each body placed for itself, somewhere in a toolbar that moves with everything else the toolbar holds. The header is the one part of the surface that does not move, which is where the trail belongs. `Modal` takes an optional `trail` of levels below its title, rendered as a breadcrumb whose ancestors are the way back. Declarative on purpose: callers of this depth already hold the state that says where they are, so the dialog reads it rather than owning a stack they would have to push and pop in step with it. Escape follows the trail. Leaving a level is what someone deep in a dialog means by it, and closing the whole surface throws away the navigating they did to get there; at the root it closes as before. That only works if a dialog can tell it is the surface being addressed, so `Disposable` now answers `isTopmost()` and the dialog asks before acting: it keeps Escape for itself, so nothing else was arbitrating between it and a drawer opened from inside it, and both were acting on one key press. Evals is the first caller: its runs list is the root, a run is a level in it, and the back button that used to sit above the table is gone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: portal dialogs out of wherever they were opened from A dialog rendered in place inherits whatever the calling component happens to sit inside. One `transform`, `filter` or `overflow` anywhere above it makes its `fixed` positioning resolve against that ancestor instead of the viewport, and a surface meant to cover the app is then confined to a box it never asked for: the nav rail paints over it and its own edges are clipped. Drawers have always portalled for this reason. Dialogs only did so when an enclosing pane claimed them, and rendered in place otherwise, so the same screen could show a drawer over everything and a dialog trapped behind the nav. They now portal the same way: to the pane when one claims it, to `body` otherwise. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: make the dialog's title the first step of its trail The trail listed levels below the title, so a dialog one level deep read "Evals > All runs > Run 20 · v6": three steps for two places, the first two of them the same place under different names. The title is the root, so it is the root's own segment, and the trail a dialog is given is now the whole path with that segment at its head. Its height stopped moving too. A heading carries a line-height of its own, so a header holding only an h3 stood six pixels shorter than one holding segments as well, and the dialog's whole top edge stepped as you navigated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: sharpen the evals controls around where you are standing Each screen now offers what belongs to it. The list starts runs; a run is a record, so it offers only the one thing that acts on the record itself, which is measuring the answers it already stored. Starting a fresh run from inside one asked which agent and which dataset from the screen least about either, and scoring an existing run was offered from the list, where there is no run to score. Which run and what it is read against are one question asked twice, so they sit together rather than at opposite ends of a row. Choosing what to run is now a toggle over the two states worth naming, the draft and the saved agent, with every earlier version one click further: running an old version is deliberate, and a list made all three look alike. The draft is read when the dialog opens rather than taken from the caller's polled copy, which could be seconds behind an agent edited a moment ago and would leave the option out exactly when it is the reason for opening the dialog. The dataset field carries its path under it and its edit button on hover, as a resource picker does, so the closed field says what the open list said. Edits waiting on an agent are a "draft" here as everywhere else in Windmill, rather than "+ edits". The dialog runs an evaluation rather than "the agent", which is what it was already called everywhere it is recorded. An agent being edited keeps its evals button on a line of its own, clear of the decision to save or discard. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: settle the evals controls on the patterns Windmill already has The version choice uses ToggleButtonMore, as the AI provider picker does: the two states worth naming stay in the group, the rest are behind the overflow menu, and the one you pick joins the group rather than appearing in a second control below it. The deployed one says which version it resolves to. A run offers nothing to start. Scoring an existing run again was the last thing left there, and it was one button explaining a distinction that the run and the dataset already make between them. The warning that a run executed an earlier draft is about the run on screen, so it goes when the run does rather than following you back to the list, and it sits against the table instead of inside a frame of its own. A dataset just created stays open for its scorers and cases: those are what a dataset is, they can only be added to one that exists, and closing on create sent you to find it again to add them. Scorer settings are a cog rather than a word, now that the row holds three actions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: close the gap in the version toggle and say what naming a dataset does The overflow trigger is not a pill, so the room it reserves showed as a gap between it and the button before it; it is pulled in by that much. The dataset field gets its clear button, which is also the slot the edit button is positioned against, so the two now sit where a resource picker puts them. Naming a new dataset said nothing about what happens next, and the drawer looked like it was missing the rest of itself. It says so instead: a scorer and a case both belong to a dataset, so there is nothing to attach either to until this one exists, and creating it leaves the drawer open on them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: choose a dataset's scorers while naming it A scorer is a reference to a runnable, not a child of the dataset, so it needs the dataset's name but not its row. The list is collected in the drawer while the dataset is being named and sent with the create, which already accepts one, so a dataset arrives holding the columns that were chosen for it rather than being made empty and then edited to hold them. Cases stay where they were: a case *is* a row of the dataset, so there is nothing for it to be a row of until one exists. The drawer says which of the two is which instead of leaving the screen looking like it is missing the rest of itself. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: level the version toggle and name the dataset in its own field The overflow trigger stands a row taller than a toggle button, so the group grew to its height and left the sunken background showing under every pill beside it. Every child of the group is the same height now, which is why the AI provider picker never had the band: it sizes them all alike. The dataset field says the summary with the path after it rather than carrying the path on a line below. The list stacks the two, which a one-line field cannot do, so it says both the other way round. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: tidy the evals forms and the run's own controls Picking a scorer that exists chooses between two sources rather than showing both: the ones already measuring something, and everything else in the workspace. The first list says what each is called with its path under it and what it already measures on the right, instead of three columns that were the same path truncated three ways whenever a scorer had no name of its own. A dataset's drawer says what it is for on the page rather than under an icon, and its summary is sized like the field beneath it. The run's own row lines up with the table under it, the warning above that table is spaced off the rule rather than sitting on it, and adding a case is gone from a run: a run is a record of cases that were answered, so curating them from it is editing what it measured. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: create a dataset holding the cases written for it Creating a dataset takes the cases to create it with, so one can be assembled in a single act instead of made empty and then filled in. The drawer holds them while the dataset is being named, gives them ids of its own to be edited by, and sends them with the create. Every case is checked before the dataset is written. `eval_case` grants users no write, so the rows cannot be inserted in the transaction that creates the dataset under the caller's own policies; validating first is what keeps "created holding these cases" from becoming "created, holding some of them", and the rows that do follow go in one transaction of their own. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: name the button for what it opens, and say what each version is Starting an evaluation asks which state of the agent and which dataset, and both cost a provider bill, so a button that read as spending one on the way past was lying about the click. It opens something, and says so. Running one case from the panel keeps its own name and its play icon, because that one does run on click. The version options say what they are rather than what they are not: what a flow step would or would not run is a fact about somewhere else, and someone choosing what to evaluate is not standing in it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: give the editing card two rows and mark evals as beta At the width of a step panel the card's one row wrapped: the line naming the agent, the line saying what saving does, and the two buttons deciding the edits' fate all fought for it. Deciding gets a row of its own, and evals sits against the line it is about, since evals of an agent being edited run the edits. Evals is named wherever it is offered. It read as a word in one state of the card and as an icon in the other, which is two things to recognise for one door. The dialog carries a beta badge against its own name, before any level below it: every way in lands there, so it is said once and stays put as you navigate. The version toggle spells out which is which. Both are the agent at v2 and the difference between them is the whole choice, so it is worth the width. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: name a new dataset, and lay the scorer's settings out like a step's inputs A new dataset arrives called "Dataset 1", which the path follows as it follows any summary: a dataset with none was one every table could only call by its path, and the two seeds are what the summary rule already produces. Scorer settings put each field's description between its label and its input, where a step's inputs put theirs, and its inputs are the size the rest of the drawer uses. The runnable behind the column is a link to it with its kind's icon, since it is a resource of its own and the one thing about it these fields cannot change. The line explaining that a pass line re-reads recorded scores went: the threshold is a number to set, and how it is applied is not a decision being made here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: curate a dataset in the drawer and save it in one act The drawer holds the cases while they are edited and writes them when it is saved: added, changed and dropped, whichever it is. Typing no longer writes, so a set is never half saved while someone is still deciding what is in it, and Save means the same thing whether the dataset exists yet or not. A case panel offers reading rather than acting. Running one case now and editing one from a run were the last two ways to change a record from the screen showing it, and the machinery behind the first went with it. The answer is rendered as the prose it is, under what it is: the case's result, whichever run is selected above it. The rest is what the run's table was doing to its own edges: a column name is clipped to its column rather than running into the next, the table squares off against an open panel, and that panel closes with the run it belonged to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: one border above a table, and a link to the run's job The row above the table drew a bottom border and the table draws its own top edge, so every table sat under two lines. The row keeps its spacing and the table keeps its edge. A column header no longer spins while its scores arrive: the cells under it are where the numbers are missing, and they say so themselves. The beta badge is the height of the word beside it rather than of the line it sits on. A run is one flow and therefore one job, so the run says where that job is: what it is doing, what it cost and what it logged are all there rather than reconstructed from the table. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: stream scores as each scorer finishes, and show them per case A scorer runs after the agent inside the case's own iteration, so its verdict can be read as soon as its step is done. Waiting for the iteration to end held every column of a case back until the last of them finished, which is why answers arrived one at a time and scores all at once. Reading a job that is still running needs one guard: a module with nothing in it is a step that has not run, not one that produced nothing, and recording the second makes a failure that never goes away. The panel beside the table shows what each column made of the case and why. The reason a judge gave was stored and never shown, which is the half of a score that says anything. It stops repeating the question the header already asks, and a case still running reads as waiting rather than as an answer that says "Running". A run is a number beside a dataset, so the list puts the two together. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: score a case with every scorer at once The scorers of a case read the answer and never each other, so they ran one after another for no reason: measuring a case now takes as long as its slowest column rather than as long as all of them. Each is a branch of its own, kept from failing the others, so a judge that errors costs its own column and no more. An iteration is three steps again — answer, payload, scores — rather than one per scorer, and each branch is named for the column it produces, so the graph of a run says which scorer did what instead of spelling out an id. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: read a judge's score out of the JSON it nearly wrote A judge quoting the agent inside its own reason writes those quotes unescaped, which is invalid JSON and also the most ordinary sentence for it to produce. The whole verdict was being thrown away over it, so a column that had a number reported having none. The number and the reason are now read straight out of such text. Deliberately not a second JSON parser: it finds the two keys and takes what follows, which is what survives a quote in the middle of a sentence. A case still running says so with a spinner rather than with the word "Running" sitting where its answer goes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: ask a judge for a shape instead of trusting it to write one A new judge carries an output schema, so the provider holds it to `{score, reason}` rather than the prompt asking it to. Windmill already delivers a schema whichever way the model takes it, a tool for Claude and Bedrock and the native parameter elsewhere, so there is no list of models to keep here. An agent with no runs offers its first one where the first row would be, rather than from a toolbar above a table that has nothing in it. Starting a run no longer picks a dataset for you. It fell back to whichever came first, which on an agent that has never run means offering another agent's set as though it were the obvious one; and with no dataset at all it says so and offers the one move there is. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: report a column that failed throughout, and hold the run dialog The runs overview dropped any column that produced no number, so a judge that failed on every case of a run vanished from the row and read as a column nobody had asked for. The aggregate now reports every column that has cells, with the count of the ones it failed on, and the badge says "failed" where there is nothing to average. A column with no cells at all is still left out: that one was added after the run and has nothing to say about it. Creating a dataset closes the drawer rather than turning it into an edit of what it just made: scorers and cases already ship with the create, so there is nothing left to stay open for. Reached from the run dialog, it gives the screen back with the new dataset selected, and the dialog keeps the version you had already chosen. Also: - the case panel's job link moves to the panel's own header, where its scope is: the job is the whole iteration, not the answer it sat over - one action in the scorer drawer's header, as its neighbours have. The reuse list picks rather than adds, and says which dataset each column already measures - adding a case is the last row of the list it lands in - the pane shows what it has read rather than an empty state it has not earned yet, and its rows say they open - the linked agent card loses a border it had inside another one * fix: keep the linked agent card's outline The card is a thing inside the step's inputs rather than a section of them, and the outline is what says so. Only the rule inside it goes: the detail it separates is already set apart by being detail. * refactor: fit the eval surface to the shipped design * feat: give a nested dialog a back control and the runs list its own moves Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: put a dialog's description under its title Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor: fold a dialog's back control into the crumb it returns to Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: edit a dataset's cases as a table rather than a list beside a form Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: edit a dataset's cases in the grid the data tables are edited in Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: edit a grid cell of prose in place, and cap a dataset at one page Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor: keep the cell editor's styles beside it, not in the vendored theme Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: keep an empty cell empty and cap the editor's growth Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: name the step that assembles a run for the scorers Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: run the payload step natively, and say so when nothing serves that tag Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: report an answer as answered while its scorers are still running Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: let a scorer say a case is not one it measures Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: score the answer, and leave a case with no expected answer unmeasured Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor: split the evals backend into modules Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: record what a run produced so it outlives its jobs * fix: read only the agent step's own tool jobs into the payload * fix: pin a run's configuration and give the judge the attachments * feat: write a dataset's cases in one transaction * chore: refresh the sqlx cache for the eval queries * fix: drop results a newer selection has superseded * fix: keep a draft the agent editor never opened on * feat: let a run record what it produced instead of waiting to be read * fix: serialize the replacements of a dataset's cases * fix: stop the poller from superseding a read slower than its interval * chore: refresh the sqlx cache * fix: keep a failed read from settling a cell as a case with no answer * fix: hold the case grid while its save is in flight * fix: keep a failed collect step from failing the run it recorded * chore: refresh the sqlx cache * fix: commit an open cell into the save that reads it * refactor: size the eval buttons with unifiedSize * docs: describe a run as the one flow it is * fix: show a run's recorded rows when part of it cannot be collected * refactor: size the remaining PR-added buttons with unifiedSize * fix: save the dataset name that was submitted, not the one typed after * fix: force an open cell into the save that was pressed for it * fix: refuse to score a run whose evidence could not be read * fix: hold one lock over a dataset's case count and its writes * fix: keep one unreadable run from costing the whole runs list * refactor: drop the banned bindable-default from the eval props * fix: hold the scorer controls while the dataset is written * fix: read only the caller's own draft of an agent * docs: say in the contract that a run pins its configuration * fix: say a scorer did not run rather than blaming a missing answer * feat: resume the agent draft you already had when you press Edit * refactor: build the trail and dataset controls from Button * fix: clear the open-cell flag when the drawer reopens * chore: refresh the sqlx cache * fix: read a run's configuration and its version from one snapshot * fix: refuse a dataset path or summary the column cannot hold * refactor: handle the agent draft the way the resource editor does * fix: run only a configuration the launch actually read * docs: bound dataset path and summary where they are submitted * fix: surface a stalled agent draft instead of claiming it is kept * fix: stop claiming a draft holds edits a failed write never sent * fix: word a missing score only once the run says whether the case answered * fix: let a breadcrumb crumb shrink so its truncation applies * docs: describe where an agent's unsaved edits live and what drops them * fix: keep harvesting scores when the run cannot yet word a missing one * fix: report a refused draft write the card was reading as a save * fix: drop the refused draft write when the server copy is taken instead * refactor: build the scorer and dataset pickers from the design system * fix: say what removing a scorer column actually does * fix: drop a refused draft write wherever the server copy is read * fix: let a picker row be as tall as the two lines it holds * docs: record what removing a scorer column does to recorded runs * fix: send a queued draft write before reopening, and drop only what it refuses * refactor: write the agent draft at commit points instead of mirroring keystrokes Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor: run an agent's edits from the step instead of keeping them as a draft Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: make the diff badge keyboard operable and refuse an edits run without its edits Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: drop the dataset icon from the scorer picker rows Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: size the evals buttons like the rest of windmill and call a run of edits edits Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: count a brain expression as an edit of the linked agent * fix: cap scorers per dataset and report a launched run as launched * fix: harvest scores in one read, refuse duplicate case ids, allow group paths * fix: mint scorer ids server-side, save a dataset edit in one request, check attachments * fix: write a dataset edit and its cases in one transaction * fix: atomic dataset create/edit, reset eval pane per agent, stable pending scorer ids * refactor: govern eval_case writes by RLS so a dataset edit is one transaction * fix: pin launch snapshot, order case locks, cap dataset size, guard stale load * fix: cap dataset bytes on single-case writes, reset run-dialog flag on load failure * feat: migrate eval datasets on username change, settle unspawned cases, drop unused case endpoints * fix: resolve scorer scripts as the caller and pin their hash; migrate scorer paths on rename * fix: bound a failed tool call's error to the payload truncation cap * fix: pin scorer hash as a hex string, reject missing judges, migrate eval authorship * fix: record an out-of-range scorer result as an error, not a score * fix: resolve judges in one caller-scoped read, pin deployed scripts, bound pass_if * fix: settle unspawned cases only when the run completes, and their score cells too * feat: reassign eval datasets and their path references when offboarding a user * fix: use the regex backreference in offboarding eval path rewrites * fix: register eval datasets in offboarding registries, keep resource-version param name * refactor: name the resource-version path param id, since it is the row id not the version * fix: validate dataset paths canonically, clone eval data on fork, surface eval load and launch failures * docs: note MCP tool results are not yet surfaced to eval scorers * fix: show the eval error state on any load failure, not only an empty dataset list * fix: preserve eval case order across a batched save Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * docs: scope the eval launch delete-safety guarantee to the assembly window Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: only offer deployed scripts as eval scorers, drop unbuilt rescore claim Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: enforce 0-1 scorer threshold in the settings drawer and clear stale eval load errors Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: scope subject version/hash reads to the caller and keep a 0 pass threshold Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: select the saved dataset when creating or renaming from the Run dialog Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: gate eval dataset rename on path ownership, not just write access Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: tolerate a malformed agent config when resolving the deployed label Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * refactor: trim eval code and comments, fix shared select and modal paths Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: drop the rename warning when editing an eval dataset path Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: add eval dataset delete, keep summary on partial edits, settle resultless scorer cells Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: cover parseThreshold and subjectLabel Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: hold dataset Save during a scorer write, derive draft_hash only from the carried draft Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c09de594b6 |
feat: version resource values with history, diff and restore (#10596)
* feat: version resource values with history, diff and restore Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: record resource versions in a trigger so direct writes are covered Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: show the selected version's value and tighten history write access Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * perf: gate resource version recording in trigger WHEN clauses Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: clear a resource's past versions, and address review nits Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: restore the displayed version and keep author attribution on pooled writes Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: scope history to the selected workspace and gate clearing on ownership Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: gate restore on write access and clearing on the signed-in workspace Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(frontend): share the version-history row between script and resource drawers Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * perf: trim resource version history in the monitor sweep, not on write Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(frontend): match the script versions drawer shell for resource history Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * perf(frontend): highlight version values instead of mounting monaco Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(frontend): match the script drawer's code preview presentation Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: rank version trim in one windowed pass instead of a correlated delete Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(frontend): treat the newest version as current by position Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * perf: gate the resource version trim to an hourly sweep Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: unnest the version row action and correct the trim cadence docs * perf: cap the history listing and use sets for reference lookup * feat: warn when a resource is written more than 60 times a minute * fix: lower the resource write advisory to 20 per minute * fix: discard stale history loads and never diff against an unread value * fix: correct the write advisory boundary and document the eviction lock * fix: read history and the live value from one snapshot * refactor: read the drawer's diff baseline from versions, not the live resource * fix: open the history drawer with no version selected * fix: disarm the clear confirmation and clear the pane when the selection moves * fix: explain the missing diff and drop a guard that can no longer fire --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3b95a2d096 |
feat: reusable AI agent steps with rigid linking and edit/fork (#9825)
* feat: reusable AI agent steps with hybrid linking and evals Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: make linked AI agents rigid (read-only) with unlink-to-fork Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: show inherited agent config read-only on linked step Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: edit/update a saved agent in place via upsert Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: bind linked AI agent tool inputs to host flow context Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: rebind linked AI agent tool inputs via graph tool nodes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: linked AI agent tool nodes, step test, and read-only card Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: remove ai_agent resource type migration, sync from hub instead Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: remove AI agent eval suite and run endpoint, defer to later Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: unwire eval routes, types and UI (completes eval removal) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: update reusable AI agents guide for eval removal and tool rebinding Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: regenerate system prompts for AIAgent agent/tool_inputs schema Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: strip brain transforms on link, avoid dirtying flow on tool open Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: flow-local test form and linked-agent marker in read-only graph Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: store linked tool overrides as diff from resource base Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: resolve linked agent tools in read-only viewer with fallback Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: use operating workspace, block non-static provider, warn on unbound tool inputs Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: resolve linked parent's tools from resource for nested agent tool lookup Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: scope linked-agent tools by flow path, thread workspace to path check and embedded viewer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: strip flow-context tool inputs on agent save, drop unbound-inputs warning Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: persist agent edit mode across tool selection, show linked tool code read-only Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: show linked agent resource path in node definition panel Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor: edit linked tool inputs in step panel, make tool nodes display-only Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor: wire step-panel tool bindings (completes display-only pivot) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: single scroll for linked card, agent path as node label, drop fill-inputs in tool cards Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * style: align linked-agent UI with design tokens and components Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: separate linked tool select target from module id to unbreak agent clicks Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@aanthropic.com> * fix: save agent tool inputs verbatim, host flows override via tool_inputs Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: scope agent edit state by flow path, require linked-tools scope at init Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: block saving an agent whose static provider is incomplete Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: type errors in agent tool bindings and save drawer input Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: key agent edit state by workspace, resync tool bindings on external changes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: include workspace in linked-tools scope and tool schema fingerprint Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: remove unused workspace prop from FlowModuleSchemaMap Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor: drop linked-agent placeholder tool node, path label suffices Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: workspace-qualified resource links, guard stale tool schema loads Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: keep flow tool overrides out of the agent on edit, fold only on unlink Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: fold preserved tool overrides into the step on edit cancel Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: refuse overwriting non-agent resources on save, show memory kind on linked card Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: consume picker value, invalidate edit state on undo/reinit, cap nested agent tools Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: guard in-flight edit fork against restores, migrate edit state on rename Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor: validate agent edit state by fork identity instead of path keys Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: key agent edit entries by fork marker alone, immune to editor nesting Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: keep agent edit state across structural graph edits and flow renames Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B6kq9PYqNdc5q7ubidBYAs * fix: centralize agent edit reanchor, guard in-flight saves, seed rename scope from flow path Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B6kq9PYqNdc5q7ubidBYAs * fix: ancestry-keyed edit reanchor and doc-scope sweep for republished linked tools Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B6kq9PYqNdc5q7ubidBYAs * fix: guard stale linked-tool fetches and resolve while-loop nested linked agents Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B6kq9PYqNdc5q7ubidBYAs * fix: drop empty tool override entries on revert and correct stale viewer comment Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B6kq9PYqNdc5q7ubidBYAs * docs: drop stale eval mention from the linked-agent comment Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: deploy linked agent resource, guard viewer fetches, align tools schema Address review findings on the reusable-agent branch: - Cross-workspace deploy never collected a linked step's `agent` resource, so the deployed flow failed at runtime unless the agent already existed there. - The read-only viewer published resolved tools without the generation guard flowState uses, letting a superseded link's tools win a race. Share one guarded publisher (`publishLinkedAgentTools`) between both call sites. - `tools` was still required in the OpenFlow AiAgent schema while the deserializer defaults it, rejecting hand-authored linked steps; make it optional and narrow the call sites. - Overlay `tool_inputs` in the non-linked branch too, so a flow persisted while a step sits in "Editing" mode still binds tools to this flow. - Cap the linked-tools store's scope map; nothing evicted it before. - Drop the orphaned `.sqlx` entry left by the eval removal, regenerate the copilot OpenFlow schema, and fix the generator's nested-`z.record` arity. - Move `refreshFlowStateStore` out of `agentEditStore` into its own module. - Document that linked agents' tool scripts are outside the lock pipeline. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: regenerate system prompts for optional AIAgent tools Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: follow saved-agent deps on deploy, accept the linked shape in the schema Round-18 review findings: - Deploying a linked flow queued only the outer ai_agent resource. Follow `$res:` refs inside a resource value (every UI-saved agent has a provider resource) and the agent's own tools, which reference scripts, flows, MCP resources and nested linked agents by bare path. - The AiAgent input_transforms schema still required provider/output_type, so it rejected the very shape linking persists (brain transforms stripped, flow-local inputs kept). Only user_message is always present. - dfs traversed `value.tools` unconditionally through a cast, which throws on a linked module that omits it now that the field is optional. - Trim the flow-refresh invariant comment to the 4-line limit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: recurse into inline nested agent tools on deploy, require provider when unlinked Round-19 review findings: - The deploy walk only inspected a saved agent's top-level tools, so an inline nested agent tool's own scripts, flows and MCP resources were skipped. Recurse into it; a linked one is still queued as a resource instead. - Normalize a `$res:`-prefixed MCP tool resource_path like other refs. - Dropping provider/output_type from the schema's required list also let a standalone providerless agent validate, which deploys clean and then fails on every run. The constraint can't go in the schema: an `anyOf` makes AiAgent a union, which breaks the FlowModuleValue discriminated union it belongs to (verified: zod throws "Invalid discriminated union option"). Enforce it in validateFlowModules instead, next to the other cross-module checks, via a shared collectProviderlessAgentIds. - Correct the deploy paragraph in the docs: provider resources are traversed now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: follow linked tool_inputs overrides on deploy, untrack vitest artifact Round-20 review findings: - A linked step's `tool_inputs` override replaces the resource tool's default at runtime, so a static `$res:`/`$var:` override is the dependency the flow actually uses. The deploy walk queued only the saved agent, leaving runs in an empty target workspace to fail on the missing override target. It also never scanned an aiagent module's own input_transforms, since the scan was gated to script/rawscript/flow. - Extract the pure walkers to deployDependencies.ts and cover them: three rounds have each found a further gap in this one function. - Untrack a vitest cache artifact committed by accident, and ignore a repo-root node_modules/ (only per-package paths were listed). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: collect inline agent provider and tool deps, correct tool_inputs docs Round-21 review findings: - An inline agent's provider credential sits inside an object-valued static transform, so the top-level string check missed it and such a flow deployed without its provider. Walk transform values instead of string-matching them. - An inline agent's own tools were only partly reachable: getAllModules drops MCP and websearch tools, so their resources were never queued. A standalone agent module now recurses through agentResourceDependencies, and the module's own input_transforms are scanned inside aiAgentModuleDependencies so one function owns the whole step rather than splitting it with the caller. - `tool_inputs` was documented as empty/absent for non-linked steps, which contradicts the runtime applying it when `agent` is unset so a flow persisted mid-Edit keeps its bindings. Describe that case in both the Rust doc and the OpenFlow description. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep linked steps brain-free on load, gate stale agent fetches, log linked tools Round-22 review findings: - loadSchemaFromModule filled every AI agent schema key with a placeholder transform, re-adding provider/memory to a linked step that deliberately carries none — persisted on the next save and rejected by the generated Copilot schema. Fill only the flow-local keys when the step is linked. - The linked-resource fetch was neither aborted nor tagged, so switching a step from agent A to B could publish A's tools under B and show A's brain next to B's link. Tag each result with the (workspace, path) it was fetched for and drop the ones that no longer match. - "Test this step" passed no tools for a linked agent, and the log viewer drops tool_call entries it cannot resolve to a definition, so the agent's invocations vanished from the log. Pass the resolved resource tools. - Correct the cancel-edit comment: the runtime does apply tool_inputs on an unlinked step, and folding is what leaves nothing for it to overlay. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: pin the edit session across saves, resolve linked tools in the run viewer Round-23 review findings: - Cancel stays enabled while a save awaits its requests, and it keeps the `tools` array identity, so the old guard passed and the completing save relinked the step and cleared the edits Cancel had just kept. It also accepted any replacement edit marker. Pin the path being saved and require the marker to still hold it, which still tolerates a content-preserving refresh re-anchoring the marker onto a clone. - Resolve linked agents' tools in the run/status viewer too: it reads module.value.tools straight from raw_flow, which is empty for a linked step, so AIAgentLogViewer dropped every tool_call it could not match and the graph drew the agent with no tool nodes. Same gap the previous commit closed for "Test this step" only. - Drop the overlay call-site comment: it claimed resource defaults are discarded and unmatched keys ignored, while overlay_tool_inputs preserves defaults and inserts new keys, as its own test asserts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: scope linked tools without the trigger-node path, keep the standalone save guard Round-24 review findings, both regressions from the previous commit: - Passing `path` to the run viewer's graph also switched on its Trigger node (`triggerNode ? path : undefined`), which reads a TriggerContext that /run/[...run] does not provide — the page threw "Cannot read properties of undefined (reading 'triggersCount')". Give the graph a separate `linkedToolsPath` for the tools bucket so the two stay independent. - The rewritten save guard tracked only the edit path, so a plain "Save as agent" no longer noticed the step being replaced mid-request (undo, session sync): the replacement has no edit path either, so the stale completion relinked it and stripped its brain. Keep the array-identity check when there is no edit session, and use path re-anchoring only when there is one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep recorded tool calls in run history, send tool_inputs from step previews Round-25 review findings: - The agent log viewer dropped any recorded tool_call whose definition it could not find among the supplied tools, so renaming or removing a tool — or losing read access to a linked agent's resource — erased calls that had actually run. Render the recorded call labelled by its function name; its args, logs and result come from the child job, not the definition. - "Test this step" sent tool_inputs only for a linked step, but a step forked for editing has no `agent` while still carrying the flow's bindings, which the runtime overlays. The preview ran resource-authored defaults instead of the bindings under test. Send them from both branches. - Polling a running flow replaces `job` every tick, so the run viewer re-read every linked agent's resource each time. Key the fetch on the set of linked steps instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: never discard edits made during a save, isolate the run viewer tools bucket Round-26 review findings: - The agent editor stays live while a save is in flight, so edits made after the snapshot were not in the resource yet linking stripped them from the step too, losing them outright. Compare the config against the snapshot on completion and, if it moved, leave the step alone and tell the user to save again. - The run viewer published into the editor's `${ws}:${flow path}` bucket, so opening an older run in the preview pane could flip the edited flow's tool nodes to that run's agent. Key it by job instead. - Drop the now-unreachable undefined filter in the agent log viewer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: claim the linked-tools generation on direct publishes and clears Round-27 review findings: - The step editor wrote resolved tools (and cleared them on unlink) straight into the store, leaving the fetch generation untouched. An older in-flight load for the previous agent then still passed its own check and overwrote them, so the graph and binding editor could show agent A while the step links to B. Claim the generation before those writes. - Correct two comments that still described unmatched tool calls as dropped; they are kept and labelled by their recorded name. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: retain the loaded linked agent, rebuild run logs when tools resolve Round-28 review findings: - Rejecting a superseded resource response left the card with nothing: a late reply for a previous agent replaces `linkedResource.current` and no refetch follows, so the linked step lost its brain, tools and provider warning until remount. Retain the last response that matched the current link instead. - The agent log viewer built its module list on mount only, so a linked agent's asynchronously resolved tools never replaced the placeholders, and switching between completed runs reused the first snapshot. Rebuild on a value key — callers rebuild the agentJob object each render, so tracking its identity would reload in a loop. - Refresh a linked-tools scope's recency when it is read, not only when it is published: a run viewer opens one bucket per nested job, which could otherwise evict the bucket a still-displayed run is using. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: supersede stale log reloads and stale tools on a link change Round-29 review findings, both on the reloads added last round: - Every prop change starts another loadToolCalls, and it awaits child-job requests before writing the shared view, so a slower reload for a previous run could restore its logs and tool states over the run now selected — or replace newly resolved definitions with an earlier empty-tools snapshot. Build the states locally and let only the newest load publish, including the parent's index-keyed job cache. - While a newly linked agent resolves, the previous agent's tools stayed in the store, so its bindings were editable against a step already linked elsewhere, and a failed load left them indefinitely. Clear them once the link moves away from what this component published; tools resolved at flow load are untouched, so selecting a step still doesn't flicker. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: resolve a run's linked agents in the run's own workspace Round-30 review finding: the run viewer fetched linked agent resources with the navigation workspace, but session and fork previews render it with `workspaceId` pointing elsewhere. Those runs resolved nothing — or an unrelated resource sharing the path — losing tool nodes and log definitions. Prefer the explicit override, then the job's own workspace. The store scope stays keyed on `workspace` so it still matches what FlowGraphV2 reads; the job id in the key already makes the bucket unique. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: refetch a run viewer's linked tools if its scope is evicted Round-31 review nit: the viewer publishes one scope per mounted nested job, hidden ones included, so a loop with many loaded iterations can push a displayed scope past the store's cap. Nothing refetched it afterwards — the set of linked steps had not changed — leaving the run without tool nodes or log definitions. Track the store and republish when the bucket is gone; publishing always writes a key, so this settles instead of looping. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: retain in-use linked-tool scopes instead of refetching evicted ones Round-32 review findings. Republishing an evicted scope settles for one scope but not against the cap: with more than 32 mounted nested jobs holding linked agents, restoring one necessarily evicts another, and that mutation reran every viewer's effect — an endless round of resource requests. Hold a scope for as long as a viewer is mounted and skip retained scopes when evicting, so buckets in use are never dropped and nothing has to refetch. The cap yields to correctness when everything mounted is in use. Dropping the publish key also restores refetching when the fetch workspace changes for an otherwise unchanged job and link. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: guard non-static brain edits during save, retain every displayed scope Round-33 review findings: - The in-flight edit guard compared the saved config, which holds only static brain values. A computed system prompt, memory or temperature changed while the save was awaiting the API therefore compared equal, and linking stripped it with no warning. Compare what linking actually discards — every brain transform and the tools — leaving the flow-local inputs free to change. - Retaining run-viewer scopes made them fill the cap, and eviction then picked any unretained scope, including the editor bucket a user is looking at, with nothing to refetch it. Retain the scope each graph draws from for as long as it is mounted, so every displayed bucket is protected. - A failed agent job has no parseable action list; the loader returned early and left the previously selected step's tool tree under the new header. Clear the view instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: resolve only flow modules in viewer scans, prune scopes on release Round-34 review findings: - Both viewer scans used the default dfs, which descends into agent tools, and published each linked agent under its bare id. Tool ids imported from a resource are not flow-global, so a nested linked agent sharing an id with a top-level step superseded that step's fetch and showed its tools instead. Scan flow modules only — the graph resolves the store per module node. - Scopes skipped while retained were never reconsidered, so closing views left the store over its cap for the tab's life. Prune on release too. - Correct two comments that still argued the premises the retain mechanism and the read-recency policy replaced. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: don't report success when a save left the step unlinked Round-35 review nits: - persist warns that changes made during the save are not in the resource and leaves the step alone, but both callers then toasted success unconditionally, burying the only actionable message. Report whether the step was linked. - Condense the tool_inputs invariant to the four-line limit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: seed the published link at mount, keep run history for toolless agents Round-36 review findings: - `publishedFor` started unset, but initFlowState has already published for the step's link by then. A link change landing before this component's own request therefore skipped the clear, leaving the previous agent's tools under the new link — indefinitely if the new one fails. Seed it from the link at mount. - A standalone agent that omits `tools` kept `undefined` here, and the gate downstream then hid the AI message and tool-call history behind the generic result view. Default to an empty list like the other consumers. - A save that lands after the step was replaced writes the resource but leaves the step alone; say so instead of closing the drawer with no outcome. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: qualify nested agent tool store keys, keep an empty tools identity stable Round-37 review findings: - The step editor keyed the linked-tools store by the bare module id for nested agent tools too. Those ids come from a resource and are not flow-global, so a nested linked agent sharing an id with a top-level step read that step's tools — then overwrote them once its own fetch landed. Qualify the key by the parent agent, as the edit store already does; flow modules keep the bare id the graph looks up. - The `tools` binding handed the editor a fresh [] on every read when the module omits the field — a shape this PR made valid — so the save guard's identity check never matched and such a step could never link. Read through one shared empty array instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: accept the first tool on an agent module that omits tools Round-38 review nit: the graph's tool insert required an existing `tools` array, so a module authored without the field — valid since `tools` became optional — swallowed the insert while still pushing history and dispatching a change. Create the array on first use. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: don't evict a scope on the write that created it, and cover the store Round-39 review findings: - A rename removed the retained old key from the order but the new one is not retained until readers re-run, so eviction deleted the fresh bucket immediately. Reorder without evicting; the next publish or release enforces the cap, by which point the new key is held. - Writing the test for that surfaced the same shape in touchScope: it evicts right after appending, so once every older scope is retained the scope just published was the only eligible victim and was dropped at once. Exclude the scope being written. Add the store's first test: retention, eviction past the cap, pruning on release, and the rename handoff — four rounds landed fixes here with nothing pinning the behaviour. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: re-resolve linked agents when a wholesale edit changes the links Round-40 review findings: - Undo/redo, YAML apply, AI apply and session restore swap a step's `agent` without re-running initFlowState, and the step editor only watches the step it is mounted on — so an unselected step kept showing, and binding against, the previous agent's tools. Re-resolve from the editor whenever the set of links changes. - Document that linked resolution is live rather than pinned: an edit landing mid-run affects steps that have not started, and a nested agent tool looks its definition up by id when its own job starts, so it can run a changed definition. Pinning would mean carrying the resolved definition into the child job instead of its id; inline agents are unaffected because their tools are snapshotted with the flow value. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: per-module empty tools identity, invalidate tools when a link is replaced Both findings are over-corrections in the two preceding commits: - The shared empty-tools array made identity stable, but stable everywhere: a wholesale edit that keeps the module id reuses the component, so when both the old and the replacement module omit tools the save guard saw no change and could link and clear the replacement. Hand out one empty array per module value, which a replacement always renews. - The editor's link watcher resolved the replacement agent without dropping the previous one's tools first, so a step selected before the fetch landed still showed agent A under link B — and the freshly mounted editor seeds itself from B, so it could not tell. Clear the entry when the link for a module changes, seeding the map from the graph so the first run doesn't refetch what initFlowState just resolved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: reserve graph space for linked tools, re-resolve only changed links Round-41 review nits: - The layout reservation read the module's own `tools`, which is empty for a linked agent, so its display-only tool nodes were drawn over the node above in read-only viewers. Count the resolved tools for a linked step. - The editor's link watcher refetched every linked agent on each run. Resolve only modules whose link actually changed, and skip the pass entirely on a rename, where the scope sweep has already carried the buckets over. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: protect a renamed scope until it is retained, drop the phantom tool row Round-42 review nits: - Readers release the old scope before retaining the new one, so a migrated bucket is unretained in between and, over the cap with everything else held, was the only thing eviction could take. Protect a just-migrated scope until a reader retains it, and cover that release/retain order in the store test. - The layout reserved an add-tool row for linked agents, which have no add-tool node, leaving dead vertical space. Match computeAIToolNodes. - Re-resolving links no longer short-circuits on a rename: comparing each module still costs nothing when only the path changed, and a restore that renames and relinks in one tick now gets both. - Hoist the duplicated linked-tools lookup in the graph's store update. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: kill a scope's in-flight fetches before migrating it Round-43 review finding: fetch generations are keyed by (scope, module), so a resolution still running against the pre-rename scope keeps a valid generation there. It publishes into the old bucket after the rename, and the doc-scope sweep — which gives the source precedence — carries it forward over a link resolved since under the new scope, leaving the graph and binding editor on the previous agent's tool ids with nothing to refetch them. Invalidate the source scope's fetches before each migration, and pin the behaviour: the new test fails without the invalidation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: re-resolve links a scope sweep cancelled, and only sweep a real bucket Round-44 review findings, both on the previous commit: - Invalidating the source scope killed fetches that were perfectly current — a link still loading when the rename landed — and nothing restarted them, because the watcher already records that link. Resolve again, in the destination, every link the migration left without tools. - The doc-scope sweep ran on every store version bump, so during a draft refresh the first completed fetch cancelled the others mid-flight. Skip the sweep entirely when the source scope holds nothing. - Condense a six-line invariant to the four-line limit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: split rename from doc sweep, hide brain fields of nested linked agents Round-45 review findings: - Two reviewers disagreed about invalidating a scope whose bucket is empty, because the two callers differ. A rename is a cut-off: every fetch still running against the old scope is stale whether or not anything resolved there, so it always invalidates. The doc-scope sweep has no cut-off — those fetches belong to the refresh in progress — so it still waits until that scope holds something. - Recording the swept links as published undid the rename+relink fix: a restore that renames and swaps a link in one tick would keep the previous agent's tools with nothing to refetch them. Leave that comparison to the watcher, which compares links rather than presence. - A nested agent that is itself linked was offered the whole agent schema in the tool bindings, but the runtime overlays only its flow-local inputs, so the rest were collected and dropped. Show what actually applies. - Condense the hybrid-linking comment to the constraint. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: don't resolve a shared agent's tool defaults when loading it Round-46 review finding: the whole agent resource was interpolated before tool_inputs was overlaid, so each tool's default `$res:`/`$var:` resolved first. A host flow overriding a default that points at the author's resource still had to resolve that resource, and an unused tool whose default is unreadable in the consumer's permission context failed the agent outright — defeating the point of sharing an agent across contexts. Read the resource raw, overlay the host's overrides, and interpolate only the brain; each tool resolves its effective inputs when it executes. The nested tool lookup reads raw too, since it only needs definitions. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: interpolate the brain before overlaying caller inputs Round-47 review findings, all on the previous commit: - user_message and user_attachments were inserted before interpolation, so they went through it a second time: a user message of `$WM_TOKEN` expanded to the job token and was sent to the model provider. Interpolate the resource first, then overlay the already-resolved flow-local inputs. - The relink watcher skips tool nodes, so a linked agent nested as a tool kept the previous agent's entry through undo, YAML/AI apply or a session restore, and the step editor seeds itself from the new link and cannot tell. Emit the ancestry-qualified key for those too. - Correct the guide, which still named the interpolation path this branch replaced, and condense two invariants to the four-line limit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: deploy $jsonvar deps, key run logs by tool identity, seed only top links Round-48 review findings: - The deploy walkers recognised `$res:` and `$var:` but not `$jsonvar:`, which the worker resolves too, so a secret referenced that way by an agent brain, a saved tool default or a host override never reached the target workspace. - The run log rebuilt only when a tool's name or the tool count changed, so a refreshed resource that altered a tool's path, code or id behind the same name kept showing the old definition. Key on the array identity instead: the store swaps it exactly when the contents differ. - Nested linked agents were seeded as already published, but initFlowState resolves only top-level links, so their tools never loaded until their editor was opened. Seed what initFlowState actually publishes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: let the watcher's fetch survive the step editor's stale-clear Round-49 review nits: - On a relink the step editor claimed the fetch generation before clearing the previous agent's tools, which discarded the watcher's already-running fetch for the new link. The tool nodes then only appeared if the step stayed selected until the editor's own refetch landed. Clear without claiming: the watcher superseded the old fetch when the link changed, so nothing stale can return. Unlink still claims, since no watcher fetch covers it. - Condense the store's opening invariant to the four-line limit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: condense the stale-clear invariant Round-50 review nit. Also records why the branch deliberately doesn't claim a fetch generation: a reviewer asked for the opposite this round, but writing `agent` re-runs the editor's watcher, which supersedes the old fetch and starts one for the new link — claiming here would discard it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: guard Edit/Unlink by step identity, not just the link path Round-51 review finding: forkFromResource compared only the agent path after its fetch, so a module replaced mid-request while keeping the same link passed the check — the stale continuation then wrote the fetched brain and tools into the replacement and unlinked it. Compare the step's own `tools` array too, which is one instance per module value and so identifies the step. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: report an Edit or Unlink abandoned because the step changed Round-52 non-blocking note: forkFromResource returns undefined when the step was replaced mid-request, and both callers treated that as do-nothing, so the click looked ignored. Say what happened, as the save path already does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: hugocasa <hugo@casademont.ch> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@aanthropic.com> |