From c2b5808c7e71fc59cc4aed38b6534f7da5661ad3 Mon Sep 17 00:00:00 2001 From: centdix Date: Thu, 9 Apr 2026 12:28:52 +0200 Subject: [PATCH] docs: align ai eval benchmark docs Co-Authored-By: Claude Opus 4.5 --- ai_evals/README.md | 118 +++++++-- docs/system-prompt-testing-plan.md | 6 + docs/system-prompt-testing-status.md | 377 ++++++--------------------- 3 files changed, 182 insertions(+), 319 deletions(-) diff --git a/ai_evals/README.md b/ai_evals/README.md index d5bfca4789..e242371652 100644 --- a/ai_evals/README.md +++ b/ai_evals/README.md @@ -1,25 +1,17 @@ # AI Evals -Minimal benchmark runner for the four Windmill AI generation modes: +Small benchmark runner for the four Windmill AI generation modes: - `cli` - `flow` - `script` - `app` -Each case is just: - -- a `prompt` -- an optional `initial` fixture -- an optional `expected` fixture - -For `flow` mode, an `initial` fixture can also include a benchmark workspace catalog of -existing scripts and flows. That lets the real `search_workspace` / -`get_runnable_details` tools discover reusable building blocks during evals. +The benchmark always tests the current production prompts, tools, and guidance in this checkout. Each attempt runs: -1. the real production prompt/tool/guidance path +1. the real production path 2. deterministic validation 3. LLM judging @@ -30,14 +22,21 @@ cd ai_evals bun install ``` -Frontend runs also require frontend dependencies: +Frontend modes also require frontend dependencies: ```bash cd frontend bun install ``` -## CLI +## Commands + +List model aliases: + +```bash +cd ai_evals +bun run cli -- models +``` List cases: @@ -47,32 +46,107 @@ bun run cli -- cases bun run cli -- cases flow ``` -Run a mode: +Run benchmarks: ```bash cd ai_evals bun run cli -- run flow -bun run cli -- run flow flow-test5-simple-modification --runs 3 +bun run cli -- run flow flow-test4-order-processing-loop --model opus +bun run cli -- run flow flow-test0-sum-two-numbers --runs 3 --verbose bun run cli -- run cli bun-hello-script ``` -`run` always writes a JSON result file under `ai_evals/results/` unless you pass -`--output`. +Public CLI surface: -It also writes generated artifacts next to that summary file, for example: +- `models` +- `cases [mode]` +- `run [caseIds...]` -- summary: `ai_evals/results/2026-04-08T13-00-00.000Z__flow.json` -- artifacts: `ai_evals/results/2026-04-08T13-00-00.000Z__flow//attempt-1/flow.json` +`run` options: + +- `--runs `: repeat each case `n` times +- `--output `: custom result JSON path +- `--model `: choose the model under test +- `--verbose`: stream assistant output for frontend runs + +## Models + +Use `bun run cli -- models` to see the current aliases. + +Today: + +- `haiku` +- `sonnet` +- `opus` +- `4o` + +Notes: + +- the command also prints accepted alias spellings such as `gpt-4o`, `claude-opus-4.6`, and `claude-haiku-4.5` +- frontend modes (`flow`, `script`, `app`) can use Anthropic and OpenAI-backed aliases +- `cli` mode always uses the Anthropic agent SDK, so only Anthropic aliases are valid there +- the judge model is separate and currently defaults to `claude-sonnet-4-6` + +## Case Format + +Cases live in one YAML file per mode under `ai_evals/cases/`. + +Minimal shape: + +```yaml +- id: flow-test0-sum-two-numbers + prompt: |- + Create a flow that takes two numbers, `a` and `b`, and returns their sum. + initial: ai_evals/fixtures/... + expected: ai_evals/fixtures/... +``` + +Optional fields: + +- `initial`: starting state fixture +- `expected`: expected artifact fixture +- `validate`: extra deterministic validation rules + +For `flow` mode, `validate` can express requirements such as: + +- accepted input schema shapes +- required `results.*` reference validity +- required module/code/input characteristics + +For `flow` mode, an `initial` fixture can also include a benchmark workspace catalog of +existing scripts and flows. That lets the real `search_workspace` and +`get_runnable_details` tools discover reusable workspace runnables during evals. + +## Results And Artifacts + +Every run writes: + +- a summary JSON under `ai_evals/results/` +- generated artifacts in a sibling directory + +Example: + +- summary: `ai_evals/results/2026-04-09T09-40-33.051Z__flow.json` +- artifacts: `ai_evals/results/2026-04-09T09-40-33.051Z__flow/` + +Typical artifacts by mode: + +- `flow`: `flow.json` +- `script`: `script.json` plus the generated script file +- `app`: `app.json` plus frontend/backend files +- `cli`: `assistant-output.txt` plus generated workspace files ## Layout - `cases/`: one YAML file per mode - `fixtures/`: initial and expected fixtures -- `core/`: shared case loading, validation, judging, and result writing +- `core/`: shared loading, model resolution, validation, judging, and result writing - `modes/`: one runner per mode +- `results/`: local benchmark output and artifacts ## Notes - Frontend modes reuse the production frontend chat code through the Vitest bridge. -- CLI mode creates an isolated workspace, writes the current checkout guidance into it, and benchmarks the real skills / AGENTS flow. +- CLI mode creates an isolated workspace, writes the current checkout guidance into it, and benchmarks the real skills / `AGENTS.md` flow. - Frontend progress streams live while the benchmark is running. +- Deterministic validators should stay focused on real correctness constraints, not one exact implementation shape. diff --git a/docs/system-prompt-testing-plan.md b/docs/system-prompt-testing-plan.md index 97e94c4615..9b12f1c5e0 100644 --- a/docs/system-prompt-testing-plan.md +++ b/docs/system-prompt-testing-plan.md @@ -1,5 +1,11 @@ # System Prompt And Skill Output Testing Plan +Historical note: + +- This file is a planning document and no longer matches the current benchmark CLI in every detail. +- The current source of truth is [ai_evals/README.md](/home/farhad/windmill__worktrees/prompt-testing-plan/ai_evals/README.md) and [system-prompt-testing-status.md](/home/farhad/windmill__worktrees/prompt-testing-plan/docs/system-prompt-testing-status.md). +- In particular, the current tool no longer has the old variants, compare, or history workflow described below. + ## Goal Build a single testing strategy that answers one question reliably: diff --git a/docs/system-prompt-testing-status.md b/docs/system-prompt-testing-status.md index c8d72c6464..f093e1cbaa 100644 --- a/docs/system-prompt-testing-status.md +++ b/docs/system-prompt-testing-status.md @@ -1,337 +1,120 @@ # System Prompt Testing Status -This document tracks implementation progress against -[system-prompt-testing-plan.md](/home/farhad/windmill__worktrees/prompt-testing-plan/docs/system-prompt-testing-plan.md). +This document describes the benchmark tool that exists today. It is the current +truth for `ai_evals/`. -It is intentionally operational: - -- what has been implemented already -- what is only scaffolded -- what is still missing -- what should be built next - -## Current Summary - -The suite has made meaningful progress on the CLI-first path from the plan. - -What is true today: - -- there is one repo-level benchmark CLI under `ai_evals/cli` -- the CLI surface is the first real benchmark adapter behind it -- frontend flow, app, and script are now exposed through that same benchmark CLI -- frontend AI benchmark logic now lives under `ai_evals`, not under `frontend/__tests__` -- frontend benchmarks reuse production prompt builders, tool definitions, and `runChatLoop` -- CLI cases are now skill-sensitive rather than only artifact-sensitive -- the benchmark harness and `wmill init` now share the same AI-guidance writer -- variants can be frozen as named snapshots through the benchmark CLI -- repeated-run CLI benchmarking exists through `--runs` -- basic CLI reliability metrics now exist -- benchmark history writing and reading now exist in the benchmark CLI -- the first official git-tracked `frontend-flow` benchmark snapshot has been recorded -- the CLI base is strong enough to stop blocking the frontend adapter - -What is not true yet: - -- frontend still has only baseline prompt variants wired by default -- frontend still uses a thin Vitest bridge for module/runtime loading -- frontend `script` only has initial benchmark coverage today -- frontend `script` still treats `test_run_script` usage as a diagnostic instead of a hard gate -- CLI still needs broader case coverage and richer efficiency metrics -- token and cost metrics are not implemented -- no UI studio exists yet - -## Comparison To The Plan - -The plan in +The longer planning document in [system-prompt-testing-plan.md](/home/farhad/windmill__worktrees/prompt-testing-plan/docs/system-prompt-testing-plan.md) -defined a seven-phase delivery order. +still contains useful background, but parts of its workflow are now historical +because the old variants/history system was removed. -### Phase 1: Stabilize the benchmark model +## Current Tool -Planned: +There is one repo-level benchmark CLI under `ai_evals/` with three commands: -- shared case schema -- shared result schema -- initial core benchmark set +- `bun run cli -- models` +- `bun run cli -- cases [mode]` +- `bun run cli -- run [caseIds...]` -Status: +Supported modes: -- partially done +- `cli` +- `flow` +- `script` +- `app` -Implemented: +Public `run` options: -- shared case storage under `ai_evals/cases/` -- initial frontend case manifests under `ai_evals/cases/frontend/` -- initial CLI case manifests under `ai_evals/cases/cli/` -- a practical result shape emitted by the benchmark CLI for `run` and `compare` +- `--runs ` +- `--output ` +- `--model ` +- `--verbose` -Still missing: +There is no variant workflow, no benchmark history workflow, and no compare +command in the current tool. -- a formally versioned shared result schema for all surfaces -- a bigger representative benchmark set -- richer shared aggregation/reporting conventions across surfaces +## How It Works -### Phase 2: Build the benchmark CLI shell +Each attempt runs: -Planned: +1. the current production prompts, tools, and guidance from this checkout +2. deterministic validation +3. LLM judging -- repo-level benchmark CLI entrypoint -- `run`, `compare`, and `history` command skeletons -- adapter selection layer -- temporary wiring to the first CLI adapter +Results are written locally under `ai_evals/results/` as: -Status: +- a summary JSON file +- a sibling artifacts directory containing the generated flow/script/app/workspace -- mostly done +## Current Architecture -Implemented: +- `ai_evals/cases/`: one YAML manifest per mode +- `ai_evals/fixtures/`: initial and expected fixtures +- `ai_evals/core/`: shared case loading, model resolution, validation, judging, and result writing +- `ai_evals/modes/`: one runner per mode -- repo-level benchmark CLI in [index.ts](/home/farhad/windmill__worktrees/prompt-testing-plan/ai_evals/cli/index.ts) -- `list-cases` -- `list-variants` -- `run` -- `compare` -- `snapshot-variant` -- `history` -- CLI adapter selection through `--surface cli` -- frontend adapter selection through `--surface frontend-flow|frontend-app|frontend-script` +Execution model: -Still missing: +- `flow`, `script`, and `app` reuse the production frontend chat loop and production tool definitions through the frontend Vitest bridge +- `cli` creates a temp workspace, writes the current checkout guidance into it, and runs the Anthropic agent SDK against that workspace -- no major CLI-shell gap remains here beyond broader ergonomics and docs polish +## Case Model -### Phase 3: Replace the CLI smoke suite with real artifact evaluation +Each case is intentionally small: -Planned: +- `prompt` +- optional `initial` +- optional `expected` +- optional `validate` -- temp-workspace runner -- automatic skill-bundle materialization -- artifact scoring -- repeated-run support -- baseline vs candidate skill-bundle comparison +`validate` is mainly used for stronger deterministic checks where exact fixture +matching would be too strict, especially for `flow` creation cases. -Status: +Examples of current deterministic checks: -- mostly done +- schema contains one of several accepted input shapes +- `results.*` references resolve +- required code/input characteristics exist in some module +- expected workspace files are created in `cli` mode -Implemented: +## Model Selection -- legacy `cli/test-skills` harness removed -- temp-workspace runner in [artifact-eval.ts](/home/farhad/windmill__worktrees/prompt-testing-plan/ai_evals/adapters/cli/artifact-eval.ts) -- automatic guidance materialization into the temp workspace -- required artifact checks for CLI cases -- required skill invocation checks for CLI cases -- required “next commands” guidance checks for CLI cases -- baseline vs candidate comparison through `compare` -- named variant snapshotting through `snapshot-variant` -- repeated-run support through `--runs` -- aggregate CLI reliability reporting: - - pass rate - - average duration - - average assistant messages - - average tool calls - - average skill invocations - - aggregated required-check failures +Model aliases are resolved through a shared registry in `ai_evals/core/models.ts`. -Most important proof point: +Current aliases: -- a deliberately broken skill variant produced a meaningful regression: - - baseline: `2/2` CLI cases passed - - broken variant: `0/2` CLI cases passed +- `haiku` +- `sonnet` +- `opus` +- `4o` -Still missing: +Notes: -- richer case coverage beyond the current initial script and flow cases -- richer efficiency metrics such as token and cost reporting +- the `models` command also shows accepted alias spellings such as `gpt-4o` and `claude-opus-4.6` +- frontend modes can use Anthropic and OpenAI-backed aliases +- `cli` mode is Anthropic-only because it runs through the Anthropic agent SDK +- the judge model is separate and currently defaults to `claude-sonnet-4-6` -### Phase 4: Add shared reporting and benchmark history around the CLI path +## What Is Working Well -Planned: +- one simple local benchmark CLI +- real production execution paths instead of synthetic prompt variants +- local result and artifact persistence by default +- live frontend progress output +- reusable flow/script/app/cli runners under one tool +- deterministic validation can now catch real runtime-invalid flow wiring -- baseline vs candidate reports -- pass-rate summaries -- worst-failure reports -- official run schema -- git-tracked benchmark summary file -- history snapshot writer -- rollup generation for trend charts +## What Still Needs Work -Status: +- broader case coverage across all four modes +- stronger deterministic validators for more cases, especially app/script semantics +- clearer per-case validation metadata as the corpus grows +- CI automation for smoke and nightly runs -- mostly done +## Recommended Next Focus -Implemented: +The next high-value work is: -- history scaffold under `ai_evals/history/` -- official run schema scaffold in `benchmark-run.schema.json` -- `summary.jsonl` -- rollup placeholders -- shared history writer under `ai_evals/history/writer.mjs` -- snapshot writer script under `ai_evals/scripts/append-official-run.mjs` -- benchmark CLI wiring for `compare --write-history` -- benchmark CLI `history --view latest|summary|surface|variant|model` -- official run snapshots generated from repeated CLI benchmark results -- history rollups rebuilt from real benchmark writes -- shared history writing path generalized so frontend surfaces can use the same official run format -- frontend history writes validated through `compare --write-history` on `frontend-app` -- first official tracked `frontend-flow` snapshot written into `ai_evals/history/` - -Still missing: - -- pass-rate summaries across repeated runs -- worst-failure reporting -- richer history filtering and reporting around those snapshots - -### Phase 5: Finish the frontend black-box harness on top of the shared model - -Planned: - -- convert current flow and app evals into proper scored reliability tests -- add script eval support -- add repeated-run support -- add prompt-variant loading from files -- align frontend outputs with the shared result and history format -- expose frontend runs through the same benchmark CLI - -Status: - -- partially done - -Implemented: - -- shared frontend case scaffolding exists -- frontend benchmark fixtures were moved under `ai_evals/fixtures/frontend/` -- repo-level benchmark CLI support for `frontend-flow`, `frontend-app`, and `frontend-script` -- ai_evals-owned frontend benchmark runner and adapter executed through the frontend Vitest environment -- frontend benchmark helpers now write to temp workspaces instead of mutating in-memory UI state -- repeated runs for frontend flow/app/script through the shared benchmark CLI -- frontend prompt variants loaded from files under `ai_evals/variants/frontend/` -- frontend run/compare payloads aligned to the shared benchmark result shape -- frontend judge score aggregation wired into compare output and official run generation -- frontend official history writes validated through the shared benchmark CLI -- legacy frontend AI chat eval tests removed from `frontend/src/lib/components/copilot/chat/__tests__` -- frontend script benchmark helper now exercises the production script tools with a file-backed script workspace -- validated a real `frontend-script` case end to end through `run` and `compare` - -Still missing: - -- more than the baseline prompt variant for frontend flow/app/script -- broader frontend reliability coverage across more cases in routine use -- stronger deterministic validators, especially for script execution semantics and deeper app/flow checks -- stronger workflow-policy checks, such as consistent `test_run_script` usage in frontend script mode -- regular use of frontend history writes in routine benchmark workflows -- optional future cleanup to remove or reduce the thin frontend Vitest bridge if production chat modules become directly importable - -### Phase 6: Add CI tiers - -Planned: - -- fast PR smoke benchmark -- fuller nightly benchmark -- official history updates on `main` and scheduled runs -- manual benchmark mode for prompt authors - -Status: - -- not done - -Implemented: - -- local manual benchmark mode exists through the benchmark CLI - -Still missing: - -- PR CI integration -- nightly integration -- automated official history writes - -### Phase 7: Build the UI studio - -Planned: - -- run selector -- variant selector -- per-case comparison view -- artifact diff view -- reliability dashboard -- trend dashboard backed by git-tracked benchmark history - -Status: - -- not started - -This matches the plan’s intended order. The UI should still come last. - -## What Was Actually Done - -The most important implemented changes so far are: - -- Added the repo-level benchmark CLI in `ai_evals/cli` -- Removed the old duplicate CLI harness so there is one benchmark entrypoint -- Added CLI `run`, `compare`, `list-cases`, and `list-variants` -- Added CLI `snapshot-variant` to freeze a named candidate bundle -- Added initial CLI benchmark cases: - - `bun-hello-script` - - `bun-hello-flow` -- Made CLI evals skill-sensitive instead of allowing silent skill bypass -- Added repeated-run CLI benchmarking with aggregate reporting -- Added frontend flow/app benchmark adapters behind the shared benchmark CLI -- Added the frontend script benchmark adapter behind the shared benchmark CLI -- Added frontend run/compare support with shared aggregation and judge score reporting -- Validated frontend `run`, `compare`, and `compare --write-history` through the shared benchmark CLI -- Moved frontend AI benchmark fixtures and runner ownership into `ai_evals` -- Removed the old frontend AI benchmark test tree -- Kept the production frontend chat loop and tool-definition path as the shared execution core -- Swapped frontend benchmark helpers from in-memory state to temp-workspace files -- Added a file-backed script helper hook so `frontend-script` can benchmark production script tool usage without UI state -- Shared AI-guidance generation between: - - benchmark temp workspaces - - `wmill init` -- Moved `wmill init` testing overrides to internal env vars instead of public flags -- Added docs for variant workflows and benchmark usage -- Added official benchmark history writing and reading - -## What Is Left To Do - -The highest-priority remaining work is: - -1. Expand the frontend and CLI case corpora to cover more real prompt and skill behavior. -2. Add stronger deterministic validators, especially for frontend script/app/flow artifact correctness. -3. Add richer aggregate metrics: - - pass rate - - flake rate - - latency - - tool-call count - - token and cost metrics if available -4. Add stronger official history summaries: - - worst-failure views - - better pass-rate rollups - - more filtering -5. Add CI tiers. -6. Build the UI last. -7. Optionally reduce the thin frontend Vitest bridge once production chat modules are easier to load headlessly. - -## Recommended Next Step - -The best next implementation step is: - -- harden validators and add more frontend/CLI benchmark cases - -Reason: - -- all three frontend surfaces now run through the shared benchmark CLI -- the biggest remaining gap is benchmark depth, not benchmark plumbing -- stronger validators and broader cases will improve signal faster than adding another interface layer - -## Relevant Files - -- Plan: [system-prompt-testing-plan.md](/home/farhad/windmill__worktrees/prompt-testing-plan/docs/system-prompt-testing-plan.md) -- Benchmark CLI: [index.ts](/home/farhad/windmill__worktrees/prompt-testing-plan/ai_evals/cli/index.ts) -- CLI adapter: [artifact-eval.ts](/home/farhad/windmill__worktrees/prompt-testing-plan/ai_evals/adapters/cli/artifact-eval.ts) -- Frontend adapter runtime: [runtime.ts](/home/farhad/windmill__worktrees/prompt-testing-plan/ai_evals/adapters/frontend/runtime.ts) -- Frontend benchmark runner: [benchmarkRunner.ts](/home/farhad/windmill__worktrees/prompt-testing-plan/ai_evals/adapters/frontend/benchmarkRunner.ts) -- Frontend script runner: [scriptEvalRunner.ts](/home/farhad/windmill__worktrees/prompt-testing-plan/ai_evals/adapters/frontend/core/script/scriptEvalRunner.ts) -- Frontend Vitest bridge: [vitestAdapter.test.ts](/home/farhad/windmill__worktrees/prompt-testing-plan/ai_evals/adapters/frontend/vitestAdapter.test.ts) -- CLI variants: [variants.ts](/home/farhad/windmill__worktrees/prompt-testing-plan/ai_evals/adapters/cli/variants.ts) -- Shared guidance writer: [writer.ts](/home/farhad/windmill__worktrees/prompt-testing-plan/cli/src/guidance/writer.ts) -- CLI benchmark docs: [README.md](/home/farhad/windmill__worktrees/prompt-testing-plan/ai_evals/cli/README.md) +1. add more realistic benchmark cases +2. keep simplifying deterministic validators so they check correctness, not one exact implementation +3. add CI only after the local benchmark signal is trustworthy