Commit Graph

3 Commits

Author SHA1 Message Date
Guilhem e866b68cdf feat: surface execution usage in the sidebar and explain what an execution is (#10760)
* feat: surface execution usage in the sidebar and explain what an execution is

Users read "executions" as a job count and are surprised by the real number,
which meters a second of compute. Every place the UI prints an execution count
now says so, and the sidebar carries a usage meter for the quota that will bind
first.

Adds SidebarUsage at the bottom of both sidebar surfaces: a ring in the
collapsed rail, a labelled bar when expanded, and a modal breaking down every
quota. On the free tier it meters the per-user and per-workspace 1000-execution
caps; on a paid plan it meters workspace usage against the executions the
workspace's seats already include.

Item.tooltip was inert on disabled dropdown rows: DropdownSubmenuItem rendered
the info icon inside the disabled button, which swallows hover, and the row's
own title attribute shadowed any wrapper title. Both renderers now fall back to
a wrapper title the way DropdownV2Inner already intended.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep the usage meter tied to the workspace it describes

isPremiumStore held the previous workspace's tier across a switch, which no
consumer noticed while it only gated affordances — the usage meter is the first
surface to render a number from it, and would have shown a paid seat quota for a
free workspace. It is now undefined until the active workspace's tier is known,
and a superseded response no longer writes.

The seat fetch had the same shape: a slow response for the workspace we left
overwrote the current count and stayed wrong until the next switch.

The usage wrapper also carried the padding the brand-mark row used to own, which
shifted the sidebar bottom by 4px on every instance where the meter renders
nothing. The component owns its own padding instead.

Names the collapsed ring for assistive tech, which otherwise saw an unlabelled
button whose only signal was the arc's color.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: scope the usage meter to the billing workspace and a known tier

A fork's usage, tier and bill all resolve to its billing root, but its member
list is deliberately a subset of the root's, so counting fork members metered
root usage against a fork-sized cap and invented billed-seat overages. Seats now
come from the billing root, and the paid meter stays hidden when that root is
not visible from the fork.

The tier was cleared only after the user-store round-trip, so the meter rendered
the previous workspace's tier for the length of it — a free→paid switch showed
the 1000-execution hard cap on a paid workspace, not a race but every time. The
clear now happens before the first await.

Workspace usage had neither guard: a superseded response overwrote the store
permanently, and the meter is the first surface to print that number as its
headline rather than bury it in a dropdown.

The free-tier counters keyed off `!$isPremiumStore`, which reads an unknown tier
as free and flashed the free-tier blocks during a paid-to-paid switch. They wait
for a known tier instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: never render an unresolved execution count as zero

The workspace-usage clear wrote 0, which is a real usage value: an in-flight or
failed fetch rendered as a green "0/1,000" bar, and a rejection left it there for
the session because loadUsage had no failure path. Usage is now undefined until
it resolves, each endpoint is assigned on its own so one failing leaves the
other's number intact, and a quota is listed only once its own usage, tier and
cap are known. The legacy counters show an em dash rather than a fabricated 0.

The fork gates read an unknown tier as not-premium, so clearing the tier on
switch made the fork entry point disappear for the length of the fetch on a
paid-to-paid switch. They hold while the tier is unknown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: read the usage endpoints as numbers, and fall back to the free tier

Both usage endpoints serve text/plain, so the client hands back a string despite
the generated `number` type. Interpolation and arithmetic coerced it, which is
why nothing noticed before, but `toLocaleString` on a string returns it
unchanged — a five-figure count rendered without its thousands separator against
a formatted cap.

A failed tier fetch left the tier unknown for the session, and consumers hold
premium-only affordances through the unknown window so a free workspace kept
offering them. It falls back to the free tier instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep an unknown tier unknown, and refresh the seat cap on demand

Falling back to the free tier on a failed tier fetch fixed the affordance gates
by lying to the meter: a paid workspace's real five-figure usage rendered
against the 1000 hard cap, red, under "jobs stop running for the rest of the
month". The tier stays unknown instead, and the two consumers get what each
needs — the meter hides, while affordances read `maybePremium`, which holds
through the pending window but fails closed once the fetch has failed.

Membership changes elsewhere don't reach this component, so the seat cap could
show an overage against a cap that had since grown. It re-resolves when the
modal opens, which is when the number is read rather than glanced at.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: let anything showing executions re-read them

The counters were written in one place, the root layout, on a workspace change
only — so a tab left open all day showed the count from whenever the workspace
was opened, and the modal-open refresh could only reach the seat cap, leaving a
freshly computed denominator over a stale numerator.

Moves the fetch to lib/usage.ts, next to the stores it writes, so the meter can
refresh both numbers when its modal opens. Seats follow a membership signal that
WorkspaceUserSettings bumps where it already refetches after every mutation, so
the cap stops lagging a role change without either side owning the other.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: order concurrent usage and seat refreshes

The workspace id doesn't order two requests for the same workspace, and both
refreshes can now have two in flight: usage through A→B→A or a modal-open
refresh landing on one already running, seats through a membership bump
arriving mid-request. An older response could win and restore the count it
replaced. Each refresh takes a generation and only writes if it is still the
newest.

The membership signal also fired on a plain read, so opening the users tab made
every consumer re-fetch a list identical to the one it held. It bumps on an
observed change to the member set instead, never on the first read.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: count only billable seats, and order the tier requests

The cap counted every member row, while the backend bills
`NOT disabled AND NOT is_service_account` — a workspace with service accounts
got an inflated included quota, which hides a real overage rather than inventing
one. The seat basis matches `count_paid_seats` now, and the membership signature
carries both fields so enabling or disabling a member re-resolves the cap.

The tier fetch was the one refresh still ordered by workspace id alone, so a
late failure for a workspace could raise the failure flag over a tier a newer
request had already resolved. It takes a generation like the other two.

The membership signature is keyed by workspace: this page survives a workspace
switch, and comparing one workspace's members against another's reported a
membership change where only the workspace had changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: compare the member set only against the same workspace's

Qualifying the signature with the workspace put the workspace inside the value
being compared, so a switch made every comparison unequal and bumped the version
unconditionally — the opposite of the intent, and worse than before the key. The
workspace is the key now, not part of the payload: a different one has nothing
to compare against and re-baselines silently.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: hold the usage and tier fetches in resources

Every one of these values belongs to a workspace but lived in a bare store, so
each writer and reader re-derived "does this still describe what I'm rendering?"
by hand. Nine sites did, and the ones that forgot were most of this branch's
review findings: three stale-workspace overwrites, three A→B→A races, and two
placeholders (`0` executions, `false` tier) that read as real data because an
in-band value was standing in for "not known".

`resource` from runed — which frontend/AGENTS.md prescribes for async data, and
which ~80 files here already use — supplies all three properties as behaviour
rather than convention: a superseded fetch is discarded, the value resets when
its key changes, and loading and error are states instead of magic values. The
seat count keys on the billing root and the membership version, so both a
workspace switch and an added member re-resolve it.

That removes three generation counters, two workspace trackers, and the manual
clear-and-compare around each fetch. What remains is one publish site that
asserts the value still carries the active workspace before it reaches a store.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: order the resource responses that runed does not

The refactor claimed `resource` discards a superseded fetch. It does not: its
only ordering is an AbortController whose signal the generated client cannot
consume, and `current = result` runs unconditionally once a fetch resolves. So a
late answer for a workspace we had left still landed in `current`, and the
publish site — which trusted `current` — cleared the value on screen for the
workspace we were on. That reinstated the races the generation counters had
covered.

`loading` was standing in for the missing ordering, and it cannot: it is also
true during a `refetch()`, when `current` is still the right value. Gating on it
meant every re-read blanked the meter, and clicking it unmounted the modal that
same click had opened, since both sit behind the quota it had just cleared.

Values now carry the scope they describe and `scopedValue` keeps the newest one
matching the active scope, so a superseded answer neither publishes nor erases,
and a re-read leaves the display alone. The account-wide user counter keys on
the account, so a workspace switch no longer clears it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: order responses within a scope, not just across scopes

The tag carried what a value described but not when it was asked for, so two
fetches for one scope — a refetch landing on an in-flight load, or a second
membership invalidation — were indistinguishable and the older won if it landed
last. That left the seat cap reading the pre-change number until the next bump
or switch, which is the stale cap the generation counters had covered.

Widening the tag to the resource key would have fixed it by blanking the bar on
every membership change, so the issue order travels alongside the scope instead:
`tagged` stamps each request as it is issued, and only a strictly newer answer
for the current scope replaces the held one.

The unit tests now cover the same-key case they missed; both new ones fail
against the key-only guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: ignore a user list a newer read has overtaken

`lastSeen` was written unconditionally after the await, so a response for a
workspace already left overwrote the baseline for the workspace on screen. The
next real membership change there then compared against a baseline that was
never taken for it, re-baselined silently, and never bumped
`workspaceMembershipVersion` — leaving the sidebar on the old seat cap. The
`users` assignment had the same hole: an overtaken list could paint over a
newer one.

Both now go through a single check: a read whose issue order is behind the last
applied one is dropped before it touches either.

Also trims the two `scopedValue` docstrings and the membership rationale to the
four lines AGENTS.md allows, and records there that a failed refresh keeps the
last successful value rather than blanking.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: do not claim a plan before the tier resolves

Widening `isPremiumStore` to `boolean | undefined` left `UserMenu`'s `{:else}`
catching the unresolved state: with the tier still in flight, or after the
request failed, a free workspace was told it was on the "Premium plan". Both
branches under that block assert a plan, so the block now renders only once the
tier is known — which also keeps the bordered divider from appearing empty
while it resolves.

Verified against the running instance with the tier stubbed slow: unresolved
shows neither branch, `false` shows the free counters, `true` shows "Premium
plan". Reverting the guard reproduces the wrong label at 300ms.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: pin how a late answer orders against the read that replaced it

Returning to a scope whose earlier read is still in flight is the one case the
guard resolves by scope rather than by sequence, and the suite only covered it
with nothing outstanding. It now covers the late answer itself: it stands while
it is the only value describing the scope, the read issued on returning
supersedes it, and it cannot come back afterwards.

Also gives the meter the explicit `type="button"` the sibling sidebar rows use.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: size the modal's plan button with unifiedSize

`size` is deprecated on `Button`. `unifiedSize="sm"` renders the plan button at
the same height and weight as the modal's own Close button.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: match the plan button to the modal's own action button

`unifiedSize="sm"` is `h-7`, and the Cancel button `Modal` renders beside it is
`px-3 py-[7px]`, i.e. 32px — so the two sat 4px apart. `md` is the unified size
that lands on 32px, which pairs them without putting a deprecated prop back.

Measured both boxes rather than the new one alone: 32px and 32px, same top.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: instrument the execution meter, and bill-align PremiumInfo's seats

The meter's only interaction is opening the modal, so that is what it counts:
`usage_meter/opened`, keyed by the plan tier and the quota that was tightest —
`free:user`, `free:workspace`, `paid:workspace`. The full set is a type next to
the call site so the vocabulary stays readable in one place.

The pair is registered in `FEATURE_USAGE_KINDS` (windmill-ee-private), without
which the post is dropped with a 204 and records nothing. Verified both halves:
the browser posts
`{"feature":"usage_meter","kind":"opened","key":"free:user","value":1}`, the
running EE image drops it because its registry predates the entry, and
`is_recordable_event` accepts it once the entry is there.

`PremiumInfo` computed its seats from an unfiltered user list, so the billing
page counted disabled members and service accounts that `count_paid_seats` does
not bill. Same filter as the sidebar's cap now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: point ee-repo-ref at the usage_meter registration

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: read the member list before the seat rows that depend on it

`loadPremiumInfo` reads `users` after its own await and nothing recomputes the
seat rows when the list lands, so whenever `premium_info` won the race the page
rendered zero developers, zero operators and zero seats and kept them. The list
is now fetched first, and a failure to read it no longer costs the rest of the
page.

Also refreshes the registered-action inventory in `docs/feature-telemetry.md`,
which the new pair makes 21 across nine features.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: scope the seat comment to the counter it matches

The comment claimed parity with the seats actually charged, which nothing in
this repo computes: `count_paid_seats` documents itself as counting provisioned
members rather than billing's active-user population, and the Stripe quantity
is not derived here. What the filter buys is agreement with that counter.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref to c6902ec2c51dc0ce30962afbfab3e456c5d9b831

This commit updates the EE repository reference after PR #735 was merged in windmill-ee-private.

Previous ee-repo-ref: bbc48fae6b73b6d72fe2e125e6003794a4ece167

New ee-repo-ref: c6902ec2c51dc0ce30962afbfab3e456c5d9b831

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
2026-08-20 22:17:15 +02:00
hugocasa 53eb94659b feat(telemetry): extend feature-usage tracking beyond AI features (#10681)
* feat(telemetry): extend feature-usage tracking to long-tail features

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: describe telemetry as product feature usage rather than AI usage

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(telemetry): trim disclosure copy and drop unused pick origin

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(telemetry): count trigger fires per run and key hub picks from hub data

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(telemetry): slugify hub keys and order both writers' upserts

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(telemetry): key native trigger adoption by service so it matches fires

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref for native trigger adoption fix

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(telemetry): move feature-usage collection into the ee crate

* docs: point feature-telemetry at the moved registry and rust writer

* docs: correct the trigger-fire gate comment to match measured step counts

* docs: put the private-build caveat on the verification step

* chore: update ee-repo-ref to f079db9e7962a413b349c4ff8036080894f30771

This commit updates the EE repository reference after PR #725 was merged in windmill-ee-private.

Previous ee-repo-ref: 055adb80416f9339c9a28ae7fbaeadad30d74959

New ee-repo-ref: f079db9e7962a413b349c4ff8036080894f30771

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-08-14 18:50:38 +02:00
hugocasa c3b2275864 docs(agents): rework agent context, fix dev-env docs, vendor skills (#10667)
* docs(agents): scope agent guidance to where it loads

AGENTS.md loads in every session. Three of its sections only ever applied to
one directory, and docs/autonomous-mode.md was unreferenced by anything in the
repo, so none of its content was in effect.

- Move "Verifying Backend Changes" to backend/CLAUDE.md, "Verifying Frontend
  Changes" and "Banned Patterns" to frontend/CLAUDE.md. They now load when
  working under those directories, which is when they apply.
- Update the two cross-references that pointed at the moved sections (pr and
  svelte-frontend skills).
- Delete docs/autonomous-mode.md. Its "don't stop early" half is already in
  .webmux.yaml's oneshot system prompt, which actually loads; its trigger was
  bypassPermissions, which does not imply an absent user; and it restated
  AGENTS.md and the pr skill with copies that had drifted (hardcoded ports,
  relative screenshot paths). Salvaged the UI traps it uniquely documented
  into frontend/CLAUDE.md and dropped the three stale profile references.

AGENTS.md drops ~3.6k characters with no guidance lost.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(agents): guidance for building a feature — reuse, telemetry, live verification

Three recurring gaps, all cases where a pointer existed but nothing triggered
on it.

Component reuse. The svelte-frontend skill documented three components with
props, which reads as the whole catalog; the barrel exports 23 and common/ has
34 subdirectories against those 23. So "never use raw HTML elements" was an
instruction agents could not follow. Added a mandatory discovery step: read
the barrel, grep the tree, and treat the documented three as examples.

Brand guidelines. frontend/brand-guidelines.md is 34k characters referenced by
bare path, which nothing opens speculatively. Added a table mapping what you
are building to the section that governs it, entered with grep rather than a
full read.

Product telemetry. feature_usage has 14 registered actions across three
features, and an unregistered (feature, kind) pair is dropped by
valid_feature_usage_event with a bare continue — no error, still a 204 — so
frontend-only instrumentation silently records nothing. New
docs/feature-telemetry.md carries the criteria for when to instrument, the
four-step recipe including the allowlist and the InstanceSettings disclosure,
and the privacy rules. Raised in the plan for user-facing work, not as a
separate question, and not at all for bugfixes or refactors.

Also: validation now ends at exercising the change on the running instance,
with standing permission to spin up whatever that takes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(dev): correct the worktree dev-environment guidance

Several things agents were told to do did not match what the machine does.

- Env discovery pointed at .env / .env.local / backend/.env. In a webmux
  worktree the real values are in $(git rev-parse --git-dir)/webmux/runtime.env
  (BACKEND_PORT, FRONTEND_PORT, DATABASE_URL, CARGO_FEATURES, WM_DB_NAME),
  sourced by every pane and undocumented. Reading it is also not blocked by the
  Read(**/.env) deny rules, which the old instruction walked straight into.
- The database name rule said branch-with-underscores. worktree-common.sh uses
  the worktree directory basename, and Postgres truncates at 63 characters, so
  branch hugo/win-2340-… resolves to windmill_win_2340_…_and_eval with no hugo_
  prefix and the tail chopped. A wrong DATABASE_URL guts the sqlx cache.
- The restart procedure said "tmux pane 1" and sent keys to an undefined
  <pane1>. Pane 1 is the backend under the full profile and the frontend under
  frontendOnly. Replaced with finding the pane by pane_current_command,
  recovering the live feature set from the running process (CARGO_FEATURES in
  runtime.env only records what the pane started with), and restarting in place.
- Added recovery for an orphaned backend holding the port: it reparents to
  systemd when its shell dies, so it survives anything that looks like cleanup.
  Three checks before killing a single pid, because pkill -f windmill takes out
  every sibling worktree.
- Agents spawned their own servers because AGENTS.md opened by telling them to.
  Now it checks for the existing panes first; the spawn commands are scoped to
  a plain checkout.
- New EE worktrees branched from the EE repo's local main, which nothing
  fast-forwards, so they started behind the commit pinned in
  backend/ee-repo-ref.txt — the one CI builds against. They now base on the pin,
  falling back to main only when it is unreadable.
- Enabled webmux autoPull so local main stays current; new worktrees are
  branched from it. Documented what WM_CLONE_DB does, including that it
  terminates every connection to the base windmill database.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(skills): vendor grilling/architecture skills; tighten PR ready and review rounds

Vendors five skills from https://github.com/mattpocock/skills (MIT, pinned at
84fdeffd12f2ee307994d1eb6feb48173b6e0502). They are one dependency closure:
grill-me is a stub that runs grilling, and improve-codebase-architecture draws
its vocabulary from codebase-design and its CONTEXT.md upkeep from
domain-modeling. .agents/skills/UPSTREAM.md records the license, the pin, and
the four local deltas so a refresh stays a diff:

- flattened the upstream engineering/ and productivity/ split
- rewrote bundled-file links to repo-root paths, since relative links break
  when read through the .claude/skills symlink
- dropped the upstream agents/openai.yaml packaging metadata
- removed every ADR path. This repo has not adopted ADRs, and a skill that
  offers to create them is how the practice arrives by side effect rather than
  by decision.

PR workflow changes, all in the pr skill:

- A round that never starts is usually a conflict with main, not a CI outage.
  Resolve by merging, not rebasing — a rebase rewrites the head SHA that round
  verdicts and the clean-round marker are keyed to. If the merge advances
  backend/ee-repo-ref.txt, the EE worktree has to follow or
  cargo check --features private compiles a tree neither the author nor CI
  intends.
- A clean round no longer means an automatic flip to ready. Wide blast radius
  (*_ee.rs, migrations, OpenAPI or the generated client, auth paths, shared
  worker infrastructure, a new public surface) asks first; self-contained
  changes flip. Unattended, the judgement holds and the action degrades: flip
  the small ones, leave the rest at a clean draft with the reason in the PR
  body.
- Rounds that never converge are usually structural. After three without
  convergence, stop, name the module the findings cluster around, and suggest
  improve-codebase-architecture rather than burning more CI.

AGENTS.local.md (gitignored, with CLAUDE.local.md importing it) holds the
ready/ask calibration, recorded as dated observations rather than a rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(dev): state that each worktree gets its own fresh database

The per-worktree section warned which DATABASE_URL to use but never said where
the database comes from: the post-create hook creates and migrates a new one
per worktree, so it starts with none of the main instance's workspaces, scripts
or flows. WM_CLONE_DB was documented only as a comment in .webmux.yaml, which
reads as how things work rather than as a per-project opt-in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(sqlx): script the cache backup/restore instead of documenting it

The update-sqlx skill spelled out a cp/comm/rm dance around `cargo sqlx
prepare`, which empties backend/.sqlx before regenerating — a failed run leaves
the cache gutted (observed: 2350 -> 142 entries), and a --all-targets run in a
CE checkout fails that way every time. Three problems with documenting it:

- The backup path was the literal /tmp/sqlx_backup, shared by every worktree.
  Two concurrent runs overwrite each other's backup, which is the only thing
  standing between a failed prepare and a gutted cache.
- The restore was a copy-pasted `rm -rf .sqlx && cp -r ... && cp ...` chain.
- Skipping the backup is what turns a routine failure into a lost cache, and a
  convention is easier to skip than a command.

sqlx-cache.sh has backup / newq / restore, keeps state in a per-worktree
directory, and leaves the judgement call where it belongs: `newq` prints each
added entry's query field for review, and only `restore` writes them in.

Also adds the general rule that scratch files belong outside the checkout —
anything written into the tree has to be deleted again, and rm prompts each
time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(agents): state why a routine cleanup prompts, and where scratch goes

The guard hook already auto-allows a plain rm whose operands are under /tmp or
inside a git checkout in $HOME, so deleting a temp dir or a stale .sqlx entry
costs nothing. What prompts is the command shape: the hook's tokenizer defers on
&&, ;, redirects, quotes and $VAR, so a chained cleanup falls through to the
Bash(rm:*) ask rule.

That was recorded only inside a paragraph about screenshot file paths in
frontend/CLAUDE.md, where nobody looking for it would find it. Stated in Core
Principles instead, alongside the rule that scratch belongs outside the tree —
for the reason that actually applies, which is not committing junk rather than
avoiding prompts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(security): deny agent edits to the permission hooks and project settings

.claude/hooks/guard-rm-outside-tmp.sh and guard-main-branch.sh are the
enforcement points for everything the permission rules are meant to catch, and
nothing stopped an agent editing them. One sed -i disables the guard for every
later command, silently, and the deny list in .claude/settings.json has the same
exposure.

Defence in depth rather than a boundary: an agent with arbitrary bash can still
delete, and this may only close the Edit-tool path if Bash writes are not
covered by Edit deny rules. It costs nothing and removes the cheapest way to
turn the guards off. Changing them now means editing the files by hand, which is
the intent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address review round findings on head 3f47dc1

- backend/ and frontend/ guidance was Claude-only. Codex and Pi read AGENTS.md,
  not CLAUDE.md, so moving "Verifying Backend/Frontend Changes" and the
  $bindable ban out of the root AGENTS.md made them invisible to two of the
  three CLIs this repo supports. Renamed both to AGENTS.md with a one-line
  @AGENTS.md CLAUDE.md beside them, matching what the repo already does at the
  root and in ai_evals/, and retargeted the four references.

- sqlx-cache.sh aborted with exit 2 and no output when .sqlx was empty:
  list_entries ran `ls -1 ./*.json`, and an unmatched glob under
  `set -euo pipefail` killed the script. An empty cache is precisely what a
  failed prepare leaves behind, so it broke in the one case it exists for.
  Replaced with a glob loop; reproduced the failure and verified the fix.

- The oneshot prompt ("never leave the PR sitting in draft") contradicted the
  "Flip, or ask first" rule added in the same PR, which tells unattended runs to
  leave wide-blast-radius changes as clean drafts. The prompt now defers to the
  skill for the flip decision and keeps only "never stop at an unreviewed
  draft".

- Bundled-resource references in the vendored skills were markdown links to
  `.agents/skills/...`, which resolve relative to the file, not the repo root.
  Replaced with inline paths stating they are repo-root relative.

- The PR-ready calibration file was write-only: the skill said to record
  answers there but never to read it. It is now consulted before deciding.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Revert "chore(security): deny agent edits to the permission hooks and project settings"

This reverts commit 3f47dc1692.

* fix: address round 2 nits

- backend/AGENTS.md told agents to persist CARGO_FEATURES in runtime.env, but
  webmux regenerates that file from metadata and .env.local every time the
  worktree is opened, so the setting is lost on the next reopen. The persistent
  source is .env.local, which scripts/post-create.sh already writes.

- UPSTREAM.md still described the vendoring delta as rewriting bundled-file
  *links* to repo-root paths. 555f063 replaced them with plain paths in prose,
  because a markdown target resolves relative to the file — a repo-root link is
  just as broken as a sibling-relative one through the symlink. Replaying the
  old wording on a refresh would reintroduce the bug UPSTREAM.md exists to
  prevent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(skills): correct the UPSTREAM.md link-rewrite delta

The delta note still described rewriting bundled-file *links* to repo-root
paths. 555f063 replaced them with plain paths in prose, because a markdown
target resolves relative to the file containing it — a repo-root link is as
broken as a sibling-relative one read through the symlink. Replaying the old
wording on a refresh would reintroduce exactly the bug UPSTREAM.md exists to
prevent.

The preceding commit's message claimed this fix; the edit had failed on a
stale anchor and only the backend/AGENTS.md half landed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(dev): describe what a fresh worktree database actually contains

Exercising a real worktree creation showed the previous wording ("none of your
workspaces, scripts or flows") reads as an empty database. It is a bootstrap
instance: the admins workspace, the admin@windmill.dev superadmin, the license
key copied from the base database, and the migration seeds — observed as
u/admin/hub_sync and the default app theme resource.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 11:12:20 +00:00