Retain complete status plugin files during replacement, symlinks and existing permissions, including legacy Windows directory permission recovery.
Fixes#24121. Continues #24131; permission-recovery review credited to pullfrog.
Co-authored-by: Ahmed Nagy <ahmednagy25t@gmail.com>
* fix(native-chat): plain wording for chat errors and status rows
Replaces Orca-internal words (host, journal, transcript, outbox, process,
unverifiable, "no contact", byte budgets) in native chat refusals, status
rows, the skills menu, subagent and background-task state labels and the
history-load error with plain language, and routes the two hard-coded
English composer errors through translate(). Copy only; no behaviour change.
* fix(native-chat): match chat copy to what happens and to the sidebar's words
- The held-message row says Orca keeps checking, which it does: the idle
sweep retries the unproven stop and a landed retry sends what waited.
- An unsettled earlier message reads as unconfirmed, not undelivered.
- The skills-unavailable line names SSH chats, its only cause.
- Subagent and background-task rows say "no recent update", the sidebar's
words for the same state, and "status unavailable" after a count; the
row no longer repeats the state as a reason.
* test(native-chat): find the repair row by its own text, not the old wording
* fix(native-chat): word the history-repair and too-large rows in the reader's language
Both rows were finished English the host wrote into the chat, so nothing could
translate them. Each now names itself with a presentation, the way the
compaction row does, and the chat says it through translate(). The English text
stays on the row for clients that predate these presentations and for the phone.
* fix(native-chat): say composer send errors with the chat's notice sentences
The composer's two errors had their own wording file beside the sentence table
every other chat notice uses. The send outcome now carries notice parts, worded
by the same function as the rest. A refused redelivery says "Orca couldn't
confirm your message reached the agent. Check the chat, then send it again if
needed." (the same sentence the failed-send rework uses), and a message this
client couldn't store says "Couldn't save your message. Try again."
* fix(native-chat): drop the retry line, keep one name for a lost task, and say only true causes
- The history error pane no longer adds "Orca keeps trying to load this chat."
under its title: the read still retries on its own, but the pane says only
that the chat didn't load.
- A write refused as unsupported asks for an Orca update only when no reason
came back, which means the host is older. A named reason (a location or agent
that can't run there, no chat host, a client missing the capability) now reads
"This isn't available in this chat.", since updating doesn't fix it.
- A task Orca lost track of is "no recent update" everywhere; after a count it
reads "2 agents with no recent update" instead of a second name.
- The skills menu announces the same sentence it shows, including in SSH chats.
- The row for messages held behind a previous agent reads "{{agent}} from before
may still be running. Your messages will send once it stops."
* fix(native-chat): a chat whose host can't run it says its history didn't load
A history read refused as unsupported with a named reason left the error pane
saying only "This isn't available in this chat.", which never said the chat
failed to load and named nothing the reader asked for. A read now says "This
chat's history couldn't be loaded."; other writes keep the shorter sentence.
* fix(native-chat): remove the "still starting" notice that flashed on chat launch
Every structured chat passes through a starting phase, and the pane showed
"<agent> is still starting. Messages wait until it is ready; close this chat to
give up on it." for it. A 5 s grace period only narrowed the flash to starts
that finish just past it. A message sent while the agent starts already counts
as working, so the chat's working indicator and Stop cover that state; the
notice only repeated it.
Removes the notice, its copy in every locale, the delayed-status hook that
existed only for it, and the per-child key the chat read only to reset it.
* docs(agents): no pop-up notices for transient or internal states
Records the rule the removed startup notice broke, so agents building UI
show transient states through existing surfaces instead of new messages.
* docs(agents): allow the common delayed loading placeholder
The rule against delaying a flashing message should not forbid a quiet
skeleton or spinner that waits a moment before appearing.
* fix(native-chat): stop publishing which provider child is starting
hostExecutionChild existed only to re-key the removed starting notice's delay.
Older clients read it as optional, so omitting it only changes when their
notice appears after a still-starting child is replaced.
* docs(agents): narrow the status-notice rule's wording
Ban 'execution host' rather than ordinary host copy, scope 'confirming' to a
connection, and keep the no-delay rule where the common pattern delays.
* docs(agents): let a control's own busy state wait, per the style guide
The status-notice rule covers notices and status lines; a delayed in-place
label swap like "Saving…" stays as STYLEGUIDE.md describes.
* docs(agents): keep the status-message guideline out of the repo
* fix(native-chat): a host admits structured sessions by client capability, not its own chat setting
A host's experimentalStructuredNativeChat decided whether any paired client could reach
agentSession.* at all, and whether session.tabs.* showed it structured tabs. That setting is the
host user's own launch preference: whether a new agent opens as a chat or a terminal is decided by
whoever launches it. Using it as admission control meant a client whose own preference was
"structured chat" was refused on a host whose preference was "terminal", and chats opened while
the setting was on were withheld from mobile once it was turned off.
The gate now asks one thing: did the client advertise agent-session.structured.v1 (in-process
callers negotiate nothing and are always admitted). Tab projection and restore follow the same
rule. With the setting no longer gating anything, the separate cleanup gate (close, cancel,
unsubscribe, release), which existed only so those kept working after the setting was switched
off, is identical to the main gate and is folded into it. The settings listener that republished
tabs when the setting changed is removed, since projection no longer depends on it.
The host setting still picks the default for launches that start on the host itself
(agent.launch from mobile, orchestration worker-start).
* fix(native-chat): the desktop declares structured chat support to paired hosts
The desktop renderer advertised agent-session.structured.v1 (and the Claude, turn-item and
background-task capabilities that go with it) to its own main process but not to a paired Orca
server. The server therefore refused every agentSession.* call from the desktop and stripped
structured chat tabs out of the tab list it published to it, so a structured chat running on a
paired server never appeared on the desktop, even though the renderer already mirrors a host's
agent-session tabs and drives each one against the server that owns its workspace.
The same renderer reads structured chats on either host, so the remote Electron list now carries
the same structured-session capabilities as the local one, and the capability test pins that
nothing is advertised only locally.
* fix(native-chat): negotiate client-chosen launch mode so released phones and old servers keep terminals
Hosts advertise agent-session.structured.client-launch-mode.v1: they admit
structured sessions by client capability alone. A remote client that does
not advertise it (phones released before agent.launch) asks createSupport
to pick the launch mode, so the host keeps answering that with its own
setting, exactly as before. Cleanup methods keep their own named gate so a
future admission condition cannot make close or cancel refusable.
* refactor(runtime): keep the Electron client capability list in its own module
protocol-version.ts is at its line budget; the list is what the desktop
advertises to paired hosts, not the host's own contract.
* fix(native-chat): the desktop declares it picks each launch mode itself
Paired hosts and the desktop's own main process then answer createSupport
by the workspace rather than by their own chat setting.
* chore(native-chat): justify the two type assertions this change's lines touch
* fix(native-chat): chats that already exist keep showing whatever the chat setting says
The structured chat setting decides only what new agents open as. With it
off, this machine's structured chats used to be hidden while the host,
which no longer reads the setting, still reported them to the workspace
activation gate, so a workspace holding only a chat opened empty. The
local chat mirror and its startup restore now run whatever the setting
says, the continue-after-restart offer follows the chats that exist, and
the setting's copy says it applies to new agents.
* test(native-chat): pin that a host advertises the client-chosen launch mode
* fix(native-chat): mirror this machine's chats only where it holds them
Round 1 ran the local chat mirror for everyone so existing chats show
whatever the setting says. That gave every desktop a permanent
session-tabs listener, which turns on the runtime's phone replication
paths, plus two full session-tab censuses at startup, and made the
browser client mirror its remote host a second time.
The runtime now says whether it holds structured chats: its structured
host is built only when saved chats were restored at startup or a client
created one here, and it announces the moment one is built. The mirror,
the startup restore and the continue-after-restart offer run only when
the setting launches chats or the host holds some, and never in the
browser client. A chat a paired client creates here with the setting off
still appears at once. The chat behaviour settings show wherever chats
exist, and the setting's copy says it picks what new agents open as. The
toggle-off teardown this made dead is removed.
* test(native-chat): record install listeners without a cast
* fix(native-chat): mirror this machine's chats only once it holds one, not once its host is built
Session history, resume preparation, terminal resume commands and replay-safe phone launches all
build the structured host for users who never had a chat, which turned on the chat mirror and the
structured-only settings rows until the next restart. The signal is now derived from the host's
records (or a records file still owed its import) and pushed when the first chat is restored or
created. A throwing listener no longer fails the install that fired it.
* feat(native-chat): createSupport reports the saved selection a new chat on this host starts with
A chat on a paired server starts with the server's saved model and options, which the desktop could
not read, so its picker showed a guess. createSupport's answer, which the desktop already waits for
before a paired launch, now also carries that seed as a new optional field (older clients ignore it).
Create and createSupport read it through one resolver so they cannot drift.
* refactor(protocol): move the Electron remote client capability list into its own module
Merging main left protocol-version.ts one line over the max-lines limit on this branch. The list of
capabilities the desktop advertises to a paired host moves, unchanged, into
electron-remote-runtime-client-capabilities.ts, the module the next PR in the stack already uses
for it; importers point there.
* test(cross-version): stub the launch seed resolver createSupport now reads
* test(protocol): pin the desktop capability divergence against what a paired server receives
Every paired transport sends the shared remote base plus the Electron list, so the
divergence test now compares that union with the renderer's local list instead of
the declared Electron list. A capability added only to the shared base can no
longer slip past it. The two base-only capabilities it surfaced are recorded:
skills.install-result.v2 has no local caller; the authoritative-inventory label is
read by the local tabs sync but dropped by main, and is marked unsettled.
The turn-item and both background-task-stop capabilities were already sent through
the shared base, so the Electron list no longer repeats them. The wire set is
unchanged; this PR's real change on the wire is structured.v1, the Claude
structured capability and the client launch-mode capability.
* fix(native-chat): the desktop tells its own host it picks each launch mode, so retrying an existing chat works with the setting off
* docs(native-chat): name the real exit for the released-phone createSupport rule
* test(cross-version): a released client still gets the host-setting createSupport answer; a launch-mode client gets supported plus the seed
* fix(jira): expire cached attachment images while idle
* test(jira): preserve cache bounds and site expiry after clears
---------
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(worktrees): delete removed checkouts in git, not in Orca's file pool
Local worktree removal renamed the checkout into a sibling trash root and
deleted it in the background with a recursive fs.rm in the main process.
That queued one request per entry on libuv's shared 4-thread file pool, so
for minutes every other async fs call in the main process (the agent-session
store behind chat sends, file explorer reads) waited behind the delete.
`git worktree remove` now deletes the checkout inline in git's own process
again, so the card stays in its Deleting state for the length of the delete
while Orca's file pool stays free. No timeout applies to the call, so a
large delete is never killed halfway.
If git reports success but the path still exists (Git for Windows leaves
junctions and their parent directories in place), the leftover is deleted
with the existing removeHostTree; WSL checkouts stay with the distro.
Nothing creates trash any more: the scheduling queue, rename/restore
helpers and the trash_rename span are gone. The startup sweep stays to
drain entries older releases left behind, and now removes each emptied
trash root so the obligation ends.
* fix(worktrees): let Git delete Windows checkouts with long paths enabled
Removal now always runs Git's own recursive delete, and worktree creation
checks out with core.longpaths on Windows, so a deep checkout Orca created
could fail to delete with "Filename too long" (#6433). The Windows recovery
then finishes the delete but keeps the branch. Pass the same command-scoped
core.longpaths option to `git worktree remove` so Git can delete what it
created.
Also point the CI shard timing entry at the renamed real-git removal suite.
* fix(worktrees): keep an inherited GIT_ASK_YESNO out of the worktree delete
Git for Windows asks $GIT_ASK_YESNO whether to retry when a file stays
locked during a recursive delete. Orca's git env inherits the user's
environment, so an inherited value would run an arbitrary prompt program
in the middle of a removal. Drop it for the removal call only.
* perf(worktrees): run worktree deletes under their own limit, outside git admission
`git worktree remove` now deletes the whole checkout in Git's own process,
which takes 20-35 s on a large tree. It took a general git admission slot at
status tier for that whole time, and that cap is as small as two slots on a
machine with six or fewer cores, so two deletes blocked every status read.
Deletes now skip general admission and queue under their own limit of two
per host instead: two concurrent deletes already saturate one disk, and more
only slow each other down. Leftover cleanup runs inside the same slot.
* fix(worktrees): delete removed checkouts in the background and mark them removing
Since the checkout is deleted by `git worktree remove` in Git's own process,
a large delete takes 20-35 s. Answering the request only after that made web
and mobile (30 s), paired desktop (60/180 s) and the CLI (60 s) report a
failure for a delete that was still going, and mobile silently re-showed the
row.
The request now does everything that can refuse (lock, cleanliness, archive
hook, watcher/terminal gate, terminal stop, shared-link unlink), records the
removal in an in-memory table on the host and answers `removing: true`. The
delete, branch cleanup and metadata purge run after it in the same order as
before, and the watcher/terminal gate stays held until they finish.
- Listings mark rows in the table `removing` for clients that advertise
`worktree.background-removal.v1` (the desktop renderer, paired desktop and
web), and leave them out for everyone else (older clients, mobile, the
CLI), which already dropped the row when the request answered.
- The outcome (removed, with any preserved branch, or the error) rides the
existing worktrees-changed event as an optional field, sent after the row
has left the table.
- A repeat delete while Git runs joins it. A create at the same path or with
the same branch is refused with "Cleanup is pending; try again shortly";
create's name search skips the path, so generated names move on.
- Nothing is persisted: after a quit or crash Git still lists the checkout
and it can be deleted again. WSL checkouts still delete inline.
- `orca worktree rm` says the checkout is still being deleted.
* fix(worktrees): keep the existing Deleting card until the host's Git finishes
The host now answers a local worktree delete on acceptance and deletes in the
background. The renderer keeps the existing delete state set until the host
publishes how it ended:
- The delete that asked waits for the outcome on the worktrees-changed event
(local IPC or the paired runtime's client event), then runs the same
teardown, preserved-branch toast and card error an inline delete did. If
that event is lost to a dropped connection, a listing that shows the row
gone after it was marked removing finishes the wait, and one that shows it
back without the marker fails it.
- Any other renderer (a reload, a paired desktop, web) sets the same delete
state from the host's `removing` marker and clears it when the marker goes.
A failure the host publishes lands on that card's existing error.
- Web advertises `worktree.background-removal.v1` so the host sends it the
marker; paired desktop does through the Electron capability list.
No new component, style or state: the card reads the delete state it always
did. A host that predates this answers when done without `removing`, and the
renderer takes that as finished, as before.
* test(worktrees): type the removal harness and projection for the node typecheck
* fix(worktrees): don't fail a delete retry with an earlier attempt's buffered failure
A background removal's outcome that reached this renderer with no waiter (another client's
delete, a host-marked card, or one already settled from listings) was buffered for 60 s and
consumed by the next delete of the same workspace, so retrying a failed delete failed at once
with the old error while the host was deleting. Drop the buffered outcome before sending the
request; only an outcome that arrives after it can belong to it.
* fix(worktrees): let only a gap in host events settle a background delete from listings
Git unlists the checkout before the host deletes the branch, cleans the push target and purges
metadata, and the worktree-directory watcher refetches within 250 ms. The renderer read the
missing row as a finished delete, so the waiter resolved without the preserved branch (no
toast) and a failure in those last steps showed as success; the real outcome was then dropped.
The listing fallback exists only for a lost outcome event, so it now applies only after this
host's event stream had a gap: a new subscription or a replay after reconnect.
* perf(worktrees): let a bulk delete start each same-repo checkout delete once the host accepts the last
A bulk delete ran one worktree at a time per repo (#2259, for packed-refs and ref-lock races in
branch cleanup). With Git now deleting each checkout for 20-35 s before the request settles, N
worktrees in one repo took N times that. The renderer now queues same-repo deletes only until
the host accepts each one; a parent still waits for its nested children to finish. The host
serializes the branch cleanup step per repo itself, which also covers removals started by
different clients.
* test(worktrees): pin the host platform in the mocked removal suites so they pass on Windows
Removal now passes -c core.longpaths=true on Windows, so the exact-argv
assertions and command-keyed mocks never matched there (17 failures on a
Windows host). Pin darwin as the add-worktree suites already do, and drive
the one Windows-specific case through the same spy.
* test(worktrees): type the blocked git remove result instead of a broad object
The anti-slop static-analysis gate rejects `object` parameters.
* test(worktrees): clear the changed-code quality gate in the removal suites
Merge the duplicate node:fs import, build the mock child without a cast, read
worktrees:list rows through one typed helper, and give the remaining casts a SAFETY line.
* fix(worktrees): record each background delete durably and finish it after a quit or crash
A quit mid-delete left git to finish the checkout on its own while the branch
delete and metadata purge never ran; a crash left a normal-looking row. Each
accepted local removal now writes a record beside the profile state before git
starts, clears it on success or failure, and the host runs the same delete
again for any record left at startup, re-deriving what remains from git and
disk. An orderly quit stops the checkout delete without waiting for it.
* test(worktrees): type the interrupted-removal assertions for the node typecheck
* fix(worktrees): finish an interrupted delete that already removed the checkout's .git file
Quit stops git worktree remove mid-delete, and Git deletes the checkout's .git
file wherever it falls in directory order. Git then refuses the checkout
("validation failed ... .git does not exist") on every retry, so the startup
finish failed and the row could never be deleted from Orca. A registered
checkout this record owns that has lost its .git file now finishes like an
unregistered one: leftover files, prune, then the branch.
* fix(worktrees): let Git finish an interrupted delete, and never take a different checkout
A quit or crash that stops `git worktree remove` after it deleted the checkout's
.git file left a registered checkout Git refuses to remove. The previous fix
deleted that leftover inside Orca's process, which is the bulk delete this
change exists to avoid (and on Windows the leftover can be most of the
checkout). The startup finish now rewrites the missing .git file from Git's
own admin entry for that path and lets `git worktree remove --force` delete
it. `git worktree repair` is not used: it also re-points every other
registered path, including a checkout another repository now owns there.
Orca deletes the leftover itself only when no admin entry claims the path.
The startup finish forces, so it now leaves the path alone when the checkout
there is not the one recorded: a registered worktree on a different branch or
head, or a `.git` at a path Git already unregistered. The record is dropped and
the card shows why.
The record write before Git starts is now bounded (2 s, logged when exceeded)
so a stalled disk cannot hold the delete, and the outcome is published before
the record's clear reaches disk.
* test(worktrees): compare worktree paths by value and tear down with Windows lock retries
Git prints forward slashes in `git worktree list` on Windows, so the real-Git
removal suites never found a joined path there: positive checks failed and
negative ones passed without proving anything. They now compare Git's parsed
rows by value. Teardown uses the shared retrying removeTree, since Windows can
hold the deleted checkout busy for a moment after Git exits. Adds a
relative-path worktree case for the .git restore (skipped before Git 2.48).
* fix(worktrees): reply to a worktree delete when it has finished, not on a broadcast event
A current client's delete request now waits for the host's background delete and gets its real
result (removed, a preserved branch, or the error) as the reply, the way it did before the delete
moved off the request. A request that arrives while the delete runs joins it and gets the same
result. Every other view keeps reading the host's `removing` marker: the row leaving means the
delete finished, and the row listed again without the marker shows "The delete did not finish.
Try again." on a card that view had marked Deleting. A request whose reply is lost (a timeout or a
dropped connection) settles the same way from a fresh listing instead of reporting a failure.
Clients without the background-removal capability (mobile, the CLI, older desktops) are still
answered on acceptance and have rows under removal left out of their listings.
This removes the outcome on worktreesChanged and everything it needed: the renderer's outcome
waiters, early-outcome buffer and TTL, per-host event-gap generations, the request pre-registration,
and the accept callback bulk delete used. Bulk delete runs same-repo deletes in parallel only on
this machine, whose host serializes branch cleanup per repo; SSH and paired hosts stay serialized.
* test(worktrees): type the pending-removal host id in the background-removal suite
* fix(worktrees): answer a delete request even when a concurrent removal of the same worktree replaced its record
The desktop app's removal and the runtime removal (CLI, paired clients) coalesce separately, so
both can be accepted for one worktree. The second replaced the first's record, and the first
delete then finished without resolving the request waiting on it, leaving the desktop card on
Deleting indefinitely. Each delete now settles the request it was started for.
* fix(worktrees): run same-repo removal archive hooks and teardown one at a time on the host
Local bulk delete now sends same-repo removals in parallel, so their archive hooks, terminal
teardown and preflight ran at once; a hook that writes refs can race the repo's ref locks
(#2259). The host now serializes each local removal up to acceptance per repo, for every
client; Git's checkout delete still runs in parallel under the delete limit.
* fix(runtime): keep waiting worktree deletes out of a host's foreground call slots
worktree.rm now replies only after Git deletes the checkout (up to minutes), so on paired
desktop and web each waiting delete held one of the host's 8 foreground call slots, and a
bulk delete queued listing refreshes and every other foreground call behind it. Deletes now
run in their own lane with the same bound; the 2-slot background lane stays for status polls.
* fix(worktrees): join a same-worktree delete accepted while a removal waited its repo turn
The desktop app and the runtime (CLI, paired clients, web) check for a running delete before
they queue for the repo's acceptance turn. A delete of the same worktree from the other path,
accepted while this one queued, was missed: this request re-ran the archive hook, stopped the
terminals again and started a second `git worktree remove` on the directory Git was deleting.
The queued acceptance now re-checks and joins the running delete.
* fix(worktrees): fence a resumed delete's checkout from startup, and drop rows a listing read before the delete finished
A delete a quit or crash interrupted took its terminal and file-watcher gate only when the resume
job ran, after the first window was shown; session restore could open a shell or watcher inside the
half-deleted checkout first, and on Windows that handle can fail the resumed git delete. Loading the
records now fences each recorded path, and the resumed job takes the fence over in the same tick it
takes its own gate.
A listing that read git's registration before a delete finished, and replied after the removal
record cleared, returned the row unmarked, so other views briefly showed "The delete did not
finish". Listings now capture the pending removals before reading git and leave out a row whose
delete finished successfully since; a row whose delete failed stays listed as before.
* test(worktrees): keep git's auto-maintenance out of the real-git removal suite
CI's Git 2.55 failed the file-pool test in teardown with ENOTEMPTY on the scratch repo's
objects/pack after the test body passed: the 3,000-file commit's detached auto-maintenance was
still writing a pack. The scratch repo now disables auto-maintenance and auto-gc.
* fix(worktrees): one archive-hook approval covers a same-repo bulk delete again
Local same-repo deletes now start together, so each queued its trust prompt with a state snapshot
taken before the first prompt was answered; approving the first still showed the same prompt once
per remaining worktree. The queued check now reads the store when its turn comes.
* fix(worktrees): a delete Git fails partway stays listed with its error; Delete retries it
`git worktree remove --force` drops the checkout's registration even when it
cannot delete a file (root-owned files, `chflags uchg`, a read-only Windows
directory). Orca lists workspaces from Git, so the row vanished after the error,
leaving the checkout, the branch and Orca's metadata with no way to retry.
- A background delete that fails with the checkout still on disk, unregistered,
and still the removed checkout's own leftover keeps its durable removal record
with the error (`failure`) instead of clearing it. Every other failure clears
it as before.
- Local listings (desktop list/list-all/detected, runtime list/ps/detected)
add a row for each such record, carrying `removalError`, and for a pending
removal whose checkout Git no longer lists (shown as removing).
- Delete on that row (desktop IPC and runtime worktree.rm) runs the recorded
removal again: terminal teardown, then the leftover, prune, branch and
metadata, under the per-host delete limit.
- The record ends on a successful retry, when the checkout is gone (listing or
startup), when a different checkout takes the path, or on forget-local.
Startup never retries a failed record.
- The finish's unregistered-path rule accepts a `.git` file naming the admin
entry Git removed (the leftover's own) and still refuses any other `.git`.
The removal table and listing projection move out of the background removal
module into worktree-removal-table.ts and worktree-removal-listing.ts.
* fix(renderer): show a failed delete's host error on its card
A row the host lists with removalError gets the existing delete-state error
(no new element), cleared when the host stops listing it failed. A row this
view marked Deleting that comes back failed, and a lost delete reply settled
from the listing, report the host's error instead of the generic one.
* test(worktrees): type the failed-removal listing and refresh mocks
* fix(worktrees): a failed delete's retry never removes a checkout Git registers at the path again
The retry replays the recorded choices (force, branch deletion) that were made for the unregistered
leftover. If the user removed the leftover and `git worktree add`ed the same branch at the path, the
new checkout matched the record's branch and head, so Delete force-removed it with its uncommitted
files, skipping the normal delete's cleanliness check. The retry now refuses a registered checkout
and lets the record go, so the next Delete takes the normal path.
* fix(worktrees): a failed delete's record ends at startup once its repo is removed from Orca
* fix(renderer): a failed delete's row offers Remove from Orca
* test(worktrees): type the failed-removal IPC test's module mocks without casts
* fix(worktrees): Delete picks retry or a normal delete from Git's current listing
* fix(worktrees): a failed delete's retry checks the leftover again right before deleting it
* fix(worktrees): Remove from Orca reaches paired clients and matches the failed row's own host
* fix(worktrees): a failed delete's retry re-lists only its own repo's authorized roots
* fix(worktrees): a second Delete joins a retry already running, and the startup finish keeps its last-resort delete
* fix(renderer): a failed delete's card says it failed, and Delete keeps the error for its dialog
* fix(renderer): a failed delete's dialog shows the host's error, and its card label keeps the full error a hover away
* style(worktrees): import the removal result types in one statement
* test(worktrees): a runtime listing right after a failed delete shows the failed row, not the cached scan
* fix(worktrees): pass the runtime retry's PTY-stop waiver in the shape the waiver invariant pins
* fix(worktrees): drop Remove from Orca from failed-delete rows
A failed delete stays listed with its error and Delete retries it; the
separate forget item, its dialog copy and forget's failed-record clearing
are removed. The startup clear for repos no longer in Orca keeps matching
the local copy only.
* test(worktrees): wait for the dropped record's write before the failed-removal suite tears down
* fix(antigravity): bound Windows hook stdin before posting status
* fix(antigravity): publish Windows hook companion before core
* test(antigravity): name Windows hook payload cases explicitly
* Bound AI Vault cache loading and cooperative atomic saves
Preserve schema 3 caches across compatible releases while limiting bytes, JSON structure, and newest unique rows. Keep in-process entries authoritative and retain a valid prior snapshot when the newest row cannot fit.
Credits @AmethystLiang for the original PR10708 cache bounds and cooperative persistence intent.
* Use checked cache JSON properties in cooperative serialization
Preserves lazy own-property access and all serializer bounds, yields and errors.
When git lists the same folder twice (a leftover worktree registration that points at the main checkout), Orca's runtime listing turned each line into its own worktree with the same id, so `orca worktree current`, `active` and `branch:` failed with selector_ambiguous, and paired clients saw a duplicate row. The runtime scan now keeps git's first row per folder, the rule the desktop sidebar already uses. Separately, for a bare or separate-git-dir repo added through a linked worktree, the scan no longer relabels the main row with that worktree's folder (it relabels only when the folder's git dir is the common git dir), so the worktree keeps its own row and branch in the CLI and the sidebar. No extra git command runs.
Part of #23631: the "Profile state writer command timed out" toast in that issue has a separate cause.
* fix(claude): settle a queued send the CLI withdrew from its own cancelled frame
Claude reports each uuid-stamped command's lifecycle (queued, started,
completed, cancelled). A send it withdraws from its queue gets `cancelled`
before the interrupt or cancel_async_message answer, so a lost or failed
answer no longer leaves that send pending: it settles as withdrawn, with the
same reason and words as the receipt path.
A command the CLI already started also ends `cancelled` when its turn is
interrupted or fails, so `cancelled` after `started` is not a withdrawal;
an echoed send has left the waiter lists and is never reached.
Tests replay real 2.1.280 captures, scrubbed.
* fix(claude): release a doubted send when the CLI reports its session idle
A Claude send whose write ended in doubt is recorded `unknown`, and a live
`unknown` reads as work still owed, so the chat showed Working until the
child exited. Claude sends `session_state_changed idle` only once its whole
queue has drained, so it can no longer be holding that send. The runtime now
routes that report to the host's existing release, the same one Codex's
thread-stopped report uses; it retires `unknown` only, never `pending`.
* fix(claude): keep a command's started mark when a redelivery re-emits queued; fixtures name msg_lifecycle_v1
* fix(claude): settle every terminal lifecycle state of a send the CLI never echoed
A send the CLI started, then cancelled before any echo, stayed pending: it may
already be in the conversation, so it is released as doubt (unknown, recovered),
never withdrawn and never re-sent. A late echo still accepts it.
The 2.1.280 schema has two more terminal states. `discarded` (the CLI ended its
session with the send still queued) settles as not delivered; `refused`
(declined before it queued) settles as not accepted by the provider. After
`started`, either one is doubt, as `cancelled` is.
The late-settlement path gains an `unknown` outcome, which the host records as
released doubt.
* fix(claude): release a send the CLI took but left unanswered when it goes idle
`session_state_changed idle` comes only once the CLI's queue has drained, so a
send it took that is still unanswered there got no echo and never will: a turn
that throws can leave `started` with no terminal state. Idle releases it as
doubt.
What proves the CLI took a send is its lifecycle frame. On a CLI that reports no
lifecycle, it is the send's place on stdin: one whose write finished before an
interrupt went out was read before the interrupt was, so the first idle after
that interrupt releases it too. A send armed ahead of the interrupt but written
after it is left alone, since the CLI may still run it.
* fix(native-chat): keep the idle sweep off a Claude child that holds a send
A Claude retrying a rate-limited request has taken the send but echoes nothing,
so no turn row exists yet and the sweep rested the child after the idle window,
turning the send into doubt. The adapter now reports whether the CLI holds a
send (lifecycle `queued` or `started`, not yet echoed or ended), derived from
the live waiters, and owed work counts it.
Nothing is stored: every held send leaves the live set on its echo, its
terminal lifecycle state, the CLI's idle, or the child's exit, so the hold ends
with the send.
* fix(claude): count only a started send at idle and as a held send
2.1.280's end-of-turn cleanup can report idle before it re-reads its queue, so
a send read in that window goes queued, idle, started. Releasing every taken
send at idle doubted that live send and dropped Working. Only a `started` send
is released at idle or keeps the child from the idle sweep; a `queued` one ends
by starting and echoing, by a terminal lifecycle frame, or with the child.
The stdin-order path for CLIs without lifecycle frames is removed: a doubted
send retired there disables content matching on CLIs that mint their own echo
ids, and no Orca failure called for it. Those CLIs keep the earlier behaviour.
Comments that said only a failed write or child exit ends a waiter, or that
idle comes only once the queue has drained, now say what ends one.
* docs(claude): say only what the CLI's lifecycle frames and idle actually prove
* fix(claude): hold the idle sweep while Claude has a send queued, not only started
The sweep rested a child whose CLI had queued a follow-up behind a turn, dropping
the send it had already taken. The hold now spans the CLI reporting it took the
send until its echo, a terminal lifecycle state, or the child's exit. The idle
release still covers only started sends: 2.1.280 can report idle before it
re-reads its queue.
* refactor(native-chat): give provider-proven late dispatch settlement its own module
* test(claude): pin a steer a Stop interrupts after it started as doubt, not withdrawn
* fix(native-chat): one ordered journal writer
Every journal write, streamed or direct, lands in the chat's one write queue
in the order it is issued, and has landed in the fold when its call returns
(except while an owed import is paid, when it lands in queue order). The event
sink stops being a queue ahead of it; journal-write coalescing, which moved a
replaced write to the tail, is removed; closing a sink no longer loses writes
already handed over. The ten flushStreamedEvents patches go. A streamed text
item takes its place at its first delta, so a direct write inside the
coalescing window never lands above text already streamed. A Stop interrupts
whatever its bookkeeping writes do.
* fix(native-chat): a failed Stop holds the lane until its queue pause lands
A Stop whose note write or stop call failed skipped the wait for its
withdrawal and pause, so the lane freed first; with writes queued behind
owed work, the drain could hand the waiting card to the agent after Stop.
* fix(native-chat): a journal write body is typed synchronous
The queue's ordering rule needs every write body to finish before it returns;
serialize still took a promise-returning body, so an await inside one would
let a later write land first. The body type now refuses a promise.
* fix(native-chat): a stream's first window snapshot always lands
The first-delta write counted as the growth checkpoint, so the window's
snapshot was skipped until 32 more characters arrived: a stream showed only
its first token, and a Codex reasoning row could sit as an empty aside. The
coalescer now marks the row-creating emit and the first snapshot after it,
and both providers' growth throttles write those.
* test(native-chat): streamed text rewritten in place above a Stop note
Covers a stream whose later deltas wait in the window across a Stop, for
Codex and Claude, and a Codex item completed inside the window.
* chore(native-chat): drop comments that still describe coalesced journal writes
* fix(native-chat): type the draft-table write body synchronous too
* fix(native-chat): an empty delta owes no streamed-text emit
The opening snapshot is forced past the growth rule, so an empty second
Codex delta rewrote the row with identical text. The coalescer now marks a
stream dirty only when a delta adds text.
* fix(native-chat): a mutation's open pays an owed import first
With the flushes gone, a Stop naming no turn, a goal set or a /clear could
read the fold while provider rows still waited behind a restore's owed
import: the Stop answered cancelled:false and interrupted nothing. The open
every mutation shares now waits for the import, as a reader's does; a failed
import is reported and never refuses the mutation.
* chore(native-chat): whenImported says mutations await it too
* test(native-chat): a Stop interrupts before its withdrawal or pause settles
Pins the order for a write that is held and then fails, for a Stop naming
its turn and one naming none.
* test(native-chat): a Stop ends a starting child before its withdrawal or pause settles
* test(native-chat): Stop order tests wait for the interrupt, not a 50 ms timer
* test(native-chat): settlement-order test reads rows through the journal row parser
Replaces a Reflect.get field walk, which the low-evidence audit rejects, with
parseJournalRow and the named row types.
* fix(native-chat): an empty delta still owes its emit, just not a forced one
The previous fix stopped an empty delta from marking the stream dirty, which
broke the pinned contract that an empty stream is snapshotted and flushed.
The coalescer now marks a stream dirty on every delta and tracks separately
whether its text changed since the last emit; only a changed snapshot is
the opening one that skips the growth throttle.
* revert(native-chat): drop the first-text row write from the delta coalescer
The immediate first-delta emit (and the opening flag and counters it needed)
had no Orca-observed failure behind it and added a journal commit per
streamed item. The coalescer, the Claude checkpoints and the tests that
pinned the first-text row go back to main's behaviour; the one ordered
writer, the sink hand-off and the Stop rules stay.
* fix(cursor): show the primary usage pool without hiding exhaustion
Use Cursor's reported plan percentage when its base allowance disagrees.
Select Cursor Models for compact display while preserving maximum-pool
warning, overflow, sorting and collapse behavior.
Adopts the plan mapping and primary headline intent from PR23531.
Co-authored-by: DakaAlvarez <149860458+Dacadev97@users.noreply.github.com>
* fix(cursor): describe compact usage summaries
Correct the existing six translations and fallback to describe one summary per provider after the compact Cursor primary-pool adoption. Preserve quota mapping, selectors, alerts, sorting and detailed presentation.
Validated source patch: d2a233e5f90a2cf8ccfb02dfafe99e0e03add48d with normal installed commit hooks, localization catalogs and changed-code quality. Standalone copy uses a private index and preserves the reviewed candidate ancestry.
Co-authored-by: DakaAlvarez <149860458+Dacadev97@users.noreply.github.com>
---------
Co-authored-by: DakaAlvarez <149860458+Dacadev97@users.noreply.github.com>
* Add Qoder session history and search with real CLI coverage
* Allow the real Qoder marker file to end with a newline
* Keep Qoder tool output out of history previews and search
* Keep Qoder search pages readable by older clients
* Verify persisted Qoder history after a real generated and resumed task
* Negotiate Qoder filters before searching an older execution host
* Combine search client imports for the CI plugin gate
* Keep the relay search oracle aligned with legacy agent filtering
* test(qoder): align search capability contracts and pin old-host fencing
* refactor(ai-vault): delete the unused session-scanner worker thread
Production always scans through the forked session-scanner service process;
the worker thread was reachable only under NODE_ENV=test or the
undocumented ORCA_AI_VAULT_SERVICE_PROCESS=0 switch, and nothing fell back
to it on service failure. Remove the thread (spawn, client, protocol, entry,
tests), its build entry, knip and plain-node-guard listings, and the
backend switch, so session-scanner-background always routes to the service.
- Move the scan options type to the service protocol as
AiVaultServiceScanOptions.
- Tests now mock session-scanner-service-spawn, the seam production calls.
- Repoint the hot-path listing reliability gate from the worker-client test
to the service-client test, which covers the same bounded-queue,
cancellation, and fault-restart properties for the real executor.
STA-9122
* test(ai-vault): cover the service's Claude-vs-OMP subagent lister choice
Runs the real service entry and subagent reader, replacing only the two
per-agent listers, so a swapped lister choice fails.
STA-9122
* fix(opencode): read the binder's session store on the foreign SQLite reader worker (STA-9122)
Before: the OpenCode session binder listed new sessions from opencode.db with
node:sqlite on the main thread every 60 s (and on SessionStart kicks), so a
large or contended store could stall the app the same way Cursor's did.
After: the read is a pure openCodeBinderSessions reader in
foreign-sqlite-readers/readers/, run only on the worker. The binder's
correlation, pane snapshot and process sweep stay where they were.
- The binder round awaits listSessions and re-checks its generation right
after, so a stop() during the read discards the round before it touches the
unbound map or the watermark.
- The client's in-flight dedupe key now includes the cursor, so a stale round
from before a restart cannot hand its rows to the restarted round.
- Idle teardown is per reader. The binder lane keeps its thread for 120 s,
longer than its 60 s poll, so the thread is not respawned every round.
- A timeout, crash, malformed reply or unstartable worker resolves to [] (no
sessions), the value the old read already returned on failure.
- An absent store still reads as [] without a log line, and a permission or
corrupt-file failure still logs (kept from #24577, now in the reader: it
stats the path and throws anything but ENOENT/ENOTDIR to the client's log).
- The binder lane inherits #24572's limits from the shared lane: no respawn
until a timed-out worker has exited, 2 consecutive deaths, a queue cap of
8. Its timeout stays 60 s, matching its poll.
- dispatch switches on the destructured kind, so a new kind without a case
still fails to compile.
orcad: the hook server runs there too, so orcad now ships
foreign-sqlite-reader-entry.js beside orcad.js (ORCAD_ARTIFACTS, built as an
orcad child). build-orcad runs a smoke check that starts the built worker
under the build's Node and under the pinned runtime, and does a real binder
read on a fixture DB, a Cursor read of a missing file and an OpenCode history
list. The OpenCode history scanner uses the same entry and was bundled into
orcad without it, so on orcad it always failed closed; it can now run.
Tests: reader (cursor, same-ms ids, OpenCode 2 rows, missing then created,
corrupt, inaccessible directory), retirement gate for the binder lane, dispatch
routing, client lane (rows, failure -> [], dedupe per cursor, own thread, idle
teardown default and override), binder loop with an async listSessions
(failure -> [], stop during the read), orcad path resolution through orcad's
host adapters, artifact list, and the smoke check against good, missing and
non-reading entries.
* test(opencode): cover the binder read deadline with fake timers and name the failure test accurately (STA-9122)
* refactor(sqlite): rename the OpenCode SQLite worker entry to foreign-sqlite-reader (STA-9122)
The worker thread that reads OpenCode's database off the main thread is about
to read other apps' databases too, so its entry is renamed to what it is:
src/main/foreign-sqlite-readers/foreign-sqlite-reader-entry.ts, built as
out/main/foreign-sqlite-reader-entry.js.
Why now: #24572 fixed the Cursor focus freeze with a second, dedicated
worker. Rather than grow one worker per foreign app, the next commit moves
Cursor onto this entry and deletes that worker. This commit is the rename
only; #24572's cursor-desktop-profile-worker-entry lines stay until then.
It moves out of ai-vault/ into a new foreign-sqlite-readers/ module because
it will no longer be session-scanner code; the module will own the readers,
their dispatch, protocol and main-process client.
The entry still routes only OpenCode kinds in this commit. The OpenCode
dispatch, protocol and process entry stay in ai-vault/ and stay OpenCode-only,
because the SSH/WSL relay reader bundles them (build-relay.mjs).
Every reference is updated: electron.vite.config.ts input key, knip entry,
the plain-node entry guard and its test, the asarUnpack list (the scanner
service still spawns this entry under ELECTRON_RUN_AS_NODE), and the
electron-builder test that reads the filename. The filename and the
beside-or-one-up (Rollup chunks) lookup now live in
foreign-sqlite-reader-entry-path.ts, which the OpenCode spawn reuses, plus an
Electron-main resolver that uses the packaged app.asar path.
* fix(cursor): move the desktop-login read from its dedicated worker onto the foreign SQLite reader (STA-9122)
#24572 fixed the Cursor focus freeze (#24360) with a dedicated worker
(rate-limits/cursor-desktop-profile-worker*.ts). Orca already runs OpenCode's
database reads on a worker, and more foreign-app SQLite reads are coming, so
keeping one worker per app means one entry, build input, asarUnpack line, knip
entry and guard line each. This keeps one pattern instead: Cursor's
state.vscdb read runs on the shared foreign SQLite reader entry, and the
dedicated worker, its entry and its config lines are deleted.
What moves:
- The read itself is a pure cursorProfile reader in
foreign-sqlite-readers/readers/ (was rate-limits/cursor-desktop-state-db.ts),
run only on the worker. A separate dispatch owns the new kinds and refuses
an unknown kind. The entry routes OpenCode kinds to the untouched OpenCode
dispatch, so the relay's OpenCode reader stays byte-identical.
- ForeignSqliteReaderClient gives each reader its own WorkerThreadRequestQueue
lane (own lazily started, idle-torn-down thread; one shared factory) with
in-flight dedupe per database path. Any failure resolves to the reader's
existing failure value and never falls back to the main thread.
Kept from #24572, so every reader gets them:
- Await worker retirement before respawning. Worker.terminate() cannot
interrupt a native SQLite call (e.g. a WAL-index rebuild), so the old
thread lives on until that call returns; respawning at once stacked a new
thread on the same work for every timed-out read (#24572 measured three
live workers). This belongs in the shared host, which fire-and-forgot
terminate(): LazyWorkerThreadHost now takes awaitRetirement and refuses to
spawn until the terminated worker settles, and the queue fails calls closed
meanwhile. Opt-in, because pure-JS clients (session scanner abort, port
scan) respawn right after an abort. A rejected terminate() also ends
retirement, so it cannot latch the reader off (raised in #24572's review).
- 10 s Cursor timeout, 2 consecutive deaths, a queue cap of 8.
- #24572's worker tests, rewritten against the shared client: responsive
caller plus coalesced probes, unavailable worker without path leaks,
stalled-worker recovery, no respawn before retirement, dispose settles.
Tests: reader, dispatch, client (timeout, 10 s default, crash, malformed,
unavailable without a main-thread read, dedupe, queue cap, own thread per
reader), queue retirement (stalled and rejected terminate), import boundary,
and an event-loop test reading a ~50 MB WAL with no -shm on a real worker.
* test(sqlite): walk the reader import boundary with the shared source-tree scan (STA-9122)
* fix(sqlite): key reader dedupe on a caller-supplied key, not the path alone (STA-9122)
An unreadable state DB was read as 'not busy', so the trust grant ran a
short app-server RPC that can refresh an abandoned backfill lease. A
tri-state pending check lets the trust grant take its fallback, while other
callers keep their existing boolean behaviour.
* fix(opencode-go): stop before OPENCODE_API_KEY when the credential database is unreadable
An unreadable OpenCode credential database read as 'no key', so the Go key
resolver fell through to OPENCODE_API_KEY, which OpenCode shares with its
Zen provider and can show usage for the wrong key. The database read now
reports unreadable separately; with an env key set the resolver stops, and
the usage fetch uses a configured cookie or shows a readable error.
* fix(opencode-go): treat a denied credential-database listing as unreadable, not missing