Commit Graph
13016 Commits
Author SHA1 Message Date
Brennan Benson 995ef11ce7 feat(native-chat): say in the chat why Orca stopped a reply, and offer Continue (#25675)
* feat(native-chat): say in the chat why Orca stopped a reply, and offer Continue

When the Orca that runs a structured chat (this computer or a paired server) quits, updates or
crashes mid-reply, the chat's stopped row now names the cause and the machine, and a Continue
button sends the existing restart continuation for that cut turn, with or without a restart offer.

- Host: a quit/update writes one turn-scoped row for the turn its stop cut, in today's words, with
  an optional `orcaStop` cause on the providerExited fact; restart adjudication stamps how the
  previous runtime ended on the deaths it proves (crash, or the quit/update it began), so the
  crash row names it too. Older clients keep their single row.
- Host: agentSession.continueInterrupted, capability-gated, rechecks under the session lock that
  the chat still sits on that cut, so a second click or a retry sends nothing.
- Client: the row's copy names the cause and machine; Continue sits above the composer.

* test(native-chat): Continue is not held by a recovery file that never answers

* fix(native-chat): bind an Orca stop's cause to the runtime that held the agent; neutral row, Continue explains itself

- The cause now rides on the host's row itself (`orcaStop` on the status row, beside today's
  words), the same row family and id scheme as the reopen's death row.
- Each recorded owner is stamped with the Orca runtime that holds it; a death proven later (at
  restart, or when recovery stops a survivor) names how that runtime ended: the quit or update it
  began, else a crash. Owners an older build recorded, a terminal's claim, an agent that died while
  its Orca ran, and unreadable quit records all keep the generic words.
- The quitting runtime's word is written first in teardown, before the recovery wait, through a
  bounded asynchronous writer apart from the chat database.
- Row copy: one neutral sentence naming the machine and cause; it drops "You can continue in this
  conversation." while Continue is offered, and Continue's tooltip says what it does.

* test(native-chat): type the Orca-stop test fixtures so the typecheck passes

The cut turn's outcome takes the journal's outcome type, and the stand-in close reads the
provider sink through a checked lookup instead of an index that may be absent.

* fix(native-chat): call a cut a crash only when Orca's runtime started and never ended

A chat said "Orca stopped unexpectedly" whenever its runtime left no quit record, and only the
desktop quit wrote one, so a headless server's restart or update, the Settings relaunch, and a
Windows logoff all read as crashes.

Each runtime now records its own start when its chat store opens, and every graceful exit records
its end through one synchronous entry point: the desktop quit's teardown, the headless server's
stop, the in-app relaunch, the GPU-fallback restarts, the update-install watchdog, and Windows
session end. A crash is a runtime that started and never ended; a runtime with no readable record
(never written, pruned, unreadable) names no cause, so the chat keeps its generic words. One file
per runtime, written durably and only by that runtime, so a damaged file never blocks a later
write and two processes never lose each other's record.

* fix(native-chat): Continue answers once Orca accepts it, not once the agent has started

On a paired server, Continue waited for the agent to start before answering, and the client
gives a paired call 15 s. A slow start (account switch, login shell, a long resume) showed
"Couldn't continue this chat" while the agent was in fact continuing.

Continue now answers when Orca has accepted the message, as a send does. The agent's start and
answer settle afterwards, and a start that fails is the chat's own note, as before. The restart
dialog's batch still waits for the handover, which is where it counts a start as done. The
verdict helpers move to their own module to keep the continuation file within its size limit.

* fix(native-chat): a reply the user steered, or a command run after the cut, still offers Continue

The cut detector stopped at the first user message after the cut turn, so a steer the turn had
taken, or a conversation command such as /context run after the cut, removed Continue while the
row still named the cause.

The rule for what is no request of its own (a conversation command, a row its turn produced, a
send handed into a running turn) moves out of the latest-request reader into one shared
predicate, which both that reader and the cut detector use. The host's "still wanted?" check
reads the same detector, so the client and host agree.

* fix(native-chat): a death proven after an earlier settle explains the turn it ends

When a chat was read before the restart proved its old agent dead, the read could only call the
turn unverifiable. The proof then revised the turn to interrupted, but the death row was scoped
by the turn still marked running, and none was, so it landed on the conversation instead of the
turn. That cut never offered Continue, and the row did not name it as the turn's explanation.

The row is now scoped to the newest root turn the settle actually ends, running or revised.

* fix(native-chat): the cause row's words stay put, and Continue waits out a resume already running

The row's "You can continue in this conversation." came and went with the button: it showed
while a paired host's answer was still on its way, vanished when the button appeared, and came
back the moment Continue was clicked. Continue also appeared on chats the restart prompt or the
launch's own resume was already carrying on.

The row now drops that sentence wherever the chat's host can continue a cut, counting a host
that has not answered yet as able (a host that writes cause rows has Continue), so its words
never change on screen. Continue is hidden while a resume is carrying that chat on.

* fix(native-chat): a refused Continue says so once, in the composer

A Continue the host refused before accepting anything wrote a red note into the chat and brought
the button back, so each retry added another identical note; a chat the host had no record of
was refused with no word at all.

Continue now reports every refusal the same way as a failed request: the existing composer line
"Couldn't continue this chat. Try again, or send a message.", which a retry replaces rather than
repeats. A refusal before acceptance writes no note. A failure after the message was accepted
(the agent could not start) is still the chat's own note, as for any send.

* fix(native-chat): the row naming Orca's stop carries its own presentation and never folds

A client that re-words host rows it cannot name (the draft that makes these cuts read as
interruptions) treated the cause row as an older red row and replaced its words, so the cause
never showed there. The row was also folded away under its collapsed turn once shown muted.

The host's cause row now names the presentation 'orca-stop' beside today's words, failure fact and
red tone, so a client that predates both changes still prints exactly today's row, red and on
screen, and a client that re-words unnamed rows passes it through. This build shows it muted,
counts it as no failure (the reply it cut stays the turn's answer), and never folds it; the fold
field becomes `explainsTurn`, as the other change names it.

* test(native-chat): pass the session-end event without a type assertion

* test(native-chat): the row naming Orca's stop renders neutral, whoever re-presented it

Pins the rendered tone on this build: the stored red row, and the same row after a reader
re-presents it in the neutral tone with its presentation and cause kept, both render muted and
never fold. The phone draws chat rows without tone styling, so it needs no change.

* fix(native-chat): the "Couldn't continue" line goes once the chat is continued

The composer line a failed or refused Continue set stayed on screen while the agent carried on:
after an answer lost in transit, or once another client or the restart prompt continued the
chat. Only the next Continue click or the user's own send cleared it, and a click also wiped an
unrelated composer error.

The line is now derived: shown only while the chat still sits on the cut that Continue failed
on, so it goes as soon as the journal shows the chat continued, from anywhere. A Continue click
clears only its own line, and a retry answered "already continued" leaves none.

* fix(native-chat): Continue waits while an opted-in launch may still resume the chat

With "resume automatically" on, Continue showed on a quit or update cut while the launch was still
waiting for its settings and reading the restart offer, then vanished when the launch's own
resume began; a click in between sent a competing continuation.

The launch's one decision (nothing offered, ask, or resume) is now published, and the chats it
resumes are named the moment it decides, with no gap. Until it decides, and while the setting
has not loaded or is on, Continue stays hidden on this machine's chats; a paired server's chats
are not the launch's to resume and keep it.

* fix(native-chat): a runtime's end survives a late reinstall, a failed write and any clean exit

Three ways the runtime record could still read a graceful stop as a crash:

- A chat host reinstalled during the quit (a request landing after teardown began) recorded the
  runtime's start again and erased the end it had just written. A second start of the same
  runtime now keeps that end.
- When the end could not be written (a full disk), the start alone stayed and read as a crash.
  The runtime now removes its record, so its chats name no cause.
- Each `app.exit(0)` had to remember to record the end. A process 'exit' with code 0 now records
  a quit when nothing else did: Electron emits it on every quit and exit once its loop runs
  (`app.exit` -> Browser::Shutdown -> the app's 'quit' -> process 'exit'), and Node on every
  `process.exit`. The relaunch and GPU-fallback calls it covers are dropped; the quit teardown,
  the headless server's stop, the update watchdog and Windows session end keep theirs, which run
  earlier or say more.

* test(native-chat): build the re-presented row as the plain status item it is

* fix(native-chat): a Continue click clears the composer's old error, so its own failure shows

Since the "Couldn't continue" line became derived, an older composer error (such as "Remove
attachments before using a chat-session command.") outranked it: a failed Continue showed the
old error instead, and a Continue that went through left the old error on screen.

A Continue click is the user's newer action, so it clears the composer's error again, as before;
the line then shows the Continue's own failure, if any. That failure still goes away by itself
once the chat is continued, and nothing but the user's own Continue click clears an unrelated
composer error.

* fix(native-chat): a chat start compares the owner process, not the runtime stamped on it

A chat start checks that the process it just started is the one the record names, by a deep
comparison of the stored owner with the adapter's process. The store stamps that owner with the
Orca runtime holding it, so the check passed only because the store happened to return the record
from before the stamp; returning the published record would have refused every chat start with
agent_session_ownership_unknown.

The start now compares the process identity without the runtime stamp, which says who holds the
process rather than which process it is.

* test(native-chat): count agentSession.continueInterrupted among the structured methods

* refactor(native-chat): derive the structured chat's transcript session in its own hook

Main's appearance work and this branch's Continue wiring together put NativeChatStructuredSession
past the 400-line limit for components. The session the transcript reads moves, unchanged, to
use-structured-chat-live-session.ts.

* refactor(native-chat): keep the Continue capability in its own module

Main grew protocol-version.ts to its line limit; the Continue capability moves to its own module,
as other capability groups have, and the runtime list still names it.

* refactor(native-chat): keep two shared files within their line limit after the main merge

Main left agent-session-record.ts and structured-agent-session-params.ts just under 300 lines,
and this branch's additions put them over. The account-home shape check moves next to the
account-home type it checks (written without a type assertion), and the Continue params move to
their own contract module; the params catalog is regenerated. No behavior change.

* test(native-chat): compare the store directory's files without depending on listing order

The corruption test checks that no file was created or removed by comparing two recursive
listings. Their order is the runtime's: with the per-runtime record directory nested under the
store, Bun returns the same entries in a different order than Node. Both listings are now sorted.

* fix: share the path bound main's launch-directory check needs

* test: give the stop-row fold rows the draws flag main's fold now reads

* refactor: mark the launch's resume decision where the resume begins

* test: count main's new structured method alongside agentSession.continueInterrupted

* fix: the journal database keeps its folder, where runtime end records live

Main's #26038 dropped stateDirectory from JournalHostDatabase; this PR's
runtime end records are read from and written beside it.
2026-10-07 14:48:28 -07:00
OrcaWinandm4air e08e0e0703 ci(e2e): move apt off the Azure mirror on runner images that use a mirror list (#26335)
Ubuntu 24.04 runner images resolve apt sources through /etc/apt/apt-mirrors.txt, so rewriting only the source lists left package downloads on azure.archive.ubuntu.com, which intermittently fails ('Ign:' on every package). Rewrite the mirror list too, and retry the install once with --fix-missing.

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-10-07 14:40:42 -07:00
Neil 31a5b42b0d perf(test): advance terminal probe policy clocks explicitly (#26327) 2026-10-07 14:38:56 -07:00
Brennan Benson 661cba999e fix(native-chat): a local chat this computer can't run opens the agent in a terminal (#25947)
* fix(native-chat): a local chat the host declines opens the agent in a terminal

A local structured launch opened its chat tab before this machine said whether it could create
the chat, so a "no" (a Claude account bound to WSL, for example) left a failed chat whose Retry
failed the same way and no terminal. Local launches now ask first through the same admission a
paired server's launch already uses: nothing of the chat exists until the host answers, and a
"no" runs the caller's own launch as a terminal.

* fix(native-chat): no stray shell while a local chat waits on its host; bound that wait

Round 1 review of the ask-first local launch:
- A launch waiting on its host now counts as a pending chat create, so the first-terminal
  watcher no longer seeds "Terminal 1" beside a local chat (watched creates and the
  empty-workspace default chat).
- This machine's answer is waited on for at most 3 s; past it, or while a new workspace is
  not resolvable yet, the chat opens and its own create reports, as before. Only a real
  "no" opens a terminal.
- The empty-workspace default chat seeds its plain shell on a "no" instead of starting the
  agent in a terminal nobody asked for, through a narrow launcher option.

* test(native-chat): type the declined-create worktree fixture's owner fields

* fix(native-chat): one first-surface claim for a launch waiting on its host

The wait was held in a separate record only the passive watcher read, so re-activating the
still-empty workspace during the wait seeded "Terminal 1" through the activation's reseed. The
wait now lives in the empty-workspace default-surface claims both seeders already read, and the
separate record is gone.
2026-10-07 14:33:08 -07:00
Brennan Benson 1617ff32ef feat(mobile): show chat visuals inline in the phone's native chat (#26071)
* feat(native-chat): visual directive grammar and host read for a chat's visuals folder

A shared grammar for the ::orca-visual{file="..." title="..."} reply line,
the per-chat visuals folder location on the owning host, and the
agentSession.readVisual runtime method that reads one visual with lexical and
canonical containment, a 512 KiB bounded read and UTF-8 refusal.

* feat(native-chat): shared frame document for chat visuals

One string builder every client wraps a visual's HTML with: the policy
(CDN assets only, no fetch, frames, workers, forms or base rewrites),
the theme variables, and a prelude that reports height, routes links to
the parent and refuses navigation. Also the validated frame-to-parent
message reader and the live theme message.

* feat(mobile): render native-chat visuals inline in the phone's chat

A finished assistant reply's ::orca-visual line now shows the visual
inline, read from the chat's owning host through agentSession.readVisual.
The visual runs in an opaque sandboxed child of a trusted host document
inside the WebView; the app accepts only a token-checked height and an
http(s) link opened under user activation. Navigation away from the
visual's own document is refused, a dead web process reloads once, and
the visual opens full screen. The hybrid shell's page renders it sealed.

* feat(native-chat): render chat visuals inline and in the right sidebar

Native-chat assistant replies render a ::orca-visual{...} line as the chat's
HTML visual in an opaque, scripts-only sandboxed frame: CSP first, the host
frame navigation guard registered before content runs, live theme without a
reload, fitted height, links opened in the viewer's browser only from a real
gesture, lazy mount, and one muted line when the visual cannot be shown.
Open in sidebar shows the same frame in the right sidebar, widened while it
is open and restored after.

* fix(mobile): chat visuals use the shared shell; only the host page may message the app

Review round 1:
- use PR 1's shared visual shell and height governor instead of a second
  builder; the app decides heights and pushes them to the host page
- react-native-webview patch: the message channel accepts only string
  messages from the main frame on iOS and Android, so a visual in its
  sandboxed child cannot reach it (or crash Android with a non-string)
- structured replies grow in place: while a turn works, a row holds back a
  directive still being typed at its tail; finished lines mount
- links need child focus + activation, one per activation window
- a refused read takes the visual down and drops its cached bytes
- in-page anchors load; text spelling a placeholder renders no visuals

* refactor(native-chat): move the visual height governor to src/shared

The phone bundles only src/shared, so the governor the desktop frame uses
moves there unchanged and the mobile frame shares it.

* fix(native-chat): point the desktop and mobile frames at the moved governor

* fix(mobile): review round 2 for chat visuals

- host page relays only messages on the visual's own channel, as its own
  copy, so a visual cannot push oversized fields through it; relay rate
  halved; a link is validated before it uses up the link window
- the frame denies camera, microphone, geolocation, clipboard and display
  capture; iOS media capture requests are denied
- react-native-webview patch: the iOS history-shim handler also accepts
  only main-frame string messages
- only the newest assistant row of a working turn holds back an unfinished
  directive; earlier finished rows show their visuals
- an error reply or older host is not a verdict: it keeps a visual on
  screen and its cache, and retries; only a host refusal takes it down

* fix(mobile): a drag that starts on an inline visual scrolls the chat

iOS gives a scrollable frame its own scroll view, which took the drag;
the inline frame is sized to its content, so it no longer scrolls.
Full screen still does.

* fix(mobile): review round 3 for chat visuals

- hold back only the last block of the newest assistant row, and not
  while a question or approval is open, so a visual followed by a tool
  call or a pending question shows at once
- an older host (method_not_found) reads as unavailable without retries
- drop a stray @pnpm/exe lockfile block; only the patch hash changes

* test(mobile): list the visual frame's web sibling; mock it in the prompt-controller harness

* fix(mobile): an older desktop's mobile allowlist refusal reads as unavailable at once

* fix(native-chat): visual CI fixes, shared height governor, live-turn streaming hold

Registers agentSession.readVisual from the methods index so the structured
method file stays under its line budget, replaces reflective reads with checked
narrowing, moves the pure height governor to src/shared for mobile, and holds a
half-written directive tail while the turn works (structured text rows carry no
running state).

* fix(native-chat): harden the visual read and link opening

Re-checks after the open that the chat's visuals folder is still the real
directory at Orca's path, reports unexpected filesystem faults by code without
host paths, and lets one click in a visual open at most one page.

* fix(native-chat): keep visual lines out of plain-text reply surfaces; review fixes

One shared helper drops visual lines (outside fenced code) from reply text where
it becomes plain text: the structured status summary that feeds the sidebar row,
dashboard, notifications, phone rows and handoffs, and AI Vault reply previews.
Review fixes: height also counts a pinned body's overflow, only the live
frontier row holds a half-written visual line, the runaway-height stop needs the
same step repeated, and any host refusal evicts the cached revision.

* fix(native-chat): resolve the visuals folder without the removed journal-paths helper

Main removed the per-chat journal paths and the journal database's state
directory; the visuals folder keeps the same sha256 layout on its own and the
read method uses the profile state directory the chat host is opened in.

* fix(mobile): keep visual lines out of the worktree row and copied message text

A reply's ::orca-visual line renders only in the transcript. The
worktree list's agent row and the message actions sheet's copy text now
drop it with the shared helper; a reply that is only a visual falls back
to the prompt, as an empty one does.

* fix(native-chat): review round 2 fixes; copy a reply without its visual lines

Reply previews in Agent Session History drop visual lines per text part before
lines are folded; the frame adds a body's overflow only when the body really
overflows; fence tracking follows CommonMark closers and openers; the copy
button copies a reply without visual lines; a coded read fault keeps its cause.

* fix(native-chat): update the frame's theme ref after render; read the visuals folder pair at once

* test(native-chat): declare agentSession.readVisual on the cross-version agent-session surface

* fix(native-chat): copying a reply keeps its code blocks and indentation

Removing visual lines now closes only the gap each removal leaves, instead of
collapsing blank lines across the whole reply and trimming its indentation; the
visuals folder is checked parent first again so a broken path answers the same
way every time.
2026-10-07 14:21:05 -07:00
Brennan Benson 0f9f199322 fix(native-chat): no saved outbox on the desktop; one send at a time, and the host owns what it accepted (#25959)
* fix(native-chat): the host owns the send queue; the window keeps no saved outbox

The desktop kept each structured chat's unsent messages in localStorage and
sent them in order, so one message whose fate was unknown froze every later
send, Retry dropped it silently, failed sends could not be discarded, and an
offline chat could send hours later. The host already records every message
and owns the queue; the window now only sends.

- One in-memory sender for composer, launch prompts and messages sent from
  outside the chat. One send in flight per chat; a transport failure resends
  the same id for up to 30 s; a refusal that proves nothing was recorded puts
  the text back in the composer with the reason; a send that went out and was
  never answered shows an in-doubt line with Send again and holds nothing up.
- The host's "unknown" rows get the same in-doubt line and Send again, from
  the journal, in every window.
- A message an older build left in localStorage is never sent: the host's
  conversation outline decides what goes back to the composer, and the copy is
  deleted once that is saved.

* fix(native-chat): send through the structured chat RPC wrapper, with its timeouts

The sender called the runtime RPC directly, skipping the per-method timeouts
every other structured chat call gets. Only a remote host's request skips the
compatibility check the sender already ran.

* fix(native-chat): an unconfirmed send goes back to the composer, with no new row line

The common pattern draws nothing extra on a message whose delivery is in
doubt once its turn is over, and puts a failed send's text back in the
composer with the reason. So a send nobody answered in time, or one a host
answers in a way that proves nothing, goes back to the composer worded as
unconfirmed, and a host "unknown" row shows nothing extra. Removes the in-doubt
phase, Send again, and its strings.

A host's made-up row for an id its journal lost now reads as unconfirmed, not
as recorded, so that message comes back instead of vanishing.

* test(native-chat): pin that nothing resends a send given back as unconfirmed

* fix(native-chat): sends survive a tab close, never resend after a Stop, and keep remote images

- A normal tab close lets sends on their way settle; what the host never took
  goes back to the conversation's draft. Only a cancelled launch or a worktree
  purge drops them.
- After any Stop, a send already on its way is never sent again under its id:
  a doubtful answer, or a resend that was due, hands its text back worded as
  unconfirmed.
- A returned image keeps the SSH connection it lives on; the connection never
  goes to the host.
- An older build's saved message the host recorded and then rejected is left
  to the host's own row, never handed back.

* fix(native-chat): a Stop or a tab close never resends a send already out, and a kept card is the card's

A Stop that landed while a same-id resend was being readied (its timer
fired, its request not out yet) let that resend go out after the Stop.
The sender now tracks whether a request is awaiting its answer: a Stop
hands back every send between attempts as unconfirmed, lets one whose
request is out settle from its answer, and never issues a request after
it. A normal tab close withdraws the same way instead of doing nothing,
so nothing more goes out and nothing is dropped.

A send the host rejected but kept as a card (keptAsQueuedMessageId) is
the card's, from its reply or the journal, and never goes back to the
composer.

Adds the freeze tests: a send whose fate is unknown holds later sends
no longer than its deadline, and one the host can neither confirm nor
deny releases the next at once.

* fix(native-chat): read a resend's turned-away call or reused id as unproven

Ports the send-answer proof contract. A call the host turned away before
running it (method_not_found, invalid_argument, unauthorized) proves only
that this request wrote nothing, so it reads as never sent on a first
attempt only; on a resend an earlier attempt may have landed, and it goes
again under the same id. A resent id the host says was already used for
other content (messageIdReused) proves nothing either, like an expired or
conflicting id.

Pins the rest of the contract: any row the host returns is its own, a
thrown error is no answer whatever its code, and an older build's
refused, held or outlived-Stop copy is handed back once and never sent.

* refactor(native-chat): type the structured composer's send with its attachment type

Keeps NativeChatStructuredSession.tsx within the file length limit.

* refactor(native-chat): drop the kept-card guards the sender never needed

A kept send's row is a rejection that was never a Stop's, so the sender
already reads it as recorded, from its reply or the journal. The test
that pins it stays; the two extra checks only covered a kept row that is
also a withdrawal, which the host never writes.

* test(native-chat): route launch tests' sends through the client wrapper the sender calls

The sender sends through callStructuredAgentSession, but these launch
tests replaced that module with a factory that answered nothing (and still
named a probe that no longer exists), so every launch prompt resent until
its 30 s deadline and the tests timed out. Each factory now hands sends to
the runtime RPC mock the tests already answer, and expectations of a local
send no longer ask for the remote-only compatibility option.

* fix(native-chat): hand a message back without importing the composer's attachment hook

The worktree purge reaches the launch prompt, which hands text back, and
the attachment hook's imports reach the store. A test that builds the
real store behind a mocked one then waited on itself and hung. Handing
back now writes images to the draft store directly, as the hook's helper
did, and a test pins that each returned image keeps its SSH connection.

* fix(native-chat): one send per chat, with no line of sends behind it

A chat with a send out took further messages into an in-memory line and
sent them one by one. A message typed behind one in doubt then hit its
own 30 s deadline and came back as not sent without ever going out.

Now, as the common pattern does, a chat takes one send at a time: while
it is out, Send is disabled and Enter leaves the text in the box. The
sender refuses a second send instead of lining it up, so the 30 s
deadline always runs from the send itself. A Stop or a tab close stops
the one send: settled from its answer if its request is out, handed back
as unconfirmed between attempts, or silently if it never went out.

Notes sent from outside the chat while its send is out stay with their
sender (not ready); a launch prompt that meets the person's own first
message waits in the composer instead of being lost.

* fix(native-chat): give a Stop-withdrawn note back to its chat once its notes were cleared

Notes sent from outside a chat clear once their message is recorded. A
message the host recorded as pending and a Stop then withdrew came back
only to its sender, which had already let go of it, so the text was
lost. The chat's draft now takes it, as an earlier build's outbox did.

* test(native-chat): type the composer-actions probe without a cast

* chore(native-chat): drop outbox wording left in comments and an empty locale group

* fix(native-chat): pace failed checks, keep the host's reason, and hold the chat for its launch prompt

- A send whose checks failed before its request went out (an unreachable
  or incompatible host, an unreadable history) was tried again at once,
  about a thousand times a second for 30 s. Attempts are now paced by
  the attempts made, whether or not their request went out.
- A host that refuses every resend by throwing (native chat turned off,
  a journal that won't open) gave back only "couldn't confirm". The
  host's reason now comes first, still without claiming not sent.
- A launch's prompt now holds the chat's one send from the click, drawn
  as sending, so a message typed while the chat starts can't overtake
  it; it goes out once the chat exists, and a cancelled launch frees it.
- Notes whose send nobody could confirm say so instead of "did not
  accept", and notes launched into a new chat let go of their text once
  that chat's composer holds it.
- An open chat keeps drawing a recorded send until its row arrives, so a
  reply that beats the history no longer makes the message flicker out.

* fix(native-chat): hold a chat's sends until it has started, and send a failed chat's message with its restart

A message sent to a chat still starting went out at once to a host that
had no record of the chat yet, read for a fence it could not get, and
came back as not sent. A message sent to a chat whose start failed
restarted it but no longer went with the restart.

While a chat starts, Send stays off and Enter leaves the text in the box,
as with a send already out. A text message sent to a chat whose start
failed restarts it and goes as the restart's first message, through the
same staged-prompt path a launch prompt takes: it holds the chat's one
send slot, drawn as sending, until the chat exists, and comes back to
the composer if the restart fails again. Notes wait while a chat starts
and ride a failed chat's restart, keeping their text until it is sent.

* chore(native-chat): test the queue request, rejection words and card hand-offs; drop an unused clear

- Pin which sends ask the host to queue, the moved rejection wording,
  and that a queue send whose card was handed off and then refused or
  withdrawn, or whose replay names a withdrawn card, is never handed
  back.
- The send-at-most-once gate now says what the desktop promises after a
  reload: it never sends the id again, and hands the text back.
- The legacy read names when it goes, and drops the notice clear nothing
  called.

* fix(native-chat): keep a sent launch prompt's entry, and give back at once what can't go out

- Cleaning up a launch prompt released its send slot even after the
  prompt had gone out, which deleted the entry the sender keeps once the
  host records it. A launch prompt recorded as pending and then
  withdrawn by a Stop was lost from both the chat and the box, and an
  open chat dropped the new chat's first message before its row arrived.
  The slot now gives back only a reservation that was never sent.
- A send stopped before its request went out by something trying again
  won't clear (this client and the server can't talk, or the host
  refused the history read) kept Send off for 30 s and then said Orca
  couldn't reach the agent. It now comes back at once with its own
  cause: not sent, since nothing went out, or unconfirmed if an earlier
  attempt did. Transport errors keep the paced retries.

* fix(native-chat): notes keep their own text through a new agent's launch

Notes sent to a New agent whose start failed came back to the notes and
also sat in the new chat's composer, so sending both delivered the text
twice. The notes keep their text until it goes out (they hold it from
the click), so the launch now says so and no composer gets a copy, on a
failed start or a refused prompt alike. The tests that asserted the copy
in the composer pinned the old double ownership and now assert the
notes are its only owner.

Notes that rode a failed chat's restart and ended unconfirmed now say
Orca couldn't confirm them, as notes sent directly do, instead of that
the agent did not accept them.

* chore(native-chat): the send-once gate says what the desktop keeps across a reload

The desktop keeps a send's id in memory only: a send still unsettled at a
reload or crash is not resent and not handed back. The coverage notes no
longer credit the renderer tests with durable identity across a remount.

* fix(native-chat): say once why notes sent to a new agent did not go

Since notes keep their own text through a new agent's launch, a prompt
the host refused, nobody could confirm, or that found the chat's send
taken came back to the notes with nothing said anywhere: the chat shows
no notice for text its caller keeps. The notes menu now reports it once,
with the toast a send to an existing chat already uses: not accepted,
couldn't confirm, or not ready. A start that failed still says so in the
new chat instead.

* fix(native-chat): retry a history refusal the host says clears, and name it at the deadline

A send stopped before its request went out came back at once for any
refusal of the history read, including ones the host names as clearing
(its journal briefly unavailable, a chat detached while the host quits),
so the automatic retry was lost. Only a version mismatch and refusals
that won't clear come back at once now; the rest go again on the paced
schedule, and if the deadline still finds nothing sent, the words are
the host's refusal rather than Orca couldn't reach the agent.

* feat(native-chat): a send makes one request, and nothing ever sends it again

The desktop resent a message under its own id for up to 30 s when the
answer was lost. A send now makes exactly one request, as the common
pattern's clients do:
- An answer that is lost, dropped or proves nothing hands the text back
  at once with "couldn't confirm… check the chat".
- A failure before the request goes out (the environment check, the
  history read for the fence, a refused connection) means nothing went
  out: the text comes back at once with its own reason, or as not sent.
- The 30 s cap stays on the one request, so a host that never answers
  can't hold Send.

Nothing ever resends, so nothing can go out after a Stop: a Stop only
takes back a send still in its pre-send checks. The resend state goes
with it (tries, generation, awaiting, stopped, the resend timer, the
send-answers-proof probe, and the first-attempt/resend split in the
evidence), and the send-once gate says the desktop never resends.

* fix(native-chat): notes keep the chat's line, and a send held behind a rewind says it was not sent

- Only a send whose text belongs to the chat's composer clears the chat's line. Notes sent from
  outside the chat no longer wipe a "couldn't confirm... Check the chat" line that explains text
  already back in the box.
- A send the host turns away behind a rewind it could not confirm went back as "not sent" but said
  "couldn't confirm what happened. Check the chat". It now gives the rewind's reason and says the
  message was not sent.
- Reliability gate names the single-request test and records a fresh evidence run; the sender
  test drops its leftover resend mocks and a duplicate Stop test.

* fix(native-chat): a send ends even when putting its text back fails

If writing the returned text into the chat's draft threw, the send never settled: it stayed
"sending", the chat refused every later send until a reload. The send now always ends after a
hand-back, the failure is logged, and the chat's line adds "Couldn't save your message."

* fix(native-chat): a message typed during /clear stays in the box and follows the chat

A /clear moves the chat to a new conversation, and its host refuses any send while it runs. A
message sent then went out, came back refused into the old conversation's draft with its line, and
vanished when the chat moved on: the box and the line now belong to the new conversation.

- A /clear holds the chat's one send slot while it runs: Enter does nothing, Send shows busy and
  the text stays in the box, as for a send that is out. Notes sent from outside get "busy".
- Once it moves the chat, the old conversation keeps taking no send until the view leaves it, and
  what is left of its draft moves into the new conversation's draft, after anything there: when the
  composer's /clear settles, and again when the view moves.

* fix(native-chat): no "Send message?" or queue clear while the chat's send is out

While a send is out (or a /clear runs) the chat takes no message, yet Enter over a held queue still
opened "Send message?", and Clear queue deleted every card before its message was refused. Enter now
does nothing there and the text stays; Clear queue re-checks and deletes nothing if a send went out
after the question opened.

* refactor(native-chat): take the /clear draft carry out of this change

Moving the old conversation's draft into the one a /clear replaces it with fixes a bug main has too
(text left in the box during a /clear stays with the old conversation), so it goes in its own change.
Kept here: a /clear holds the chat's send slot while it runs, and the conversation it moved away from
takes no send until the view leaves it.

* fix(native-chat): a /clear's hold on sends always ends

- A view that unmounted while a /clear that moves the chat was out left the old conversation
  holding its sends until a reload: the late reply kept the hold for a view that was gone. The
  reply now releases it.
- The local call for a conversation command has no deadline of its own, so a /clear that never
  answers kept Send off for good. The hold now also ends at the command's deadline (195 s, the one
  the remote call already uses), whichever comes first.

* test(native-chat): opening a chat an older build left stuck; hand back its copy in send order

Pins, through the real chat hook, sends, legacy recovery and draft store,
what opening such a chat does: the queued messages behind a message the
host recorded in doubt come back to the composer once, in order, with the
"not sent" or "couldn't confirm" wording; the chat is free to send under
its read's fence; nothing happens while the chat has no fence here.

The legacy reader now hands entries back in the order they were sent
(queuedAt), as the older build's reader did, instead of array order.

* fix(native-chat): a send this window turned away for a re-paired server comes back as not sent

A managed server's update rotates its pairing, and this window's main
process then answers the next call itself, before forwarding anything,
with runtime_environment_changed. The send read that thrown answer as
proving nothing, so the person was told Orca couldn't confirm a message
that never left. It is now handed back as not sent, with the reason.
Every other thrown answer still reads as unconfirmed once the request
may have gone out.

* test(native-chat): a message refused for an expired attachment comes back with its file and why

A paired server checks every stored file a message names when it admits it,
and refuses the whole message before recording it when one has expired. The
sender hands such a message back to the chat's draft, file included, with the
refusal's words, whether or not a view shows the chat, so it can be removed
and attached again. Ported from the saved-outbox test that came with
attaching files to a structured chat on a paired server.
2026-10-07 14:07:30 -07:00
Jinwoo Hong 7b054e6494 fix(browser): keep the desktop drawing while a phone streams a browser tab (#26284)
While the desktop was in the screen saver or minimized, a paired phone opening a browser tab spun forever: the throttled main window stopped compositing, so the guest's captures hung. Each page stream now holds the existing renderer-throttle lease for its lifetime, released by the stream's own cleanup on every exit. A 10 s no-frame deadline reports the existing timeout error so a viewer never spins forever.
2026-10-07 17:04:09 -04:00
Jinwoo Hong e05c69fe8f fix(session): at startup, a local copy never overrides an SSH-owned workspace's own copy (#26098)
* fix(session): at startup a local copy never overrides an SSH-owned workspace's own copy

Rows the local partition holds for a workspace whose repo the catalog places on
an SSH target are residue (pre-#19572 builds, relay reattach). Startup used to
keep them whenever they held a tab and skip the SSH partition for that
workspace, dropping live tabs, restoring closed ones and erasing agent-resume
records on the first save. Boot hydration now drops those local rows (keeping
open files and visit recency) so adoption takes the SSH partition whole.

* refactor(session): let adoption take a catalog-owned SSH workspace instead of pre-filtering local

Replaces the sentinel-partition split with one rule in adoption: the base's
terminal tabs no longer keep out the partition the repo catalog places the
workspace on (contested ids keep today's behavior).

* fix(session): an SSH-owned workspace's records win under shared keys; unsaved local drafts survive

Addresses review: a stale local tab sharing an id with the SSH tab kept its
layout and resume records (fill-only), and a replaced open-files row dropped
local-only unsaved drafts.

* fix(session): an SSH copy with no tabs never replaces local tabs

Addresses review: an owned workspace the SSH partition holds no tabs for keeps
today's rule, so its empty tab row cannot wipe the local tabs.

* refactor(session): drop a superseded local copy before adoption; publish path passes owned ids

Simplifies the rule: for a workspace the catalog places on this SSH host,
uncontested and with host tabs, the base's rows are dropped and the existing
gap-fill adoption runs unchanged; unsaved local drafts survive. One resolver
says which workspace each scoped entry belongs to, shared by the census, the
drop and adoption. The upload path (persistedSessionForTarget) now passes
owned ids too, so main cannot publish the stale local copy to the host.

* fix(session): match host tab rows by workspace id; a local draft beats a clean host entry

Addresses review: a host tab row under a workspace key now counts toward
superseding the local copy, and a local unsaved draft for a path the host
holds clean is kept instead of dropped.

* fix(session): scope catalog-owned ids to the ssh partition the catalog names

A workspace homed on one SSH target with leftover rows in another target's
partition no longer has the owner's adopted rows removed by the leftover's
pass.
2026-10-07 17:00:09 -04:00
Neil 7b26725ff4 Keep native watcher contracts in Node and reset runtime test caches (#26315)
* test: keep native filesystem watcher contract under Node

* test: reset canonical repository keys with runtime mocks
2026-10-07 13:39:28 -07:00
Neil 00984ccb25 ci: run full pull-request unit tests across ten shards (#26295)
* ci: run full pull-request unit tests across ten shards

* Bound required Linux package tooling setup
2026-10-07 13:12:02 -07:00
Kelvin Amoaba 42a40f861d feat(markdown): render GitHub-style callouts (#26255) 2026-10-07 12:35:19 -07:00
Kelvin Amoaba 49e6cc1d71 fix(native-chat): show right-click Copy and Paste only where they apply (#26207)
Copy was always listed, greyed out or copying text selected earlier.
Paste was offered on messages, and did nothing in a structured
question's answer field.
2026-10-07 12:33:22 -07:00
Kelvin Amoaba 726eaf117e fix(native-chat): recall every sent prompt with Up/Down in the composer (#26044)
* fix(native-chat): recall every sent prompt with Up/Down in the composer

Recall read a list kept in the composer, so it forgot everything on reload or remount.
It now reads the prompts the chat shows.

* fix(native-chat): put the caret at the end of a recalled prompt

It stayed where the last prompt had it, so Up skipped past a recalled multi-line prompt.

* test(native-chat): read the editor in the recall spec through a type guard
2026-10-07 12:03:02 -07:00
Brennan Benson 990c62e6c6 feat(native-chat): attach files to a structured chat on a paired server (#25146)
* feat(native-chat): a paired server keeps a store for chat attachments

A structured chat on a paired Orca server had nowhere to put a file the client
attached. The server now keeps a per-chat attachment store under its userData,
filled through agentSessionAttachment.uploadStart/Append/Commit/Abort and read
back for previews through agentSessionAttachment.read, behind the structured
session gate and advertised as agent-session.attachments.v1.

Nothing records cleanup as owed: a sweep re-derives what may go from the host's
chat records and journal (unfinished uploads after an hour, uploads for a chat
that never existed or that its journal never mentions after a day). The clipboard
RPC's in-flight bookkeeping moves into a shared ChunkedUploadRegistry both use.

* feat(native-chat): upload attached files into a paired server's chat store

Main streams dropped or picked files (fs:uploadPathsToAgentSessionAttachments)
and pasted images (clipboard:saveImageAsTempFile with agentSessionAttachment)
into the chat's store on its paired server, reusing the file-import slice
streamer and pinning every call to the pairing revision and server process.
The browser client's paste takes the same route.

* feat(native-chat): attach files to a structured chat on a paired server

Pasted images, files dropped from Finder/Explorer and the file picker no longer
refuse with "Local attachments are not available for remote sessions." in a
structured chat on a paired server: they upload into that server's chat store and
the chat gets only the stored path. Dropped files show as pending chips at once
so Send waits for them; every attached item carries the server, pairing and chat
it was stored for, and a send to anywhere else drops it with the existing
"changed hosts" notice. Chips and transcript images read stored files back
through the server, never from this machine's disk. An older server gets an
update notice instead of an upload. The terminal-backed chat keeps refusing.

* test(native-chat): type the attachment test doubles without bare casts

* refactor(native-chat): keep the clipboard RPC on its own upload bookkeeping

The chunked upload registry stays for the chat attachment store only, beside it.

* feat(native-chat): claim chat attachments when the host admits a message

A message's references into the host's attachment store are claimed in the same
journal transaction that makes it durable (a direct send, a queued draft, a /clear
carry). The sweep deletes an upload only after marking it in that database while
no claim exists, so a send and a sweep can never both win: a client's send that
names an expired or foreign upload is refused whole with the new attachmentExpired
reason, before anything is recorded. This replaces scanning chat transcripts for
paths.

The store is now <root>/<upload id>/<name>, refuses uploads for chats the host
does not hold, caps stored names in UTF-8 bytes and avoids Windows device names.
The preview read goes through the protected bounded read and the request's reply
budget, so an image too large for one reply is refused instead of closing the
connection.

* feat(native-chat): let structured Claude read the host's chat attachment store

Attached non-image files live outside the workspace; the store root is passed as an
added directory so the agent reads them without asking, the common pattern.

* fix(native-chat): skip a dropped folder before staging walks into it

* fix(native-chat): pin the browser client's chat paste to its server

Every upload call goes to the environment the paste was meant for and checks its
pairing and server process, as the desktop upload does; a re-pair ends the upload
instead of storing the image on another server.

* refactor(native-chat): let the host's claim be the only attachment check

The composer no longer records which server and pairing each attachment came from,
and Send no longer strips attachments it judges foreign: the host refuses an
expired or unknown stored path when it admits the message, which also covers
retries, Stop-restores and queued edits the client check missed. Previews of
stored files route through the chat's own server by path.

Removing a file's chip while it uploads now keeps its @path out of the draft. A
dropped or picked file shows its name and kind on its chip while it uploads, with
no separate progress toast, and files that did not attach are named in one notice.
A rich-text paste into a chat on an older server no longer shows the update notice
beside the pasted text.

* chore(native-chat): keep the claim hook beside the draft consume and fit the line budgets

The submission's claim runs in the same append hook as a draft's consume, owned by
the queued-message collaborator, so the journal store stays within its size limit.
Formats the new files and regenerates the runtime-required English catalog.

* test(native-chat): pin the attachment store grant in the Claude launch, restore upload result types

* fix(native-chat): create the attachment store at install and claim only exact store paths

Claude drops an added directory that does not exist when it starts, so the store
root is created when the host installs it, best effort. A commit keeps its upload
in flight until the rename lands, so a sweep never takes a slow upload's part file.
Only a path under this host's exact store root is claimed and required; any other
mention of a store path is plain text and never refuses the message.

* refactor(native-chat): reuse the newer-Orca notice and drop leftover attach options

An older server's refusal to store attachments uses the notice every other write
it refuses already shows, so one string fewer in every catalog; the attach
callbacks lose an options parameter no caller passes since the provenance check
went.

* fix(native-chat): give a message refused for an expired attachment back to the composer

Sending the same message again can never bring the attachment back, so instead of
a Retry that always fails the message's text and images return to the composer,
as a Stop's withdrawn message does, and the notice says what to remove.

* fix(native-chat): keep uploads across a prompt, say why a file did not attach

An upload that finishes while a prompt card has replaced the composer lands in the
scope's attachment and draft caches, which the composer reads back when it
returns, instead of the unmounted composer. The single failure notice names the
cause the files share, such as the size limit. Rich text pasted into a chat on an
older server no longer flashes an image chip: the chip waits for the server to
take the image. A picked command waits, as Send does, while an attachment is still
uploading.

* fix(native-chat): keep a paste's upload across a prompt; pin the Claude grant hand-off

A paste uploading into a paired server's store keeps its chip live when a prompt
card unmounts the composer, and its result is filed for the composer's return,
as a drop's is. A failure cause ending in full-width punctuation gets no extra
full stop. A runtime test pins that the store root reaches Claude's launch
resolver through the adapter.

* test(native-chat): type the grant hand-off captures without casts

* test(native-chat): pin that the composer files a paste's upload across a prompt

* fix(native-chat): retry a held rename at commit and keep cut names Windows-safe

A commit renames the part file with the Windows retry the app uses elsewhere, so
an antivirus or indexer holding it briefly no longer loses the upload. A name cut
to the byte limit is stripped of trailing dots and spaces again. The upload
pump's result cast states why it holds.

* fix(native-chat): keep attachments still on their way in the pane's attachment cache

A prompt card unmounts the composer, and an attachment still saving or uploading
lived only in that composer: one that came back showed no pending chip, so Send
went out without the file and the file then landed in the next draft. Pending
chips now live in the pane's attachment cache beside settled ones and settle
there whichever composer is showing, so Send waits for them after a remount too.
A stored file's @path goes into the pane's draft cache, which keeps it through an
input-method composition and a remount. A rich-text paste's image is registered
as a hidden pending chip before the server is asked, so Send waits for it from
the start, and it shows only once the server takes it.

* test(native-chat): a dropped file still uploading stays pending across a remount

* fix(native-chat): reveal a rich-text image in a composer that came back; keep caret inserts

A rich-text paste's held image is revealed through the pane's attachment cache,
so a composer a prompt card remounted during the server's answer shows it rather
than holding Send on an invisible chip. A stored file's @path goes in at the
caret while the composer is showing and not composing, as every other attach
does; only mid-composition or after a remount does it go to the pane's draft.

* fix(native-chat): never evict a pane's attachments while a composer shows them or one is on its way

The pane attachment cache is now where chips live, so its 128-scope bound passes
over a scope a mounted composer subscribes to or that still holds a pending
attachment, and evicts the oldest scope nobody uses instead. Protected scopes are
bounded by mounted composers and attachments in flight.

* fix(native-chat): never evict the scope just written when every older one is in use

* fix(native-chat): keep the attachment store out of the RPC method table's imports

The method table is imported far and wide (the SSH relay's CLI included), and the
attachment RPCs pulled in the store, whose preview read loads modules that read
fs constants at load: any test that mocks fs/promises without them failed to
import. The installed store now lives in a small registry module; only the host
wiring loads the store itself.

* refactor(native-chat): the submission hook gets its own module

Merging main added the operation receipt to the queued-message insert, which
put journal-queued-messages.ts over the 300-line limit. The submission hook
only uses the queued-messages object's public methods, so it moves out.

* refactor(native-chat): fit the attachment sweep and paste upload into main's line budgets

After merging main, the record store and the clipboard handlers each ran
a few lines over the 300-line budget. The sweep now asks the record store
whether a chat is recorded through the two lookups it already has
(readable or unreadable) instead of a new id-listing method, and the
paired-server paste upload lives beside the other attachment uploads.
No behavior changes.

* fix(native-chat): pending chips sit beside main's saved draft, which keeps server-stored images

Main now saves a composer's draft (text and settled images) in one store
that survives a reload or quit. This branch kept every chip, settled or
not, in its own pane cache. After the merge the saved draft owns settled
images and the pane cache holds only chips still on their way: they come
back pending in a composer a prompt card remounted, hold Send, settle into
the saved draft whichever composer is showing, and are never saved
themselves, so a restored draft cannot bring back an upload as if it were
attached.

An image a paired server stored for the chat is saved with the draft as
the real image (a pasted one is no longer turned into a "Not kept"
placeholder), and the restore check leaves it to that server, whose claim
at Send refuses one it no longer holds, instead of asking this machine's
disk for a path that only exists on the server.

Also: previews and the pending subscription move into their own small
hooks to fit the line budget, the "local attachments" notice moves to the
composer-target module so the attachments hook no longer pulls in the
upload module's imports, and tests follow main's new mocks.

* test(native-chat): give the claims test main's provider handle and message source

* fix(native-chat): a paste still uploading dies with its tab or workspace

A pasted image still uploading to a paired server sits in the pending
chip cache, which by design outlives the composer and its tab. Removing
the workspace deleted its saved drafts but left those chips, so the
upload finishing afterwards recreated and saved the deleted draft. With
the workspace's tabs gone it had no owner either, so no later workspace
cleanup could ever remove it.

A pending chip now records the owner its draft would have, resolved the
way the draft store resolves one when the chip is added. Workspace
removal and a user's tab close drop pending chips by the same owner,
conversation and tab matches they delete drafts by, and a chip that
settles after its chat's tab closed writes its draft under the owner it
was begun in.

* fix(native-chat): the claim type uses the journal's own SQLite type

* refactor(native-chat): read the pending-image handlers off the composer's attachments

Main's multi-file picker (#23956) and this branch together took NativeChatComposer over the 400-line limit; reading the two pending-image handlers the way the neighbouring ones already are keeps it at the limit.

* test(native-chat): pass the launch-args resolver main now requires in the attachment grant test

* fix(native-chat): grant the attachment store beside folders the saved Claude Arguments add

Main (#25721) now builds additionalDirectories from the user's --add-dir arguments; this branch's attachment-store grant replaced that list instead of adding to it. Both are granted now. Moves the Claude session-id derivation to its own module to keep the resolver under the line limit.

* test(native-chat): run the attachment store's SQLite tests in the Node runtime project

* test(native-chat): follow main's attach ownership recheck (#25749) in the attachment tests

* test(native-chat): a named upload chip still shows when the agent takes no images

* fix(native-chat): a paired-server image paste failure keeps its error apart, as main's notice card does

* fix(native-chat): the attachment sweep reads the chat list directly now that records have no import still owed
2026-10-07 11:56:55 -07:00
Jinwoo Hong d68f3bb5b2 fix(ssh): bind reattached SSH panes into the target's own session partition (#26088)
* fix(ssh): bind reattached SSH panes into the target's own session partition

A relay reattach persisted the pane binding into the local partition, minting a
minimal copy of every reattached SSH tab there. When nothing saved over it (a
reconnect with no window open), the next startup kept that copy and skipped the
SSH partition's rows for the workspace: its open editor tabs were missing for a
launch and its agent-resume records were dropped (STA-9544).

Bind into ssh:<target>, where the spawn bound the same pane.

* test(e2e): run the reattach home-partition spec only in the Docker SSH lane

* test(e2e): keep Electron running after its last window closes on Linux in the reattach spec

* test(e2e): sync the SSH reattach spec on state, not sleeps

Waits for the SSH partition to persist the open file and for main's SSH
state to report the reattached connection, instead of fixed 3s/30s sleeps
that could pass vacuously on a slow reconnect.
2026-10-07 14:47:53 -04:00
Jinwoo Hong 7bea20d83a test(terminal): pin layout invariants the mirror refactor must keep (#26073) 2026-10-07 14:46:12 -04:00
Jinwoo Hong a0026588e4 terminal: window says where each new terminal goes (no behavior change) (#26078)
* refactor(terminal): renderer senders attach placement to pty spawn, report-only

The window now says where each fresh terminal goes: the first pane of a tab
sends new-tab with the tab's creation fields, a split sends the parent leaf,
direction, ratio and the tree after the split, background launches send the
same, and a Codex restart that rebuilds a rootless layout sends root. A
reattach sends nothing. Main still only records whether placement names the
tab its binding write picks; its 'absent' value now counts only new leaves
without placement, so it reads the fallback mint and graft directly.

Extends the shared placement type with optional row, ratio and proposedRoot,
each dropped alone if malformed, so older and newer peers keep working.

* fix(terminal): background agent placement row claims no launchAgent the tab lacks

The adopted tab is created without launchAgent, so a row naming one would
diverge from the renderer's tab once main applies placement.

* fix(terminal): a before-split publishes its pane with the final tree order

wrapInSplit now places the new pane first itself, so the pane-created
handler's proposedRoot sees the real post-split tree instead of the order
before the subtree split moved it.

* test(terminal): type the background-terminal tab patch precisely
2026-10-07 14:45:26 -04:00
Neil 7f43d9b2c2 test(startup): advance mocked display readiness polling (#26283) 2026-10-07 11:12:25 -07:00
Brennan Benson 98f0193afb feat(native-chat): show agent-written visuals inline in structured chats (#26103)
* feat(native-chat): visual directive grammar and host read for a chat's visuals folder

A shared grammar for the ::orca-visual{file="..." title="..."} reply line,
the per-chat visuals folder location on the owning host, and the
agentSession.readVisual runtime method that reads one visual with lexical and
canonical containment, a 512 KiB bounded read and UTF-8 refusal.

* feat(native-chat): render chat visuals inline and in the right sidebar

Native-chat assistant replies render a ::orca-visual{...} line as the chat's
HTML visual in an opaque, scripts-only sandboxed frame: CSP first, the host
frame navigation guard registered before content runs, live theme without a
reload, fitted height, links opened in the viewer's browser only from a real
gesture, lazy mount, and one muted line when the visual cannot be shown.
Open in sidebar shows the same frame in the right sidebar, widened while it
is open and restored after.

* fix(native-chat): visual CI fixes, shared height governor, live-turn streaming hold

Registers agentSession.readVisual from the methods index so the structured
method file stays under its line budget, replaces reflective reads with checked
narrowing, moves the pure height governor to src/shared for mobile, and holds a
half-written directive tail while the turn works (structured text rows carry no
running state).

* fix(native-chat): harden the visual read and link opening

Re-checks after the open that the chat's visuals folder is still the real
directory at Orca's path, reports unexpected filesystem faults by code without
host paths, and lets one click in a visual open at most one page.

* fix(native-chat): keep visual lines out of plain-text reply surfaces; review fixes

One shared helper drops visual lines (outside fenced code) from reply text where
it becomes plain text: the structured status summary that feeds the sidebar row,
dashboard, notifications, phone rows and handoffs, and AI Vault reply previews.
Review fixes: height also counts a pinned body's overflow, only the live
frontier row holds a half-written visual line, the runaway-height stop needs the
same step repeated, and any host refusal evicts the cached revision.

* fix(native-chat): resolve the visuals folder without the removed journal-paths helper

Main removed the per-chat journal paths and the journal database's state
directory; the visuals folder keeps the same sha256 layout on its own and the
read method uses the profile state directory the chat host is opened in.

* fix(native-chat): review round 2 fixes; copy a reply without its visual lines

Reply previews in Agent Session History drop visual lines per text part before
lines are folded; the frame adds a body's overflow only when the body really
overflows; fence tracking follows CommonMark closers and openers; the copy
button copies a reply without visual lines; a coded read fault keeps its cause.

* fix(native-chat): update the frame's theme ref after render; read the visuals folder pair at once

* test(native-chat): declare agentSession.readVisual on the cross-version agent-session surface

* fix(native-chat): copying a reply keeps its code blocks and indentation

Removing visual lines now closes only the gap each removal leaves, instead of
collapsing blank lines across the whole reply and trimming its indentation; the
visuals folder is checked parent first again so a broken path answers the same
way every time.
2026-10-07 11:11:35 -07:00
Brennan Benson 80d45095d2 fix(native-chat): a message kept after a quit or close waits in order and follows the chat's next turn (#25960)
* fix(native-chat): drop the queue-paused header and Resume button

A Stop, a restart or /clear holds the queued cards. The hold stays; only the
header row naming why, and its Resume button, go. A held card shows no
caption, and its own Steer, or any new message, releases the queue.

* test(native-chat): type the unknown hold reason a newer host may publish

* fix(native-chat): a held card offers Send, not Steer, when no turn runs

Steer vs Send now follows whether a turn is running, not the card's hold,
so a card held after a Stop, a restart or /clear reads Send.

* fix(native-chat): the queue sends past held cards instead of stalling behind them

A card queued after a Stop (or written after a restart or /clear) sent only
once the cards held before it were released; with no header to explain or
release the hold, it sat silently. The next sendable card now skips held
cards; a returned card still blocks what is behind it.

* fix(native-chat): the queue's send of a card is the person's turn, so held cards follow it

After a Stop, a card queued later sent past the held cards, but the queue
recorded that send as Orca's own turn. It never ended the Stop's pause, so
the held cards then waited forever with nothing on the card saying why.

A queued card is always something the person wrote: only the client send
RPC may now create one. The queue's send of it is therefore recorded as the
person's turn, which ends the Stop's pause once the agent takes it, and the
held cards then drain in order.

* fix(native-chat): a queued card carries its author, so the queue's send of it is that author's turn

Main now lets Orca's own sends ask to queue (sendAgentTurn's 'queue' delivery),
so "every card is a person's" no longer holds by refusing host sends. Each card
records who wrote it (the submission's client/host vocabulary) in a new nullable
column; the drain records that origin, so a person's card ends a Stop's pause
and Orca's does not. /clear carries the author. Rows from before the column
read as a person's. The userSend-only admission gate is removed.

* docs(native-chat): state why an unrecorded card author reads as a person's

* fix(native-chat): a restart holds only cards written before it, and an idle held queue offers Resume

A restart's pause held every waiting card, including one a person typed after the restart while
Orca's own continuation ran, and nothing released it except a per-card Send. It now holds only
cards another host process wrote, the same way a Stop holds only cards queued before it.

The composer's primary button becomes Resume (Play) while nothing is typed, no turn runs and the
host holds a card Resume would send, whatever held it (Stop, restart or /clear). It calls the
existing agentSession.queuedMessagesResume, guarded against a second press in flight.

A card nothing holds keeps the run going between a turn's end and the queue's send of it, so its
Steer no longer flips to Send for the frame in between.

* fix(native-chat): the host publishes which pause holds each queued card

The host published one pause for the whole queue, so a client held every waiting card while it was
set. Between a turn's end and the queue's send of a card queued after a Stop or restart, the
composer could flash Resume and the cards Send, and a card queued after a Stop lost its
"Waiting for your answer" caption.

Each published card now carries an optional `heldBy`: the pause holding it, or null, derived from
the same rule the drain reads. A client holds only those cards; against a host without the field
it falls back to the queue-level pause.

* test(native-chat): Resume needs the queue capability and is disabled whenever Send is

* docs(native-chat): describe per-card holds in the queue contract and table comments

* fix(native-chat): the composer goes from Resume straight to Stop, and Resume returns focus

After Resume, the host lifts the hold in one update and sends the first card in a later one. In
between nothing was running, so the composer's button flashed a disabled Send. A card nothing
holds now keeps the queue's run going for the button too: an empty composer shows Stop, disabled
until the turn starts. Not when the host refuses every send (a rewind whose outcome is unknown,
read from its status), where nothing is coming. The same fix removes the Stop, Send, Stop flip
between queued turns.

Resume disables the button, which dropped keyboard focus; focus now returns to the composer.

* fix(native-chat): the host names the card its queue sends next, so the chat stays working across the gap

A turn's end, or a Resume, and the queue's send of the next card commit as two host updates. In
between nothing was running, so the working status, timer, pickers and composer button flipped
for one update. The client guessed the drain from its own copy of the host's gates, which missed a
/clear-replaced source and covered only the button.

The queue publication now carries `nextQueuedMessageId`: the drain's own next card through the
drain's own gate (`nextStructuredQueuedMessage`, which the drain step now calls), null whenever the
host would refuse the send. The client derives one fact, the queue is about to send, and every
working reader follows it; Stop stays disabled until a turn can be stopped. The client-side copy of
the gates and the status-feed rewind read are removed.

* test(native-chat): the queue's next card survives the coalescer, the reducer and a history page

* test(native-chat): build the snapshot that names the next card through its helper

* feat(native-chat): a held queue keeps its header row, and a new message asks before passing it

The queue's header row ("Queue paused because you interrupted", or Orca
restarted, or you cleared the conversation) comes back above the cards it
holds, with Resume; it names the oldest held card's pause, as the host
publishes it per card, and hides over cards held only on their own or
returned. The header's Resume and the composer's share one in-flight guard.

A held card reads Steer again whether or not a turn runs; a card held on its
own or returned keeps Send.

Sending a message while the header shows (Enter or the button) first asks
"Send message?": Clear queue deletes every card and then sends (a failed
delete sends nothing), Send message sends and keeps the cards, which follow
the new turn, and dismissing sends nothing and keeps the draft. Host
commands send as they are.

* fix(native-chat): the paused row goes while your own message is on its way to lift it

After "Send message" over a held queue, the row kept saying "Queue paused…"
until the agent accepted the new turn. The chat now reads that gap from the
outbox: while this composer's direct send is recorded by the host and not yet
accepted, the controller shows no paused row (and so no Resume or
confirmation). A refusal settles the entry and the row comes back, since the
hold did not lift. Orca's own sends never enter this outbox, and the queue's
send of a card goes under a fresh id, so neither hides it. Nothing is stored.

* fix(native-chat): a "Send message?" choice is taken once, and a failed Clear queue is one toast

The closing dialog stays mounted and clickable through its exit animation,
and a double-click or a held Enter lands twice before any re-render, so
Send message (or Clear queue) could send the captured message twice. The
pending send now lives in a ref that the first choice takes; a second one
finds nothing.

Clear queue deletes one card at a time and stops at the first failure, so a
failed press shows one toast instead of one per card.

The dialog keeps its compact width at desktop sizes and the primitive's
narrow-window gutter (`max-w-sm sm:max-w-sm`, as the other compact
confirmations).

* fix(native-chat): Clear queue's message goes out once, and keeps text typed while it waits

After Clear queue, the message waited in the composer while the cards were
deleted one by one. A second Enter in that window sent it again, and text
typed meanwhile was wiped when the chained send was accepted.

From the Clear queue choice until its message has gone out, the composer's
structured send does nothing. The chained send (and Send message's) now
carries the composition it was taken from, and the composer is cleared on
acceptance only if it still holds exactly that, as host commands already do.

Also: the v1 contract comment names `nextQueuedMessageId` and its absent-
means-null fallback, and the own-send check returns at once on an empty
outbox.

* fix(native-chat): the queue carries on after any turn, in order, and a restart sends nothing by itself

- Any accepted turn ends a Stop's or a /clear's pause, whoever sent it (a person,
  Orca's own messages, or the queue), and so does Resume. The card and submission
  author fields that only fed the old person-only rule are gone.
- The queue sends strictly in order: a card never overtakes a held one.
- After a restart nothing sends by itself and no paused row shows: the chat's next
  turn (the carry-on, or the person's own message) runs first, then the cards.
- Resume and "Send message?" are offered only while nothing runs and no prompt waits.

* fix(native-chat): after a restart no queue pause shows, and a card written before the next turn waits for it too

* fix(native-chat): a quit hands no queued card off, and the paused row goes while any turn that will lift it is on its way

- The queue stops handing cards off when the host tears down. A card sent during
  a quit was refused at close, and that refused send withdrew the chat's restart
  offer, so resuming after the relaunch sent nothing.
- The host publishes no pause while a turn sent after it (your message, Steer, or
  Orca's own) waits for the agent; a refusal shows it again. This replaces the
  client's own-send check.
- A card written after a restart is an ordinary card again: it waits while any
  card from before the restart still waits.

* refactor(native-chat): the host's paused-row-while-a-turn-is-on-its-way check in one expression

* refactor(native-chat): the composer's queue Resume rides the structured transport beside the held queue

* fix(native-chat): the "Send message?" choice ends with the pause it asked about; tests follow main's draft props

- The open dialog closes when the queue's pause lifts under it (Orca's mail, another client's
  Resume, any accepted turn): nothing is sent, the draft stays, and the next Enter sends as
  usual. The pending choice records the hold it was asked under; nothing new is stored.
- The composer-field Resume test passes main's dropScopeKey/draftScopeKey.
- The dialog test expects main's rule: only the sent text leaves the composer.

* fix(native-chat): a message kept after a quit or close waits like every other card, and follows the chat's next turn

A message Orca accepted but never handed to the agent before a quit, crash or
close came back as a card held on its own ("Not sent yet — press Send"): the
queue skipped past it, so later cards sent first, and only the person's own
Send released it. It is now an ordinary card at the head of the queue that
waits, with every card the chat closed with, for the chat's next accepted turn.

One rule for a chat that was not running, derived from the journal: when this
host first opens a chat (after a restart or crash), or a person closes it, and
cards are waiting, a reopen mark is written (a tombstone carrier key, like the
Stop and Resume marks). The cards queued before it wait until a turn is
accepted or Resume comes after it; nothing sends by itself, and no paused row
shows, a Stop's included. The idle sweep's own eviction writes nothing and
changes nothing the person sees. A mark that cannot be written leaves the open
working and holds from the open itself until the next turn; the next open marks
again. A rewind restates the mark. /clear's carried cards also wait unshown.

This replaces the host-instance comparison and its adoption write, and the
'kept' hold (stored 'kept' and legacy 'stopped' holds now read as none).

* fix(native-chat): mark the reopen in the one open path, and keep an idle chat with waiting cards open

Review round 1: the startup restore opened chats past the per-host first-open
mark, so their cards could send by themselves after a quit or crash; a card
mid-hand-off at the open got no mark; a close that left the chat open re-marked
after every new send.

- Every open marks when a card waits, or is mid-hand-off with its send unanswered.
- The idle sweep keeps a chat's handle while cards wait, so its eviction never
  reopens one and stays invisible; it drops it on the next sweep once they leave.
- A person's close marks once; its delivery re-check marks only when it settled
  a send.
- A failed mark holds from where the mark would have gone.
- A Stop made after a reopen shows its row.
- The rig's restart is a real quit and relaunch.

* fix(native-chat): review round 2: restore the dropped Resume and failed-Stop tests; a late mark starts where the chat stopped

- Restores eight tests the previous commit dropped by mistake.
- A mark the delivery loop or a close's re-check writes after a later send starts
  where the chat stopped, so that send still lifts it; a later mark never narrows
  an earlier, wider one.
- Only the idle sweep's own close keeps a chat with a card waiting (or mid-hand-off)
  open, and it still releases an ended child's lease first; a person's close
  drops it as before.

* test(mobile): a host-kept card's test stands in a Stop's pause, as this host publishes no restart pause

Main's #24660 test published queuePause 'restarted', which this branch's wire
type no longer lists, so the mobile tests typecheck ratchet failed.

* refactor(native-chat): settle a restart's leftovers and mark them in one host-lifetime step

Keeps the delivery loop under its line limit after main's Stopping change; no
behaviour change.

* test(native-chat): a kept card's Resume and failed-mark tests quit through the held start's release

Main's #25152 holds the start these tests send into; quitting without releasing it left the quit waiting.
2026-10-07 10:29:52 -07:00
Neil 68c493fa72 test(runtime): advance injected teardown policy clocks (#26267) 2026-10-07 10:28:30 -07:00
Brennan Benson c2f3229362 fix(worktrees): a phone or CLI create keeps a sparse preset only when its folders match (#26104)
* fix(worktrees): the host's create keeps a sparse preset only when its directories match

The window's create records a sparse preset on the new worktree only when the
checked-out directories are exactly that preset's, so an edited selection is
not shown as the preset it began as. The host's create wrote whatever preset
id it was sent. Both now share one check (moved out of the two copies in the
window's create path).

* fix(worktrees): read sparse presets only when a preset was chosen, and never fail the create on them

The shared preset check took the preset list up front, so every create read
the presets even with no preset chosen (the window's create used to skip the
read), and a read that threw failed the create - on the host after
`git worktree add` had already run. The check now takes a reader, returns early
when no preset was sent, and reads inside its try so unreadable presets record
no preset instead of blocking the create. Repeated checked-out directories no
longer match a preset with a different set of directories.
2026-10-07 10:23:29 -07:00
Neil 136c039896 test: import Activity functions directly in eight suites (#26237) 2026-10-07 10:22:35 -07:00
Brennan Benson 91d6ed0f3c fix(worktrees): a phone or CLI create puts the agent in the first orca.yaml default tab, as the desktop does (#26106)
* fix(worktrees): the host's create puts its agent in the first default tab, as the window does

With orca.yaml default tabs, the window's create starts the agent in the first
default tab: that tab takes the template's title and color, and the template's
command does not run beside the agent. The host's create started the agent in
its own untitled tab and then created every default tab, the first one running
its command too, so the workspace had one tab more than the window's create.
Now the agent's tab takes the first template's place.

* fix(worktrees): title the agent's default tab as a tab, so the phone keeps the agent's status

The host's create titled the agent's tab by renaming the terminal, which
writes the pane's own title stamped to outrank every title the agent sets
later. On a headless host the phone reads status from those titles, so the
agent tab showed the template title forever and lost its working/idle
status. The window's create only sets the tab's custom title.

The host now titles the tab the same way: the window gets the tab's custom
title (not counted as a user rename), and a headless host saves it on the
tab and shows it on the phone until the agent titles itself. The pane's
title is left to the agent.

The host also finds a terminal's tab itself from its handle, so the create
over SSH, which never passed the tab, colors the agent's tab too. Every
default tab is dressed by the same step, so the other tabs' titles are now
kept on a headless host as well, and a failed title no longer skips the
color.

* fix(worktrees): find a default tab through the window's leaf once the window adopts its handle

The host found a provisioned terminal's tab only through the runtime's own
pty record. Once the window's graph sync adopts the handle, that lookup
misses, and the default tab's title and color would be dropped with only a
log line. No create reaches it today (the tab is dressed before the window's
first graph sync), but any wait added before provisioning would. It now falls
back to the window's leaf, as renaming a terminal already does.

Tests: the restart case now saves the title to a real session and rebuilds
the phone's tabs from it; a create through createManagedWorktree with an
agent and default tabs spawns the agent and the second template only and
titles the agent's tab "Dev" without counting it as a user rename.

* test(worktrees): drop type assertions from the default-tab tests

Use a runtime subclass for the protected launch scope and provisioning host,
a typed store and notifier, and typed mocks in place of casts, so the
changed-code quality gate passes.
2026-10-07 10:22:29 -07:00
Neil 7d6d7ca6f8 fix(skills): observe archive abort errors before reading (#26246) 2026-10-07 09:20:57 -07:00
Neil a576c0bc4e test(persistence): reuse first constructors in six more suites (#26250) 2026-10-07 08:59:23 -07:00
Neil 76d1b7bf29 test: reuse first SQLite fixture construction in five suites (#26231) 2026-10-07 07:39:43 -07:00
Neil 39710c79b8 test: load renderer fixture builders without the full Store (#26217) 2026-10-07 07:24:25 -07:00
Neil 4a0418d9c1 test: import Activity fixture builders from their existing owners (#26229) 2026-10-07 07:21:16 -07:00
Neil 40d35fb7bf test: run real SSH command contracts under Node (#26226) 2026-10-07 06:53:40 -07:00
Neil e381443886 test: advance retry policy clocks without real waits (#26220) 2026-10-07 06:43:57 -07:00
github-actions[bot] 0d82899972 Update README downloads badge 2026-10-07 12:42:41 +00:00
Neil 7f8b4a8dd4 test: isolate remote upload cases from bootstrap writes and clock rollback (#26202) 2026-10-07 05:36:43 -07:00
Neil 2dc98c93de perf(tests): reuse composer decisions without loading the hook (#26179) 2026-10-07 04:43:34 -07:00
Neil b8c967196b test(ssh): isolate permission probes from real PID reuse (#26190) 2026-10-07 04:09:51 -07:00
Neil e4da27e1e9 perf(tests): reuse the first UI-state store constructor per case (#26185) 2026-10-07 04:08:28 -07:00
2584ed906e feat: play video and music files in mobile previews (#26148)
Play workspace video and music files on mobile using bounded authenticated downloads and native media controls.

Adversarial review fixes cover rewritten files and Android policy tests. Includes iOS simulator screenshots and a playback recording in PR #26148.

Co-authored-by: ChangJun Park <40492343+ckdwns9121@users.noreply.github.com>
Co-authored-by: Lirone Levy <lirone88@outlook.fr>
Co-authored-by: dupi <david.li.du@gmail.com>
Co-authored-by: John Cusack <5961784+John-Cusack@users.noreply.github.com>
2026-10-07 03:41:42 -07:00
Brennan Benson 520033b0e1 fix(native-chat): fold thoughts and tool calls between replies into one row (#26048)
* fix(native-chat): fold thoughts and tool calls between replies into one row

A model that thinks before every tool call drew a 'Thought for Ns' row between
each call, and each thought also stopped the calls from merging, so a turn read
as dozens of alternating rows. The transcript now draws an unbroken stretch of
tool calls and thoughts within a turn as one tool run; thoughts read in order
inside it when opened and are not counted in its label.

* fix(native-chat): keep work runs whole while a turn edits files

A live turn that had edited a file kept splitting its newest call or
thought off the run (the row carrying the turn's diff rollup could not
join), so a lone "Thought for Ns" row came back at the bottom.

- Run boundaries are now one pure pass in src/shared
  (nativeChatWorkRunSpans), applied over the built transcript slots.
  Section heads end runs by sitting between rows, and rows inside an
  expanded subagent section are grouped the same way.
- The row carrying the turn's diff rollup joins as the run's last
  member; the run takes the rollup from it and the bar from its head.
- A collapsed run reserves one tool row plus a lead's words, not every
  thought's text.
- Lone rows and runs draw through the same MessageRow, with one flat
  keyed line list, so a row becoming a run no longer remounts (a
  revealed diff card no longer re-fires its scroll).
- A run opens by default while the reader holds one of its thoughts
  open; the run's own choice still wins.
- A drawn thought with no visible text joins a run instead of
  splitting it.

* fix(native-chat): pair each row's result with its own call; keep open thoughts out of runs

- A structured journal row's call and result now share one id (the
  provider's call id, else the row's own id), so a diff or output can no
  longer pair with an earlier unanswered call in the same work run or
  folded message. A quiet command followed by edits now shows each diff
  card under its own call, and the rollup's reveal lands on it.
- A thought still being thought, while its turn or subagent works, keeps
  its own row after the run and joins once it ends.
- Removed the rule that opened a run because a thought inside it was
  open; a thought keeps its own open state inside its run.

* fix(native-chat): an open thought the live line lets go of joins its run

Only a subagent section keeps a still-running thought out of its run.
At the top level the live line owns the open thought; when it lets go
(a Stop in flight, a prompt waiting on the reader) the thought now joins
the run at once instead of flashing as its own row until the turn ends.

* test(native-chat): guard the run row memo; name the slot post-pass test after it

Settled work runs must not rebuild their lines and thoughts on every stream
frame of the live run. Renames the slot-level run test so it no longer
shares a name with the shared span test.
2026-10-07 03:35:08 -07:00
5cafefe726 Phase 3: every SSH host runs a managed Orca server (orcad), replacing the relay (#24863)
* Revert "revert: take the 26 Phase 3 (#16741 port) PRs back out of main (#24559)"

This reverts commit 5f308bfa9c.

* feat(orcad): Windows remote primitives for managed orcad hosts (W1) (#24525)

* feat(orcad): Windows remote primitives for managed orcad hosts (W1)

* refactor(orcad): run Windows host ops as node.exe with plain argv, no PowerShell hop

* fix(orcad): refuse secret-shaped names on the breakaway launcher's --env

* fix(orcad): name the secret env guard for its role

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9) (#24521)

* feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9)

The destination half of a catalog migration: an orcad stages a T6-7 manifest
(repositories, project groups, folder workspaces, dormant session, client,
automation and worktree metadata, retired names, scrollback snapshots) with
exclusive claims, then commits it with a receipt so a retried commit returns
the same receipt and never imports twice. Served as orcad.migration.* runtime
RPC behind the orcad.migration-catalog.v1 capability; the client refuses a
host without it or with method-not-found, and any other failure is left for
the caller to recheck. Dormant only: no live PTY projection. Inert on the
desktop until T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(rpc): catalog the orcad.migration params in the shared contract; name the catalog-import install target

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(ssh): journal and fence an SSH host for dormant migration, gated on proven terminal exit (#16741 T8-c1+c2) (#24522)

A migration from a relay-hosted SSH target into a managed orcad now starts with
a journal in its own sidecar directory, then the target's managed-owner fence,
then a profile flush, before any remote call. A fence with no journal is
unverifiable and never released; a journal whose fence is gone is stale and
grants nothing; an unreadable journal fails closed. The fence requires every
terminal the target ever leased to be proven exited, checked before the fence
(with the relay's process list) and again under it. Same-owner claims now need
the durable record that explains them. Inert until T8-c4/T6-10.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): run orcad itself on Windows hosts (W2) (#24529)

* feat(orcad): run orcad itself on Windows hosts (W2)

* test(orcad): load the ConPTY smoke's addon from out/orcad so the temp slot can be removed

* test(orcad): skip the foreign-uid lock case when running as root

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a connected but unused SSH host previews as movable (#24609)

The untransferred-dependency census counted activeConnectionIdsAtShutdown naming the
target as workspace-session state. The renderer rewrites that list on every connection
change, so merely connecting to an empty host blocked the move. It is a reconnect hint;
the remote work it can stand for is counted on its own. The empty-target claim check
likewise ignores global-field copies inside the host's session partition.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3) (#24523)

* feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3)

The coordinator re-checks before every stage and commit that the fenced source
still exports the journaled manifest, carries no untransferable state and
started no terminal. A lost answer is re-read from the destination's catalog
state; only a committed read whose receipt matches the journal advances it,
and the journal is on disk before anything returns. Abort releases the fence
only on proof the destination holds nothing, or on an unsupported destination
before anything was staged, and never once the destination committed. Codes
against T6-9's catalog client through an injected interface. Inert until
T8-c4/T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(ssh): import the T6-9 client's unsupported refusal instead of mirroring it

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3) (#24563)

* feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3)

* test(orcad): exhaustive op switch in the Windows lifecycle fake

* test(ssh): narrow the Windows host-cell descriptor by lane before building a relay cell

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11) (#24608)

* feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11)

* fix(serve): keep orcad selection app-side and wait out Windows temp cleanup

* refactor(orcad): move the data-root privacy check out of the instance lock

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch (#24619)

* test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch

* ci(e2e): install ripgrep for the serve mode-switch job's window-manager wait

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4) (#24562)

* feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4)

The conversion entry resumes or takes the fence, deploys and pairs the
managed server into it, marks the server as migrated, then stages and commits
the dormant catalog. Every step is keyed by the journal, so a repeat after a
crash, deferral or lost reply resumes the same migration. Status reports an
unfinished migration, and rollback is refused while one runs or when the
rollback snapshot predates the migrated catalog. Inert until T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* style: oxfmt the c4 conversion and maintenance files

* fix(ssh): name the fake migration destination's type so declarations stay portable

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4) (#24570)

* feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4)

* fix(orcad): accept a managed stop request whose lock path is spelled with Windows client separators

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5) (#24565)

* feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5)

Once the journal records destination-committed, the source profile drops the
manifest's repositories, folder workspaces and unreferenced project groups,
its dormant session, automation, client and worktree state, and the leases
the fence proved exited. The profile flushes, the retirement is verified,
the journal moves to source-retired and compacts once the server matches.
A retry after any crash repeats idempotent work. The fenced target stays: it
carries the managed server's tunnel. Conversion now ends retired. Inert until
T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(orcad): retirement drops the migrated host from the reconnect hint

The census no longer treats activeConnectionIdsAtShutdown as untransferable (#24609), so
retirement must remove the target from it; otherwise a restart dials a host that is now a
managed server.

* style: oxfmt the c5 conversion file

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1) (#24579)

* feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1)

* test(orcad): start the Windows lane's exec spy after the relay gate prelude restores its own

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI) (#24590)

* feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI)

Adds a Managed servers section under Remote servers (deploy an empty server,
status with deferred-update and migration states, update, rollback, recover,
stop and cancel-stop, and SSH access for paired servers), and a Move to managed
server action on connected macOS and Linux SSH hosts with a preflight summary,
a terminals-closed confirmation and a resumable progress view. Main wires the
conversion to the relay's process list, the direct session and the T6-9 catalog
client. Everything is hidden until the new experimental setting is turned on,
and Windows SSH hosts are never offered. Merges the T6-9 branch (#24521) until
it lands on the integration branch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settings): align managed-server form controls and name the section the setting reveals

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(orcad): ask the relay with an absolute deadline via the W5 terminal-gate lister

The conversion wiring passed a relative 10 s as listProcesses' deadlineMs, which the
provider reads as an absolute time, so every relay inventory timed out after 1 ms and
the terminal gate could never prove exit. Adopt #24579's lister verbatim so the stacks
merge cleanly.

* feat(settings): name blocking saved state in plain, localized words

The move preview listed internal dependency ids such as workspace-session; each kind
now has its own catalog entry.

* feat(settings): offer managed servers and the move on Windows SSH hosts

W1-W5 are on the integration branch, so a Windows relay-hosted host can deploy, convert
and retire like a POSIX one. The move still waits for a connected relay that reported
its platform.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix (#24865)

* test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix

* test(ci): expect the orcad Windows host cells in the SSH Windows hosts workflow

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(orcad): wait for killed terminal daemons to exit before removing their temp profiles (#24871)

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(updater): read rollout kill switches from the update-campaign payload, all inactive (#24867)

The nudge request Orca already polls may now carry an optional versioned rollout block
naming the Node runtime flips. A typed reader resolves each flip with kill-switch, version
range and install-id-bucketed percent semantics, falling back to the baked value (every flip
inactive) when the block is absent, invalid or never read. No consumer reads it yet.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(telemetry): report Windows security-software refusals and unverifiable or failed runtime checks (#24866)

ssh_remote_runtime_resolved dropped the self-test's security_software refusal to 'none' and
sent nothing when a self-test was unverifiable or failed, because those attempts throw before
a rung settles. Add the refusal value, self_test 'unverifiable', and an outcome field
(resolved | unverifiable | failed) deduplicated per host and outcome per session.

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(settings): move WorktreeVisibilityDefaults into its own module (#24884)

global-settings-types.ts sits at the 300-line max-lines ceiling; merging main's two new
agent-state-rules settings with experimentalManagedServers put it at 301. The worktree
visibility defaults type moves next to the other visibility types and is re-exported so its
38 importers are unchanged.

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(watcher): move the supervisor's child, terminating child and canary into a child slot (#24887)

* refactor(watcher): move the supervisor's child, terminating child and canary into a child slot

* fix(watcher,runtime): take the child slot's child type from the shared wrapper, and stub main's title-display clear in the projection test

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* ci(adhoc): build and ship the orcad template in adhoc macOS and Windows builds (#24969)

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy (#24973)

* feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy

* fix(orcad): keep the in-use slot plus the two most recent others, and prove same-version reuse survives eviction

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* Phase 3: managed orcad on every SSH connect, with downgrade-safe fence (#24975, #24979)

* feat(orcad): fence managed hosts outside owner and keep converted projects for downgrades

Shipped builds hide any SSH target with an owner, so a downgrade made a converted host and its
projects vanish. The managed fence moves to a new orcadFence field (older builds keep but ignore
it), phase-3 owner fences migrate on load, and managed hosts stay visible but refuse a direct
relay. Conversion now stops at destination-committed with sourceRetainedAt; source retirement
waits on the baked-off orcad-source-retirement rollout flag. This build hides retained rows, and
a start that finds an older build changed them marks the host sourceChangedAt (relay, needs a
new move) instead of merging a second manifest.

* feat(ssh): every SSH host runs managed orcad, decided on each connect (#24979)

* feat(settings): managed servers are no longer experimental

Remove experimentalManagedServers and its gates; the Managed servers section always shows.
Loading drops a stored value, which an older build reads back as its own default (off).

* feat(ssh): every SSH host runs managed orcad, decided on each connect

Before a relay is started, the connect decides the host's server: a converted host connects
through its tunnel (retiring a retained source once orcad-source-retirement is on), an empty
host deploys orcad, and a host with Orca state converts through the journaled migration. A host
with live or unproven relay terminals keeps the relay this session and converts later. A
refusal for any other reason keeps the relay and names the blocker. A host orcad can't run on
(unsupported target, no template, native preflight, runtime self-test) releases any claim or
fence it took, records why with this app version, and keeps the pinned-relay ladder.

Progress and the decision ride an optional SshConnectionState.managedServer field (dropped by
older clients' admission). SSH Hosts shows each host's server status; the manual move dialog
is removed.

* fix(ssh): always await the connect's server decision, after the provider authority rotates

The decision now runs after the old session and transport are torn down and the authority has
rotated synchronously, so concurrent connects still join one attempt; a shutdown that began
during the decision wins over the rotation. The IPC tests use an async double, no sync branch.

* test(renderer): the IPC events store double carries its SSH connection states

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* test(ssh): prove connect-time conversion to managed orcad on real Linux and Windows hosts (#24981)

* test(ssh): prove connect-time conversion to managed orcad on real Linux and Windows hosts, and the downgrade view

* fix(e2e): read the SSH host's session partition, provision the convert cell's account, and keep rollout file overrides out of packaged builds

* test(e2e): report the full relay state when the convert poll times out

* test(e2e): require the relay to list no terminals before the converting connect

* test(e2e): require no running-terminal lease before the converting connect

* test(e2e): report the host's terminal leases when the converting connect keeps the relay

* test(e2e): end the relay era with no relay shell left to respawn

* test(e2e): settle before the converting connect and name the tabs a blocking shell belongs to

* test(e2e): use the exited relay tab as the session tab, since any mounted tab starts a shell

* test(e2e): carry an editor tab through the conversion instead of a terminal tab

* test(e2e): log the conversion census inputs before the converting connect

* fix(orcad): log which state blocked a refused conversion

* test(e2e): log both session partitions before the converting connect

* fix(orcad): a source partition's copy of focus on another host no longer blocks conversion

* fix(ssh): an ssh2 forward on port 0 reports the port it bound, so managed tunnels pair

* test(e2e): give the conversion its full budget again

* test(e2e): report the migration journal phase when the conversion stalls

* test(e2e): report the connect's own result and main's state when the conversion stalls

* fix(ssh): a converted host's managed state reaches the renderer instead of staying on 'converting'

* test(e2e): print a failed server call's response

* test(e2e): give server calls the budget a fresh server's first inventory needs

* test(e2e): read the converted worktree's tabs with a scoped session.tabs.list

* test(e2e): log the converted worktree's tabs instead of asserting them, pending the server-side fix

* test(e2e): prove retirement by the dropped source rows; the journal compacts away after it

* test(e2e): drop the conversion diagnostics now the cell passes

* test(e2e): fail a hung disconnect or connect with main's state instead of the whole budget

* test(e2e): convert an upgraded relay-era profile's host on its first connect, on Docker and Windows

* test(e2e): seed the relay-era target the way addTarget registers it

* fix(ssh): a stale ssh2 forward drops a late connection instead of crashing main on 'Not connected'

* fix(orcad): the active-slot readiness probe reads Windows hosts through the host script

* ci(ssh-windows): let only the convert cell's account open the SSH local forward its managed server needs

* test(e2e): match server paths in their JSON-escaped form, for Windows backslashes

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): move what an older build added to a converted host, or keep the server's version (#24980)

* feat(orcad): move what an older build added to a converted host, or keep the server's version

A host an older build changed stays on the relay with two actions. 'Move the new projects' runs a
fresh journaled conversion of only what the host's earlier migrations didn't move: the source is
viewed with those migrations retired from it, so nothing that overlaps the server is merged, and
a row the server already holds fails the whole move at stage, before any commit. Its journal
supersedes the chain head; retained-source checks compare against its baseline, and retirement
retires every manifest in the chain only after the newest committed. 'Keep the server's version'
records the current source as the baseline and returns the host to its managed server.

* test(ipc): the runtime environment handler contract lists the delta-move channels

* test(native-chat): snapshot the journal directory after the attach's restart-offer lock is released

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): converted hosts keep their editor tabs, forwarding-refusing hosts stay on the relay, listAll settles (#25099)

* fix(orcad): publish the headless graph so session.tabs.listAll settles instead of hanging

* fix(ssh): a system SSH forward on port 0 picks a free port first and reports it

* fix(runtime): a headless host lists and closes the editor tabs its session holds, so migrated editors reach clients

* fix(ssh): keep a host that refuses TCP forwarding on the relay, and release a conversion it stranded

* test(e2e): assert the migrated editor tab and a settled listAll, and keep a forwarding-refusing host on the relay

* fix: restore the journal import after rebase, type the probe's failure code, and update headless-graph test seams

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): client focus and 'local'-stamped tabs no longer refuse a host's move to its managed server (#25110)

* fix(orcad): client focus and 'local'-stamped tabs no longer refuse a host's move to its managed server

* fix(orcad): a v1.4.218 profile focused on the SSH worktree converts, and a refusal names what blocks it

The debounced session writer never patched activeWorkspaceKey or activeWorkspaceExecutionHostId,
so the first focus stayed on disk. The all-dependency census also counted global focus copies in
every non-source partition.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): offer to move a host whose open terminals keep it on the relay (#25097)

A host with any open SSH terminal never converted: the connect gate keeps the relay while relay
terminals are live, and an open tab respawns them on every connect. The first such connect per
host per app version now marks the relay status with offerMove (recorded as
managedServerMoveOffered), which toasts "Move <host> ... Its N open terminals will restart." The
SSH Hosts status line keeps a "Move to managed server" action while terminals are live.

Confirming runs ssh:moveToManagedServer: stop the host's relay terminals through the extracted
ssh:terminateSessions path, re-run the connect gate's terminal census, refuse on anything but
exited, then reconnect so the connect-time decision runs the journaled conversion.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): reach managed orcad over an SSH stdio bridge where sshd refuses port forwarding (#25120)

* feat(orcad): reach managed orcad over an SSH stdio bridge where sshd refuses port forwarding

Hosts with AllowTcpForwarding no were kept on the relay. The managed tunnel now
probes forwarding each time it starts and, on refusal, serves the same local port
through a second provider: each accepted socket opens one SSH exec channel running
a small bridge on the host's pinned Node, which dials orcad's loopback port.
POSIX hosts run it with node -e; Windows hosts run it as the content-addressed
host script's stdio-bridge op with base64 line framing. Bridges are capped at 8
per connection under sshd's MaxSessions default, and a lost channel only drops its
socket. Only a host where even the bridge cannot run keeps the relay, recorded as
ssh_tunnel_unavailable; the older tcp_forwarding_refused record is retried.

* test(e2e): prove the stdio bridge by refused forwarding plus a working call

* test(e2e): connect the refused-forwarding host without the relay-only repo step

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): report each connect's server decision, and show orcad.log's last lines when setup fails (#25118)

* telemetry: one enum-only event per connect decision (outcome, reason, tunnel transport, host platform, duration), plus conversion start/commit/fail, deploy failures, and the per-host move offer and its result
* errors: deploy, launch and rollback failures carry the last 40 redacted lines of the host's orcad.log, read over SSH (Windows through the node.exe host script)
* SSH Hosts: a deferred or failed setup offers its reason, log tail included, under Details

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(startup): Windows never crashes resolving userData when roaming AppData is unavailable (#25113)

* fix(startup): pin Windows appData and userData before anything resolves them

A Windows session without a loaded profile (e.g. orca serve over SSH) can fail the
roaming AppData known-folder lookup. Electron 43 then falls through to Chromium's
userData provider and crashes natively. Resolve appData first (falling back to
APPDATA, then USERPROFILE\AppData\Roaming), and set userData explicitly so
Electron's provider never runs.

* test(startup): remove the AppData fixture through the retrying helper

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): keep a terminal the previous Orca version's relay still runs instead of replacing it (#25124)

After an app update the new relay answers "not found" for a PTY the previous build's relay still
runs, because the old relay refuses this build's handshake. The client read that as absence: it
expired the lease and the pane cold-restored into an empty shell while the user's shell kept
running, unreachable. Each deploy now takes a census of this target's older relay endpoints; while
one is live or unverifiable, a not-found reattach keeps the lease and the pane binding, and the
pane says the terminal is still running under the previous Orca version. A detached lease also
keeps blocking managed-server conversion until that terminal exits.

The cross-version harness now extracts src/relay, and a new test drives v1.4.218's relay socket and
grace lifecycle with this build's endpoint probe: the probe leaves no grace deadline and reads the
old relay as live work.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): update a managed server on connect when it runs an older build and is idle (#25122)

* feat(orcad): update a managed server on connect when it runs an older build and is idle

A connect to a managed SSH host now runs the Managed servers update when the host's
orcad differs from this app's bundled build, the template carries the host's target,
and the update planner finds no live or uncounted terminals. A rejected candidate is
restored through the activation journal; the reason is recorded per app version so
later connects don't retry it. A host a newer Orca activated is never downgraded:
the activation record now names the app version behind each build, and an explicit
rollback holds the build it left.

* test(e2e): connect without a racing disconnect after relaunch, and report each attempt

On launch the app already reaches the managed server through its tunnel; a disconnect racing that
restore cancelled the connect that runs the update.

* feat(orcad): run the update check when the launch restores a managed server's tunnel

An auto-restored host may never see an SSH connect, so its server would never update. The tunnel
restore now runs the same check, once per server per session and off the caller's path, through
the shared update-check module; a server mid-migration is left alone.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): retry the pinned Node download through transient network failures (#25154)

A Chromium network change (net::ERR_NETWORK_CHANGED) during the pinned Node
download failed the deploy and sent the host back to the relay. The archive
download now retries up to three more times, after 1s, 3s and 9s, on dropped
connections, timeouts, stalls and retryable HTTP statuses, removing the partial
file first; checksum mismatches, other HTTP errors and cancels stay final. The
transient-error classifier moves from the speech download to
src/main/network/transient-download-error.ts so both share it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(serve): run orca serve on orcad by default, with ORCA_SERVE_RUNTIME=electron as the opt-out (#24972)

* feat(serve): run orca serve on orcad by default, with ORCA_SERVE_RUNTIME=electron as the opt-out

* test(serve): read Electron serve's pretty-printed readiness in the CLI mode-switch e2e

* feat(serve): gate D7 on Windows and serve on orcad there by default

* test(serve): take the profile lock in the CLI mode-switch e2e, and keep Windows profile logs on failure

* test(serve): tell Electron and orcad serve apart by readiness health, and trace Windows startup

* ci(e2e): dump Electron's native log and stack on the Windows serve mode-switch job

* fix(serve): keep Windows on Electron serve until it can adopt orcad's daemon

The Windows D7 job shows Electron serve exiting before its window when it relaunches
onto a terminal daemon orcad forked. Restore the win32 fallback and skip that case
there as a known gap; the follow-up PR fixes it and re-flips Windows.

* fix(serve): let ORCA_SERVE_RUNTIME=orcad opt in on Windows while Electron stays the default

* test(serve): skip the Windows D7 cases where orcad forks the daemon, and stop cleanup hiding a failed relaunch

Test 3 hits the same Windows gap as test 2: orcad forks its own daemon there, and Electron
crashes at startup beside it. A failed relaunch also made dispose close the old, already
closed app, whose throw replaced the launch error.

* test(serve): run every Windows D7 case, with the isolated home's AppData in place

Electron 43 crashes natively (0xFFFF7003) when it resolves userData and Windows cannot find
roaming AppData. The e2e home isolation points USERPROFILE at a fresh folder with no AppData,
so later launches hit that. The harness now creates it, every D7 case runs on Windows again,
and a new case proves Electron serve starts beside another profile's live orcad daemon.

* test(serve): retry removing a Windows e2e profile while a killed daemon releases it

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): resume a terminal the previous Orca version's relay still runs, through that relay's own bridge (#25170)

* feat(ssh): resume a terminal the previous Orca version's relay still runs, through that relay's own bridge

After an app update the previous relay keeps the user's shells alive but refuses this build's
handshake. Its own relay.js --connect, run from its own version directory, presents its own bundle
hash, so on POSIX hosts the client now reaches it that way: a pane whose reattach the current relay
held for an older relay opens a route through the old bridge, takes the PTY owner role without
output flow control, reattaches the PTY with its replay, and routes every later operation on that
id to the old relay. When the last pane a route serves exits, the route hangs up and the old relay's
own idle grace retires it. Windows hosts, relocated short sockets and unreachable bridges keep the
held-pane behaviour.

The cross-version harness now builds v1.4.218's relay from its tagged sources and runs it as the
real detached daemon: a shipped client leaves a shell in it, this build resumes the pane through the
old bridge, types into it, sees its output, and watches the old relay exit on its own after the
shell does.

* fix(ssh): read the legacy relay router through its instance

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a host re-upgraded after a downgrade can move what the older build added (#25179)

The delta view left the moved projects' session state in place whenever it held anything
unmovable, so it then counted against the delta. Relay PTY bindings, shutdown markers and the
relay consumer's recovery record blocked every move although the terminal gate already proves
those terminals exited before any move commits; the move now drops them instead.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(worktree-ps): report a lost-contact host's terminals as unverifiable, not live:0 pty:no (#25167)

During a network drop to a relay-served SSH host, every PTY on it reads as an
unconfirmed exit, so worktree ps skipped them and printed live:0 pty:no for a
terminal that was still running. The listing now counts terminals whose liveness
verdict is unverifiable into a new optional unverifiableTerminalCount, and the CLI
prints live:unverifiable pty:unverifiable (JSON carries the same word) instead of
zero. A host-confirmed exit still reads as zero.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): a held pane shows no client OS or shell, and the boundary doc says POSIX hosts resume it (#25194)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): tunnel to the port managed orcad bound and verify it is ours (#25182)

When another Orca already listens on 6768, orcad binds a different port. The
tunnel kept forwarding to 6768, reached the other runtime, was rejected with
4001, and the connect still reported a managed server.

The tunnel now reads orcad's bound port from its active slot's readiness
(falling back to the persisted port for slots without one) and, after the
forward is up, proves the server answering is the paired runtime. On a
mismatch it re-reads the port once and fails with orcad_identity_mismatch
instead of reporting managed.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a converted host shows only its managed server's rows, and its old editor tabs load there (#25193)

Retained relay-era project groups and the per-host SSH catalog now hide like repos and folders.
On the managed transition the renderer reloads server names, groups, folders and worktrees, then
drops the host's relay-era rows a local refresh would keep. A restored tab with no host stamp in
a workspace now owned by a managed server takes that server as owner, in place.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(orcad): ship the port-scan worker so managed servers detect workspace ports (#25197)

orcad resolves port-scan-command-worker-entry.js beside orcad.js, but the orcad
build never emitted it and ORCAD_ARTIFACTS never listed it, so every managed
server logged 'probe worker unavailable' and had no port detection. The build now
emits it with the other children, and the artifact list carries it, so it is
uploaded, hashed and covered by the template contract.

A new test bundles orcad and its children and fails when the bundle names a
worker or child entry the slot does not ship. Two existing gaps it found,
session-scanner-service-entry.js and wsl-transcript-fs-process-entry.js, are
listed as known and may only shrink.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): keep retrying when a reconnect loses the socket before the SSH banner (#25195)

After a network outage, a port forwarder or NAT can accept the TCP
connection and then close it before the server sends its banner. ssh2
reports that as 'Connection lost before handshake' with no errno, so the
reconnect ladder classified it as permanent and published 'error' with
no retry scheduled. Treat it as recoverable on the bounded ladder only;
the initial connect keeps its narrow classifier.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(serve): serve on orcad by default on Windows too (#25162)

D7 now runs every case on Windows (orcad-serve-mode-switch-windows), including Electron
serve adopting a daemon orcad forked, so Windows no longer needs the Electron default.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): the terminal gate asks the relays before treating a detached terminal as running (#25200)

* fix(orcad): the terminal gate asks the relays before treating a detached terminal as running, and retiring a host drops its relay recovery record

* fix(ssh): earlier-relay census gaps and an unanswered relay stay unverifiable; asking the terminal gate changes nothing

- the gate is read-only; the conversion and delta move retire proven detached leases themselves
- an expired lease also needs every earlier-build relay to answer before it reads as exited
- a failed, input-less or truncated census marks its older relays unverifiable, so reattach holds
- a legacy relay route closes only when no attach or listing still awaits it
- a disposed session forgets the census it started

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): recover interrupted conversions and delta moves, and trim dead migration code (#25226)

* refactor(migration): drop the unused full dependency census

* refactor(migration): one table for the routed UI fields a migration carries

* refactor(migration): share the session focus field list and fix stale claim comments

* fix(session): keep a fenced SSH host's source session partition intact through renderer saves

The renderer cannot see a fenced host's repos or folders, so hydration drops their
worktree rows; a save that still routes any row to ssh:<target> (a folder workspace
key stays valid) rewrote that partition without them. That changes the migration's
source manifest and loses the session a downgraded build reads back. Main now
ignores renderer writes to a fenced, unchanged host's partition.

* fix(migration): resume or back out unfinished conversions and delta moves, serialize delta moves, clean stale journals

- On connect, a registered server whose conversion never committed resumes the commit; a failure aborts it through the destination, unregisters the server and releases the fence.
- An interrupted delta move keeps its journal and gets its changed mark back; the next move resumes it from the journaled manifest or aborts it when the source has moved on.
- Delta moves run under the target lifecycle queue and re-check the head, so two concurrent moves cannot write two journal heads.
- Delta checks re-read the live source instead of a frozen copy.
- Stale journals from a stopped or never-registered server are cleared.
- Retiring a chain goes newest first and compacts only at the end.
- The journal schema tolerates fields a newer build adds.

* fix(migration): hide source rows only after commit, and retire what an older build added to a moved project

- Source rows and the renderer session guard share one rule: hidden once the server committed, shown while a migration is still in flight.
- A delta move a crash interrupted gets its changed mark back on the next start.
- Retirement removes worktree metadata and automations by moved-project scope, so an older build's additions no longer fail the leftover check forever; that check ignores rows of other migrations in the chain.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): survive a slow first daemon start and close review gaps (#25213)

- Give the terminal daemon 30 s to start on Windows and retry the spawn once before
  falling back, so a cold first deploy is not refused as daemonless.
- Say in the activation refusal that orcad.log holds the daemon's error and that the
  next connect retries.
- Managed stop completion now waits out daemon retirement plus the shutdown deadline.
- Publish the instance lock atomically and reclaim an abandoned torn lock.
- Fix the stop listener closing before it was defined on an already-present request.
- `orca serve`'s cache prune keeps the slot a running local orcad uses.
- A timed-out daemon retirement reopens admission once it is refused; a retirement that
  may have reached the daemon stays fenced.
- A headless host keeps an editor tab's unsaved draft unless the close is forced.
- Unverifiable terminals are attributed by the host's PTY record, like live ones.
- An explicit --user-data-dir is no longer overridden on Windows.
- Read Windows daemon process identity from the process table, not PowerShell.
- Remove the unread data-event incarnationId and daemon health runtime fields, share
  one process-alive and error-code check, and fold the websocket limits file back.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(cli): stop, update, roll back and recover managed Orca servers from the CLI (#25201)

* feat(cli): stop, update, roll back and recover managed Orca servers from the CLI

orca environment status|update|rollback|recover|stop|cancel-stop call the same managed-server
actions as Settings > Managed servers, over new managedServer.* runtime RPC methods. The desktop
main process registers those actions; the runtime advertises managedServer.v1 only then, and the
CLI refuses on a runtime without it or one that answers method_not_found. stop requires --yes.

* fix(ipc): keep a missing managed-server selector a rejection, not a synchronous throw

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move to managed server converts the host it is offered for (#25196)

The move stops the relay terminals and passes the terminal census; with #25179 the conversion no
longer refuses on the stopped terminal's saved tab, layout leaf and pane incarnation. A new test
drives one live relay terminal through Move to a conversion whose manifest carries the tab without
its relay PTY, so it spawns a fresh shell on orcad.

ssh:terminateSessions also no longer records a not-found shutdown as terminated while an older
Orca build's relay may still run the PTY (#25124): that terminal is reported unverifiable and its
lease kept, so the move refuses instead of converting over a running shell.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): deploy the glibc 2.17 runtime on old-glibc hosts, and fall back to the pinned relay (#25180)

* fix(orcad): deploy the glibc 2.17 runtime on old-glibc hosts, and fall back to the pinned relay

Managed orcad picked its runtime from the host's libc flavour alone, so a CentOS 7 host
(glibc 2.17) got the default linux-x64-glibc Node, whose self-test fails there. The deploy
now picks the runtime by glibc the same way the relay ladder picks rung A or B, and the
host-side slot checks accept the compat target it ships.

When orcad still can't run, the host is recorded as such and the relay connect that
follows now runs the pinned-runtime ladder instead of defaulting to the host-Node relay,
so a host with no Node lands on rung B rather than failing.

The hostile-host lane deploys managed orcad on a fresh CentOS 7 host and asserts it runs on
the compat runtime with no default runtime uploaded.

* test(ssh): CentOS 7 cell asserts the compat runtime pick, then the refusal and relay rung B fallback

The compat template still ships the base @parcel/watcher binary, which needs GLIBCXX_3.4.20; CentOS 7's
libstdc++ stops at 3.4.19, so the candidate's preflight refuses. The cell now pins that end-to-end
behaviour: compat runtime picked and uploaded alone, refusal classified native_preflight, and the relay
that follows settles on rung B with no runtime setting.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(serve): prove D7 on an installed Windows app whose daemon runs from the relocated daemon-host (#24976)

* test(serve): prove D7 on an installed Windows app whose daemon runs from the relocated daemon-host

* test(serve): read the pre-switch scrollback best-effort and wait for it after reattach

* test(serve): opt the packaged Windows serve switch into orcad explicitly

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): a retirement that fails after deleting rows resumes instead of reading as changed (#25234)

The chain head records sourceRetiringAt before any row is retired; the startup change check and the chain's own comparison skip a head that carries it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): managed orcad idles out after 15 minutes and is started again whenever it is down (#25121)

* feat(orcad): a managed orcad stops after 15 idle minutes and starts again on the next connect

A client-launched orcad now exits, like the relay, once no client, terminal,
working agent, staged migration or activation fence has been seen for 15
minutes. The exit is the normal graceful shutdown, which leaves the terminal
daemon running; the daemon retires only if it proves itself empty. A record in
the data root tells the next start (and its readiness health) that the stop
was an idle one rather than a crash.

On connect and after host resume, a fresh tunnel whose server does not answer
starts the activated slot under the activation fence, but only on a proven
exit, so a stopped server reads as not running rather than a failure.

* fix(orcad): keep orcad-entry under max-lines; idle e2e connects without a relay repo

* fix(orcad): deploy and rollback launches carry the managed idle-exit fence

The candidate launch in activation and the rollback launch built their
own launch spec without the activation root, so a freshly deployed orcad
never enabled idle exit; only the wake path did. The field is now required
on every launch spec, so the type system covers each launch site.

* feat(ssh): start a stopped managed orcad on connect, on restore and after resume

Once orcad stopped (idle, kill or host reboot), a connect still resolved
managed over a forward to a dead port and every call failed. Every connect
now checks the server behind its tunnel, as does a call through a restored
environment; a server proven stopped is started from its activated slot
under the activation fence, adopting a surviving daemon and its terminals.
The status line shows the start, and a start that fails keeps the host
managed with the reason and orcad.log's tail, never as a terminal verdict.

* fix(ssh): reuse a serving verdict only on the same SSH transport, for 5s

A reconnect right after a reboot was answered from the previous
transport's cached verdict, so the stopped server was never started.

* fix(ssh): key the serving verdict on the tunnel's remote port too

* test(ssh): a stopped server starts before the update counts its terminals

* feat(ssh): check serving at the bound port, and follow a restarted orcad to a new one

The serving check uses the port the tunnel forwards to (the one orcad bound).
A restart that binds a different port drops the forward and rebuilds it at
the new port, within the same ensure or on the explicit connect check. The
tunnel manager class moves to its own file to stay under max-lines.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect with no leases asks the host's relay endpoints before converting (#25223)

* fix(ssh): a connect with no leases and no relay session asks the host's relay endpoints before converting

* refactor(ssh): one isLiveSshPtyLease for every lease-liveness check

* fix(ssh): the host relay census asks a relay its PTYs before reading it as live work

An accepting relay whose holders or children the probe could not read (no lsof, an unrecognised
service child) read as live work, so a relay-era host that had exited its last shell never
converted. The relay's own bridge answers pty.listProcesses without the owner role.

* fix(ssh): the host relay census runs each relay's probe and bridge on the runtime it runs on

Pinned-ladder hosts often have no Node on PATH, so a PATH Node read every relay as unverifiable.
Each daemon's own argv names its pinned runtime, or the host Node a legacy relay started with.

* fix(ssh): an older relay's bridge runs on that relay's own runtime

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): build @parcel/watcher into the glibc 2.17 compat slot, so CentOS 7 runs managed orcad (#25199)

The compat target swapped in only node-pty, so it shipped the base target's upstream
watcher.node, which needs GLIBCXX_3.4.20; CentOS 7 stops at 3.4.19. orcad's preflight
refused it, and relay rung B lost file watching without saying so.

The compat slot now compiles @parcel/watcher from the package's own sources against the
pinned headers with the C++ runtime static, like node-pty. The slot gates (glibc 2.17
symbol floor, no shared libstdc++, N-API 8) and the smoke load cover it, the template
stages it into the compat target, and both orcad and relay rung B pick it up from there.

COMPAT_SLOT_ADDONS names a compat slot's own addons: a compat slot missing one fails
--require-slots, and a compat template target that would ship any native file without a
compat build fails the template build. The CentOS 7 cell expects an activated managed
server again, and every launched relay cell loads its watcher directly.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(phase3): close e2e routing holes, localize managed-server outcomes, trim harness and CI (#25206)

* fix(phase3): close e2e routing holes, localize managed-server outcomes, trim harness and CI

- Route the auto-convert spec into needs_build and route the new orcad/serve
  e2e helpers to the specs that use them, with routing tests.
- Run the auto-convert Docker lane only when routed; fold the missing-AppData
  check into the Windows mode-switch job and drop crash-hunt diagnostic env.
- Delta-move dialog keys its preview on the target id and cannot close or
  resubmit mid-move; Resume has an in-flight guard.
- Managed-server toasts and status lines show localized messages instead of
  raw codes or main-process English; add singular and zero-count wording.
- Settings style fixes (quiet Cancel, section header, labelled fields,
  progress labels); delete dead i18n keys and the unused previewConversion.
- Harness: kill the serve child on readiness timeout, share the isolated
  profile and spawn-until-ready helpers, hooks and a condition wait in the
  auto-convert spec, shared cross-version exec helper.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(phase3): localize the refusal toast, share serve readiness and liveness helpers, correct D7 docs

- The refusal toast no longer shows main's English blocker detail; the SSH
  Hosts status line keeps it under Details. A missing terminal count is no
  longer defaulted to 0.
- startOrcadServe uses spawnUntilReady, so a readiness timeout kills orcad;
  one isPidAlive replaces the spec's copy and the lock-holder loop.
- The port-6768 auto-convert test uses the shared hooks.
- The docs say the Windows D7 job checks rather than gates.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ci): give the idle-exit spec its app build and route the convert harness to it

Co-authored-by: m4air <m4air@Mac.localdomain>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* ci(adhoc): pin every job of an adhoc build to the commit resolved at dispatch (#25311)

Each job read the requested branch name and checked out whatever it pointed at when that
job started. A push mid-run mixed commits: in run 37234548638 the glibc217 slot lane
built 2ddea8736e, then #25199 merged, and desktop_template checked out 70948d597f, whose
merge step requires the compat watcher that lane never built.

A first resolve-ref job (no secrets) resolves the branch, tag or full SHA once, and every
job checks out that commit. The mac job still vets it for reachability before signing;
the requested name stays the concurrency key and the release label.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): report a confirmed relay PTY exit as exited and give a cold Windows PTY probe more time (#25304)

A worktree terminal close returned as soon as the relay confirmed the stop, but the
exit frame reaches the runtime record only after the SSH output intake drains, so the
verdict read straight after saw a still-connected PTY and answered unverifiable. The
stop now waits, bounded by its deadline or 10 s, for that record.

The bundled runtime's PTY probe gives the first Windows spawn 15 s and retries once
after a timeout only; spawn errors and non-zero exits still fail at once.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay,orcad): keep new wire messages forward compatible, drop unshipped relay.reset, fix shutdown and serve gaps (#25298)

* fix(orcad-migration): keep migration wire forward compatible with newer peers

* docs(pty): record why pty.resumeClient negotiates by method-not-found

* fix(relay): remove the unused relay.reset method and keep a deferred shutdown serving

* refactor: drop dead hold API, release gate and migration pass-throughs; fix serve and delegation gaps

* fix(lint): keep reopen hooks within file size limits

* revert type dedupe in orcad-incumbent-recovery to avoid a parallel conflict

* fix(lint): name stop-reply request fields for their role

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(ssh): one managed-tunnel ownership check and forward bookkeeping (#25316)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a stuck managed server can recover, cancels finish their change, unknown stays unknown; delete unused reset/drain code (#25227)

* refactor(ssh): delete the unused connection reset and drain machinery

* fix(ssh): surface a stuck managed server, finish activations past their first change, and fall back from Windows launch refusals

- A rejected build that changed profile state now refuses as orcad_recovery_changed_state
  (unverifiable) with a Recover path that restores the snapshot once the operator accepts;
  an interrupted activation is no longer reported as a quiet update deferral.
- Activation and rollback drop the abort signal after their first journaled change.
- Windows launch refusals and unsafe command lines send the host back to the relay.
- Fence refusals keep unverifiable blockers unverifiable; the POSIX liveness probe reads
  kill errors in the C locale and treats permission errors as unknown.
- A host with no relay fallback surfaces its real connect error; disconnect and removal
  clear the setting-up status, and a cancelled decision's progress is dropped.
- Remove tcp_forwarding_refused leftovers, dedupe incumbent stop/restore, exec-or-empty,
  census client, errorMessage and the blocking-blocker predicate; split activation and
  snapshot files under 300 lines.

* fix(ssh): one import per module in the rollback transition

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): committed reads trust the receipt, conversions survive UI churn, journals are cached, and round-2 low items (#25327)

* fix(ssh): keep a newer build's fence and note fields, and validate a fence beside a legacy owner

Normalization now validates known fields but passes unknown ones through on orcadFence and the managed-server notes. A legacy managed owner next to a malformed fence falls back to the owner's environment id instead of keeping the malformed fence.

* fix(migration): delete a retired migration's scrollback files once retirement is durable

Retirement dropped the moved terminals' scrollback refs from session state but kept the files forever. After the journal records source-retired, the files the manifest names are deleted, except refs any session partition or pending export still names. A crash before the delete repeats it on the next retirement pass.

* fix(orcad): a committed migration reads as committed from its receipt, survives receipt eviction, and abandoned stages expire

- Committed-state reads trusted only a byte-identical copy of every moved row and snapshot, so a live server that had been used could never confirm its own commit. The receipt alone now proves it; full equality stays inside the commit.
- A receipt that ages out of the 64-entry list keeps a compact record, so the commit never reads as absent.
- A stage no client returns to within a week stops holding the server awake and is dropped at the next stage.

* fix(migration): freeze a host's session from the fence on, ignore UI churn in the frozen-source check, and cache parsed journals

- Renderer writes to a fenced host's ssh: partition are skipped from fence time, not only after commit, so tab work mid-conversion cannot change the source.
- The frozen-source check leaves out workspace session and client routing state, which the UI rewrites as the user works (including through the local partition and UI state); the server gets them as journaled.
- Parsed journals are cached per directory, keyed by each file's inode, size and mtime and dropped on every write and remove, so list and session calls stop re-parsing every manifest.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): an older relay a census cannot rule out stays held, and a reconnect resumes held PTYs through it (#25331)

- A census that does not know the host platform is unverifiable, never "no older relay". On Windows
  it now probes each older version directory's pipe for the target, the way relay GC does, so a
  live older Windows relay keeps its PTYs held instead of respawning their panes; those endpoints
  are held, never bridged.
- On reconnect, a PTY the current relay disowned while an older relay holds it is reattached through
  that relay's own bridge, with its runtime restored and its replay forwarded. When no route serves
  it, it is left for recovery like an exhausted reattach.
- A superseded or disposed deploy no longer starts the census, so it cannot replace the current
  attempt's.
- A route stops holding an unserved PTY that exits.
- Drop the unused censusPreviousRelays export.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(phase3): round-2 ui-infra review fixes (#25325)

- An unverifiable move refusal with no count says so instead of "0 terminals".
- One move per host: the dialog stays mounted while open, joins a run already
  in flight on remount, and the toast shares the same guard.
- A managed server's start, wake or update no longer reloads every host's
  catalog; only a new environment for the host loads, scoped to that host
  and the local catalog.
- A failed host-partition session write is no longer acknowledged as written,
  so the writer re-queues those fields.
- The deploy picker lists only hosts with no managed server or pending move.
- Harness: orcad and the released relay daemon are stopped when startup fails.
- Windows host CI: the convert cell always runs last and fails on a failed
  native switch or build; ssh-host-server unit tests no longer trigger it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay,cli): never report an unreachable terminal as exited; keep PTYs through a deferred shutdown (#25328)

* fix(relay,cli): never read an unreachable relay or terminal as exited; keep PTYs on a deferred shutdown

* fix: stop an unrecorded breakaway child, route orcad serve through the spawn chokepoint, and drop review-flagged leftovers

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): reclaim a crashed desktop lock, close the stop-cancel race, drop dead daemon and RPC code (#25338)

* fix(orcad): reclaim a crashed desktop lock, close the stop-cancel race, drop dead daemon and RPC code

- A desktop reclaims a stale desktop lock record (Electron's lock proves it), and a real
  orcad hold shows a dialog instead of exiting silently.
- A failing quit handler no longer skips closing observability.
- Completion withdraws its request when a cancel lands mid-write; the listener removes a
  cancelled leftover.
- An idle stop's clean record is retracted when the shutdown fails or overruns.
- The managed-stop request tolerates unknown fields from a newer client.
- Shared error-code checks, one win32 coverage rule, no redundant isAlive filters.
- Remove the recovery-only daemon provider, requirePinnedWsPort/strictPort and the
  superseded orcad.migration.importCatalog RPC.

* fix(startup): keep main-process-preflight under the line cap

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): fail fast on a held fence, restore untouched rollbacks, atomic Windows host script, honest tunnel ensure (#25332)

* fix(ssh): fail fast on a held activation fence, restore a rollback the target never touched, stage the Windows host script atomically

- withOrcadActivationLock no longer waits up to 15 minutes inside the target lifecycle: a held
  fence answers at once (orcad_activation_recovery_required); deploy probes it before uploading.
- A rejected rollback target that left the restored snapshot untouched puts the newer build back
  unasked; only a real change waits for the operator. Crash recovery applies the same rule.
- The Windows host script is written only when missing, through a partial file and a rename; a
  bridge exit without a sentinel is unverifiable unless the shell could not find the command.
- The managed tunnel's ensure() throws when superseded, and a caller arriving after close()
  builds a fresh run instead of joining the doomed one.
- An unparseable stop answer after a stop was sent keeps the fence (new 'unconfirmed' outcome).
- stopRemote starts over instead of reporting live when another run settled the journal.
- Releasing an unreachable setup stops the orcad it activated and clears active on proven exit.
- A failed startup is torn down quietly so its error stays published; port-forward listeners
  keep an error handler; a linked SSH access connect is cancelled when its window closes.
- Shared errorMessage, one SFTP transfer helper, getConnectGeneration only, test-cell names.

* fix(ssh): stage the Windows host script through the pinned node.exe, which cmd.exe and PowerShell both run

* fix(ssh): host-script staging runs as host-script ops; a joiner of an overtaken tunnel run builds its own

- The presence check and the install are ops in the uploaded host script (script-present, and
  script-install run from the partial upload itself), invoked as node.exe <script> <op> <args>
  like every other Windows host op: no inline code.
- ensure() callers that joined an in-flight run no longer inherit its 'superseded' end; they
  start a fresh run, which also covers a close() in between.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect re-checks unverifiable relay terminals once its relay session can answer (#25385)

The server decision runs before any relay session exists, so a terminal this desktop left
detached could never be asked about and read unverifiable for as long as it ran. On Windows no
endpoint census can fill that gap. Once the session is up, the relay that holds the PTY answers
and a terminal it still runs is reported live.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(settings): a managed server whose status never loaded reads unknown, not "Not running" (#25389)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay,runtime-env,cli): keep a deferred relay's AI Vault and skill uploads, make managed re-pair crash-safe, show unverifiable in worktree ps text (#25393)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): keep a surviving daemon's slot through the cache prune, and drop an idle record a signal stop took over (#25387)

- The slot prune also protects any slot a live terminal daemon's PID record points into, and
  evicts nothing while a daemon record is unreadable.
- The shutdown trigger reports whether it took ownership; an idle stop that another source
  took over discards its clean-idle record.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): edit a managed host's connection, show rows no migration owns, drop the unused quit-drain predicate (#25403)

- SSH settings may edit a managed host's connection fields; its fence and generation are never written from the renderer, and its tunnel is closed so the next use redials. Removing it is refused with a pointer to Stop under Managed servers, which stops and removes the server; the renderer no longer ends the host's terminals before that refusal.
- A fenced host with no journal (an empty host's deploy, or one whose journal compacted after retirement) no longer hides its rows: no committed move owns them, so projects an older build added stay visible. An unreadable journal still hides them.
- The quit drain's mayDetach predicate and disconnectAll's filter had no production caller and are removed.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): input to a terminal an older relay holds but no route serves is refused, not dropped (#25404)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): refresh a changed host's status once its delta move or keep-server choice lands (#25416)

The status line kept saying the host was changed on an older Orca until a manual reconnect. Main now publishes the host as managed by its server after a successful move or keep, with a disconnected state when the move released the relay session.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a stopped managed server stops showing on its SSH host without a reconnect (#25409)

* fix(ssh): a stopped managed server stops showing on its SSH host without a reconnect

On unlink, forget the host's managed-server decision and republish its connection state. The
republish also drops the host's cached worktree scans, so worktree.ps stops naming the removed
server.

* fix(runtime-environments): a removed server's workspace session partition goes with it

Host-scoped listings enumerate session partitions as known hosts, so worktree ps kept naming a
stopped or removed server as an omitted, unselectable host. Unlinking or removing a server now
drops its runtime:<id> partition.

* test: give the removal-storage fake store casts their SAFETY rationale

* fix(runtime-environments): drop orphaned runtime workspace sessions at startup

A crash between unlinking a server and dropping its session, or a build that unlinked before the
drop existed, left a runtime:<id> session listings name as an unselectable host. Startup now drops
runtime sessions whose server is not in the environment store; it never touches local or ssh
sessions, and skips entirely when that store is missing or unreadable.

* test(migration): retirement re-aims focus without creating the destination's session

Locks in the ordering the startup reconcile relies on: a runtime:<id> session is never written
before that server is registered.

* docs(runtime-environments): name the ordering invariant the startup reconcile relies on

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect whose relays prove its leases ended retires them, so the next connect converts (#25405)

Every connect decides before a relay session exists, so a detached or expired lease reads
unverifiable there. The post-session re-check proved such leases ended but left them in place, so
the host stayed unverifiable on every connect and each respawned pane added another lease.

Co-authored-by: m4air <m4air@Mac.localdomain>

* ci(ssh-windows): check out preload and renderer for the orcad-convert e2e build

Main's #25359 narrowed this workflow's checkout to the server and test trees. Phase 3's
orcad-convert cell builds the full e2e app with electron-vite, which also needs src/preload
and src/renderer, so the x64 inbox cell failed with UNRESOLVED_ENTRY.

* test(orcad): model a really converted host in the v1.4.218 downgrade wire test

#25403 shows a fenced host's rows when no migration journal explains the fence. The downgrade
test fenced the host without a journal, so this build showed the retained project it is meant
to hide. Write the destination-committed cutover journal a real conversion leaves.

* fix(ssh): a briefly held fence waits and is retried, not recorded as a failed update; terminate keeps held leases (#25420)

- The activation fence is retried for a few seconds. A fence still held answers
  orcad_activation_fence_busy (a waiting deferral, never an update failure) unless a journal or
  a lock past its stale age shows an interrupted run, which stays recovery-required.
- acquireInstallLock reports Busy only when a holder answered; a lock command that only ever
  failed surfaces its own error.
- Terminate detaches instead of disposing when any PTY was unverifiable, so the final teardown
  never bulk-marks a lease an older relay may hold as terminated.
- A wake whose connection dropped while holding the fence releases that fence on this client's
  next wake (no journal, slot proven exited), so a relaunch-then-connect is not left fenced.
- The orcad e2e reconnect helper surfaces a connect's error text instead of a JSON parse error.

Co-authored-by: m4air <m4air@Mac.localdomain>

* ci(ssh-windows): check out all of src for the orcad-convert cell

Its e2e global setup also compiles the bundled CLI from src/cli, which the narrowed checkout
left out.

* fix(orcad): idle stop — drop the record when a signal stop wins, and read the activation lock, not its root (#25464)

* fix(orcad): drop the idle-stop record when a signal stop finishes first

Every stop's cleanup now discards the record unless the idle trigger owns the stop, so a
takeover that exits before the idle request runs no longer leaves a false idle stop.

* fix(orcad): idle check reads the activation lock, not the transaction root

An interrupted acquire can leave the root without a lock; the client already treats that as
unfenced, and managed orcad now does too instead of never idling out.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): main refuses a managed server host's removal before ending its terminals (#25465)

The remove flow skipped ending terminals for a managed host only when the renderer's cached target list already showed the fence, so a fence that landed after the list loaded still lost the host's terminals before main refused the removal. The removal's terminate call now carries forRemoval, and main refuses it for a managed host with the Stop… message before touching anything; the renderer no longer decides.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): closing an SSH workspace counts a host-confirmed terminal exit as stopped (#25479)

* fix(terminal): a workspace close counts an SSH terminal's confirmed exit as stopped

The close's verdict treated any SSH terminal whose record was still present as
unconfirmed, even when that record held a host-confirmed exit. SSH records outlive
their exit, so every successful close of a relay terminal answered unverifiable.
Also, a relay reattach that finishes registering after the stop no longer revives an
incarnation whose exit is already recorded.

* test(terminal): cover a reattached relay exit confirmed without an incarnation

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(cli,ssh): managed-server actions outwait their own deadlines; cap the unverifiable serving detail (#25473)

orca environment update/rollback/recover/stop/cancel-stop waited the 60 s RPC default while the
desktop runs the whole action inline, so a slow host printed a timeout failure for an action that
kept going. They now wait 20 minutes, and a timeout says the action may still be running and points
at `orca environment status` instead of reporting failure. status keeps the default.

The retained managed-server state now clamps serving.detail to the same byte limit as its sibling
detail fields.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): keep the previous orcad.log on Windows across a restart (#25481)

The Windows breakaway launcher's addon recreates orcad.log on every start, while POSIX
appends, so a crash's log was gone once orcad restarted. The orcad launch now asks the
launcher to move the last run's log to orcad.log.1 first, capped to its last 1 MiB.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(sidebar): a stopped managed server leaves no empty project group behind (#25488)

The removed-runtime purge dropped the server's repos, setups and worktree rows but kept the project
groups and folder workspaces the renderer fetched from it, so an empty heading stayed in the
sidebar. It now drops those runtime-stamped rows too.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): run automations on orcad and keep a managed host up while they can fire (#25475)

* fix(orcad): run automations on orcad and keep a managed host up while they can fire

orcad never built an AutomationService, so with orca serve defaulting to orcad scheduled
runs never dispatched and Run now threw runtime_unavailable. The headless service setup
moves out of Electron startup into automations/runtime-automation-service.ts; orcad
installs, binds, starts and stops it, and managed idle exit counts an enabled schedule or
an unsettled run as busy.

* docs(orcad): list automations among the idle-exit conditions

* fix(automations): precheck reads the SSH manager from its registry, keeping electron out of orcad

The precheck imported getSshConnectionManager through ipc/ssh, which pulled 32 electron
modules and node:sqlite into the orcad bundle and failed build:orcad.

* test(orcad): stub the automation wiring in the push-startup runtime harness

That harness stubs OrcaRuntimeService without an automation surface, so starting the real
service threw setAutomationService is not a function and timed out the next test.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move to managed server survives a relay that hangs up on its last terminal exit (#25487)

BUG-15: with two live terminals, one served through an older build's relay, Move failed with
"Failed to terminate SSH host sessions: …: Multiplexer disposed" although both shells died. The
old relay reports the exit before the shutdown reply; that exit closes the route (its last served
PTY), and the disposed mux rejected the shutdown still awaiting its reply.

- The legacy relay route settles a shutdown whose PTY exit it already observed.
- Move no longer aborts on a failed stop: the terminal census (what the conversion trusts) decides.
  Exited closes the relay session and converts; live or unverifiable refuses and republishes the
  relay status, so the stale terminal count is replaced.
- The move dialog offers Try again after a refusal or failure.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): closing a terminal a previous Orca relay runs stops it there and confirms the exit (#25471)

A stop on a PTY an older relay runs reached the current relay whenever no route served it at that
moment (after a reconnect with the pane unmounted, or a shell no pane ever resumed), and the current
relay answers a stop for an id it never minted as done, so the PTY read stopped while its shell
kept running. A stop on a served PTY failed instead: its exit arrived before the stop's reply,
closing the route under the pending request.

- A served PTY stops on its route; a PTY no route serves is stopped through a short-lived route to
  the older relay that lists it, which hangs up once that PTY exits.
- When an older relay may hold the PTY but cannot be asked (incomplete census, Windows pipe, a
  bridge that will not open), or its bridge drops mid-stop, the stop is unverifiable, never reported
  done; terminate keeps the lease.
- A route stays open until its in-flight requests settle.
- The provider's exit stream includes the exits older relays report, so a stop observes the PTY it
  stopped exit on the relay that ran it.

The cross-version harness runs what terminal close --all runs per PTY against a real v1.4.218 relay,
for a pane resumed this connection and for a shell no pane resumed, and sees the old shell exit there
and the old relay retire on its own grace.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a live run's journal reads busy, a wake releases only its own fence, a cancelled re-check never reports connected (#25470)

* fix(ssh): a live run's journal reads busy, a wake releases only its own fence, a cancelled re-check never reports connected

- A held fence asks for Recover only when its lock is stale or a journal has no fence over it;
  a journal under a fresh fence is a run still working and reads orcad_activation_fence_busy.
- A wake writes an owner token into the fence it takes and later releases only a fence carrying
  that token; observing the fence gone forgets it.
- The connect re-checks ownership after the relay-terminal re-check, and the re-check itself
  neither records a decision nor retires leases for a cancelled attempt.

* fix(ssh): a tunnel caller that joined a run a disconnect cancelled builds its own

The launch-time restore's tunnel run connects over SSH; a disconnect then cancels that connect.
An explicit connect that had joined the run inherited its SshConnectAttemptCancelledError and
failed (seen as the idle-exit e2e's 'connect threw: ... cancelled'). A joiner now builds once
anew after any end of the joined run except an auth failure, which it shares rather than prompt again.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(runtime-env): a late reply from a replaced pairing never overwrites the re-paired device identity (#25490)

* fix(runtime-env): a reply from a replaced pairing never overwrites the re-paired device identity

* fix(types): narrow identity fields in markEnvironmentUsed

* fix(ssh): managed tunnel proves its server by runtime id; SSH access linking keeps the strict device check

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): retained scrollback survives a closed tab, local acknowledgements stop blocking, close-intent retirement retries, mobile selections merge per workspace (#25508)

- A transfer reads a retained snapshot straight from storage once its tab closes, and releasing the retention deletes a ref no session names; the frozen-source check leaves the snapshot list to the journaled manifest.
- Only acknowledgements on panes the source host owns count toward the ui-routing blocker.
- Retiring close intents treats an absent source with an identical destination entry as already done.
- Importing a device's mobile selections keeps its selections for other workspaces.

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): an expired lease an older relay still lists is never retired (#25449)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): keep each host's state its own, count only saved commits, and never wedge connect on a partial session (#25516)

* fix(migration): a host-qualified owner key belongs only to its own host

Owner matching stripped a key's host qualifier before matching its repo id, so converting host A claimed, moved and on retirement removed host B's session state when the two hosts share a repo id, and destination-qualified focus written by retarget read as leftover source state, failing retirement with orcad_migration_source_ui_routing_reappeared. A qualified key now matches only its own host, and an unqualified key in another host's session partition belongs to that host.

* fix(migration): a partially written session partition never wedges connect, and a marker two partitions agree on stops blocking the move

Real-host BUG-14: the renderer's per-host snapshot leaves out maps a host has no rows in, and main stored host partitions exactly as sent, so a runtime partition lacked tabsByWorktree and the dormant-state collector threw on every connect. Main now fills the required maps on every host-partition write, the migration collectors tolerate a partition persisted without them, and an unreadable session blocks the move instead of failing connect.

The same profile's workspace-session blocker was a false positive: the local and host partitions both carried defaultTerminalTabsApplied for the moved worktree with the same value, and the fragment merge refused any shared worktree key. It now refuses only when the partitions disagree.

* fix(migration): only a flushed commit acknowledgement moves a migration to committed

A committed state read may come from a receipt the server holds in memory but failed to flush. A retry took that read as proof, journaled destination-committed and went on to retire the source, so a later server restart could lose the catalog on both sides. A committed read in stage and in abort is now confirmed through the idempotent commit(), which flushes before it answers; a failure leaves the journal and the fence where they were.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): relay shells no lease here knows count as live, Move stops them, and Windows asks every relay pipe (#25518)

* fix(ssh): a relay shell no lease here knows counts as live, and terminate stops it

A CLI-created terminal has no lease, so with an attached lease the gate answered from leases alone,
a live decision was never re-counted once the relay could answer, and terminate stopped only the
shells it held leases or panes for. The relay's own listing is the authority on what runs.

* fix(ssh): a Windows connect asks every relay version's pipe for its PTYs before converting

Windows pipes cannot be listed, so the connect-time census answered 'unenumerable' and a shell no
lease here knew let the host convert under it. Each version directory's pipe for this target is
derived from its path, so the census probes them all, current included, and asks a live one
through its own bridge; a live pipe it cannot ask is unverifiable.

* fix(ssh): the terminal gate counts what earlier relays still run, leased or not

After an app update a shell a respawn superseded on its tab keeps running on the previous relay
with no live lease here, so a decision counted only the leased shells and Move could not see it.

* fix(ssh): a Windows relay folder with a live pipe it cannot ask stays unverifiable

Each pipe is probed and asked on its own, so one that answered with no PTYs can no longer stand in
for a live sibling the census could not reach.

* fix(ssh): terminate also stops shells only an earlier relay lists

provider.shutdown routes a held id to the older relay that runs it, so the terminate set now takes
listPreviousRelayPtyIds too; one it cannot reach is reported unverifiable as before.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a refused Move to managed server reconnects the host on its relay (#25543)

Stopping the terminals closes the relay session, so a move the census refused (another desktop's
terminals, or an older relay it can't rule out) left the host and its workspaces disconnected
until a manual Connect. The refusal now reconnects the host; the connect-time decision reads the
same census and keeps the relay. A reconnect that converts after all reports the move; a
reconnect that fails still reports the refusal.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay): an agent exec ends on its child's exit, not on pipes a background process still holds (#25544)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): stop the automation scheduler first when a managed stop is dispatched (#25548)

A dispatch could otherwise race the daemon retirement census or write a run record that a
rollback restore then silently discards.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): count a degraded daemon's in-process terminals in the terminal census (#25545)

In degraded mode fresh terminals run on the local fallback inside orcad, but the census read
only daemon adapters, so an update or stop saw 0 live sessions and killed running agents.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a retained source whose drafts or settings an older build changed is never hidden as unchanged or retired (#25549)

* fix(orcad): a retained source whose drafts or settings an older build changed is never hidden as unchanged or retired

The retained-source fingerprint covered only repo, folder and group identity, so an unsaved draft
edited on an older build read unchanged: the host kept serving the server's older draft and
retirement deleted the newer one. Retention now also records a versioned fingerprint of the source's
drafts and user-authored names and settings; a mismatch marks the host changed, and a journal
without one is never retired automatically.

* fix(orcad): a retained source's automations are part of its state fingerprint

An older build can edit an automation the source keeps; retirement would delete that edit as if
the server held it.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a rollback restores the older snapshot only behind orcad's terminal barrier (#25551)

The rollback's census is taken while orcad still admits work, so a terminal or automation that
starts before the stop had its state wiped while its PTY survived. The incumbent is now stopped
through its managed stop with idle-daemon retirement: orcad closes terminal admission on every
daemon generation, counts live sessions under that fence, and retires the daemon only when none
exist. Only 'retired' lets the older snapshot replace state; live, unverifiable or a missing
answer refuses and relaunches the incumbent on its untouched state. A build without managed stop
is refused before anything changes.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): failed managed setup leaves 'connecting', edited managed host redials, idled-out server starts before its census (#25474)

* fix(ssh): a failed managed setup leaves 'connecting', an edited managed host redials, an idled-out server starts before its census

- doConnect publishes the error and clears the 'setting up' status when the managed-server
  decision throws for a still-current attempt; a cancelled one still reports cancellation.
- Editing a managed host's connection fields closes its tunnel and disconnects its transport,
  serialized with the target's lifecycle, so the next use dials the edited target.
- The terminal census starts a server that idled out behind a forward still up, so Stop, Update,
  Rollback and status no longer refuse with 'census unavailable' on every retry.

* fix(ssh): a fenced failed setup publishes its cause, never-launched slots are collectable, a reused PID is not orcad

- doConnect publishes the relay decision's setup failure (and clears 'setting up') when a failed
  managed setup kept the host fenced, instead of throwing a bare 'serves a managed server'.
- The liveness probe answers NEVER_LAUNCHED for a slot with no process record and no readiness
  file; GC removes such a slot, and every other reader still reads it as UNKNOWN.
- On POSIX a PID whose command line does not run the slot's orcad.js reads DEAD, so a stopped
  orcad behind a reused PID is woken instead of reported serving or unverifiable.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a legacy relay route a new pane started serving stays open when a pending stop settles (#25542)

A served PTY's exit that lands before its stop's reply defers the route's hang-up until the stop
settles. A second pane the same older relay holds could start serving through the route in that
window, and the deferred hang-up then closed it anyway, sending that pane's input and stops to the
current relay. Serving a pane now cancels the deferred hang-up, and a settling request hangs up only
a route that serves nothing.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): the journal holds a migration's scrollback until commit or abort, across failed uploads and restarts (#25550)

Retention was scoped to one transfer call, and its release in finally deleted a closed tab's
snapshot after an interrupted upload, so every retry failed with source_snapshot_changed.
Inline buffers had no file for the retained read at all. Retention now follows the cutover
journal: held from the journaled export through staging, rebuilt at startup, released on
commit or a removed journal. Inline bytes are written to their ref while held.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(migration): a repo id two hosts share never lets legacy keys cross hosts, and a dangling identity alias stops blocking the move (#25558)

- Repo ids are not unique across hosts: the same id may be registered on two SSH hosts. Stores that are not session partitions (worktree metadata, automations, lineage, client state, sparse presets, retired names) can hold legacy keys with no host qualifier, so an id both hosts register said nothing about whose a key was. The scope now records such shared ids; an unqualified key or bare repo id for one only matches with its row's own host evidence (worktree metadata's hostId, an automation's ssh target), so another host's rows are never moved, counted or retired. An automation's target generation now matches only alongside its target id.
- An identity alias whose identities hold no metadata (worktreeMetaByIdentity lost them, as on the B4 profile) is nothing to move rather than a worktree-metadata blocker; retirement drops it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(cli): orca environment recover --accept-changed-state --yes restores over changed state (#25597)

Recover refused when a rejected build changed profile state, and its refusal told the user to run
Recover, which the CLI could not do. --accept-changed-state (confirmed with --yes) maps to the same
acceptChangedState the Managed servers settings pass, and the refusal now names the flags.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): refuse an update that would end a degraded host's in-process terminals (#25598)

The census now reports inProcessSessions separately. Those terminals run inside orcad and
end with any restart, so planOrcadUpdate defers with a non-forceable
orcad_update_ends_in_process_terminals instead of claiming they survive on the daemon.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): a delta move protects what it imported from rolling back across it (#25599)

A delta move never advanced the server's migration mark, so rolling back an update taken before the delta was admitted and dropped the delta's projects. The delta now records the mark before any commit can land, resumed commits included; the mark keeps the latest migration and never moves back; and the rollback gate also counts every journal into the server, so deltas finished before this change stay protected.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a missing relay inventory never proves terminals exited, and the Windows census covers every desktop's relays (#25611)

- The migration terminal gate returned `exited` whenever this desktop held no unresolved lease, even
  when the current relay or the earlier relays could not be asked. A failed or incomplete inventory
  is now `unverifiable` regardless of local leases. With no relay session at all the gate asks for a
  host census (`needsHostCensus`) instead of reading the silence as exit; the connect, conversion and
  delta move pass that census in, and the connect hands its own census result to the conversion it
  starts. A census that cannot list endpoints (`unenumerable`) is unverifiable too.
- The Windows connect-time census derived pipe names from this desktop's target id only, so another
  desktop's relay on the same account was never probed. It now lists every `orca-relay-*` pipe on
  the machine and maps each to the relay instance that owns it through that instance's credential
  file or active-pipe marker, asking each with its own credential. A pipe no version directory
  accounts for is unverifiable unless the host proves it another account's (or gone), and an
  inventory that could not be read is unverifiable.

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor: drop the unshipped pty.resumeClient relay method and unused SSH provider unregister guards (#25595)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect still deciding its server holds the raw 'connected', closes a transport its cancelled decision opened, and an edit keeps a relay host's session (#25641)

- handleSshConnectionStateChange holds a raw 'connected' while a connect is in flight even before
  any relay session exists (published as 'connecting'), so the census, deploy or conversion that
  dials the pool no longer reports the host up with no providers.
- priorConnection is captured before the server decision; a connect cancelled after the decision
  closes a transport the decision opened, unless a newer connect is using it.
- Editing a fenced host an older build changed (it runs on the relay directly) no longer
  disconnects its transport; only a host reached through its managed server redials.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a retained source is unchanged only against a pre-commit baseline of everything a user wrote (#25602)

The state fingerprint read drafts, automations and workspace metadata from what a move could
carry, so a session a move refuses hid an older build's draft edit, and retention hashed the source
after the commit, so a crash before retention blessed whatever an older build changed in between.
The baseline is now written with the fence, before any commit is possible, from the source read
directly; a session that cannot be read, or a journal without that baseline, is unverified.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a delta move refuses a source that changed while it checked terminals (#25691)

The plan and manifest are taken before the terminal check and session release are awaited, but the
journal took its baseline after them, so a draft typed in between became the baseline while the
server received the older one, and retirement deleted the newer draft. The baseline now comes from
the plan's own snapshot, and a source that changed since it refuses the move.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): a save landing mid-retirement no longer defers retirement (#25692)

Retirement removed the source rows, then awaited the profile flush, then checked nothing came back. A session save that landed during that flush re-added a source-owned row, the check failed, and the journal stayed committed until a later connect. Retirement is idempotent, so it now runs one more pass before deferring; a row back after that is reported with the partition and owner key it reappeared under.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): a headless run never reads completed when its agent never ran (#25700)

* fix(automations): a headless run never reads completed when its agent never ran

orcad (and Electron serve) finished a dispatched run on a satisfied
tui-idle wait, and a ready shell prompt satisfies it: a run whose agent is
not installed read 'completed' within seconds. Like the desktop runner,
completion now needs the agent's own status for the run's pane after
dispatch; without it the run fails after the agent-start window with the
reason, instead of claiming completion.

* fix(automations): keep idle-means-done for agents without status; fail only a refused command

Not every automation agent reports status on orcad (no hooks on the host, no
recognised title), so requiring it would fail their runs. A run completes on
the agent's own status, fails when the shell refused the agent's command
(bash, zsh, dash, fish, PowerShell, cmd), and otherwise keeps the old
idle-means-done rule after the agent-start window.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): the retained-source fingerprint sees all worktree metadata retirement owns (#25694)

* fix(migration): the retained-source fingerprint sees all worktree metadata retirement owns

The state view read only legacy worktreeMeta keys and attributed them without meta.hostId, so
identity-backed metadata and unqualified rows a shared repository id leaves to hostId were missing
from the fingerprint: an older build's edit read as unchanged and retirement deleted it. The view
now uses the same attribution as export and retirement, and metadata that claims the source host
but cannot be attributed leaves the source unverified. The fingerprint version moves to v2.

* fix(migration): the retained-source fingerprint skips automations on a repo id another host shares

The state view matched automations by scope.repoIds, which ignores sharedRepoIds, so host A's
fingerprint included host B's automation on a shared repository id; editing it marked A changed and
routed it back to the relay. The view now uses the move's automationTouchesScope.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a failed readiness read no longer kills a healthy candidate; log unsettled activations (BUG-17a) (#25701)

* fix(orcad): a failed readiness read no longer fails a candidate's launch, and log every unsettled activation

BUG-17: the candidate went ready and the client SIGTERMed it ~1 s later through its reject path,
yet that readiness passes the gate, so the launch itself failed: one readiness-wait exec that
errored failed the launch outright. Retry such reads until the readiness deadline; an
unconfirmed termination still fails at once. The update's outcome never reached the app log,
so every update or rollback that does not go through now logs its code and reason.

* fix(orcad): the host-side readiness wait ends on a wall-clock deadline

On a loaded host each poll's reads outlasted its sleep, so the step-counted loop ran past
the client's 30 s exec timeout. That timeout failed the launch, the reject path SIGTERMed a
candidate still starting, and orcad, which defers a stop until startup completes, published
readiness and then exited (BUG-17).

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): stop a converted host's stale tabs landing in local, and retry retirement only on an exact replay (#25712)

* fix(migration): retirement re-removes only an exact replay of moved session state

#25692's second pass re-ran retirement on any row that reappeared during the flush, which could delete a tab or draft written after the move. Retirement now records the session rows it removes before it runs; a row that reappears is removed again only when it is byte-identical to one of those (in any partition). A new or changed row defers retirement and stays.

* fix(ssh): a converted host's leftover session rows stay in its own partition, never local

After conversion the renderer drops the SSH host's projects and worktrees but keeps their session rows. With no catalog owner left, the next save routed those rows to the local partition, where retirement read them as moved source state reappearing and deferred. Converting now pins each dropped worktree's session key to the host's partition; main's fence guard keeps that partition frozen, so the stale rows are never written.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(test): restore the codexProviderHandle import main's #25078 dropped again

#25722 restored it, then #25078 removed it, so pnpm tc fails on main's tip.

* chore(sync): keep main's own cloud and mobile files byte-identical to main

Earlier syncs added lint-only brace and template fixes to these main-owned files; reverting
them keeps #24863's diff against main free of files Phase 3 does not own.

* fix(terminal): a remote pane's reattach error no longer shows this client's OS and shell (#25693)

* fix(terminal): a remote pane's reattach error no longer shows this client's OS and shell

The error toast appended the client's environment for every non-SSH error. It now shows it only
when the pane's known execution host is this client; an SSH, managed or not-yet-known host omits it.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(terminal): type the pane-host fixture as runtime owner state

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(orcad): a retained activation fence is ownerless, so recovery can always take it over (BUG-17) (#25698)

* fix(orcad): a retained activation fence is ownerless, so recovery can always take it over

BUG-17: a recovery that took a stale fence over, failed and retained it left a fresh lock, so
every later recovery read it as still fresh and the host could never be recovered. A run that
keeps the fence once it is done now backdates the lock; one whose remote command may still be
running keeps it fresh.

Also run the in-process terminal deferral before the forced protocol check, so a degraded host
whose daemon is empty names its in-process terminals instead of an unreported protocol.

* test(orcad): the CLI's accepting recover takes over a fence a refused recover just retained

The fake host now answers a stale-only takeover busy while a recovery's own takeover is
fresh, which reproduces BUG-17's 'still fresh' loop without the ownerless mark.

* test(orcad): keep the fake host's fence acquisition void where callers expect it

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a cancelled connect closes only the transport its own decision opened and nothing newer adopted (#25696)

* fix(ssh): a cancelled connect closes only the transport its own decision opened and nothing newer adopted

#25641's cleanup took any transport that differed from the pre-decision one as the cancelled
decision's, guarded only by connectInFlight. A replacement connect that completed (and left
connectInFlight) then had its live transport disconnected by the stale attempt.

The pool now attributes a transport it opens inside a connect's server decision to that attempt
(AsyncLocalStorage), and the latest user adopts it: a connect that connects or publishes a managed
route, or a managed tunnel that records a forward. A cancelled attempt closes the transport only
when it opened it, nothing newer adopted it, and no current replacement is in flight.

* fix(ssh): a still-current connect whose server decision fails closes the transport that decision dialed

The decision's own failure (a throw, or a fenced relay refusal) published 'error' but left the
transport it dialed open, so getPublicSshState read 'connected' and a later auto-reconnect
broadcast a plain 'connected' with no relay. Both branches now close exactly the decision-owned
transport through abandonDecisionTransport, which treats the attempt's own in-flight entry as
no newer owner while that attempt is still current.

* test(ssh): name the stand-in transport type

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): Windows readiness identifies itself by one PID when the process-table snapshot times out (#25723)

* fix(orcad): Windows readiness identifies itself by one PID when the process-table snapshot times out

A loaded Windows runner timed the whole-table snapshot out during the bundled runtime's
readiness preflight, so the candidate failed to start and activation rejected it.

* fix(orcad): fall back to the one-PID query only for a slow process table, never an unreadable one

An unreadable table (EDR-hooked snapshot, restricted token) must still fail qualification.
The table now rejects slowness with a typed WindowsProcessTableTimeoutError, and only that
falls back. Review by win-serve.

* build(cli): list the process-table timeout error in the CLI project's file list

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connected conversion proves the host idle with the account-wide census, not this target's lists (#25697)

* fix(ssh): a connected conversion proves the host idle with the account-wide census, not this target's lists

The migration terminal gate asked the host-wide census only when no relay session existed. With a
session, it trusted this target's relay listing and its earlier-relay census, both of which name
only this target's instances, and returned `exited` when they were empty, so another desktop's live
shell on the same account, under a different target id, let the host convert under it.

`exited` now always needs a complete host-wide census: this target's lists can prove `live`, but
empty lists only pass the question to the census, and a gate given none answers `unverifiable`
with `needsHostCensus`. The connect-time refinement passes the census too, so a connect retires its
leases only when no relay on the account holds work. The Windows host lane now runs a second
desktop's relay with a live shell and expects the connected gate to read the host live.

* test(ssh): a connected relay's empty lists still ask the account-wide census

* refactor(ssh): drop the gate's unread needsHostCensus flag; its unverifiable reason says why

* test(ssh): the delta snapshot fixtures give their account-wide census

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a reconnected transport starts managed orcad itself instead of inheriting a dropped start (#25689)

* fix(ssh): a reconnected transport runs its own orcad start instead of inheriting a dropped one

The serving check deduplicated in-flight starts by environment only. When a
connect dropped mid-start (as when the launch-time auto-connect is replaced
by a reconnect), the caller on the new transport joined the start bound to
the dead one, got its failure, and reported the host managed with no server
running. In-flight checks now join only on the same connection, connect
generation and port, and a wake's own fence token is cleared only by that
wake.

* fix(ssh): a reconnected wake releases the fence its dropped wake held at any point

The flake's real verdict was 'fenced': the dropped launch-time wake held the
activation fence, and the reconnected wake could not prove it its own. A
wake now claims its token before its first remote step, an absent owner
record under a held token is still its own, and a reconnected wake waits for
this client's dropped wake to settle before reading the fence.

* fix(ssh): release only a fence carrying this process's own wake token

A fence with no owner file could be another client's fresh one. Releasing it
now requires the owner token this process wrote; the token is claimed before
the write so a drop after it still proves ownership.

* fix(ssh): type the wake's fenced fallback

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): stop the exec-stdin test double from failing on EPIPE (#25739)

The truncation test's fake exec channel forwarded the local shell's
EPIPE (or 'Cannot call end after a stream was destroyed') as a channel
error. Whether that error or the shell's exit code won depended on
scheduling, so the test failed under full-suite load. ssh2 silently
drops writes once the remote stops reading; the double now does the
same, and a 1 MB payload makes the early-stop path deterministic.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move lets the reconnect's decision run its census, and a census outside a connect holds 'connected' and closes its own transport (#25735)

- Move to managed server no longer runs a separate host census after tearing the relay down,
  which dialed the pool with no connect in flight and broadcast a raw 'connected' with no
  session or providers. It reconnects, and reads the decision the reconnect's census recorded.
  A relay a failed stop left up is detached (leases kept), not disposed, before the reconnect.
- The CLI and delta-move census (censusHostRelayTerminalsFor) runs outside a connect under its
  own owner: the raw 'connected' it causes is held, and a transport it opened that nothing
  adopted is closed afterwards. Reusing a pooled transport inside a scope now adopts it.
- Drop the now-unused publishRelayTerminalsStatus.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(test): drop the restored codexProviderHandle import now that main restored it

* refactor(migration): keep a converted host's source rows instead of retiring them automatically (#25768)

Automatic source retirement leaves Phase 3: nothing deletes a converted host's retained rows on connect, delta move, keep-server's-version or restart. They stay hidden and are removed only by stopping the server, removing the host or uninstalling. Change detection goes back to the catalog-identity fingerprint, so an older build's edits inside an already-moved project stay preserved in the retained rows without marking the host changed. The copy-only helpers the delta view uses are renamed to subtract, and the converted-host session pin now also overrides a boot primary of local.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): a headless run is watched past tui-idle timeouts by the one run observer (#25733)

* fix(automations): a headless run is watched past tui-idle timeouts by the one run observer

The headless dispatcher awaited a single tui-idle wait, which rejects after
its 5-minute default, so a healthy agent working longer was published as
dispatch_failed and never observed again. The dispatcher now hands the run
to its completion watcher, whose runtime observer already re-arms wait
timeouts, honours cancellation and bounds total observation; the agent
status and missing-command checks fold into that observer, and the separate
completion loop is gone.

* fix(automations): an already-idle pane completes when the start window passes

Real-host: a stub that exited before the window left an idle shell with no
agent status, and the observer re-armed a tui-idle wait that never resolves
for an already-idle shell, so it timed out instead of completing. The
observer now keeps judging the pane while its output is unchanged, and only
waits again once the pane changes.

* fix(automations): resolve a watched headless run by its launch handle first

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(runtime-env): a re-paired managed server's subscribers recover without a reload (#25752)

* fix(runtime-env): a re-paired managed server's subscribers recover without a reload

The renderer kept the pairing revision it last read, so after an on-connect update re-paired a
managed server every subscribe and request was refused as 'pairing changed' until a reload. The
first refusal now re-reads the environment catalog, so revision-keyed subscriptions resubscribe
and requests carry the new pairing. A managed server re-pairing for the same host registration is
the same peer, so its workspaces and tabs are no longer purged as a replaced environment.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): a re-paired managed server is the same machine only when its host proves the same identity

Same SSH target registration is not proof: a reinstalled host or a target now pointing elsewhere
keeps it. The runtime id the pairing handshake verifies must be known and unchanged; otherwise the
re-pair retires the environment as before. A proven runtime id change also counts as replaced.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): decide a managed re-pair's same machine by the host's proven key, not its runtime id

The runtime id is minted per process start, so every orcad restart would read as a new host.
The host's E2EE public key persists in its own profile across updates and its pairing handshake
proves it; main now lists a digest of it, and the renderer keeps a re-paired managed server only
when that digest is known and unchanged under the same SSH target registration.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): defer a managed re-pair's same-machine decision until the host key is known

A re-read that lands before the new pairing's host key is listed no longer purges: the decision
waits for a catalog that carries the key and retires only if it differs. Adds the update-flow
store test: same registration and key keeps workspaces and tabs, including a re-read that runs
before the key is known.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): watch for a pairing refusal on a side branch so requests settle on the same tick

Chaining .catch onto every subscribe and request delayed each success by a microtask, which let a
StrictMode cleanup run before a client-event subscription resolved, so its unsubscribe landed late.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* refactor: consolidate pane-ownership, migration-catalog, activation-launch and parser helpers (#25738)

* refactor(orcad): one launch-and-judge helper for activation and rollback

* refactor(orcad-migration): one copy each of the destination projections, selectNewRows, assertSameValue, compareKeys and slotLiveness

* refactor(orcad-migration): one string-list validator and one uniqueness check, error codes passed in

* refactor: one shared pane-ownership and terminal-layout module for migration, profile transfer and split layout

* fix(orcad-migration): row-identity helpers in a leaf module (no import cycle); key order in its own module

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): reopen a managed tunnel to a host another desktop restarted (#25800)

The tunnel's identity check pinned the saved runtime id, which orcad mints per process. A host
updated or woken by another desktop, or restarted while this one was away, failed every reconnect
with orcad_identity_mismatch, and nothing could refresh the id because that needs the tunnel. The
E2EE handshake with the pinned host key and our accepted token now prove the server; the first
authenticated status reply records the new id. A different host is still refused.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(orcad): create the readiness file owner-only so its pairing token is not world-readable (#25809)

orcadLaunchCommand truncated .orcad-readiness before setting umask 077, so under a
login umask of 022 the file that receives the pairing offer (with a runtime-scope
device token) came out 0644. umask 077 now runs first, the readiness file is
chmod 600 after the truncate (a redirect keeps an earlier build's 0644), the pid
and log files are tightened too, and the slot dir and ~/.orca-remote are chmod
700 so files earlier builds left readable are no longer reachable. The state
snapshot capture also sets its umask before creating the snapshot directory.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): explain a pane whose saved session another host connection owns (#25814)

terminal_pane_owner_host_mismatch reached the user raw, with an issue link. It now reads as a
plain explanation with the open-a-new-terminal action, like the reattach failure.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* test(topology): allow phase3's headless editor-tab retirement in main's boundary ratchet (#25823)

Main's #25329 added the ratchet; phase3's mobile-session-editor-projection.ts writes the host's
own session through setWorkspaceSessionForWorktree, the same way the listed headless
mobile-session tab writers do.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): a headless host closes finished run terminals, keeping the newest few (#25831)

* fix(automations): a headless host closes finished run terminals, keeping the newest few

The desktop closes a run's terminal when the run completes; orcad had no
renderer to do it, so hourly automations left a shell and PTY per run open
forever (28 after ~6h on a real host). The headless service now closes a
finished run's terminal after a 10-minute grace, keeps the newest three per
automation viewable, and never touches a run that has not finished.

* fix(automations): never close a run terminal a client typed into or is viewing

Mirrors the desktop's take-over rule on headless hosts: a finished run's
terminal stays open when any client drove input to it since spawn, is
attached to or viewing it, or when this process cannot tell (it adopted
the PTY rather than spawned it).

* test(runtime): register a viewer through the public subscribe API

* fix(automations): close only completed runs' own panes

A failed run can still hold a live agent (blocked on a prompt, past the
watch window, or after an observer error), so like the desktop only a
completed run's terminal is closed. And only the run's own pane closes, so
a pane a user split into the same tab survives.

* fix(automations): close a run pane only while it still holds the run's PTY

The use check read run.terminalPtyId, but the close hit whatever PTY now
occupies the run's pane. Restart-exited-pane and the Codex account-switch
restart put a new PTY there, so a terminal a user was using could be killed.
The close now resolves the pane's current PTY and closes only when it is the
run's own; otherwise it closes nothing and only clears the run's terminal.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a snapshot restore past the exec timeout keeps its fence instead of racing a second restore (#25811)

* fix(orcad): a snapshot restore past the exec timeout keeps its fence instead of racing a second restore

A capture, restore or clear ran under the generic 30s exec timeout. On ssh2 the timeout closes the
channel and reads as a confirmed failure, but sshd leaves a pty-less command running, so rollback
ran its rescue restore and recover orphaned the fence for a second one, both in the same stage.

- State mutations run through execOrcadStateMutation: no abort, a client wait past the host's
  deadline, and any closed channel or busy/deadline answer is unconfirmed, so the fence stays fresh.
- POSIX hosts wrap each one in `timeout -s KILL` (where present) and a pid-checked lock dir under
  ~/.orca-remote; the Windows host script takes the same lock.

* fix(orcad): a running state mutation keeps the activation fence fresh

The fence goes stale by its lock dir's mtime after 20 minutes, and a capture, restore or clear can
now run up to 15 under it, so a rollback's rescue capture plus restore could outlast the window and
let a recovery steal the fence from a live run. While a mutation runs, the host now touches the
fence every 60s (POSIX: a background beat that stops with its shell; Windows: an interval in the
host script, whose mutations are now async so the timer runs). A dead process stops refreshing, so
stale takeover still recovers it.

* fix(orcad): the state-mutation fence heartbeat never refreshes a wake's fence

A wake writes .orca-wake-owner into the fence dir and lets its fence age toward takeover; the
heartbeat now skips a fence that holds that token, and only ever changes the dir's mtime.

* fix(orcad): a state mutation releases its host lock before answering, and names its holder by pid and start time

On Windows answer() exits in the stdout write callback, so an op that answered before its first
await (MISSING, EMPTY, FAILED) exited before the wrapper's finally and leaked the lock; a reused
pid then read as alive and every later capture, restore and clear answered busy. Ops now return
their token and the wrapper answers after releasing the lock. The holder is pid plus creation time
(the slot's process-tree addon); one that cannot be identified is stale once its lock misses five
heartbeats. POSIX gets the same heartbeat-age check for a reused pid.

* fix(orcad): a state mutation's host lock is owned by its whole process group

The lock named only the shell's pid, so a shell killed while its rm or tar ran let the next
mutation take the lock and race that child. Each mutation now runs in its own process group
(setsid, or perl setpgrp on macOS), with timeout inside it so a deadline KILL reaches the children
too. The lock records the group, and is taken over only once no member is alive; a host that can
start no group records none, and its lock is never taken over. Windows ops run in-process, with no
children to outlive the holder.

* fix(orcad): record a state mutation's process group without ps -p, and never hold a groupless lock forever

BusyBox ps has no -p, so Alpine hosts recorded no group and their lock read busy forever after a
timeout kill, reboot or OOM. The group now comes from /proc/<pid>/stat (read after the comm field's
last paren), with ps -o pgid= -p as the fallback. A lock that still names no group is taken over
once its pid is dead and its heartbeat has missed three beats.

* fix(orcad): only proof of exit frees a state-mutation lock

A Windows holder whose creation time could not be read was taken over after five quiet minutes
though its pid was alive, so a suspended clear could resume and delete freshly restored profiles.
Both platforms now free the lock only on proof of exit: a dead pid, a different creation time, or
(POSIX) a group with no live member. A live holder of unknown identity stays busy until it exits.
The owner record is written exclusively, so a run that resumes after a takeover backs off.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): never offer Move for terminals another Orca desktop or session runs (#25815)

On a host where another desktop held a live relay shell, this desktop read relay_terminals_live
with offerMove, and its copy ("Its N open terminals will restart") implied they were its own. The
census already attributes them: terminals counted only by the host-wide census, with no lease or
listing of this target naming one, run under another target or session. That verdict now carries
elsewhere / terminalsElsewhere; no move is offered (no toast, no status-line action) and the status
line says the terminals belong to another Orca desktop or session.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): an update releases finished, unused automation shells before counting terminals (#25844)

Hosts with schedules kept completed run shells (the newest three, and any not
yet past their grace), which counted as running terminals and deferred every
on-connect update with orcad_update_terminals_running. The update and
rollback census now ask the server to close completed automation run
terminals no client used, with no grace or keep rule, and count after the
daemon drops them. Used, unknown, failed and running ones still count.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): every activation fence holder carries a generation token its steps and release must match (#25834)

* fix(orcad): clear a bare stale activation fence instead of asking for Recover, and report a restarting update from status

BUG-21: a wake cut short leaves a stale fence with no journal. Every update then answered
'Recover it first' while Recover answered 'none'. The fence-hold check now takes such a fence
over and drops it, and the update retries once; Recover is asked for only over a journal.
The CLI also treats a connection closed by the server's own restart during update or rollback
as expected and reports what status shows once the runtime answers.

* fix(cli): type the reconnect status response explicitly

* fix(ssh): a wake's fence carries its owner token from the moment the lock exists

The idle-exit e2e still read 'fenced' on reconnect: the launch-time wake's
lock landed on the host but its connection dropped before the client saw OK,
so the wake body never ran and never wrote its owner token, leaving a fence
nothing could prove. The token is now claimed before the lock and written by
the same command that creates it, and a wake registers itself before any
remote step so a reconnected wake waits for it instead of racing it.

* fix(orcad): every activation fence holder carries a generation token its steps and release must still match

Astra pass 8: a holder suspended past the stale window resumed, kept acting, and its
unconditional release deleted the successor's fence and recovery journal mid-update. Every
holder (activation, rollback, stop, recover, wake) now writes a token into the lock it creates or
takes over. Each remote step it issues checks that token on the host, in the same command on
POSIX and inside the host script for Windows host ops; release is conditional on the token and
moves the lock aside instead of removing the root. A superseded holder aborts with
OrcadFenceLostError and its release is a no-op.

* fix(orcad): state mutations check the fence token before their lock, and refresh only a fence they still own

On POSIX the fence guard runs outermost in serializedStateMutationCommand, before the mutation
lock and the work, and the heartbeat touches the fence only while the token is still this run's.
The Windows host script records the --fence token and refreshFence compares it. A fence-lost
answer from a state mutation is a refusal, never a FAILED fallback. One owner-file constant
replaces the wake-owner copies.

* refactor(orcad): a state mutation's heartbeat touches the fence directory it checked ownership of (review)

* fix(orcad): a release moves the journal and lock aside and keeps only its own generation's

Astra pass 9: a release that passed its token check and stalled before deleting could, once a
takeover and a successor came and went, delete the successor's journal and lock. The journal is
now stamped with the writing run's fence token (a recovery takeover re-stamps the journal it
adopts), and the release renames the journal and the lock aside, deletes each only if it carries
this run's token, and otherwise moves it straight back.

* test(orcad): a successor restore stays busy beside a paused clear on the Windows host script

* fix(orcad): classify a lost fence from the step's exit and stdout, never the error message

The real exec error quotes the command, and every fenced command carries the
guard's marker text, so any failed or timed-out fenced step read as a lost
fence and dropped its unconfirmed flag. execCommand now attaches exitCode and
stdout to its exit error; execOrcadRemote rethrows unconfirmed terminations
before any reclassification.

* fix(orcad): a wake keeps its fence token until the fence is released

A disconnect fails a wake's next step without the unconfirmed flag, and the
fence release then fails over the dead connection. Forgetting the token on
that error left a fence the reconnected wake could not prove its own, so it
reported the host as held by an update.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(recovery): keep Phase 3's recovery-lifetime test on main's legacy-worker ports

Main added a required hasRequestedReleases port and now skips persist when a pass resolves
nothing, so the test mocks the new port and holds the pass at workspace resolution instead.

* fix(ssh): say "1 terminal" when another Orca desktop runs one on the host (#25853)

The terminalsElsewhere status line had no plural forms, so B9 read "while 1 terminals another
Orca desktop … are running". It gains _one/_other entries like the other terminal-count strings
on that line and in the move offer, which were already pluralized.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): close failed and exited runs' terminals once the shell is proven alone (#25859)

* fix(automations): close failed and exited runs' terminals once the shell is proven alone

Failed (command not found, timeout) and forever-dispatched runs kept one
shell per run, unbounded and counted by the update gate. Their terminals now
close like completed ones (unused, past the grace, outside the newest few)
but only on fresh execution-host proof that the spawned shell is alone at
its prompt; a live or unprovable agent keeps its terminal. A still-dispatched
run closed this way is marked failed. Dead terminals no longer take one of
the newest-three keep slots. The update drain follows the same rules.

* fix(automations): prove a run shell alone from the process table, not the daemon's ownership flag

On a real daemon session the daemon's confirmShellForeground stays false
after a plain 'command not found' and after an agent that exited, because its
ownership flag only turns 'shell' after a full-screen command; failed runs
would never have closed. The proof now also reads the host's process table:
on POSIX the PTY's root shell must own the terminal foreground group with
nothing stopped under it, on Windows the host's job-based child census must
be empty. Anything unobservable still keeps the terminal.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): extensive orca CLI matrix on Windows hosts (#25114)

* test(ssh): extensive orca CLI matrix on Windows hosts

Adds three dispatch-only app cells to the ssh-windows-hosts lane that drive the
e2e build and the bundled orca CLI against the provisioned Win32-OpenSSH host:
empty-host deploy/terminal/reconnect/orcad-restart/decommission, seeded
relay-era conversion, and an open relay terminal keeping the host on the relay.

* test(ssh): pin the relay-kept cell's runtime; keep cleanup from masking failures

* test(ssh): run decommission before the orcad restart in the managed cell

* test(ssh): decommission through orca environment stop; accept an unverifiable relay close

* test(ssh): require a confirmed relay close; app cells must run last

* test(ssh): log and accept either relay-kept census reason; keep app-cell test results

* test(ssh): relay-kept requires a live census and its status line again

* test(ssh): match the pluralized relay-kept status line

* test(ssh): orcad restart proves a new process, terminal adoption, and a kill-then-connect relaunch

* test(ssh): restart kills only the orcad server, not its terminal daemon; wait for a released profile

* test(ssh): collect orcad.log.1 so a restarted orcad's previous run is kept

* test(ssh): the managed cell proves a workspace listener is detected and attributed

* test(ssh): start the port listener without $, so a PowerShell terminal doesn't expand it

* test(ssh): the port check proves Windows command-line attribution; retry a dropped version read

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(runtime): move run-terminal client-use and shell-alone checks out of the branch-cleanup runtime (#25876)

* refactor(runtime): move run-terminal client-use and shell-alone checks out of the branch-cleanup runtime

orca-runtime-preserved-branch-cleanup.ts had grown past max-lines (303) with
the headless run-terminal helpers. Their logic now lives in
run-terminal-client-use.ts and the runtime keeps one-line delegators, with
behavior unchanged.

* fix(ci): the runtime Electron ratchet bundles its entry points once, not 2.5k times

check-runtime-electron-ratchet bundled ~2,532 entry points each in full (format cjs, no
splitting), so esbuild held thousands of copies of the runtime graph: about 2.2GB RSS and 11s per
run, twice per test file. It was in flight in every unit shard that died with "The runner has
received a shutdown signal" (#25815 5/5 twice, #25876 2/5 twice). With esm + splitting the shared
modules land in one chunk: same metafile, about 200MB and 2s.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): wait for the busy relay's child before probing it (#25916)

The fake relay's spawn is not visible to pgrep at READY on Linux under Bun, so
the probe could count zero children. The sibling cases already wait.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a quit that aborts an upload whose read already ended no longer crashes main (BUG-23) (#25922)

Quitting while an on-connect orcad update was uploading its bundle aborted the connection's
teardown signal. sftp-upload's abort handler destroyed the local read stream with the signal's
reason, but once that read had ended, 'finished' had already removed its listeners, so the
stream emitted an unhandled 'error': [main_uncaught_exception] AbortError: This operation was
aborted. Electron's error dialog then blocked the main thread and the app never exited.

The read stream now always has a no-op error listener; the transfer's outcome still comes from
'finished'. Both the bare upload and the connection-level teardown abort are covered.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a launch reads readiness at least once; fake hosts match the capture's tar flag, not any -cf (#25918)

A random fence token contains `-cf` about 1 time in 125, and the fake hosts
read any command containing it as a snapshot capture, so a rollback's restore
answered CAPTURED and the rollback never launched. Separately, a client
descheduled between computing the readiness deadline and checking it skipped
every read and failed a ready launch.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a fence this desktop's exited process left is cleared without the 20-minute wait (#25941)

* fix(orcad): a fence this desktop's exited process left is cleared without the 20-minute wait

The client records the fence tokens its processes hold beside the profile. On
a later launch, a POSIX fence carrying a token from a process that has exited,
quiet for three heartbeats and with no live state mutation, is backdated so
the existing stale rules clear it or hand it to Recover at once. Another
desktop's fence, a live holder's, or one with a mutation still running keeps
the normal stale window.

* fix(orcad): held fence tokens are best effort, pinned to this machine and boot, and pruned after a day

A token-file write that fails no longer breaks a fence operation; an entry
recorded on another machine sharing the profile, or before a reboot, never
proves its holder exited; entries older than 24 hours are dropped. Tests cover
a journal kept for Recover, a successor freshened back after a racing backdate,
a Windows host, and the record across release, supersession, busy and a lost
connection.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): an install lock this desktop's exited process left mid-upload is taken over without the 20-minute wait (#25991)

A quit during the bundle upload leaves the version dir's install lock, not the
activation fence. The lock now carries this desktop's token, recorded in the
held-token store, and is forgotten only once its removal is confirmed. On a
later attempt, before each stale check, a POSIX lock whose token belongs to an
exited process of this machine and boot, quiet for three minutes, is backdated
so the existing stale takeover claims it at once. The fence path now shares
the same helper.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a desktop that met another desktop's update fence clears its note once the host answers (#25995)

The serving note "holds this host" and a fence-busy update deferral stayed until a reconnect,
minutes after the other desktop's update finished. The connect now rechecks serving and the
update every 45s while the fence holds, and publishes the first answer without it. A recorded
deferral is dropped once the host runs its candidate or a newer release.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* test(ci): run Phase 3's SQLite-backed tests in the Node runtime project

Main's #25967/#25998 boundary requires every test that opens real SQLite to be listed. This adds
Phase 3's eleven orcad and SSH migration tests, plus main's own agent-launch-instant-tab test
(#25430), which main's tip also leaves unlisted.

* fix(ci): keep Electron probes out of the node-server suites again (#26046)

The runner excluded *.electron.test.ts with a CLI --exclude, but main's switch to Vitest inline
projects (#25967) gave each project its own exclude list, which overrides the CLI one. The
directory selectors then pulled profile-state-writer-stall.electron.test.ts into the glibc-floor
and musl orcad-template jobs, which have no xvfb. Resolve the exact files with vitest list and
drop Electron and cross-runtime ones before running.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): keep the SSH host card quiet while its managed server is healthy (#26072)

The card showed "Runs a managed Orca server" under every healthy host. A managed server is the default, so the status line now appears only for setup progress, updates, the relay, or failures.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move to managed server keeps the host's terminal tabs (#26077)

* fix(ssh): Move to managed server keeps the host's terminal tabs

Move stops the relay shells; their exits read as a user exit and closed
the tabs before the conversion copied them to the server. Suppress those
exits for the move, restart stopped shells on the relay when the host
stays, re-home the open workspace onto the server, and report stopped
shells to the runtime so terminal list stops calling them connected.

* fix(ssh): mark Move's relay stops in main's intentional-stop register

The renderer-only exit suppression left main retiring the stopped tab from
the saved SSH session before the conversion copied it, left other viewers
unprotected, and swallowed real exits for the whole request. Register
exactly the shells the move stops, from just before each shutdown, as a
'replaced' stop with their incarnation; main keeps the surface and labels
the exit for every viewer, while a confirmed death stays 'exited'. Move
now returns the shells it stopped, and a host that stays on the relay
restarts only those tabs, discarding any buffered exit first.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(hosts): show an SSH host and its managed Orca server as one host (#26076)

* fix(hosts): show an SSH host and its managed Orca server as one host

Phase 3 registers the Orca server it deploys over SSH as its own runtime environment, so every
host list built from the execution-host registry listed the machine twice under the same name.
The registry now folds the pair into one row named after the SSH host. The row routes to the
server, since a managed host has no relay, unless main reports the host back on its relay; the
other id stays as an alias so selections and renames saved under it still resolve. A retired id
that workspaces still point at keeps its own row, and servers no configured SSH host deployed
(manual pairings, orphans) are untouched.

* fix(hosts): keep both ids of a merged SSH host and dedupe only in pickers

Deleting the merged-away id from the registry broke every consumer that matches hosts by exact
id: Add Project fell back to local after a connect, the composer lost ready projects and drafts
(and could swap in an unrelated local project), and a host scope hid folder-only workspaces.

The registry now keeps both entries and marks the pair (aliasHostIds on the row pickers show,
mergedIntoHostId on the other). Pickers show one row per machine, and a choice of that row
expands to both ids: sidebar host scope, jump palette filter, notification toggles, run-target
and repository host offers. Add Project resolves a saved SSH id to its server row and blocks the
actions while that server comes up instead of choosing local. The composer's resolver now fails
closed when a named draft repo isn't actionable rather than picking another project.

* fix(hosts): widen saved host scopes, both-way palette aliases, guard Add Project host

- A sidebar or agents host scope saved by an older build (or before a route flip) can hold one id
  of a merged SSH host; a background gate widens it to both ids so exact-id filters match either
  owner.
- The palette host filter now resolves a saved id to both owners whichever id it names.
- Add Project's create and clone refuse to run while the chosen host is unresolved, and their
  submit buttons stay disabled, instead of falling through to this computer.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(sync): reconcile main's ratchet bundling and cold-serve hydrate with phase3

The Electron-import ratchet keeps main's single-stdin bundle (cjs); the
auto-merge had also kept phase3's esm splitting, which broke main's
import-graph test. Editor tabs now follow the windowless full-seed rule
from #26022, so a cold serve restart lists persisted editors too.

* fix(ssh): reclaim this desktop's own exited lock on Windows hosts too (#26087)

* fix(ssh): reclaim this desktop's own exited lock on Windows hosts too

The relaunch after a quit mid-update now frees the activation fence and the
version-dir install lock on a Windows SSH host the same way it does on POSIX,
instead of waiting out the 20-minute stale window. The host script ages the
lock only when its token belongs to a desktop process proven exited, it has
been quiet for three heartbeats, and (for the fence) no state mutation is
live, where a mutation holder counts as gone only by pid plus creation time.

* fix(ssh): take an exited holder's lock only through the steal arbitration

Review found the reclaim backdated the lock by path after checking it, so a
live successor that replaced the lock in between could be aged and then
stolen, and an interrupted or failed restore left it aged for good.

The exited-holder check is now read-only. The steal command itself accepts
the proven token and, inside its steal claim and identity recheck, also takes
a lock whose owner file still names that token and that has been quiet for
three heartbeats. Nothing is written to a lock before the steal owns it.
POSIX uses the same path.

* fix(ssh): never take an exited holder's fence while a state mutation can start

Review round 2 found the fence's live-mutation guard ran only in the read-only
proof, so a mutation admitted after the proof, or one whose first heartbeat
landed after the steal sampled the fence's age, kept running under a fence
the steal had replaced.

For the fence, the steal now takes the state-mutation lock inside its claim
(mkdir on POSIX, the exclusive owner.json on Windows) and holds it until the
takeover is done; it refuses when any mutation lock exists. Holding it, it
rereads the owner and only then re-samples the fence identity. A mutation now
rechecks its fence token right after it takes the mutation lock and stops with
the fence-lost marker if it changed. The Windows proof also falls back to the
stale window when its command line would not fit cmd.exe.

* fix(ssh): record the exited-owner steal as a real mutation-lock holder

Review round 3 found the POSIX steal held the state-mutation lock as an empty
directory, which a mutation reclaims after a minute without any liveness
check; a steal stalled that long lost its exclusion and could replace the
fence under a running mutation.

The steal now writes its pid (and group, under the same rule) with the
mutation's own noclobber owner writer, so only proof of its exit frees the
lock, and it removes the lock only while the lock still names it. On
Windows the owner record is moved into place whole, so it never exists
empty, and is removed only while it still names the steal's pid.

* refactor(ssh): keep the relay lock commands off the orcad host-script graph

The mutation-lock owner writers moved into a leaf module, so the relay's
install-lock commands no longer import orcad-state-snapshot and, through it,
the Windows host script, orcad-instance-lock and the daemon process query.
Those modules evaluate imports at load time that existing suites mock
partially. No behavior change.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-07 03:15:06 -07:00
Neil 5b0d38749d test: avoid Store imports in terminal session fixtures (#26166) 2026-10-07 03:03:14 -07:00
Neil 81ba5c3125 test: reuse scratch cells in terminal replay oracle scans (#26162) 2026-10-07 02:32:26 -07:00
Neil e6bcc5a6e0 Keep draft reload tests on the existing cache module (#26145) 2026-10-07 01:54:03 -07:00
Neil d58ea0c362 Advance mocked scheduler clocks without real sleeps (#26135) 2026-10-07 01:53:07 -07:00
Neil 6fd09185d1 Stop pane migration fixture hooks before closing stores (#26131) 2026-10-07 01:49:35 -07:00
25e0a29701 feat: stream desktop audio and video previews with native controls (#26120)
Open local and SSH video and music files with native playback controls. Stream bounded byte ranges through a scoped Electron URL instead of loading entire files into the editor.

Fixes #24859

Co-authored-by: ChangJun Park <40492343+ckdwns9121@users.noreply.github.com>
Co-authored-by: Lirone Levy <lirone88@outlook.fr>
Co-authored-by: dupi <david.li.du@gmail.com>
Co-authored-by: John Cusack <5961784+John-Cusack@users.noreply.github.com>
2026-10-07 01:29:55 -07:00
Brennan Benson 535c83a687 refactor(composer): delete the full-create path no caller reaches (#26052)
* refactor(composer): delete the full-create path no caller reaches

NewWorkspaceComposerModal is the only renderer of the composer card, and it overrides the card
props' onCreate (the full-create submit) with its quick create and never reads useComposerState's
submit. The full path (source and submit preparation, creation execution and its finalization,
issue-command, startup and structured-launch helpers, the orchestration) was reachable only from
its own tests. Delete it with them, and drop onCreate and submit from the composer contracts.

* refactor(composer): delete code the full-create path left orphaned

Removing the full-create path left code whose only readers were gone:
- applyWorktreeMeta and the updateWorktreeMeta plumbing that fed it
- createWorktree and setSidebarOpen on the composer target store
- currentIssueCommand on the composer model
- buildAgentPromptWithContext and getLinkedWorkItemPromptContext (only
  their own tests still called them) and those test cases
- the "Selected agent is disabled" locale string in every catalog
- the export on confirmRuntimeIssueCommandRead

The 'full' create-gate mode goes too. Its only caller passed 'quick', so
the default named a mode with no create action. That removes the option,
the full gate, the issue-automation wait flag and the renderer issue
command preload that ran only in 'full' mode. Quick create is unchanged.

The quick path's name-retirement comment pointed at the deleted full
submit path; it now carries that reasoning itself, including the mobile
counterpart that must change with it. An e2e comment no longer names
applyWorktreeMeta.
2026-10-07 01:26:38 -07:00
Brennan Benson e6c298e8a9 fix(native-chat): a Stop Codex took settles the message whose turn never opened (follow-up to #25217) (#26105)
* fix(native-chat): a Stop the agent took settles the message whose turn never opened

When Codex takes a person's Stop on a turn it never opened, no turn-ended
event follows, so the message stayed pending: the chat read Working with
Stop shown, and the next turn's clock counted from that message.

The Stop's settle now handles that case from its own answer: when the
provider names the turn it took and that turn has no record, the sends it
was for are withdrawn by the same rule a Codex child's end after a Stop
already uses. The send's own row then says it was stopped before the agent
started, and Working, the clock and the opening-send hold follow from it
being settled.

Deletes the opening-send hold's special case for a taken Stop's note, which
this makes dead, and moves the Stop note key back next to its only users.

* test(native-chat): a Claude Stop before the echo settles the send through the CLI's end, and the next message goes out

* fix(native-chat): an interrupted Codex turn that never started records nothing, and the withdrawal reads from the handover

Codex can abort a turn before it starts: it answers the interrupt, then
sends turn/completed for a turn that never sent turn/started. The
translator wrote an empty interrupted turn for it, so the person saw that
turn beside the message's own "Stopped before the agent started" row, and
when the end was read before the interrupt's answer the Stop's settle
found a record and withdrew nothing. Codex records a turn's prompt only
once the turn starts, so an interrupted turn this child never started,
with no item or prompt read, now gets no record. Failed ends keep theirs.

The withdrawal now asks whether a turn opened since the send's handover
row, the same point the opening-send hold reads, rather than since its
acceptance: a turn record written in between is not the send's turn.

* test(native-chat): say why the item-read case pins only that the send is not taken back
2026-10-07 01:25:50 -07:00
Brennan Benson 47f4d3f527 feat(agent-launch): the desktop AI buttons start their agent through agent.launch (#25624)
* feat(agent-launch): host-assigned caller identity and a launch record written when the surface exists

Step 1 of the agent-launch unification, on main.

- The dispatcher stamps every request's caller from what its connection proved (runtime socket:
  the local CLI; the desktop's IPC: the desktop; a paired socket: its device). Params never set it.
- The launch record is written twice: once when the tab exists (what creation settled: on the
  launch command, a draft, or a submit still `unconfirmed`), and again once the prompt's fate is
  known. A restart in between finds the running agent instead of answering "unknown".
- A replay re-derives its terminal handle from the pane key in the running host, and shows
  `unconfirmed` only to callers that advertise agent.launch.prompt-unconfirmed.v1.
- The record store opens in its own slot, without building the chat host; the chat host is built
  on that same store.

Rebuilt from this PR's own commits (b9adf88da0, 0d7d9b5b30, c83f44dd72, 39d15d52c0) onto
main, without #24080/#24081. Conflicts: the delivery doc table (main's "line fits" row plus the
`unconfirmed` row), and main's journal-database open in install(), which now goes through the
record-store slot.

* feat(agent-launch): show an agent's tab at once, where the caller asked for it

An agent.launch now shows its terminal tab before admission and spawn, in
the requested placement (group and/or anchor tab), and the pane attaches to
the agent as soon as it runs. A pane whose agent can't start, or whose start
can't be confirmed, says so instead of becoming a plain shell, and keeps
saying so across restarts. A user's close of the tab or its pane during the
launch stops it and answers agent_launch_tab_closed. Whose view moves is
unchanged from main for every caller.

Rebuilt on main (with #24934) from the previous branch head 8f62858e60.

* feat(agent-launch): the desktop AI buttons start their agent through agent.launch

Part 3 of 3 of the agent-launch unification, on main. The AI buttons (Fix
checks, the source-control actions, commit/push recovery, Explain commit,
notes and annotation sends, session continuation) started their agent from
the window. They now ask the host's agent.launch to start it, with no
prompt, into the tab the window makes at the click, under one operation id
per click; the window's pane waits for the host's agent instead of spawning
a shell. The window then pastes the prompt with main's own paste, moved out
verbatim and started only once the host's agent holds the tab, so delivery
and follow-ups are exactly main's. A typed new-tab prompt and launches while
chat is the default keep main's path. The source-control dialog drops its
launch-command preview, which the host builds now.

* style(runtime): one-line the tab-order map so the headless browser-tabs runtime stays under its line limit

The merge of main (#25724's emitMobileSessionTabsSnapshot metadata) plus this branch's close mark put
the file one line over max-lines. Formatting only.

* fix(agent-launch): watch an AI button's agent for readiness from its first output

The window started watching for the agent's ready signal only after the host
answered agent.launch, so an agent that had already turned on bracketed paste
by then was never seen ready and its prompt waited out the readiness budget.

Readiness is now watched from the moment the tab's terminal exists, as main
watches it; the paste is written, and the chat copy seeded, only once the host
says its agent started in that tab.
2026-10-07 01:18:57 -07:00
Neil 385fe8ab4e perf(tests): avoid repeated persistence imports in automation fixtures (#26111)
* perf(tests): cache Vitest module transforms between runs

* test(native-chat): align resume fixture with mention menu props

* perf(tests): reuse the SQLite store fixture in automation suites
2026-10-07 01:04:36 -07:00
Brennan Benson e5ade4b868 fix(worktrees): when the host can't fetch a remote base, create from the local branch and say so (#26009)
* fix(worktrees): create from the usable local base when there is no tracking ref, and say so, on the host's create

The host's create (agent.launch, worktree.create, the phone and the CLI) asked for a remote base like
origin/main with no tracking ref yet fetched it, and threw offline, even when the local branch existed;
it also never reported a base it fell back to. Like the desktop create, it now uses the local branch
without fetching and returns baseFallback, which the desktop window already turns into a notice.

* fix(worktrees): fetch the remote base first and fall back to the local branch only when the fetch fails

The host's create skipped the fetch whenever a local branch matched a
remote base with no tracking ref, so online phone/CLI/agent.launch creates
silently built on a possibly stale local branch, every time. Fetch first,
as main does; only when that fetch fails use the local branch the remote
names and report it with baseFallback. With no local branch the existing
error is unchanged.

Decide the base before naming so branch reuse and conflict checks, the add
and the persisted baseRef all see the base actually used, and return
baseFallback as its own field of the create result instead of riding on the
git add result.

* fix(worktrees): report the fallback for a local ref of the requested name, as the desktop create does

The tracking ref is still missing there; only the fetch is skipped, as before.

* test(worktrees): pin that the host create names only after the base fetch settles

Also pins the existing not-found-after-fetching error, now owned by the base decision.

* fix(worktrees): fall back to the local base only for creates a person asked for; keep the exact-name ref silent

Scheduled automations, orchestration workers and federation keep main's
"Check your network" error offline: nobody is watching to be told the
workspace was built on a possibly stale local branch. The fallback is a
host-side create option (allowLocalBaseFallback, default off, not on the
wire) that the worktree.create RPC and agent.launch set unless the request
carries automation provenance.

A local ref of the exact requested name goes back to silent, as on main's
host: reporting it showed "may not include the latest remote changes"
online for plain local branches.

* test(worktrees): pin the local-base opt-in at agent.launch and its absence for an automation's worktree.create
2026-10-07 00:55:21 -07:00