mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 16:02:37 +00:00
* Revert "revert: take the 26 Phase 3 (#16741 port) PRs back out of main (#24559)" This reverts commit5f308bfa9c. * feat(orcad): Windows remote primitives for managed orcad hosts (W1) (#24525) * feat(orcad): Windows remote primitives for managed orcad hosts (W1) * refactor(orcad): run Windows host ops as node.exe with plain argv, no PowerShell hop * fix(orcad): refuse secret-shaped names on the breakaway launcher's --env * fix(orcad): name the secret env guard for its role --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9) (#24521) * feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9) The destination half of a catalog migration: an orcad stages a T6-7 manifest (repositories, project groups, folder workspaces, dormant session, client, automation and worktree metadata, retired names, scrollback snapshots) with exclusive claims, then commits it with a receipt so a retried commit returns the same receipt and never imports twice. Served as orcad.migration.* runtime RPC behind the orcad.migration-catalog.v1 capability; the client refuses a host without it or with method-not-found, and any other failure is left for the caller to recheck. Dormant only: no live PTY projection. Inert on the desktop until T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(rpc): catalog the orcad.migration params in the shared contract; name the catalog-import install target Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(ssh): journal and fence an SSH host for dormant migration, gated on proven terminal exit (#16741 T8-c1+c2) (#24522) A migration from a relay-hosted SSH target into a managed orcad now starts with a journal in its own sidecar directory, then the target's managed-owner fence, then a profile flush, before any remote call. A fence with no journal is unverifiable and never released; a journal whose fence is gone is stale and grants nothing; an unreadable journal fails closed. The fence requires every terminal the target ever leased to be proven exited, checked before the fence (with the relay's process list) and again under it. Same-owner claims now need the durable record that explains them. Inert until T8-c4/T6-10. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(orcad): run orcad itself on Windows hosts (W2) (#24529) * feat(orcad): run orcad itself on Windows hosts (W2) * test(orcad): load the ConPTY smoke's addon from out/orcad so the temp slot can be removed * test(orcad): skip the foreign-uid lock case when running as root --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a connected but unused SSH host previews as movable (#24609) The untransferred-dependency census counted activeConnectionIdsAtShutdown naming the target as workspace-session state. The renderer rewrites that list on every connection change, so merely connecting to an empty host blocked the move. It is a reconnect hint; the remote work it can stand for is counted on its own. The empty-target claim check likewise ignores global-field copies inside the host's session partition. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3) (#24523) * feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3) The coordinator re-checks before every stage and commit that the fenced source still exports the journaled manifest, carries no untransferable state and started no terminal. A lost answer is re-read from the destination's catalog state; only a committed read whose receipt matches the journal advances it, and the journal is on disk before anything returns. Abort releases the fence only on proof the destination holds nothing, or on an unsupported destination before anything was staged, and never once the destination committed. Codes against T6-9's catalog client through an injected interface. Inert until T8-c4/T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(ssh): import the T6-9 client's unsupported refusal instead of mirroring it --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3) (#24563) * feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3) * test(orcad): exhaustive op switch in the Windows lifecycle fake * test(ssh): narrow the Windows host-cell descriptor by lane before building a relay cell --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11) (#24608) * feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11) * fix(serve): keep orcad selection app-side and wait out Windows temp cleanup * refactor(orcad): move the data-root privacy check out of the instance lock --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch (#24619) * test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch * ci(e2e): install ripgrep for the serve mode-switch job's window-manager wait --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4) (#24562) * feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4) The conversion entry resumes or takes the fence, deploys and pairs the managed server into it, marks the server as migrated, then stages and commits the dormant catalog. Every step is keyed by the journal, so a repeat after a crash, deferral or lost reply resumes the same migration. Status reports an unfinished migration, and rollback is refused while one runs or when the rollback snapshot predates the migrated catalog. Inert until T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * style: oxfmt the c4 conversion and maintenance files * fix(ssh): name the fake migration destination's type so declarations stay portable --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4) (#24570) * feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4) * fix(orcad): accept a managed stop request whose lock path is spelled with Windows client separators --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5) (#24565) * feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5) Once the journal records destination-committed, the source profile drops the manifest's repositories, folder workspaces and unreferenced project groups, its dormant session, automation, client and worktree state, and the leases the fence proved exited. The profile flushes, the retirement is verified, the journal moves to source-retired and compacts once the server matches. A retry after any crash repeats idempotent work. The fenced target stays: it carries the managed server's tunnel. Conversion now ends retired. Inert until T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(orcad): retirement drops the migrated host from the reconnect hint The census no longer treats activeConnectionIdsAtShutdown as untransferable (#24609), so retirement must remove the target from it; otherwise a restart dials a host that is now a managed server. * style: oxfmt the c5 conversion file --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1) (#24579) * feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1) * test(orcad): start the Windows lane's exec spy after the relay gate prelude restores its own --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI) (#24590) * feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI) Adds a Managed servers section under Remote servers (deploy an empty server, status with deferred-update and migration states, update, rollback, recover, stop and cancel-stop, and SSH access for paired servers), and a Move to managed server action on connected macOS and Linux SSH hosts with a preflight summary, a terminals-closed confirmation and a resumable progress view. Main wires the conversion to the relay's process list, the direct session and the T6-9 catalog client. Everything is hidden until the new experimental setting is turned on, and Windows SSH hosts are never offered. Merges the T6-9 branch (#24521) until it lands on the integration branch. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(settings): align managed-server form controls and name the section the setting reveals Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(orcad): ask the relay with an absolute deadline via the W5 terminal-gate lister The conversion wiring passed a relative 10 s as listProcesses' deadlineMs, which the provider reads as an absolute time, so every relay inventory timed out after 1 ms and the terminal gate could never prove exit. Adopt #24579's lister verbatim so the stacks merge cleanly. * feat(settings): name blocking saved state in plain, localized words The move preview listed internal dependency ids such as workspace-session; each kind now has its own catalog entry. * feat(settings): offer managed servers and the move on Windows SSH hosts W1-W5 are on the integration branch, so a Windows relay-hosted host can deploy, convert and retire like a POSIX one. The move still waits for a connected relay that reported its platform. --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix (#24865) * test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix * test(ci): expect the orcad Windows host cells in the SSH Windows hosts workflow --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(orcad): wait for killed terminal daemons to exit before removing their temp profiles (#24871) Co-authored-by: m4air <m4air@Mac.localdomain> * feat(updater): read rollout kill switches from the update-campaign payload, all inactive (#24867) The nudge request Orca already polls may now carry an optional versioned rollout block naming the Node runtime flips. A typed reader resolves each flip with kill-switch, version range and install-id-bucketed percent semantics, falling back to the baked value (every flip inactive) when the block is absent, invalid or never read. No consumer reads it yet. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(telemetry): report Windows security-software refusals and unverifiable or failed runtime checks (#24866) ssh_remote_runtime_resolved dropped the self-test's security_software refusal to 'none' and sent nothing when a self-test was unverifiable or failed, because those attempts throw before a rung settles. Add the refusal value, self_test 'unverifiable', and an outcome field (resolved | unverifiable | failed) deduplicated per host and outcome per session. Co-authored-by: m4air <m4air@Mac.localdomain> * refactor(settings): move WorktreeVisibilityDefaults into its own module (#24884) global-settings-types.ts sits at the 300-line max-lines ceiling; merging main's two new agent-state-rules settings with experimentalManagedServers put it at 301. The worktree visibility defaults type moves next to the other visibility types and is re-exported so its 38 importers are unchanged. Co-authored-by: m4air <m4air@Mac.localdomain> * refactor(watcher): move the supervisor's child, terminating child and canary into a child slot (#24887) * refactor(watcher): move the supervisor's child, terminating child and canary into a child slot * fix(watcher,runtime): take the child slot's child type from the shared wrapper, and stub main's title-display clear in the projection test --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: m4air <m4air@m4airs-Air.localdomain> * ci(adhoc): build and ship the orcad template in adhoc macOS and Windows builds (#24969) Co-authored-by: m4air <m4air@m4airs-Air.localdomain> * feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy (#24973) * feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy * fix(orcad): keep the in-use slot plus the two most recent others, and prove same-version reuse survives eviction --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> * Phase 3: managed orcad on every SSH connect, with downgrade-safe fence (#24975, #24979) * feat(orcad): fence managed hosts outside owner and keep converted projects for downgrades Shipped builds hide any SSH target with an owner, so a downgrade made a converted host and its projects vanish. The managed fence moves to a new orcadFence field (older builds keep but ignore it), phase-3 owner fences migrate on load, and managed hosts stay visible but refuse a direct relay. Conversion now stops at destination-committed with sourceRetainedAt; source retirement waits on the baked-off orcad-source-retirement rollout flag. This build hides retained rows, and a start that finds an older build changed them marks the host sourceChangedAt (relay, needs a new move) instead of merging a second manifest. * feat(ssh): every SSH host runs managed orcad, decided on each connect (#24979) * feat(settings): managed servers are no longer experimental Remove experimentalManagedServers and its gates; the Managed servers section always shows. Loading drops a stored value, which an older build reads back as its own default (off). * feat(ssh): every SSH host runs managed orcad, decided on each connect Before a relay is started, the connect decides the host's server: a converted host connects through its tunnel (retiring a retained source once orcad-source-retirement is on), an empty host deploys orcad, and a host with Orca state converts through the journaled migration. A host with live or unproven relay terminals keeps the relay this session and converts later. A refusal for any other reason keeps the relay and names the blocker. A host orcad can't run on (unsupported target, no template, native preflight, runtime self-test) releases any claim or fence it took, records why with this app version, and keeps the pinned-relay ladder. Progress and the decision ride an optional SshConnectionState.managedServer field (dropped by older clients' admission). SSH Hosts shows each host's server status; the manual move dialog is removed. * fix(ssh): always await the connect's server decision, after the provider authority rotates The decision now runs after the old session and transport are torn down and the authority has rotated synchronously, so concurrent connects still join one attempt; a shutdown that began during the decision wins over the rotation. The IPC tests use an async double, no sync branch. * test(renderer): the IPC events store double carries its SSH connection states --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> * test(ssh): prove connect-time conversion to managed orcad on real Linux and Windows hosts (#24981) * test(ssh): prove connect-time conversion to managed orcad on real Linux and Windows hosts, and the downgrade view * fix(e2e): read the SSH host's session partition, provision the convert cell's account, and keep rollout file overrides out of packaged builds * test(e2e): report the full relay state when the convert poll times out * test(e2e): require the relay to list no terminals before the converting connect * test(e2e): require no running-terminal lease before the converting connect * test(e2e): report the host's terminal leases when the converting connect keeps the relay * test(e2e): end the relay era with no relay shell left to respawn * test(e2e): settle before the converting connect and name the tabs a blocking shell belongs to * test(e2e): use the exited relay tab as the session tab, since any mounted tab starts a shell * test(e2e): carry an editor tab through the conversion instead of a terminal tab * test(e2e): log the conversion census inputs before the converting connect * fix(orcad): log which state blocked a refused conversion * test(e2e): log both session partitions before the converting connect * fix(orcad): a source partition's copy of focus on another host no longer blocks conversion * fix(ssh): an ssh2 forward on port 0 reports the port it bound, so managed tunnels pair * test(e2e): give the conversion its full budget again * test(e2e): report the migration journal phase when the conversion stalls * test(e2e): report the connect's own result and main's state when the conversion stalls * fix(ssh): a converted host's managed state reaches the renderer instead of staying on 'converting' * test(e2e): print a failed server call's response * test(e2e): give server calls the budget a fresh server's first inventory needs * test(e2e): read the converted worktree's tabs with a scoped session.tabs.list * test(e2e): log the converted worktree's tabs instead of asserting them, pending the server-side fix * test(e2e): prove retirement by the dropped source rows; the journal compacts away after it * test(e2e): drop the conversion diagnostics now the cell passes * test(e2e): fail a hung disconnect or connect with main's state instead of the whole budget * test(e2e): convert an upgraded relay-era profile's host on its first connect, on Docker and Windows * test(e2e): seed the relay-era target the way addTarget registers it * fix(ssh): a stale ssh2 forward drops a late connection instead of crashing main on 'Not connected' * fix(orcad): the active-slot readiness probe reads Windows hosts through the host script * ci(ssh-windows): let only the convert cell's account open the SSH local forward its managed server needs * test(e2e): match server paths in their JSON-escaped form, for Windows backslashes --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): move what an older build added to a converted host, or keep the server's version (#24980) * feat(orcad): move what an older build added to a converted host, or keep the server's version A host an older build changed stays on the relay with two actions. 'Move the new projects' runs a fresh journaled conversion of only what the host's earlier migrations didn't move: the source is viewed with those migrations retired from it, so nothing that overlaps the server is merged, and a row the server already holds fails the whole move at stage, before any commit. Its journal supersedes the chain head; retained-source checks compare against its baseline, and retirement retires every manifest in the chain only after the newest committed. 'Keep the server's version' records the current source as the baseline and returns the host to its managed server. * test(ipc): the runtime environment handler contract lists the delta-move channels * test(native-chat): snapshot the journal directory after the attach's restart-offer lock is released --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): converted hosts keep their editor tabs, forwarding-refusing hosts stay on the relay, listAll settles (#25099) * fix(orcad): publish the headless graph so session.tabs.listAll settles instead of hanging * fix(ssh): a system SSH forward on port 0 picks a free port first and reports it * fix(runtime): a headless host lists and closes the editor tabs its session holds, so migrated editors reach clients * fix(ssh): keep a host that refuses TCP forwarding on the relay, and release a conversion it stranded * test(e2e): assert the migrated editor tab and a settled listAll, and keep a forwarding-refusing host on the relay * fix: restore the journal import after rebase, type the probe's failure code, and update headless-graph test seams --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): client focus and 'local'-stamped tabs no longer refuse a host's move to its managed server (#25110) * fix(orcad): client focus and 'local'-stamped tabs no longer refuse a host's move to its managed server * fix(orcad): a v1.4.218 profile focused on the SSH worktree converts, and a refusal names what blocks it The debounced session writer never patched activeWorkspaceKey or activeWorkspaceExecutionHostId, so the first focus stayed on disk. The all-dependency census also counted global focus copies in every non-source partition. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): offer to move a host whose open terminals keep it on the relay (#25097) A host with any open SSH terminal never converted: the connect gate keeps the relay while relay terminals are live, and an open tab respawns them on every connect. The first such connect per host per app version now marks the relay status with offerMove (recorded as managedServerMoveOffered), which toasts "Move <host> ... Its N open terminals will restart." The SSH Hosts status line keeps a "Move to managed server" action while terminals are live. Confirming runs ssh:moveToManagedServer: stop the host's relay terminals through the extracted ssh:terminateSessions path, re-run the connect gate's terminal census, refuse on anything but exited, then reconnect so the connect-time decision runs the journaled conversion. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): reach managed orcad over an SSH stdio bridge where sshd refuses port forwarding (#25120) * feat(orcad): reach managed orcad over an SSH stdio bridge where sshd refuses port forwarding Hosts with AllowTcpForwarding no were kept on the relay. The managed tunnel now probes forwarding each time it starts and, on refusal, serves the same local port through a second provider: each accepted socket opens one SSH exec channel running a small bridge on the host's pinned Node, which dials orcad's loopback port. POSIX hosts run it with node -e; Windows hosts run it as the content-addressed host script's stdio-bridge op with base64 line framing. Bridges are capped at 8 per connection under sshd's MaxSessions default, and a lost channel only drops its socket. Only a host where even the bridge cannot run keeps the relay, recorded as ssh_tunnel_unavailable; the older tcp_forwarding_refused record is retried. * test(e2e): prove the stdio bridge by refused forwarding plus a working call * test(e2e): connect the refused-forwarding host without the relay-only repo step --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): report each connect's server decision, and show orcad.log's last lines when setup fails (#25118) * telemetry: one enum-only event per connect decision (outcome, reason, tunnel transport, host platform, duration), plus conversion start/commit/fail, deploy failures, and the per-host move offer and its result * errors: deploy, launch and rollback failures carry the last 40 redacted lines of the host's orcad.log, read over SSH (Windows through the node.exe host script) * SSH Hosts: a deferred or failed setup offers its reason, log tail included, under Details Co-authored-by: m4air <m4air@Mac.localdomain> * fix(startup): Windows never crashes resolving userData when roaming AppData is unavailable (#25113) * fix(startup): pin Windows appData and userData before anything resolves them A Windows session without a loaded profile (e.g. orca serve over SSH) can fail the roaming AppData known-folder lookup. Electron 43 then falls through to Chromium's userData provider and crashes natively. Resolve appData first (falling back to APPDATA, then USERPROFILE\AppData\Roaming), and set userData explicitly so Electron's provider never runs. * test(startup): remove the AppData fixture through the retrying helper --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): keep a terminal the previous Orca version's relay still runs instead of replacing it (#25124) After an app update the new relay answers "not found" for a PTY the previous build's relay still runs, because the old relay refuses this build's handshake. The client read that as absence: it expired the lease and the pane cold-restored into an empty shell while the user's shell kept running, unreachable. Each deploy now takes a census of this target's older relay endpoints; while one is live or unverifiable, a not-found reattach keeps the lease and the pane binding, and the pane says the terminal is still running under the previous Orca version. A detached lease also keeps blocking managed-server conversion until that terminal exits. The cross-version harness now extracts src/relay, and a new test drives v1.4.218's relay socket and grace lifecycle with this build's endpoint probe: the probe leaves no grace deadline and reads the old relay as live work. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): update a managed server on connect when it runs an older build and is idle (#25122) * feat(orcad): update a managed server on connect when it runs an older build and is idle A connect to a managed SSH host now runs the Managed servers update when the host's orcad differs from this app's bundled build, the template carries the host's target, and the update planner finds no live or uncounted terminals. A rejected candidate is restored through the activation journal; the reason is recorded per app version so later connects don't retry it. A host a newer Orca activated is never downgraded: the activation record now names the app version behind each build, and an explicit rollback holds the build it left. * test(e2e): connect without a racing disconnect after relaunch, and report each attempt On launch the app already reaches the managed server through its tunnel; a disconnect racing that restore cancelled the connect that runs the update. * feat(orcad): run the update check when the launch restores a managed server's tunnel An auto-restored host may never see an SSH connect, so its server would never update. The tunnel restore now runs the same check, once per server per session and off the caller's path, through the shared update-check module; a server mid-migration is left alone. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): retry the pinned Node download through transient network failures (#25154) A Chromium network change (net::ERR_NETWORK_CHANGED) during the pinned Node download failed the deploy and sent the host back to the relay. The archive download now retries up to three more times, after 1s, 3s and 9s, on dropped connections, timeouts, stalls and retryable HTTP statuses, removing the partial file first; checksum mismatches, other HTTP errors and cancels stay final. The transient-error classifier moves from the speech download to src/main/network/transient-download-error.ts so both share it. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(serve): run orca serve on orcad by default, with ORCA_SERVE_RUNTIME=electron as the opt-out (#24972) * feat(serve): run orca serve on orcad by default, with ORCA_SERVE_RUNTIME=electron as the opt-out * test(serve): read Electron serve's pretty-printed readiness in the CLI mode-switch e2e * feat(serve): gate D7 on Windows and serve on orcad there by default * test(serve): take the profile lock in the CLI mode-switch e2e, and keep Windows profile logs on failure * test(serve): tell Electron and orcad serve apart by readiness health, and trace Windows startup * ci(e2e): dump Electron's native log and stack on the Windows serve mode-switch job * fix(serve): keep Windows on Electron serve until it can adopt orcad's daemon The Windows D7 job shows Electron serve exiting before its window when it relaunches onto a terminal daemon orcad forked. Restore the win32 fallback and skip that case there as a known gap; the follow-up PR fixes it and re-flips Windows. * fix(serve): let ORCA_SERVE_RUNTIME=orcad opt in on Windows while Electron stays the default * test(serve): skip the Windows D7 cases where orcad forks the daemon, and stop cleanup hiding a failed relaunch Test 3 hits the same Windows gap as test 2: orcad forks its own daemon there, and Electron crashes at startup beside it. A failed relaunch also made dispose close the old, already closed app, whose throw replaced the launch error. * test(serve): run every Windows D7 case, with the isolated home's AppData in place Electron 43 crashes natively (0xFFFF7003) when it resolves userData and Windows cannot find roaming AppData. The e2e home isolation points USERPROFILE at a fresh folder with no AppData, so later launches hit that. The harness now creates it, every D7 case runs on Windows again, and a new case proves Electron serve starts beside another profile's live orcad daemon. * test(serve): retry removing a Windows e2e profile while a killed daemon releases it --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): resume a terminal the previous Orca version's relay still runs, through that relay's own bridge (#25170) * feat(ssh): resume a terminal the previous Orca version's relay still runs, through that relay's own bridge After an app update the previous relay keeps the user's shells alive but refuses this build's handshake. Its own relay.js --connect, run from its own version directory, presents its own bundle hash, so on POSIX hosts the client now reaches it that way: a pane whose reattach the current relay held for an older relay opens a route through the old bridge, takes the PTY owner role without output flow control, reattaches the PTY with its replay, and routes every later operation on that id to the old relay. When the last pane a route serves exits, the route hangs up and the old relay's own idle grace retires it. Windows hosts, relocated short sockets and unreachable bridges keep the held-pane behaviour. The cross-version harness now builds v1.4.218's relay from its tagged sources and runs it as the real detached daemon: a shipped client leaves a shell in it, this build resumes the pane through the old bridge, types into it, sees its output, and watches the old relay exit on its own after the shell does. * fix(ssh): read the legacy relay router through its instance --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a host re-upgraded after a downgrade can move what the older build added (#25179) The delta view left the moved projects' session state in place whenever it held anything unmovable, so it then counted against the delta. Relay PTY bindings, shutdown markers and the relay consumer's recovery record blocked every move although the terminal gate already proves those terminals exited before any move commits; the move now drops them instead. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(worktree-ps): report a lost-contact host's terminals as unverifiable, not live:0 pty:no (#25167) During a network drop to a relay-served SSH host, every PTY on it reads as an unconfirmed exit, so worktree ps skipped them and printed live:0 pty:no for a terminal that was still running. The listing now counts terminals whose liveness verdict is unverifiable into a new optional unverifiableTerminalCount, and the CLI prints live:unverifiable pty:unverifiable (JSON carries the same word) instead of zero. A host-confirmed exit still reads as zero. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(terminal): a held pane shows no client OS or shell, and the boundary doc says POSIX hosts resume it (#25194) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): tunnel to the port managed orcad bound and verify it is ours (#25182) When another Orca already listens on 6768, orcad binds a different port. The tunnel kept forwarding to 6768, reached the other runtime, was rejected with 4001, and the connect still reported a managed server. The tunnel now reads orcad's bound port from its active slot's readiness (falling back to the persisted port for slots without one) and, after the forward is up, proves the server answering is the paired runtime. On a mismatch it re-reads the port once and fails with orcad_identity_mismatch instead of reporting managed. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a converted host shows only its managed server's rows, and its old editor tabs load there (#25193) Retained relay-era project groups and the per-host SSH catalog now hide like repos and folders. On the managed transition the renderer reloads server names, groups, folders and worktrees, then drops the host's relay-era rows a local refresh would keep. A restored tab with no host stamp in a workspace now owned by a managed server takes that server as owner, in place. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * fix(orcad): ship the port-scan worker so managed servers detect workspace ports (#25197) orcad resolves port-scan-command-worker-entry.js beside orcad.js, but the orcad build never emitted it and ORCAD_ARTIFACTS never listed it, so every managed server logged 'probe worker unavailable' and had no port detection. The build now emits it with the other children, and the artifact list carries it, so it is uploaded, hashed and covered by the template contract. A new test bundles orcad and its children and fails when the bundle names a worker or child entry the slot does not ship. Two existing gaps it found, session-scanner-service-entry.js and wsl-transcript-fs-process-entry.js, are listed as known and may only shrink. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): keep retrying when a reconnect loses the socket before the SSH banner (#25195) After a network outage, a port forwarder or NAT can accept the TCP connection and then close it before the server sends its banner. ssh2 reports that as 'Connection lost before handshake' with no errno, so the reconnect ladder classified it as permanent and published 'error' with no retry scheduled. Treat it as recoverable on the bounded ladder only; the initial connect keeps its narrow classifier. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(serve): serve on orcad by default on Windows too (#25162) D7 now runs every case on Windows (orcad-serve-mode-switch-windows), including Electron serve adopting a daemon orcad forked, so Windows no longer needs the Electron default. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): the terminal gate asks the relays before treating a detached terminal as running (#25200) * fix(orcad): the terminal gate asks the relays before treating a detached terminal as running, and retiring a host drops its relay recovery record * fix(ssh): earlier-relay census gaps and an unanswered relay stay unverifiable; asking the terminal gate changes nothing - the gate is read-only; the conversion and delta move retire proven detached leases themselves - an expired lease also needs every earlier-build relay to answer before it reads as exited - a failed, input-less or truncated census marks its older relays unverifiable, so reattach holds - a legacy relay route closes only when no attach or listing still awaits it - a disposed session forgets the census it started --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): recover interrupted conversions and delta moves, and trim dead migration code (#25226) * refactor(migration): drop the unused full dependency census * refactor(migration): one table for the routed UI fields a migration carries * refactor(migration): share the session focus field list and fix stale claim comments * fix(session): keep a fenced SSH host's source session partition intact through renderer saves The renderer cannot see a fenced host's repos or folders, so hydration drops their worktree rows; a save that still routes any row to ssh:<target> (a folder workspace key stays valid) rewrote that partition without them. That changes the migration's source manifest and loses the session a downgraded build reads back. Main now ignores renderer writes to a fenced, unchanged host's partition. * fix(migration): resume or back out unfinished conversions and delta moves, serialize delta moves, clean stale journals - On connect, a registered server whose conversion never committed resumes the commit; a failure aborts it through the destination, unregisters the server and releases the fence. - An interrupted delta move keeps its journal and gets its changed mark back; the next move resumes it from the journaled manifest or aborts it when the source has moved on. - Delta moves run under the target lifecycle queue and re-check the head, so two concurrent moves cannot write two journal heads. - Delta checks re-read the live source instead of a frozen copy. - Stale journals from a stopped or never-registered server are cleared. - Retiring a chain goes newest first and compacts only at the end. - The journal schema tolerates fields a newer build adds. * fix(migration): hide source rows only after commit, and retire what an older build added to a moved project - Source rows and the renderer session guard share one rule: hidden once the server committed, shown while a migration is still in flight. - A delta move a crash interrupted gets its changed mark back on the next start. - Retirement removes worktree metadata and automations by moved-project scope, so an older build's additions no longer fail the leftover check forever; that check ignores rows of other migrations in the chain. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): survive a slow first daemon start and close review gaps (#25213) - Give the terminal daemon 30 s to start on Windows and retry the spawn once before falling back, so a cold first deploy is not refused as daemonless. - Say in the activation refusal that orcad.log holds the daemon's error and that the next connect retries. - Managed stop completion now waits out daemon retirement plus the shutdown deadline. - Publish the instance lock atomically and reclaim an abandoned torn lock. - Fix the stop listener closing before it was defined on an already-present request. - `orca serve`'s cache prune keeps the slot a running local orcad uses. - A timed-out daemon retirement reopens admission once it is refused; a retirement that may have reached the daemon stays fenced. - A headless host keeps an editor tab's unsaved draft unless the close is forced. - Unverifiable terminals are attributed by the host's PTY record, like live ones. - An explicit --user-data-dir is no longer overridden on Windows. - Read Windows daemon process identity from the process table, not PowerShell. - Remove the unread data-event incarnationId and daemon health runtime fields, share one process-alive and error-code check, and fold the websocket limits file back. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(cli): stop, update, roll back and recover managed Orca servers from the CLI (#25201) * feat(cli): stop, update, roll back and recover managed Orca servers from the CLI orca environment status|update|rollback|recover|stop|cancel-stop call the same managed-server actions as Settings > Managed servers, over new managedServer.* runtime RPC methods. The desktop main process registers those actions; the runtime advertises managedServer.v1 only then, and the CLI refuses on a runtime without it or one that answers method_not_found. stop requires --yes. * fix(ipc): keep a missing managed-server selector a rejection, not a synchronous throw --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): Move to managed server converts the host it is offered for (#25196) The move stops the relay terminals and passes the terminal census; with #25179 the conversion no longer refuses on the stopped terminal's saved tab, layout leaf and pane incarnation. A new test drives one live relay terminal through Move to a conversion whose manifest carries the tab without its relay PTY, so it spawns a fresh shell on orcad. ssh:terminateSessions also no longer records a not-found shutdown as terminated while an older Orca build's relay may still run the PTY (#25124): that terminal is reported unverifiable and its lease kept, so the move refuses instead of converting over a running shell. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): deploy the glibc 2.17 runtime on old-glibc hosts, and fall back to the pinned relay (#25180) * fix(orcad): deploy the glibc 2.17 runtime on old-glibc hosts, and fall back to the pinned relay Managed orcad picked its runtime from the host's libc flavour alone, so a CentOS 7 host (glibc 2.17) got the default linux-x64-glibc Node, whose self-test fails there. The deploy now picks the runtime by glibc the same way the relay ladder picks rung A or B, and the host-side slot checks accept the compat target it ships. When orcad still can't run, the host is recorded as such and the relay connect that follows now runs the pinned-runtime ladder instead of defaulting to the host-Node relay, so a host with no Node lands on rung B rather than failing. The hostile-host lane deploys managed orcad on a fresh CentOS 7 host and asserts it runs on the compat runtime with no default runtime uploaded. * test(ssh): CentOS 7 cell asserts the compat runtime pick, then the refusal and relay rung B fallback The compat template still ships the base @parcel/watcher binary, which needs GLIBCXX_3.4.20; CentOS 7's libstdc++ stops at 3.4.19, so the candidate's preflight refuses. The cell now pins that end-to-end behaviour: compat runtime picked and uploaded alone, refusal classified native_preflight, and the relay that follows settles on rung B with no runtime setting. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(serve): prove D7 on an installed Windows app whose daemon runs from the relocated daemon-host (#24976) * test(serve): prove D7 on an installed Windows app whose daemon runs from the relocated daemon-host * test(serve): read the pre-switch scrollback best-effort and wait for it after reattach * test(serve): opt the packaged Windows serve switch into orcad explicitly --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): a retirement that fails after deleting rows resumes instead of reading as changed (#25234) The chain head records sourceRetiringAt before any row is retired; the startup change check and the chain's own comparison skip a head that carries it. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): managed orcad idles out after 15 minutes and is started again whenever it is down (#25121) * feat(orcad): a managed orcad stops after 15 idle minutes and starts again on the next connect A client-launched orcad now exits, like the relay, once no client, terminal, working agent, staged migration or activation fence has been seen for 15 minutes. The exit is the normal graceful shutdown, which leaves the terminal daemon running; the daemon retires only if it proves itself empty. A record in the data root tells the next start (and its readiness health) that the stop was an idle one rather than a crash. On connect and after host resume, a fresh tunnel whose server does not answer starts the activated slot under the activation fence, but only on a proven exit, so a stopped server reads as not running rather than a failure. * fix(orcad): keep orcad-entry under max-lines; idle e2e connects without a relay repo * fix(orcad): deploy and rollback launches carry the managed idle-exit fence The candidate launch in activation and the rollback launch built their own launch spec without the activation root, so a freshly deployed orcad never enabled idle exit; only the wake path did. The field is now required on every launch spec, so the type system covers each launch site. * feat(ssh): start a stopped managed orcad on connect, on restore and after resume Once orcad stopped (idle, kill or host reboot), a connect still resolved managed over a forward to a dead port and every call failed. Every connect now checks the server behind its tunnel, as does a call through a restored environment; a server proven stopped is started from its activated slot under the activation fence, adopting a surviving daemon and its terminals. The status line shows the start, and a start that fails keeps the host managed with the reason and orcad.log's tail, never as a terminal verdict. * fix(ssh): reuse a serving verdict only on the same SSH transport, for 5s A reconnect right after a reboot was answered from the previous transport's cached verdict, so the stopped server was never started. * fix(ssh): key the serving verdict on the tunnel's remote port too * test(ssh): a stopped server starts before the update counts its terminals * feat(ssh): check serving at the bound port, and follow a restarted orcad to a new one The serving check uses the port the tunnel forwards to (the one orcad bound). A restart that binds a different port drops the forward and rebuilds it at the new port, within the same ensure or on the explicit connect check. The tunnel manager class moves to its own file to stay under max-lines. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a connect with no leases asks the host's relay endpoints before converting (#25223) * fix(ssh): a connect with no leases and no relay session asks the host's relay endpoints before converting * refactor(ssh): one isLiveSshPtyLease for every lease-liveness check * fix(ssh): the host relay census asks a relay its PTYs before reading it as live work An accepting relay whose holders or children the probe could not read (no lsof, an unrecognised service child) read as live work, so a relay-era host that had exited its last shell never converted. The relay's own bridge answers pty.listProcesses without the owner role. * fix(ssh): the host relay census runs each relay's probe and bridge on the runtime it runs on Pinned-ladder hosts often have no Node on PATH, so a PATH Node read every relay as unverifiable. Each daemon's own argv names its pinned runtime, or the host Node a legacy relay started with. * fix(ssh): an older relay's bridge runs on that relay's own runtime --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): build @parcel/watcher into the glibc 2.17 compat slot, so CentOS 7 runs managed orcad (#25199) The compat target swapped in only node-pty, so it shipped the base target's upstream watcher.node, which needs GLIBCXX_3.4.20; CentOS 7 stops at 3.4.19. orcad's preflight refused it, and relay rung B lost file watching without saying so. The compat slot now compiles @parcel/watcher from the package's own sources against the pinned headers with the C++ runtime static, like node-pty. The slot gates (glibc 2.17 symbol floor, no shared libstdc++, N-API 8) and the smoke load cover it, the template stages it into the compat target, and both orcad and relay rung B pick it up from there. COMPAT_SLOT_ADDONS names a compat slot's own addons: a compat slot missing one fails --require-slots, and a compat template target that would ship any native file without a compat build fails the template build. The CentOS 7 cell expects an activated managed server again, and every launched relay cell loads its watcher directly. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(phase3): close e2e routing holes, localize managed-server outcomes, trim harness and CI (#25206) * fix(phase3): close e2e routing holes, localize managed-server outcomes, trim harness and CI - Route the auto-convert spec into needs_build and route the new orcad/serve e2e helpers to the specs that use them, with routing tests. - Run the auto-convert Docker lane only when routed; fold the missing-AppData check into the Windows mode-switch job and drop crash-hunt diagnostic env. - Delta-move dialog keys its preview on the target id and cannot close or resubmit mid-move; Resume has an in-flight guard. - Managed-server toasts and status lines show localized messages instead of raw codes or main-process English; add singular and zero-count wording. - Settings style fixes (quiet Cancel, section header, labelled fields, progress labels); delete dead i18n keys and the unused previewConversion. - Harness: kill the serve child on readiness timeout, share the isolated profile and spawn-until-ready helpers, hooks and a condition wait in the auto-convert spec, shared cross-version exec helper. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(phase3): localize the refusal toast, share serve readiness and liveness helpers, correct D7 docs - The refusal toast no longer shows main's English blocker detail; the SSH Hosts status line keeps it under Details. A missing terminal count is no longer defaulted to 0. - startOrcadServe uses spawnUntilReady, so a readiness timeout kills orcad; one isPidAlive replaces the spec's copy and the lock-holder loop. - The port-6768 auto-convert test uses the shared hooks. - The docs say the Windows D7 job checks rather than gates. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ci): give the idle-exit spec its app build and route the convert harness to it Co-authored-by: m4air <m4air@Mac.localdomain> --------- Co-authored-by: m4air <m4air@Mac.localdomain> * ci(adhoc): pin every job of an adhoc build to the commit resolved at dispatch (#25311) Each job read the requested branch name and checked out whatever it pointed at when that job started. A push mid-run mixed commits: in run 37234548638 the glibc217 slot lane built2ddea8736e, then #25199 merged, and desktop_template checked out70948d597f, whose merge step requires the compat watcher that lane never built. A first resolve-ref job (no secrets) resolves the branch, tag or full SHA once, and every job checks out that commit. The mac job still vets it for reachability before signing; the requested name stays the concurrency key and the release label. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(terminal): report a confirmed relay PTY exit as exited and give a cold Windows PTY probe more time (#25304) A worktree terminal close returned as soon as the relay confirmed the stop, but the exit frame reaches the runtime record only after the SSH output intake drains, so the verdict read straight after saw a still-connected PTY and answered unverifiable. The stop now waits, bounded by its deadline or 10 s, for that record. The bundled runtime's PTY probe gives the first Windows spawn 15 s and retries once after a timeout only; spawn errors and non-zero exits still fail at once. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(relay,orcad): keep new wire messages forward compatible, drop unshipped relay.reset, fix shutdown and serve gaps (#25298) * fix(orcad-migration): keep migration wire forward compatible with newer peers * docs(pty): record why pty.resumeClient negotiates by method-not-found * fix(relay): remove the unused relay.reset method and keep a deferred shutdown serving * refactor: drop dead hold API, release gate and migration pass-throughs; fix serve and delegation gaps * fix(lint): keep reopen hooks within file size limits * revert type dedupe in orcad-incumbent-recovery to avoid a parallel conflict * fix(lint): name stop-reply request fields for their role --------- Co-authored-by: m4air <m4air@Mac.localdomain> * refactor(ssh): one managed-tunnel ownership check and forward bookkeeping (#25316) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a stuck managed server can recover, cancels finish their change, unknown stays unknown; delete unused reset/drain code (#25227) * refactor(ssh): delete the unused connection reset and drain machinery * fix(ssh): surface a stuck managed server, finish activations past their first change, and fall back from Windows launch refusals - A rejected build that changed profile state now refuses as orcad_recovery_changed_state (unverifiable) with a Recover path that restores the snapshot once the operator accepts; an interrupted activation is no longer reported as a quiet update deferral. - Activation and rollback drop the abort signal after their first journaled change. - Windows launch refusals and unsafe command lines send the host back to the relay. - Fence refusals keep unverifiable blockers unverifiable; the POSIX liveness probe reads kill errors in the C locale and treats permission errors as unknown. - A host with no relay fallback surfaces its real connect error; disconnect and removal clear the setting-up status, and a cancelled decision's progress is dropped. - Remove tcp_forwarding_refused leftovers, dedupe incumbent stop/restore, exec-or-empty, census client, errorMessage and the blocking-blocker predicate; split activation and snapshot files under 300 lines. * fix(ssh): one import per module in the rollback transition --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): committed reads trust the receipt, conversions survive UI churn, journals are cached, and round-2 low items (#25327) * fix(ssh): keep a newer build's fence and note fields, and validate a fence beside a legacy owner Normalization now validates known fields but passes unknown ones through on orcadFence and the managed-server notes. A legacy managed owner next to a malformed fence falls back to the owner's environment id instead of keeping the malformed fence. * fix(migration): delete a retired migration's scrollback files once retirement is durable Retirement dropped the moved terminals' scrollback refs from session state but kept the files forever. After the journal records source-retired, the files the manifest names are deleted, except refs any session partition or pending export still names. A crash before the delete repeats it on the next retirement pass. * fix(orcad): a committed migration reads as committed from its receipt, survives receipt eviction, and abandoned stages expire - Committed-state reads trusted only a byte-identical copy of every moved row and snapshot, so a live server that had been used could never confirm its own commit. The receipt alone now proves it; full equality stays inside the commit. - A receipt that ages out of the 64-entry list keeps a compact record, so the commit never reads as absent. - A stage no client returns to within a week stops holding the server awake and is dropped at the next stage. * fix(migration): freeze a host's session from the fence on, ignore UI churn in the frozen-source check, and cache parsed journals - Renderer writes to a fenced host's ssh: partition are skipped from fence time, not only after commit, so tab work mid-conversion cannot change the source. - The frozen-source check leaves out workspace session and client routing state, which the UI rewrites as the user works (including through the local partition and UI state); the server gets them as journaled. - Parsed journals are cached per directory, keyed by each file's inode, size and mtime and dropped on every write and remove, so list and session calls stop re-parsing every manifest. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): an older relay a census cannot rule out stays held, and a reconnect resumes held PTYs through it (#25331) - A census that does not know the host platform is unverifiable, never "no older relay". On Windows it now probes each older version directory's pipe for the target, the way relay GC does, so a live older Windows relay keeps its PTYs held instead of respawning their panes; those endpoints are held, never bridged. - On reconnect, a PTY the current relay disowned while an older relay holds it is reattached through that relay's own bridge, with its runtime restored and its replay forwarded. When no route serves it, it is left for recovery like an exhausted reattach. - A superseded or disposed deploy no longer starts the census, so it cannot replace the current attempt's. - A route stops holding an unserved PTY that exits. - Drop the unused censusPreviousRelays export. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(phase3): round-2 ui-infra review fixes (#25325) - An unverifiable move refusal with no count says so instead of "0 terminals". - One move per host: the dialog stays mounted while open, joins a run already in flight on remount, and the toast shares the same guard. - A managed server's start, wake or update no longer reloads every host's catalog; only a new environment for the host loads, scoped to that host and the local catalog. - A failed host-partition session write is no longer acknowledged as written, so the writer re-queues those fields. - The deploy picker lists only hosts with no managed server or pending move. - Harness: orcad and the released relay daemon are stopped when startup fails. - Windows host CI: the convert cell always runs last and fails on a failed native switch or build; ssh-host-server unit tests no longer trigger it. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(relay,cli): never report an unreachable terminal as exited; keep PTYs through a deferred shutdown (#25328) * fix(relay,cli): never read an unreachable relay or terminal as exited; keep PTYs on a deferred shutdown * fix: stop an unrecorded breakaway child, route orcad serve through the spawn chokepoint, and drop review-flagged leftovers --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): reclaim a crashed desktop lock, close the stop-cancel race, drop dead daemon and RPC code (#25338) * fix(orcad): reclaim a crashed desktop lock, close the stop-cancel race, drop dead daemon and RPC code - A desktop reclaims a stale desktop lock record (Electron's lock proves it), and a real orcad hold shows a dialog instead of exiting silently. - A failing quit handler no longer skips closing observability. - Completion withdraws its request when a cancel lands mid-write; the listener removes a cancelled leftover. - An idle stop's clean record is retracted when the shutdown fails or overruns. - The managed-stop request tolerates unknown fields from a newer client. - Shared error-code checks, one win32 coverage rule, no redundant isAlive filters. - Remove the recovery-only daemon provider, requirePinnedWsPort/strictPort and the superseded orcad.migration.importCatalog RPC. * fix(startup): keep main-process-preflight under the line cap --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): fail fast on a held fence, restore untouched rollbacks, atomic Windows host script, honest tunnel ensure (#25332) * fix(ssh): fail fast on a held activation fence, restore a rollback the target never touched, stage the Windows host script atomically - withOrcadActivationLock no longer waits up to 15 minutes inside the target lifecycle: a held fence answers at once (orcad_activation_recovery_required); deploy probes it before uploading. - A rejected rollback target that left the restored snapshot untouched puts the newer build back unasked; only a real change waits for the operator. Crash recovery applies the same rule. - The Windows host script is written only when missing, through a partial file and a rename; a bridge exit without a sentinel is unverifiable unless the shell could not find the command. - The managed tunnel's ensure() throws when superseded, and a caller arriving after close() builds a fresh run instead of joining the doomed one. - An unparseable stop answer after a stop was sent keeps the fence (new 'unconfirmed' outcome). - stopRemote starts over instead of reporting live when another run settled the journal. - Releasing an unreachable setup stops the orcad it activated and clears active on proven exit. - A failed startup is torn down quietly so its error stays published; port-forward listeners keep an error handler; a linked SSH access connect is cancelled when its window closes. - Shared errorMessage, one SFTP transfer helper, getConnectGeneration only, test-cell names. * fix(ssh): stage the Windows host script through the pinned node.exe, which cmd.exe and PowerShell both run * fix(ssh): host-script staging runs as host-script ops; a joiner of an overtaken tunnel run builds its own - The presence check and the install are ops in the uploaded host script (script-present, and script-install run from the partial upload itself), invoked as node.exe <script> <op> <args> like every other Windows host op: no inline code. - ensure() callers that joined an in-flight run no longer inherit its 'superseded' end; they start a fresh run, which also covers a close() in between. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a connect re-checks unverifiable relay terminals once its relay session can answer (#25385) The server decision runs before any relay session exists, so a terminal this desktop left detached could never be asked about and read unverifiable for as long as it ran. On Windows no endpoint census can fill that gap. Once the session is up, the relay that holds the PTY answers and a terminal it still runs is reported live. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(settings): a managed server whose status never loaded reads unknown, not "Not running" (#25389) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(relay,runtime-env,cli): keep a deferred relay's AI Vault and skill uploads, make managed re-pair crash-safe, show unverifiable in worktree ps text (#25393) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): keep a surviving daemon's slot through the cache prune, and drop an idle record a signal stop took over (#25387) - The slot prune also protects any slot a live terminal daemon's PID record points into, and evicts nothing while a daemon record is unreadable. - The shutdown trigger reports whether it took ownership; an idle stop that another source took over discards its clean-idle record. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): edit a managed host's connection, show rows no migration owns, drop the unused quit-drain predicate (#25403) - SSH settings may edit a managed host's connection fields; its fence and generation are never written from the renderer, and its tunnel is closed so the next use redials. Removing it is refused with a pointer to Stop under Managed servers, which stops and removes the server; the renderer no longer ends the host's terminals before that refusal. - A fenced host with no journal (an empty host's deploy, or one whose journal compacted after retirement) no longer hides its rows: no committed move owns them, so projects an older build added stay visible. An unreadable journal still hides them. - The quit drain's mayDetach predicate and disconnectAll's filter had no production caller and are removed. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): input to a terminal an older relay holds but no route serves is refused, not dropped (#25404) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): refresh a changed host's status once its delta move or keep-server choice lands (#25416) The status line kept saying the host was changed on an older Orca until a manual reconnect. Main now publishes the host as managed by its server after a successful move or keep, with a disconnected state when the move released the relay session. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a stopped managed server stops showing on its SSH host without a reconnect (#25409) * fix(ssh): a stopped managed server stops showing on its SSH host without a reconnect On unlink, forget the host's managed-server decision and republish its connection state. The republish also drops the host's cached worktree scans, so worktree.ps stops naming the removed server. * fix(runtime-environments): a removed server's workspace session partition goes with it Host-scoped listings enumerate session partitions as known hosts, so worktree ps kept naming a stopped or removed server as an omitted, unselectable host. Unlinking or removing a server now drops its runtime:<id> partition. * test: give the removal-storage fake store casts their SAFETY rationale * fix(runtime-environments): drop orphaned runtime workspace sessions at startup A crash between unlinking a server and dropping its session, or a build that unlinked before the drop existed, left a runtime:<id> session listings name as an unselectable host. Startup now drops runtime sessions whose server is not in the environment store; it never touches local or ssh sessions, and skips entirely when that store is missing or unreadable. * test(migration): retirement re-aims focus without creating the destination's session Locks in the ordering the startup reconcile relies on: a runtime:<id> session is never written before that server is registered. * docs(runtime-environments): name the ordering invariant the startup reconcile relies on --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a connect whose relays prove its leases ended retires them, so the next connect converts (#25405) Every connect decides before a relay session exists, so a detached or expired lease reads unverifiable there. The post-session re-check proved such leases ended but left them in place, so the host stayed unverifiable on every connect and each respawned pane added another lease. Co-authored-by: m4air <m4air@Mac.localdomain> * ci(ssh-windows): check out preload and renderer for the orcad-convert e2e build Main's #25359 narrowed this workflow's checkout to the server and test trees. Phase 3's orcad-convert cell builds the full e2e app with electron-vite, which also needs src/preload and src/renderer, so the x64 inbox cell failed with UNRESOLVED_ENTRY. * test(orcad): model a really converted host in the v1.4.218 downgrade wire test #25403 shows a fenced host's rows when no migration journal explains the fence. The downgrade test fenced the host without a journal, so this build showed the retained project it is meant to hide. Write the destination-committed cutover journal a real conversion leaves. * fix(ssh): a briefly held fence waits and is retried, not recorded as a failed update; terminate keeps held leases (#25420) - The activation fence is retried for a few seconds. A fence still held answers orcad_activation_fence_busy (a waiting deferral, never an update failure) unless a journal or a lock past its stale age shows an interrupted run, which stays recovery-required. - acquireInstallLock reports Busy only when a holder answered; a lock command that only ever failed surfaces its own error. - Terminate detaches instead of disposing when any PTY was unverifiable, so the final teardown never bulk-marks a lease an older relay may hold as terminated. - A wake whose connection dropped while holding the fence releases that fence on this client's next wake (no journal, slot proven exited), so a relaunch-then-connect is not left fenced. - The orcad e2e reconnect helper surfaces a connect's error text instead of a JSON parse error. Co-authored-by: m4air <m4air@Mac.localdomain> * ci(ssh-windows): check out all of src for the orcad-convert cell Its e2e global setup also compiles the bundled CLI from src/cli, which the narrowed checkout left out. * fix(orcad): idle stop — drop the record when a signal stop wins, and read the activation lock, not its root (#25464) * fix(orcad): drop the idle-stop record when a signal stop finishes first Every stop's cleanup now discards the record unless the idle trigger owns the stop, so a takeover that exits before the idle request runs no longer leaves a false idle stop. * fix(orcad): idle check reads the activation lock, not the transaction root An interrupted acquire can leave the root without a lock; the client already treats that as unfenced, and managed orcad now does too instead of never idling out. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): main refuses a managed server host's removal before ending its terminals (#25465) The remove flow skipped ending terminals for a managed host only when the renderer's cached target list already showed the fence, so a fence that landed after the list loaded still lost the host's terminals before main refused the removal. The removal's terminate call now carries forRemoval, and main refuses it for a managed host with the Stop… message before touching anything; the renderer no longer decides. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(terminal): closing an SSH workspace counts a host-confirmed terminal exit as stopped (#25479) * fix(terminal): a workspace close counts an SSH terminal's confirmed exit as stopped The close's verdict treated any SSH terminal whose record was still present as unconfirmed, even when that record held a host-confirmed exit. SSH records outlive their exit, so every successful close of a relay terminal answered unverifiable. Also, a relay reattach that finishes registering after the stop no longer revives an incarnation whose exit is already recorded. * test(terminal): cover a reattached relay exit confirmed without an incarnation --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(cli,ssh): managed-server actions outwait their own deadlines; cap the unverifiable serving detail (#25473) orca environment update/rollback/recover/stop/cancel-stop waited the 60 s RPC default while the desktop runs the whole action inline, so a slow host printed a timeout failure for an action that kept going. They now wait 20 minutes, and a timeout says the action may still be running and points at `orca environment status` instead of reporting failure. status keeps the default. The retained managed-server state now clamps serving.detail to the same byte limit as its sibling detail fields. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): keep the previous orcad.log on Windows across a restart (#25481) The Windows breakaway launcher's addon recreates orcad.log on every start, while POSIX appends, so a crash's log was gone once orcad restarted. The orcad launch now asks the launcher to move the last run's log to orcad.log.1 first, capped to its last 1 MiB. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(sidebar): a stopped managed server leaves no empty project group behind (#25488) The removed-runtime purge dropped the server's repos, setups and worktree rows but kept the project groups and folder workspaces the renderer fetched from it, so an empty heading stayed in the sidebar. It now drops those runtime-stamped rows too. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): run automations on orcad and keep a managed host up while they can fire (#25475) * fix(orcad): run automations on orcad and keep a managed host up while they can fire orcad never built an AutomationService, so with orca serve defaulting to orcad scheduled runs never dispatched and Run now threw runtime_unavailable. The headless service setup moves out of Electron startup into automations/runtime-automation-service.ts; orcad installs, binds, starts and stops it, and managed idle exit counts an enabled schedule or an unsettled run as busy. * docs(orcad): list automations among the idle-exit conditions * fix(automations): precheck reads the SSH manager from its registry, keeping electron out of orcad The precheck imported getSshConnectionManager through ipc/ssh, which pulled 32 electron modules and node:sqlite into the orcad bundle and failed build:orcad. * test(orcad): stub the automation wiring in the push-startup runtime harness That harness stubs OrcaRuntimeService without an automation surface, so starting the real service threw setAutomationService is not a function and timed out the next test. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): Move to managed server survives a relay that hangs up on its last terminal exit (#25487) BUG-15: with two live terminals, one served through an older build's relay, Move failed with "Failed to terminate SSH host sessions: …: Multiplexer disposed" although both shells died. The old relay reports the exit before the shutdown reply; that exit closes the route (its last served PTY), and the disposed mux rejected the shutdown still awaiting its reply. - The legacy relay route settles a shutdown whose PTY exit it already observed. - Move no longer aborts on a failed stop: the terminal census (what the conversion trusts) decides. Exited closes the relay session and converts; live or unverifiable refuses and republishes the relay status, so the stale terminal count is replaced. - The move dialog offers Try again after a refusal or failure. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): closing a terminal a previous Orca relay runs stops it there and confirms the exit (#25471) A stop on a PTY an older relay runs reached the current relay whenever no route served it at that moment (after a reconnect with the pane unmounted, or a shell no pane ever resumed), and the current relay answers a stop for an id it never minted as done, so the PTY read stopped while its shell kept running. A stop on a served PTY failed instead: its exit arrived before the stop's reply, closing the route under the pending request. - A served PTY stops on its route; a PTY no route serves is stopped through a short-lived route to the older relay that lists it, which hangs up once that PTY exits. - When an older relay may hold the PTY but cannot be asked (incomplete census, Windows pipe, a bridge that will not open), or its bridge drops mid-stop, the stop is unverifiable, never reported done; terminate keeps the lease. - A route stays open until its in-flight requests settle. - The provider's exit stream includes the exits older relays report, so a stop observes the PTY it stopped exit on the relay that ran it. The cross-version harness runs what terminal close --all runs per PTY against a real v1.4.218 relay, for a pane resumed this connection and for a shell no pane resumed, and sees the old shell exit there and the old relay retire on its own grace. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a live run's journal reads busy, a wake releases only its own fence, a cancelled re-check never reports connected (#25470) * fix(ssh): a live run's journal reads busy, a wake releases only its own fence, a cancelled re-check never reports connected - A held fence asks for Recover only when its lock is stale or a journal has no fence over it; a journal under a fresh fence is a run still working and reads orcad_activation_fence_busy. - A wake writes an owner token into the fence it takes and later releases only a fence carrying that token; observing the fence gone forgets it. - The connect re-checks ownership after the relay-terminal re-check, and the re-check itself neither records a decision nor retires leases for a cancelled attempt. * fix(ssh): a tunnel caller that joined a run a disconnect cancelled builds its own The launch-time restore's tunnel run connects over SSH; a disconnect then cancels that connect. An explicit connect that had joined the run inherited its SshConnectAttemptCancelledError and failed (seen as the idle-exit e2e's 'connect threw: ... cancelled'). A joiner now builds once anew after any end of the joined run except an auth failure, which it shares rather than prompt again. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(runtime-env): a late reply from a replaced pairing never overwrites the re-paired device identity (#25490) * fix(runtime-env): a reply from a replaced pairing never overwrites the re-paired device identity * fix(types): narrow identity fields in markEnvironmentUsed * fix(ssh): managed tunnel proves its server by runtime id; SSH access linking keeps the strict device check --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): retained scrollback survives a closed tab, local acknowledgements stop blocking, close-intent retirement retries, mobile selections merge per workspace (#25508) - A transfer reads a retained snapshot straight from storage once its tab closes, and releasing the retention deletes a ref no session names; the frozen-source check leaves the snapshot list to the journaled manifest. - Only acknowledgements on panes the source host owns count toward the ui-routing blocker. - Retiring close intents treats an absent source with an identical destination entry as already done. - Importing a device's mobile selections keeps its selections for other workspaces. Co-authored-by: m4air <m4air@Mac.localdomain> * test(ssh): an expired lease an older relay still lists is never retired (#25449) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): keep each host's state its own, count only saved commits, and never wedge connect on a partial session (#25516) * fix(migration): a host-qualified owner key belongs only to its own host Owner matching stripped a key's host qualifier before matching its repo id, so converting host A claimed, moved and on retirement removed host B's session state when the two hosts share a repo id, and destination-qualified focus written by retarget read as leftover source state, failing retirement with orcad_migration_source_ui_routing_reappeared. A qualified key now matches only its own host, and an unqualified key in another host's session partition belongs to that host. * fix(migration): a partially written session partition never wedges connect, and a marker two partitions agree on stops blocking the move Real-host BUG-14: the renderer's per-host snapshot leaves out maps a host has no rows in, and main stored host partitions exactly as sent, so a runtime partition lacked tabsByWorktree and the dormant-state collector threw on every connect. Main now fills the required maps on every host-partition write, the migration collectors tolerate a partition persisted without them, and an unreadable session blocks the move instead of failing connect. The same profile's workspace-session blocker was a false positive: the local and host partitions both carried defaultTerminalTabsApplied for the moved worktree with the same value, and the fragment merge refused any shared worktree key. It now refuses only when the partitions disagree. * fix(migration): only a flushed commit acknowledgement moves a migration to committed A committed state read may come from a receipt the server holds in memory but failed to flush. A retry took that read as proof, journaled destination-committed and went on to retire the source, so a later server restart could lose the catalog on both sides. A committed read in stage and in abort is now confirmed through the idempotent commit(), which flushes before it answers; a failure leaves the journal and the fence where they were. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): relay shells no lease here knows count as live, Move stops them, and Windows asks every relay pipe (#25518) * fix(ssh): a relay shell no lease here knows counts as live, and terminate stops it A CLI-created terminal has no lease, so with an attached lease the gate answered from leases alone, a live decision was never re-counted once the relay could answer, and terminate stopped only the shells it held leases or panes for. The relay's own listing is the authority on what runs. * fix(ssh): a Windows connect asks every relay version's pipe for its PTYs before converting Windows pipes cannot be listed, so the connect-time census answered 'unenumerable' and a shell no lease here knew let the host convert under it. Each version directory's pipe for this target is derived from its path, so the census probes them all, current included, and asks a live one through its own bridge; a live pipe it cannot ask is unverifiable. * fix(ssh): the terminal gate counts what earlier relays still run, leased or not After an app update a shell a respawn superseded on its tab keeps running on the previous relay with no live lease here, so a decision counted only the leased shells and Move could not see it. * fix(ssh): a Windows relay folder with a live pipe it cannot ask stays unverifiable Each pipe is probed and asked on its own, so one that answered with no PTYs can no longer stand in for a live sibling the census could not reach. * fix(ssh): terminate also stops shells only an earlier relay lists provider.shutdown routes a held id to the older relay that runs it, so the terminate set now takes listPreviousRelayPtyIds too; one it cannot reach is reported unverifiable as before. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a refused Move to managed server reconnects the host on its relay (#25543) Stopping the terminals closes the relay session, so a move the census refused (another desktop's terminals, or an older relay it can't rule out) left the host and its workspaces disconnected until a manual Connect. The refusal now reconnects the host; the connect-time decision reads the same census and keeps the relay. A reconnect that converts after all reports the move; a reconnect that fails still reports the refusal. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(relay): an agent exec ends on its child's exit, not on pipes a background process still holds (#25544) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): stop the automation scheduler first when a managed stop is dispatched (#25548) A dispatch could otherwise race the daemon retirement census or write a run record that a rollback restore then silently discards. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): count a degraded daemon's in-process terminals in the terminal census (#25545) In degraded mode fresh terminals run on the local fallback inside orcad, but the census read only daemon adapters, so an update or stop saw 0 live sessions and killed running agents. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a retained source whose drafts or settings an older build changed is never hidden as unchanged or retired (#25549) * fix(orcad): a retained source whose drafts or settings an older build changed is never hidden as unchanged or retired The retained-source fingerprint covered only repo, folder and group identity, so an unsaved draft edited on an older build read unchanged: the host kept serving the server's older draft and retirement deleted the newer one. Retention now also records a versioned fingerprint of the source's drafts and user-authored names and settings; a mismatch marks the host changed, and a journal without one is never retired automatically. * fix(orcad): a retained source's automations are part of its state fingerprint An older build can edit an automation the source keeps; retirement would delete that edit as if the server held it. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a rollback restores the older snapshot only behind orcad's terminal barrier (#25551) The rollback's census is taken while orcad still admits work, so a terminal or automation that starts before the stop had its state wiped while its PTY survived. The incumbent is now stopped through its managed stop with idle-daemon retirement: orcad closes terminal admission on every daemon generation, counts live sessions under that fence, and retires the daemon only when none exist. Only 'retired' lets the older snapshot replace state; live, unverifiable or a missing answer refuses and relaunches the incumbent on its untouched state. A build without managed stop is refused before anything changes. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): failed managed setup leaves 'connecting', edited managed host redials, idled-out server starts before its census (#25474) * fix(ssh): a failed managed setup leaves 'connecting', an edited managed host redials, an idled-out server starts before its census - doConnect publishes the error and clears the 'setting up' status when the managed-server decision throws for a still-current attempt; a cancelled one still reports cancellation. - Editing a managed host's connection fields closes its tunnel and disconnects its transport, serialized with the target's lifecycle, so the next use dials the edited target. - The terminal census starts a server that idled out behind a forward still up, so Stop, Update, Rollback and status no longer refuse with 'census unavailable' on every retry. * fix(ssh): a fenced failed setup publishes its cause, never-launched slots are collectable, a reused PID is not orcad - doConnect publishes the relay decision's setup failure (and clears 'setting up') when a failed managed setup kept the host fenced, instead of throwing a bare 'serves a managed server'. - The liveness probe answers NEVER_LAUNCHED for a slot with no process record and no readiness file; GC removes such a slot, and every other reader still reads it as UNKNOWN. - On POSIX a PID whose command line does not run the slot's orcad.js reads DEAD, so a stopped orcad behind a reused PID is woken instead of reported serving or unverifiable. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a legacy relay route a new pane started serving stays open when a pending stop settles (#25542) A served PTY's exit that lands before its stop's reply defers the route's hang-up until the stop settles. A second pane the same older relay holds could start serving through the route in that window, and the deferred hang-up then closed it anyway, sending that pane's input and stops to the current relay. Serving a pane now cancels the deferred hang-up, and a settling request hangs up only a route that serves nothing. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): the journal holds a migration's scrollback until commit or abort, across failed uploads and restarts (#25550) Retention was scoped to one transfer call, and its release in finally deleted a closed tab's snapshot after an interrupted upload, so every retry failed with source_snapshot_changed. Inline buffers had no file for the retained read at all. Retention now follows the cutover journal: held from the journaled export through staging, rebuilt at startup, released on commit or a removed journal. Inline bytes are written to their ref while held. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * fix(migration): a repo id two hosts share never lets legacy keys cross hosts, and a dangling identity alias stops blocking the move (#25558) - Repo ids are not unique across hosts: the same id may be registered on two SSH hosts. Stores that are not session partitions (worktree metadata, automations, lineage, client state, sparse presets, retired names) can hold legacy keys with no host qualifier, so an id both hosts register said nothing about whose a key was. The scope now records such shared ids; an unqualified key or bare repo id for one only matches with its row's own host evidence (worktree metadata's hostId, an automation's ssh target), so another host's rows are never moved, counted or retired. An automation's target generation now matches only alongside its target id. - An identity alias whose identities hold no metadata (worktreeMetaByIdentity lost them, as on the B4 profile) is nothing to move rather than a worktree-metadata blocker; retirement drops it. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(cli): orca environment recover --accept-changed-state --yes restores over changed state (#25597) Recover refused when a rejected build changed profile state, and its refusal told the user to run Recover, which the CLI could not do. --accept-changed-state (confirmed with --yes) maps to the same acceptChangedState the Managed servers settings pass, and the refusal now names the flags. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): refuse an update that would end a degraded host's in-process terminals (#25598) The census now reports inProcessSessions separately. Those terminals run inside orcad and end with any restart, so planOrcadUpdate defers with a non-forceable orcad_update_ends_in_process_terminals instead of claiming they survive on the daemon. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): a delta move protects what it imported from rolling back across it (#25599) A delta move never advanced the server's migration mark, so rolling back an update taken before the delta was admitted and dropped the delta's projects. The delta now records the mark before any commit can land, resumed commits included; the mark keeps the latest migration and never moves back; and the rollback gate also counts every journal into the server, so deltas finished before this change stay protected. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a missing relay inventory never proves terminals exited, and the Windows census covers every desktop's relays (#25611) - The migration terminal gate returned `exited` whenever this desktop held no unresolved lease, even when the current relay or the earlier relays could not be asked. A failed or incomplete inventory is now `unverifiable` regardless of local leases. With no relay session at all the gate asks for a host census (`needsHostCensus`) instead of reading the silence as exit; the connect, conversion and delta move pass that census in, and the connect hands its own census result to the conversion it starts. A census that cannot list endpoints (`unenumerable`) is unverifiable too. - The Windows connect-time census derived pipe names from this desktop's target id only, so another desktop's relay on the same account was never probed. It now lists every `orca-relay-*` pipe on the machine and maps each to the relay instance that owns it through that instance's credential file or active-pipe marker, asking each with its own credential. A pipe no version directory accounts for is unverifiable unless the host proves it another account's (or gone), and an inventory that could not be read is unverifiable. Co-authored-by: m4air <m4air@Mac.localdomain> * refactor: drop the unshipped pty.resumeClient relay method and unused SSH provider unregister guards (#25595) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a connect still deciding its server holds the raw 'connected', closes a transport its cancelled decision opened, and an edit keeps a relay host's session (#25641) - handleSshConnectionStateChange holds a raw 'connected' while a connect is in flight even before any relay session exists (published as 'connecting'), so the census, deploy or conversion that dials the pool no longer reports the host up with no providers. - priorConnection is captured before the server decision; a connect cancelled after the decision closes a transport the decision opened, unless a newer connect is using it. - Editing a fenced host an older build changed (it runs on the relay directly) no longer disconnects its transport; only a host reached through its managed server redials. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a retained source is unchanged only against a pre-commit baseline of everything a user wrote (#25602) The state fingerprint read drafts, automations and workspace metadata from what a move could carry, so a session a move refuses hid an older build's draft edit, and retention hashed the source after the commit, so a crash before retention blessed whatever an older build changed in between. The baseline is now written with the fence, before any commit is possible, from the source read directly; a session that cannot be read, or a journal without that baseline, is unverified. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a delta move refuses a source that changed while it checked terminals (#25691) The plan and manifest are taken before the terminal check and session release are awaited, but the journal took its baseline after them, so a draft typed in between became the baseline while the server received the older one, and retirement deleted the newer draft. The baseline now comes from the plan's own snapshot, and a source that changed since it refuses the move. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): a save landing mid-retirement no longer defers retirement (#25692) Retirement removed the source rows, then awaited the profile flush, then checked nothing came back. A session save that landed during that flush re-added a source-owned row, the check failed, and the journal stayed committed until a later connect. Retirement is idempotent, so it now runs one more pass before deferring; a row back after that is reported with the partition and owner key it reappeared under. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(automations): a headless run never reads completed when its agent never ran (#25700) * fix(automations): a headless run never reads completed when its agent never ran orcad (and Electron serve) finished a dispatched run on a satisfied tui-idle wait, and a ready shell prompt satisfies it: a run whose agent is not installed read 'completed' within seconds. Like the desktop runner, completion now needs the agent's own status for the run's pane after dispatch; without it the run fails after the agent-start window with the reason, instead of claiming completion. * fix(automations): keep idle-means-done for agents without status; fail only a refused command Not every automation agent reports status on orcad (no hooks on the host, no recognised title), so requiring it would fail their runs. A run completes on the agent's own status, fails when the shell refused the agent's command (bash, zsh, dash, fish, PowerShell, cmd), and otherwise keeps the old idle-means-done rule after the agent-start window. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): the retained-source fingerprint sees all worktree metadata retirement owns (#25694) * fix(migration): the retained-source fingerprint sees all worktree metadata retirement owns The state view read only legacy worktreeMeta keys and attributed them without meta.hostId, so identity-backed metadata and unqualified rows a shared repository id leaves to hostId were missing from the fingerprint: an older build's edit read as unchanged and retirement deleted it. The view now uses the same attribution as export and retirement, and metadata that claims the source host but cannot be attributed leaves the source unverified. The fingerprint version moves to v2. * fix(migration): the retained-source fingerprint skips automations on a repo id another host shares The state view matched automations by scope.repoIds, which ignores sharedRepoIds, so host A's fingerprint included host B's automation on a shared repository id; editing it marked A changed and routed it back to the relay. The view now uses the move's automationTouchesScope. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a failed readiness read no longer kills a healthy candidate; log unsettled activations (BUG-17a) (#25701) * fix(orcad): a failed readiness read no longer fails a candidate's launch, and log every unsettled activation BUG-17: the candidate went ready and the client SIGTERMed it ~1 s later through its reject path, yet that readiness passes the gate, so the launch itself failed: one readiness-wait exec that errored failed the launch outright. Retry such reads until the readiness deadline; an unconfirmed termination still fails at once. The update's outcome never reached the app log, so every update or rollback that does not go through now logs its code and reason. * fix(orcad): the host-side readiness wait ends on a wall-clock deadline On a loaded host each poll's reads outlasted its sleep, so the step-counted loop ran past the client's 30 s exec timeout. That timeout failed the launch, the reject path SIGTERMed a candidate still starting, and orcad, which defers a stop until startup completes, published readiness and then exited (BUG-17). --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): stop a converted host's stale tabs landing in local, and retry retirement only on an exact replay (#25712) * fix(migration): retirement re-removes only an exact replay of moved session state #25692's second pass re-ran retirement on any row that reappeared during the flush, which could delete a tab or draft written after the move. Retirement now records the session rows it removes before it runs; a row that reappears is removed again only when it is byte-identical to one of those (in any partition). A new or changed row defers retirement and stays. * fix(ssh): a converted host's leftover session rows stay in its own partition, never local After conversion the renderer drops the SSH host's projects and worktrees but keeps their session rows. With no catalog owner left, the next save routed those rows to the local partition, where retirement read them as moved source state reappearing and deferred. Converting now pins each dropped worktree's session key to the host's partition; main's fence guard keeps that partition frozen, so the stale rows are never written. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(test): restore the codexProviderHandle import main's #25078 dropped again #25722 restored it, then #25078 removed it, so pnpm tc fails on main's tip. * chore(sync): keep main's own cloud and mobile files byte-identical to main Earlier syncs added lint-only brace and template fixes to these main-owned files; reverting them keeps #24863's diff against main free of files Phase 3 does not own. * fix(terminal): a remote pane's reattach error no longer shows this client's OS and shell (#25693) * fix(terminal): a remote pane's reattach error no longer shows this client's OS and shell The error toast appended the client's environment for every non-SSH error. It now shows it only when the pane's known execution host is this client; an SSH, managed or not-yet-known host omits it. Co-Authored-By: Claude <noreply@anthropic.com> * test(terminal): type the pane-host fixture as runtime owner state Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * fix(orcad): a retained activation fence is ownerless, so recovery can always take it over (BUG-17) (#25698) * fix(orcad): a retained activation fence is ownerless, so recovery can always take it over BUG-17: a recovery that took a stale fence over, failed and retained it left a fresh lock, so every later recovery read it as still fresh and the host could never be recovered. A run that keeps the fence once it is done now backdates the lock; one whose remote command may still be running keeps it fresh. Also run the in-process terminal deferral before the forced protocol check, so a degraded host whose daemon is empty names its in-process terminals instead of an unreported protocol. * test(orcad): the CLI's accepting recover takes over a fence a refused recover just retained The fake host now answers a stale-only takeover busy while a recovery's own takeover is fresh, which reproduces BUG-17's 'still fresh' loop without the ownerless mark. * test(orcad): keep the fake host's fence acquisition void where callers expect it --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a cancelled connect closes only the transport its own decision opened and nothing newer adopted (#25696) * fix(ssh): a cancelled connect closes only the transport its own decision opened and nothing newer adopted #25641's cleanup took any transport that differed from the pre-decision one as the cancelled decision's, guarded only by connectInFlight. A replacement connect that completed (and left connectInFlight) then had its live transport disconnected by the stale attempt. The pool now attributes a transport it opens inside a connect's server decision to that attempt (AsyncLocalStorage), and the latest user adopts it: a connect that connects or publishes a managed route, or a managed tunnel that records a forward. A cancelled attempt closes the transport only when it opened it, nothing newer adopted it, and no current replacement is in flight. * fix(ssh): a still-current connect whose server decision fails closes the transport that decision dialed The decision's own failure (a throw, or a fenced relay refusal) published 'error' but left the transport it dialed open, so getPublicSshState read 'connected' and a later auto-reconnect broadcast a plain 'connected' with no relay. Both branches now close exactly the decision-owned transport through abandonDecisionTransport, which treats the attempt's own in-flight entry as no newer owner while that attempt is still current. * test(ssh): name the stand-in transport type --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): Windows readiness identifies itself by one PID when the process-table snapshot times out (#25723) * fix(orcad): Windows readiness identifies itself by one PID when the process-table snapshot times out A loaded Windows runner timed the whole-table snapshot out during the bundled runtime's readiness preflight, so the candidate failed to start and activation rejected it. * fix(orcad): fall back to the one-PID query only for a slow process table, never an unreadable one An unreadable table (EDR-hooked snapshot, restricted token) must still fail qualification. The table now rejects slowness with a typed WindowsProcessTableTimeoutError, and only that falls back. Review by win-serve. * build(cli): list the process-table timeout error in the CLI project's file list --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a connected conversion proves the host idle with the account-wide census, not this target's lists (#25697) * fix(ssh): a connected conversion proves the host idle with the account-wide census, not this target's lists The migration terminal gate asked the host-wide census only when no relay session existed. With a session, it trusted this target's relay listing and its earlier-relay census, both of which name only this target's instances, and returned `exited` when they were empty, so another desktop's live shell on the same account, under a different target id, let the host convert under it. `exited` now always needs a complete host-wide census: this target's lists can prove `live`, but empty lists only pass the question to the census, and a gate given none answers `unverifiable` with `needsHostCensus`. The connect-time refinement passes the census too, so a connect retires its leases only when no relay on the account holds work. The Windows host lane now runs a second desktop's relay with a live shell and expects the connected gate to read the host live. * test(ssh): a connected relay's empty lists still ask the account-wide census * refactor(ssh): drop the gate's unread needsHostCensus flag; its unverifiable reason says why * test(ssh): the delta snapshot fixtures give their account-wide census --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a reconnected transport starts managed orcad itself instead of inheriting a dropped start (#25689) * fix(ssh): a reconnected transport runs its own orcad start instead of inheriting a dropped one The serving check deduplicated in-flight starts by environment only. When a connect dropped mid-start (as when the launch-time auto-connect is replaced by a reconnect), the caller on the new transport joined the start bound to the dead one, got its failure, and reported the host managed with no server running. In-flight checks now join only on the same connection, connect generation and port, and a wake's own fence token is cleared only by that wake. * fix(ssh): a reconnected wake releases the fence its dropped wake held at any point The flake's real verdict was 'fenced': the dropped launch-time wake held the activation fence, and the reconnected wake could not prove it its own. A wake now claims its token before its first remote step, an absent owner record under a held token is still its own, and a reconnected wake waits for this client's dropped wake to settle before reading the fence. * fix(ssh): release only a fence carrying this process's own wake token A fence with no owner file could be another client's fresh one. Releasing it now requires the owner token this process wrote; the token is claimed before the write so a drop after it still proves ownership. * fix(ssh): type the wake's fenced fallback --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(ssh): stop the exec-stdin test double from failing on EPIPE (#25739) The truncation test's fake exec channel forwarded the local shell's EPIPE (or 'Cannot call end after a stream was destroyed') as a channel error. Whether that error or the shell's exit code won depended on scheduling, so the test failed under full-suite load. ssh2 silently drops writes once the remote stops reading; the double now does the same, and a 1 MB payload makes the early-stop path deterministic. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): Move lets the reconnect's decision run its census, and a census outside a connect holds 'connected' and closes its own transport (#25735) - Move to managed server no longer runs a separate host census after tearing the relay down, which dialed the pool with no connect in flight and broadcast a raw 'connected' with no session or providers. It reconnects, and reads the decision the reconnect's census recorded. A relay a failed stop left up is detached (leases kept), not disposed, before the reconnect. - The CLI and delta-move census (censusHostRelayTerminalsFor) runs outside a connect under its own owner: the raw 'connected' it causes is held, and a transport it opened that nothing adopted is closed afterwards. Reusing a pooled transport inside a scope now adopts it. - Drop the now-unused publishRelayTerminalsStatus. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(test): drop the restored codexProviderHandle import now that main restored it * refactor(migration): keep a converted host's source rows instead of retiring them automatically (#25768) Automatic source retirement leaves Phase 3: nothing deletes a converted host's retained rows on connect, delta move, keep-server's-version or restart. They stay hidden and are removed only by stopping the server, removing the host or uninstalling. Change detection goes back to the catalog-identity fingerprint, so an older build's edits inside an already-moved project stay preserved in the retained rows without marking the host changed. The copy-only helpers the delta view uses are renamed to subtract, and the converted-host session pin now also overrides a boot primary of local. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(automations): a headless run is watched past tui-idle timeouts by the one run observer (#25733) * fix(automations): a headless run is watched past tui-idle timeouts by the one run observer The headless dispatcher awaited a single tui-idle wait, which rejects after its 5-minute default, so a healthy agent working longer was published as dispatch_failed and never observed again. The dispatcher now hands the run to its completion watcher, whose runtime observer already re-arms wait timeouts, honours cancellation and bounds total observation; the agent status and missing-command checks fold into that observer, and the separate completion loop is gone. * fix(automations): an already-idle pane completes when the start window passes Real-host: a stub that exited before the window left an idle shell with no agent status, and the observer re-armed a tui-idle wait that never resolves for an already-idle shell, so it timed out instead of completing. The observer now keeps judging the pane while its output is unchanged, and only waits again once the pane changes. * fix(automations): resolve a watched headless run by its launch handle first --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(runtime-env): a re-paired managed server's subscribers recover without a reload (#25752) * fix(runtime-env): a re-paired managed server's subscribers recover without a reload The renderer kept the pairing revision it last read, so after an on-connect update re-paired a managed server every subscribe and request was refused as 'pairing changed' until a reload. The first refusal now re-reads the environment catalog, so revision-keyed subscriptions resubscribe and requests carry the new pairing. A managed server re-pairing for the same host registration is the same peer, so its workspaces and tabs are no longer purged as a replaced environment. Co-Authored-By: Claude <noreply@anthropic.com> * fix(runtime-env): a re-paired managed server is the same machine only when its host proves the same identity Same SSH target registration is not proof: a reinstalled host or a target now pointing elsewhere keeps it. The runtime id the pairing handshake verifies must be known and unchanged; otherwise the re-pair retires the environment as before. A proven runtime id change also counts as replaced. Co-Authored-By: Claude <noreply@anthropic.com> * fix(runtime-env): decide a managed re-pair's same machine by the host's proven key, not its runtime id The runtime id is minted per process start, so every orcad restart would read as a new host. The host's E2EE public key persists in its own profile across updates and its pairing handshake proves it; main now lists a digest of it, and the renderer keeps a re-paired managed server only when that digest is known and unchanged under the same SSH target registration. Co-Authored-By: Claude <noreply@anthropic.com> * fix(runtime-env): defer a managed re-pair's same-machine decision until the host key is known A re-read that lands before the new pairing's host key is listed no longer purges: the decision waits for a catalog that carries the key and retires only if it differs. Adds the update-flow store test: same registration and key keeps workspaces and tabs, including a re-read that runs before the key is known. Co-Authored-By: Claude <noreply@anthropic.com> * fix(runtime-env): watch for a pairing refusal on a side branch so requests settle on the same tick Chaining .catch onto every subscribe and request delayed each success by a microtask, which let a StrictMode cleanup run before a client-event subscription resolved, so its unsubscribe landed late. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * refactor: consolidate pane-ownership, migration-catalog, activation-launch and parser helpers (#25738) * refactor(orcad): one launch-and-judge helper for activation and rollback * refactor(orcad-migration): one copy each of the destination projections, selectNewRows, assertSameValue, compareKeys and slotLiveness * refactor(orcad-migration): one string-list validator and one uniqueness check, error codes passed in * refactor: one shared pane-ownership and terminal-layout module for migration, profile transfer and split layout * fix(orcad-migration): row-identity helpers in a leaf module (no import cycle); key order in its own module --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): reopen a managed tunnel to a host another desktop restarted (#25800) The tunnel's identity check pinned the saved runtime id, which orcad mints per process. A host updated or woken by another desktop, or restarted while this one was away, failed every reconnect with orcad_identity_mismatch, and nothing could refresh the id because that needs the tunnel. The E2EE handshake with the pinned host key and our accepted token now prove the server; the first authenticated status reply records the new id. A different host is still refused. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * fix(orcad): create the readiness file owner-only so its pairing token is not world-readable (#25809) orcadLaunchCommand truncated .orcad-readiness before setting umask 077, so under a login umask of 022 the file that receives the pairing offer (with a runtime-scope device token) came out 0644. umask 077 now runs first, the readiness file is chmod 600 after the truncate (a redirect keeps an earlier build's 0644), the pid and log files are tightened too, and the slot dir and ~/.orca-remote are chmod 700 so files earlier builds left readable are no longer reachable. The state snapshot capture also sets its umask before creating the snapshot directory. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(terminal): explain a pane whose saved session another host connection owns (#25814) terminal_pane_owner_host_mismatch reached the user raw, with an issue link. It now reads as a plain explanation with the open-a-new-terminal action, like the reattach failure. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * test(topology): allow phase3's headless editor-tab retirement in main's boundary ratchet (#25823) Main's #25329 added the ratchet; phase3's mobile-session-editor-projection.ts writes the host's own session through setWorkspaceSessionForWorktree, the same way the listed headless mobile-session tab writers do. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(automations): a headless host closes finished run terminals, keeping the newest few (#25831) * fix(automations): a headless host closes finished run terminals, keeping the newest few The desktop closes a run's terminal when the run completes; orcad had no renderer to do it, so hourly automations left a shell and PTY per run open forever (28 after ~6h on a real host). The headless service now closes a finished run's terminal after a 10-minute grace, keeps the newest three per automation viewable, and never touches a run that has not finished. * fix(automations): never close a run terminal a client typed into or is viewing Mirrors the desktop's take-over rule on headless hosts: a finished run's terminal stays open when any client drove input to it since spawn, is attached to or viewing it, or when this process cannot tell (it adopted the PTY rather than spawned it). * test(runtime): register a viewer through the public subscribe API * fix(automations): close only completed runs' own panes A failed run can still hold a live agent (blocked on a prompt, past the watch window, or after an observer error), so like the desktop only a completed run's terminal is closed. And only the run's own pane closes, so a pane a user split into the same tab survives. * fix(automations): close a run pane only while it still holds the run's PTY The use check read run.terminalPtyId, but the close hit whatever PTY now occupies the run's pane. Restart-exited-pane and the Codex account-switch restart put a new PTY there, so a terminal a user was using could be killed. The close now resolves the pane's current PTY and closes only when it is the run's own; otherwise it closes nothing and only clears the run's terminal. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a snapshot restore past the exec timeout keeps its fence instead of racing a second restore (#25811) * fix(orcad): a snapshot restore past the exec timeout keeps its fence instead of racing a second restore A capture, restore or clear ran under the generic 30s exec timeout. On ssh2 the timeout closes the channel and reads as a confirmed failure, but sshd leaves a pty-less command running, so rollback ran its rescue restore and recover orphaned the fence for a second one, both in the same stage. - State mutations run through execOrcadStateMutation: no abort, a client wait past the host's deadline, and any closed channel or busy/deadline answer is unconfirmed, so the fence stays fresh. - POSIX hosts wrap each one in `timeout -s KILL` (where present) and a pid-checked lock dir under ~/.orca-remote; the Windows host script takes the same lock. * fix(orcad): a running state mutation keeps the activation fence fresh The fence goes stale by its lock dir's mtime after 20 minutes, and a capture, restore or clear can now run up to 15 under it, so a rollback's rescue capture plus restore could outlast the window and let a recovery steal the fence from a live run. While a mutation runs, the host now touches the fence every 60s (POSIX: a background beat that stops with its shell; Windows: an interval in the host script, whose mutations are now async so the timer runs). A dead process stops refreshing, so stale takeover still recovers it. * fix(orcad): the state-mutation fence heartbeat never refreshes a wake's fence A wake writes .orca-wake-owner into the fence dir and lets its fence age toward takeover; the heartbeat now skips a fence that holds that token, and only ever changes the dir's mtime. * fix(orcad): a state mutation releases its host lock before answering, and names its holder by pid and start time On Windows answer() exits in the stdout write callback, so an op that answered before its first await (MISSING, EMPTY, FAILED) exited before the wrapper's finally and leaked the lock; a reused pid then read as alive and every later capture, restore and clear answered busy. Ops now return their token and the wrapper answers after releasing the lock. The holder is pid plus creation time (the slot's process-tree addon); one that cannot be identified is stale once its lock misses five heartbeats. POSIX gets the same heartbeat-age check for a reused pid. * fix(orcad): a state mutation's host lock is owned by its whole process group The lock named only the shell's pid, so a shell killed while its rm or tar ran let the next mutation take the lock and race that child. Each mutation now runs in its own process group (setsid, or perl setpgrp on macOS), with timeout inside it so a deadline KILL reaches the children too. The lock records the group, and is taken over only once no member is alive; a host that can start no group records none, and its lock is never taken over. Windows ops run in-process, with no children to outlive the holder. * fix(orcad): record a state mutation's process group without ps -p, and never hold a groupless lock forever BusyBox ps has no -p, so Alpine hosts recorded no group and their lock read busy forever after a timeout kill, reboot or OOM. The group now comes from /proc/<pid>/stat (read after the comm field's last paren), with ps -o pgid= -p as the fallback. A lock that still names no group is taken over once its pid is dead and its heartbeat has missed three beats. * fix(orcad): only proof of exit frees a state-mutation lock A Windows holder whose creation time could not be read was taken over after five quiet minutes though its pid was alive, so a suspended clear could resume and delete freshly restored profiles. Both platforms now free the lock only on proof of exit: a dead pid, a different creation time, or (POSIX) a group with no live member. A live holder of unknown identity stays busy until it exits. The owner record is written exclusively, so a run that resumes after a takeover backs off. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): never offer Move for terminals another Orca desktop or session runs (#25815) On a host where another desktop held a live relay shell, this desktop read relay_terminals_live with offerMove, and its copy ("Its N open terminals will restart") implied they were its own. The census already attributes them: terminals counted only by the host-wide census, with no lease or listing of this target naming one, run under another target or session. That verdict now carries elsewhere / terminalsElsewhere; no move is offered (no toast, no status-line action) and the status line says the terminals belong to another Orca desktop or session. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(automations): an update releases finished, unused automation shells before counting terminals (#25844) Hosts with schedules kept completed run shells (the newest three, and any not yet past their grace), which counted as running terminals and deferred every on-connect update with orcad_update_terminals_running. The update and rollback census now ask the server to close completed automation run terminals no client used, with no grace or keep rule, and count after the daemon drops them. Used, unknown, failed and running ones still count. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): every activation fence holder carries a generation token its steps and release must match (#25834) * fix(orcad): clear a bare stale activation fence instead of asking for Recover, and report a restarting update from status BUG-21: a wake cut short leaves a stale fence with no journal. Every update then answered 'Recover it first' while Recover answered 'none'. The fence-hold check now takes such a fence over and drops it, and the update retries once; Recover is asked for only over a journal. The CLI also treats a connection closed by the server's own restart during update or rollback as expected and reports what status shows once the runtime answers. * fix(cli): type the reconnect status response explicitly * fix(ssh): a wake's fence carries its owner token from the moment the lock exists The idle-exit e2e still read 'fenced' on reconnect: the launch-time wake's lock landed on the host but its connection dropped before the client saw OK, so the wake body never ran and never wrote its owner token, leaving a fence nothing could prove. The token is now claimed before the lock and written by the same command that creates it, and a wake registers itself before any remote step so a reconnected wake waits for it instead of racing it. * fix(orcad): every activation fence holder carries a generation token its steps and release must still match Astra pass 8: a holder suspended past the stale window resumed, kept acting, and its unconditional release deleted the successor's fence and recovery journal mid-update. Every holder (activation, rollback, stop, recover, wake) now writes a token into the lock it creates or takes over. Each remote step it issues checks that token on the host, in the same command on POSIX and inside the host script for Windows host ops; release is conditional on the token and moves the lock aside instead of removing the root. A superseded holder aborts with OrcadFenceLostError and its release is a no-op. * fix(orcad): state mutations check the fence token before their lock, and refresh only a fence they still own On POSIX the fence guard runs outermost in serializedStateMutationCommand, before the mutation lock and the work, and the heartbeat touches the fence only while the token is still this run's. The Windows host script records the --fence token and refreshFence compares it. A fence-lost answer from a state mutation is a refusal, never a FAILED fallback. One owner-file constant replaces the wake-owner copies. * refactor(orcad): a state mutation's heartbeat touches the fence directory it checked ownership of (review) * fix(orcad): a release moves the journal and lock aside and keeps only its own generation's Astra pass 9: a release that passed its token check and stalled before deleting could, once a takeover and a successor came and went, delete the successor's journal and lock. The journal is now stamped with the writing run's fence token (a recovery takeover re-stamps the journal it adopts), and the release renames the journal and the lock aside, deletes each only if it carries this run's token, and otherwise moves it straight back. * test(orcad): a successor restore stays busy beside a paused clear on the Windows host script * fix(orcad): classify a lost fence from the step's exit and stdout, never the error message The real exec error quotes the command, and every fenced command carries the guard's marker text, so any failed or timed-out fenced step read as a lost fence and dropped its unconfirmed flag. execCommand now attaches exitCode and stdout to its exit error; execOrcadRemote rethrows unconfirmed terminations before any reclassification. * fix(orcad): a wake keeps its fence token until the fence is released A disconnect fails a wake's next step without the unconfirmed flag, and the fence release then fails over the dead connection. Forgetting the token on that error left a fence the reconnected wake could not prove its own, so it reported the host as held by an update. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(recovery): keep Phase 3's recovery-lifetime test on main's legacy-worker ports Main added a required hasRequestedReleases port and now skips persist when a pass resolves nothing, so the test mocks the new port and holds the pass at workspace resolution instead. * fix(ssh): say "1 terminal" when another Orca desktop runs one on the host (#25853) The terminalsElsewhere status line had no plural forms, so B9 read "while 1 terminals another Orca desktop … are running". It gains _one/_other entries like the other terminal-count strings on that line and in the move offer, which were already pluralized. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(automations): close failed and exited runs' terminals once the shell is proven alone (#25859) * fix(automations): close failed and exited runs' terminals once the shell is proven alone Failed (command not found, timeout) and forever-dispatched runs kept one shell per run, unbounded and counted by the update gate. Their terminals now close like completed ones (unused, past the grace, outside the newest few) but only on fresh execution-host proof that the spawned shell is alone at its prompt; a live or unprovable agent keeps its terminal. A still-dispatched run closed this way is marked failed. Dead terminals no longer take one of the newest-three keep slots. The update drain follows the same rules. * fix(automations): prove a run shell alone from the process table, not the daemon's ownership flag On a real daemon session the daemon's confirmShellForeground stays false after a plain 'command not found' and after an agent that exited, because its ownership flag only turns 'shell' after a full-screen command; failed runs would never have closed. The proof now also reads the host's process table: on POSIX the PTY's root shell must own the terminal foreground group with nothing stopped under it, on Windows the host's job-based child census must be empty. Anything unobservable still keeps the terminal. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(ssh): extensive orca CLI matrix on Windows hosts (#25114) * test(ssh): extensive orca CLI matrix on Windows hosts Adds three dispatch-only app cells to the ssh-windows-hosts lane that drive the e2e build and the bundled orca CLI against the provisioned Win32-OpenSSH host: empty-host deploy/terminal/reconnect/orcad-restart/decommission, seeded relay-era conversion, and an open relay terminal keeping the host on the relay. * test(ssh): pin the relay-kept cell's runtime; keep cleanup from masking failures * test(ssh): run decommission before the orcad restart in the managed cell * test(ssh): decommission through orca environment stop; accept an unverifiable relay close * test(ssh): require a confirmed relay close; app cells must run last * test(ssh): log and accept either relay-kept census reason; keep app-cell test results * test(ssh): relay-kept requires a live census and its status line again * test(ssh): match the pluralized relay-kept status line * test(ssh): orcad restart proves a new process, terminal adoption, and a kill-then-connect relaunch * test(ssh): restart kills only the orcad server, not its terminal daemon; wait for a released profile * test(ssh): collect orcad.log.1 so a restarted orcad's previous run is kept * test(ssh): the managed cell proves a workspace listener is detected and attributed * test(ssh): start the port listener without $, so a PowerShell terminal doesn't expand it * test(ssh): the port check proves Windows command-line attribution; retry a dropped version read --------- Co-authored-by: m4air <m4air@Mac.localdomain> * refactor(runtime): move run-terminal client-use and shell-alone checks out of the branch-cleanup runtime (#25876) * refactor(runtime): move run-terminal client-use and shell-alone checks out of the branch-cleanup runtime orca-runtime-preserved-branch-cleanup.ts had grown past max-lines (303) with the headless run-terminal helpers. Their logic now lives in run-terminal-client-use.ts and the runtime keeps one-line delegators, with behavior unchanged. * fix(ci): the runtime Electron ratchet bundles its entry points once, not 2.5k times check-runtime-electron-ratchet bundled ~2,532 entry points each in full (format cjs, no splitting), so esbuild held thousands of copies of the runtime graph: about 2.2GB RSS and 11s per run, twice per test file. It was in flight in every unit shard that died with "The runner has received a shutdown signal" (#25815 5/5 twice, #25876 2/5 twice). With esm + splitting the shared modules land in one chunk: same metafile, about 200MB and 2s. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(ssh): wait for the busy relay's child before probing it (#25916) The fake relay's spawn is not visible to pgrep at READY on Linux under Bun, so the probe could count zero children. The sibling cases already wait. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a quit that aborts an upload whose read already ended no longer crashes main (BUG-23) (#25922) Quitting while an on-connect orcad update was uploading its bundle aborted the connection's teardown signal. sftp-upload's abort handler destroyed the local read stream with the signal's reason, but once that read had ended, 'finished' had already removed its listeners, so the stream emitted an unhandled 'error': [main_uncaught_exception] AbortError: This operation was aborted. Electron's error dialog then blocked the main thread and the app never exited. The read stream now always has a no-op error listener; the transfer's outcome still comes from 'finished'. Both the bare upload and the connection-level teardown abort are covered. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a launch reads readiness at least once; fake hosts match the capture's tar flag, not any -cf (#25918) A random fence token contains `-cf` about 1 time in 125, and the fake hosts read any command containing it as a snapshot capture, so a rollback's restore answered CAPTURED and the rollback never launched. Separately, a client descheduled between computing the readiness deadline and checking it skipped every read and failed a ready launch. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a fence this desktop's exited process left is cleared without the 20-minute wait (#25941) * fix(orcad): a fence this desktop's exited process left is cleared without the 20-minute wait The client records the fence tokens its processes hold beside the profile. On a later launch, a POSIX fence carrying a token from a process that has exited, quiet for three heartbeats and with no live state mutation, is backdated so the existing stale rules clear it or hand it to Recover at once. Another desktop's fence, a live holder's, or one with a mutation still running keeps the normal stale window. * fix(orcad): held fence tokens are best effort, pinned to this machine and boot, and pruned after a day A token-file write that fails no longer breaks a fence operation; an entry recorded on another machine sharing the profile, or before a reboot, never proves its holder exited; entries older than 24 hours are dropped. Tests cover a journal kept for Recover, a successor freshened back after a racing backdate, a Windows host, and the record across release, supersession, busy and a lost connection. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): an install lock this desktop's exited process left mid-upload is taken over without the 20-minute wait (#25991) A quit during the bundle upload leaves the version dir's install lock, not the activation fence. The lock now carries this desktop's token, recorded in the held-token store, and is forgotten only once its removal is confirmed. On a later attempt, before each stale check, a POSIX lock whose token belongs to an exited process of this machine and boot, quiet for three minutes, is backdated so the existing stale takeover claims it at once. The fence path now shares the same helper. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a desktop that met another desktop's update fence clears its note once the host answers (#25995) The serving note "holds this host" and a fence-busy update deferral stayed until a reconnect, minutes after the other desktop's update finished. The connect now rechecks serving and the update every 45s while the fence holds, and publishes the first answer without it. A recorded deferral is dropped once the host runs its candidate or a newer release. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * test(ci): run Phase 3's SQLite-backed tests in the Node runtime project Main's #25967/#25998 boundary requires every test that opens real SQLite to be listed. This adds Phase 3's eleven orcad and SSH migration tests, plus main's own agent-launch-instant-tab test (#25430), which main's tip also leaves unlisted. * fix(ci): keep Electron probes out of the node-server suites again (#26046) The runner excluded *.electron.test.ts with a CLI --exclude, but main's switch to Vitest inline projects (#25967) gave each project its own exclude list, which overrides the CLI one. The directory selectors then pulled profile-state-writer-stall.electron.test.ts into the glibc-floor and musl orcad-template jobs, which have no xvfb. Resolve the exact files with vitest list and drop Electron and cross-runtime ones before running. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): keep the SSH host card quiet while its managed server is healthy (#26072) The card showed "Runs a managed Orca server" under every healthy host. A managed server is the default, so the status line now appears only for setup progress, updates, the relay, or failures. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): Move to managed server keeps the host's terminal tabs (#26077) * fix(ssh): Move to managed server keeps the host's terminal tabs Move stops the relay shells; their exits read as a user exit and closed the tabs before the conversion copied them to the server. Suppress those exits for the move, restart stopped shells on the relay when the host stays, re-home the open workspace onto the server, and report stopped shells to the runtime so terminal list stops calling them connected. * fix(ssh): mark Move's relay stops in main's intentional-stop register The renderer-only exit suppression left main retiring the stopped tab from the saved SSH session before the conversion copied it, left other viewers unprotected, and swallowed real exits for the whole request. Register exactly the shells the move stops, from just before each shutdown, as a 'replaced' stop with their incarnation; main keeps the surface and labels the exit for every viewer, while a confirmed death stays 'exited'. Move now returns the shells it stopped, and a host that stays on the relay restarts only those tabs, discarding any buffered exit first. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(hosts): show an SSH host and its managed Orca server as one host (#26076) * fix(hosts): show an SSH host and its managed Orca server as one host Phase 3 registers the Orca server it deploys over SSH as its own runtime environment, so every host list built from the execution-host registry listed the machine twice under the same name. The registry now folds the pair into one row named after the SSH host. The row routes to the server, since a managed host has no relay, unless main reports the host back on its relay; the other id stays as an alias so selections and renames saved under it still resolve. A retired id that workspaces still point at keeps its own row, and servers no configured SSH host deployed (manual pairings, orphans) are untouched. * fix(hosts): keep both ids of a merged SSH host and dedupe only in pickers Deleting the merged-away id from the registry broke every consumer that matches hosts by exact id: Add Project fell back to local after a connect, the composer lost ready projects and drafts (and could swap in an unrelated local project), and a host scope hid folder-only workspaces. The registry now keeps both entries and marks the pair (aliasHostIds on the row pickers show, mergedIntoHostId on the other). Pickers show one row per machine, and a choice of that row expands to both ids: sidebar host scope, jump palette filter, notification toggles, run-target and repository host offers. Add Project resolves a saved SSH id to its server row and blocks the actions while that server comes up instead of choosing local. The composer's resolver now fails closed when a named draft repo isn't actionable rather than picking another project. * fix(hosts): widen saved host scopes, both-way palette aliases, guard Add Project host - A sidebar or agents host scope saved by an older build (or before a route flip) can hold one id of a merged SSH host; a background gate widens it to both ids so exact-id filters match either owner. - The palette host filter now resolves a saved id to both owners whichever id it names. - Add Project's create and clone refuse to run while the chosen host is unresolved, and their submit buttons stay disabled, instead of falling through to this computer. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(sync): reconcile main's ratchet bundling and cold-serve hydrate with phase3 The Electron-import ratchet keeps main's single-stdin bundle (cjs); the auto-merge had also kept phase3's esm splitting, which broke main's import-graph test. Editor tabs now follow the windowless full-seed rule from #26022, so a cold serve restart lists persisted editors too. * fix(ssh): reclaim this desktop's own exited lock on Windows hosts too (#26087) * fix(ssh): reclaim this desktop's own exited lock on Windows hosts too The relaunch after a quit mid-update now frees the activation fence and the version-dir install lock on a Windows SSH host the same way it does on POSIX, instead of waiting out the 20-minute stale window. The host script ages the lock only when its token belongs to a desktop process proven exited, it has been quiet for three heartbeats, and (for the fence) no state mutation is live, where a mutation holder counts as gone only by pid plus creation time. * fix(ssh): take an exited holder's lock only through the steal arbitration Review found the reclaim backdated the lock by path after checking it, so a live successor that replaced the lock in between could be aged and then stolen, and an interrupted or failed restore left it aged for good. The exited-holder check is now read-only. The steal command itself accepts the proven token and, inside its steal claim and identity recheck, also takes a lock whose owner file still names that token and that has been quiet for three heartbeats. Nothing is written to a lock before the steal owns it. POSIX uses the same path. * fix(ssh): never take an exited holder's fence while a state mutation can start Review round 2 found the fence's live-mutation guard ran only in the read-only proof, so a mutation admitted after the proof, or one whose first heartbeat landed after the steal sampled the fence's age, kept running under a fence the steal had replaced. For the fence, the steal now takes the state-mutation lock inside its claim (mkdir on POSIX, the exclusive owner.json on Windows) and holds it until the takeover is done; it refuses when any mutation lock exists. Holding it, it rereads the owner and only then re-samples the fence identity. A mutation now rechecks its fence token right after it takes the mutation lock and stops with the fence-lost marker if it changed. The Windows proof also falls back to the stale window when its command line would not fit cmd.exe. * fix(ssh): record the exited-owner steal as a real mutation-lock holder Review round 3 found the POSIX steal held the state-mutation lock as an empty directory, which a mutation reclaims after a minute without any liveness check; a steal stalled that long lost its exclusion and could replace the fence under a running mutation. The steal now writes its pid (and group, under the same rule) with the mutation's own noclobber owner writer, so only proof of its exit frees the lock, and it removes the lock only while the lock still names it. On Windows the owner record is moved into place whole, so it never exists empty, and is removed only while it still names the steal's pid. * refactor(ssh): keep the relay lock commands off the orcad host-script graph The mutation-lock owner writers moved into a leaf module, so the relay's install-lock commands no longer import orcad-state-snapshot and, through it, the Windows host script, orcad-instance-lock and the daemon process query. Those modules evaluate imports at load time that existing suites mock partially. No behavior change. --------- Co-authored-by: m4air <m4air@Mac.localdomain> --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
380 lines
20 KiB
JavaScript
380 lines
20 KiB
JavaScript
import process from 'node:process'
|
|
import { pathToFileURL } from 'node:url'
|
|
|
|
const isProductSource = (file) => !/\.test\.tsx?$/.test(file)
|
|
|
|
// Why config/patches: the xterm fork owns the helper textarea an input method attaches to, so a
|
|
// patch edit can break composition without touching a file named "ime".
|
|
const NATIVE_IME_PRODUCT_SOURCE =
|
|
/^(?:config\/patches\/|src\/shared\/terminal-unicode-provider\.ts$|src\/renderer\/src\/lib\/pane-manager\/terminal-ime-|src\/renderer\/src\/components\/terminal-pane\/(?:terminal-ime-|terminal-ios-hangul-|xterm-bypass-policy))/
|
|
|
|
/** The harness itself: the session runner, the boundary probes, and the native specs. */
|
|
const NATIVE_IME_HARNESS =
|
|
/^(?:config\/scripts\/focus-nested-wayland-terminal\.sh$|config\/scripts\/(?:run-terminal-ibus-hangul-e2e|terminal-ime-engagement-receipt)\.mjs$|tests\/e2e\/terminal-ime-(?:boundary-probe|byte-reader|engagement-receipt)\.ts$|tests\/e2e\/terminal-(?:ibus-hangul|hangul-terminating-digit|macos-2set-korean)-native\.spec\.ts$)/
|
|
|
|
export const PR_E2E_SOURCE_ROUTES = [
|
|
{
|
|
id: 'serve.orcad-mode-switch',
|
|
specs: ['tests/e2e/orcad-serve-mode-switch.spec.ts'],
|
|
matches: (file) =>
|
|
/^tests\/e2e\/helpers\/(?:orca-serve-cli-host|headless-paired-runtime-host)\.ts$/.test(
|
|
file
|
|
) ||
|
|
(isProductSource(file) &&
|
|
/^src\/(?:cli\/runtime\/(?:launch|serve-)|main\/orcad\/(?:main|orcad-entry|orcad-instance-lock|orcad-command-arguments|orcad-lifecycle)\.ts$|main\/startup\/desktop-profile-instance-lock\.ts$|main\/daemon\/daemon-(?:spawner|endpoint-adoption|init)|main\/server\/serve-)/.test(
|
|
file
|
|
))
|
|
},
|
|
{
|
|
id: 'startup.windows-missing-appdata',
|
|
specs: ['tests/e2e/windows-missing-appdata-startup.spec.ts'],
|
|
matches: (file) =>
|
|
file === 'tests/e2e/helpers/orca-serve-cli-host.ts' ||
|
|
(isProductSource(file) &&
|
|
/^src\/main\/startup\/(?:windows-app-data-path|main-process-preflight)\.ts$/.test(file))
|
|
},
|
|
{
|
|
id: 'ssh.orcad-auto-convert',
|
|
specs: ['tests/e2e/ssh-orcad-auto-convert.spec.ts'],
|
|
matches: (file) =>
|
|
/^tests\/e2e\/helpers\/(?:orcad-convert-(?:flow|host)|orcad-template-variant|orcad-upgrade-profile)\.ts$/.test(
|
|
file
|
|
) ||
|
|
(isProductSource(file) &&
|
|
/^src\/main\/(?:ipc\/ssh-host-server-|ssh\/(?:ssh-host-server-|orcad-runtime-conversion|orcad-migration-|orcad-retained-source|orcad-runtime-deployment))/.test(
|
|
file
|
|
))
|
|
},
|
|
{
|
|
id: 'ssh.orcad-idle-exit',
|
|
specs: ['tests/e2e/ssh-orcad-idle-exit.spec.ts'],
|
|
matches: (file) =>
|
|
/^tests\/e2e\/helpers\/orcad-convert-(?:flow|host)\.ts$/.test(file) ||
|
|
(isProductSource(file) &&
|
|
/^src\/(?:main\/(?:orcad\/orcad-(?:idle-|managed-idle-)|ssh\/orcad-(?:managed-wake|managed-tunnel|recovery-slot|remote-launch))|shared\/orcad-idle-exit)/.test(
|
|
file
|
|
))
|
|
},
|
|
{
|
|
id: 'ssh.localhost-agent-hooks',
|
|
specs: ['tests/e2e/ssh-localhost.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^src\/(?:relay\/(?:agent-hook|relay-agent-hook-runtime|plugin-overlay)|main\/(?:agent-hooks\/|ssh\/ssh-relay-session\.ts$)|shared\/agent-hook)/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
id: 'browser-network.ssh-docker-route',
|
|
specs: ['tests/e2e/ssh-browser-network-execution-route.docker.unit.test.ts'],
|
|
matches: (file) =>
|
|
file === 'tests/e2e/ssh-browser-network-execution-route.docker.unit.test.ts' ||
|
|
/^tests\/e2e\/helpers\/docker-ssh-relay-(?:image|target)\.ts$/.test(file) ||
|
|
(isProductSource(file) &&
|
|
/^src\/main\/(?:browser\/(?:ssh-browser-network-execution-route|browser-network-deferred-socket|browser-network-execution-route|system-ssh-socks-client-socket)|ssh\/system-ssh-dynamic-forward-process)\.ts$/.test(
|
|
file
|
|
))
|
|
},
|
|
{
|
|
// Why the host-connection phase: the route gate waits on it, so a phase change can strand the
|
|
// SSH-unavailable card without touching a browser file.
|
|
id: 'browser.local-ssh-workspace-route',
|
|
specs: ['tests/e2e/local-ssh-browser-routing.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^src\/(?:main\/browser\/local-ssh-browser-(?:route|partitions)\.ts|renderer\/src\/(?:components\/browser-pane\/(?:use-ssh-workspace-browser-route\.ts|assemble-chrome\/ssh-routed-browser-page-gate\.tsx)|lib\/worktree-host-connection-phase\.ts))$/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
id: 'terminal.windows-wsl-launch-and-paste',
|
|
specs: [
|
|
'tests/e2e/golden-tab-bar-agent-launch.spec.ts',
|
|
'tests/e2e/terminal-windows-shell-paste-ownership.spec.ts'
|
|
],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^(?:config\/scripts\/(?:verify-wsl-e2e-participation|verify-playwright-participation)\.mjs$|src\/main\/(?:wsl[/-]|pty\/.*wsl|providers\/wsl)|src\/shared\/(?:wsl-|windows-terminal-shell)|src\/renderer\/src\/.*(?:terminal-paste|pty-paste)|tests\/e2e\/(?:golden-tab-bar-agent-launch\.spec|terminal-windows-shell-paste-ownership\.spec|helpers\/(?:wsl-golden-stub-agent|golden-stub-agent))|\.github\/(?:actions\/setup-wsl-test-runtime\/|workflows\/windows-wsl-e2e\.yml))/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
id: 'ephemeral-vm-runtime.rollback-readable-sidecar',
|
|
specs: ['tests/e2e/ephemeral-vm-provisioned-root.spec.ts'],
|
|
matches: (file) =>
|
|
/^(?:src\/main\/ephemeral-vm-(?:runtime-(?:service|provisioning-persistence)|failed-start-cleanup)|src\/shared\/(?:ephemeral-vm-runtime-(?:store|feature-store|rollback-projection|runtimes)|ephemeral-vm-recipes|orca-yaml-hook-types))\.ts$/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
id: 'ssh-terminal-source',
|
|
specs: [
|
|
'tests/e2e/pty-input-write-queue-ssh.spec.ts',
|
|
'tests/e2e/ssh-codex-display-artifacts-repro.spec.ts',
|
|
'tests/e2e/ssh-cold-activation-restore.spec.ts',
|
|
'tests/e2e/ssh-docker-half-open-link.spec.ts',
|
|
'tests/e2e/ssh-docker-reconnect-pane-restore.spec.ts',
|
|
'tests/e2e/ssh-docker-relay-stall-credential.spec.ts',
|
|
'tests/e2e/ssh-docker-resource-accumulation.spec.ts',
|
|
'tests/e2e/ssh-docker-transport-drop-recovery.spec.ts',
|
|
'tests/e2e/ssh-port-forward-lifecycle.spec.ts',
|
|
'tests/e2e/ssh-reconnect-tab-destruction.spec.ts',
|
|
'tests/e2e/ssh-startup-exec-readiness.spec.ts',
|
|
'tests/e2e/ssh-terminal-window-wake-stale-grid-repro.spec.ts'
|
|
],
|
|
// Why the store/startup/shared additions: the SSH-named authorities stop at the main
|
|
// process and the pane component, but the reconnect ledgers and retained-payload
|
|
// admission that decide whether a pane rebinds live in the renderer store.
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^(?:src\/main\/ssh\/|src\/main\/providers\/ssh-|src\/main\/ipc\/(?:ssh-|pty)|src\/main\/runtime\/(?:public-ssh-state|ssh-file-explorer-chunk-read)\.ts|src\/relay\/|src\/shared\/(?:ssh-|skill-ssh-relay-contract)|src\/renderer\/src\/startup\/(?:ssh-startup-reconnect|startup-ssh-connection-restore)\.ts|src\/renderer\/src\/store\/slices\/(?:ssh|direct-ssh-)|src\/renderer\/src\/components\/terminal-pane\/(?:pty-|ssh-|remote-runtime-|terminal-parked-pty))/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
// Why a sibling route rather than more paths on ssh-terminal-source: these modules carry
|
|
// no "ssh" in their names, and only the two restore specs gate them. Folding them in
|
|
// would run the whole SSH terminal list for a tab-tombstone edit.
|
|
id: 'ssh-workspace-session-restore',
|
|
specs: [
|
|
'tests/e2e/ssh-cold-activation-restore.spec.ts',
|
|
'tests/e2e/ssh-reconnect-tab-destruction.spec.ts'
|
|
],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
!file.endsWith('-test-harness.ts') &&
|
|
/^(?:src\/main\/ipc\/remote-workspace|src\/shared\/remote-workspace-|src\/renderer\/src\/hooks\/remote-workspace-|src\/renderer\/src\/lib\/worktree-(?:initial-terminal-seeding|default-terminal-tabs)\.ts|src\/renderer\/src\/components\/terminal\/initial-terminal)/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
id: 'terminal-input.ime-and-synthetic-forwarding',
|
|
specs: [
|
|
'tests/e2e/terminal-cjk-ime-committed-text.spec.ts',
|
|
'tests/e2e/terminal-hangul-wrap-boundary-bytes.spec.ts',
|
|
'tests/e2e/terminal-ime-exact-byte.spec.ts',
|
|
'tests/e2e/terminal-korean-composing-chord-order.spec.ts',
|
|
'tests/e2e/terminal-korean-endofrow-preedit-cell-span.spec.ts',
|
|
'tests/e2e/terminal-korean-midline-preedit-occlusion.spec.ts',
|
|
'tests/e2e/terminal-korean-preedit-visibility.spec.ts'
|
|
],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^(?:config\/patches\/|src\/renderer\/src\/components\/terminal-pane\/(?:terminal-ime-|use-terminal-pane-lifecycle|xterm-bypass-policy|terminal-option-shortcut-policy))/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
// Why a route beside terminal-input.ime-and-synthetic-forwarding rather than more specs on
|
|
// it: that route selects the CDP-synthetic specs, which drive composition through
|
|
// Input.imeSetComposition and so prove Orca's handling without an input method existing.
|
|
// This one names the surface only a real ibus-hangul session can judge, and is the sole
|
|
// trigger that puts the real-IME lane on a PR.
|
|
id: 'terminal-ime.native-input-method',
|
|
specs: ['tests/e2e/terminal-ibus-hangul-native.spec.ts'],
|
|
matches: (file) =>
|
|
(isProductSource(file) && NATIVE_IME_PRODUCT_SOURCE.test(file)) ||
|
|
NATIVE_IME_HARNESS.test(file)
|
|
},
|
|
{
|
|
id: 'terminal-startup.quick-command-pre-bind-recovery',
|
|
specs: ['tests/e2e/terminal-quick-command-pre-bind-recovery.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^(?:src\/renderer\/src\/components\/tab-bar\/TabBarQuickCommandsMenu\.tsx|src\/renderer\/src\/hooks\/use-terminal-quick-command-hosts\.ts|src\/renderer\/src\/components\/terminal-pane\/(?:pty-connection|pty-transport|terminal-pty-pre-spawn-e2e-barrier)\.ts|src\/renderer\/src\/components\/terminal-pane\/pty-connection\/(?:connect-pane-pty|fresh-spawn-start|pane-pty-visibility-bind|pty-input-recovery)\.ts|src\/renderer\/src\/components\/terminal-pane\/(?:TerminalPane|use-terminal-pane-lifecycle)\.tsx?|src\/renderer\/src\/store\/slices\/terminals\.ts)$/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
id: 'quick-open.paired-host-path-search',
|
|
specs: ['tests/e2e/paired-quick-open-large-tree.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^(?:src\/main\/ipc\/filesystem-(?:list-files|search-file-paths)\.ts|src\/main\/ripgrep\/bundled-ripgrep-path\.ts|src\/main\/providers\/(?:filesystem-provider-contract|ssh-filesystem-provider(?:-capabilities)?)\.ts|src\/main\/runtime\/(?:orca-runtime-files|rpc\/methods\/files)\.ts|src\/relay\/(?:fs-handler(?:-install-rg|-list-files|-ripgrep-fallback)?|fs-list-files-fallback-chain|relay-bundled-ripgrep)\.ts|src\/renderer\/src\/(?:components\/(?:QuickOpen|quick-open-file-list|quick-open-search)\.tsx?|runtime\/(?:runtime-file-client|runtime-legacy-quick-open-inventory)\.ts)|src\/shared\/(?:quick-open-(?:install-rg|path-search|transport-budget)|ripgrep-process-availability|bundled-ripgrep)\.ts)$/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
id: 'terminal-session.host-cold-park-stream-continuity',
|
|
specs: ['tests/e2e/host-parked-pane-remote-viewer.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^(?:src\/renderer\/src\/components\/terminal-pane\/(?:terminal-hidden-view-parking|terminal-tab-park-candidates|terminal-tab-activation-order|terminal-parked-pty-watcher|terminal-parked-tab-watchers|terminal-parked-watcher-registry)\.ts|src\/renderer\/src\/runtime\/sync-runtime-graph\.ts)$/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
// Why a route of its own: every other terminal-pane route names what BINDS a pane — the pty
|
|
// transports, the ssh reconnect ledgers, the park watchers. Nothing named what unbinds one,
|
|
// so the close/retire lifecycle reached main with e2e skipped outright. Unbinding is the half
|
|
// that can strand a PTY or leave a retired leaf mounted as a blank pane.
|
|
//
|
|
// Deliberately absent: src/renderer/src/runtime/runtime-rpc-client.ts, the transport these
|
|
// retirements call out through. It carries no close decision and churns ~3x these files, so
|
|
// routing on it would run this lane on unrelated runtime work.
|
|
id: 'terminal-pane.close-and-retirement',
|
|
specs: [
|
|
// Closing a tab whose pane is parked (never mounted) must retire that exact PTY.
|
|
'tests/e2e/terminal-parked-close-retirement.spec.ts',
|
|
// Closing one leaf of a split must leave root leaves, leaf→pty bindings, and live panes
|
|
// agreeing — the ghost-blank-pane shape a bad unbind produces.
|
|
'tests/e2e/terminal-pane-close-layout-consistency.spec.ts',
|
|
// The runtime half: a leaf the host retires must stop being mounted on a paired client.
|
|
'tests/e2e/paired-remote-split-pane-host-retired-ghost.spec.ts'
|
|
],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^(?:src\/renderer\/src\/components\/terminal-pane\/(?:retire-unbound-(?:ipc|runtime)-terminal-pane|terminal-pane-(?:close-admission|close-identity|lifecycle-close|pane-closed|retirement-ownership)|use-terminal-pane-close-actions)|src\/renderer\/src\/store\/(?:terminals\/terminal-tab-close(?:-providers)?|slices\/(?:terminal-tab-retirement|terminal-retirement-teardown-reservation|retired-terminal-tab-state-sweep)))\.ts$/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
id: 'terminal-session.parked-cli-split',
|
|
specs: ['tests/e2e/terminal-parked-cli-split.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^(?:src\/main\/window\/attach-main-window-services\.ts|src\/preload\/(?:index|api\/ui-command-event-api)\.ts|src\/renderer\/src\/components\/terminal-pane\/(?:terminal-pane-split-request-routing|use-terminal-pane-lifecycle|use-terminal-tab-cold-parking)\.ts|src\/renderer\/src\/hooks\/ipc-events\/terminal-ui-routing-ipc-bridge\.ts)$/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
id: 'terminal-session.paired-serve-restart-binding-continuity',
|
|
specs: ['tests/e2e/paired-remote-terminal-serve-restart-binding.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^(?:src\/main\/daemon\/(?:daemon-attach-only-retirement|daemon-pty-applied-size|daemon-pty-session-control|daemon-pty-spawn-result)\.ts|src\/renderer\/src\/components\/terminal-pane\/(?:remote-runtime-pty-transport|terminal-error-accumulation)\.ts|src\/renderer\/src\/runtime\/(?:web-runtime-session|web-session-tabs-sync|web-session-terminal-orphan-(?:topology|recovery(?:-(?:adoption|surface|inventory|inventory-validation|cache|queue|rpc-lane|pane))?))\.ts)$/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
id: 'terminal-provider.ssh-remote-reattach-contract',
|
|
specs: ['tests/e2e/paired-remote-terminal-materialization-reconnect.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
!file.endsWith('-test-harness.ts') &&
|
|
/^(?:src\/renderer\/src\/components\/terminal-pane\/remote-runtime-pty-transport(?:-[a-z0-9-]+)?\.ts|src\/renderer\/src\/runtime\/remote-runtime-terminal-multiplexer\.ts)$/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
// Why: layout resolution is the only place a split direction can be invented, and the
|
|
// loss is one-way — the guess is published and written back over the real tree.
|
|
id: 'terminal-session.split-orientation-resolution',
|
|
specs: ['tests/e2e/desktop-published-split-orientation-legacy-leaf.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^src\/renderer\/src\/runtime\/(?:remote-terminal-layout-resolution\.ts|sync-runtime-graph\/(?:graph-publication|mobile-session-terminal-tabs|mobile-session-surfaces)\.ts|web-session-tabs-sync\/terminal-surfaces\.ts)$/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
id: 'terminal-session.remote-pane-layout-retry',
|
|
specs: ['tests/e2e/paired-remote-pane-layout-retry.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^(?:src\/renderer\/src\/components\/terminal-pane\/(?:remote-pane-layout-push|TerminalPane)\.tsx?|src\/renderer\/src\/lib\/terminal-layout-equality\.ts|src\/renderer\/src\/runtime\/web-session-tabs-sync\.ts|src\/renderer\/src\/store\/slices\/terminals\.ts)$/.test(
|
|
file
|
|
)
|
|
},
|
|
{
|
|
// Why: the host's row for a client-rendered page only exists across two real Electron
|
|
// apps, so this spec is the only gate on it. The high-churn seams it also rides
|
|
// (ipc/runtime, useIpcEvents, preload) are left out deliberately: routing on those runs a
|
|
// two-app e2e on most PRs, and their client-hosted share is already covered by the
|
|
// main-process integration test.
|
|
id: 'client-hosted-browser.host-strip',
|
|
specs: ['tests/e2e/paired-client-hosted-browser-host-strip.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^src\/.*(?:[Cc]lient-?[Hh]osted-?[Bb]rowser|BrowserPaneOverlayLayer)/.test(file)
|
|
},
|
|
{
|
|
// Why a second, wider pattern: restart survival breaks from seams that never say
|
|
// "client-hosted" - page adoption, the host lease/reconciliation plan, the session-tab
|
|
// snapshot the client culls rows against. orca-runtime.ts is included despite its churn: it
|
|
// publishes the snapshot flag the client holds its rows on, and no narrower path names that
|
|
// seam.
|
|
id: 'client-hosted-browser.restart-survival',
|
|
specs: ['tests/e2e/paired-client-hosted-browser-restart-survival.spec.ts'],
|
|
matches: (file) =>
|
|
isProductSource(file) &&
|
|
/^src\/.*(?:[Cc]lient-?[Hh]osted|browser-host-(?:lease|page|client-page)|browser-client-(?:host|page)|runtime-browser-(?:client-)?page|session-tabs-sync|host-session-snapshot-authority|orca-runtime(?:-browser)?\.ts|\/runtime-(?:status|types)\.ts)/.test(
|
|
file
|
|
)
|
|
}
|
|
]
|
|
|
|
export function selectPrE2eSpecs(changedPaths, reportRoute = () => undefined) {
|
|
const specs = new Set(changedPaths.filter((file) => /^tests\/e2e\/.*\.spec\.ts$/.test(file)))
|
|
for (const route of PR_E2E_SOURCE_ROUTES) {
|
|
const matchedFiles = changedPaths.filter(route.matches)
|
|
if (matchedFiles.length === 0) {
|
|
continue
|
|
}
|
|
route.specs.forEach((spec) => specs.add(spec))
|
|
reportRoute(`[pr-e2e] ${route.id}: ${route.specs.join(', ')}`)
|
|
}
|
|
return [...specs].sort((left, right) => left.localeCompare(right))
|
|
}
|
|
|
|
/** Routes whose authorities are SSH execution source, and so require the Docker-SSH lane. */
|
|
export const SSH_SOURCE_ROUTE_IDS = ['ssh-terminal-source', 'ssh-workspace-session-restore']
|
|
|
|
// Why derive this from the routes instead of a second path list: the Docker-SSH lane used to
|
|
// trigger only because one route happened to list a startup-readiness spec, so pruning that
|
|
// spec would have silently retired the lane. Two lists that must agree is how that drifted.
|
|
export function hasSshSourceChange(changedPaths) {
|
|
return PR_E2E_SOURCE_ROUTES.filter((route) => SSH_SOURCE_ROUTE_IDS.includes(route.id)).some(
|
|
(route) => changedPaths.some(route.matches)
|
|
)
|
|
}
|
|
|
|
/** Routes whose authorities a real input method can judge, and so require the native IME lane. */
|
|
export const NATIVE_IME_SOURCE_ROUTE_IDS = ['terminal-ime.native-input-method']
|
|
|
|
// Why derived from the routes, like hasSshSourceChange: the native lane must trigger on IME
|
|
// source, not on the native spec surviving in some route's spec list.
|
|
export function hasNativeImeSourceChange(changedPaths) {
|
|
return PR_E2E_SOURCE_ROUTES.filter((route) =>
|
|
NATIVE_IME_SOURCE_ROUTE_IDS.includes(route.id)
|
|
).some((route) => changedPaths.some(route.matches))
|
|
}
|
|
|
|
export function shouldRunReusablePrE2e(changedPaths) {
|
|
// Native IME has its own workflow; SSH still runs inside the reusable workflow.
|
|
return (
|
|
hasSshSourceChange(changedPaths) ||
|
|
selectPrE2eSpecs(changedPaths).some(
|
|
(spec) => spec !== 'tests/e2e/terminal-ibus-hangul-native.spec.ts'
|
|
)
|
|
)
|
|
}
|
|
|
|
export function hasWslSourceChange(changedPaths) {
|
|
const route = PR_E2E_SOURCE_ROUTES.find(
|
|
(candidate) => candidate.id === 'terminal.windows-wsl-launch-and-paste'
|
|
)
|
|
return changedPaths.some(route.matches)
|
|
}
|
|
|
|
if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) {
|
|
let input = ''
|
|
process.stdin.setEncoding('utf8')
|
|
for await (const chunk of process.stdin) {
|
|
input += chunk
|
|
}
|
|
const changedPaths = input.split(/\r?\n/).filter(Boolean)
|
|
if (process.argv.includes('--ssh-source')) {
|
|
process.stdout.write(`${hasSshSourceChange(changedPaths)}\n`)
|
|
} else if (process.argv.includes('--reusable-workflow')) {
|
|
process.stdout.write(`${shouldRunReusablePrE2e(changedPaths)}\n`)
|
|
} else if (process.argv.includes('--wsl-source')) {
|
|
process.stdout.write(`${hasWslSourceChange(changedPaths)}\n`)
|
|
} else if (process.argv.includes('--native-ime-source')) {
|
|
process.stdout.write(`${hasNativeImeSourceChange(changedPaths)}\n`)
|
|
} else {
|
|
const specs = selectPrE2eSpecs(changedPaths, (message) => console.error(message))
|
|
process.stdout.write(`${JSON.stringify(specs)}\n`)
|
|
}
|
|
}
|