mirror of
https://github.com/stablyai/orca.git
synced 2026-10-09 00:02:39 +00:00
Phase 3: every SSH host runs a managed Orca server (orcad), replacing the relay (#24863)
* Revert "revert: take the 26 Phase 3 (#16741 port) PRs back out of main (#24559)" This reverts commit5f308bfa9c. * feat(orcad): Windows remote primitives for managed orcad hosts (W1) (#24525) * feat(orcad): Windows remote primitives for managed orcad hosts (W1) * refactor(orcad): run Windows host ops as node.exe with plain argv, no PowerShell hop * fix(orcad): refuse secret-shaped names on the breakaway launcher's --env * fix(orcad): name the secret env guard for its role --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9) (#24521) * feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9) The destination half of a catalog migration: an orcad stages a T6-7 manifest (repositories, project groups, folder workspaces, dormant session, client, automation and worktree metadata, retired names, scrollback snapshots) with exclusive claims, then commits it with a receipt so a retried commit returns the same receipt and never imports twice. Served as orcad.migration.* runtime RPC behind the orcad.migration-catalog.v1 capability; the client refuses a host without it or with method-not-found, and any other failure is left for the caller to recheck. Dormant only: no live PTY projection. Inert on the desktop until T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(rpc): catalog the orcad.migration params in the shared contract; name the catalog-import install target Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(ssh): journal and fence an SSH host for dormant migration, gated on proven terminal exit (#16741 T8-c1+c2) (#24522) A migration from a relay-hosted SSH target into a managed orcad now starts with a journal in its own sidecar directory, then the target's managed-owner fence, then a profile flush, before any remote call. A fence with no journal is unverifiable and never released; a journal whose fence is gone is stale and grants nothing; an unreadable journal fails closed. The fence requires every terminal the target ever leased to be proven exited, checked before the fence (with the relay's process list) and again under it. Same-owner claims now need the durable record that explains them. Inert until T8-c4/T6-10. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(orcad): run orcad itself on Windows hosts (W2) (#24529) * feat(orcad): run orcad itself on Windows hosts (W2) * test(orcad): load the ConPTY smoke's addon from out/orcad so the temp slot can be removed * test(orcad): skip the foreign-uid lock case when running as root --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a connected but unused SSH host previews as movable (#24609) The untransferred-dependency census counted activeConnectionIdsAtShutdown naming the target as workspace-session state. The renderer rewrites that list on every connection change, so merely connecting to an empty host blocked the move. It is a reconnect hint; the remote work it can stand for is counted on its own. The empty-target claim check likewise ignores global-field copies inside the host's session partition. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3) (#24523) * feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3) The coordinator re-checks before every stage and commit that the fenced source still exports the journaled manifest, carries no untransferable state and started no terminal. A lost answer is re-read from the destination's catalog state; only a committed read whose receipt matches the journal advances it, and the journal is on disk before anything returns. Abort releases the fence only on proof the destination holds nothing, or on an unsupported destination before anything was staged, and never once the destination committed. Codes against T6-9's catalog client through an injected interface. Inert until T8-c4/T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(ssh): import the T6-9 client's unsupported refusal instead of mirroring it --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3) (#24563) * feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3) * test(orcad): exhaustive op switch in the Windows lifecycle fake * test(ssh): narrow the Windows host-cell descriptor by lane before building a relay cell --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11) (#24608) * feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11) * fix(serve): keep orcad selection app-side and wait out Windows temp cleanup * refactor(orcad): move the data-root privacy check out of the instance lock --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch (#24619) * test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch * ci(e2e): install ripgrep for the serve mode-switch job's window-manager wait --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4) (#24562) * feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4) The conversion entry resumes or takes the fence, deploys and pairs the managed server into it, marks the server as migrated, then stages and commits the dormant catalog. Every step is keyed by the journal, so a repeat after a crash, deferral or lost reply resumes the same migration. Status reports an unfinished migration, and rollback is refused while one runs or when the rollback snapshot predates the migrated catalog. Inert until T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * style: oxfmt the c4 conversion and maintenance files * fix(ssh): name the fake migration destination's type so declarations stay portable --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4) (#24570) * feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4) * fix(orcad): accept a managed stop request whose lock path is spelled with Windows client separators --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5) (#24565) * feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5) Once the journal records destination-committed, the source profile drops the manifest's repositories, folder workspaces and unreferenced project groups, its dormant session, automation, client and worktree state, and the leases the fence proved exited. The profile flushes, the retirement is verified, the journal moves to source-retired and compacts once the server matches. A retry after any crash repeats idempotent work. The fenced target stays: it carries the managed server's tunnel. Conversion now ends retired. Inert until T6-10. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(orcad): retirement drops the migrated host from the reconnect hint The census no longer treats activeConnectionIdsAtShutdown as untransferable (#24609), so retirement must remove the target from it; otherwise a restart dials a host that is now a managed server. * style: oxfmt the c5 conversion file --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1) (#24579) * feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1) * test(orcad): start the Windows lane's exec spy after the relay gate prelude restores its own --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI) (#24590) * feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI) Adds a Managed servers section under Remote servers (deploy an empty server, status with deferred-update and migration states, update, rollback, recover, stop and cancel-stop, and SSH access for paired servers), and a Move to managed server action on connected macOS and Linux SSH hosts with a preflight summary, a terminals-closed confirmation and a resumable progress view. Main wires the conversion to the relay's process list, the direct session and the T6-9 catalog client. Everything is hidden until the new experimental setting is turned on, and Windows SSH hosts are never offered. Merges the T6-9 branch (#24521) until it lands on the integration branch. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(settings): align managed-server form controls and name the section the setting reveals Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(orcad): ask the relay with an absolute deadline via the W5 terminal-gate lister The conversion wiring passed a relative 10 s as listProcesses' deadlineMs, which the provider reads as an absolute time, so every relay inventory timed out after 1 ms and the terminal gate could never prove exit. Adopt #24579's lister verbatim so the stacks merge cleanly. * feat(settings): name blocking saved state in plain, localized words The move preview listed internal dependency ids such as workspace-session; each kind now has its own catalog entry. * feat(settings): offer managed servers and the move on Windows SSH hosts W1-W5 are on the integration branch, so a Windows relay-hosted host can deploy, convert and retire like a POSIX one. The move still waits for a connected relay that reported its platform. --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix (#24865) * test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix * test(ci): expect the orcad Windows host cells in the SSH Windows hosts workflow --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(orcad): wait for killed terminal daemons to exit before removing their temp profiles (#24871) Co-authored-by: m4air <m4air@Mac.localdomain> * feat(updater): read rollout kill switches from the update-campaign payload, all inactive (#24867) The nudge request Orca already polls may now carry an optional versioned rollout block naming the Node runtime flips. A typed reader resolves each flip with kill-switch, version range and install-id-bucketed percent semantics, falling back to the baked value (every flip inactive) when the block is absent, invalid or never read. No consumer reads it yet. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(telemetry): report Windows security-software refusals and unverifiable or failed runtime checks (#24866) ssh_remote_runtime_resolved dropped the self-test's security_software refusal to 'none' and sent nothing when a self-test was unverifiable or failed, because those attempts throw before a rung settles. Add the refusal value, self_test 'unverifiable', and an outcome field (resolved | unverifiable | failed) deduplicated per host and outcome per session. Co-authored-by: m4air <m4air@Mac.localdomain> * refactor(settings): move WorktreeVisibilityDefaults into its own module (#24884) global-settings-types.ts sits at the 300-line max-lines ceiling; merging main's two new agent-state-rules settings with experimentalManagedServers put it at 301. The worktree visibility defaults type moves next to the other visibility types and is re-exported so its 38 importers are unchanged. Co-authored-by: m4air <m4air@Mac.localdomain> * refactor(watcher): move the supervisor's child, terminating child and canary into a child slot (#24887) * refactor(watcher): move the supervisor's child, terminating child and canary into a child slot * fix(watcher,runtime): take the child slot's child type from the shared wrapper, and stub main's title-display clear in the projection test --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: m4air <m4air@m4airs-Air.localdomain> * ci(adhoc): build and ship the orcad template in adhoc macOS and Windows builds (#24969) Co-authored-by: m4air <m4air@m4airs-Air.localdomain> * feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy (#24973) * feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy * fix(orcad): keep the in-use slot plus the two most recent others, and prove same-version reuse survives eviction --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> * Phase 3: managed orcad on every SSH connect, with downgrade-safe fence (#24975, #24979) * feat(orcad): fence managed hosts outside owner and keep converted projects for downgrades Shipped builds hide any SSH target with an owner, so a downgrade made a converted host and its projects vanish. The managed fence moves to a new orcadFence field (older builds keep but ignore it), phase-3 owner fences migrate on load, and managed hosts stay visible but refuse a direct relay. Conversion now stops at destination-committed with sourceRetainedAt; source retirement waits on the baked-off orcad-source-retirement rollout flag. This build hides retained rows, and a start that finds an older build changed them marks the host sourceChangedAt (relay, needs a new move) instead of merging a second manifest. * feat(ssh): every SSH host runs managed orcad, decided on each connect (#24979) * feat(settings): managed servers are no longer experimental Remove experimentalManagedServers and its gates; the Managed servers section always shows. Loading drops a stored value, which an older build reads back as its own default (off). * feat(ssh): every SSH host runs managed orcad, decided on each connect Before a relay is started, the connect decides the host's server: a converted host connects through its tunnel (retiring a retained source once orcad-source-retirement is on), an empty host deploys orcad, and a host with Orca state converts through the journaled migration. A host with live or unproven relay terminals keeps the relay this session and converts later. A refusal for any other reason keeps the relay and names the blocker. A host orcad can't run on (unsupported target, no template, native preflight, runtime self-test) releases any claim or fence it took, records why with this app version, and keeps the pinned-relay ladder. Progress and the decision ride an optional SshConnectionState.managedServer field (dropped by older clients' admission). SSH Hosts shows each host's server status; the manual move dialog is removed. * fix(ssh): always await the connect's server decision, after the provider authority rotates The decision now runs after the old session and transport are torn down and the authority has rotated synchronously, so concurrent connects still join one attempt; a shutdown that began during the decision wins over the rotation. The IPC tests use an async double, no sync branch. * test(renderer): the IPC events store double carries its SSH connection states --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> * test(ssh): prove connect-time conversion to managed orcad on real Linux and Windows hosts (#24981) * test(ssh): prove connect-time conversion to managed orcad on real Linux and Windows hosts, and the downgrade view * fix(e2e): read the SSH host's session partition, provision the convert cell's account, and keep rollout file overrides out of packaged builds * test(e2e): report the full relay state when the convert poll times out * test(e2e): require the relay to list no terminals before the converting connect * test(e2e): require no running-terminal lease before the converting connect * test(e2e): report the host's terminal leases when the converting connect keeps the relay * test(e2e): end the relay era with no relay shell left to respawn * test(e2e): settle before the converting connect and name the tabs a blocking shell belongs to * test(e2e): use the exited relay tab as the session tab, since any mounted tab starts a shell * test(e2e): carry an editor tab through the conversion instead of a terminal tab * test(e2e): log the conversion census inputs before the converting connect * fix(orcad): log which state blocked a refused conversion * test(e2e): log both session partitions before the converting connect * fix(orcad): a source partition's copy of focus on another host no longer blocks conversion * fix(ssh): an ssh2 forward on port 0 reports the port it bound, so managed tunnels pair * test(e2e): give the conversion its full budget again * test(e2e): report the migration journal phase when the conversion stalls * test(e2e): report the connect's own result and main's state when the conversion stalls * fix(ssh): a converted host's managed state reaches the renderer instead of staying on 'converting' * test(e2e): print a failed server call's response * test(e2e): give server calls the budget a fresh server's first inventory needs * test(e2e): read the converted worktree's tabs with a scoped session.tabs.list * test(e2e): log the converted worktree's tabs instead of asserting them, pending the server-side fix * test(e2e): prove retirement by the dropped source rows; the journal compacts away after it * test(e2e): drop the conversion diagnostics now the cell passes * test(e2e): fail a hung disconnect or connect with main's state instead of the whole budget * test(e2e): convert an upgraded relay-era profile's host on its first connect, on Docker and Windows * test(e2e): seed the relay-era target the way addTarget registers it * fix(ssh): a stale ssh2 forward drops a late connection instead of crashing main on 'Not connected' * fix(orcad): the active-slot readiness probe reads Windows hosts through the host script * ci(ssh-windows): let only the convert cell's account open the SSH local forward its managed server needs * test(e2e): match server paths in their JSON-escaped form, for Windows backslashes --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): move what an older build added to a converted host, or keep the server's version (#24980) * feat(orcad): move what an older build added to a converted host, or keep the server's version A host an older build changed stays on the relay with two actions. 'Move the new projects' runs a fresh journaled conversion of only what the host's earlier migrations didn't move: the source is viewed with those migrations retired from it, so nothing that overlaps the server is merged, and a row the server already holds fails the whole move at stage, before any commit. Its journal supersedes the chain head; retained-source checks compare against its baseline, and retirement retires every manifest in the chain only after the newest committed. 'Keep the server's version' records the current source as the baseline and returns the host to its managed server. * test(ipc): the runtime environment handler contract lists the delta-move channels * test(native-chat): snapshot the journal directory after the attach's restart-offer lock is released --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): converted hosts keep their editor tabs, forwarding-refusing hosts stay on the relay, listAll settles (#25099) * fix(orcad): publish the headless graph so session.tabs.listAll settles instead of hanging * fix(ssh): a system SSH forward on port 0 picks a free port first and reports it * fix(runtime): a headless host lists and closes the editor tabs its session holds, so migrated editors reach clients * fix(ssh): keep a host that refuses TCP forwarding on the relay, and release a conversion it stranded * test(e2e): assert the migrated editor tab and a settled listAll, and keep a forwarding-refusing host on the relay * fix: restore the journal import after rebase, type the probe's failure code, and update headless-graph test seams --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): client focus and 'local'-stamped tabs no longer refuse a host's move to its managed server (#25110) * fix(orcad): client focus and 'local'-stamped tabs no longer refuse a host's move to its managed server * fix(orcad): a v1.4.218 profile focused on the SSH worktree converts, and a refusal names what blocks it The debounced session writer never patched activeWorkspaceKey or activeWorkspaceExecutionHostId, so the first focus stayed on disk. The all-dependency census also counted global focus copies in every non-source partition. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): offer to move a host whose open terminals keep it on the relay (#25097) A host with any open SSH terminal never converted: the connect gate keeps the relay while relay terminals are live, and an open tab respawns them on every connect. The first such connect per host per app version now marks the relay status with offerMove (recorded as managedServerMoveOffered), which toasts "Move <host> ... Its N open terminals will restart." The SSH Hosts status line keeps a "Move to managed server" action while terminals are live. Confirming runs ssh:moveToManagedServer: stop the host's relay terminals through the extracted ssh:terminateSessions path, re-run the connect gate's terminal census, refuse on anything but exited, then reconnect so the connect-time decision runs the journaled conversion. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): reach managed orcad over an SSH stdio bridge where sshd refuses port forwarding (#25120) * feat(orcad): reach managed orcad over an SSH stdio bridge where sshd refuses port forwarding Hosts with AllowTcpForwarding no were kept on the relay. The managed tunnel now probes forwarding each time it starts and, on refusal, serves the same local port through a second provider: each accepted socket opens one SSH exec channel running a small bridge on the host's pinned Node, which dials orcad's loopback port. POSIX hosts run it with node -e; Windows hosts run it as the content-addressed host script's stdio-bridge op with base64 line framing. Bridges are capped at 8 per connection under sshd's MaxSessions default, and a lost channel only drops its socket. Only a host where even the bridge cannot run keeps the relay, recorded as ssh_tunnel_unavailable; the older tcp_forwarding_refused record is retried. * test(e2e): prove the stdio bridge by refused forwarding plus a working call * test(e2e): connect the refused-forwarding host without the relay-only repo step --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): report each connect's server decision, and show orcad.log's last lines when setup fails (#25118) * telemetry: one enum-only event per connect decision (outcome, reason, tunnel transport, host platform, duration), plus conversion start/commit/fail, deploy failures, and the per-host move offer and its result * errors: deploy, launch and rollback failures carry the last 40 redacted lines of the host's orcad.log, read over SSH (Windows through the node.exe host script) * SSH Hosts: a deferred or failed setup offers its reason, log tail included, under Details Co-authored-by: m4air <m4air@Mac.localdomain> * fix(startup): Windows never crashes resolving userData when roaming AppData is unavailable (#25113) * fix(startup): pin Windows appData and userData before anything resolves them A Windows session without a loaded profile (e.g. orca serve over SSH) can fail the roaming AppData known-folder lookup. Electron 43 then falls through to Chromium's userData provider and crashes natively. Resolve appData first (falling back to APPDATA, then USERPROFILE\AppData\Roaming), and set userData explicitly so Electron's provider never runs. * test(startup): remove the AppData fixture through the retrying helper --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): keep a terminal the previous Orca version's relay still runs instead of replacing it (#25124) After an app update the new relay answers "not found" for a PTY the previous build's relay still runs, because the old relay refuses this build's handshake. The client read that as absence: it expired the lease and the pane cold-restored into an empty shell while the user's shell kept running, unreachable. Each deploy now takes a census of this target's older relay endpoints; while one is live or unverifiable, a not-found reattach keeps the lease and the pane binding, and the pane says the terminal is still running under the previous Orca version. A detached lease also keeps blocking managed-server conversion until that terminal exits. The cross-version harness now extracts src/relay, and a new test drives v1.4.218's relay socket and grace lifecycle with this build's endpoint probe: the probe leaves no grace deadline and reads the old relay as live work. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): update a managed server on connect when it runs an older build and is idle (#25122) * feat(orcad): update a managed server on connect when it runs an older build and is idle A connect to a managed SSH host now runs the Managed servers update when the host's orcad differs from this app's bundled build, the template carries the host's target, and the update planner finds no live or uncounted terminals. A rejected candidate is restored through the activation journal; the reason is recorded per app version so later connects don't retry it. A host a newer Orca activated is never downgraded: the activation record now names the app version behind each build, and an explicit rollback holds the build it left. * test(e2e): connect without a racing disconnect after relaunch, and report each attempt On launch the app already reaches the managed server through its tunnel; a disconnect racing that restore cancelled the connect that runs the update. * feat(orcad): run the update check when the launch restores a managed server's tunnel An auto-restored host may never see an SSH connect, so its server would never update. The tunnel restore now runs the same check, once per server per session and off the caller's path, through the shared update-check module; a server mid-migration is left alone. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): retry the pinned Node download through transient network failures (#25154) A Chromium network change (net::ERR_NETWORK_CHANGED) during the pinned Node download failed the deploy and sent the host back to the relay. The archive download now retries up to three more times, after 1s, 3s and 9s, on dropped connections, timeouts, stalls and retryable HTTP statuses, removing the partial file first; checksum mismatches, other HTTP errors and cancels stay final. The transient-error classifier moves from the speech download to src/main/network/transient-download-error.ts so both share it. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(serve): run orca serve on orcad by default, with ORCA_SERVE_RUNTIME=electron as the opt-out (#24972) * feat(serve): run orca serve on orcad by default, with ORCA_SERVE_RUNTIME=electron as the opt-out * test(serve): read Electron serve's pretty-printed readiness in the CLI mode-switch e2e * feat(serve): gate D7 on Windows and serve on orcad there by default * test(serve): take the profile lock in the CLI mode-switch e2e, and keep Windows profile logs on failure * test(serve): tell Electron and orcad serve apart by readiness health, and trace Windows startup * ci(e2e): dump Electron's native log and stack on the Windows serve mode-switch job * fix(serve): keep Windows on Electron serve until it can adopt orcad's daemon The Windows D7 job shows Electron serve exiting before its window when it relaunches onto a terminal daemon orcad forked. Restore the win32 fallback and skip that case there as a known gap; the follow-up PR fixes it and re-flips Windows. * fix(serve): let ORCA_SERVE_RUNTIME=orcad opt in on Windows while Electron stays the default * test(serve): skip the Windows D7 cases where orcad forks the daemon, and stop cleanup hiding a failed relaunch Test 3 hits the same Windows gap as test 2: orcad forks its own daemon there, and Electron crashes at startup beside it. A failed relaunch also made dispose close the old, already closed app, whose throw replaced the launch error. * test(serve): run every Windows D7 case, with the isolated home's AppData in place Electron 43 crashes natively (0xFFFF7003) when it resolves userData and Windows cannot find roaming AppData. The e2e home isolation points USERPROFILE at a fresh folder with no AppData, so later launches hit that. The harness now creates it, every D7 case runs on Windows again, and a new case proves Electron serve starts beside another profile's live orcad daemon. * test(serve): retry removing a Windows e2e profile while a killed daemon releases it --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain> * feat(ssh): resume a terminal the previous Orca version's relay still runs, through that relay's own bridge (#25170) * feat(ssh): resume a terminal the previous Orca version's relay still runs, through that relay's own bridge After an app update the previous relay keeps the user's shells alive but refuses this build's handshake. Its own relay.js --connect, run from its own version directory, presents its own bundle hash, so on POSIX hosts the client now reaches it that way: a pane whose reattach the current relay held for an older relay opens a route through the old bridge, takes the PTY owner role without output flow control, reattaches the PTY with its replay, and routes every later operation on that id to the old relay. When the last pane a route serves exits, the route hangs up and the old relay's own idle grace retires it. Windows hosts, relocated short sockets and unreachable bridges keep the held-pane behaviour. The cross-version harness now builds v1.4.218's relay from its tagged sources and runs it as the real detached daemon: a shipped client leaves a shell in it, this build resumes the pane through the old bridge, types into it, sees its output, and watches the old relay exit on its own after the shell does. * fix(ssh): read the legacy relay router through its instance --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a host re-upgraded after a downgrade can move what the older build added (#25179) The delta view left the moved projects' session state in place whenever it held anything unmovable, so it then counted against the delta. Relay PTY bindings, shutdown markers and the relay consumer's recovery record blocked every move although the terminal gate already proves those terminals exited before any move commits; the move now drops them instead. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(worktree-ps): report a lost-contact host's terminals as unverifiable, not live:0 pty:no (#25167) During a network drop to a relay-served SSH host, every PTY on it reads as an unconfirmed exit, so worktree ps skipped them and printed live:0 pty:no for a terminal that was still running. The listing now counts terminals whose liveness verdict is unverifiable into a new optional unverifiableTerminalCount, and the CLI prints live:unverifiable pty:unverifiable (JSON carries the same word) instead of zero. A host-confirmed exit still reads as zero. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(terminal): a held pane shows no client OS or shell, and the boundary doc says POSIX hosts resume it (#25194) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): tunnel to the port managed orcad bound and verify it is ours (#25182) When another Orca already listens on 6768, orcad binds a different port. The tunnel kept forwarding to 6768, reached the other runtime, was rejected with 4001, and the connect still reported a managed server. The tunnel now reads orcad's bound port from its active slot's readiness (falling back to the persisted port for slots without one) and, after the forward is up, proves the server answering is the paired runtime. On a mismatch it re-reads the port once and fails with orcad_identity_mismatch instead of reporting managed. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a converted host shows only its managed server's rows, and its old editor tabs load there (#25193) Retained relay-era project groups and the per-host SSH catalog now hide like repos and folders. On the managed transition the renderer reloads server names, groups, folders and worktrees, then drops the host's relay-era rows a local refresh would keep. A restored tab with no host stamp in a workspace now owned by a managed server takes that server as owner, in place. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * fix(orcad): ship the port-scan worker so managed servers detect workspace ports (#25197) orcad resolves port-scan-command-worker-entry.js beside orcad.js, but the orcad build never emitted it and ORCAD_ARTIFACTS never listed it, so every managed server logged 'probe worker unavailable' and had no port detection. The build now emits it with the other children, and the artifact list carries it, so it is uploaded, hashed and covered by the template contract. A new test bundles orcad and its children and fails when the bundle names a worker or child entry the slot does not ship. Two existing gaps it found, session-scanner-service-entry.js and wsl-transcript-fs-process-entry.js, are listed as known and may only shrink. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): keep retrying when a reconnect loses the socket before the SSH banner (#25195) After a network outage, a port forwarder or NAT can accept the TCP connection and then close it before the server sends its banner. ssh2 reports that as 'Connection lost before handshake' with no errno, so the reconnect ladder classified it as permanent and published 'error' with no retry scheduled. Treat it as recoverable on the bounded ladder only; the initial connect keeps its narrow classifier. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(serve): serve on orcad by default on Windows too (#25162) D7 now runs every case on Windows (orcad-serve-mode-switch-windows), including Electron serve adopting a daemon orcad forked, so Windows no longer needs the Electron default. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): the terminal gate asks the relays before treating a detached terminal as running (#25200) * fix(orcad): the terminal gate asks the relays before treating a detached terminal as running, and retiring a host drops its relay recovery record * fix(ssh): earlier-relay census gaps and an unanswered relay stay unverifiable; asking the terminal gate changes nothing - the gate is read-only; the conversion and delta move retire proven detached leases themselves - an expired lease also needs every earlier-build relay to answer before it reads as exited - a failed, input-less or truncated census marks its older relays unverifiable, so reattach holds - a legacy relay route closes only when no attach or listing still awaits it - a disposed session forgets the census it started --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): recover interrupted conversions and delta moves, and trim dead migration code (#25226) * refactor(migration): drop the unused full dependency census * refactor(migration): one table for the routed UI fields a migration carries * refactor(migration): share the session focus field list and fix stale claim comments * fix(session): keep a fenced SSH host's source session partition intact through renderer saves The renderer cannot see a fenced host's repos or folders, so hydration drops their worktree rows; a save that still routes any row to ssh:<target> (a folder workspace key stays valid) rewrote that partition without them. That changes the migration's source manifest and loses the session a downgraded build reads back. Main now ignores renderer writes to a fenced, unchanged host's partition. * fix(migration): resume or back out unfinished conversions and delta moves, serialize delta moves, clean stale journals - On connect, a registered server whose conversion never committed resumes the commit; a failure aborts it through the destination, unregisters the server and releases the fence. - An interrupted delta move keeps its journal and gets its changed mark back; the next move resumes it from the journaled manifest or aborts it when the source has moved on. - Delta moves run under the target lifecycle queue and re-check the head, so two concurrent moves cannot write two journal heads. - Delta checks re-read the live source instead of a frozen copy. - Stale journals from a stopped or never-registered server are cleared. - Retiring a chain goes newest first and compacts only at the end. - The journal schema tolerates fields a newer build adds. * fix(migration): hide source rows only after commit, and retire what an older build added to a moved project - Source rows and the renderer session guard share one rule: hidden once the server committed, shown while a migration is still in flight. - A delta move a crash interrupted gets its changed mark back on the next start. - Retirement removes worktree metadata and automations by moved-project scope, so an older build's additions no longer fail the leftover check forever; that check ignores rows of other migrations in the chain. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): survive a slow first daemon start and close review gaps (#25213) - Give the terminal daemon 30 s to start on Windows and retry the spawn once before falling back, so a cold first deploy is not refused as daemonless. - Say in the activation refusal that orcad.log holds the daemon's error and that the next connect retries. - Managed stop completion now waits out daemon retirement plus the shutdown deadline. - Publish the instance lock atomically and reclaim an abandoned torn lock. - Fix the stop listener closing before it was defined on an already-present request. - `orca serve`'s cache prune keeps the slot a running local orcad uses. - A timed-out daemon retirement reopens admission once it is refused; a retirement that may have reached the daemon stays fenced. - A headless host keeps an editor tab's unsaved draft unless the close is forced. - Unverifiable terminals are attributed by the host's PTY record, like live ones. - An explicit --user-data-dir is no longer overridden on Windows. - Read Windows daemon process identity from the process table, not PowerShell. - Remove the unread data-event incarnationId and daemon health runtime fields, share one process-alive and error-code check, and fold the websocket limits file back. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(cli): stop, update, roll back and recover managed Orca servers from the CLI (#25201) * feat(cli): stop, update, roll back and recover managed Orca servers from the CLI orca environment status|update|rollback|recover|stop|cancel-stop call the same managed-server actions as Settings > Managed servers, over new managedServer.* runtime RPC methods. The desktop main process registers those actions; the runtime advertises managedServer.v1 only then, and the CLI refuses on a runtime without it or one that answers method_not_found. stop requires --yes. * fix(ipc): keep a missing managed-server selector a rejection, not a synchronous throw --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): Move to managed server converts the host it is offered for (#25196) The move stops the relay terminals and passes the terminal census; with #25179 the conversion no longer refuses on the stopped terminal's saved tab, layout leaf and pane incarnation. A new test drives one live relay terminal through Move to a conversion whose manifest carries the tab without its relay PTY, so it spawns a fresh shell on orcad. ssh:terminateSessions also no longer records a not-found shutdown as terminated while an older Orca build's relay may still run the PTY (#25124): that terminal is reported unverifiable and its lease kept, so the move refuses instead of converting over a running shell. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): deploy the glibc 2.17 runtime on old-glibc hosts, and fall back to the pinned relay (#25180) * fix(orcad): deploy the glibc 2.17 runtime on old-glibc hosts, and fall back to the pinned relay Managed orcad picked its runtime from the host's libc flavour alone, so a CentOS 7 host (glibc 2.17) got the default linux-x64-glibc Node, whose self-test fails there. The deploy now picks the runtime by glibc the same way the relay ladder picks rung A or B, and the host-side slot checks accept the compat target it ships. When orcad still can't run, the host is recorded as such and the relay connect that follows now runs the pinned-runtime ladder instead of defaulting to the host-Node relay, so a host with no Node lands on rung B rather than failing. The hostile-host lane deploys managed orcad on a fresh CentOS 7 host and asserts it runs on the compat runtime with no default runtime uploaded. * test(ssh): CentOS 7 cell asserts the compat runtime pick, then the refusal and relay rung B fallback The compat template still ships the base @parcel/watcher binary, which needs GLIBCXX_3.4.20; CentOS 7's libstdc++ stops at 3.4.19, so the candidate's preflight refuses. The cell now pins that end-to-end behaviour: compat runtime picked and uploaded alone, refusal classified native_preflight, and the relay that follows settles on rung B with no runtime setting. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(serve): prove D7 on an installed Windows app whose daemon runs from the relocated daemon-host (#24976) * test(serve): prove D7 on an installed Windows app whose daemon runs from the relocated daemon-host * test(serve): read the pre-switch scrollback best-effort and wait for it after reattach * test(serve): opt the packaged Windows serve switch into orcad explicitly --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): a retirement that fails after deleting rows resumes instead of reading as changed (#25234) The chain head records sourceRetiringAt before any row is retired; the startup change check and the chain's own comparison skip a head that carries it. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): managed orcad idles out after 15 minutes and is started again whenever it is down (#25121) * feat(orcad): a managed orcad stops after 15 idle minutes and starts again on the next connect A client-launched orcad now exits, like the relay, once no client, terminal, working agent, staged migration or activation fence has been seen for 15 minutes. The exit is the normal graceful shutdown, which leaves the terminal daemon running; the daemon retires only if it proves itself empty. A record in the data root tells the next start (and its readiness health) that the stop was an idle one rather than a crash. On connect and after host resume, a fresh tunnel whose server does not answer starts the activated slot under the activation fence, but only on a proven exit, so a stopped server reads as not running rather than a failure. * fix(orcad): keep orcad-entry under max-lines; idle e2e connects without a relay repo * fix(orcad): deploy and rollback launches carry the managed idle-exit fence The candidate launch in activation and the rollback launch built their own launch spec without the activation root, so a freshly deployed orcad never enabled idle exit; only the wake path did. The field is now required on every launch spec, so the type system covers each launch site. * feat(ssh): start a stopped managed orcad on connect, on restore and after resume Once orcad stopped (idle, kill or host reboot), a connect still resolved managed over a forward to a dead port and every call failed. Every connect now checks the server behind its tunnel, as does a call through a restored environment; a server proven stopped is started from its activated slot under the activation fence, adopting a surviving daemon and its terminals. The status line shows the start, and a start that fails keeps the host managed with the reason and orcad.log's tail, never as a terminal verdict. * fix(ssh): reuse a serving verdict only on the same SSH transport, for 5s A reconnect right after a reboot was answered from the previous transport's cached verdict, so the stopped server was never started. * fix(ssh): key the serving verdict on the tunnel's remote port too * test(ssh): a stopped server starts before the update counts its terminals * feat(ssh): check serving at the bound port, and follow a restarted orcad to a new one The serving check uses the port the tunnel forwards to (the one orcad bound). A restart that binds a different port drops the forward and rebuilds it at the new port, within the same ensure or on the explicit connect check. The tunnel manager class moves to its own file to stay under max-lines. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a connect with no leases asks the host's relay endpoints before converting (#25223) * fix(ssh): a connect with no leases and no relay session asks the host's relay endpoints before converting * refactor(ssh): one isLiveSshPtyLease for every lease-liveness check * fix(ssh): the host relay census asks a relay its PTYs before reading it as live work An accepting relay whose holders or children the probe could not read (no lsof, an unrecognised service child) read as live work, so a relay-era host that had exited its last shell never converted. The relay's own bridge answers pty.listProcesses without the owner role. * fix(ssh): the host relay census runs each relay's probe and bridge on the runtime it runs on Pinned-ladder hosts often have no Node on PATH, so a PATH Node read every relay as unverifiable. Each daemon's own argv names its pinned runtime, or the host Node a legacy relay started with. * fix(ssh): an older relay's bridge runs on that relay's own runtime --------- Co-authored-by: m4air <m4air@Mac.localdomain> * feat(orcad): build @parcel/watcher into the glibc 2.17 compat slot, so CentOS 7 runs managed orcad (#25199) The compat target swapped in only node-pty, so it shipped the base target's upstream watcher.node, which needs GLIBCXX_3.4.20; CentOS 7 stops at 3.4.19. orcad's preflight refused it, and relay rung B lost file watching without saying so. The compat slot now compiles @parcel/watcher from the package's own sources against the pinned headers with the C++ runtime static, like node-pty. The slot gates (glibc 2.17 symbol floor, no shared libstdc++, N-API 8) and the smoke load cover it, the template stages it into the compat target, and both orcad and relay rung B pick it up from there. COMPAT_SLOT_ADDONS names a compat slot's own addons: a compat slot missing one fails --require-slots, and a compat template target that would ship any native file without a compat build fails the template build. The CentOS 7 cell expects an activated managed server again, and every launched relay cell loads its watcher directly. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(phase3): close e2e routing holes, localize managed-server outcomes, trim harness and CI (#25206) * fix(phase3): close e2e routing holes, localize managed-server outcomes, trim harness and CI - Route the auto-convert spec into needs_build and route the new orcad/serve e2e helpers to the specs that use them, with routing tests. - Run the auto-convert Docker lane only when routed; fold the missing-AppData check into the Windows mode-switch job and drop crash-hunt diagnostic env. - Delta-move dialog keys its preview on the target id and cannot close or resubmit mid-move; Resume has an in-flight guard. - Managed-server toasts and status lines show localized messages instead of raw codes or main-process English; add singular and zero-count wording. - Settings style fixes (quiet Cancel, section header, labelled fields, progress labels); delete dead i18n keys and the unused previewConversion. - Harness: kill the serve child on readiness timeout, share the isolated profile and spawn-until-ready helpers, hooks and a condition wait in the auto-convert spec, shared cross-version exec helper. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(phase3): localize the refusal toast, share serve readiness and liveness helpers, correct D7 docs - The refusal toast no longer shows main's English blocker detail; the SSH Hosts status line keeps it under Details. A missing terminal count is no longer defaulted to 0. - startOrcadServe uses spawnUntilReady, so a readiness timeout kills orcad; one isPidAlive replaces the spec's copy and the lock-holder loop. - The port-6768 auto-convert test uses the shared hooks. - The docs say the Windows D7 job checks rather than gates. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ci): give the idle-exit spec its app build and route the convert harness to it Co-authored-by: m4air <m4air@Mac.localdomain> --------- Co-authored-by: m4air <m4air@Mac.localdomain> * ci(adhoc): pin every job of an adhoc build to the commit resolved at dispatch (#25311) Each job read the requested branch name and checked out whatever it pointed at when that job started. A push mid-run mixed commits: in run 37234548638 the glibc217 slot lane built2ddea8736e, then #25199 merged, and desktop_template checked out70948d597f, whose merge step requires the compat watcher that lane never built. A first resolve-ref job (no secrets) resolves the branch, tag or full SHA once, and every job checks out that commit. The mac job still vets it for reachability before signing; the requested name stays the concurrency key and the release label. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(terminal): report a confirmed relay PTY exit as exited and give a cold Windows PTY probe more time (#25304) A worktree terminal close returned as soon as the relay confirmed the stop, but the exit frame reaches the runtime record only after the SSH output intake drains, so the verdict read straight after saw a still-connected PTY and answered unverifiable. The stop now waits, bounded by its deadline or 10 s, for that record. The bundled runtime's PTY probe gives the first Windows spawn 15 s and retries once after a timeout only; spawn errors and non-zero exits still fail at once. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(relay,orcad): keep new wire messages forward compatible, drop unshipped relay.reset, fix shutdown and serve gaps (#25298) * fix(orcad-migration): keep migration wire forward compatible with newer peers * docs(pty): record why pty.resumeClient negotiates by method-not-found * fix(relay): remove the unused relay.reset method and keep a deferred shutdown serving * refactor: drop dead hold API, release gate and migration pass-throughs; fix serve and delegation gaps * fix(lint): keep reopen hooks within file size limits * revert type dedupe in orcad-incumbent-recovery to avoid a parallel conflict * fix(lint): name stop-reply request fields for their role --------- Co-authored-by: m4air <m4air@Mac.localdomain> * refactor(ssh): one managed-tunnel ownership check and forward bookkeeping (#25316) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a stuck managed server can recover, cancels finish their change, unknown stays unknown; delete unused reset/drain code (#25227) * refactor(ssh): delete the unused connection reset and drain machinery * fix(ssh): surface a stuck managed server, finish activations past their first change, and fall back from Windows launch refusals - A rejected build that changed profile state now refuses as orcad_recovery_changed_state (unverifiable) with a Recover path that restores the snapshot once the operator accepts; an interrupted activation is no longer reported as a quiet update deferral. - Activation and rollback drop the abort signal after their first journaled change. - Windows launch refusals and unsafe command lines send the host back to the relay. - Fence refusals keep unverifiable blockers unverifiable; the POSIX liveness probe reads kill errors in the C locale and treats permission errors as unknown. - A host with no relay fallback surfaces its real connect error; disconnect and removal clear the setting-up status, and a cancelled decision's progress is dropped. - Remove tcp_forwarding_refused leftovers, dedupe incumbent stop/restore, exec-or-empty, census client, errorMessage and the blocking-blocker predicate; split activation and snapshot files under 300 lines. * fix(ssh): one import per module in the rollback transition --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): committed reads trust the receipt, conversions survive UI churn, journals are cached, and round-2 low items (#25327) * fix(ssh): keep a newer build's fence and note fields, and validate a fence beside a legacy owner Normalization now validates known fields but passes unknown ones through on orcadFence and the managed-server notes. A legacy managed owner next to a malformed fence falls back to the owner's environment id instead of keeping the malformed fence. * fix(migration): delete a retired migration's scrollback files once retirement is durable Retirement dropped the moved terminals' scrollback refs from session state but kept the files forever. After the journal records source-retired, the files the manifest names are deleted, except refs any session partition or pending export still names. A crash before the delete repeats it on the next retirement pass. * fix(orcad): a committed migration reads as committed from its receipt, survives receipt eviction, and abandoned stages expire - Committed-state reads trusted only a byte-identical copy of every moved row and snapshot, so a live server that had been used could never confirm its own commit. The receipt alone now proves it; full equality stays inside the commit. - A receipt that ages out of the 64-entry list keeps a compact record, so the commit never reads as absent. - A stage no client returns to within a week stops holding the server awake and is dropped at the next stage. * fix(migration): freeze a host's session from the fence on, ignore UI churn in the frozen-source check, and cache parsed journals - Renderer writes to a fenced host's ssh: partition are skipped from fence time, not only after commit, so tab work mid-conversion cannot change the source. - The frozen-source check leaves out workspace session and client routing state, which the UI rewrites as the user works (including through the local partition and UI state); the server gets them as journaled. - Parsed journals are cached per directory, keyed by each file's inode, size and mtime and dropped on every write and remove, so list and session calls stop re-parsing every manifest. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): an older relay a census cannot rule out stays held, and a reconnect resumes held PTYs through it (#25331) - A census that does not know the host platform is unverifiable, never "no older relay". On Windows it now probes each older version directory's pipe for the target, the way relay GC does, so a live older Windows relay keeps its PTYs held instead of respawning their panes; those endpoints are held, never bridged. - On reconnect, a PTY the current relay disowned while an older relay holds it is reattached through that relay's own bridge, with its runtime restored and its replay forwarded. When no route serves it, it is left for recovery like an exhausted reattach. - A superseded or disposed deploy no longer starts the census, so it cannot replace the current attempt's. - A route stops holding an unserved PTY that exits. - Drop the unused censusPreviousRelays export. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(phase3): round-2 ui-infra review fixes (#25325) - An unverifiable move refusal with no count says so instead of "0 terminals". - One move per host: the dialog stays mounted while open, joins a run already in flight on remount, and the toast shares the same guard. - A managed server's start, wake or update no longer reloads every host's catalog; only a new environment for the host loads, scoped to that host and the local catalog. - A failed host-partition session write is no longer acknowledged as written, so the writer re-queues those fields. - The deploy picker lists only hosts with no managed server or pending move. - Harness: orcad and the released relay daemon are stopped when startup fails. - Windows host CI: the convert cell always runs last and fails on a failed native switch or build; ssh-host-server unit tests no longer trigger it. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(relay,cli): never report an unreachable terminal as exited; keep PTYs through a deferred shutdown (#25328) * fix(relay,cli): never read an unreachable relay or terminal as exited; keep PTYs on a deferred shutdown * fix: stop an unrecorded breakaway child, route orcad serve through the spawn chokepoint, and drop review-flagged leftovers --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): reclaim a crashed desktop lock, close the stop-cancel race, drop dead daemon and RPC code (#25338) * fix(orcad): reclaim a crashed desktop lock, close the stop-cancel race, drop dead daemon and RPC code - A desktop reclaims a stale desktop lock record (Electron's lock proves it), and a real orcad hold shows a dialog instead of exiting silently. - A failing quit handler no longer skips closing observability. - Completion withdraws its request when a cancel lands mid-write; the listener removes a cancelled leftover. - An idle stop's clean record is retracted when the shutdown fails or overruns. - The managed-stop request tolerates unknown fields from a newer client. - Shared error-code checks, one win32 coverage rule, no redundant isAlive filters. - Remove the recovery-only daemon provider, requirePinnedWsPort/strictPort and the superseded orcad.migration.importCatalog RPC. * fix(startup): keep main-process-preflight under the line cap --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): fail fast on a held fence, restore untouched rollbacks, atomic Windows host script, honest tunnel ensure (#25332) * fix(ssh): fail fast on a held activation fence, restore a rollback the target never touched, stage the Windows host script atomically - withOrcadActivationLock no longer waits up to 15 minutes inside the target lifecycle: a held fence answers at once (orcad_activation_recovery_required); deploy probes it before uploading. - A rejected rollback target that left the restored snapshot untouched puts the newer build back unasked; only a real change waits for the operator. Crash recovery applies the same rule. - The Windows host script is written only when missing, through a partial file and a rename; a bridge exit without a sentinel is unverifiable unless the shell could not find the command. - The managed tunnel's ensure() throws when superseded, and a caller arriving after close() builds a fresh run instead of joining the doomed one. - An unparseable stop answer after a stop was sent keeps the fence (new 'unconfirmed' outcome). - stopRemote starts over instead of reporting live when another run settled the journal. - Releasing an unreachable setup stops the orcad it activated and clears active on proven exit. - A failed startup is torn down quietly so its error stays published; port-forward listeners keep an error handler; a linked SSH access connect is cancelled when its window closes. - Shared errorMessage, one SFTP transfer helper, getConnectGeneration only, test-cell names. * fix(ssh): stage the Windows host script through the pinned node.exe, which cmd.exe and PowerShell both run * fix(ssh): host-script staging runs as host-script ops; a joiner of an overtaken tunnel run builds its own - The presence check and the install are ops in the uploaded host script (script-present, and script-install run from the partial upload itself), invoked as node.exe <script> <op> <args> like every other Windows host op: no inline code. - ensure() callers that joined an in-flight run no longer inherit its 'superseded' end; they start a fresh run, which also covers a close() in between. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a connect re-checks unverifiable relay terminals once its relay session can answer (#25385) The server decision runs before any relay session exists, so a terminal this desktop left detached could never be asked about and read unverifiable for as long as it ran. On Windows no endpoint census can fill that gap. Once the session is up, the relay that holds the PTY answers and a terminal it still runs is reported live. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(settings): a managed server whose status never loaded reads unknown, not "Not running" (#25389) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(relay,runtime-env,cli): keep a deferred relay's AI Vault and skill uploads, make managed re-pair crash-safe, show unverifiable in worktree ps text (#25393) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): keep a surviving daemon's slot through the cache prune, and drop an idle record a signal stop took over (#25387) - The slot prune also protects any slot a live terminal daemon's PID record points into, and evicts nothing while a daemon record is unreadable. - The shutdown trigger reports whether it took ownership; an idle stop that another source took over discards its clean-idle record. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): edit a managed host's connection, show rows no migration owns, drop the unused quit-drain predicate (#25403) - SSH settings may edit a managed host's connection fields; its fence and generation are never written from the renderer, and its tunnel is closed so the next use redials. Removing it is refused with a pointer to Stop under Managed servers, which stops and removes the server; the renderer no longer ends the host's terminals before that refusal. - A fenced host with no journal (an empty host's deploy, or one whose journal compacted after retirement) no longer hides its rows: no committed move owns them, so projects an older build added stay visible. An unreadable journal still hides them. - The quit drain's mayDetach predicate and disconnectAll's filter had no production caller and are removed. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): input to a terminal an older relay holds but no route serves is refused, not dropped (#25404) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): refresh a changed host's status once its delta move or keep-server choice lands (#25416) The status line kept saying the host was changed on an older Orca until a manual reconnect. Main now publishes the host as managed by its server after a successful move or keep, with a disconnected state when the move released the relay session. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a stopped managed server stops showing on its SSH host without a reconnect (#25409) * fix(ssh): a stopped managed server stops showing on its SSH host without a reconnect On unlink, forget the host's managed-server decision and republish its connection state. The republish also drops the host's cached worktree scans, so worktree.ps stops naming the removed server. * fix(runtime-environments): a removed server's workspace session partition goes with it Host-scoped listings enumerate session partitions as known hosts, so worktree ps kept naming a stopped or removed server as an omitted, unselectable host. Unlinking or removing a server now drops its runtime:<id> partition. * test: give the removal-storage fake store casts their SAFETY rationale * fix(runtime-environments): drop orphaned runtime workspace sessions at startup A crash between unlinking a server and dropping its session, or a build that unlinked before the drop existed, left a runtime:<id> session listings name as an unselectable host. Startup now drops runtime sessions whose server is not in the environment store; it never touches local or ssh sessions, and skips entirely when that store is missing or unreadable. * test(migration): retirement re-aims focus without creating the destination's session Locks in the ordering the startup reconcile relies on: a runtime:<id> session is never written before that server is registered. * docs(runtime-environments): name the ordering invariant the startup reconcile relies on --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a connect whose relays prove its leases ended retires them, so the next connect converts (#25405) Every connect decides before a relay session exists, so a detached or expired lease reads unverifiable there. The post-session re-check proved such leases ended but left them in place, so the host stayed unverifiable on every connect and each respawned pane added another lease. Co-authored-by: m4air <m4air@Mac.localdomain> * ci(ssh-windows): check out preload and renderer for the orcad-convert e2e build Main's #25359 narrowed this workflow's checkout to the server and test trees. Phase 3's orcad-convert cell builds the full e2e app with electron-vite, which also needs src/preload and src/renderer, so the x64 inbox cell failed with UNRESOLVED_ENTRY. * test(orcad): model a really converted host in the v1.4.218 downgrade wire test #25403 shows a fenced host's rows when no migration journal explains the fence. The downgrade test fenced the host without a journal, so this build showed the retained project it is meant to hide. Write the destination-committed cutover journal a real conversion leaves. * fix(ssh): a briefly held fence waits and is retried, not recorded as a failed update; terminate keeps held leases (#25420) - The activation fence is retried for a few seconds. A fence still held answers orcad_activation_fence_busy (a waiting deferral, never an update failure) unless a journal or a lock past its stale age shows an interrupted run, which stays recovery-required. - acquireInstallLock reports Busy only when a holder answered; a lock command that only ever failed surfaces its own error. - Terminate detaches instead of disposing when any PTY was unverifiable, so the final teardown never bulk-marks a lease an older relay may hold as terminated. - A wake whose connection dropped while holding the fence releases that fence on this client's next wake (no journal, slot proven exited), so a relaunch-then-connect is not left fenced. - The orcad e2e reconnect helper surfaces a connect's error text instead of a JSON parse error. Co-authored-by: m4air <m4air@Mac.localdomain> * ci(ssh-windows): check out all of src for the orcad-convert cell Its e2e global setup also compiles the bundled CLI from src/cli, which the narrowed checkout left out. * fix(orcad): idle stop — drop the record when a signal stop wins, and read the activation lock, not its root (#25464) * fix(orcad): drop the idle-stop record when a signal stop finishes first Every stop's cleanup now discards the record unless the idle trigger owns the stop, so a takeover that exits before the idle request runs no longer leaves a false idle stop. * fix(orcad): idle check reads the activation lock, not the transaction root An interrupted acquire can leave the root without a lock; the client already treats that as unfenced, and managed orcad now does too instead of never idling out. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): main refuses a managed server host's removal before ending its terminals (#25465) The remove flow skipped ending terminals for a managed host only when the renderer's cached target list already showed the fence, so a fence that landed after the list loaded still lost the host's terminals before main refused the removal. The removal's terminate call now carries forRemoval, and main refuses it for a managed host with the Stop… message before touching anything; the renderer no longer decides. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(terminal): closing an SSH workspace counts a host-confirmed terminal exit as stopped (#25479) * fix(terminal): a workspace close counts an SSH terminal's confirmed exit as stopped The close's verdict treated any SSH terminal whose record was still present as unconfirmed, even when that record held a host-confirmed exit. SSH records outlive their exit, so every successful close of a relay terminal answered unverifiable. Also, a relay reattach that finishes registering after the stop no longer revives an incarnation whose exit is already recorded. * test(terminal): cover a reattached relay exit confirmed without an incarnation --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(cli,ssh): managed-server actions outwait their own deadlines; cap the unverifiable serving detail (#25473) orca environment update/rollback/recover/stop/cancel-stop waited the 60 s RPC default while the desktop runs the whole action inline, so a slow host printed a timeout failure for an action that kept going. They now wait 20 minutes, and a timeout says the action may still be running and points at `orca environment status` instead of reporting failure. status keeps the default. The retained managed-server state now clamps serving.detail to the same byte limit as its sibling detail fields. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): keep the previous orcad.log on Windows across a restart (#25481) The Windows breakaway launcher's addon recreates orcad.log on every start, while POSIX appends, so a crash's log was gone once orcad restarted. The orcad launch now asks the launcher to move the last run's log to orcad.log.1 first, capped to its last 1 MiB. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(sidebar): a stopped managed server leaves no empty project group behind (#25488) The removed-runtime purge dropped the server's repos, setups and worktree rows but kept the project groups and folder workspaces the renderer fetched from it, so an empty heading stayed in the sidebar. It now drops those runtime-stamped rows too. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): run automations on orcad and keep a managed host up while they can fire (#25475) * fix(orcad): run automations on orcad and keep a managed host up while they can fire orcad never built an AutomationService, so with orca serve defaulting to orcad scheduled runs never dispatched and Run now threw runtime_unavailable. The headless service setup moves out of Electron startup into automations/runtime-automation-service.ts; orcad installs, binds, starts and stops it, and managed idle exit counts an enabled schedule or an unsettled run as busy. * docs(orcad): list automations among the idle-exit conditions * fix(automations): precheck reads the SSH manager from its registry, keeping electron out of orcad The precheck imported getSshConnectionManager through ipc/ssh, which pulled 32 electron modules and node:sqlite into the orcad bundle and failed build:orcad. * test(orcad): stub the automation wiring in the push-startup runtime harness That harness stubs OrcaRuntimeService without an automation surface, so starting the real service threw setAutomationService is not a function and timed out the next test. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): Move to managed server survives a relay that hangs up on its last terminal exit (#25487) BUG-15: with two live terminals, one served through an older build's relay, Move failed with "Failed to terminate SSH host sessions: …: Multiplexer disposed" although both shells died. The old relay reports the exit before the shutdown reply; that exit closes the route (its last served PTY), and the disposed mux rejected the shutdown still awaiting its reply. - The legacy relay route settles a shutdown whose PTY exit it already observed. - Move no longer aborts on a failed stop: the terminal census (what the conversion trusts) decides. Exited closes the relay session and converts; live or unverifiable refuses and republishes the relay status, so the stale terminal count is replaced. - The move dialog offers Try again after a refusal or failure. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): closing a terminal a previous Orca relay runs stops it there and confirms the exit (#25471) A stop on a PTY an older relay runs reached the current relay whenever no route served it at that moment (after a reconnect with the pane unmounted, or a shell no pane ever resumed), and the current relay answers a stop for an id it never minted as done, so the PTY read stopped while its shell kept running. A stop on a served PTY failed instead: its exit arrived before the stop's reply, closing the route under the pending request. - A served PTY stops on its route; a PTY no route serves is stopped through a short-lived route to the older relay that lists it, which hangs up once that PTY exits. - When an older relay may hold the PTY but cannot be asked (incomplete census, Windows pipe, a bridge that will not open), or its bridge drops mid-stop, the stop is unverifiable, never reported done; terminate keeps the lease. - A route stays open until its in-flight requests settle. - The provider's exit stream includes the exits older relays report, so a stop observes the PTY it stopped exit on the relay that ran it. The cross-version harness runs what terminal close --all runs per PTY against a real v1.4.218 relay, for a pane resumed this connection and for a shell no pane resumed, and sees the old shell exit there and the old relay retire on its own grace. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a live run's journal reads busy, a wake releases only its own fence, a cancelled re-check never reports connected (#25470) * fix(ssh): a live run's journal reads busy, a wake releases only its own fence, a cancelled re-check never reports connected - A held fence asks for Recover only when its lock is stale or a journal has no fence over it; a journal under a fresh fence is a run still working and reads orcad_activation_fence_busy. - A wake writes an owner token into the fence it takes and later releases only a fence carrying that token; observing the fence gone forgets it. - The connect re-checks ownership after the relay-terminal re-check, and the re-check itself neither records a decision nor retires leases for a cancelled attempt. * fix(ssh): a tunnel caller that joined a run a disconnect cancelled builds its own The launch-time restore's tunnel run connects over SSH; a disconnect then cancels that connect. An explicit connect that had joined the run inherited its SshConnectAttemptCancelledError and failed (seen as the idle-exit e2e's 'connect threw: ... cancelled'). A joiner now builds once anew after any end of the joined run except an auth failure, which it shares rather than prompt again. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(runtime-env): a late reply from a replaced pairing never overwrites the re-paired device identity (#25490) * fix(runtime-env): a reply from a replaced pairing never overwrites the re-paired device identity * fix(types): narrow identity fields in markEnvironmentUsed * fix(ssh): managed tunnel proves its server by runtime id; SSH access linking keeps the strict device check --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): retained scrollback survives a closed tab, local acknowledgements stop blocking, close-intent retirement retries, mobile selections merge per workspace (#25508) - A transfer reads a retained snapshot straight from storage once its tab closes, and releasing the retention deletes a ref no session names; the frozen-source check leaves the snapshot list to the journaled manifest. - Only acknowledgements on panes the source host owns count toward the ui-routing blocker. - Retiring close intents treats an absent source with an identical destination entry as already done. - Importing a device's mobile selections keeps its selections for other workspaces. Co-authored-by: m4air <m4air@Mac.localdomain> * test(ssh): an expired lease an older relay still lists is never retired (#25449) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): keep each host's state its own, count only saved commits, and never wedge connect on a partial session (#25516) * fix(migration): a host-qualified owner key belongs only to its own host Owner matching stripped a key's host qualifier before matching its repo id, so converting host A claimed, moved and on retirement removed host B's session state when the two hosts share a repo id, and destination-qualified focus written by retarget read as leftover source state, failing retirement with orcad_migration_source_ui_routing_reappeared. A qualified key now matches only its own host, and an unqualified key in another host's session partition belongs to that host. * fix(migration): a partially written session partition never wedges connect, and a marker two partitions agree on stops blocking the move Real-host BUG-14: the renderer's per-host snapshot leaves out maps a host has no rows in, and main stored host partitions exactly as sent, so a runtime partition lacked tabsByWorktree and the dormant-state collector threw on every connect. Main now fills the required maps on every host-partition write, the migration collectors tolerate a partition persisted without them, and an unreadable session blocks the move instead of failing connect. The same profile's workspace-session blocker was a false positive: the local and host partitions both carried defaultTerminalTabsApplied for the moved worktree with the same value, and the fragment merge refused any shared worktree key. It now refuses only when the partitions disagree. * fix(migration): only a flushed commit acknowledgement moves a migration to committed A committed state read may come from a receipt the server holds in memory but failed to flush. A retry took that read as proof, journaled destination-committed and went on to retire the source, so a later server restart could lose the catalog on both sides. A committed read in stage and in abort is now confirmed through the idempotent commit(), which flushes before it answers; a failure leaves the journal and the fence where they were. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): relay shells no lease here knows count as live, Move stops them, and Windows asks every relay pipe (#25518) * fix(ssh): a relay shell no lease here knows counts as live, and terminate stops it A CLI-created terminal has no lease, so with an attached lease the gate answered from leases alone, a live decision was never re-counted once the relay could answer, and terminate stopped only the shells it held leases or panes for. The relay's own listing is the authority on what runs. * fix(ssh): a Windows connect asks every relay version's pipe for its PTYs before converting Windows pipes cannot be listed, so the connect-time census answered 'unenumerable' and a shell no lease here knew let the host convert under it. Each version directory's pipe for this target is derived from its path, so the census probes them all, current included, and asks a live one through its own bridge; a live pipe it cannot ask is unverifiable. * fix(ssh): the terminal gate counts what earlier relays still run, leased or not After an app update a shell a respawn superseded on its tab keeps running on the previous relay with no live lease here, so a decision counted only the leased shells and Move could not see it. * fix(ssh): a Windows relay folder with a live pipe it cannot ask stays unverifiable Each pipe is probed and asked on its own, so one that answered with no PTYs can no longer stand in for a live sibling the census could not reach. * fix(ssh): terminate also stops shells only an earlier relay lists provider.shutdown routes a held id to the older relay that runs it, so the terminate set now takes listPreviousRelayPtyIds too; one it cannot reach is reported unverifiable as before. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a refused Move to managed server reconnects the host on its relay (#25543) Stopping the terminals closes the relay session, so a move the census refused (another desktop's terminals, or an older relay it can't rule out) left the host and its workspaces disconnected until a manual Connect. The refusal now reconnects the host; the connect-time decision reads the same census and keeps the relay. A reconnect that converts after all reports the move; a reconnect that fails still reports the refusal. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(relay): an agent exec ends on its child's exit, not on pipes a background process still holds (#25544) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): stop the automation scheduler first when a managed stop is dispatched (#25548) A dispatch could otherwise race the daemon retirement census or write a run record that a rollback restore then silently discards. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): count a degraded daemon's in-process terminals in the terminal census (#25545) In degraded mode fresh terminals run on the local fallback inside orcad, but the census read only daemon adapters, so an update or stop saw 0 live sessions and killed running agents. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a retained source whose drafts or settings an older build changed is never hidden as unchanged or retired (#25549) * fix(orcad): a retained source whose drafts or settings an older build changed is never hidden as unchanged or retired The retained-source fingerprint covered only repo, folder and group identity, so an unsaved draft edited on an older build read unchanged: the host kept serving the server's older draft and retirement deleted the newer one. Retention now also records a versioned fingerprint of the source's drafts and user-authored names and settings; a mismatch marks the host changed, and a journal without one is never retired automatically. * fix(orcad): a retained source's automations are part of its state fingerprint An older build can edit an automation the source keeps; retirement would delete that edit as if the server held it. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a rollback restores the older snapshot only behind orcad's terminal barrier (#25551) The rollback's census is taken while orcad still admits work, so a terminal or automation that starts before the stop had its state wiped while its PTY survived. The incumbent is now stopped through its managed stop with idle-daemon retirement: orcad closes terminal admission on every daemon generation, counts live sessions under that fence, and retires the daemon only when none exist. Only 'retired' lets the older snapshot replace state; live, unverifiable or a missing answer refuses and relaunches the incumbent on its untouched state. A build without managed stop is refused before anything changes. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): failed managed setup leaves 'connecting', edited managed host redials, idled-out server starts before its census (#25474) * fix(ssh): a failed managed setup leaves 'connecting', an edited managed host redials, an idled-out server starts before its census - doConnect publishes the error and clears the 'setting up' status when the managed-server decision throws for a still-current attempt; a cancelled one still reports cancellation. - Editing a managed host's connection fields closes its tunnel and disconnects its transport, serialized with the target's lifecycle, so the next use dials the edited target. - The terminal census starts a server that idled out behind a forward still up, so Stop, Update, Rollback and status no longer refuse with 'census unavailable' on every retry. * fix(ssh): a fenced failed setup publishes its cause, never-launched slots are collectable, a reused PID is not orcad - doConnect publishes the relay decision's setup failure (and clears 'setting up') when a failed managed setup kept the host fenced, instead of throwing a bare 'serves a managed server'. - The liveness probe answers NEVER_LAUNCHED for a slot with no process record and no readiness file; GC removes such a slot, and every other reader still reads it as UNKNOWN. - On POSIX a PID whose command line does not run the slot's orcad.js reads DEAD, so a stopped orcad behind a reused PID is woken instead of reported serving or unverifiable. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a legacy relay route a new pane started serving stays open when a pending stop settles (#25542) A served PTY's exit that lands before its stop's reply defers the route's hang-up until the stop settles. A second pane the same older relay holds could start serving through the route in that window, and the deferred hang-up then closed it anyway, sending that pane's input and stops to the current relay. Serving a pane now cancels the deferred hang-up, and a settling request hangs up only a route that serves nothing. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): the journal holds a migration's scrollback until commit or abort, across failed uploads and restarts (#25550) Retention was scoped to one transfer call, and its release in finally deleted a closed tab's snapshot after an interrupted upload, so every retry failed with source_snapshot_changed. Inline buffers had no file for the retained read at all. Retention now follows the cutover journal: held from the journaled export through staging, rebuilt at startup, released on commit or a removed journal. Inline bytes are written to their ref while held. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * fix(migration): a repo id two hosts share never lets legacy keys cross hosts, and a dangling identity alias stops blocking the move (#25558) - Repo ids are not unique across hosts: the same id may be registered on two SSH hosts. Stores that are not session partitions (worktree metadata, automations, lineage, client state, sparse presets, retired names) can hold legacy keys with no host qualifier, so an id both hosts register said nothing about whose a key was. The scope now records such shared ids; an unqualified key or bare repo id for one only matches with its row's own host evidence (worktree metadata's hostId, an automation's ssh target), so another host's rows are never moved, counted or retired. An automation's target generation now matches only alongside its target id. - An identity alias whose identities hold no metadata (worktreeMetaByIdentity lost them, as on the B4 profile) is nothing to move rather than a worktree-metadata blocker; retirement drops it. Co-authored-by: m4air <m4air@Mac.localdomain> * feat(cli): orca environment recover --accept-changed-state --yes restores over changed state (#25597) Recover refused when a rejected build changed profile state, and its refusal told the user to run Recover, which the CLI could not do. --accept-changed-state (confirmed with --yes) maps to the same acceptChangedState the Managed servers settings pass, and the refusal now names the flags. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): refuse an update that would end a degraded host's in-process terminals (#25598) The census now reports inProcessSessions separately. Those terminals run inside orcad and end with any restart, so planOrcadUpdate defers with a non-forceable orcad_update_ends_in_process_terminals instead of claiming they survive on the daemon. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): a delta move protects what it imported from rolling back across it (#25599) A delta move never advanced the server's migration mark, so rolling back an update taken before the delta was admitted and dropped the delta's projects. The delta now records the mark before any commit can land, resumed commits included; the mark keeps the latest migration and never moves back; and the rollback gate also counts every journal into the server, so deltas finished before this change stay protected. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a missing relay inventory never proves terminals exited, and the Windows census covers every desktop's relays (#25611) - The migration terminal gate returned `exited` whenever this desktop held no unresolved lease, even when the current relay or the earlier relays could not be asked. A failed or incomplete inventory is now `unverifiable` regardless of local leases. With no relay session at all the gate asks for a host census (`needsHostCensus`) instead of reading the silence as exit; the connect, conversion and delta move pass that census in, and the connect hands its own census result to the conversion it starts. A census that cannot list endpoints (`unenumerable`) is unverifiable too. - The Windows connect-time census derived pipe names from this desktop's target id only, so another desktop's relay on the same account was never probed. It now lists every `orca-relay-*` pipe on the machine and maps each to the relay instance that owns it through that instance's credential file or active-pipe marker, asking each with its own credential. A pipe no version directory accounts for is unverifiable unless the host proves it another account's (or gone), and an inventory that could not be read is unverifiable. Co-authored-by: m4air <m4air@Mac.localdomain> * refactor: drop the unshipped pty.resumeClient relay method and unused SSH provider unregister guards (#25595) Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a connect still deciding its server holds the raw 'connected', closes a transport its cancelled decision opened, and an edit keeps a relay host's session (#25641) - handleSshConnectionStateChange holds a raw 'connected' while a connect is in flight even before any relay session exists (published as 'connecting'), so the census, deploy or conversion that dials the pool no longer reports the host up with no providers. - priorConnection is captured before the server decision; a connect cancelled after the decision closes a transport the decision opened, unless a newer connect is using it. - Editing a fenced host an older build changed (it runs on the relay directly) no longer disconnects its transport; only a host reached through its managed server redials. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a retained source is unchanged only against a pre-commit baseline of everything a user wrote (#25602) The state fingerprint read drafts, automations and workspace metadata from what a move could carry, so a session a move refuses hid an older build's draft edit, and retention hashed the source after the commit, so a crash before retention blessed whatever an older build changed in between. The baseline is now written with the fence, before any commit is possible, from the source read directly; a session that cannot be read, or a journal without that baseline, is unverified. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a delta move refuses a source that changed while it checked terminals (#25691) The plan and manifest are taken before the terminal check and session release are awaited, but the journal took its baseline after them, so a draft typed in between became the baseline while the server received the older one, and retirement deleted the newer draft. The baseline now comes from the plan's own snapshot, and a source that changed since it refuses the move. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): a save landing mid-retirement no longer defers retirement (#25692) Retirement removed the source rows, then awaited the profile flush, then checked nothing came back. A session save that landed during that flush re-added a source-owned row, the check failed, and the journal stayed committed until a later connect. Retirement is idempotent, so it now runs one more pass before deferring; a row back after that is reported with the partition and owner key it reappeared under. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(automations): a headless run never reads completed when its agent never ran (#25700) * fix(automations): a headless run never reads completed when its agent never ran orcad (and Electron serve) finished a dispatched run on a satisfied tui-idle wait, and a ready shell prompt satisfies it: a run whose agent is not installed read 'completed' within seconds. Like the desktop runner, completion now needs the agent's own status for the run's pane after dispatch; without it the run fails after the agent-start window with the reason, instead of claiming completion. * fix(automations): keep idle-means-done for agents without status; fail only a refused command Not every automation agent reports status on orcad (no hooks on the host, no recognised title), so requiring it would fail their runs. A run completes on the agent's own status, fails when the shell refused the agent's command (bash, zsh, dash, fish, PowerShell, cmd), and otherwise keeps the old idle-means-done rule after the agent-start window. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): the retained-source fingerprint sees all worktree metadata retirement owns (#25694) * fix(migration): the retained-source fingerprint sees all worktree metadata retirement owns The state view read only legacy worktreeMeta keys and attributed them without meta.hostId, so identity-backed metadata and unqualified rows a shared repository id leaves to hostId were missing from the fingerprint: an older build's edit read as unchanged and retirement deleted it. The view now uses the same attribution as export and retirement, and metadata that claims the source host but cannot be attributed leaves the source unverified. The fingerprint version moves to v2. * fix(migration): the retained-source fingerprint skips automations on a repo id another host shares The state view matched automations by scope.repoIds, which ignores sharedRepoIds, so host A's fingerprint included host B's automation on a shared repository id; editing it marked A changed and routed it back to the relay. The view now uses the move's automationTouchesScope. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a failed readiness read no longer kills a healthy candidate; log unsettled activations (BUG-17a) (#25701) * fix(orcad): a failed readiness read no longer fails a candidate's launch, and log every unsettled activation BUG-17: the candidate went ready and the client SIGTERMed it ~1 s later through its reject path, yet that readiness passes the gate, so the launch itself failed: one readiness-wait exec that errored failed the launch outright. Retry such reads until the readiness deadline; an unconfirmed termination still fails at once. The update's outcome never reached the app log, so every update or rollback that does not go through now logs its code and reason. * fix(orcad): the host-side readiness wait ends on a wall-clock deadline On a loaded host each poll's reads outlasted its sleep, so the step-counted loop ran past the client's 30 s exec timeout. That timeout failed the launch, the reject path SIGTERMed a candidate still starting, and orcad, which defers a stop until startup completes, published readiness and then exited (BUG-17). --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(migration): stop a converted host's stale tabs landing in local, and retry retirement only on an exact replay (#25712) * fix(migration): retirement re-removes only an exact replay of moved session state #25692's second pass re-ran retirement on any row that reappeared during the flush, which could delete a tab or draft written after the move. Retirement now records the session rows it removes before it runs; a row that reappears is removed again only when it is byte-identical to one of those (in any partition). A new or changed row defers retirement and stays. * fix(ssh): a converted host's leftover session rows stay in its own partition, never local After conversion the renderer drops the SSH host's projects and worktrees but keeps their session rows. With no catalog owner left, the next save routed those rows to the local partition, where retirement read them as moved source state reappearing and deferred. Converting now pins each dropped worktree's session key to the host's partition; main's fence guard keeps that partition frozen, so the stale rows are never written. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(test): restore the codexProviderHandle import main's #25078 dropped again #25722 restored it, then #25078 removed it, so pnpm tc fails on main's tip. * chore(sync): keep main's own cloud and mobile files byte-identical to main Earlier syncs added lint-only brace and template fixes to these main-owned files; reverting them keeps #24863's diff against main free of files Phase 3 does not own. * fix(terminal): a remote pane's reattach error no longer shows this client's OS and shell (#25693) * fix(terminal): a remote pane's reattach error no longer shows this client's OS and shell The error toast appended the client's environment for every non-SSH error. It now shows it only when the pane's known execution host is this client; an SSH, managed or not-yet-known host omits it. Co-Authored-By: Claude <noreply@anthropic.com> * test(terminal): type the pane-host fixture as runtime owner state Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * fix(orcad): a retained activation fence is ownerless, so recovery can always take it over (BUG-17) (#25698) * fix(orcad): a retained activation fence is ownerless, so recovery can always take it over BUG-17: a recovery that took a stale fence over, failed and retained it left a fresh lock, so every later recovery read it as still fresh and the host could never be recovered. A run that keeps the fence once it is done now backdates the lock; one whose remote command may still be running keeps it fresh. Also run the in-process terminal deferral before the forced protocol check, so a degraded host whose daemon is empty names its in-process terminals instead of an unreported protocol. * test(orcad): the CLI's accepting recover takes over a fence a refused recover just retained The fake host now answers a stale-only takeover busy while a recovery's own takeover is fresh, which reproduces BUG-17's 'still fresh' loop without the ownerless mark. * test(orcad): keep the fake host's fence acquisition void where callers expect it --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a cancelled connect closes only the transport its own decision opened and nothing newer adopted (#25696) * fix(ssh): a cancelled connect closes only the transport its own decision opened and nothing newer adopted #25641's cleanup took any transport that differed from the pre-decision one as the cancelled decision's, guarded only by connectInFlight. A replacement connect that completed (and left connectInFlight) then had its live transport disconnected by the stale attempt. The pool now attributes a transport it opens inside a connect's server decision to that attempt (AsyncLocalStorage), and the latest user adopts it: a connect that connects or publishes a managed route, or a managed tunnel that records a forward. A cancelled attempt closes the transport only when it opened it, nothing newer adopted it, and no current replacement is in flight. * fix(ssh): a still-current connect whose server decision fails closes the transport that decision dialed The decision's own failure (a throw, or a fenced relay refusal) published 'error' but left the transport it dialed open, so getPublicSshState read 'connected' and a later auto-reconnect broadcast a plain 'connected' with no relay. Both branches now close exactly the decision-owned transport through abandonDecisionTransport, which treats the attempt's own in-flight entry as no newer owner while that attempt is still current. * test(ssh): name the stand-in transport type --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): Windows readiness identifies itself by one PID when the process-table snapshot times out (#25723) * fix(orcad): Windows readiness identifies itself by one PID when the process-table snapshot times out A loaded Windows runner timed the whole-table snapshot out during the bundled runtime's readiness preflight, so the candidate failed to start and activation rejected it. * fix(orcad): fall back to the one-PID query only for a slow process table, never an unreadable one An unreadable table (EDR-hooked snapshot, restricted token) must still fail qualification. The table now rejects slowness with a typed WindowsProcessTableTimeoutError, and only that falls back. Review by win-serve. * build(cli): list the process-table timeout error in the CLI project's file list --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a connected conversion proves the host idle with the account-wide census, not this target's lists (#25697) * fix(ssh): a connected conversion proves the host idle with the account-wide census, not this target's lists The migration terminal gate asked the host-wide census only when no relay session existed. With a session, it trusted this target's relay listing and its earlier-relay census, both of which name only this target's instances, and returned `exited` when they were empty, so another desktop's live shell on the same account, under a different target id, let the host convert under it. `exited` now always needs a complete host-wide census: this target's lists can prove `live`, but empty lists only pass the question to the census, and a gate given none answers `unverifiable` with `needsHostCensus`. The connect-time refinement passes the census too, so a connect retires its leases only when no relay on the account holds work. The Windows host lane now runs a second desktop's relay with a live shell and expects the connected gate to read the host live. * test(ssh): a connected relay's empty lists still ask the account-wide census * refactor(ssh): drop the gate's unread needsHostCensus flag; its unverifiable reason says why * test(ssh): the delta snapshot fixtures give their account-wide census --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a reconnected transport starts managed orcad itself instead of inheriting a dropped start (#25689) * fix(ssh): a reconnected transport runs its own orcad start instead of inheriting a dropped one The serving check deduplicated in-flight starts by environment only. When a connect dropped mid-start (as when the launch-time auto-connect is replaced by a reconnect), the caller on the new transport joined the start bound to the dead one, got its failure, and reported the host managed with no server running. In-flight checks now join only on the same connection, connect generation and port, and a wake's own fence token is cleared only by that wake. * fix(ssh): a reconnected wake releases the fence its dropped wake held at any point The flake's real verdict was 'fenced': the dropped launch-time wake held the activation fence, and the reconnected wake could not prove it its own. A wake now claims its token before its first remote step, an absent owner record under a held token is still its own, and a reconnected wake waits for this client's dropped wake to settle before reading the fence. * fix(ssh): release only a fence carrying this process's own wake token A fence with no owner file could be another client's fresh one. Releasing it now requires the owner token this process wrote; the token is claimed before the write so a drop after it still proves ownership. * fix(ssh): type the wake's fenced fallback --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(ssh): stop the exec-stdin test double from failing on EPIPE (#25739) The truncation test's fake exec channel forwarded the local shell's EPIPE (or 'Cannot call end after a stream was destroyed') as a channel error. Whether that error or the shell's exit code won depended on scheduling, so the test failed under full-suite load. ssh2 silently drops writes once the remote stops reading; the double now does the same, and a 1 MB payload makes the early-stop path deterministic. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): Move lets the reconnect's decision run its census, and a census outside a connect holds 'connected' and closes its own transport (#25735) - Move to managed server no longer runs a separate host census after tearing the relay down, which dialed the pool with no connect in flight and broadcast a raw 'connected' with no session or providers. It reconnects, and reads the decision the reconnect's census recorded. A relay a failed stop left up is detached (leases kept), not disposed, before the reconnect. - The CLI and delta-move census (censusHostRelayTerminalsFor) runs outside a connect under its own owner: the raw 'connected' it causes is held, and a transport it opened that nothing adopted is closed afterwards. Reusing a pooled transport inside a scope now adopts it. - Drop the now-unused publishRelayTerminalsStatus. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(test): drop the restored codexProviderHandle import now that main restored it * refactor(migration): keep a converted host's source rows instead of retiring them automatically (#25768) Automatic source retirement leaves Phase 3: nothing deletes a converted host's retained rows on connect, delta move, keep-server's-version or restart. They stay hidden and are removed only by stopping the server, removing the host or uninstalling. Change detection goes back to the catalog-identity fingerprint, so an older build's edits inside an already-moved project stay preserved in the retained rows without marking the host changed. The copy-only helpers the delta view uses are renamed to subtract, and the converted-host session pin now also overrides a boot primary of local. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(automations): a headless run is watched past tui-idle timeouts by the one run observer (#25733) * fix(automations): a headless run is watched past tui-idle timeouts by the one run observer The headless dispatcher awaited a single tui-idle wait, which rejects after its 5-minute default, so a healthy agent working longer was published as dispatch_failed and never observed again. The dispatcher now hands the run to its completion watcher, whose runtime observer already re-arms wait timeouts, honours cancellation and bounds total observation; the agent status and missing-command checks fold into that observer, and the separate completion loop is gone. * fix(automations): an already-idle pane completes when the start window passes Real-host: a stub that exited before the window left an idle shell with no agent status, and the observer re-armed a tui-idle wait that never resolves for an already-idle shell, so it timed out instead of completing. The observer now keeps judging the pane while its output is unchanged, and only waits again once the pane changes. * fix(automations): resolve a watched headless run by its launch handle first --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(runtime-env): a re-paired managed server's subscribers recover without a reload (#25752) * fix(runtime-env): a re-paired managed server's subscribers recover without a reload The renderer kept the pairing revision it last read, so after an on-connect update re-paired a managed server every subscribe and request was refused as 'pairing changed' until a reload. The first refusal now re-reads the environment catalog, so revision-keyed subscriptions resubscribe and requests carry the new pairing. A managed server re-pairing for the same host registration is the same peer, so its workspaces and tabs are no longer purged as a replaced environment. Co-Authored-By: Claude <noreply@anthropic.com> * fix(runtime-env): a re-paired managed server is the same machine only when its host proves the same identity Same SSH target registration is not proof: a reinstalled host or a target now pointing elsewhere keeps it. The runtime id the pairing handshake verifies must be known and unchanged; otherwise the re-pair retires the environment as before. A proven runtime id change also counts as replaced. Co-Authored-By: Claude <noreply@anthropic.com> * fix(runtime-env): decide a managed re-pair's same machine by the host's proven key, not its runtime id The runtime id is minted per process start, so every orcad restart would read as a new host. The host's E2EE public key persists in its own profile across updates and its pairing handshake proves it; main now lists a digest of it, and the renderer keeps a re-paired managed server only when that digest is known and unchanged under the same SSH target registration. Co-Authored-By: Claude <noreply@anthropic.com> * fix(runtime-env): defer a managed re-pair's same-machine decision until the host key is known A re-read that lands before the new pairing's host key is listed no longer purges: the decision waits for a catalog that carries the key and retires only if it differs. Adds the update-flow store test: same registration and key keeps workspaces and tabs, including a re-read that runs before the key is known. Co-Authored-By: Claude <noreply@anthropic.com> * fix(runtime-env): watch for a pairing refusal on a side branch so requests settle on the same tick Chaining .catch onto every subscribe and request delayed each success by a microtask, which let a StrictMode cleanup run before a client-event subscription resolved, so its unsubscribe landed late. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * refactor: consolidate pane-ownership, migration-catalog, activation-launch and parser helpers (#25738) * refactor(orcad): one launch-and-judge helper for activation and rollback * refactor(orcad-migration): one copy each of the destination projections, selectNewRows, assertSameValue, compareKeys and slotLiveness * refactor(orcad-migration): one string-list validator and one uniqueness check, error codes passed in * refactor: one shared pane-ownership and terminal-layout module for migration, profile transfer and split layout * fix(orcad-migration): row-identity helpers in a leaf module (no import cycle); key order in its own module --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): reopen a managed tunnel to a host another desktop restarted (#25800) The tunnel's identity check pinned the saved runtime id, which orcad mints per process. A host updated or woken by another desktop, or restarted while this one was away, failed every reconnect with orcad_identity_mismatch, and nothing could refresh the id because that needs the tunnel. The E2EE handshake with the pinned host key and our accepted token now prove the server; the first authenticated status reply records the new id. A different host is still refused. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * fix(orcad): create the readiness file owner-only so its pairing token is not world-readable (#25809) orcadLaunchCommand truncated .orcad-readiness before setting umask 077, so under a login umask of 022 the file that receives the pairing offer (with a runtime-scope device token) came out 0644. umask 077 now runs first, the readiness file is chmod 600 after the truncate (a redirect keeps an earlier build's 0644), the pid and log files are tightened too, and the slot dir and ~/.orca-remote are chmod 700 so files earlier builds left readable are no longer reachable. The state snapshot capture also sets its umask before creating the snapshot directory. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(terminal): explain a pane whose saved session another host connection owns (#25814) terminal_pane_owner_host_mismatch reached the user raw, with an issue link. It now reads as a plain explanation with the open-a-new-terminal action, like the reattach failure. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * test(topology): allow phase3's headless editor-tab retirement in main's boundary ratchet (#25823) Main's #25329 added the ratchet; phase3's mobile-session-editor-projection.ts writes the host's own session through setWorkspaceSessionForWorktree, the same way the listed headless mobile-session tab writers do. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(automations): a headless host closes finished run terminals, keeping the newest few (#25831) * fix(automations): a headless host closes finished run terminals, keeping the newest few The desktop closes a run's terminal when the run completes; orcad had no renderer to do it, so hourly automations left a shell and PTY per run open forever (28 after ~6h on a real host). The headless service now closes a finished run's terminal after a 10-minute grace, keeps the newest three per automation viewable, and never touches a run that has not finished. * fix(automations): never close a run terminal a client typed into or is viewing Mirrors the desktop's take-over rule on headless hosts: a finished run's terminal stays open when any client drove input to it since spawn, is attached to or viewing it, or when this process cannot tell (it adopted the PTY rather than spawned it). * test(runtime): register a viewer through the public subscribe API * fix(automations): close only completed runs' own panes A failed run can still hold a live agent (blocked on a prompt, past the watch window, or after an observer error), so like the desktop only a completed run's terminal is closed. And only the run's own pane closes, so a pane a user split into the same tab survives. * fix(automations): close a run pane only while it still holds the run's PTY The use check read run.terminalPtyId, but the close hit whatever PTY now occupies the run's pane. Restart-exited-pane and the Codex account-switch restart put a new PTY there, so a terminal a user was using could be killed. The close now resolves the pane's current PTY and closes only when it is the run's own; otherwise it closes nothing and only clears the run's terminal. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a snapshot restore past the exec timeout keeps its fence instead of racing a second restore (#25811) * fix(orcad): a snapshot restore past the exec timeout keeps its fence instead of racing a second restore A capture, restore or clear ran under the generic 30s exec timeout. On ssh2 the timeout closes the channel and reads as a confirmed failure, but sshd leaves a pty-less command running, so rollback ran its rescue restore and recover orphaned the fence for a second one, both in the same stage. - State mutations run through execOrcadStateMutation: no abort, a client wait past the host's deadline, and any closed channel or busy/deadline answer is unconfirmed, so the fence stays fresh. - POSIX hosts wrap each one in `timeout -s KILL` (where present) and a pid-checked lock dir under ~/.orca-remote; the Windows host script takes the same lock. * fix(orcad): a running state mutation keeps the activation fence fresh The fence goes stale by its lock dir's mtime after 20 minutes, and a capture, restore or clear can now run up to 15 under it, so a rollback's rescue capture plus restore could outlast the window and let a recovery steal the fence from a live run. While a mutation runs, the host now touches the fence every 60s (POSIX: a background beat that stops with its shell; Windows: an interval in the host script, whose mutations are now async so the timer runs). A dead process stops refreshing, so stale takeover still recovers it. * fix(orcad): the state-mutation fence heartbeat never refreshes a wake's fence A wake writes .orca-wake-owner into the fence dir and lets its fence age toward takeover; the heartbeat now skips a fence that holds that token, and only ever changes the dir's mtime. * fix(orcad): a state mutation releases its host lock before answering, and names its holder by pid and start time On Windows answer() exits in the stdout write callback, so an op that answered before its first await (MISSING, EMPTY, FAILED) exited before the wrapper's finally and leaked the lock; a reused pid then read as alive and every later capture, restore and clear answered busy. Ops now return their token and the wrapper answers after releasing the lock. The holder is pid plus creation time (the slot's process-tree addon); one that cannot be identified is stale once its lock misses five heartbeats. POSIX gets the same heartbeat-age check for a reused pid. * fix(orcad): a state mutation's host lock is owned by its whole process group The lock named only the shell's pid, so a shell killed while its rm or tar ran let the next mutation take the lock and race that child. Each mutation now runs in its own process group (setsid, or perl setpgrp on macOS), with timeout inside it so a deadline KILL reaches the children too. The lock records the group, and is taken over only once no member is alive; a host that can start no group records none, and its lock is never taken over. Windows ops run in-process, with no children to outlive the holder. * fix(orcad): record a state mutation's process group without ps -p, and never hold a groupless lock forever BusyBox ps has no -p, so Alpine hosts recorded no group and their lock read busy forever after a timeout kill, reboot or OOM. The group now comes from /proc/<pid>/stat (read after the comm field's last paren), with ps -o pgid= -p as the fallback. A lock that still names no group is taken over once its pid is dead and its heartbeat has missed three beats. * fix(orcad): only proof of exit frees a state-mutation lock A Windows holder whose creation time could not be read was taken over after five quiet minutes though its pid was alive, so a suspended clear could resume and delete freshly restored profiles. Both platforms now free the lock only on proof of exit: a dead pid, a different creation time, or (POSIX) a group with no live member. A live holder of unknown identity stays busy until it exits. The owner record is written exclusively, so a run that resumes after a takeover backs off. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): never offer Move for terminals another Orca desktop or session runs (#25815) On a host where another desktop held a live relay shell, this desktop read relay_terminals_live with offerMove, and its copy ("Its N open terminals will restart") implied they were its own. The census already attributes them: terminals counted only by the host-wide census, with no lease or listing of this target naming one, run under another target or session. That verdict now carries elsewhere / terminalsElsewhere; no move is offered (no toast, no status-line action) and the status line says the terminals belong to another Orca desktop or session. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(automations): an update releases finished, unused automation shells before counting terminals (#25844) Hosts with schedules kept completed run shells (the newest three, and any not yet past their grace), which counted as running terminals and deferred every on-connect update with orcad_update_terminals_running. The update and rollback census now ask the server to close completed automation run terminals no client used, with no grace or keep rule, and count after the daemon drops them. Used, unknown, failed and running ones still count. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): every activation fence holder carries a generation token its steps and release must match (#25834) * fix(orcad): clear a bare stale activation fence instead of asking for Recover, and report a restarting update from status BUG-21: a wake cut short leaves a stale fence with no journal. Every update then answered 'Recover it first' while Recover answered 'none'. The fence-hold check now takes such a fence over and drops it, and the update retries once; Recover is asked for only over a journal. The CLI also treats a connection closed by the server's own restart during update or rollback as expected and reports what status shows once the runtime answers. * fix(cli): type the reconnect status response explicitly * fix(ssh): a wake's fence carries its owner token from the moment the lock exists The idle-exit e2e still read 'fenced' on reconnect: the launch-time wake's lock landed on the host but its connection dropped before the client saw OK, so the wake body never ran and never wrote its owner token, leaving a fence nothing could prove. The token is now claimed before the lock and written by the same command that creates it, and a wake registers itself before any remote step so a reconnected wake waits for it instead of racing it. * fix(orcad): every activation fence holder carries a generation token its steps and release must still match Astra pass 8: a holder suspended past the stale window resumed, kept acting, and its unconditional release deleted the successor's fence and recovery journal mid-update. Every holder (activation, rollback, stop, recover, wake) now writes a token into the lock it creates or takes over. Each remote step it issues checks that token on the host, in the same command on POSIX and inside the host script for Windows host ops; release is conditional on the token and moves the lock aside instead of removing the root. A superseded holder aborts with OrcadFenceLostError and its release is a no-op. * fix(orcad): state mutations check the fence token before their lock, and refresh only a fence they still own On POSIX the fence guard runs outermost in serializedStateMutationCommand, before the mutation lock and the work, and the heartbeat touches the fence only while the token is still this run's. The Windows host script records the --fence token and refreshFence compares it. A fence-lost answer from a state mutation is a refusal, never a FAILED fallback. One owner-file constant replaces the wake-owner copies. * refactor(orcad): a state mutation's heartbeat touches the fence directory it checked ownership of (review) * fix(orcad): a release moves the journal and lock aside and keeps only its own generation's Astra pass 9: a release that passed its token check and stalled before deleting could, once a takeover and a successor came and went, delete the successor's journal and lock. The journal is now stamped with the writing run's fence token (a recovery takeover re-stamps the journal it adopts), and the release renames the journal and the lock aside, deletes each only if it carries this run's token, and otherwise moves it straight back. * test(orcad): a successor restore stays busy beside a paused clear on the Windows host script * fix(orcad): classify a lost fence from the step's exit and stdout, never the error message The real exec error quotes the command, and every fenced command carries the guard's marker text, so any failed or timed-out fenced step read as a lost fence and dropped its unconfirmed flag. execCommand now attaches exitCode and stdout to its exit error; execOrcadRemote rethrows unconfirmed terminations before any reclassification. * fix(orcad): a wake keeps its fence token until the fence is released A disconnect fails a wake's next step without the unconfirmed flag, and the fence release then fails over the dead connection. Forgetting the token on that error left a fence the reconnected wake could not prove its own, so it reported the host as held by an update. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(recovery): keep Phase 3's recovery-lifetime test on main's legacy-worker ports Main added a required hasRequestedReleases port and now skips persist when a pass resolves nothing, so the test mocks the new port and holds the pass at workspace resolution instead. * fix(ssh): say "1 terminal" when another Orca desktop runs one on the host (#25853) The terminalsElsewhere status line had no plural forms, so B9 read "while 1 terminals another Orca desktop … are running". It gains _one/_other entries like the other terminal-count strings on that line and in the move offer, which were already pluralized. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(automations): close failed and exited runs' terminals once the shell is proven alone (#25859) * fix(automations): close failed and exited runs' terminals once the shell is proven alone Failed (command not found, timeout) and forever-dispatched runs kept one shell per run, unbounded and counted by the update gate. Their terminals now close like completed ones (unused, past the grace, outside the newest few) but only on fresh execution-host proof that the spawned shell is alone at its prompt; a live or unprovable agent keeps its terminal. A still-dispatched run closed this way is marked failed. Dead terminals no longer take one of the newest-three keep slots. The update drain follows the same rules. * fix(automations): prove a run shell alone from the process table, not the daemon's ownership flag On a real daemon session the daemon's confirmShellForeground stays false after a plain 'command not found' and after an agent that exited, because its ownership flag only turns 'shell' after a full-screen command; failed runs would never have closed. The proof now also reads the host's process table: on POSIX the PTY's root shell must own the terminal foreground group with nothing stopped under it, on Windows the host's job-based child census must be empty. Anything unobservable still keeps the terminal. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(ssh): extensive orca CLI matrix on Windows hosts (#25114) * test(ssh): extensive orca CLI matrix on Windows hosts Adds three dispatch-only app cells to the ssh-windows-hosts lane that drive the e2e build and the bundled orca CLI against the provisioned Win32-OpenSSH host: empty-host deploy/terminal/reconnect/orcad-restart/decommission, seeded relay-era conversion, and an open relay terminal keeping the host on the relay. * test(ssh): pin the relay-kept cell's runtime; keep cleanup from masking failures * test(ssh): run decommission before the orcad restart in the managed cell * test(ssh): decommission through orca environment stop; accept an unverifiable relay close * test(ssh): require a confirmed relay close; app cells must run last * test(ssh): log and accept either relay-kept census reason; keep app-cell test results * test(ssh): relay-kept requires a live census and its status line again * test(ssh): match the pluralized relay-kept status line * test(ssh): orcad restart proves a new process, terminal adoption, and a kill-then-connect relaunch * test(ssh): restart kills only the orcad server, not its terminal daemon; wait for a released profile * test(ssh): collect orcad.log.1 so a restarted orcad's previous run is kept * test(ssh): the managed cell proves a workspace listener is detected and attributed * test(ssh): start the port listener without $, so a PowerShell terminal doesn't expand it * test(ssh): the port check proves Windows command-line attribution; retry a dropped version read --------- Co-authored-by: m4air <m4air@Mac.localdomain> * refactor(runtime): move run-terminal client-use and shell-alone checks out of the branch-cleanup runtime (#25876) * refactor(runtime): move run-terminal client-use and shell-alone checks out of the branch-cleanup runtime orca-runtime-preserved-branch-cleanup.ts had grown past max-lines (303) with the headless run-terminal helpers. Their logic now lives in run-terminal-client-use.ts and the runtime keeps one-line delegators, with behavior unchanged. * fix(ci): the runtime Electron ratchet bundles its entry points once, not 2.5k times check-runtime-electron-ratchet bundled ~2,532 entry points each in full (format cjs, no splitting), so esbuild held thousands of copies of the runtime graph: about 2.2GB RSS and 11s per run, twice per test file. It was in flight in every unit shard that died with "The runner has received a shutdown signal" (#25815 5/5 twice, #25876 2/5 twice). With esm + splitting the shared modules land in one chunk: same metafile, about 200MB and 2s. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * test(ssh): wait for the busy relay's child before probing it (#25916) The fake relay's spawn is not visible to pgrep at READY on Linux under Bun, so the probe could count zero children. The sibling cases already wait. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): a quit that aborts an upload whose read already ended no longer crashes main (BUG-23) (#25922) Quitting while an on-connect orcad update was uploading its bundle aborted the connection's teardown signal. sftp-upload's abort handler destroyed the local read stream with the signal's reason, but once that read had ended, 'finished' had already removed its listeners, so the stream emitted an unhandled 'error': [main_uncaught_exception] AbortError: This operation was aborted. Electron's error dialog then blocked the main thread and the app never exited. The read stream now always has a no-op error listener; the transfer's outcome still comes from 'finished'. Both the bare upload and the connection-level teardown abort are covered. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a launch reads readiness at least once; fake hosts match the capture's tar flag, not any -cf (#25918) A random fence token contains `-cf` about 1 time in 125, and the fake hosts read any command containing it as a snapshot capture, so a rollback's restore answered CAPTURED and the rollback never launched. Separately, a client descheduled between computing the readiness deadline and checking it skipped every read and failed a ready launch. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a fence this desktop's exited process left is cleared without the 20-minute wait (#25941) * fix(orcad): a fence this desktop's exited process left is cleared without the 20-minute wait The client records the fence tokens its processes hold beside the profile. On a later launch, a POSIX fence carrying a token from a process that has exited, quiet for three heartbeats and with no live state mutation, is backdated so the existing stale rules clear it or hand it to Recover at once. Another desktop's fence, a live holder's, or one with a mutation still running keeps the normal stale window. * fix(orcad): held fence tokens are best effort, pinned to this machine and boot, and pruned after a day A token-file write that fails no longer breaks a fence operation; an entry recorded on another machine sharing the profile, or before a reboot, never proves its holder exited; entries older than 24 hours are dropped. Tests cover a journal kept for Recover, a successor freshened back after a racing backdate, a Windows host, and the record across release, supersession, busy and a lost connection. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): an install lock this desktop's exited process left mid-upload is taken over without the 20-minute wait (#25991) A quit during the bundle upload leaves the version dir's install lock, not the activation fence. The lock now carries this desktop's token, recorded in the held-token store, and is forgotten only once its removal is confirmed. On a later attempt, before each stale check, a POSIX lock whose token belongs to an exited process of this machine and boot, quiet for three minutes, is backdated so the existing stale takeover claims it at once. The fence path now shares the same helper. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(orcad): a desktop that met another desktop's update fence clears its note once the host answers (#25995) The serving note "holds this host" and a fence-busy update deferral stayed until a reconnect, minutes after the other desktop's update finished. The connect now rechecks serving and the update every 45s while the fence holds, and publishes the first answer without it. A recorded deferral is dropped once the host runs its candidate or a newer release. Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude <noreply@anthropic.com> * test(ci): run Phase 3's SQLite-backed tests in the Node runtime project Main's #25967/#25998 boundary requires every test that opens real SQLite to be listed. This adds Phase 3's eleven orcad and SSH migration tests, plus main's own agent-launch-instant-tab test (#25430), which main's tip also leaves unlisted. * fix(ci): keep Electron probes out of the node-server suites again (#26046) The runner excluded *.electron.test.ts with a CLI --exclude, but main's switch to Vitest inline projects (#25967) gave each project its own exclude list, which overrides the CLI one. The directory selectors then pulled profile-state-writer-stall.electron.test.ts into the glibc-floor and musl orcad-template jobs, which have no xvfb. Resolve the exact files with vitest list and drop Electron and cross-runtime ones before running. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): keep the SSH host card quiet while its managed server is healthy (#26072) The card showed "Runs a managed Orca server" under every healthy host. A managed server is the default, so the status line now appears only for setup progress, updates, the relay, or failures. Co-authored-by: m4air <m4air@Mac.localdomain> * fix(ssh): Move to managed server keeps the host's terminal tabs (#26077) * fix(ssh): Move to managed server keeps the host's terminal tabs Move stops the relay shells; their exits read as a user exit and closed the tabs before the conversion copied them to the server. Suppress those exits for the move, restart stopped shells on the relay when the host stays, re-home the open workspace onto the server, and report stopped shells to the runtime so terminal list stops calling them connected. * fix(ssh): mark Move's relay stops in main's intentional-stop register The renderer-only exit suppression left main retiring the stopped tab from the saved SSH session before the conversion copied it, left other viewers unprotected, and swallowed real exits for the whole request. Register exactly the shells the move stops, from just before each shutdown, as a 'replaced' stop with their incarnation; main keeps the surface and labels the exit for every viewer, while a confirmed death stays 'exited'. Move now returns the shells it stopped, and a host that stays on the relay restarts only those tabs, discarding any buffered exit first. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(hosts): show an SSH host and its managed Orca server as one host (#26076) * fix(hosts): show an SSH host and its managed Orca server as one host Phase 3 registers the Orca server it deploys over SSH as its own runtime environment, so every host list built from the execution-host registry listed the machine twice under the same name. The registry now folds the pair into one row named after the SSH host. The row routes to the server, since a managed host has no relay, unless main reports the host back on its relay; the other id stays as an alias so selections and renames saved under it still resolve. A retired id that workspaces still point at keeps its own row, and servers no configured SSH host deployed (manual pairings, orphans) are untouched. * fix(hosts): keep both ids of a merged SSH host and dedupe only in pickers Deleting the merged-away id from the registry broke every consumer that matches hosts by exact id: Add Project fell back to local after a connect, the composer lost ready projects and drafts (and could swap in an unrelated local project), and a host scope hid folder-only workspaces. The registry now keeps both entries and marks the pair (aliasHostIds on the row pickers show, mergedIntoHostId on the other). Pickers show one row per machine, and a choice of that row expands to both ids: sidebar host scope, jump palette filter, notification toggles, run-target and repository host offers. Add Project resolves a saved SSH id to its server row and blocks the actions while that server comes up instead of choosing local. The composer's resolver now fails closed when a named draft repo isn't actionable rather than picking another project. * fix(hosts): widen saved host scopes, both-way palette aliases, guard Add Project host - A sidebar or agents host scope saved by an older build (or before a route flip) can hold one id of a merged SSH host; a background gate widens it to both ids so exact-id filters match either owner. - The palette host filter now resolves a saved id to both owners whichever id it names. - Add Project's create and clone refuse to run while the chosen host is unresolved, and their submit buttons stay disabled, instead of falling through to this computer. --------- Co-authored-by: m4air <m4air@Mac.localdomain> * fix(sync): reconcile main's ratchet bundling and cold-serve hydrate with phase3 The Electron-import ratchet keeps main's single-stdin bundle (cjs); the auto-merge had also kept phase3's esm splitting, which broke main's import-graph test. Editor tabs now follow the windowless full-seed rule from #26022, so a cold serve restart lists persisted editors too. * fix(ssh): reclaim this desktop's own exited lock on Windows hosts too (#26087) * fix(ssh): reclaim this desktop's own exited lock on Windows hosts too The relaunch after a quit mid-update now frees the activation fence and the version-dir install lock on a Windows SSH host the same way it does on POSIX, instead of waiting out the 20-minute stale window. The host script ages the lock only when its token belongs to a desktop process proven exited, it has been quiet for three heartbeats, and (for the fence) no state mutation is live, where a mutation holder counts as gone only by pid plus creation time. * fix(ssh): take an exited holder's lock only through the steal arbitration Review found the reclaim backdated the lock by path after checking it, so a live successor that replaced the lock in between could be aged and then stolen, and an interrupted or failed restore left it aged for good. The exited-holder check is now read-only. The steal command itself accepts the proven token and, inside its steal claim and identity recheck, also takes a lock whose owner file still names that token and that has been quiet for three heartbeats. Nothing is written to a lock before the steal owns it. POSIX uses the same path. * fix(ssh): never take an exited holder's fence while a state mutation can start Review round 2 found the fence's live-mutation guard ran only in the read-only proof, so a mutation admitted after the proof, or one whose first heartbeat landed after the steal sampled the fence's age, kept running under a fence the steal had replaced. For the fence, the steal now takes the state-mutation lock inside its claim (mkdir on POSIX, the exclusive owner.json on Windows) and holds it until the takeover is done; it refuses when any mutation lock exists. Holding it, it rereads the owner and only then re-samples the fence identity. A mutation now rechecks its fence token right after it takes the mutation lock and stops with the fence-lost marker if it changed. The Windows proof also falls back to the stale window when its command line would not fit cmd.exe. * fix(ssh): record the exited-owner steal as a real mutation-lock holder Review round 3 found the POSIX steal held the state-mutation lock as an empty directory, which a mutation reclaims after a minute without any liveness check; a steal stalled that long lost its exclusion and could replace the fence under a running mutation. The steal now writes its pid (and group, under the same rule) with the mutation's own noclobber owner writer, so only proof of its exit frees the lock, and it removes the lock only while the lock still names it. On Windows the owner record is moved into place whole, so it never exists empty, and is removed only while it still names the steal's pid. * refactor(ssh): keep the relay lock commands off the orcad host-script graph The mutation-lock owner writers moved into a leaf module, so the relay's install-lock commands no longer import orcad-state-snapshot and, through it, the Windows host script, orcad-instance-lock and the daemon process query. Those modules evaluate imports at load time that existing suites mock partially. No behavior change. --------- Co-authored-by: m4air <m4air@Mac.localdomain> --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
This commit is contained in:
co-authored by
m4air
Claude Opus 5.5
m4air
parent
5b0d38749d
commit
5cafefe726
@@ -6,7 +6,9 @@ name: Adhoc macOS + Windows Dev Build
|
||||
#
|
||||
# Deliberately narrow scope:
|
||||
# - macOS and Windows desktop installers. Linux keeps using RC/stable.
|
||||
# - No tests, no lint, no e2e. PR CI and release-cut remain the gates.
|
||||
# - No tests, no lint, no e2e. PR CI and release-cut remain the gates. The one
|
||||
# exception is the orcad template: like release-cut, it is merged from the
|
||||
# node-server lanes that build and qualify each SSH target's slot.
|
||||
# - macOS is signed and notarized so TCC grants survive updates.
|
||||
# - Windows is unsigned; the published release notes explain the one-time
|
||||
# SmartScreen/manual-install requirement.
|
||||
@@ -81,21 +83,116 @@ env:
|
||||
ADHOC_RETAIN_DAYS: 30
|
||||
|
||||
jobs:
|
||||
# Why: every job below checks out its own copy, and a branch name read per job builds whatever
|
||||
# the branch points at when that job starts. A push mid-run then mixes commits, e.g. slot lanes
|
||||
# from one commit merged by a template step from the next. Resolving once pins the whole run.
|
||||
# The mac job still vets this commit's reachability before anything is signed.
|
||||
resolve-ref:
|
||||
if: github.repository == 'stablyai/orca'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 5
|
||||
permissions:
|
||||
contents: read
|
||||
outputs:
|
||||
sha: ${{ steps.resolve.outputs.sha }}
|
||||
steps:
|
||||
- name: Resolve the requested ref to one commit
|
||||
id: resolve
|
||||
shell: bash
|
||||
env:
|
||||
REQUESTED_REF: ${{ inputs.ref || github.ref_name }}
|
||||
REPO_URL: https://github.com/${{ github.repository }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
case "$REQUESTED_REF" in
|
||||
refs/pull/*|pull/*)
|
||||
echo "::error::Refusing to build PR ref '$REQUESTED_REF'; push the code to a branch of stablyai/orca instead."
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
if [[ "$REQUESTED_REF" =~ ^[0-9a-f]{40}$ ]]; then
|
||||
sha="$REQUESTED_REF"
|
||||
else
|
||||
# Branch first, as the mac job's vet does; a peeled annotated tag names its commit.
|
||||
refs="$(git ls-remote "$REPO_URL" "refs/heads/$REQUESTED_REF" "refs/tags/$REQUESTED_REF" "refs/tags/$REQUESTED_REF^{}")"
|
||||
sha=""
|
||||
for name in "refs/heads/$REQUESTED_REF" "refs/tags/$REQUESTED_REF^{}" "refs/tags/$REQUESTED_REF"; do
|
||||
sha="$(awk -v ref="$name" '$2 == ref { print $1; exit }' <<<"$refs")"
|
||||
[[ -n "$sha" ]] && break
|
||||
done
|
||||
fi
|
||||
if [[ -z "$sha" ]]; then
|
||||
echo "::error::'$REQUESTED_REF' is not a branch, tag, or full commit SHA of stablyai/orca."
|
||||
exit 1
|
||||
fi
|
||||
echo "Resolved $REQUESTED_REF -> $sha"
|
||||
echo "sha=$sha" >>"$GITHUB_OUTPUT"
|
||||
|
||||
# Why its own job: this package ships relays for Windows SSH hosts, and only a
|
||||
# Windows runner compiles the addon that launches one outside sshd's job.
|
||||
# Why the unvetted ref is safe here: this job holds no secrets, and the mac job
|
||||
# vets the same ref before it downloads anything, so fork code never gets signed.
|
||||
relay-windows-process-tree:
|
||||
if: github.repository == 'stablyai/orca'
|
||||
needs: resolve-ref
|
||||
permissions:
|
||||
contents: read
|
||||
uses: ./.github/workflows/relay-windows-process-tree.yml
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref_name }}
|
||||
ref: ${{ needs.resolve-ref.outputs.sha }}
|
||||
|
||||
# Why: a branch cut before the orcad template landed has none to build or ship.
|
||||
orcad-template-support:
|
||||
needs: resolve-ref
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 5
|
||||
permissions:
|
||||
contents: read
|
||||
outputs:
|
||||
ships: ${{ steps.detect.outputs.ships }}
|
||||
steps:
|
||||
- name: Checkout the template packager only
|
||||
uses: actions/checkout@v6
|
||||
with:
|
||||
ref: ${{ needs.resolve-ref.outputs.sha }}
|
||||
sparse-checkout: |
|
||||
/config/scripts/packaged-orcad-template.cjs
|
||||
sparse-checkout-cone-mode: false
|
||||
persist-credentials: false
|
||||
- name: Detect whether the ref ships the orcad template
|
||||
id: detect
|
||||
shell: bash
|
||||
run: |
|
||||
if [[ -f config/scripts/packaged-orcad-template.cjs ]]; then
|
||||
echo "ships=true" >>"$GITHUB_OUTPUT"
|
||||
else
|
||||
echo "ships=false" >>"$GITHUB_OUTPUT"
|
||||
echo "::notice::This ref predates the orcad template; the build ships without it."
|
||||
fi
|
||||
|
||||
# Design D2, as release-cut does: every desktop build ships the orcad template (server JS plus
|
||||
# every target's addons), merged from the node-server lanes that qualify each slot at this ref.
|
||||
# Without it an adhoc build cannot deploy managed orcad to an SSH host. No secrets, like the
|
||||
# relay job, and the mac job vets the ref before anything is signed.
|
||||
orcad-template:
|
||||
needs: [resolve-ref, orcad-template-support]
|
||||
if: needs.orcad-template-support.outputs.ships == 'true'
|
||||
permissions:
|
||||
contents: read
|
||||
uses: ./.github/workflows/node-server-tests.yml
|
||||
with:
|
||||
ref: ${{ needs.resolve-ref.outputs.sha }}
|
||||
build_template: true
|
||||
|
||||
build-adhoc-mac:
|
||||
needs: relay-windows-process-tree
|
||||
if: github.repository == 'stablyai/orca'
|
||||
needs: [resolve-ref, relay-windows-process-tree, orcad-template-support, orcad-template]
|
||||
# Why not the implicit success(): a ref without the orcad template skips that job on purpose.
|
||||
if: >-
|
||||
!cancelled() && github.repository == 'stablyai/orca' &&
|
||||
needs.relay-windows-process-tree.result == 'success' &&
|
||||
needs.orcad-template-support.result == 'success' &&
|
||||
(needs.orcad-template.result == 'success' ||
|
||||
(needs.orcad-template.result == 'skipped' &&
|
||||
needs.orcad-template-support.outputs.ships == 'false'))
|
||||
# Why an environment: it gives the signing/notary/App secrets somewhere to
|
||||
# live that a stale copy of this workflow on an old branch cannot reach.
|
||||
# Referencing it is a no-op until repo settings give it teeth; the intended
|
||||
@@ -109,6 +206,7 @@ jobs:
|
||||
version: ${{ steps.adhoc.outputs.version }}
|
||||
head_sha: ${{ steps.adhoc.outputs.head_sha }}
|
||||
published: ${{ steps.publish_live.outcome == 'success' && 'true' || 'false' }}
|
||||
ships_orcad_template: ${{ needs.orcad-template-support.outputs.ships }}
|
||||
runs-on: blacksmith-6vcpu-macos-15
|
||||
# Why 150: it must exceed the worst case the retry budgets below can produce
|
||||
# (install 3x10 + publish 2x45 = 120, plus ~25 for checkout/build/verify), or
|
||||
@@ -128,7 +226,8 @@ jobs:
|
||||
id: vetted
|
||||
shell: bash
|
||||
env:
|
||||
REQUESTED_REF: ${{ inputs.ref || github.ref_name }}
|
||||
# The commit resolve-ref pinned for every job; vetting it keeps the reachability test.
|
||||
REQUESTED_REF: ${{ needs.resolve-ref.outputs.sha }}
|
||||
REPO_URL: https://github.com/${{ github.repository }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
@@ -292,6 +391,14 @@ jobs:
|
||||
# sshd's job for a standard user on a Windows SSH host.
|
||||
ORCA_REQUIRE_RELAY_NATIVE_ADDONS: x64,arm64
|
||||
|
||||
# After the app build so nothing that cleans out/ can drop it; electron-builder ships it.
|
||||
- name: Download the orcad deployment template
|
||||
if: needs.orcad-template-support.outputs.ships == 'true'
|
||||
uses: actions/download-artifact@v8
|
||||
with:
|
||||
name: orcad-template
|
||||
path: out/orcad-template
|
||||
|
||||
# Why the token is minted here and not at the top: installation tokens live
|
||||
# one hour, everything before this point writes nothing, and the notary round
|
||||
# trip inside the publish step can be tens of minutes. Minting after the build
|
||||
@@ -366,6 +473,8 @@ jobs:
|
||||
GH_TOKEN: ${{ steps.app_token.outputs.token }}
|
||||
ORCA_ADHOC_BUILD_VERSION: ${{ steps.adhoc.outputs.version }}
|
||||
ORCA_BUILD_COMMIT: ${{ steps.adhoc.outputs.commit }}
|
||||
# beforePack and afterPack fail the package when the template is absent.
|
||||
ORCA_REQUIRE_ORCAD_TEMPLATE: ${{ needs.orcad-template-support.outputs.ships == 'true' && '1' || '' }}
|
||||
CSC_LINK: ${{ secrets.MAC_CERTS }}
|
||||
CSC_KEY_PASSWORD: ${{ secrets.MAC_CERTS_PASSWORD }}
|
||||
# Why all three: electron-builder's notarize step authenticates to the
|
||||
@@ -514,3 +623,4 @@ jobs:
|
||||
tag: ${{ needs.build-adhoc-mac.outputs.tag }}
|
||||
ref: ${{ needs.build-adhoc-mac.outputs.head_sha }}
|
||||
version: ${{ needs.build-adhoc-mac.outputs.version }}
|
||||
orcad_template: ${{ needs.build-adhoc-mac.outputs.ships_orcad_template == 'true' }}
|
||||
|
||||
@@ -58,6 +58,11 @@ on:
|
||||
description: Version to package, without the leading v
|
||||
required: true
|
||||
type: string
|
||||
orcad_template:
|
||||
description: Ship the orcad-template artifact the calling run built (adhoc does; hourly and daily do not yet)
|
||||
required: false
|
||||
type: boolean
|
||||
default: false
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
channel:
|
||||
@@ -264,6 +269,14 @@ jobs:
|
||||
# is correct for unvetted artifacts. Same as the mac dev channels.
|
||||
ORCA_DIAGNOSTICS_TOKEN_URL: https://www.onorca.dev/diagnostics/token
|
||||
|
||||
# After the app build so nothing that cleans out/ can drop it; electron-builder ships it.
|
||||
- name: Download the orcad deployment template
|
||||
if: inputs.orcad_template
|
||||
uses: actions/download-artifact@v8
|
||||
with:
|
||||
name: orcad-template
|
||||
path: out/orcad-template
|
||||
|
||||
# Why the token is minted here and not at the top: installation tokens live
|
||||
# one hour and nothing before this point writes anything.
|
||||
- name: Mint dev channel repo token
|
||||
@@ -301,6 +314,8 @@ jobs:
|
||||
env:
|
||||
GH_TOKEN: ${{ steps.app_token.outputs.token }}
|
||||
ORCA_BUILD_COMMIT: ${{ inputs.ref }}
|
||||
# beforePack and afterPack fail the package when the template is absent.
|
||||
ORCA_REQUIRE_ORCAD_TEMPLATE: ${{ inputs.orcad_template && '1' || '' }}
|
||||
# Why: electron-publish refuses to upload into a release published more
|
||||
# than two hours ago (gitHubPublisher.getOrCreateRelease). The mac leg
|
||||
# publishes the draft live as soon as *it* finishes, so a slow notary
|
||||
|
||||
+156
-1
@@ -341,7 +341,11 @@ jobs:
|
||||
. != "tests/e2e/paired-startup-exec-readiness.spec.ts" and
|
||||
. != "tests/e2e/ssh-browser-network-execution-route.docker.unit.test.ts" and
|
||||
. != "tests/e2e/ssh-localhost.spec.ts" and
|
||||
. != "tests/e2e/terminal-ibus-hangul-native.spec.ts"
|
||||
. != "tests/e2e/terminal-ibus-hangul-native.spec.ts" and
|
||||
. != "tests/e2e/orcad-serve-mode-switch.spec.ts" and
|
||||
. != "tests/e2e/ssh-orcad-auto-convert.spec.ts" and
|
||||
. != "tests/e2e/windows-missing-appdata-startup.spec.ts" and
|
||||
. != "tests/e2e/ssh-orcad-idle-exit.spec.ts"
|
||||
)' <<<"$TEST_FILES_JSON" > "$RUNNER_TEMP/general-e2e-specs"
|
||||
fi
|
||||
mapfile -t TEST_FILES < "$RUNNER_TEMP/general-e2e-specs"
|
||||
@@ -510,6 +514,157 @@ jobs:
|
||||
ORCA_RUN_DOCKER_SSH_BROWSER_E2E: '1'
|
||||
run: node_modules/.bin/vitest run --config config/vitest.config.ts tests/e2e/ssh-browser-network-execution-route.docker.unit.test.ts
|
||||
|
||||
orcad-serve-mode-switch:
|
||||
name: orca serve Electron/orcad mode switch (D7)
|
||||
needs: [build, prepare-native-cache]
|
||||
if: inputs.test_files == '' || contains(inputs.test_files, 'tests/e2e/orcad-serve-mode-switch.spec.ts')
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
env:
|
||||
ORCA_BACKGROUND_LAUNCH: '1'
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
persist-credentials: false
|
||||
- name: Install headless tools
|
||||
run: sudo apt-get update && sudo apt-get install -y build-essential ripgrep xvfb openbox x11-utils
|
||||
- uses: ./.github/actions/install-node-dependencies
|
||||
with:
|
||||
native-runtime: electron
|
||||
- uses: actions/download-artifact@v8
|
||||
with:
|
||||
name: e2e-build-out
|
||||
path: out/
|
||||
# The packaged orcad slot the T6-11 launcher runs: its own node-pty under the pinned Node.
|
||||
# The template is what `orca serve`'s default selection materializes that slot from.
|
||||
- name: Build this runner's orcad slot and template
|
||||
run: |
|
||||
slot="$(node config/scripts/build-orcad-prebuilds.mjs --print-slot)"
|
||||
pnpm build:orcad-prebuilds
|
||||
pnpm build:orcad-prebuilds --require-slots "$slot"
|
||||
pnpm build:orcad
|
||||
node config/scripts/build-orcad-template.mjs --targets "$slot"
|
||||
- name: Switch serve hosts on one profile
|
||||
env:
|
||||
SKIP_BUILD: '1'
|
||||
ORCA_E2E_ORCAD_SERVE: '1'
|
||||
ORCA_E2E_FORWARD_APP_LOGS: '1'
|
||||
run: xvfb-run --auto-servernum bash .github/scripts/e2e-with-window-manager.sh pnpm exec playwright test --config tests/playwright.config.ts tests/e2e/orcad-serve-mode-switch.spec.ts --project=electron-headless --workers=1
|
||||
- uses: actions/upload-artifact@v7
|
||||
if: failure()
|
||||
with:
|
||||
name: orcad-serve-mode-switch-traces
|
||||
path: test-results/
|
||||
retention-days: 7
|
||||
if-no-files-found: ignore
|
||||
|
||||
# #24979 on a real host: a relay-era Docker host converts to managed orcad on connect, and a managed
|
||||
# orcad idles out and restarts on the next connect. Needs the
|
||||
# orcad template for the fixture's target (Debian, linux-x64-glibc), which only this job builds.
|
||||
orcad-auto-convert-docker:
|
||||
name: ssh host auto-converts to managed orcad (Docker)
|
||||
needs: [build, prepare-native-cache]
|
||||
if: inputs.test_files == '' || contains(inputs.test_files, 'tests/e2e/ssh-orcad-auto-convert.spec.ts') || contains(inputs.test_files, 'tests/e2e/ssh-orcad-idle-exit.spec.ts')
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 45
|
||||
env:
|
||||
ORCA_BACKGROUND_LAUNCH: '1'
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
persist-credentials: false
|
||||
- name: Install headless tools
|
||||
run: sudo apt-get update && sudo apt-get install -y build-essential ripgrep xvfb openbox x11-utils
|
||||
- uses: ./.github/actions/install-node-dependencies
|
||||
with:
|
||||
native-runtime: electron
|
||||
- uses: actions/download-artifact@v8
|
||||
with:
|
||||
name: e2e-build-out
|
||||
path: out/
|
||||
# Built to a path the app does not look at by default, so the relay phase has no template.
|
||||
- name: Build the linux-x64-glibc orcad template
|
||||
run: |
|
||||
pnpm build:orcad-prebuilds
|
||||
pnpm build:orcad-prebuilds --require-slots linux-x64-glibc
|
||||
pnpm build:orcad
|
||||
node config/scripts/build-orcad-template.mjs --targets linux-x64-glibc
|
||||
mv out/orcad-template "$RUNNER_TEMP/orcad-convert-template"
|
||||
- name: Convert a relay-era Docker host
|
||||
env:
|
||||
SKIP_BUILD: '1'
|
||||
ORCA_E2E_SSH_DOCKER: '1'
|
||||
ORCA_E2E_ORCAD_CONVERT_HOST: docker
|
||||
ORCA_E2E_FORWARD_APP_LOGS: '1'
|
||||
ORCA_RELAY_PATH: ${{ github.workspace }}/out/relay
|
||||
run: |
|
||||
export ORCA_E2E_ORCAD_CONVERT_TEMPLATE="$RUNNER_TEMP/orcad-convert-template"
|
||||
xvfb-run --auto-servernum bash .github/scripts/e2e-with-window-manager.sh pnpm exec playwright test --config tests/playwright.config.ts tests/e2e/ssh-orcad-auto-convert.spec.ts tests/e2e/ssh-orcad-idle-exit.spec.ts --project=electron-headless --workers=1
|
||||
- uses: actions/upload-artifact@v7
|
||||
if: failure()
|
||||
with:
|
||||
name: orcad-auto-convert-docker-traces
|
||||
path: test-results/
|
||||
retention-days: 7
|
||||
if-no-files-found: ignore
|
||||
|
||||
# D7 on Windows: both hosts share <userData>\daemon and so one daemon pipe. Built here rather than
|
||||
# from the Linux e2e build because the slot, template and addons are Windows-native. Also runs the
|
||||
# missing-AppData startup check, which needs the same Windows e2e build.
|
||||
orcad-serve-mode-switch-windows:
|
||||
name: orca serve on Windows (D7 mode switch, missing AppData)
|
||||
if: inputs.test_files == '' || contains(inputs.test_files, 'tests/e2e/orcad-serve-mode-switch.spec.ts') || contains(inputs.test_files, 'tests/e2e/windows-missing-appdata-startup.spec.ts')
|
||||
runs-on: windows-2022
|
||||
timeout-minutes: 45
|
||||
env:
|
||||
NODE_OPTIONS: --max-old-space-size=4096
|
||||
ORCA_BACKGROUND_LAUNCH: '1'
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
persist-credentials: false
|
||||
- uses: ./.github/actions/install-node-dependencies
|
||||
with:
|
||||
native-runtime: electron
|
||||
- uses: ./.github/actions/prepare-orcad-prebuilds
|
||||
with:
|
||||
resolve-windows-cache: 'true'
|
||||
restore-windows-cache: 'true'
|
||||
- name: Build the e2e app, CLI, orcad slot and template
|
||||
shell: bash
|
||||
run: |
|
||||
pnpm run build:relay
|
||||
pnpm exec electron-vite build --mode e2e
|
||||
pnpm run build:cli
|
||||
pnpm build:orcad
|
||||
node config/scripts/build-windows-process-tree-relay-addon.mjs --arch=x64
|
||||
node config/scripts/build-orcad-template.mjs --targets win32-x64
|
||||
# A profile-less Windows session has no roaming AppData; Electron 43 used to crash natively there.
|
||||
- name: Start serve with no AppData folder
|
||||
if: inputs.test_files == '' || contains(inputs.test_files, 'tests/e2e/windows-missing-appdata-startup.spec.ts')
|
||||
env:
|
||||
SKIP_BUILD: '1'
|
||||
ORCA_E2E_FORWARD_APP_LOGS: '1'
|
||||
run: pnpm exec playwright test --config tests/playwright.config.ts tests/e2e/windows-missing-appdata-startup.spec.ts --project=electron-headless --workers=1
|
||||
- name: Switch serve hosts on one profile
|
||||
if: ${{ !cancelled() && (inputs.test_files == '' || contains(inputs.test_files, 'tests/e2e/orcad-serve-mode-switch.spec.ts')) }}
|
||||
env:
|
||||
SKIP_BUILD: '1'
|
||||
ORCA_E2E_ORCAD_SERVE: '1'
|
||||
ORCA_E2E_FORWARD_APP_LOGS: '1'
|
||||
ORCA_E2E_PRESERVE_PROFILE_LOGS_DIR: ${{ github.workspace }}\test-results\profile-logs
|
||||
run: pnpm exec playwright test --config tests/playwright.config.ts tests/e2e/orcad-serve-mode-switch.spec.ts --project=electron-headless --workers=1
|
||||
- uses: actions/upload-artifact@v7
|
||||
if: failure()
|
||||
with:
|
||||
name: orcad-serve-mode-switch-windows-traces
|
||||
path: test-results/
|
||||
retention-days: 7
|
||||
if-no-files-found: ignore
|
||||
|
||||
ssh-localhost:
|
||||
name: localhost SSH terminal and hooks
|
||||
needs: [build, prepare-native-cache]
|
||||
|
||||
@@ -1269,6 +1269,7 @@ jobs:
|
||||
src/main/cursor/hook-service.test.ts
|
||||
src/main/orca-profiles/profile-index-store.test.ts
|
||||
src/main/startup/windows-install-dir-acl-repair.win32.test.ts
|
||||
src/main/startup/windows-app-data-path.test.ts
|
||||
src/main/runtime/repo-worktree-admin-fingerprint.test.ts
|
||||
src/main/runtime/worktree-scan-admin-fingerprint-gate.test.ts
|
||||
src/shared/secure-file-fsync-flags.test.ts
|
||||
|
||||
@@ -60,6 +60,23 @@ jobs:
|
||||
node config/scripts/build-windows-process-tree-relay-addon.mjs --arch=x64
|
||||
node config/scripts/build-windows-process-tree-relay-addon.mjs --arch=arm64
|
||||
|
||||
# orcad's instance lock and stop proof read one PID's creation time through this addon on
|
||||
# Windows SSH hosts; prove the x64 binary this runner can load answers it, and answers
|
||||
# nothing for a PID that does not exist rather than a stale or zero time.
|
||||
- name: Assert the addon reads process creation time
|
||||
shell: bash
|
||||
run: |
|
||||
node -e "
|
||||
const { assertWindowsProcessTreeCreationTime } = require('./config/scripts/windows-process-tree-creation-time.cjs')
|
||||
const addon = require('./.build/windows-process-tree/x64/windows-process-tree.node')
|
||||
assertWindowsProcessTreeCreationTime({ module: addon })
|
||||
const own = addon.getProcessCreationTime(process.pid)
|
||||
const started = Date.now() - process.uptime() * 1000
|
||||
if (typeof own !== 'number' || Math.abs(own - started) > 5000) throw new Error('creation time ' + own + ' vs ' + started)
|
||||
if (addon.getProcessCreationTime(0x7ffffff0) !== undefined) throw new Error('a missing PID must read undefined')
|
||||
console.log('creation time ok:', own)
|
||||
"
|
||||
|
||||
- name: Upload relay addons
|
||||
uses: actions/upload-artifact@v7
|
||||
with:
|
||||
|
||||
@@ -31,8 +31,18 @@ on:
|
||||
- 'config/scripts/relay-windows-process-tree-prepared-addon*.mjs'
|
||||
- 'config/scripts/windows-process-tree-gyp-rebuild.mjs'
|
||||
- 'src/shared/relay-windows-breakaway-launch.ts'
|
||||
- 'src/shared/windows-breakaway-launch*.ts'
|
||||
- 'src/main/ipc/ssh-host-server-*.ts'
|
||||
- '!src/**/*.test.ts'
|
||||
- 'src/main/ssh/ssh-relay-windows-host-lane.test.ts'
|
||||
- 'src/main/ssh/orcad-windows-host-lane.test.ts'
|
||||
- 'tests/e2e/ssh-orcad-auto-convert.spec.ts'
|
||||
- 'tests/e2e/helpers/orcad-convert-host.ts'
|
||||
- 'tests/e2e/helpers/orcad-convert-flow.ts'
|
||||
- 'tests/e2e/helpers/orcad-upgrade-profile.ts'
|
||||
- 'tests/e2e/ssh-orcad-windows-cli-matrix.spec.ts'
|
||||
- 'tests/e2e/helpers/compiled-orca-cli.ts'
|
||||
- 'tests/e2e/helpers/windows-host-orcad-processes.ts'
|
||||
- 'config/ci/windows-ssh-provider/**'
|
||||
- '.github/workflows/ssh-windows-hosts.yml'
|
||||
- '.github/actions/prepare-orcad-prebuilds/**'
|
||||
@@ -41,7 +51,7 @@ on:
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
cells:
|
||||
description: Comma-separated cell ids from src/main/ssh/ssh-windows-host-cells.ts; empty runs all.
|
||||
description: Comma-separated cell ids from src/main/ssh/ssh-windows-host-cells.ts; empty runs the default set (orcad-cli-* cells run only when named).
|
||||
required: false
|
||||
default: ''
|
||||
|
||||
@@ -75,8 +85,9 @@ jobs:
|
||||
server: preview
|
||||
archive: OpenSSH-ARM64.zip
|
||||
runs-on: ${{ matrix.runner }}
|
||||
# Installing the inbox capability alone can take several minutes on a fresh image.
|
||||
timeout-minutes: 75
|
||||
# Installing the inbox capability alone can take several minutes on a fresh image; the CLI matrix
|
||||
# cells launch the e2e app several times each.
|
||||
timeout-minutes: ${{ contains(github.event.inputs.cells || '', 'orcad-cli') && 160 || 75 }}
|
||||
env:
|
||||
ORCA_BACKGROUND_LAUNCH: '1'
|
||||
ORCA_ISOLATED_SSH_CI: '1'
|
||||
@@ -85,16 +96,14 @@ jobs:
|
||||
with:
|
||||
persist-credentials: false
|
||||
# Keep complete server/test trees; missing future imports fail the unchanged builds.
|
||||
# All of src: the orcad-convert cell builds the full e2e app and the bundled CLI.
|
||||
sparse-checkout: |
|
||||
.github
|
||||
config
|
||||
native
|
||||
resources
|
||||
tests
|
||||
src/main
|
||||
src/shared
|
||||
src/relay
|
||||
src/types
|
||||
src
|
||||
- name: Self-test the provisioning scripts before touching the machine
|
||||
shell: pwsh
|
||||
run: |
|
||||
@@ -149,7 +158,7 @@ jobs:
|
||||
- wait: inbox-capability
|
||||
- name: Run the Windows host cells against a private ${{ matrix.server }} sshd
|
||||
shell: pwsh
|
||||
timeout-minutes: 50
|
||||
timeout-minutes: ${{ contains(github.event.inputs.cells || '', 'orcad-cli') && 130 || 50 }}
|
||||
env:
|
||||
CELLS: ${{ github.event.inputs.cells || '' }}
|
||||
run: |
|
||||
@@ -158,7 +167,13 @@ jobs:
|
||||
$receipts=Join-Path $pwd '.build/ssh-windows-host-receipts'
|
||||
New-Item -ItemType Directory -Force -Path $receipts | Out-Null
|
||||
$cells=@($env:CELLS -split ',' | ForEach-Object {$_.Trim()} | Where-Object {$_})
|
||||
if(-not $cells.Count){$cells=@('pinned-cmd','pinned-powershell','legacy-opt-out')}
|
||||
if(-not $cells.Count){$cells=@('pinned-cmd','pinned-powershell','legacy-opt-out','orcad-cmd','orcad-powershell')}
|
||||
# One Windows host runs the app-level conversion; it is last because it switches native modules.
|
||||
if(-not $env:CELLS -and '${{ matrix.arch }}' -eq 'x64' -and '${{ matrix.server }}' -eq 'inbox'){$cells+='orcad-convert'}
|
||||
# App cells reach the managed server through an SSH local forward, which provisioning grants to the last accounts only.
|
||||
$appCellIds=@('orcad-convert','orcad-cli-managed','orcad-cli-convert','orcad-cli-relay-kept')
|
||||
$cells=@($cells | Where-Object {$appCellIds -notcontains $_})+@($cells | Where-Object {$appCellIds -contains $_})
|
||||
$forwarding=@($cells | Where-Object {$appCellIds -contains $_}).Count
|
||||
$archive=''
|
||||
$preparation=''
|
||||
if('${{ matrix.server }}' -eq 'inbox' -and '${{ matrix.arch }}' -eq 'arm64'){$preparation=Join-Path $receipts 'inbox-capability-preparation.json'}
|
||||
@@ -171,7 +186,7 @@ jobs:
|
||||
$callback={param($context)
|
||||
& (Join-Path $tools 'invoke-pinned-relay-cells.ps1') -SourceRoot $sourceRoot -Context $context -Target 'win32-${{ matrix.arch }}' -ReceiptRoot $receipts -Cells $cells
|
||||
}.GetNewClosure()
|
||||
& (Join-Path $tools 'preview-ssh/prove-preview-openssh.ps1') -Archive $archive -Arch '${{ matrix.arch }}' -Server '${{ matrix.server }}' -Receipt (Join-Path $receipts 'provider-server.json') -InboxPreparationReceipt $preparation -Accounts $cells.Count -HiddenTools @('npm','npx','node-gyp','gcc','g++','cc','c++','make','cl','clang','clang++','msbuild','cmake') -HostCellProbe $callback 2>&1 | Tee-Object (Join-Path $receipts 'provision.log')
|
||||
& (Join-Path $tools 'preview-ssh/prove-preview-openssh.ps1') -Archive $archive -Arch '${{ matrix.arch }}' -Server '${{ matrix.server }}' -Receipt (Join-Path $receipts 'provider-server.json') -InboxPreparationReceipt $preparation -Accounts $cells.Count -ForwardingAccounts $forwarding -HiddenTools @('npm','npx','node-gyp','gcc','g++','cc','c++','make','cl','clang','clang++','msbuild','cmake') -HostCellProbe $callback 2>&1 | Tee-Object (Join-Path $receipts 'provision.log')
|
||||
- uses: actions/upload-artifact@v7
|
||||
if: always()
|
||||
with:
|
||||
|
||||
@@ -0,0 +1,104 @@
|
||||
name: Windows packaged orcad serve switch E2E
|
||||
|
||||
# Why: proves D7 on a real Windows install. The installed desktop app forks its terminal daemon
|
||||
# from the relocated %LOCALAPPDATA%\Orca\daemon-host; `orca serve` then serves the same profile
|
||||
# on orcad and must adopt that daemon, and the desktop must reattach the terminal afterwards and
|
||||
# still write its settings under orcad's data-root ACL. Unpackaged e2e cannot exercise the
|
||||
# relocation, which only packaged builds do. A CI runner is the only safe place to install.
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
types: [opened, synchronize, reopened, ready_for_review]
|
||||
paths:
|
||||
- 'src/main/daemon/**'
|
||||
- 'src/main/orcad/**'
|
||||
- 'src/cli/runtime/**'
|
||||
- 'src/shared/orcad-local-serve-selection.ts'
|
||||
- 'tests/tools/win-update-e2e/**'
|
||||
- '.github/workflows/win-orcad-serve-switch-e2e.yml'
|
||||
workflow_dispatch:
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
concurrency:
|
||||
group: win-orcad-serve-switch-e2e-${{ github.event.pull_request.number || github.ref }}
|
||||
cancel-in-progress: true
|
||||
|
||||
jobs:
|
||||
serve-switch:
|
||||
name: packaged Windows Electron/orcad serve switch (D7)
|
||||
if: github.event_name != 'pull_request' || github.event.pull_request.draft != true
|
||||
runs-on: windows-2022
|
||||
timeout-minutes: 75
|
||||
env:
|
||||
NODE_OPTIONS: --max-old-space-size=4096
|
||||
ORCA_BACKGROUND_LAUNCH: '1'
|
||||
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v6
|
||||
with:
|
||||
persist-credentials: false
|
||||
|
||||
- uses: ./.github/actions/install-node-dependencies
|
||||
with:
|
||||
native-runtime: node
|
||||
|
||||
# Same key inputs as win-update-survival-e2e: a harness-only edit skips the installer build.
|
||||
- name: Cache branch installer
|
||||
id: cache-installer
|
||||
uses: actions/cache/restore@v4
|
||||
with:
|
||||
path: dist/orca-windows-setup.exe
|
||||
key: serve-switch-installer-${{ hashFiles('src/**', 'config/**', 'native/**', 'resources/**', 'mobile/**', 'patches/**', 'package.json', 'pnpm-lock.yaml', 'pnpm-workspace.yaml') }}
|
||||
|
||||
- uses: ./.github/actions/install-mobile-dependencies
|
||||
if: steps.cache-installer.outputs.cache-hit != 'true'
|
||||
|
||||
# Before the template exists: packaging verifies a present template for every target.
|
||||
- name: Build Windows installer (unsigned)
|
||||
if: steps.cache-installer.outputs.cache-hit != 'true'
|
||||
shell: bash
|
||||
run: |
|
||||
node config/scripts/ensure-native-runtime.mjs --runtime=electron
|
||||
pnpm run build:desktop
|
||||
pnpm exec electron-builder --config config/electron-builder.config.cjs --win --publish never
|
||||
|
||||
- name: Cache the built installer
|
||||
if: steps.cache-installer.outputs.cache-hit != 'true'
|
||||
uses: actions/cache/save@v4
|
||||
with:
|
||||
path: dist/orca-windows-setup.exe
|
||||
key: ${{ steps.cache-installer.outputs.cache-primary-key }}
|
||||
|
||||
# The win32-x64 template the harness drops into the install, which is where serve looks.
|
||||
- uses: ./.github/actions/prepare-orcad-prebuilds
|
||||
with:
|
||||
resolve-windows-cache: 'true'
|
||||
restore-windows-cache: 'true'
|
||||
- name: Build the win32-x64 orcad template
|
||||
shell: bash
|
||||
run: |
|
||||
node config/scripts/ensure-native-runtime.mjs --runtime=node
|
||||
pnpm build:orcad
|
||||
node config/scripts/build-windows-process-tree-relay-addon.mjs --arch=x64
|
||||
node config/scripts/build-orcad-template.mjs --targets win32-x64
|
||||
|
||||
- name: Switch serve hosts on the installed app's profile
|
||||
shell: bash
|
||||
run: |
|
||||
mkdir -p artifacts
|
||||
node tests/tools/win-update-e2e/serve-switch.mjs \
|
||||
--installer dist/orca-windows-setup.exe \
|
||||
--template out/orcad-template 2>&1 | tee artifacts/serve-switch.log
|
||||
exit "${PIPESTATUS[0]}"
|
||||
|
||||
- name: Upload serve switch output
|
||||
if: always()
|
||||
uses: actions/upload-artifact@v7
|
||||
with:
|
||||
name: win-orcad-serve-switch-output
|
||||
path: artifacts/
|
||||
retention-days: 7
|
||||
if-no-files-found: warn
|
||||
@@ -43,6 +43,7 @@ const PLAIN_NODE_ENTRY_NAMES = [
|
||||
'parcel-watcher-process-entry',
|
||||
'computer-sidecar',
|
||||
'wsl-transcript-fs-process-entry',
|
||||
'orcad/orcad-local-serve-selection-entry',
|
||||
...CLI_MAIN_ENTRY_NAMES
|
||||
] as const
|
||||
|
||||
|
||||
@@ -6,12 +6,17 @@ param(
|
||||
[Parameter(Mandatory=$true)][hashtable]$Context,
|
||||
[Parameter(Mandatory=$true)][ValidateSet('win32-arm64','win32-x64')][string]$Target,
|
||||
[Parameter(Mandatory=$true)][string]$ReceiptRoot,
|
||||
[ValidateSet('pinned-cmd','pinned-powershell','legacy-opt-out')][string[]]$Cells=@('pinned-cmd','pinned-powershell','legacy-opt-out')
|
||||
[ValidateSet('pinned-cmd','pinned-powershell','legacy-opt-out','orcad-cmd','orcad-powershell','orcad-convert','orcad-cli-managed','orcad-cli-convert','orcad-cli-relay-kept')][string[]]$Cells=@('pinned-cmd','pinned-powershell','legacy-opt-out','orcad-cmd','orcad-powershell')
|
||||
)
|
||||
$ErrorActionPreference='Stop'
|
||||
if($env:GITHUB_ACTIONS -ne 'true' -or $env:ORCA_ISOLATED_SSH_CI -ne '1'){throw 'Disposable CI only'}
|
||||
$shells=@{'pinned-cmd'='cmd';'pinned-powershell'='powershell';'legacy-opt-out'='cmd'}
|
||||
$shells=@{'pinned-cmd'='cmd';'pinned-powershell'='powershell';'legacy-opt-out'='cmd';'orcad-cmd'='cmd';'orcad-powershell'='powershell';'orcad-convert'='cmd';'orcad-cli-managed'='cmd';'orcad-cli-convert'='cmd';'orcad-cli-relay-kept'='cmd'}
|
||||
# App-level cells drive the e2e build (and the bundled CLI) against the host; each greps one tagged test.
|
||||
$appCells=@{'orcad-convert'=@('tests/e2e/ssh-orcad-auto-convert.spec.ts','');'orcad-cli-managed'=@('tests/e2e/ssh-orcad-windows-cli-matrix.spec.ts','@orcad-cli-managed');'orcad-cli-convert'=@('tests/e2e/ssh-orcad-windows-cli-matrix.spec.ts','@orcad-cli-convert');'orcad-cli-relay-kept'=@('tests/e2e/ssh-orcad-windows-cli-matrix.spec.ts','@orcad-cli-relay-kept')}
|
||||
$electronBuilt=$false
|
||||
if($Context.accounts.Count -lt $Cells.Count){throw 'Each cell needs its own private account'}
|
||||
$seenApp=$false
|
||||
foreach($id in $Cells){if($appCells.ContainsKey($id)){$seenApp=$true}elseif($seenApp){throw 'App cells must run last: they switch native modules to Electron'}}
|
||||
if(-not $Context.forbiddenToolLog){throw 'Run the provisioning with -HiddenTools so toolchain calls are logged'}
|
||||
$openSshKey='HKLM:\SOFTWARE\OpenSSH'
|
||||
$windowsPowerShell=Join-Path $env:WINDIR 'System32\WindowsPowerShell\v1.0\powershell.exe'
|
||||
@@ -53,6 +58,37 @@ function Test-PrivateWmiLaunch([string]$Account) {
|
||||
if($match.Success){return $match.Groups[1].Value}else{return 'no-output'}
|
||||
}
|
||||
|
||||
# App-level cells (tests/e2e/ssh-orcad-auto-convert.spec.ts, ssh-orcad-windows-cli-matrix.spec.ts). Last
|
||||
# in the run: they switch native modules to Electron's ABI, which the vitest cells cannot load.
|
||||
function Invoke-AppCell($Account,[string]$Descriptor,[string]$Log,[string]$Spec,[string]$Grep) {
|
||||
$ready=Invoke-PrivateSsh $Account.name 'git init -q orca-convert-repo && git -C orca-convert-repo -c user.name=orca -c user.email=orca@example.invalid commit -q --allow-empty -m init && echo ORCA_REPO_READY'
|
||||
if($ready -notmatch 'ORCA_REPO_READY'){throw 'Could not create the convert cell repository as the account'}
|
||||
# Out of the app's default lookup, so the relay phase runs without a template.
|
||||
$template=Join-Path $env:RUNNER_TEMP 'orcad-convert-template'
|
||||
# Why guarded: Copy-Item into an existing folder nests the copy instead of replacing it.
|
||||
if(-not (Test-Path -LiteralPath $template)){Copy-Item -LiteralPath 'out\orcad-template' -Destination $template -Recurse -Force}
|
||||
Rename-Item -LiteralPath 'out\orcad-template' -NewName 'orcad-template.convert-hidden'
|
||||
try {
|
||||
if(-not $script:electronBuilt){
|
||||
& node config/scripts/ensure-native-runtime.mjs --runtime=electron 2>&1 | Tee-Object -FilePath $Log | Out-Host
|
||||
if($global:LASTEXITCODE -ne 0){Write-Host 'Switching native modules to Electron failed';return $global:LASTEXITCODE}
|
||||
& pnpm exec electron-vite build --mode e2e 2>&1 | Tee-Object -FilePath $Log -Append | Out-Host
|
||||
if($global:LASTEXITCODE -ne 0){Write-Host 'The e2e app build failed';return $global:LASTEXITCODE}
|
||||
$script:electronBuilt=$true
|
||||
}
|
||||
$env:ORCA_E2E_ORCAD_CONVERT_HOST=$Descriptor;$env:ORCA_E2E_ORCAD_CONVERT_TEMPLATE=$template;$env:SKIP_BUILD='1'
|
||||
$grepArgs=if($Grep){@('--grep',$Grep)}else{@()}
|
||||
& pnpm exec playwright test --config tests/playwright.config.ts $Spec @grepArgs --project=electron-headless --workers=1 2>&1 | Tee-Object -FilePath $Log -Append | Out-Host
|
||||
# Functions return uncaptured output, so only the exit code may reach the caller.
|
||||
return $global:LASTEXITCODE
|
||||
} finally {
|
||||
# Screenshots, traces and error context: Playwright clears test-results on the next cell's run.
|
||||
if(Test-Path -LiteralPath 'test-results'){Copy-Item -LiteralPath 'test-results' -Destination ($Log -replace '\.log$','.test-results') -Recurse -Force}
|
||||
Remove-Item Env:ORCA_E2E_ORCAD_CONVERT_HOST,Env:ORCA_E2E_ORCAD_CONVERT_TEMPLATE,Env:SKIP_BUILD -ErrorAction SilentlyContinue
|
||||
Rename-Item -LiteralPath 'out\orcad-template.convert-hidden' -NewName 'orcad-template'
|
||||
}
|
||||
}
|
||||
|
||||
New-Item -ItemType Directory -Force -Path $ReceiptRoot | Out-Null
|
||||
Push-Location $SourceRoot
|
||||
try {
|
||||
@@ -78,11 +114,18 @@ try {
|
||||
@{cell=$cell;target=$Target;host='127.0.0.1';port=[int]$Context.port;username=$account.name;identityFile=$Context.identityFile;home=$account.home;forbiddenToolLog=$Context.forbiddenToolLog;receipt=(Join-Path $ReceiptRoot "$cell.json")} | ConvertTo-Json | Set-Content -LiteralPath $descriptor -Encoding utf8NoBOM
|
||||
$env:ORCA_RUN_SSH_WINDOWS_HOST='1';$env:ORCA_SSH_WINDOWS_HOST_CELL=$descriptor
|
||||
Write-Host "Windows host cell $cell ($Target, DefaultShell $shell, account $($account.name))"
|
||||
& node node_modules/vitest/vitest.mjs run --config config/vitest.config.ts src/main/ssh/ssh-relay-windows-host-lane.test.ts --reporter=verbose 2>&1 | Tee-Object -FilePath (Join-Path $ReceiptRoot "$cell.log")
|
||||
# Why global: under the workflow's GetNewClosure callback, bare $LASTEXITCODE reads a stale captured copy.
|
||||
$code=$global:LASTEXITCODE
|
||||
if($appCells.ContainsKey($cell)){
|
||||
$code=Invoke-AppCell $account $descriptor (Join-Path $ReceiptRoot "$cell.log") $appCells[$cell][0] $appCells[$cell][1]
|
||||
} else {
|
||||
# orcad cells deploy managed orcad instead of the relay; same account and descriptor shape.
|
||||
$lane=if($cell.StartsWith('orcad-')){'src/main/ssh/orcad-windows-host-lane.test.ts'}else{'src/main/ssh/ssh-relay-windows-host-lane.test.ts'}
|
||||
& node node_modules/vitest/vitest.mjs run --config config/vitest.config.ts $lane --reporter=verbose 2>&1 | Tee-Object -FilePath (Join-Path $ReceiptRoot "$cell.log")
|
||||
# Why global: under the workflow's GetNewClosure callback, bare $LASTEXITCODE reads a stale captured copy.
|
||||
$code=$global:LASTEXITCODE
|
||||
}
|
||||
# The relay's own log is the only record of why it closed a client.
|
||||
foreach($log in @(Get-ChildItem -Path (Join-Path $account.home '.orca-remote\relay-*\relay*.log') -File -ErrorAction SilentlyContinue)){Copy-Item -LiteralPath $log.FullName -Destination (Join-Path $ReceiptRoot "$cell.$($log.Directory.Name).$($log.Name)")}
|
||||
# orcad.log holds only the last launch on Windows; orcad.log.1 is the one before a restart.
|
||||
foreach($log in @(Get-ChildItem -Path (Join-Path $account.home '.orca-remote\relay-*\relay*.log'),(Join-Path $account.home '.orca-remote\orcad-*\orcad.log'),(Join-Path $account.home '.orca-remote\orcad-*\orcad.log.1') -File -ErrorAction SilentlyContinue)){Copy-Item -LiteralPath $log.FullName -Destination (Join-Path $ReceiptRoot "$cell.$($log.Directory.Name).$($log.Name)")}
|
||||
if(Test-Path -LiteralPath $Context.forbiddenToolLog){Copy-Item -LiteralPath $Context.forbiddenToolLog -Destination (Join-Path $ReceiptRoot "$cell.forbidden-tool-calls.log")}
|
||||
$summary.Add(@{cell=$cell;shell=$shell;account=$account.name;exitCode=$code})
|
||||
if($code -ne 0){$failed.Add($cell)}
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
# -HiddenTools: the private accounts are denied every machine PATH directory holding one of these
|
||||
# executables, and their own PATH carries logging shims for them, so SSH sessions have no host toolchain.
|
||||
# -HostCellProbe receives a context hashtable (accounts, port, keys, shim log) once provisioning passes.
|
||||
param([Parameter(Mandatory=$true)][string]$Receipt,[string]$Archive,[Parameter(Mandatory=$true)][ValidateSet('arm64','x64')][string]$Arch,[ValidateSet('preview','inbox')][string]$Server='preview',[scriptblock]$ProductionRouteProbe,[ValidateRange(1,4)][int]$Accounts=1,[string[]]$HiddenTools=@(),[scriptblock]$HostCellProbe,[string]$InboxPreparationReceipt)
|
||||
param([Parameter(Mandatory=$true)][string]$Receipt,[string]$Archive,[Parameter(Mandatory=$true)][ValidateSet('arm64','x64')][string]$Arch,[ValidateSet('preview','inbox')][string]$Server='preview',[scriptblock]$ProductionRouteProbe,[ValidateRange(1,6)][int]$Accounts=1,[ValidateRange(0,6)][int]$ForwardingAccounts=0,[string[]]$HiddenTools=@(),[scriptblock]$HostCellProbe,[string]$InboxPreparationReceipt)
|
||||
$ErrorActionPreference = 'Stop'
|
||||
. (Join-Path $PSScriptRoot 'windows-ssh-capability.ps1')
|
||||
$target=@{arm64=@{os='Arm64';folder='OpenSSH-ARM64';machine='0xAA64';archive='698c6aec31c1dd0fb996206e8741f4531a97355686b5431ef347d531b07fcd42'};x64=@{os='X64';folder='OpenSSH-Win64';machine='0x8664';archive='23f50f3458c4c5d0b12217c6a5ddfde0137210a30fa870e98b29827f7b43aba5'}}[$Arch]
|
||||
@@ -246,6 +246,7 @@ PermitTunnel no
|
||||
PermitTTY no
|
||||
Subsystem sftp "$sftpServerPosix"
|
||||
LogLevel DEBUG1
|
||||
$(@($accountNames | Select-Object -Last $ForwardingAccounts | ForEach-Object {"Match User $_`n AllowTcpForwarding local"}) -join "`n")
|
||||
"@ | Set-Content -LiteralPath $config -Encoding ascii
|
||||
Write-Stage 'server-config-validate-start'
|
||||
Invoke-Bounded $sshd @('-t','-f',$config) | Out-Null
|
||||
|
||||
@@ -0,0 +1,93 @@
|
||||
// An adhoc run builds one commit: every job checks out what resolve-ref pinned at dispatch, so a
|
||||
// push mid-run can never mix slot lanes from one commit with a template merge from the next.
|
||||
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs'
|
||||
import { tmpdir } from 'node:os'
|
||||
import { join } from 'node:path'
|
||||
import { pathToFileURL } from 'node:url'
|
||||
import { afterAll, beforeAll, describe, expect, it } from 'vitest'
|
||||
import { parse } from 'yaml'
|
||||
import { runProcess } from '../../src/shared/child-process/run-process'
|
||||
|
||||
const workflow = parse(readFileSync('.github/workflows/adhoc-mac-build.yml', 'utf8'))
|
||||
const PINNED = '${{ needs.resolve-ref.outputs.sha }}'
|
||||
const resolveStep = workflow.jobs['resolve-ref'].steps.find((step) => step.id === 'resolve')
|
||||
const directory = mkdtempSync(join(tmpdir(), 'adhoc-pinned-commit-'))
|
||||
const repository = join(directory, 'remote.git')
|
||||
const identity = {
|
||||
...process.env,
|
||||
GIT_AUTHOR_NAME: 'Ref test',
|
||||
GIT_AUTHOR_EMAIL: 'ref-test@example.com',
|
||||
GIT_COMMITTER_NAME: 'Ref test',
|
||||
GIT_COMMITTER_EMAIL: 'ref-test@example.com'
|
||||
}
|
||||
let branchTip, tagged
|
||||
|
||||
async function git(args) {
|
||||
const result = await runProcess({ program: 'git', args, env: identity })
|
||||
expect(result.code, result.stderr).toBe(0)
|
||||
return result.stdout.trim()
|
||||
}
|
||||
|
||||
async function resolve(ref) {
|
||||
const scratch = mkdtempSync(join(directory, 'attempt-'))
|
||||
const script = join(scratch, 'resolve.sh')
|
||||
writeFileSync(script, resolveStep.run)
|
||||
const output = join(scratch, 'output')
|
||||
writeFileSync(output, '')
|
||||
const result = await runProcess({
|
||||
program: 'bash',
|
||||
args: [script],
|
||||
env: {
|
||||
...identity,
|
||||
REPO_URL: pathToFileURL(repository).href,
|
||||
GITHUB_OUTPUT: output,
|
||||
REQUESTED_REF: ref
|
||||
}
|
||||
})
|
||||
return { code: result.code, output: readFileSync(output, 'utf8').trim() }
|
||||
}
|
||||
|
||||
beforeAll(async () => {
|
||||
await git(['init', '--bare', repository])
|
||||
const tree = await git(['-C', repository, 'mktree'])
|
||||
tagged = await git(['-C', repository, 'commit-tree', tree, '-m', 'tagged'])
|
||||
branchTip = await git(['-C', repository, 'commit-tree', tree, '-p', tagged, '-m', 'tip'])
|
||||
await git(['-C', repository, 'update-ref', 'refs/heads/feature/x', branchTip])
|
||||
await git(['-C', repository, 'tag', '-a', 'v1', tagged, '-m', 'annotated'])
|
||||
})
|
||||
|
||||
afterAll(() => rmSync(directory, { recursive: true, force: true }))
|
||||
|
||||
describe('adhoc build pins one commit for the whole run', () => {
|
||||
it('checks out the pinned commit in every job that builds from the repo', () => {
|
||||
const { jobs } = workflow
|
||||
expect(jobs['relay-windows-process-tree'].with.ref).toBe(PINNED)
|
||||
expect(jobs['orcad-template'].with.ref).toBe(PINNED)
|
||||
const support = jobs['orcad-template-support'].steps.find((step) =>
|
||||
step.uses?.startsWith('actions/checkout@')
|
||||
)
|
||||
expect(support.with.ref).toBe(PINNED)
|
||||
const vet = jobs['build-adhoc-mac'].steps.find((step) => step.id === 'vetted')
|
||||
expect(vet.env.REQUESTED_REF).toBe(PINNED)
|
||||
for (const name of [
|
||||
'relay-windows-process-tree',
|
||||
'orcad-template-support',
|
||||
'orcad-template',
|
||||
'build-adhoc-mac'
|
||||
]) {
|
||||
expect([jobs[name].needs].flat(), name).toContain('resolve-ref')
|
||||
}
|
||||
})
|
||||
|
||||
it('resolves a branch, an annotated tag to its commit, and passes a full SHA through', async () => {
|
||||
expect(await resolve('feature/x')).toEqual({ code: 0, output: `sha=${branchTip}` })
|
||||
expect(await resolve('v1')).toEqual({ code: 0, output: `sha=${tagged}` })
|
||||
expect(await resolve(tagged)).toEqual({ code: 0, output: `sha=${tagged}` })
|
||||
})
|
||||
|
||||
it('refuses PR refs and names it cannot resolve', async () => {
|
||||
for (const ref of ['refs/pull/1/head', 'pull/1/head', 'missing', tagged.slice(0, 12)]) {
|
||||
expect(await resolve(ref), ref).toEqual({ code: 1, output: '' })
|
||||
}
|
||||
})
|
||||
})
|
||||
@@ -3,7 +3,8 @@
|
||||
* Build one node-pty prebuilt for the CURRENT platform/arch/libc and file it in orcad's
|
||||
* prebuilds matrix, so a deployment target needs no C/C++ toolchain.
|
||||
*
|
||||
* node-pty is the only ABI-sensitive native module orcad requires. It is also PATCHED in
|
||||
* node-pty is the only ABI-sensitive native module every slot builds; a compat slot also
|
||||
* builds the addons in COMPAT_SLOT_ADDONS (orcad-prebuild-compat-addons.mjs). node-pty is PATCHED in
|
||||
* this repo (config/patches/node-pty@1.1.0.patch), and that patch is the glibc-floor fix:
|
||||
* `.symver` pins on openpty/forkpty/pthread_sigmask plus the `--no-as-needed` ldflags that
|
||||
* keep libutil/libpthread in DT_NEEDED. An upstream prebuilt has none of it and reproduces
|
||||
@@ -40,13 +41,14 @@ import {
|
||||
highestGlibcNeed,
|
||||
isCompatSlot,
|
||||
mergeManifest,
|
||||
prebuildCompileGypi,
|
||||
readManifest,
|
||||
sha256Of,
|
||||
slotGlibcFloor,
|
||||
slotSourceFiles,
|
||||
SLOT_NAPI_VERSION
|
||||
} from './orcad-prebuild-slot-contents.mjs'
|
||||
import { compileCompatAddons } from './orcad-prebuild-compat-addons.mjs'
|
||||
import { nodeGypRebuild, stageNodeAddonApi } from './orcad-prebuild-node-gyp.mjs'
|
||||
import { ensurePinnedNodeExecutable, preparePinnedNodeDir } from './pinned-node-downloads.mjs'
|
||||
|
||||
export { readManifest }
|
||||
@@ -203,44 +205,23 @@ async function compileNodePty(sourceDir, slot) {
|
||||
)
|
||||
const ptySourcePath = join(stagedDir, 'src', 'unix', 'pty.cc')
|
||||
writeFileSync(ptySourcePath, ptySourceForLibc(readFileSync(ptySourcePath, 'utf8'), libc))
|
||||
const addonApiDir = dirname(
|
||||
require.resolve('node-addon-api/package.json', { paths: [sourceDir] })
|
||||
)
|
||||
cpSync(addonApiDir, join(stagedDir, 'node_modules', 'node-addon-api'), {
|
||||
recursive: true,
|
||||
dereference: true
|
||||
})
|
||||
stageNodeAddonApi(sourceDir, stagedDir)
|
||||
if (process.platform === 'win32') {
|
||||
require('./node-pty-job-ownership.cjs').assertNodePtySourceDeniesMsysBreakaway({
|
||||
nodePtyDir: stagedDir
|
||||
})
|
||||
}
|
||||
const compileGypi = join(workDir, 'prebuild-compile.gypi')
|
||||
writeFileSync(compileGypi, prebuildCompileGypi({ staticCxxRuntime: isCompatSlot(slot) }))
|
||||
const nodeDir = await preparePinnedNodeDir({ target: slot, workDir: join(workDir, 'nodedir') })
|
||||
|
||||
console.log(
|
||||
`[orcad-prebuilds] compiling patched node-pty for ${slot} against Node ${NODE_RUNTIME_PIN.version} headers, N-API ${SLOT_NAPI_VERSION} ...`
|
||||
)
|
||||
const { runProcessSync } = await import('./script-child-process.mjs')
|
||||
const result = runProcessSync({
|
||||
program: process.execPath,
|
||||
args: [
|
||||
join(ROOT, 'node_modules', 'node-gyp', 'bin', 'node-gyp.js'),
|
||||
'rebuild',
|
||||
`--nodedir=${nodeDir}`,
|
||||
'--',
|
||||
'-I',
|
||||
compileGypi
|
||||
],
|
||||
cwd: stagedDir,
|
||||
stdio: 'inherit',
|
||||
timeoutMs: null
|
||||
const buildDir = await nodeGypRebuild({
|
||||
stagedDir,
|
||||
workDir,
|
||||
nodeDir,
|
||||
staticCxxRuntime: isCompatSlot(slot)
|
||||
})
|
||||
if (result.code !== 0) {
|
||||
throw new Error(`[orcad-prebuilds] node-gyp rebuild failed (status ${result.code})`)
|
||||
}
|
||||
const buildDir = join(stagedDir, 'build', 'Release')
|
||||
if (process.platform === 'win32') {
|
||||
require('./node-pty-job-ownership.cjs').assertRebuiltConptyDeniesMsysBreakaway({
|
||||
nodePtyDir: stagedDir,
|
||||
@@ -248,7 +229,7 @@ async function compileNodePty(sourceDir, slot) {
|
||||
crossHost: false
|
||||
})
|
||||
}
|
||||
return buildDir
|
||||
return { buildDir, nodeDir }
|
||||
}
|
||||
|
||||
function requireSlots(slots) {
|
||||
@@ -310,16 +291,22 @@ async function build() {
|
||||
const slot = slotName()
|
||||
assertCompatSlotHost(slot, { platform: process.platform, arch: process.arch, libc: detectLibc() })
|
||||
const slotDir = join(PREBUILDS_DIR, slot)
|
||||
const buildDir = await compileNodePty(sourceDir, slot)
|
||||
const { buildDir, nodeDir } = await compileNodePty(sourceDir, slot)
|
||||
const compatAddons = isCompatSlot(slot)
|
||||
? await compileCompatAddons({ slot, workDir: join(WORK_DIR, slot), nodeDir })
|
||||
: []
|
||||
|
||||
rmSync(slotDir, { recursive: true, force: true })
|
||||
const files = {}
|
||||
for (const [relative, source] of slotSourceFiles({
|
||||
platform: process.platform,
|
||||
arch: process.arch,
|
||||
buildDir,
|
||||
nodePtyDir: sourceDir
|
||||
})) {
|
||||
for (const [relative, source] of [
|
||||
...slotSourceFiles({
|
||||
platform: process.platform,
|
||||
arch: process.arch,
|
||||
buildDir,
|
||||
nodePtyDir: sourceDir
|
||||
}),
|
||||
...compatAddons
|
||||
]) {
|
||||
if (!existsSync(source)) {
|
||||
throw new Error(`[orcad-prebuilds] ${slot} needs ${relative}, but ${source} is missing`)
|
||||
}
|
||||
|
||||
@@ -29,7 +29,12 @@ import {
|
||||
pinnedNodeRuntimeAsset
|
||||
} from '../../src/shared/node-runtime-pin.ts'
|
||||
import { ORCAD_PREBUILDS_DIR } from './build-orcad-prebuilds.mjs'
|
||||
import { findSlotProblems, readManifest } from './orcad-prebuild-slot-contents.mjs'
|
||||
import {
|
||||
COMPAT_SLOT_ADDONS,
|
||||
findCompatAddonGaps,
|
||||
findSlotProblems,
|
||||
readManifest
|
||||
} from './orcad-prebuild-slot-contents.mjs'
|
||||
import { runProcessSync } from './script-child-process.mjs'
|
||||
import { verifyPackagedOrcadTemplate } from './verify-packaged-orcad-template.cjs'
|
||||
|
||||
@@ -121,8 +126,8 @@ export function requestedTemplateTargets(argv = process.argv) {
|
||||
}
|
||||
|
||||
/**
|
||||
* A compat target (design D6 rung B) is its base target's package with the compat node-pty
|
||||
* slot and runtime marker swapped in; everything else is target-independent or libc-static.
|
||||
* A compat target (design D6 rung B) is its base target's package with the compat slot's addons
|
||||
* and runtime marker swapped in; everything else is target-independent or libc-static.
|
||||
* Omitted, not failed, when this build has no compat slot: rung B then refuses as unavailable.
|
||||
*/
|
||||
function stageCompatTarget(compat, basePackageDir) {
|
||||
@@ -136,12 +141,21 @@ function stageCompatTarget(compat, basePackageDir) {
|
||||
return null
|
||||
}
|
||||
const destination = join(outputDir, ORCAD_TEMPLATE_TARGETS_DIR, compat)
|
||||
const slotFiles = new Map(
|
||||
orcadNodePtySlotFiles(compat).map((file) => [
|
||||
const slotFiles = new Map([
|
||||
...orcadNodePtySlotFiles(compat).map((file) => [
|
||||
`${ORCAD_NODE_PTY_DIR}/build/Release/${file}`,
|
||||
join(ORCAD_PREBUILDS_DIR, compat, ...file.split('/'))
|
||||
]),
|
||||
...Object.entries(COMPAT_SLOT_ADDONS).map(([file, shipped]) => [
|
||||
shipped,
|
||||
join(ORCAD_PREBUILDS_DIR, compat, ...file.split('/'))
|
||||
])
|
||||
)
|
||||
])
|
||||
// Why fatal: the base binary would pass every check here and fail only on a compat host.
|
||||
const gaps = findCompatAddonGaps(orcadTemplateTargetFilenames(compat), slotFiles)
|
||||
if (gaps.length > 0) {
|
||||
throw new Error(`compat target ${compat} would ship base-target addons: ${gaps.join(', ')}`)
|
||||
}
|
||||
const files = {}
|
||||
for (const filename of orcadTemplateTargetFilenames(compat)) {
|
||||
const staged = join(destination, ...filename.split('/'))
|
||||
|
||||
@@ -29,6 +29,7 @@ import { stageOrcadWindowsProcessTree } from './orcad-windows-process-tree.mjs'
|
||||
import {
|
||||
ORCAD_EMOJI_SHORTCODE_DATASET,
|
||||
ORCAD_FOREIGN_SQLITE_READER_ENTRY,
|
||||
ORCAD_PORT_SCAN_COMMAND_WORKER_ENTRY,
|
||||
ORCAD_NODE_PTY_DIR,
|
||||
ORCAD_NODE_PTY_JS_ARTIFACTS,
|
||||
ORCAD_NODE_RUNTIME_MARKER_FILENAME,
|
||||
@@ -62,6 +63,10 @@ const DAEMON_OUT_FILE = join(OUT_DIR, 'daemon-entry.js')
|
||||
// start this worker from the module dir, since orcad has no Electron resources tree.
|
||||
const FOREIGN_SQLITE_READER_ENTRY = join(ROOT, ORCAD_CHILD_ENTRY_POINTS.foreignSqliteReader)
|
||||
const FOREIGN_SQLITE_READER_OUT_FILE = join(OUT_DIR, ORCAD_FOREIGN_SQLITE_READER_ENTRY)
|
||||
// Why beside orcad.js: workspace port detection runs its probe commands on this worker thread,
|
||||
// and `resolveWorkerEntryPath` looks for it next to the running bundle.
|
||||
const PORT_SCAN_WORKER_ENTRY = join(ROOT, ORCAD_CHILD_ENTRY_POINTS.portScanCommandWorker)
|
||||
const PORT_SCAN_WORKER_OUT_FILE = join(OUT_DIR, ORCAD_PORT_SCAN_COMMAND_WORKER_ENTRY)
|
||||
const OUT_FILE = join(OUT_DIR, 'orcad.js')
|
||||
const BUILD_TARGET = process.env.ORCAD_BUILD_TARGET
|
||||
if (!BUILD_TARGET) {
|
||||
@@ -212,6 +217,7 @@ const childResults = await Promise.all([
|
||||
buildForkedChild(WATCHER_ENTRY, WATCHER_OUT_FILE),
|
||||
buildForkedChild(DAEMON_ENTRY, DAEMON_OUT_FILE),
|
||||
buildForkedChild(FOREIGN_SQLITE_READER_ENTRY, FOREIGN_SQLITE_READER_OUT_FILE),
|
||||
buildForkedChild(PORT_SCAN_WORKER_ENTRY, PORT_SCAN_WORKER_OUT_FILE),
|
||||
...['writer', 'backup'].map((role) =>
|
||||
buildForkedChild(
|
||||
join(ROOT, ORCAD_CHILD_ENTRY_POINTS[role]),
|
||||
|
||||
@@ -38,11 +38,23 @@ export const NODE_NETWORK_E2E_SPEC =
|
||||
'tests/e2e/ssh-browser-network-execution-route.docker.unit.test.ts'
|
||||
export const LOCALHOST_SSH_E2E_SPEC = 'tests/e2e/ssh-localhost.spec.ts'
|
||||
export const NATIVE_IME_E2E_SPEC = 'tests/e2e/terminal-ibus-hangul-native.spec.ts'
|
||||
// Needs the packaged orcad slot, which only its own job builds.
|
||||
export const ORCAD_SERVE_MODE_SWITCH_E2E_SPEC = 'tests/e2e/orcad-serve-mode-switch.spec.ts'
|
||||
// Needs the orcad template for its host's target, which only its own job builds.
|
||||
export const ORCAD_AUTO_CONVERT_E2E_SPEC = 'tests/e2e/ssh-orcad-auto-convert.spec.ts'
|
||||
// Windows-only; its own job runs it on a Windows runner.
|
||||
export const WINDOWS_MISSING_APPDATA_E2E_SPEC = 'tests/e2e/windows-missing-appdata-startup.spec.ts'
|
||||
// Runs in the auto-convert job, which builds the template it needs.
|
||||
export const ORCAD_IDLE_EXIT_E2E_SPEC = 'tests/e2e/ssh-orcad-idle-exit.spec.ts'
|
||||
export const DEDICATED_E2E_SPECS = [
|
||||
...DOCKER_SSH_E2E_SPECS,
|
||||
NODE_NETWORK_E2E_SPEC,
|
||||
LOCALHOST_SSH_E2E_SPEC,
|
||||
NATIVE_IME_E2E_SPEC
|
||||
NATIVE_IME_E2E_SPEC,
|
||||
ORCAD_SERVE_MODE_SWITCH_E2E_SPEC,
|
||||
ORCAD_AUTO_CONVERT_E2E_SPEC,
|
||||
WINDOWS_MISSING_APPDATA_E2E_SPEC,
|
||||
ORCAD_IDLE_EXIT_E2E_SPEC
|
||||
]
|
||||
const dedicatedSpecs = new Set(DEDICATED_E2E_SPECS)
|
||||
const dockerSpecs = new Set(DOCKER_SSH_E2E_SPECS)
|
||||
@@ -77,7 +89,14 @@ export function classifyE2eJobs(input, sshSourceChanged = 'false') {
|
||||
e2e_needs_build:
|
||||
runChanged ||
|
||||
sshSourceChanged !== 'false' ||
|
||||
specs.some((spec) => dockerSpecs.has(spec) || spec === LOCALHOST_SSH_E2E_SPEC)
|
||||
specs.some(
|
||||
(spec) =>
|
||||
dockerSpecs.has(spec) ||
|
||||
spec === LOCALHOST_SSH_E2E_SPEC ||
|
||||
spec === ORCAD_SERVE_MODE_SWITCH_E2E_SPEC ||
|
||||
spec === ORCAD_AUTO_CONVERT_E2E_SPEC ||
|
||||
spec === ORCAD_IDLE_EXIT_E2E_SPEC
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -1,9 +1,15 @@
|
||||
import { mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs'
|
||||
import { tmpdir } from 'node:os'
|
||||
import { join } from 'node:path'
|
||||
import { dirname, join } from 'node:path'
|
||||
import { afterEach, describe, expect, it } from 'vitest'
|
||||
import { mergeOrcadPrebuildTrees } from './merge-orcad-prebuilds.mjs'
|
||||
import { findSlotProblems, mergeManifest, sha256Of } from './orcad-prebuild-slot-contents.mjs'
|
||||
import {
|
||||
COMPAT_SLOT_ADDONS,
|
||||
findSlotProblems,
|
||||
isCompatSlot,
|
||||
mergeManifest,
|
||||
sha256Of
|
||||
} from './orcad-prebuild-slot-contents.mjs'
|
||||
|
||||
const dirs = []
|
||||
function temp() {
|
||||
@@ -20,15 +26,20 @@ afterEach(() => {
|
||||
/** One CI lane's `out/orcad-prebuilds`: a single slot plus its manifest. */
|
||||
function laneTree(slot, { version = '1.1.0', nodeHeaders = '24.21.0', bytes = slot } = {}) {
|
||||
const dir = temp()
|
||||
mkdirSync(join(dir, slot), { recursive: true })
|
||||
const binary = join(dir, slot, 'pty.node')
|
||||
writeFileSync(binary, bytes)
|
||||
const files = {}
|
||||
// A compat slot also carries its own addons.
|
||||
for (const file of ['pty.node', ...(isCompatSlot(slot) ? Object.keys(COMPAT_SLOT_ADDONS) : [])]) {
|
||||
const binary = join(dir, slot, ...file.split('/'))
|
||||
mkdirSync(dirname(binary), { recursive: true })
|
||||
writeFileSync(binary, bytes)
|
||||
files[file] = sha256Of(binary)
|
||||
}
|
||||
const manifest = mergeManifest(null, {
|
||||
slot,
|
||||
version,
|
||||
napi: 8,
|
||||
nodeHeaders,
|
||||
entry: { napi: 8, files: { 'pty.node': sha256Of(binary) } }
|
||||
entry: { napi: 8, files }
|
||||
})
|
||||
writeFileSync(join(dir, 'manifest.json'), JSON.stringify(manifest))
|
||||
return dir
|
||||
|
||||
@@ -14,6 +14,12 @@ export function nodeServerTestPaths({ artifact = false, crossRuntime = false } =
|
||||
'src/main/sqlite',
|
||||
'src/main/orcad/orcad-entry.test.ts',
|
||||
'src/main/orcad/orcad-push-startup.test.ts',
|
||||
// The orcad server's identity and stop path, which Windows SSH hosts rely on (W2).
|
||||
'src/main/orcad/orcad-instance-lock.test.ts',
|
||||
'src/main/orcad/orcad-process-start-time.test.ts',
|
||||
'src/main/orcad/orcad-stop-request-listener.test.ts',
|
||||
'src/main/orcad/orcad-managed-stop.test.ts',
|
||||
'src/main/orcad/orcad-managed-stop-cancellation.test.ts',
|
||||
// The directory, not a prefix: its siblings are POSIX-host unit tests pr.yml already runs.
|
||||
'src/main/daemon/pty-subprocess/',
|
||||
'src/main/daemon/pty-subprocess-spawn-file-foreground.test.ts',
|
||||
@@ -25,6 +31,9 @@ export function nodeServerTestPaths({ artifact = false, crossRuntime = false } =
|
||||
'src/main/orcad/orcad-packaged-node-pty.integration.test.ts',
|
||||
'src/main/providers/agent-foreground-process-git-bash.win32.test.ts',
|
||||
'src/main/orcad/orcad-node-launcher.integration.test.ts',
|
||||
'src/main/orcad/orcad-stop-request-shutdown.integration.test.ts',
|
||||
'src/main/orcad/orcad-windows-conpty-breakaway.integration.test.ts',
|
||||
'src/main/orcad/orcad-serve-parity.integration.test.ts',
|
||||
'config/scripts/zip-extractor-command.test.mjs'
|
||||
]
|
||||
: []),
|
||||
|
||||
@@ -0,0 +1,100 @@
|
||||
// D7: orcad's update and rollback planning must agree with the CI protocol-crossing facts.
|
||||
import { readFileSync } from 'node:fs'
|
||||
import { join, resolve } from 'node:path'
|
||||
import { describe, expect, it } from 'vitest'
|
||||
import {
|
||||
DAEMON_PROTOCOL_SOURCE_PATH,
|
||||
canAttach,
|
||||
parseDaemonProtocolFacts
|
||||
} from './daemon-protocol-facts.mjs'
|
||||
import { CURRENT_ORCAD_DAEMON_PROTOCOL } from '../../src/main/ssh/orcad-daemon-protocol-crossing'
|
||||
import { assessOrcadRollback, planOrcadUpdate } from '../../src/main/ssh/orcad-update-plan'
|
||||
|
||||
const projectDir = resolve(import.meta.dirname, '../..')
|
||||
const current = parseDaemonProtocolFacts(
|
||||
readFileSync(join(projectDir, DAEMON_PROTOCOL_SOURCE_PATH), 'utf8')
|
||||
)
|
||||
// The release before the newest protocol bump: it speaks one version lower and cannot list ours.
|
||||
const older = {
|
||||
protocolVersion: current.protocolVersion - 1,
|
||||
previousProtocolVersions: current.previousProtocolVersions.filter(
|
||||
(version) => version < current.protocolVersion - 1
|
||||
)
|
||||
}
|
||||
const record = {
|
||||
schemaVersion: 1,
|
||||
active: '0.3.0+new',
|
||||
previous: '0.2.0+old',
|
||||
activatedAt: '2026-01-01T00:00:00.000Z',
|
||||
snapshot: {
|
||||
dirName: 'pre-0.3.0+new-1',
|
||||
takenBeforeVersion: '0.3.0+new',
|
||||
readableByVersion: '0.2.0+old',
|
||||
takenAt: '2026-01-01T00:00:00.000Z'
|
||||
}
|
||||
}
|
||||
const live = (daemonProtocolVersion) => ({
|
||||
liveSessions: 2,
|
||||
startedSinceActivation: 0,
|
||||
daemonProtocolVersion
|
||||
})
|
||||
|
||||
function rollback(target, daemonProtocolVersion) {
|
||||
return assessOrcadRollback({
|
||||
record,
|
||||
snapshotPresent: true,
|
||||
census: live(daemonProtocolVersion),
|
||||
targetDaemonProtocol: target,
|
||||
stateWritesSinceActivation: false
|
||||
})
|
||||
}
|
||||
|
||||
describe('orcad daemon protocol crossing', () => {
|
||||
it('deploys exactly the protocol the working tree declares', () => {
|
||||
expect({
|
||||
protocolVersion: CURRENT_ORCAD_DAEMON_PROTOCOL.protocolVersion,
|
||||
previousProtocolVersions: [...CURRENT_ORCAD_DAEMON_PROTOCOL.previousProtocolVersions]
|
||||
}).toEqual(current)
|
||||
})
|
||||
|
||||
it('keeps terminals on rollback only when the old build lists the new protocol', () => {
|
||||
expect(canAttach(older, current)).toBe(false)
|
||||
expect(rollback(older, current.protocolVersion)).toMatchObject({
|
||||
safety: 'unsafe',
|
||||
code: 'orcad_rollback_strands_live_terminals'
|
||||
})
|
||||
// A daemon preserved from before the activation still speaks the old build's protocol.
|
||||
expect(canAttach(older, older)).toBe(true)
|
||||
expect(rollback(older, older.protocolVersion)).toMatchObject({ safety: 'clean' })
|
||||
const listing = { ...older, previousProtocolVersions: [...older.previousProtocolVersions] }
|
||||
listing.previousProtocolVersions.push(current.protocolVersion + 1)
|
||||
expect(canAttach(listing, { ...current, protocolVersion: current.protocolVersion + 1 })).toBe(
|
||||
true
|
||||
)
|
||||
expect(rollback(listing, current.protocolVersion + 1)).toMatchObject({ safety: 'clean' })
|
||||
})
|
||||
|
||||
it('updates over live terminals only when the candidate can attach their daemon', () => {
|
||||
const plan = (daemonProtocolVersion) =>
|
||||
planOrcadUpdate({
|
||||
record,
|
||||
candidateVersion: '0.4.0+next',
|
||||
census: live(daemonProtocolVersion),
|
||||
candidateDaemonProtocol: current,
|
||||
force: true
|
||||
})
|
||||
expect(canAttach(current, older)).toBe(true)
|
||||
expect(plan(older.protocolVersion)).toMatchObject({
|
||||
action: 'proceed',
|
||||
preservesLiveDaemon: true
|
||||
})
|
||||
const dropped = current.protocolVersion + 1
|
||||
expect(canAttach(current, { protocolVersion: dropped, previousProtocolVersions: [] })).toBe(
|
||||
false
|
||||
)
|
||||
expect(plan(dropped)).toMatchObject({
|
||||
action: 'defer',
|
||||
code: 'orcad_update_strands_live_terminals'
|
||||
})
|
||||
})
|
||||
})
|
||||
@@ -0,0 +1,102 @@
|
||||
import { existsSync, readFileSync } from 'node:fs'
|
||||
import { resolve } from 'node:path'
|
||||
import { expect, it } from 'vitest'
|
||||
import { parse } from 'yaml'
|
||||
import {
|
||||
classifyE2eJobs,
|
||||
ORCAD_AUTO_CONVERT_E2E_SPEC,
|
||||
ORCAD_IDLE_EXIT_E2E_SPEC,
|
||||
ORCAD_SERVE_MODE_SWITCH_E2E_SPEC,
|
||||
WINDOWS_MISSING_APPDATA_E2E_SPEC
|
||||
} from './ci-e2e-job-selection.mjs'
|
||||
import { selectPrE2eSpecs, shouldRunReusablePrE2e } from './pr-e2e-source-routing.mjs'
|
||||
|
||||
const root = resolve(import.meta.dirname, '../..')
|
||||
const jobs = parse(readFileSync(resolve(root, '.github/workflows/e2e.yml'), 'utf8')).jobs
|
||||
|
||||
function expectRouted(files, spec) {
|
||||
for (const file of files) {
|
||||
expect(existsSync(resolve(root, file)), file).toBe(true)
|
||||
expect(selectPrE2eSpecs([file]), file).toContain(spec)
|
||||
expect(shouldRunReusablePrE2e([file]), file).toBe(true)
|
||||
}
|
||||
}
|
||||
|
||||
it('routes the mode-switch spec from its serve sources and harness', () => {
|
||||
expectRouted(
|
||||
[
|
||||
'src/main/orcad/orcad-lifecycle.ts',
|
||||
'src/main/daemon/daemon-spawner.ts',
|
||||
'tests/e2e/helpers/orca-serve-cli-host.ts',
|
||||
'tests/e2e/helpers/headless-paired-runtime-host.ts'
|
||||
],
|
||||
ORCAD_SERVE_MODE_SWITCH_E2E_SPEC
|
||||
)
|
||||
})
|
||||
|
||||
it('routes the missing-AppData spec from its startup sources and harness', () => {
|
||||
expectRouted(
|
||||
['src/main/startup/windows-app-data-path.ts', 'tests/e2e/helpers/orca-serve-cli-host.ts'],
|
||||
WINDOWS_MISSING_APPDATA_E2E_SPEC
|
||||
)
|
||||
})
|
||||
|
||||
it('routes the auto-convert spec from its conversion sources and harness', () => {
|
||||
expectRouted(
|
||||
[
|
||||
'src/main/ssh/orcad-runtime-conversion.ts',
|
||||
'tests/e2e/helpers/orcad-convert-flow.ts',
|
||||
'tests/e2e/helpers/orcad-convert-host.ts',
|
||||
'tests/e2e/helpers/orcad-template-variant.ts',
|
||||
'tests/e2e/helpers/orcad-upgrade-profile.ts'
|
||||
],
|
||||
ORCAD_AUTO_CONVERT_E2E_SPEC
|
||||
)
|
||||
expect(selectPrE2eSpecs(['src/main/ssh/orcad-runtime-conversion.test.ts'])).not.toContain(
|
||||
ORCAD_AUTO_CONVERT_E2E_SPEC
|
||||
)
|
||||
})
|
||||
|
||||
it('routes the idle-exit spec from its idle sources and the shared convert harness', () => {
|
||||
expectRouted(
|
||||
[
|
||||
'src/shared/orcad-idle-exit.ts',
|
||||
'tests/e2e/helpers/orcad-convert-flow.ts',
|
||||
'tests/e2e/helpers/orcad-convert-host.ts'
|
||||
],
|
||||
ORCAD_IDLE_EXIT_E2E_SPEC
|
||||
)
|
||||
})
|
||||
|
||||
it('builds the e2e app when only a build-dependent orcad spec is requested', () => {
|
||||
for (const spec of [
|
||||
ORCAD_SERVE_MODE_SWITCH_E2E_SPEC,
|
||||
ORCAD_AUTO_CONVERT_E2E_SPEC,
|
||||
ORCAD_IDLE_EXIT_E2E_SPEC
|
||||
]) {
|
||||
expect(classifyE2eJobs(JSON.stringify([spec])), spec).toEqual({
|
||||
e2e_run_changed: false,
|
||||
e2e_needs_build: true
|
||||
})
|
||||
}
|
||||
expect(jobs['orcad-auto-convert-docker'].needs).toEqual(['build', 'prepare-native-cache'])
|
||||
})
|
||||
|
||||
it('runs the auto-convert lane only when routed, not on every SSH source change', () => {
|
||||
const condition = jobs['orcad-auto-convert-docker'].if
|
||||
for (const spec of [ORCAD_AUTO_CONVERT_E2E_SPEC, ORCAD_IDLE_EXIT_E2E_SPEC]) {
|
||||
expect(condition).toContain(`contains(inputs.test_files, '${spec}')`)
|
||||
}
|
||||
expect(condition).not.toContain('ssh_source_changed')
|
||||
})
|
||||
|
||||
it('runs both Windows serve specs on one runner, each only when requested', () => {
|
||||
expect(jobs['windows-missing-appdata-startup']).toBeUndefined()
|
||||
const job = jobs['orcad-serve-mode-switch-windows']
|
||||
for (const spec of [ORCAD_SERVE_MODE_SWITCH_E2E_SPEC, WINDOWS_MISSING_APPDATA_E2E_SPEC]) {
|
||||
expect(job.if, spec).toContain(`contains(inputs.test_files, '${spec}')`)
|
||||
const step = job.steps.find((candidate) => candidate.run?.includes(spec))
|
||||
expect(step.if, spec).toContain(`contains(inputs.test_files, '${spec}')`)
|
||||
expect(step.env.ORCA_STARTUP_DIAGNOSTICS, spec).toBeUndefined()
|
||||
}
|
||||
})
|
||||
@@ -9,7 +9,8 @@ export const ORCAD_CHILD_ENTRY_POINTS = {
|
||||
daemon: 'src/main/daemon/daemon-entry.ts',
|
||||
writer: 'src/main/persistence/profile-state/profile-state-writer-worker-entry.ts',
|
||||
backup: 'src/main/persistence/profile-state/profile-state-backup-worker-entry.ts',
|
||||
foreignSqliteReader: 'src/main/foreign-sqlite-readers/foreign-sqlite-reader-entry.ts'
|
||||
foreignSqliteReader: 'src/main/foreign-sqlite-readers/foreign-sqlite-reader-entry.ts',
|
||||
portScanCommandWorker: 'src/main/ports/port-scan-command-worker-entry.ts'
|
||||
}
|
||||
|
||||
export const ORCAD_EXTERNAL_MODULES = ['electron', 'node-pty', '@parcel/watcher', 'fsevents']
|
||||
|
||||
@@ -0,0 +1,52 @@
|
||||
/**
|
||||
* The addons a compat slot builds beside node-pty (design D6 rung B). The default slots take
|
||||
* @parcel/watcher's upstream prebuild, which needs a newer libstdc++ than a glibc 2.17 host has,
|
||||
* so the compat slot compiles it from the package's own sources with the C++ runtime static.
|
||||
*/
|
||||
import { cpSync, mkdirSync, rmSync } from 'node:fs'
|
||||
import { createRequire } from 'node:module'
|
||||
import { dirname, join } from 'node:path'
|
||||
import { NODE_RUNTIME_PIN } from '../../src/shared/node-runtime-pin.ts'
|
||||
import { nodeGypRebuild, stageNodeAddonApi } from './orcad-prebuild-node-gyp.mjs'
|
||||
import { COMPAT_SLOT_ADDONS, SLOT_NAPI_VERSION } from './orcad-prebuild-slot-contents.mjs'
|
||||
|
||||
const require = createRequire(import.meta.url)
|
||||
|
||||
const BUILDERS = {
|
||||
'parcel-watcher/watcher.node': compileParcelWatcher
|
||||
}
|
||||
|
||||
/** `[slot-relative path, built file]` for every compat addon, compiled under `workDir`. */
|
||||
export async function compileCompatAddons({ slot, workDir, nodeDir }) {
|
||||
const built = []
|
||||
for (const relative of Object.keys(COMPAT_SLOT_ADDONS)) {
|
||||
const builder = BUILDERS[relative]
|
||||
if (!builder) {
|
||||
throw new Error(`[orcad-prebuilds] no builder for compat addon ${relative}`)
|
||||
}
|
||||
built.push([relative, await builder({ slot, workDir, nodeDir })])
|
||||
}
|
||||
return built
|
||||
}
|
||||
|
||||
async function compileParcelWatcher({ slot, workDir, nodeDir }) {
|
||||
const sourceDir = dirname(require.resolve('@parcel/watcher/package.json'))
|
||||
const addonWorkDir = join(workDir, 'parcel-watcher')
|
||||
const stagedDir = join(addonWorkDir, 'watcher')
|
||||
rmSync(addonWorkDir, { recursive: true, force: true })
|
||||
mkdirSync(stagedDir, { recursive: true })
|
||||
for (const entry of ['package.json', 'binding.gyp', 'src']) {
|
||||
cpSync(join(sourceDir, entry), join(stagedDir, entry), { recursive: true })
|
||||
}
|
||||
stageNodeAddonApi(sourceDir, stagedDir)
|
||||
console.log(
|
||||
`[orcad-prebuilds] compiling @parcel/watcher for ${slot} against Node ${NODE_RUNTIME_PIN.version} headers, N-API ${SLOT_NAPI_VERSION} ...`
|
||||
)
|
||||
const buildDir = await nodeGypRebuild({
|
||||
stagedDir,
|
||||
workDir: addonWorkDir,
|
||||
nodeDir,
|
||||
staticCxxRuntime: true
|
||||
})
|
||||
return join(buildDir, 'watcher.node')
|
||||
}
|
||||
@@ -0,0 +1,44 @@
|
||||
// node-gyp rebuild of one staged addon against the pinned Node headers, shared by every slot addon.
|
||||
import { cpSync, writeFileSync } from 'node:fs'
|
||||
import { createRequire } from 'node:module'
|
||||
import { dirname, join } from 'node:path'
|
||||
import process from 'node:process'
|
||||
import { prebuildCompileGypi } from './orcad-prebuild-slot-contents.mjs'
|
||||
|
||||
const require = createRequire(import.meta.url)
|
||||
const ROOT = join(import.meta.dirname, '..', '..')
|
||||
|
||||
/** Copies node-addon-api beside a staged addon, since scratch copies leave the pnpm tree behind. */
|
||||
export function stageNodeAddonApi(sourceDir, stagedDir) {
|
||||
const addonApiDir = dirname(
|
||||
require.resolve('node-addon-api/package.json', { paths: [sourceDir] })
|
||||
)
|
||||
cpSync(addonApiDir, join(stagedDir, 'node_modules', 'node-addon-api'), {
|
||||
recursive: true,
|
||||
dereference: true
|
||||
})
|
||||
}
|
||||
|
||||
export async function nodeGypRebuild({ stagedDir, workDir, nodeDir, staticCxxRuntime }) {
|
||||
const compileGypi = join(workDir, 'prebuild-compile.gypi')
|
||||
writeFileSync(compileGypi, prebuildCompileGypi({ staticCxxRuntime }))
|
||||
const { runProcessSync } = await import('./script-child-process.mjs')
|
||||
const result = runProcessSync({
|
||||
program: process.execPath,
|
||||
args: [
|
||||
join(ROOT, 'node_modules', 'node-gyp', 'bin', 'node-gyp.js'),
|
||||
'rebuild',
|
||||
`--nodedir=${nodeDir}`,
|
||||
'--',
|
||||
'-I',
|
||||
compileGypi
|
||||
],
|
||||
cwd: stagedDir,
|
||||
stdio: 'inherit',
|
||||
timeoutMs: null
|
||||
})
|
||||
if (result.code !== 0) {
|
||||
throw new Error(`[orcad-prebuilds] node-gyp rebuild failed (status ${result.code})`)
|
||||
}
|
||||
return join(stagedDir, 'build', 'Release')
|
||||
}
|
||||
@@ -1,5 +1,6 @@
|
||||
/**
|
||||
* What goes into one orcad node-pty prebuild slot, and the manifest that records it.
|
||||
* What goes into one orcad node-pty prebuild slot (plus a compat slot's own addons), and the
|
||||
* manifest that records it.
|
||||
*
|
||||
* The manifest is the loader's contract (src/main/orcad/node-pty-prebuilt-slot.ts): per-slot
|
||||
* N-API level, libc, the highest glibc symbol version the binaries need, and a sha256 per
|
||||
@@ -33,6 +34,23 @@ export const COMPAT_SLOTS = Object.freeze({
|
||||
'linux-x64-glibc217': Object.freeze({ platform: 'linux', arch: 'x64', libc: 'glibc' })
|
||||
})
|
||||
|
||||
/**
|
||||
* Native addons a compat slot builds beside node-pty, by slot-relative path, mapped to where an
|
||||
* orcad slot ships them. A compat target ships no native file without a compat build.
|
||||
*/
|
||||
export const COMPAT_SLOT_ADDONS = Object.freeze({
|
||||
'parcel-watcher/watcher.node': 'node_modules/@parcel/watcher/watcher.node'
|
||||
})
|
||||
|
||||
/**
|
||||
* The native files of a compat target with no compat build behind them: each would be the base
|
||||
* target's binary, built for a newer glibc and libstdc++ than the compat host has.
|
||||
* `compatSources` maps orcad slot paths to the compat slot files that fill them.
|
||||
*/
|
||||
export function findCompatAddonGaps(targetFilenames, compatSources) {
|
||||
return targetFilenames.filter((file) => file.endsWith('.node') && !compatSources.has(file))
|
||||
}
|
||||
|
||||
export function isCompatSlot(slot) {
|
||||
return Object.hasOwn(COMPAT_SLOTS, slot)
|
||||
}
|
||||
@@ -217,6 +235,13 @@ export function findSlotProblems(manifest, prebuildsDir, requiredSlots) {
|
||||
problems.push(`${slot}: not built`)
|
||||
continue
|
||||
}
|
||||
if (isCompatSlot(slot)) {
|
||||
for (const addon of Object.keys(COMPAT_SLOT_ADDONS)) {
|
||||
if (!Object.hasOwn(entry.files ?? {}, addon)) {
|
||||
problems.push(`${slot}/${addon}: not built`)
|
||||
}
|
||||
}
|
||||
}
|
||||
for (const [file, expected] of Object.entries(entry.files ?? {})) {
|
||||
const path = join(prebuildsDir, slot, ...file.split('/'))
|
||||
if (!existsSync(path)) {
|
||||
|
||||
@@ -5,7 +5,9 @@ import { join } from 'node:path'
|
||||
import { afterEach, describe, expect, it } from 'vitest'
|
||||
import {
|
||||
assertCompatSlotHost,
|
||||
COMPAT_SLOT_ADDONS,
|
||||
COMPAT_SLOTS,
|
||||
findCompatAddonGaps,
|
||||
findPostBaselineNodeApiNames,
|
||||
findSharedCxxRuntimeNeeds,
|
||||
findSlotProblems,
|
||||
@@ -19,7 +21,10 @@ import {
|
||||
SLOT_NAPI_VERSION,
|
||||
windowsConptyRuntimeDir
|
||||
} from './orcad-prebuild-slot-contents.mjs'
|
||||
import { ORCAD_ADDON_NAPI_VERSION } from '../../src/shared/orcad-artifacts.ts'
|
||||
import {
|
||||
ORCAD_ADDON_NAPI_VERSION,
|
||||
orcadTemplateTargetFilenames
|
||||
} from '../../src/shared/orcad-artifacts.ts'
|
||||
|
||||
const floors = createRequire(import.meta.url)('./verify-linux-glibc-floor.cjs')
|
||||
const dirs = []
|
||||
@@ -232,4 +237,37 @@ describe('findSlotProblems', () => {
|
||||
'manifest.json is missing or not schema 2'
|
||||
])
|
||||
})
|
||||
|
||||
it('refuses a compat slot missing one of its own addons', () => {
|
||||
const dir = temp()
|
||||
const slot = 'linux-x64-glibc217'
|
||||
mkdirSync(join(dir, slot))
|
||||
writeFileSync(join(dir, slot, 'pty.node'), 'binary')
|
||||
const manifest = mergeManifest(
|
||||
null,
|
||||
next(slot, { entry: entry({ 'pty.node': sha256Of(join(dir, slot, 'pty.node')) }) })
|
||||
)
|
||||
expect(findSlotProblems(manifest, dir, [slot])).toEqual([
|
||||
`${slot}/parcel-watcher/watcher.node: not built`
|
||||
])
|
||||
})
|
||||
})
|
||||
|
||||
describe('compat addon coverage', () => {
|
||||
const nodePtySources = (target) =>
|
||||
orcadTemplateTargetFilenames(target).filter((file) => file.includes('node-pty/build/Release/'))
|
||||
|
||||
it('gives every native addon a compat target ships a compat build', () => {
|
||||
for (const compat of Object.keys(COMPAT_SLOTS)) {
|
||||
const sources = new Set([...nodePtySources(compat), ...Object.values(COMPAT_SLOT_ADDONS)])
|
||||
expect(findCompatAddonGaps(orcadTemplateTargetFilenames(compat), sources)).toEqual([])
|
||||
}
|
||||
})
|
||||
|
||||
it('names a native file that would ship as the base target build', () => {
|
||||
const target = 'linux-x64-glibc217'
|
||||
expect(
|
||||
findCompatAddonGaps(orcadTemplateTargetFilenames(target), new Set(nodePtySources(target)))
|
||||
).toEqual(['node_modules/@parcel/watcher/watcher.node'])
|
||||
})
|
||||
})
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
const { join } = require('node:path')
|
||||
const { tmpdir } = require('node:os')
|
||||
|
||||
const [nodePtyDir, expectedVersion] = process.argv.slice(2)
|
||||
const [nodePtyDir, expectedVersion, ...addons] = process.argv.slice(2)
|
||||
if (!nodePtyDir || !expectedVersion) {
|
||||
throw new Error('usage: orcad-prebuild-smoke-child.cjs <node-pty dir> <expected node version>')
|
||||
}
|
||||
@@ -11,6 +11,10 @@ if (process.version !== `v${expectedVersion}`) {
|
||||
}
|
||||
|
||||
const pty = require(nodePtyDir)
|
||||
// A compat slot's own addons: loading proves their glibc and C++ runtime needs resolve here.
|
||||
for (const addon of addons) {
|
||||
require(addon)
|
||||
}
|
||||
if (process.platform === 'win32') {
|
||||
// Loaded only by the non-DLL kill path; prove the shipped module still loads under this Node.
|
||||
const { loadNativeModule } = require(join(nodePtyDir, 'lib', 'utils'))
|
||||
|
||||
@@ -3,7 +3,12 @@ import { chmodSync, cpSync, existsSync, mkdirSync, rmSync } from 'node:fs'
|
||||
import { createRequire } from 'node:module'
|
||||
import { dirname, join } from 'node:path'
|
||||
import { NODE_RUNTIME_PIN } from '../../src/shared/node-runtime-pin.ts'
|
||||
import { findSlotProblems, readManifest } from './orcad-prebuild-slot-contents.mjs'
|
||||
import {
|
||||
COMPAT_SLOT_ADDONS,
|
||||
findSlotProblems,
|
||||
isCompatSlot,
|
||||
readManifest
|
||||
} from './orcad-prebuild-slot-contents.mjs'
|
||||
import { ensurePinnedNodeExecutable } from './pinned-node-downloads.mjs'
|
||||
import { runProcessSync } from './script-child-process.mjs'
|
||||
|
||||
@@ -28,6 +33,15 @@ export function stageSmokeNodePty({ slotDir, stageDir }) {
|
||||
return nodePtyDir
|
||||
}
|
||||
|
||||
/** stageSmokeNodePty copies the whole slot into build/Release, compat addons included. */
|
||||
function compatAddonPaths(slot, nodePtyDir) {
|
||||
return isCompatSlot(slot)
|
||||
? Object.keys(COMPAT_SLOT_ADDONS).map((file) =>
|
||||
join(nodePtyDir, 'build', 'Release', ...file.split('/'))
|
||||
)
|
||||
: []
|
||||
}
|
||||
|
||||
export async function runOrcadPrebuildSmoke({ slot, prebuildsDir }) {
|
||||
const problems = findSlotProblems(readManifest(prebuildsDir), prebuildsDir, [slot])
|
||||
if (problems.length > 0) {
|
||||
@@ -43,7 +57,8 @@ export async function runOrcadPrebuildSmoke({ slot, prebuildsDir }) {
|
||||
args: [
|
||||
join(import.meta.dirname, 'orcad-prebuild-smoke-child.cjs'),
|
||||
nodePtyDir,
|
||||
NODE_RUNTIME_PIN.version
|
||||
NODE_RUNTIME_PIN.version,
|
||||
...compatAddonPaths(slot, nodePtyDir)
|
||||
],
|
||||
timeoutMs: 60_000
|
||||
})
|
||||
|
||||
@@ -205,7 +205,9 @@ describe('SSH Windows consumers of qualified server slots', () => {
|
||||
expect(template).toBeLessThan(hosts)
|
||||
expect(sshSteps[template].env.ORCA_REQUIRE_RELAY_NATIVE_ADDONS).toBe('${{ matrix.arch }}')
|
||||
expect(sshSteps[template].run).toContain('--require-slots "win32-${{ matrix.arch }}"')
|
||||
expect(sshSteps[hosts].run).toContain("@('pinned-cmd','pinned-powershell','legacy-opt-out')")
|
||||
expect(sshSteps[hosts].run).toContain(
|
||||
"@('pinned-cmd','pinned-powershell','legacy-opt-out','orcad-cmd','orcad-powershell')"
|
||||
)
|
||||
for (const workflowPaths of [
|
||||
workflow.on.pull_request.paths,
|
||||
sshWorkflow.on.pull_request.paths
|
||||
|
||||
@@ -13,6 +13,48 @@ const NATIVE_IME_HARNESS =
|
||||
/^(?:config\/scripts\/focus-nested-wayland-terminal\.sh$|config\/scripts\/(?:run-terminal-ibus-hangul-e2e|terminal-ime-engagement-receipt)\.mjs$|tests\/e2e\/terminal-ime-(?:boundary-probe|byte-reader|engagement-receipt)\.ts$|tests\/e2e\/terminal-(?:ibus-hangul|hangul-terminating-digit|macos-2set-korean)-native\.spec\.ts$)/
|
||||
|
||||
export const PR_E2E_SOURCE_ROUTES = [
|
||||
{
|
||||
id: 'serve.orcad-mode-switch',
|
||||
specs: ['tests/e2e/orcad-serve-mode-switch.spec.ts'],
|
||||
matches: (file) =>
|
||||
/^tests\/e2e\/helpers\/(?:orca-serve-cli-host|headless-paired-runtime-host)\.ts$/.test(
|
||||
file
|
||||
) ||
|
||||
(isProductSource(file) &&
|
||||
/^src\/(?:cli\/runtime\/(?:launch|serve-)|main\/orcad\/(?:main|orcad-entry|orcad-instance-lock|orcad-command-arguments|orcad-lifecycle)\.ts$|main\/startup\/desktop-profile-instance-lock\.ts$|main\/daemon\/daemon-(?:spawner|endpoint-adoption|init)|main\/server\/serve-)/.test(
|
||||
file
|
||||
))
|
||||
},
|
||||
{
|
||||
id: 'startup.windows-missing-appdata',
|
||||
specs: ['tests/e2e/windows-missing-appdata-startup.spec.ts'],
|
||||
matches: (file) =>
|
||||
file === 'tests/e2e/helpers/orca-serve-cli-host.ts' ||
|
||||
(isProductSource(file) &&
|
||||
/^src\/main\/startup\/(?:windows-app-data-path|main-process-preflight)\.ts$/.test(file))
|
||||
},
|
||||
{
|
||||
id: 'ssh.orcad-auto-convert',
|
||||
specs: ['tests/e2e/ssh-orcad-auto-convert.spec.ts'],
|
||||
matches: (file) =>
|
||||
/^tests\/e2e\/helpers\/(?:orcad-convert-(?:flow|host)|orcad-template-variant|orcad-upgrade-profile)\.ts$/.test(
|
||||
file
|
||||
) ||
|
||||
(isProductSource(file) &&
|
||||
/^src\/main\/(?:ipc\/ssh-host-server-|ssh\/(?:ssh-host-server-|orcad-runtime-conversion|orcad-migration-|orcad-retained-source|orcad-runtime-deployment))/.test(
|
||||
file
|
||||
))
|
||||
},
|
||||
{
|
||||
id: 'ssh.orcad-idle-exit',
|
||||
specs: ['tests/e2e/ssh-orcad-idle-exit.spec.ts'],
|
||||
matches: (file) =>
|
||||
/^tests\/e2e\/helpers\/orcad-convert-(?:flow|host)\.ts$/.test(file) ||
|
||||
(isProductSource(file) &&
|
||||
/^src\/(?:main\/(?:orcad\/orcad-(?:idle-|managed-idle-)|ssh\/orcad-(?:managed-wake|managed-tunnel|recovery-slot|remote-launch))|shared\/orcad-idle-exit)/.test(
|
||||
file
|
||||
))
|
||||
},
|
||||
{
|
||||
id: 'ssh.localhost-agent-hooks',
|
||||
specs: ['tests/e2e/ssh-localhost.spec.ts'],
|
||||
|
||||
@@ -11,6 +11,9 @@ const EXPECTED_MATRIX = {
|
||||
'.github/workflows/e2e.yml#build': { contents: 'read' },
|
||||
'.github/workflows/e2e.yml#changed-e2e': { contents: 'read' },
|
||||
'.github/workflows/e2e.yml#e2e': { contents: 'read' },
|
||||
'.github/workflows/e2e.yml#orcad-auto-convert-docker': { contents: 'read' },
|
||||
'.github/workflows/e2e.yml#orcad-serve-mode-switch': { contents: 'read' },
|
||||
'.github/workflows/e2e.yml#orcad-serve-mode-switch-windows': { contents: 'read' },
|
||||
'.github/workflows/e2e.yml#prepare-native-cache': { contents: 'read' },
|
||||
'.github/workflows/e2e.yml#ssh-browser-network-route': { contents: 'read' },
|
||||
'.github/workflows/e2e.yml#ssh-localhost': { contents: 'read' },
|
||||
|
||||
@@ -4,7 +4,10 @@ import { describe, expect, it } from 'vitest'
|
||||
import { parse } from 'yaml'
|
||||
import {
|
||||
WINDOWS_FORBIDDEN_TOOLS,
|
||||
WINDOWS_HOST_CELL_IDS
|
||||
WINDOWS_HOST_CELL_IDS,
|
||||
WINDOWS_CLI_MATRIX_CELL_IDS,
|
||||
WINDOWS_CONVERT_CELL_ID,
|
||||
WINDOWS_ORCAD_CELL_IDS
|
||||
} from '../../src/main/ssh/ssh-windows-host-cells.ts'
|
||||
|
||||
const projectDir = resolve(import.meta.dirname, '../..')
|
||||
@@ -41,6 +44,10 @@ describe('SSH Windows-host workflow', () => {
|
||||
expect(paths.indexOf('src/main/ssh/ssh-relay-windows-host-lane.test.ts')).toBeGreaterThan(
|
||||
paths.indexOf('!src/**/*.test.ts')
|
||||
)
|
||||
// A later glob would re-include unit tests the Windows lane never runs.
|
||||
for (const glob of paths.slice(paths.indexOf('!src/**/*.test.ts') + 1)) {
|
||||
expect(glob.startsWith('src/') ? glob.endsWith('.test.ts') : true, glob).toBe(true)
|
||||
}
|
||||
expect(job.if).toContain('github.event.pull_request.draft != true')
|
||||
})
|
||||
|
||||
@@ -54,10 +61,18 @@ describe('SSH Windows-host workflow', () => {
|
||||
'x64/windows-2022/inbox',
|
||||
'x64/windows-2022/preview'
|
||||
])
|
||||
expect(job.env).toMatchObject({ ORCA_BACKGROUND_LAUNCH: '1', ORCA_ISOLATED_SSH_CI: '1' })
|
||||
expect(job.env).toMatchObject({
|
||||
ORCA_BACKGROUND_LAUNCH: '1',
|
||||
ORCA_ISOLATED_SSH_CI: '1'
|
||||
})
|
||||
expect(job.strategy['fail-fast']).toBe(false)
|
||||
expect(job['timeout-minutes']).toBe(75)
|
||||
expect(runStep['timeout-minutes']).toBe(50)
|
||||
// The CLI matrix cells, dispatched by name only, get a longer budget.
|
||||
expect(job['timeout-minutes']).toBe(
|
||||
"${{ contains(github.event.inputs.cells || '', 'orcad-cli') && 160 || 75 }}"
|
||||
)
|
||||
expect(runStep['timeout-minutes']).toBe(
|
||||
"${{ contains(github.event.inputs.cells || '', 'orcad-cli') && 130 || 50 }}"
|
||||
)
|
||||
})
|
||||
|
||||
it('overlaps only guarded ARM inbox capability preparation with the existing builds', () => {
|
||||
@@ -102,7 +117,10 @@ describe('SSH Windows-host workflow', () => {
|
||||
"if('${{ matrix.server }}' -eq 'inbox' -and '${{ matrix.arch }}' -eq 'arm64'){$preparation="
|
||||
)
|
||||
expect(runStep.background).toBeUndefined()
|
||||
expect(job.steps.at(-1)).toMatchObject({ if: 'always()', uses: 'actions/upload-artifact@v7' })
|
||||
expect(job.steps.at(-1)).toMatchObject({
|
||||
if: 'always()',
|
||||
uses: 'actions/upload-artifact@v7'
|
||||
})
|
||||
})
|
||||
|
||||
it('shares one capability installer without bypassing native verification or private cleanup', () => {
|
||||
@@ -171,14 +189,42 @@ describe('SSH Windows-host workflow', () => {
|
||||
it('defaults to every cell the TypeScript lane knows', () => {
|
||||
const defaults = /\{\$cells=@\(([^)]*)\)\}/.exec(runStep.run)?.[1]
|
||||
expect(defaults?.split(',').map((id) => id.trim().replaceAll("'", ''))).toEqual([
|
||||
...WINDOWS_HOST_CELL_IDS
|
||||
...WINDOWS_HOST_CELL_IDS,
|
||||
...WINDOWS_ORCAD_CELL_IDS
|
||||
])
|
||||
const invoker = readFileSync(
|
||||
join(projectDir, 'config/ci/windows-ssh-provider/invoke-pinned-relay-cells.ps1'),
|
||||
'utf8'
|
||||
)
|
||||
for (const id of WINDOWS_HOST_CELL_IDS) {
|
||||
for (const id of [...WINDOWS_HOST_CELL_IDS, ...WINDOWS_ORCAD_CELL_IDS]) {
|
||||
expect(invoker).toContain(`'${id}'`)
|
||||
}
|
||||
expect(invoker).toContain('src/main/ssh/orcad-windows-host-lane.test.ts')
|
||||
})
|
||||
|
||||
it('provisions one private account for every cell, convert cell included', () => {
|
||||
expect(runStep.run).toContain(`$cells+='${WINDOWS_CONVERT_CELL_ID}'`)
|
||||
// Only app cells reach a managed server, through an SSH local forward; they run last.
|
||||
const appCells = /\$appCellIds=@\(([^)]*)\)/.exec(runStep.run)?.[1]
|
||||
expect(appCells?.split(',').map((id) => id.trim().replaceAll("'", ''))).toEqual([
|
||||
WINDOWS_CONVERT_CELL_ID,
|
||||
...WINDOWS_CLI_MATRIX_CELL_IDS
|
||||
])
|
||||
expect(runStep.run).toContain(
|
||||
'$forwarding=@($cells | Where-Object {$appCellIds -contains $_}).Count'
|
||||
)
|
||||
expect(runStep.run).toContain('-ForwardingAccounts $forwarding')
|
||||
// A dispatched list may name them anywhere; they still run last.
|
||||
expect(runStep.run).toContain(
|
||||
'$cells=@($cells | Where-Object {$appCellIds -notcontains $_})+@($cells | Where-Object {$appCellIds -contains $_})'
|
||||
)
|
||||
const provisioner = readFileSync(
|
||||
join(projectDir, 'config/ci/windows-ssh-provider/preview-ssh/prove-preview-openssh.ps1'),
|
||||
'utf8'
|
||||
)
|
||||
const max = Number(/\[ValidateRange\(1,(\d+)\)\]\[int\]\$Accounts/.exec(provisioner)?.[1])
|
||||
expect(max).toBeGreaterThanOrEqual(
|
||||
WINDOWS_HOST_CELL_IDS.length + WINDOWS_ORCAD_CELL_IDS.length + 1
|
||||
)
|
||||
})
|
||||
})
|
||||
|
||||
@@ -33,6 +33,8 @@ export const SQLITE_RUNTIME_INCLUDE = [
|
||||
'src/main/codex/codex-structured-question-order.test.ts',
|
||||
'src/main/ipc/pty/ipc/spawn-commit-ssh-lease-cardinality.test.ts',
|
||||
'src/main/ipc/ssh-host-partition-session-export.test.ts',
|
||||
'src/main/ipc/ssh-host-server-on-connect-wiring.test.ts',
|
||||
'src/main/ipc/ssh-managed-server-move-conversion.test.ts',
|
||||
'src/main/native-chat/agent-session-wire/agent-session-history-byte-accounting.test.ts',
|
||||
'src/main/native-chat/agent-session-wire/agent-session-history-conversation-window.test.ts',
|
||||
'src/main/native-chat/agent-session-wire/agent-session-history-forward-read-budget.test.ts',
|
||||
@@ -142,6 +144,7 @@ export const SQLITE_RUNTIME_INCLUDE = [
|
||||
'src/main/native-chat/agent-session-wire/structured-agent-session-wedged-profile-migration.test.ts',
|
||||
'src/main/native-chat/agent-session-wire/structured-agent-session-wire-admission.test.ts',
|
||||
'src/main/native-chat/agent-session-wire/structured-conversation-command.test.ts',
|
||||
'src/main/orcad/orcad-automations.test.ts',
|
||||
'src/main/runtime/agent-session-conversation-clear-commit.test.ts',
|
||||
'src/main/runtime/agent-session-conversation-name-store.test.ts',
|
||||
'src/main/runtime/agent-session-death-evidence-persistence.test.ts',
|
||||
@@ -161,6 +164,7 @@ export const SQLITE_RUNTIME_INCLUDE = [
|
||||
'src/main/runtime/orchestration/orchestration-party-location.test.ts',
|
||||
'src/main/runtime/orchestration/structured-worker-journal-page.test.ts',
|
||||
'src/main/runtime/rpc/methods/agent-launch-caller-selection.test.ts',
|
||||
'src/main/runtime/rpc/methods/agent-launch-instant-tab.test.ts',
|
||||
'src/main/runtime/rpc/methods/agent-launch-pane-reservation.test.ts',
|
||||
'src/main/runtime/rpc/methods/agent-launch-prestart-failure.test.ts',
|
||||
'src/main/runtime/rpc/methods/agent-launch-replay.test.ts',
|
||||
@@ -186,6 +190,14 @@ export const SQLITE_RUNTIME_INCLUDE = [
|
||||
'src/main/runtime/structured-claude-pending-rewind.test.ts',
|
||||
'src/main/ssh-expired-lease-pane-readoption.test.ts',
|
||||
'src/main/ssh-reattach-pane-cardinality.test.ts',
|
||||
'src/main/ssh/orcad-migration-cutover-coordinator.test.ts',
|
||||
'src/main/ssh/orcad-migration-delta-move.test.ts',
|
||||
'src/main/ssh/orcad-migration-delta-snapshot.test.ts',
|
||||
'src/main/ssh/orcad-migration-snapshot-resume.test.ts',
|
||||
'src/main/ssh/orcad-migration-source-fence.test.ts',
|
||||
'src/main/ssh/orcad-runtime-conversion.test.ts',
|
||||
'src/main/ssh/orcad-unreachable-setup-release.test.ts',
|
||||
'src/main/ssh/ssh-target-orcad-preflight.test.ts',
|
||||
'src/main/worktree-identity-persistence.test.ts',
|
||||
'src/main/worktree-removal-close-records.test.ts',
|
||||
'src/main/worktree-removal-session-partition-fencing.test.ts',
|
||||
|
||||
@@ -102,6 +102,7 @@
|
||||
"../src/main/windows-process-tree-kill.ts",
|
||||
"../src/main/windows-pty-root-identity.ts",
|
||||
"../src/main/windows/windows-process-table-cim-scan.ts",
|
||||
"../src/main/windows/windows-process-table-timeout-error.ts",
|
||||
"../src/main/daemon/daemon-process-start-time.ts",
|
||||
"../src/main/daemon/daemon-process-identity-query.ts",
|
||||
"../src/main/startup/startup-diagnostics.ts",
|
||||
|
||||
@@ -66,6 +66,25 @@ the daemon shares the service cgroup and a combined-unit stop ends live terminal
|
||||
it is populated from `/proc/self/cgroup`, so it reports the isolation the daemon actually has
|
||||
rather than what the launcher intended.
|
||||
|
||||
## `orca serve` on this machine
|
||||
|
||||
`orca serve` runs on the local orcad slot by default. The CLI asks the app's
|
||||
`out/main/orcad/orcad-local-serve-selection-entry.js` (run as plain Node on the app's executable)
|
||||
which host to use. Any reason orcad cannot serve falls back to Electron serve with one
|
||||
`[serve] using Electron serve: <reason>` line on stderr. Those reasons are: no slot for this host,
|
||||
no template in the install, the pinned Node could not be fetched, or a failed native preflight.
|
||||
|
||||
- `ORCA_SERVE_RUNTIME=electron` keeps Electron serve and skips the question. `orcad` (or unset) is
|
||||
the default, and any other value falls back with a reason.
|
||||
- Packaged macOS stays on Electron: only Electron serve, supervised by the CLI, can take a remote
|
||||
app update there, and orcad has no updater. Recipe-JSON serve has no handoff and uses orcad.
|
||||
- Windows serves on orcad too. Both hosts share `<userData>\daemon`, so the daemon pipe name
|
||||
(hashed from that path) is the same, and the relocated Electron daemon host changes only the
|
||||
executable, not the pipe. The `orcad-serve-mode-switch-windows` e2e job checks D7 there, in the
|
||||
daily run and on PRs routed to it; it does not block merges.
|
||||
|
||||
The slot and its pinned Node live under the desktop's `<userData>/orcad-artifacts`.
|
||||
|
||||
## Bind policy
|
||||
|
||||
`--bind <literal-ip>`, **default `127.0.0.1`**.
|
||||
@@ -84,6 +103,16 @@ nothing can reach.
|
||||
Under the shipping design a client reaches a remote orcad over an SSH local port-forward, so
|
||||
loopback is the correct default and the pairing credential travels over SSH.
|
||||
|
||||
A host whose sshd refuses forwarding (`AllowTcpForwarding no`) is reached through the stdio
|
||||
bridge instead: the client keeps the same local port, and each connection to it opens one SSH
|
||||
exec channel running a small script on the host's pinned Node that dials orcad's loopback port.
|
||||
Windows hosts run it as the host script's `stdio-bridge` op and frame bytes as base64 lines,
|
||||
because a PowerShell DefaultShell re-decodes native output. Bridges are capped below OpenSSH's
|
||||
default `MaxSessions` of 10 per connection; further connections wait for a free one. The choice
|
||||
is made each time the tunnel starts (`orcad-managed-tunnel-transport.ts`), so nothing is
|
||||
recorded per host, and only a host where even the bridge cannot run keeps the relay, recorded as
|
||||
`ssh_tunnel_unavailable`.
|
||||
|
||||
## Data root and the instance lock
|
||||
|
||||
The data root is `$ORCA_USER_DATA`, else `$XDG_DATA_HOME/Orca`, else `~/.orca`.
|
||||
@@ -102,11 +131,16 @@ It refuses to start when:
|
||||
A root that is merely too permissive and that we own is tightened to `0700` rather than
|
||||
refused — orcad stores credentials there unsealed (no OS keyring on this host), so the goal
|
||||
is a private root, and refusing when we could just fix it helps nobody. We refuse when the
|
||||
permissions are not ours to fix. Windows is exempt from the owner and mode checks: ACLs are
|
||||
not expressible as a POSIX mode, and `statSync().mode` there reports a synthesized one.
|
||||
permissions are not ours to fix. Windows has no owner or mode check, because ACLs are not
|
||||
expressible as a POSIX mode and `statSync().mode` there reports a synthesized one. Instead
|
||||
orcad restricts the root's ACL to its own user with `icacls` (the same verified restriction
|
||||
`secure-file.ts` applies to credential files) and refuses with `orcad_data_root_shared` when
|
||||
that cannot be applied.
|
||||
|
||||
A dead holder's record is reclaimed (PID plus process start time, so a recycled PID does not
|
||||
read as alive). A record belonging to a different identity is never reclaimed.
|
||||
read as alive). On Windows the start time is the kernel creation time read through the
|
||||
process-tree addon the slot stages; without the addon it is null and the PID alone fences,
|
||||
which errs toward "held". A record belonging to a different identity is never reclaimed.
|
||||
|
||||
**The lock scopes one role — who is the runtime.** It deliberately says nothing about the
|
||||
daemon, which lives under `<data-root>/daemon` and fences its own endpoint with its own PID
|
||||
@@ -150,6 +184,47 @@ An external supervisor (systemd, launchd, a process manager). orcad conforms to
|
||||
exits with code 1 if teardown stalls. The bundled runtime also stops gracefully if its
|
||||
launcher's IPC channel closes. On POSIX, both the launcher and runtime ignore `SIGHUP`,
|
||||
so terminal hangups do not stop a headless host. Use `SIGTERM` or `SIGINT` to stop it.
|
||||
- **Stop requests.** A file stops orcad the same way `SIGTERM` does, without a PID that may
|
||||
since have been reused by another process:
|
||||
- `.orcad-stop-request` beside `orcad.js` in the running slot. orcad deletes it and stops.
|
||||
- An instance-bound request in the data root, named
|
||||
`.orcad-managed-stop-request.<sha256 of the instance lock nonce>`. orcad stops only when it
|
||||
names this orcad's version, runtime ID, PID, start time and lock nonce, and while the
|
||||
instance lock still holds that record. The file is kept as evidence.
|
||||
- `orcad --complete-managed-stop '<request JSON>'` writes that request, waits for the
|
||||
instance to exit, and prints one JSON line whose `verdict` is `live`, `unverifiable` or
|
||||
`exited`. `exited` needs proof: no process with that PID, or a PID whose start time shows
|
||||
it now belongs to another process. On `exited` it writes
|
||||
`<data-root>/orcad-stop-receipts/<transactionId>.json`. It exits 0 whenever it printed a
|
||||
verdict, 64 for a malformed invocation, and 1 for a failure before any verdict, which is
|
||||
never evidence of exit.
|
||||
- A request with `retireIdleDaemon: true` asks orcad to retire the terminal daemon too. This
|
||||
is best effort and never blocks or fails the stop:
|
||||
- The daemon is retired only when it proves it owns no live session across every
|
||||
generation.
|
||||
- A busy daemon (`live`) or one whose state cannot be proven (`unverifiable`) stays up with
|
||||
its terminals, and orcad reopens new-terminal admission before exiting.
|
||||
- The completed-stop receipt records `retirement` as `retired`, `live` or `unverifiable`.
|
||||
If orcad exits without recording an outcome, the receipt says `unverifiable`.
|
||||
- `orcad --cancel-managed-stop '<request JSON>'` withdraws a request orcad has not acted on.
|
||||
orcad and the canceller each try to create `<transactionId>.decision.json` exclusively,
|
||||
so exactly one wins. `canceled` means orcad keeps running and the request file is removed;
|
||||
`dispatched` means orcad already began stopping, and only the completion can say how it
|
||||
ended.
|
||||
- A build advertises all of the above with `health.stopRequests: 1` in its readiness line.
|
||||
Clients stop such a build through the slot request file and older builds with `SIGTERM`,
|
||||
after corroborating the PID with readiness either way. A launch clears a slot request
|
||||
that the previous process never consumed.
|
||||
- **Decommissioning a managed slot.** An Orca client decommissions through the same activation
|
||||
journal and fence as deploy and rollback. It refuses while the terminal census reports live
|
||||
or uncounted terminals, stops the instance with a managed request that also asks to retire
|
||||
the daemon, and records that no version is active only after `exited` is proven. A stop
|
||||
that did not finish is cancelled; if orcad already acted on it, or the host cannot answer,
|
||||
the fence stays for recovery.
|
||||
- **Instance lock.** `<data-root>/orcad.lock` names the running orcad. A record that is
|
||||
unreadable, malformed or over 64 KiB is never reclaimed: orcad exits 78 until an operator
|
||||
removes it. A shutdown whose teardown failed keeps the lock until the process exits, so a
|
||||
second orcad cannot start beside a writer that may still be running.
|
||||
- **Exit codes.**
|
||||
|
||||
| Code | Meaning | Supervisor should |
|
||||
@@ -202,6 +277,65 @@ above, stop orcad, then stop the daemon named by `health.terminalDaemon.pid`.
|
||||
Only report it `exited` after verification on the execution host; loss of contact is
|
||||
`unverifiable`.
|
||||
|
||||
### Windows hosts
|
||||
|
||||
What differs on a Windows SSH host, and what deliberately does not:
|
||||
|
||||
- **Stop path.** A signal is TerminateProcess on Windows: no flush, no lock release. The
|
||||
slot's `.orcad-stop-request` file (and the managed, instance-bound request) is therefore the
|
||||
only graceful stop. A detached orcad receives no console control events, so the listener
|
||||
(`fs.watch` plus a one-second poll) is what stops it; the packaged-slot test proves it exits
|
||||
cleanly within the 15 s shutdown deadline on every server lane, Windows included.
|
||||
- **Exit proof.** `--complete-managed-stop` proves a reused PID by the addon's creation time.
|
||||
Without the addon a live PID stays `live` or `unverifiable`, never `exited`.
|
||||
- **Daemon endpoint.** The terminal daemon listens on a named pipe
|
||||
(`\\?\pipe\orca-terminal-host-v<protocol>-<suffix>`), not a socket under the data root.
|
||||
- **Leaving sshd's job.** orcad is started outside the SSH session's kill-on-close job, so the
|
||||
daemon it forks inherits no such job and outlives the connection the same way.
|
||||
- **Per-PTY jobs.** Each ConPTY child gets its own job (`windows-pty-job.ts`), and Git Bash /
|
||||
MSYS panes follow [`windows-msys-job-breakaway.md`](./windows-msys-job-breakaway.md)
|
||||
unchanged. A ConPTY smoke test runs inside a process started exactly that way (breakaway,
|
||||
no window) on the Windows server lanes.
|
||||
- **No daemon-host relocation.** The desktop copies its runtime to `%LOCALAPPDATA%` because
|
||||
the NSIS updater deletes the install directory under a running daemon
|
||||
([`windows-daemon-host-relocation.md`](./windows-daemon-host-relocation.md)). orcad slots are
|
||||
versioned directories that nothing deletes while a process runs from them: Windows refuses
|
||||
to delete a running image, and GC treats an in-use slot as live.
|
||||
|
||||
## Idle exit (client-managed orcad only)
|
||||
|
||||
An orcad that a desktop client launched over SSH stops itself, like the relay, once its host has
|
||||
been unused for 15 minutes. The client's launch sets `ORCA_ORCAD_MANAGED_ACTIVATION_ROOT`; an
|
||||
orcad started by hand, by a supervisor, or as a paired server never carries it and never idles
|
||||
out.
|
||||
|
||||
"Unused" means every one of these held on every check for the whole period:
|
||||
|
||||
- no client socket open and no RPC request running;
|
||||
- no terminal in the PTY provider, and the daemon answered with zero live sessions (a daemon
|
||||
that does not answer keeps orcad up);
|
||||
- no agent reporting `working`;
|
||||
- no staged migration into this server;
|
||||
- no enabled automation and no automation run still in flight (nothing on the host would start
|
||||
orcad again for the next scheduled run, so a server with an enabled automation never idles out);
|
||||
- no activation fence on the host (an update, rollback, decommission or recovery in flight).
|
||||
|
||||
The stop is the ordinary graceful shutdown, which disconnects from the daemon and never shuts it
|
||||
down, so it cannot kill a terminal. It then asks the daemon to retire only if the daemon itself
|
||||
proves it holds no session. Before stopping, orcad writes `<data-root>/orcad-idle-stop.json`;
|
||||
the next start reports it once as `health.previousIdleStop` and removes it, so a later crash is
|
||||
never read as an idle stop. A managed start with no record reports `previousIdleStop: null`.
|
||||
|
||||
The client starts a stopped server again, whatever stopped it (an idle stop, a kill, a host
|
||||
reboot): on every connect, on every fresh tunnel (including after the client wakes from sleep),
|
||||
and before a call through an environment the client restored at launch. A server that does not
|
||||
answer is checked on the host; only a proven exit starts the activated slot, under the activation
|
||||
fence, and the status line shows "Starting managed server…". A daemon that survived is adopted
|
||||
with its terminals; after a reboot both start fresh. A process that is live or cannot be proven
|
||||
gone is left alone, and a start that fails keeps the host managed with the reason and orcad.log's
|
||||
tail, never as a verdict about its terminals. `ORCA_E2E_ORCAD_IDLE_TIMEOUT_MS` shortens the idle
|
||||
period for tests; the client forwards it to the servers it launches.
|
||||
|
||||
## Health
|
||||
|
||||
The readiness payload carries a `health` object:
|
||||
|
||||
@@ -42,9 +42,9 @@ Reconnect re-attaches to the same live PTYs and replays a bounded buffer (`REPLA
|
||||
|
||||
## Updating Orca strands relay-backed terminals
|
||||
|
||||
There is a third outcome that is neither of the two above, and the vocabulary matters: the work does not stop, it becomes permanently unreachable.
|
||||
There is a third outcome that is neither of the two above, and the vocabulary matters: the work does not stop, it becomes unreachable through the new build's own relay.
|
||||
|
||||
The relay's install directory — and therefore its socket path — is namespaced by a content hash of the relay bundle (`computeRemoteRelayDir` in `src/main/ssh/ssh-relay-versioned-install.ts`, consumed by `resolveRemoteInstallState` in `src/main/ssh/ssh-relay-deploy.ts`), and the daemon refuses any client whose bundle hash differs (`handleDaemonHandshakeFrame` in `src/relay/relay-handshake.ts`, exit `EXIT_CODE_VERSION_MISMATCH` 42). Two builds whose relay protocol is byte-identical still refuse each other. So the first reconnect after an app update deploys a new relay at a path the incumbent was never listening on, and cannot reach it even in principle. Every PTY the incumbent owns is `unverifiable` — running, unreachable, and never `exited`. The client's leases are attempted against a relay that never minted their ids, expired on the not-found answer (`handlePtyReattachFailure` in `src/main/ssh/ssh-relay-session.ts`), and the pane falls back to a cold-restore agent resume — or to a bare shell when no resumable provider session was captured for it. The old relay keeps its directory pinned against GC, because its socket really is live (`hasLiveRelaySocket` in `src/main/ssh/remote-install-gc.ts`). See #13852.
|
||||
The relay's install directory — and therefore its socket path — is namespaced by a content hash of the relay bundle (`computeRemoteRelayDir` in `src/main/ssh/ssh-relay-versioned-install.ts`, consumed by `resolveRemoteInstallState` in `src/main/ssh/ssh-relay-deploy.ts`), and the daemon refuses any client whose bundle hash differs (`handleDaemonHandshakeFrame` in `src/relay/relay-handshake.ts`, exit `EXIT_CODE_VERSION_MISMATCH` 42). Two builds whose relay protocol is byte-identical still refuse each other. So the first reconnect after an app update deploys a new relay at a path the incumbent was never listening on, and this build's own bridge cannot reach the incumbent. The incumbent's own bridge can: its version directory still holds its `relay.js` and `.version`, so `relay.js --connect` run from there presents its own hash by construction. Each deploy takes a census of this target's older endpoints (`startPreviousRelayCensus` in `src/main/ssh/ssh-previous-relay-terminals.ts`); while one is live or unverifiable, a not-found reattach keeps its lease and answers `SSH_PTY_HELD_BY_PREVIOUS_RELAY` instead of expiring. On POSIX hosts the pane is then reattached through that older bridge (`SshLegacyRelayRoute` in `src/main/ssh/ssh-legacy-relay-route.ts`), which takes the PTY owner role without output flow control, and every later operation on that PTY id is routed there (`src/main/providers/ssh-pty-legacy-relay-delegation.ts`). When the last pane it serves exits, the route hangs up and the incumbent's own grace retires it. On Windows the census probes each older version directory's pipe for the target (`src/main/ssh/ssh-previous-relay-windows-census.ts`), and a census that cannot run or does not know the host platform counts as `unverifiable`, never as "no older relay". The connect-time host census, which decides whether a host may convert to a managed server before any relay session exists, instead lists every `orca-relay-*` pipe on a Windows machine and maps each to the relay instance that owns it through that instance's credential file or pipe marker, whichever desktop launched it (`src/main/ssh/ssh-host-relay-windows-inventory.ts`); a pipe no directory accounts for counts as `unverifiable` unless the host proves it another account's. The migration terminal gate (`assessOrcadMigrationTerminals`) answers `exited` only from a complete inventory: every relay session answered, or, with no session, a host census proved the endpoints idle. A missing or failed inventory is `unverifiable` even when this desktop leases nothing. Where no route can be opened — Windows named pipes, a socket relocated under the short `/tmp` base, a bridge that fails — the PTY stays `unverifiable`: the pane says it is still running under the previous Orca version, and is never respawned. The old relay keeps its directory pinned against GC while its socket is live (`hasLiveRelaySocket` in `src/main/ssh/remote-install-gc.ts`). See #13852: on POSIX hosts that is no longer a dead end, because a terminal held by a previous relay version now resumes in its pane through that relay's own bridge.
|
||||
|
||||
The peer model does not have this failure, and that is the concrete reason behind "One host, one model" below. The daemon's endpoint is namespaced by a **semantic protocol version** rather than a build (`daemon-v<N>.sock`, from `getDaemonSocketPath` in `src/main/daemon/daemon-spawner.ts`), every earlier protocol version stays attachable (`PROTOCOL_VERSION` in `src/main/daemon/daemon-protocol-version.ts`), and a daemon holding live sessions is preserved across a version change instead of replaced (`shouldPreserveDaemonWithLiveSessions` in `src/main/daemon/daemon-replacement-preflight.ts`).
|
||||
|
||||
|
||||
@@ -396,6 +396,16 @@ built without the addon or a job that refuses breakaway, and a refusal there is
|
||||
reported as `ORCA_RELAY_LAUNCH_REFUSED`. The Windows SSH-host lanes run with no
|
||||
WMI grant and assert the breakaway route.
|
||||
|
||||
`orcad.js` exposes the same launcher (`src/shared/windows-breakaway-launcher.ts`)
|
||||
with **no** WMI fallback: a host that cannot break away refuses the orcad launch.
|
||||
Windows orcad operations start **no PowerShell**: sshd's DefaultShell runs the
|
||||
pinned `node.exe` directly with plain path arguments, against one content-addressed
|
||||
host script staged beside the slots (`src/main/ssh/orcad-windows-host-script.ts`).
|
||||
Only a `node.exe` path that itself needs quoting (a profile name with a space)
|
||||
falls back to one unencoded `powershell.exe -Command`. The launch waits for
|
||||
readiness host-side, one exec per 20 s at most, rather than re-running an exec
|
||||
every 500 ms.
|
||||
|
||||
## Signing is not the gate
|
||||
|
||||
The most useful calibration in the whole incident set came from the reporter's
|
||||
|
||||
@@ -167,6 +167,8 @@ orca-ide serve --port 6768 --pairing-address 100.64.1.20
|
||||
|
||||
Use only one host mode at a time. If the Orca desktop app is already sharing that computer, do not start a second `orca serve` process for the same setup.
|
||||
|
||||
`orca serve` runs on Orca's own Node server (orcad) by default. If orcad can't serve on this computer, it uses the desktop app's server instead and prints one line saying why. On a packaged macOS app it currently always uses the desktop app's server, so paired clients can still update it. To always use the desktop app's server, set `ORCA_SERVE_RUNTIME=electron`.
|
||||
|
||||
### Mobile from a headless server
|
||||
|
||||
For the Orca mobile app, request a mobile-scoped QR code and link:
|
||||
|
||||
@@ -25,7 +25,7 @@ The categories of behavior we observe:
|
||||
- **Agent errors** — a coarse error category and which agent kind was involved. We never see raw error messages or stack traces; per-incident detail stays in a local diagnostic trace file on your machine and only reaches Orca if you explicitly share a diagnostic bundle.
|
||||
- **Settings** — when you toggle one of a small whitelisted set of feature-flag or UX preferences. We record which preference changed and whether it's a boolean or an enum, never the raw value of any free-form setting.
|
||||
- **Privacy controls** — when you opt in or out of telemetry, so we can tell from aggregate data whether our consent UI is working.
|
||||
- **SSH remote runtime** — once per SSH host per app session, which runtime Orca used to run on that host and why: coarse host facts (operating system, CPU architecture, C library family and glibc minor version, and the host's Node.js major version when Orca runs on it), the category of any refusal, whether a runtime was uploaded, and a coarse duration bucket. Never the hostname, user name, paths, or raw error text.
|
||||
- **SSH remote runtime** — once per SSH host and outcome per app session, which runtime Orca used to run on that host and why, or that its runtime check could not be confirmed or failed: coarse host facts (operating system, CPU architecture, C library family and glibc minor version, and the host's Node.js major version when Orca runs on it), the category of any refusal (including security software blocking the runtime on Windows), whether a runtime was uploaded, and a coarse duration bucket. Never the hostname, user name, paths, or raw error text.
|
||||
|
||||
Fields include fixed enum values, version strings, numeric counts and revisions, and random local IDs. Conversation IDs allow repeated token summaries for the same conversation to be associated; they are pseudonymous. No free-form strings from any UI input ever leave your machine.
|
||||
|
||||
|
||||
@@ -10,6 +10,7 @@ import {
|
||||
CLI_MAIN_ENTRY_NAMES,
|
||||
createPlainNodeEntryGuardPlugin
|
||||
} from './config/build-plugins/plain-node-entry-guard'
|
||||
import { ORCAD_LOCAL_SERVE_SELECTION_ENTRY } from './src/shared/orcad-local-serve-selection'
|
||||
import packageJson from './package.json' with { type: 'json' }
|
||||
|
||||
const BUNDLED_MAIN_DEPENDENCIES = new Set([
|
||||
@@ -269,6 +270,10 @@ export const electronViteConfig: UserConfig = {
|
||||
// Why: forked with ELECTRON_RUN_AS_NODE so @parcel/watcher faults
|
||||
// can't take down the main process (issue #7547).
|
||||
'parcel-watcher-process-entry': resolve('src/main/ipc/parcel-watcher-process-entry.ts'),
|
||||
// Why: `orca serve` runs it under ELECTRON_RUN_AS_NODE so the CLI never bundles orcad prep.
|
||||
[ORCAD_LOCAL_SERVE_SELECTION_ENTRY]: resolve(
|
||||
'src/main/orcad/orcad-local-serve-selection-entry.ts'
|
||||
),
|
||||
// Why: a worker thread survives the macOS 26 AppKit main-thread deadlock
|
||||
// without paying for another Electron process.
|
||||
'main-thread-hang-watchdog-entry': resolve(
|
||||
|
||||
@@ -14,6 +14,17 @@ const COMMAND_SCOPED_FLAG_HELP: Record<string, Record<string, string>> = {
|
||||
reference: '--reference <name> Print one bundled reference by name',
|
||||
references: '--references List the bundled reference names for a topic'
|
||||
},
|
||||
'environment update': {
|
||||
force: '--force Restart over running terminals instead of deferring the update'
|
||||
},
|
||||
'environment recover': {
|
||||
'accept-changed-state':
|
||||
'--accept-changed-state Restore the prelaunch snapshot over state a rejected build changed',
|
||||
yes: '--yes Confirm discarding what the rejected build changed'
|
||||
},
|
||||
'environment stop': {
|
||||
yes: '--yes Confirm stopping the server and unlinking it from this machine'
|
||||
},
|
||||
'file open': {
|
||||
focus: FILE_OPEN_FOCUS_HELP
|
||||
},
|
||||
|
||||
@@ -217,6 +217,18 @@ export const HANDLER_GROUPS: readonly HandlerGroup[] = [
|
||||
],
|
||||
load: async () => (await import('./handlers/environment.js')).ENVIRONMENT_HANDLERS
|
||||
},
|
||||
{
|
||||
name: 'managed-server',
|
||||
keys: [
|
||||
'environment status',
|
||||
'environment update',
|
||||
'environment rollback',
|
||||
'environment recover',
|
||||
'environment stop',
|
||||
'environment cancel-stop'
|
||||
],
|
||||
load: async () => (await import('./handlers/managed-server.js')).MANAGED_SERVER_HANDLERS
|
||||
},
|
||||
{
|
||||
name: 'linear',
|
||||
keys: [
|
||||
|
||||
@@ -0,0 +1,27 @@
|
||||
import type { OrcadManagedRuntimeStatus } from '../../shared/orcad-managed-runtime'
|
||||
|
||||
function count(value: number | null): string {
|
||||
return value === null ? 'unverifiable' : String(value)
|
||||
}
|
||||
|
||||
export function formatManagedServerStatus(status: OrcadManagedRuntimeStatus): string {
|
||||
const lines = [
|
||||
`Active version: ${status.activeVersion ?? 'none'}`,
|
||||
`Previous version: ${status.previousVersion ?? 'none'}${status.rollbackAvailable ? ' (rollback available)' : ''}`,
|
||||
`Live terminals: ${count(status.terminals.liveSessions)}`
|
||||
]
|
||||
if (status.recovery) {
|
||||
lines.push(
|
||||
`Interrupted ${status.recovery.operation} of ${status.recovery.version} (${status.recovery.phase}); run \`orca environment recover\`.`
|
||||
)
|
||||
}
|
||||
if (status.migration) {
|
||||
lines.push(`Migration into this server: ${status.migration.phase}`)
|
||||
}
|
||||
if (status.deferredUpdate) {
|
||||
lines.push(
|
||||
`Deferred update to ${status.deferredUpdate.candidateVersion}: ${status.deferredUpdate.reason}`
|
||||
)
|
||||
}
|
||||
return lines.join('\n')
|
||||
}
|
||||
@@ -0,0 +1,50 @@
|
||||
/**
|
||||
* An update or rollback restarts the managed server, and a CLI talking to that runtime then sees
|
||||
* its connection close mid-call. That is expected, not a failure: read the host's status once the
|
||||
* runtime answers again and report what the action left behind.
|
||||
*/
|
||||
import type { OrcadManagedRuntimeStatus } from '../../shared/orcad-managed-runtime'
|
||||
import type { HandlerContext } from '../dispatch'
|
||||
import { printResult } from '../format'
|
||||
import { RuntimeClientError, type RuntimeRpcSuccess } from '../runtime-client'
|
||||
|
||||
const STATUS_ATTEMPTS = 15
|
||||
const STATUS_RETRY_MS = 2_000
|
||||
|
||||
export async function reportAfterClosedConnection(
|
||||
{ client, json }: HandlerContext,
|
||||
selector: { selector: string },
|
||||
sleep: (ms: number) => Promise<void> = (ms) => new Promise((done) => setTimeout(done, ms))
|
||||
): Promise<void> {
|
||||
for (let attempt = 0; attempt < STATUS_ATTEMPTS; attempt++) {
|
||||
await sleep(STATUS_RETRY_MS)
|
||||
let response: RuntimeRpcSuccess<OrcadManagedRuntimeStatus>
|
||||
try {
|
||||
response = await client.call<OrcadManagedRuntimeStatus>('managedServer.status', selector)
|
||||
} catch {
|
||||
continue
|
||||
}
|
||||
const status = response.result
|
||||
if (status.recovery) {
|
||||
throw new RuntimeClientError(
|
||||
'managed_server_interrupted',
|
||||
`The ${status.recovery.operation} of ${status.recovery.version} was interrupted. Run \`orca environment recover\`.`,
|
||||
status
|
||||
)
|
||||
}
|
||||
if (status.deferredUpdate) {
|
||||
throw new RuntimeClientError('managed_server_deferred', status.deferredUpdate.reason, status)
|
||||
}
|
||||
printResult(
|
||||
response,
|
||||
json,
|
||||
(value) =>
|
||||
`The server restarted during the action and now runs ${value.activeVersion ?? 'no version'}.`
|
||||
)
|
||||
return
|
||||
}
|
||||
throw new RuntimeClientError(
|
||||
'managed_server_in_progress',
|
||||
'The connection closed while the server restarted, and it has not answered since. Check `orca environment status`.'
|
||||
)
|
||||
}
|
||||
@@ -0,0 +1,289 @@
|
||||
import { afterEach, describe, expect, it, vi } from 'vitest'
|
||||
import { ORCAD_RECOVERY_CHANGED_STATE_CODE } from '../../shared/orcad-managed-runtime'
|
||||
import { MANAGED_SERVER_RUNTIME_CAPABILITY } from '../../shared/protocol-version'
|
||||
import { RuntimeClientError } from '../runtime-client'
|
||||
import { MANAGED_SERVER_ACTION_TIMEOUT_MS, MANAGED_SERVER_HANDLERS } from './managed-server'
|
||||
|
||||
function envelope(result: unknown) {
|
||||
return { id: 'r', ok: true, result, _meta: { runtimeId: 'runtime-1' } }
|
||||
}
|
||||
|
||||
function client(result: unknown, capabilities = [MANAGED_SERVER_RUNTIME_CAPABILITY]) {
|
||||
return vi.fn(async (method: string) =>
|
||||
method === 'status.get' ? envelope({ capabilities }) : envelope(result)
|
||||
)
|
||||
}
|
||||
|
||||
async function run(
|
||||
command: string,
|
||||
call: ReturnType<typeof client>,
|
||||
flags: [string, string | boolean][]
|
||||
) {
|
||||
const handler = MANAGED_SERVER_HANDLERS[command]
|
||||
if (!handler) {
|
||||
throw new Error(`no handler for ${command}`)
|
||||
}
|
||||
await handler({
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the handlers only call client.call.
|
||||
client: { call } as never,
|
||||
cwd: '/tmp',
|
||||
flags: new Map([['environment', 'build-box'], ...flags]),
|
||||
json: false
|
||||
})
|
||||
}
|
||||
|
||||
afterEach(() => vi.restoreAllMocks())
|
||||
|
||||
describe('managed server CLI verbs', () => {
|
||||
it('refuses on a runtime that does not advertise managed servers, before calling the action', async () => {
|
||||
const call = client({ outcome: 'none' }, [])
|
||||
await expect(run('environment recover', call, [])).rejects.toMatchObject({
|
||||
code: 'incompatible_runtime'
|
||||
})
|
||||
expect(call).toHaveBeenCalledTimes(1)
|
||||
})
|
||||
|
||||
it('reads an older runtime’s method_not_found as the same refusal', async () => {
|
||||
const call = vi.fn(async (method: string) => {
|
||||
if (method === 'status.get') {
|
||||
return envelope({ capabilities: [MANAGED_SERVER_RUNTIME_CAPABILITY] })
|
||||
}
|
||||
throw new RuntimeClientError('method_not_found', 'Unknown method')
|
||||
})
|
||||
await expect(run('environment status', call, [])).rejects.toMatchObject({
|
||||
code: 'incompatible_runtime'
|
||||
})
|
||||
})
|
||||
|
||||
it('needs --yes before stopping, and then calls the same stop the settings use', async () => {
|
||||
const stopped = {
|
||||
outcome: 'unlinked',
|
||||
verdict: 'exited',
|
||||
environmentId: 'env-1',
|
||||
sshTargetId: 'ssh-1',
|
||||
stoppedVersion: '1.0.0',
|
||||
retirement: null
|
||||
}
|
||||
const unconfirmed = client(stopped)
|
||||
await expect(run('environment stop', unconfirmed, [])).rejects.toMatchObject({
|
||||
code: 'confirmation_required'
|
||||
})
|
||||
expect(unconfirmed).not.toHaveBeenCalled()
|
||||
|
||||
const log = vi.spyOn(console, 'log').mockImplementation(() => undefined)
|
||||
const confirmed = client(stopped)
|
||||
await run('environment stop', confirmed, [['yes', true]])
|
||||
expect(confirmed).toHaveBeenCalledWith(
|
||||
'managedServer.stop',
|
||||
{ selector: 'build-box' },
|
||||
{ timeoutMs: MANAGED_SERVER_ACTION_TIMEOUT_MS }
|
||||
)
|
||||
expect(log).toHaveBeenCalledWith('Stopped build-box and unlinked it from this machine.')
|
||||
})
|
||||
|
||||
it('fails the command with the refusal, so scripts see a non-zero exit', async () => {
|
||||
const refusal = {
|
||||
outcome: 'refused',
|
||||
verdict: 'live',
|
||||
code: 'orcad_stop_active_environment',
|
||||
reason: 'Choose another Active Server in Advanced before stopping this server.'
|
||||
}
|
||||
await expect(run('environment stop', client(refusal), [['yes', true]])).rejects.toMatchObject({
|
||||
code: 'managed_server_refused',
|
||||
message: refusal.reason,
|
||||
data: refusal
|
||||
})
|
||||
})
|
||||
|
||||
it('passes --force to an update, and reports a deferred one as unsettled', async () => {
|
||||
const call = client({
|
||||
outcome: 'deferred',
|
||||
candidateVersion: '2.0.0',
|
||||
code: 'orcad_update_terminals_running',
|
||||
reason: 'Terminals are running.'
|
||||
})
|
||||
await expect(run('environment update', call, [])).rejects.toMatchObject({
|
||||
code: 'managed_server_deferred'
|
||||
})
|
||||
expect(call).toHaveBeenCalledWith(
|
||||
'managedServer.update',
|
||||
{ selector: 'build-box', force: false },
|
||||
{ timeoutMs: MANAGED_SERVER_ACTION_TIMEOUT_MS }
|
||||
)
|
||||
|
||||
vi.spyOn(console, 'log').mockImplementation(() => undefined)
|
||||
const forced = client({
|
||||
outcome: 'updated',
|
||||
environment: { name: 'build-box' },
|
||||
activeVersion: '2.0.0'
|
||||
})
|
||||
await run('environment update', forced, [['force', true]])
|
||||
expect(forced).toHaveBeenCalledWith(
|
||||
'managedServer.update',
|
||||
{ selector: 'build-box', force: true },
|
||||
{ timeoutMs: MANAGED_SERVER_ACTION_TIMEOUT_MS }
|
||||
)
|
||||
})
|
||||
|
||||
it('fails on an outcome a newer desktop added instead of printing undefined', async () => {
|
||||
const log = vi.spyOn(console, 'log').mockImplementation(() => undefined)
|
||||
await expect(
|
||||
run('environment cancel-stop', client({ outcome: 'future-outcome' }), [])
|
||||
).rejects.toMatchObject({ code: 'managed_server_future-outcome' })
|
||||
expect(log).not.toHaveBeenCalled()
|
||||
})
|
||||
|
||||
it('prints a readable status', async () => {
|
||||
const log = vi.spyOn(console, 'log').mockImplementation(() => undefined)
|
||||
await run(
|
||||
'environment status',
|
||||
client({
|
||||
activeVersion: '1.0.0',
|
||||
previousVersion: '0.9.0',
|
||||
rollbackAvailable: true,
|
||||
recovery: null,
|
||||
terminals: { liveSessions: null },
|
||||
migration: null,
|
||||
deferredUpdate: null
|
||||
}),
|
||||
[]
|
||||
)
|
||||
expect(log.mock.calls[0]?.[0]).toBe(
|
||||
'Active version: 1.0.0\nPrevious version: 0.9.0 (rollback available)\nLive terminals: unverifiable'
|
||||
)
|
||||
})
|
||||
|
||||
it('gives mutating actions a budget past the desktop’s own deadlines, and status the default', async () => {
|
||||
const stopped = {
|
||||
outcome: 'unlinked',
|
||||
verdict: 'exited',
|
||||
environmentId: 'env-1',
|
||||
sshTargetId: 'ssh-1',
|
||||
stoppedVersion: '1.0.0',
|
||||
retirement: null
|
||||
}
|
||||
vi.spyOn(console, 'log').mockImplementation(() => undefined)
|
||||
const stop = client(stopped)
|
||||
await run('environment stop', stop, [['yes', true]])
|
||||
expect(stop).toHaveBeenCalledWith(
|
||||
'managedServer.stop',
|
||||
{ selector: 'build-box' },
|
||||
{ timeoutMs: MANAGED_SERVER_ACTION_TIMEOUT_MS }
|
||||
)
|
||||
const status = client({
|
||||
activeVersion: '1.0.0',
|
||||
previousVersion: null,
|
||||
rollbackAvailable: false,
|
||||
recovery: null,
|
||||
terminals: { liveSessions: 0 },
|
||||
migration: null,
|
||||
deferredUpdate: null
|
||||
})
|
||||
await run('environment status', status, [])
|
||||
expect(status).toHaveBeenCalledWith(
|
||||
'managedServer.status',
|
||||
{ selector: 'build-box' },
|
||||
undefined
|
||||
)
|
||||
})
|
||||
|
||||
it('reports a timed-out action as possibly still running, never as a failure of the action', async () => {
|
||||
const call = vi.fn(async (method: string) => {
|
||||
if (method === 'status.get') {
|
||||
return envelope({ capabilities: [MANAGED_SERVER_RUNTIME_CAPABILITY] })
|
||||
}
|
||||
throw new RuntimeClientError(
|
||||
'runtime_timeout',
|
||||
'Timed out waiting for the Orca runtime to respond.'
|
||||
)
|
||||
})
|
||||
await expect(run('environment update', call, [])).rejects.toMatchObject({
|
||||
code: 'managed_server_in_progress',
|
||||
message: expect.stringContaining('orca environment status')
|
||||
})
|
||||
})
|
||||
|
||||
// BUG-21: an update restarts the server, so the call's connection closes before it answers.
|
||||
it('reports what a restarting update left behind instead of failing on the closed connection', async () => {
|
||||
vi.useFakeTimers()
|
||||
const log = vi.spyOn(console, 'log').mockImplementation(() => {})
|
||||
let updated = false
|
||||
const call = vi.fn(async (method: string) => {
|
||||
if (method === 'status.get') {
|
||||
return envelope({ capabilities: [MANAGED_SERVER_RUNTIME_CAPABILITY] })
|
||||
}
|
||||
if (method === 'managedServer.update') {
|
||||
updated = true
|
||||
throw new RuntimeClientError(
|
||||
'runtime_unavailable',
|
||||
'The Orca runtime closed the connection before responding.'
|
||||
)
|
||||
}
|
||||
return envelope({ activeVersion: '0.2.0+bb01', recovery: null, deferredUpdate: null })
|
||||
})
|
||||
try {
|
||||
const done = run('environment update', call, [])
|
||||
await vi.advanceTimersByTimeAsync(2_000)
|
||||
await done
|
||||
expect(updated).toBe(true)
|
||||
expect(log).toHaveBeenCalledWith(expect.stringContaining('now runs 0.2.0+bb01'))
|
||||
} finally {
|
||||
vi.useRealTimers()
|
||||
}
|
||||
})
|
||||
|
||||
it('reports an update the restart left interrupted as needing recover', async () => {
|
||||
vi.useFakeTimers()
|
||||
const call = vi.fn(async (method: string) => {
|
||||
if (method === 'status.get') {
|
||||
return envelope({ capabilities: [MANAGED_SERVER_RUNTIME_CAPABILITY] })
|
||||
}
|
||||
if (method === 'managedServer.update') {
|
||||
throw new RuntimeClientError('runtime_unavailable', 'closed')
|
||||
}
|
||||
return envelope({
|
||||
activeVersion: '0.1.0+aa01',
|
||||
recovery: { operation: 'activate', version: '0.2.0+bb01', phase: 'snapshot-captured' },
|
||||
deferredUpdate: null
|
||||
})
|
||||
})
|
||||
try {
|
||||
const done = run('environment update', call, []).catch((error: unknown) => error)
|
||||
await vi.advanceTimersByTimeAsync(2_000)
|
||||
expect(await done).toMatchObject({ code: 'managed_server_interrupted' })
|
||||
} finally {
|
||||
vi.useRealTimers()
|
||||
}
|
||||
})
|
||||
|
||||
it('restores changed state only with --accept-changed-state --yes, and its refusal names the flags', async () => {
|
||||
const refusal = {
|
||||
outcome: 'refused',
|
||||
verdict: 'unverifiable',
|
||||
code: ORCAD_RECOVERY_CHANGED_STATE_CODE,
|
||||
reason: 'The launched build changed profile state. Recover to restore the prelaunch snapshot.'
|
||||
}
|
||||
await expect(run('environment recover', client(refusal), [])).rejects.toMatchObject({
|
||||
code: 'managed_server_refused',
|
||||
message: expect.stringContaining('--accept-changed-state --yes')
|
||||
})
|
||||
|
||||
const unconfirmed = client({ outcome: 'none' })
|
||||
await expect(
|
||||
run('environment recover', unconfirmed, [['accept-changed-state', true]])
|
||||
).rejects.toMatchObject({ code: 'confirmation_required' })
|
||||
expect(unconfirmed).not.toHaveBeenCalled()
|
||||
|
||||
vi.spyOn(console, 'log').mockImplementation(() => undefined)
|
||||
const confirmed = client({ outcome: 'none' })
|
||||
await run('environment recover', confirmed, [
|
||||
['accept-changed-state', true],
|
||||
['yes', true]
|
||||
])
|
||||
expect(confirmed).toHaveBeenCalledWith(
|
||||
'managedServer.recover',
|
||||
{ selector: 'build-box', acceptChangedState: true },
|
||||
{ timeoutMs: MANAGED_SERVER_ACTION_TIMEOUT_MS }
|
||||
)
|
||||
})
|
||||
})
|
||||
@@ -0,0 +1,227 @@
|
||||
/**
|
||||
* `orca environment status|update|rollback|recover|stop|cancel-stop`: the Managed servers
|
||||
* actions over runtime RPC. Each call is gated on the runtime's managedServer.v1 capability, and
|
||||
* an older runtime's method_not_found reads the same as a missing capability.
|
||||
*/
|
||||
import type {
|
||||
OrcadManagedCancelStopResult,
|
||||
OrcadManagedDeployResult,
|
||||
OrcadManagedRecoveryResult,
|
||||
OrcadManagedRollbackResult,
|
||||
OrcadManagedRuntimeStatus,
|
||||
OrcadManagedStopResult
|
||||
} from '../../shared/orcad-managed-runtime'
|
||||
import { ORCAD_RECOVERY_CHANGED_STATE_CODE } from '../../shared/orcad-managed-runtime'
|
||||
import { MANAGED_SERVER_RUNTIME_CAPABILITY } from '../../shared/protocol-version'
|
||||
import type { RuntimeStatus } from '../../shared/runtime-types'
|
||||
import type { CommandHandler, HandlerContext } from '../dispatch'
|
||||
import { getRequiredStringFlag } from '../flags'
|
||||
import { printResult } from '../format'
|
||||
import { RuntimeClientError, type RuntimeRpcSuccess } from '../runtime-client'
|
||||
import { formatManagedServerStatus } from './managed-server-format'
|
||||
import { reportAfterClosedConnection } from './managed-server-reconnect'
|
||||
|
||||
const UNSUPPORTED_MESSAGE =
|
||||
'This Orca runtime cannot manage servers over SSH. Run this on the computer whose Orca desktop app deployed the server, after updating Orca there.'
|
||||
|
||||
// Why 20 minutes: the desktop runs the whole action inline, and a Windows runtime promotion (5 min)
|
||||
// plus a readiness wait (5 min) on a slow host already outlast the 60 s RPC default.
|
||||
export const MANAGED_SERVER_ACTION_TIMEOUT_MS = 20 * 60_000
|
||||
const READ_ONLY_METHODS = new Set(['managedServer.status'])
|
||||
|
||||
function unsupported(): RuntimeClientError {
|
||||
return new RuntimeClientError('incompatible_runtime', UNSUPPORTED_MESSAGE)
|
||||
}
|
||||
|
||||
async function callManagedServer<TResult>(
|
||||
{ client }: HandlerContext,
|
||||
method: string,
|
||||
params: Record<string, unknown>
|
||||
): Promise<RuntimeRpcSuccess<TResult>> {
|
||||
const status = await client.call<RuntimeStatus>('status.get')
|
||||
if (!status.result.capabilities?.includes(MANAGED_SERVER_RUNTIME_CAPABILITY)) {
|
||||
throw unsupported()
|
||||
}
|
||||
const readOnly = READ_ONLY_METHODS.has(method)
|
||||
try {
|
||||
return await client.call<TResult>(
|
||||
method,
|
||||
params,
|
||||
readOnly ? undefined : { timeoutMs: MANAGED_SERVER_ACTION_TIMEOUT_MS }
|
||||
)
|
||||
} catch (error) {
|
||||
if (error instanceof RuntimeClientError && error.code === 'method_not_found') {
|
||||
throw unsupported()
|
||||
}
|
||||
// Why not a failure: the desktop keeps running the action after this client stops waiting.
|
||||
if (!readOnly && error instanceof RuntimeClientError && error.code === 'runtime_timeout') {
|
||||
throw new RuntimeClientError(
|
||||
'managed_server_in_progress',
|
||||
'Stopped waiting, but the desktop may still be running this action. Check `orca environment status` before retrying.'
|
||||
)
|
||||
}
|
||||
throw error
|
||||
}
|
||||
}
|
||||
|
||||
/** Null once the restart's outcome was reported from status instead. */
|
||||
async function callExpectingRestart<TResult>(
|
||||
context: HandlerContext,
|
||||
method: string,
|
||||
params: { selector: string } & Record<string, unknown>
|
||||
): Promise<RuntimeRpcSuccess<TResult> | null> {
|
||||
try {
|
||||
return await callManagedServer<TResult>(context, method, params)
|
||||
} catch (error) {
|
||||
if (!(error instanceof RuntimeClientError) || error.code !== 'runtime_unavailable') {
|
||||
throw error
|
||||
}
|
||||
await reportAfterClosedConnection(context, { selector: params.selector })
|
||||
return null
|
||||
}
|
||||
}
|
||||
|
||||
function selectorOf(context: HandlerContext): { selector: string } {
|
||||
return { selector: getRequiredStringFlag(context.flags, 'environment') }
|
||||
}
|
||||
|
||||
type Outcome = { outcome: string; code?: string; reason?: string }
|
||||
|
||||
/**
|
||||
* Prints a result whose outcome is in `settled`; throws any other so scripts see a non-zero exit.
|
||||
* Why an allow-list: an outcome a newer desktop adds must fail loudly, not print `undefined`.
|
||||
*/
|
||||
function report<TResult extends Outcome, TSettled extends TResult['outcome']>(
|
||||
response: RuntimeRpcSuccess<TResult>,
|
||||
json: boolean,
|
||||
settled: readonly TSettled[],
|
||||
done: (result: Extract<TResult, { outcome: TSettled }>) => string,
|
||||
nextStep?: (result: TResult) => string | null
|
||||
): void {
|
||||
const result = response.result
|
||||
if (!isSettled(result, settled)) {
|
||||
const reason = result.reason ?? `The managed Orca server action was ${result.outcome}.`
|
||||
const step = nextStep?.(result)
|
||||
throw new RuntimeClientError(
|
||||
`managed_server_${result.outcome}`,
|
||||
step ? `${reason} ${step}` : reason,
|
||||
result
|
||||
)
|
||||
}
|
||||
printResult({ ...response, result }, json, done)
|
||||
}
|
||||
|
||||
function isSettled<TResult extends Outcome, TSettled extends TResult['outcome']>(
|
||||
result: TResult,
|
||||
settled: readonly TSettled[]
|
||||
): result is Extract<TResult, { outcome: TSettled }> {
|
||||
return settled.some((outcome) => outcome === result.outcome)
|
||||
}
|
||||
|
||||
export const MANAGED_SERVER_HANDLERS: Record<string, CommandHandler> = {
|
||||
'environment status': async (context) => {
|
||||
const response = await callManagedServer<OrcadManagedRuntimeStatus>(
|
||||
context,
|
||||
'managedServer.status',
|
||||
selectorOf(context)
|
||||
)
|
||||
printResult(response, context.json, formatManagedServerStatus)
|
||||
},
|
||||
'environment update': async (context) => {
|
||||
const params = { ...selectorOf(context), force: context.flags.get('force') === true }
|
||||
const response = await callExpectingRestart<OrcadManagedDeployResult>(
|
||||
context,
|
||||
'managedServer.update',
|
||||
params
|
||||
)
|
||||
if (!response) {
|
||||
return
|
||||
}
|
||||
report(response, context.json, ['created', 'updated', 'already-current'], (result) =>
|
||||
result.outcome === 'already-current'
|
||||
? `Already on ${result.activeVersion}.`
|
||||
: `Updated ${result.environment.name} to ${result.activeVersion}.`
|
||||
)
|
||||
},
|
||||
'environment rollback': async (context) => {
|
||||
const response = await callExpectingRestart<OrcadManagedRollbackResult>(
|
||||
context,
|
||||
'managedServer.rollback',
|
||||
selectorOf(context)
|
||||
)
|
||||
if (!response) {
|
||||
return
|
||||
}
|
||||
report(
|
||||
response,
|
||||
context.json,
|
||||
['rolled-back'],
|
||||
(result) => `Rolled ${result.environment.name} back to ${result.activeVersion}.`
|
||||
)
|
||||
},
|
||||
'environment recover': async (context) => {
|
||||
const params = selectorOf(context)
|
||||
const acceptChangedState = context.flags.get('accept-changed-state') === true
|
||||
if (acceptChangedState && context.flags.get('yes') !== true) {
|
||||
throw new RuntimeClientError(
|
||||
'confirmation_required',
|
||||
`Restoring ${params.selector}'s prelaunch snapshot discards what the rejected build changed. Re-run with --yes to confirm.`
|
||||
)
|
||||
}
|
||||
const response = await callManagedServer<OrcadManagedRecoveryResult>(
|
||||
context,
|
||||
'managedServer.recover',
|
||||
acceptChangedState ? { ...params, acceptChangedState } : params
|
||||
)
|
||||
report(
|
||||
response,
|
||||
context.json,
|
||||
['recovered', 'none'],
|
||||
(result) =>
|
||||
result.outcome === 'recovered'
|
||||
? `Recovered ${result.environment.name} (${result.resolution}); active version ${result.activeVersion ?? 'none'}.`
|
||||
: 'Nothing to recover.',
|
||||
(result) =>
|
||||
!acceptChangedState && 'code' in result && result.code === ORCAD_RECOVERY_CHANGED_STATE_CODE
|
||||
? 'To restore it from the CLI, re-run with --accept-changed-state --yes.'
|
||||
: null
|
||||
)
|
||||
},
|
||||
'environment stop': async (context) => {
|
||||
const params = selectorOf(context)
|
||||
if (context.flags.get('yes') !== true) {
|
||||
throw new RuntimeClientError(
|
||||
'confirmation_required',
|
||||
`Stopping ${params.selector} ends its terminals and unlinks it from this machine. Re-run with --yes to confirm.`
|
||||
)
|
||||
}
|
||||
const response = await callManagedServer<OrcadManagedStopResult>(
|
||||
context,
|
||||
'managedServer.stop',
|
||||
params
|
||||
)
|
||||
report(
|
||||
response,
|
||||
context.json,
|
||||
['unlinked'],
|
||||
() => `Stopped ${params.selector} and unlinked it from this machine.`
|
||||
)
|
||||
},
|
||||
'environment cancel-stop': async (context) => {
|
||||
const response = await callManagedServer<OrcadManagedCancelStopResult>(
|
||||
context,
|
||||
'managedServer.cancelStop',
|
||||
selectorOf(context)
|
||||
)
|
||||
report(response, context.json, ['canceled', 'already-stopped', 'none'], (result) => {
|
||||
switch (result.outcome) {
|
||||
case 'canceled':
|
||||
return `Stop withdrawn; the server keeps serving ${result.activeVersion}.`
|
||||
case 'already-stopped':
|
||||
return 'orcad had already exited. Run `orca environment stop --yes` to unlink it.'
|
||||
case 'none':
|
||||
return 'No stop is pending.'
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -26,6 +26,7 @@ import {
|
||||
import { isTuiAgent } from '../../shared/tui-agent-config'
|
||||
import { isWorkspaceKey, worktreeWorkspaceKey } from '../../shared/workspace-scope'
|
||||
import { printLineageSummary } from './worktree-lineage-summary'
|
||||
import { projectWorktreePsTerminalVerdict } from '../worktree-ps-terminal-verdict'
|
||||
import {
|
||||
assertWorkspaceTargetFlagsCompatible,
|
||||
hasWorkspaceProjectTarget,
|
||||
@@ -143,7 +144,10 @@ export const WORKTREE_HANDLERS: Record<string, CommandHandler> = {
|
||||
{ limit: getOptionalPositiveIntegerFlag(flags, 'limit') }
|
||||
)
|
||||
await annotateOmittedHostScope(client, result.result)
|
||||
printResult(result, json, formatWorktreePs)
|
||||
const worktrees = result.result.worktrees.map(projectWorktreePsTerminalVerdict)
|
||||
printResult({ ...result, result: { ...result.result, worktrees } }, json, () =>
|
||||
formatWorktreePs(result.result)
|
||||
)
|
||||
},
|
||||
'worktree list': async ({ flags, client, json }) => {
|
||||
const result = await client.call<WithAnnotatedHostScope<RuntimeWorktreeListResult>>(
|
||||
|
||||
@@ -0,0 +1,93 @@
|
||||
import { EventEmitter } from 'node:events'
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
|
||||
|
||||
const { spawnMock, resolveLocalServeRuntimeMock, serveWithOrcadMock } = vi.hoisted(() => ({
|
||||
spawnMock: vi.fn(),
|
||||
resolveLocalServeRuntimeMock: vi.fn(),
|
||||
serveWithOrcadMock: vi.fn()
|
||||
}))
|
||||
|
||||
vi.mock('child_process', () => ({ spawn: spawnMock, spawnSync: vi.fn() }))
|
||||
vi.mock('./serve-orcad-launch', () => ({
|
||||
resolveLocalServeRuntime: resolveLocalServeRuntimeMock,
|
||||
serveWithOrcad: serveWithOrcadMock
|
||||
}))
|
||||
|
||||
import { serveOrcaApp } from './launch'
|
||||
|
||||
class FakeChildProcess extends EventEmitter {
|
||||
stdout = new EventEmitter()
|
||||
kill = vi.fn()
|
||||
unref = vi.fn()
|
||||
pid = 4101
|
||||
}
|
||||
|
||||
/** Electron serve that exits cleanly once spawned, whenever selection gets to spawning it. */
|
||||
function electronChild(): void {
|
||||
spawnMock.mockImplementation(() => {
|
||||
const child = new FakeChildProcess()
|
||||
setTimeout(() => child.emit('exit', 0, null), 0)
|
||||
return child
|
||||
})
|
||||
}
|
||||
|
||||
describe('orca serve host selection', () => {
|
||||
let stderr: string[]
|
||||
|
||||
beforeEach(() => {
|
||||
spawnMock.mockReset()
|
||||
resolveLocalServeRuntimeMock.mockReset()
|
||||
serveWithOrcadMock.mockReset()
|
||||
process.env.ORCA_APP_EXECUTABLE = '/opt/orca/orca-ide'
|
||||
stderr = []
|
||||
vi.spyOn(process.stderr, 'write').mockImplementation((chunk) => {
|
||||
stderr.push(String(chunk))
|
||||
return true
|
||||
})
|
||||
})
|
||||
|
||||
afterEach(() => {
|
||||
vi.restoreAllMocks()
|
||||
delete process.env.ORCA_APP_EXECUTABLE
|
||||
delete process.env.ORCA_SERVE_RUNTIME
|
||||
})
|
||||
|
||||
it('serves on orcad by default', async () => {
|
||||
const selection = { kind: 'orcad', runtime: '/node', entry: '/slot/orcad.js', version: '1' }
|
||||
resolveLocalServeRuntimeMock.mockResolvedValue(selection)
|
||||
serveWithOrcadMock.mockResolvedValue(0)
|
||||
|
||||
await expect(serveOrcaApp({ json: true })).resolves.toBe(0)
|
||||
expect(serveWithOrcadMock).toHaveBeenCalledWith(
|
||||
selection,
|
||||
{ json: true },
|
||||
expect.any(String),
|
||||
expect.not.objectContaining({ ELECTRON_RUN_AS_NODE: expect.anything() })
|
||||
)
|
||||
expect(spawnMock).not.toHaveBeenCalled()
|
||||
expect(stderr.join('')).toContain('[serve] running on orcad 1')
|
||||
})
|
||||
|
||||
it('falls back to Electron and prints why when orcad cannot serve', async () => {
|
||||
resolveLocalServeRuntimeMock.mockResolvedValue({ kind: 'electron', reason: 'no template' })
|
||||
electronChild()
|
||||
|
||||
await expect(serveOrcaApp({ json: true })).resolves.toBe(0)
|
||||
expect(spawnMock).toHaveBeenCalledWith(
|
||||
'/opt/orca/orca-ide',
|
||||
expect.arrayContaining(['--serve', '--serve-json']),
|
||||
expect.any(Object)
|
||||
)
|
||||
expect(stderr.join('')).toContain('[serve] using Electron serve: no template')
|
||||
})
|
||||
|
||||
it('keeps Electron without asking orcad when ORCA_SERVE_RUNTIME=electron', async () => {
|
||||
process.env.ORCA_SERVE_RUNTIME = 'electron'
|
||||
electronChild()
|
||||
|
||||
await expect(serveOrcaApp({ json: true })).resolves.toBe(0)
|
||||
expect(resolveLocalServeRuntimeMock).not.toHaveBeenCalled()
|
||||
expect(spawnMock).toHaveBeenCalledOnce()
|
||||
expect(stderr.join('')).not.toContain('[serve]')
|
||||
})
|
||||
})
|
||||
@@ -90,10 +90,13 @@ describe('serveOrcaApp', () => {
|
||||
spawnMock.mockReset()
|
||||
spawnSyncMock.mockReset()
|
||||
process.env.ORCA_APP_EXECUTABLE = '/Applications/Orca.app/Contents/MacOS/Orca'
|
||||
// These cover Electron serve itself; the orcad default is in launch-serve-runtime.test.ts.
|
||||
process.env.ORCA_SERVE_RUNTIME = 'electron'
|
||||
})
|
||||
|
||||
afterEach(() => {
|
||||
vi.restoreAllMocks()
|
||||
delete process.env.ORCA_SERVE_RUNTIME
|
||||
delete process.env.ORCA_APP_EXECUTABLE
|
||||
delete process.env.ORCA_APP_EXECUTABLE_NEEDS_APP_ROOT
|
||||
delete process.env.ORCA_USER_DATA_PATH
|
||||
|
||||
+45
-115
@@ -1,16 +1,11 @@
|
||||
import { spawn as spawnProcess, type SpawnOptions } from 'node:child_process'
|
||||
import { existsSync } from 'node:fs'
|
||||
import { dirname, join, resolve } from 'node:path'
|
||||
import { StringDecoder } from 'node:string_decoder'
|
||||
import { runProcessSync } from '../../shared/child-process/run-process'
|
||||
import {
|
||||
SERVE_UPDATE_HANDOFF_PATH_ENV,
|
||||
getServeUpdateHandoffPath
|
||||
} from '../../shared/serve-update-handoff'
|
||||
import {
|
||||
getEphemeralVmRecipeResultConnection,
|
||||
parseEphemeralVmRecipeResult
|
||||
} from '../../shared/ephemeral-vm-recipes'
|
||||
import { getDefaultUserDataPath } from './metadata'
|
||||
import { getMacAppBundlePath } from './mac-app-update-bundle'
|
||||
import {
|
||||
@@ -19,8 +14,14 @@ import {
|
||||
superviseForegroundServe
|
||||
} from './serve-update-supervisor'
|
||||
import { RuntimeClientError } from './types'
|
||||
import { SERVE_RUNTIME_ELECTRON, SERVE_RUNTIME_ENV } from '../../shared/orcad-local-serve-selection'
|
||||
import {
|
||||
resolveLocalServeRuntime,
|
||||
serveWithOrcad,
|
||||
type ServeOrcaAppArgs
|
||||
} from './serve-orcad-launch'
|
||||
import { waitForRecipeJson } from './serve-recipe-json'
|
||||
|
||||
const IGNORED_NON_RECIPE_STDOUT = '[serve] ignored non-recipe stdout'
|
||||
const USER_NAMESPACE_PROBE_TIMEOUT_MS = 2_000
|
||||
|
||||
export function launchOrcaApp(): void {
|
||||
@@ -77,18 +78,44 @@ function spawnDetached(command: string, args: string[], options: SpawnOptions):
|
||||
child.unref()
|
||||
}
|
||||
|
||||
export function serveOrcaApp(
|
||||
args: {
|
||||
json?: boolean
|
||||
port?: string | null
|
||||
pairingAddress?: string | null
|
||||
noPairing?: boolean
|
||||
mobilePairing?: boolean
|
||||
recipeJson?: boolean
|
||||
projectRoot?: string | null
|
||||
} = {}
|
||||
): Promise<number> {
|
||||
export function serveOrcaApp(args: ServeOrcaAppArgs = {}): Promise<number> {
|
||||
const executable = resolveForegroundOrcaExecutable()
|
||||
if (args.recipeJson && !args.projectRoot) {
|
||||
throw new RuntimeClientError('invalid_argument', 'Recipe JSON output requires --project-root.')
|
||||
}
|
||||
// Why synchronous on the opt-out: it must spawn Electron exactly as before, without asking.
|
||||
if (process.env[SERVE_RUNTIME_ENV] === SERVE_RUNTIME_ELECTRON) {
|
||||
return serveWithElectron(executable, args)
|
||||
}
|
||||
return serveWithSelectedRuntime(executable, args)
|
||||
}
|
||||
|
||||
async function serveWithSelectedRuntime(
|
||||
executable: string,
|
||||
args: ServeOrcaAppArgs
|
||||
): Promise<number> {
|
||||
const selection = await resolveLocalServeRuntime({
|
||||
executable,
|
||||
appRoot: resolveAppRoot(),
|
||||
userDataPath: getDefaultUserDataPath(),
|
||||
usesMacUpdateHandoff: args.recipeJson !== true && getMacAppBundlePath(executable) !== null
|
||||
})
|
||||
if (selection.kind === 'orcad') {
|
||||
process.stderr.write(`[serve] running on orcad ${selection.version}\n`)
|
||||
return serveWithOrcad(
|
||||
selection,
|
||||
args,
|
||||
getDefaultUserDataPath(),
|
||||
stripElectronRunAsNode(process.env)
|
||||
)
|
||||
}
|
||||
if (selection.reason) {
|
||||
process.stderr.write(`[serve] using Electron serve: ${selection.reason}\n`)
|
||||
}
|
||||
return serveWithElectron(executable, args)
|
||||
}
|
||||
|
||||
function serveWithElectron(executable: string, args: ServeOrcaAppArgs): Promise<number> {
|
||||
const childArgs = [...getExecutableAppArgs(executable)]
|
||||
childArgs.push('--serve')
|
||||
if (args.json) {
|
||||
@@ -106,13 +133,7 @@ export function serveOrcaApp(
|
||||
if (args.mobilePairing) {
|
||||
childArgs.push('--serve-mobile-pairing')
|
||||
}
|
||||
if (args.recipeJson) {
|
||||
if (!args.projectRoot) {
|
||||
throw new RuntimeClientError(
|
||||
'invalid_argument',
|
||||
'Recipe JSON output requires --project-root.'
|
||||
)
|
||||
}
|
||||
if (args.recipeJson && args.projectRoot) {
|
||||
childArgs.push('--serve-recipe-json', '--serve-project-root', args.projectRoot)
|
||||
}
|
||||
|
||||
@@ -164,97 +185,6 @@ export function serveOrcaApp(
|
||||
})
|
||||
}
|
||||
|
||||
function waitForRecipeJson(child: ReturnType<typeof spawnProcess>): Promise<number> {
|
||||
return new Promise((resolve, reject) => {
|
||||
let output = ''
|
||||
let settled = false
|
||||
const timeout = setTimeout(() => {
|
||||
finish(new RuntimeClientError('runtime_serve_failed', 'Timed out waiting for recipe JSON.'))
|
||||
child.kill('SIGTERM')
|
||||
}, 60000)
|
||||
const finish = (error?: Error): void => {
|
||||
if (settled) {
|
||||
return
|
||||
}
|
||||
settled = true
|
||||
clearTimeout(timeout)
|
||||
child.stdout?.off('data', onData)
|
||||
child.off('error', onError)
|
||||
child.off('close', onClose)
|
||||
if (error) {
|
||||
reject(error)
|
||||
return
|
||||
}
|
||||
child.stdout?.destroy?.()
|
||||
child.unref()
|
||||
resolve(0)
|
||||
}
|
||||
const writeIgnoredRecipeStdout = (): void => {
|
||||
// Why: non-readiness child stdout is untrusted and cannot be safely
|
||||
// redacted, including schema-valid results with arbitrary user data.
|
||||
process.stderr.write(`${IGNORED_NON_RECIPE_STDOUT}\n`)
|
||||
}
|
||||
const processRecipeOutputLine = (line: string): void => {
|
||||
const normalizedLine = line.endsWith('\r') ? line.slice(0, -1) : line
|
||||
if (!normalizedLine.trim()) {
|
||||
return
|
||||
}
|
||||
const parsed = parseEphemeralVmRecipeResult(normalizedLine)
|
||||
if (!parsed.ok) {
|
||||
writeIgnoredRecipeStdout()
|
||||
return
|
||||
}
|
||||
if (getEphemeralVmRecipeResultConnection(parsed.result).type !== 'orca-server') {
|
||||
writeIgnoredRecipeStdout()
|
||||
return
|
||||
}
|
||||
process.stdout.write(`${normalizedLine.trim()}\n`)
|
||||
finish()
|
||||
}
|
||||
const stdoutDecoder = new StringDecoder('utf8')
|
||||
const onData = (chunk: Buffer | string): void => {
|
||||
output += typeof chunk === 'string' ? chunk : stdoutDecoder.write(chunk)
|
||||
while (!settled) {
|
||||
const newlineIndex = output.indexOf('\n')
|
||||
if (newlineIndex === -1) {
|
||||
return
|
||||
}
|
||||
const line = output.slice(0, newlineIndex)
|
||||
output = output.slice(newlineIndex + 1)
|
||||
processRecipeOutputLine(line)
|
||||
}
|
||||
}
|
||||
const onError = (error: Error): void => {
|
||||
finish(error)
|
||||
}
|
||||
const onClose = (code: number | null, signal: NodeJS.Signals | null): void => {
|
||||
if (settled) {
|
||||
return
|
||||
}
|
||||
output += stdoutDecoder.end()
|
||||
if (output.trim()) {
|
||||
processRecipeOutputLine(output)
|
||||
}
|
||||
if (settled) {
|
||||
return
|
||||
}
|
||||
finish(
|
||||
new RuntimeClientError(
|
||||
'runtime_serve_failed',
|
||||
typeof code === 'number'
|
||||
? `Orca serve exited before printing valid recipe JSON with code ${code}.`
|
||||
: `Orca serve exited before printing valid recipe JSON via ${signal}.`
|
||||
)
|
||||
)
|
||||
}
|
||||
child.stdout?.on('data', onData)
|
||||
child.once('error', onError)
|
||||
// Why: `exit` can precede the final piped stdout data. `close` waits until
|
||||
// stdio closes so a last recipe chunk is not mistaken for missing output.
|
||||
child.once('close', onClose)
|
||||
})
|
||||
}
|
||||
|
||||
export function getExecutableAppArgs(executable: string): string[] {
|
||||
const args = process.env.ORCA_APP_EXECUTABLE_NEEDS_APP_ROOT === '1' ? [resolveAppRoot()] : []
|
||||
if (shouldDisableExtractedAppImageSandbox(executable)) {
|
||||
|
||||
@@ -0,0 +1,103 @@
|
||||
import { describe, expect, it, vi } from 'vitest'
|
||||
import { formatServeRuntimeSelection } from '../../shared/orcad-local-serve-selection'
|
||||
import type { runProcess } from '../../shared/child-process/run-process'
|
||||
import { orcadServeArgs, resolveLocalServeRuntime } from './serve-orcad-launch'
|
||||
|
||||
type RunProcess = typeof runProcess
|
||||
|
||||
function answering(stdout: string, code = 0): RunProcess {
|
||||
return vi.fn<RunProcess>(async () => ({
|
||||
code,
|
||||
signal: null,
|
||||
stdout,
|
||||
stderr: '',
|
||||
timedOut: false
|
||||
}))
|
||||
}
|
||||
|
||||
const options = {
|
||||
executable: '/Applications/Orca.app/Contents/MacOS/Orca',
|
||||
appRoot: '/Applications/Orca.app/Contents/Resources/app.asar',
|
||||
userDataPath: '/Users/u/Library/Application Support/orca',
|
||||
usesMacUpdateHandoff: false
|
||||
}
|
||||
|
||||
describe('orca serve asking the app which host to run', () => {
|
||||
it("runs the app's own selection entry as plain Node and reads its answer", async () => {
|
||||
const selection = {
|
||||
kind: 'orcad' as const,
|
||||
runtime: '/rt/node',
|
||||
entry: '/slot/orcad.js',
|
||||
version: '1'
|
||||
}
|
||||
const run = answering(`noise\n${formatServeRuntimeSelection(selection)}\n`)
|
||||
expect(await resolveLocalServeRuntime(options, run)).toEqual(selection)
|
||||
expect(run).toHaveBeenCalledWith(
|
||||
expect.objectContaining({
|
||||
program: options.executable,
|
||||
args: [
|
||||
`${options.appRoot}/out/main/orcad/orcad-local-serve-selection-entry.js`,
|
||||
'--user-data',
|
||||
options.userDataPath,
|
||||
'--app-root',
|
||||
options.appRoot
|
||||
],
|
||||
env: expect.objectContaining({ ELECTRON_RUN_AS_NODE: '1' }),
|
||||
stdio: ['ignore', 'pipe', 'inherit']
|
||||
})
|
||||
)
|
||||
})
|
||||
|
||||
it('keeps packaged macOS on Electron without starting the app to ask', async () => {
|
||||
const run = answering('')
|
||||
expect(await resolveLocalServeRuntime({ ...options, usesMacUpdateHandoff: true }, run)).toEqual(
|
||||
{
|
||||
kind: 'electron',
|
||||
reason: expect.stringContaining('packaged macOS')
|
||||
}
|
||||
)
|
||||
expect(run).not.toHaveBeenCalled()
|
||||
})
|
||||
|
||||
it('serves on Electron, and says why, when the app gives no answer', async () => {
|
||||
expect(await resolveLocalServeRuntime(options, answering('', 1))).toEqual({
|
||||
kind: 'electron',
|
||||
reason: expect.stringContaining('did not answer')
|
||||
})
|
||||
const failing = vi.fn<RunProcess>(async () => {
|
||||
throw new Error('spawn ENOENT')
|
||||
})
|
||||
expect(await resolveLocalServeRuntime(options, failing)).toEqual({
|
||||
kind: 'electron',
|
||||
reason: expect.stringContaining('spawn ENOENT')
|
||||
})
|
||||
})
|
||||
|
||||
it('forwards every desktop serve flag, binding wide as Electron serve does', () => {
|
||||
expect(
|
||||
orcadServeArgs({
|
||||
json: true,
|
||||
port: '6768',
|
||||
pairingAddress: '10.0.0.5',
|
||||
noPairing: true,
|
||||
mobilePairing: true,
|
||||
recipeJson: true,
|
||||
projectRoot: '/work/app'
|
||||
})
|
||||
).toEqual([
|
||||
'--bind',
|
||||
'0.0.0.0',
|
||||
'--json',
|
||||
'--port',
|
||||
'6768',
|
||||
'--pairing-address',
|
||||
'10.0.0.5',
|
||||
'--no-pairing',
|
||||
'--mobile-pairing',
|
||||
'--recipe-json',
|
||||
'--project-root',
|
||||
'/work/app'
|
||||
])
|
||||
expect(orcadServeArgs({})).toEqual(['--bind', '0.0.0.0'])
|
||||
})
|
||||
})
|
||||
@@ -0,0 +1,135 @@
|
||||
/**
|
||||
* `orca serve` on this machine's orcad slot. Whether to (and the slot itself) is decided
|
||||
* app-side by `src/main/orcad/orcad-local-serve-selection.ts`; the CLI only asks and runs.
|
||||
*/
|
||||
import { dirname, join } from 'node:path'
|
||||
import { runProcess, spawnProcess } from '../../shared/child-process/run-process'
|
||||
import {
|
||||
ORCAD_LOCAL_SERVE_SELECTION_ENTRY,
|
||||
ORCAD_LOCAL_SERVE_SELECTION_FLAGS as FLAGS,
|
||||
parseServeRuntimeSelection,
|
||||
type ServeRuntimeSelection
|
||||
} from '../../shared/orcad-local-serve-selection'
|
||||
import { waitForRecipeJson } from './serve-recipe-json'
|
||||
import { superviseForegroundServe } from './serve-update-supervisor'
|
||||
|
||||
type SupervisorArgs = Parameters<typeof superviseForegroundServe>[0]
|
||||
|
||||
export type ServeOrcaAppArgs = {
|
||||
json?: boolean
|
||||
port?: string | null
|
||||
pairingAddress?: string | null
|
||||
noPairing?: boolean
|
||||
mobilePairing?: boolean
|
||||
recipeJson?: boolean
|
||||
projectRoot?: string | null
|
||||
}
|
||||
|
||||
/** A first run may download and verify the pinned Node; bound it well past that. */
|
||||
const SELECTION_TIMEOUT_MS = 10 * 60_000
|
||||
|
||||
/** Asks the app's own entry, run on the app's executable as plain Node, which host to serve on. */
|
||||
export async function resolveLocalServeRuntime(
|
||||
options: {
|
||||
executable: string
|
||||
appRoot: string
|
||||
userDataPath: string
|
||||
usesMacUpdateHandoff: boolean
|
||||
},
|
||||
run: typeof runProcess = runProcess
|
||||
): Promise<ServeRuntimeSelection> {
|
||||
// Why: only packaged macOS serve can take a remote app update, through Electron's updater and
|
||||
// this CLI's supervisor; orcad has no updater, so switching would drop that.
|
||||
if (options.usesMacUpdateHandoff) {
|
||||
return {
|
||||
kind: 'electron',
|
||||
reason:
|
||||
'packaged macOS serve stays on Electron so paired clients can still update it (orcad has no app updater)'
|
||||
}
|
||||
}
|
||||
const entry = join(options.appRoot, 'out', 'main', `${ORCAD_LOCAL_SERVE_SELECTION_ENTRY}.js`)
|
||||
try {
|
||||
const result = await run({
|
||||
program: options.executable,
|
||||
args: [entry, FLAGS.userData, options.userDataPath, FLAGS.appRoot, options.appRoot],
|
||||
env: { ...process.env, ELECTRON_RUN_AS_NODE: '1' },
|
||||
// Why inherit stderr: a first run may download the pinned Node, and that progress is the
|
||||
// only sign `orca serve` is not hung.
|
||||
stdio: ['ignore', 'pipe', 'inherit'],
|
||||
timeoutMs: SELECTION_TIMEOUT_MS
|
||||
})
|
||||
return (
|
||||
parseServeRuntimeSelection(result.stdout) ?? {
|
||||
kind: 'electron',
|
||||
reason: `the app did not answer which serve host to use (exit ${String(result.code)})`
|
||||
}
|
||||
)
|
||||
} catch (error) {
|
||||
return {
|
||||
kind: 'electron',
|
||||
reason: `the app could not check orcad: ${error instanceof Error ? error.message : String(error)}`
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/** The shared spawn chokepoint (windowsHide, no shell) in the supervisor's spawn shape. */
|
||||
const spawnThroughChokepoint: SupervisorArgs['spawnChild'] = (program, args, options) =>
|
||||
spawnProcess({
|
||||
program,
|
||||
args,
|
||||
cwd: typeof options.cwd === 'string' ? options.cwd : undefined,
|
||||
env: options.env,
|
||||
stdio: options.stdio,
|
||||
detached: options.detached
|
||||
})
|
||||
|
||||
/** Electron serve binds every interface (`exposeNetworkByDefault`); orcad does it on request. */
|
||||
export function serveWithOrcad(
|
||||
selection: Extract<ServeRuntimeSelection, { kind: 'orcad' }>,
|
||||
args: ServeOrcaAppArgs,
|
||||
userDataPath: string,
|
||||
/** The caller's environment without `ELECTRON_RUN_AS_NODE`. */
|
||||
baseEnv: NodeJS.ProcessEnv,
|
||||
spawnChild: SupervisorArgs['spawnChild'] = spawnThroughChokepoint
|
||||
): Promise<number> {
|
||||
const childArgs = [selection.entry, ...orcadServeArgs(args)]
|
||||
const spawnOptions: SupervisorArgs['spawnOptions'] = {
|
||||
detached: args.recipeJson === true,
|
||||
cwd: dirname(selection.entry),
|
||||
stdio: args.recipeJson === true ? ['ignore', 'pipe', 'inherit'] : 'inherit',
|
||||
env: {
|
||||
...baseEnv,
|
||||
// The desktop's profile: its instance lock makes the two refuse each other.
|
||||
ORCA_USER_DATA: userDataPath,
|
||||
ORCA_VERSION: selection.version
|
||||
}
|
||||
}
|
||||
const child = spawnChild(selection.runtime, childArgs, spawnOptions)
|
||||
if (args.recipeJson) {
|
||||
return waitForRecipeJson(child)
|
||||
}
|
||||
return superviseForegroundServe({
|
||||
executable: selection.runtime,
|
||||
childArgs,
|
||||
spawnOptions,
|
||||
spawnChild,
|
||||
child,
|
||||
handoffPath: null,
|
||||
expectedHandoff: null
|
||||
})
|
||||
}
|
||||
|
||||
export function orcadServeArgs(args: ServeOrcaAppArgs): string[] {
|
||||
return [
|
||||
'--bind',
|
||||
'0.0.0.0',
|
||||
...(args.json ? ['--json'] : []),
|
||||
...(args.port ? ['--port', args.port] : []),
|
||||
...(args.pairingAddress ? ['--pairing-address', args.pairingAddress] : []),
|
||||
...(args.noPairing ? ['--no-pairing'] : []),
|
||||
...(args.mobilePairing ? ['--mobile-pairing'] : []),
|
||||
...(args.recipeJson && args.projectRoot
|
||||
? ['--recipe-json', '--project-root', args.projectRoot]
|
||||
: [])
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,101 @@
|
||||
import type { ChildProcessHandle } from '../../shared/child-process/process-spec'
|
||||
import { StringDecoder } from 'node:string_decoder'
|
||||
import {
|
||||
getEphemeralVmRecipeResultConnection,
|
||||
parseEphemeralVmRecipeResult
|
||||
} from '../../shared/ephemeral-vm-recipes'
|
||||
import { RuntimeClientError } from './types'
|
||||
|
||||
const IGNORED_NON_RECIPE_STDOUT = '[serve] ignored non-recipe stdout'
|
||||
|
||||
/** Relays the one recipe line a detached serve child prints, then lets the CLI exit. */
|
||||
export function waitForRecipeJson(child: ChildProcessHandle): Promise<number> {
|
||||
return new Promise((resolve, reject) => {
|
||||
let output = ''
|
||||
let settled = false
|
||||
const timeout = setTimeout(() => {
|
||||
finish(new RuntimeClientError('runtime_serve_failed', 'Timed out waiting for recipe JSON.'))
|
||||
child.kill('SIGTERM')
|
||||
}, 60000)
|
||||
const finish = (error?: Error): void => {
|
||||
if (settled) {
|
||||
return
|
||||
}
|
||||
settled = true
|
||||
clearTimeout(timeout)
|
||||
child.stdout?.off('data', onData)
|
||||
child.off('error', onError)
|
||||
child.off('close', onClose)
|
||||
if (error) {
|
||||
reject(error)
|
||||
return
|
||||
}
|
||||
child.stdout?.destroy?.()
|
||||
child.unref()
|
||||
resolve(0)
|
||||
}
|
||||
const writeIgnoredRecipeStdout = (): void => {
|
||||
// Why: non-readiness child stdout is untrusted and cannot be safely
|
||||
// redacted, including schema-valid results with arbitrary user data.
|
||||
process.stderr.write(`${IGNORED_NON_RECIPE_STDOUT}\n`)
|
||||
}
|
||||
const processRecipeOutputLine = (line: string): void => {
|
||||
const normalizedLine = line.endsWith('\r') ? line.slice(0, -1) : line
|
||||
if (!normalizedLine.trim()) {
|
||||
return
|
||||
}
|
||||
const parsed = parseEphemeralVmRecipeResult(normalizedLine)
|
||||
if (!parsed.ok) {
|
||||
writeIgnoredRecipeStdout()
|
||||
return
|
||||
}
|
||||
if (getEphemeralVmRecipeResultConnection(parsed.result).type !== 'orca-server') {
|
||||
writeIgnoredRecipeStdout()
|
||||
return
|
||||
}
|
||||
process.stdout.write(`${normalizedLine.trim()}\n`)
|
||||
finish()
|
||||
}
|
||||
const stdoutDecoder = new StringDecoder('utf8')
|
||||
const onData = (chunk: Buffer | string): void => {
|
||||
output += typeof chunk === 'string' ? chunk : stdoutDecoder.write(chunk)
|
||||
while (!settled) {
|
||||
const newlineIndex = output.indexOf('\n')
|
||||
if (newlineIndex === -1) {
|
||||
return
|
||||
}
|
||||
const line = output.slice(0, newlineIndex)
|
||||
output = output.slice(newlineIndex + 1)
|
||||
processRecipeOutputLine(line)
|
||||
}
|
||||
}
|
||||
const onError = (error: Error): void => {
|
||||
finish(error)
|
||||
}
|
||||
const onClose = (code: number | null, signal: NodeJS.Signals | null): void => {
|
||||
if (settled) {
|
||||
return
|
||||
}
|
||||
output += stdoutDecoder.end()
|
||||
if (output.trim()) {
|
||||
processRecipeOutputLine(output)
|
||||
}
|
||||
if (settled) {
|
||||
return
|
||||
}
|
||||
finish(
|
||||
new RuntimeClientError(
|
||||
'runtime_serve_failed',
|
||||
typeof code === 'number'
|
||||
? `Orca serve exited before printing valid recipe JSON with code ${code}.`
|
||||
: `Orca serve exited before printing valid recipe JSON via ${signal}.`
|
||||
)
|
||||
)
|
||||
}
|
||||
child.stdout?.on('data', onData)
|
||||
child.once('error', onError)
|
||||
// Why: `exit` can precede the final piped stdout data. `close` waits until
|
||||
// stdio closes so a last recipe chunk is not mistaken for missing output.
|
||||
child.once('close', onClose)
|
||||
})
|
||||
}
|
||||
@@ -1,4 +1,4 @@
|
||||
import type { ChildProcess, SpawnOptions, spawn } from 'node:child_process'
|
||||
import type { ChildProcess, SpawnOptions } from 'node:child_process'
|
||||
import { readFileSync } from 'node:fs'
|
||||
import { readFile, rename, unlink, writeFile } from 'node:fs/promises'
|
||||
import {
|
||||
@@ -27,7 +27,7 @@ type ServeSupervisorArgs = {
|
||||
executable: string
|
||||
childArgs: string[]
|
||||
spawnOptions: SpawnOptions
|
||||
spawnChild: typeof spawn
|
||||
spawnChild: (program: string, args: string[], options: SpawnOptions) => ChildProcess
|
||||
handoffPath: string | null
|
||||
}
|
||||
|
||||
|
||||
@@ -9,6 +9,7 @@ import { PROJECT_COMMAND_SPECS } from './project'
|
||||
import { ORCHESTRATION_COMMAND_SPECS } from './orchestration'
|
||||
import { COMPUTER_COMMAND_SPECS } from './computer'
|
||||
import { ENVIRONMENT_COMMAND_SPECS } from './environment'
|
||||
import { MANAGED_SERVER_COMMAND_SPECS } from './managed-server'
|
||||
import { AGENT_HOOK_COMMAND_SPECS } from './agent-hooks'
|
||||
import { DIAGNOSTICS_COMMAND_SPECS } from './diagnostics'
|
||||
import { EMULATOR_COMMAND_SPECS } from './emulator'
|
||||
@@ -35,6 +36,7 @@ export const COMMAND_SPECS: CommandSpec[] = [
|
||||
...DIAGNOSTICS_COMMAND_SPECS,
|
||||
...INTROSPECTION_COMMAND_SPECS,
|
||||
...ENVIRONMENT_COMMAND_SPECS,
|
||||
...MANAGED_SERVER_COMMAND_SPECS,
|
||||
...LINEAR_COMMAND_SPECS,
|
||||
...VM_COMMAND_SPECS,
|
||||
...EMULATOR_COMMAND_SPECS,
|
||||
|
||||
@@ -0,0 +1,75 @@
|
||||
import type { CommandSpec } from '../args'
|
||||
import { GLOBAL_FLAGS } from '../args'
|
||||
|
||||
const SELECTOR_NOTE =
|
||||
'--environment names a managed Orca server this machine deployed over SSH (see `orca environment list`); it is a selector here, not a routing flag.'
|
||||
const DESKTOP_NOTE =
|
||||
'Runs on this machine’s Orca desktop app, the same action as Settings > Managed servers. A runtime without the SSH registry (headless `orca serve`) refuses it.'
|
||||
|
||||
export const MANAGED_SERVER_COMMAND_SPECS: CommandSpec[] = [
|
||||
{
|
||||
path: ['environment', 'status'],
|
||||
summary: 'Show a managed Orca server’s active version, terminals and pending work',
|
||||
usage: 'orca environment status --environment <selector> [--json]',
|
||||
allowedFlags: [...GLOBAL_FLAGS],
|
||||
notes: [SELECTOR_NOTE, DESKTOP_NOTE],
|
||||
examples: ['orca environment status --environment build-box']
|
||||
},
|
||||
{
|
||||
path: ['environment', 'update'],
|
||||
summary: 'Update a managed Orca server to this build’s version',
|
||||
usage: 'orca environment update --environment <selector> [--force] [--json]',
|
||||
allowedFlags: [...GLOBAL_FLAGS, 'force'],
|
||||
notes: [
|
||||
'Without --force, an update that would restart over running terminals is deferred and reported, not applied.',
|
||||
SELECTOR_NOTE,
|
||||
DESKTOP_NOTE
|
||||
],
|
||||
examples: ['orca environment update --environment build-box']
|
||||
},
|
||||
{
|
||||
path: ['environment', 'rollback'],
|
||||
summary: 'Roll a managed Orca server back to its previous version',
|
||||
usage: 'orca environment rollback --environment <selector> [--json]',
|
||||
allowedFlags: [...GLOBAL_FLAGS],
|
||||
notes: [SELECTOR_NOTE, DESKTOP_NOTE]
|
||||
},
|
||||
{
|
||||
path: ['environment', 'recover'],
|
||||
summary: 'Finish or undo a managed Orca server’s interrupted update, rollback or stop',
|
||||
usage:
|
||||
'orca environment recover --environment <selector> [--accept-changed-state --yes] [--json]',
|
||||
allowedFlags: [...GLOBAL_FLAGS, 'accept-changed-state', 'yes'],
|
||||
notes: [
|
||||
'When a rejected build changed profile state, recover refuses rather than restart the previous build over it. --accept-changed-state --yes restores the prelaunch snapshot instead, the same as Restore in Settings > Managed servers; what the rejected build changed is discarded.',
|
||||
SELECTOR_NOTE,
|
||||
DESKTOP_NOTE
|
||||
],
|
||||
examples: [
|
||||
'orca environment recover --environment build-box',
|
||||
'orca environment recover --environment build-box --accept-changed-state --yes'
|
||||
]
|
||||
},
|
||||
{
|
||||
path: ['environment', 'stop'],
|
||||
destructive: true,
|
||||
summary: 'Stop a managed Orca server and unlink it from this machine (decommission)',
|
||||
usage: 'orca environment stop --environment <selector> --yes [--json]',
|
||||
allowedFlags: [...GLOBAL_FLAGS, 'yes'],
|
||||
notes: [
|
||||
'Stops orcad on the SSH host and, once the host proves it exited, removes the server from this machine. Its terminals end. Requires --yes.',
|
||||
'If the host cannot prove orcad exited, the server stays linked and the refusal says why; nothing is removed on a guess.',
|
||||
'Pick another Active Server first if this one is active. `orca environment cancel-stop` withdraws a stop orcad has not acted on yet.',
|
||||
SELECTOR_NOTE,
|
||||
DESKTOP_NOTE
|
||||
],
|
||||
examples: ['orca environment stop --environment build-box --yes']
|
||||
},
|
||||
{
|
||||
path: ['environment', 'cancel-stop'],
|
||||
summary: 'Withdraw a managed Orca server stop that orcad has not acted on yet',
|
||||
usage: 'orca environment cancel-stop --environment <selector> [--json]',
|
||||
allowedFlags: [...GLOBAL_FLAGS],
|
||||
notes: [SELECTOR_NOTE, DESKTOP_NOTE]
|
||||
}
|
||||
]
|
||||
@@ -8,6 +8,7 @@ import type {
|
||||
} from '../shared/runtime-types'
|
||||
import type { MemorySnapshot, WorktreeMemory } from '../shared/process-stats-types'
|
||||
import { formatListingHostScope, type WithAnnotatedHostScope } from './omitted-host-scope-selectors'
|
||||
import { formatWorktreePsTerminalFields } from './worktree-ps-terminal-verdict'
|
||||
|
||||
export function formatMemorySnapshot(snapshot: MemorySnapshot): string {
|
||||
const topWorktrees = [...snapshot.worktrees].sort((a, b) => b.memory - a.memory).slice(0, 10)
|
||||
@@ -139,7 +140,7 @@ export function formatWorktreePs(result: WithAnnotatedHostScope<RuntimeWorktreeP
|
||||
const body = result.worktrees
|
||||
.map(
|
||||
(worktree) =>
|
||||
`${worktree.repo} ${worktree.branch} host=${worktree.hostId ?? 'unverifiable'} live:${worktree.liveTerminalCount} pty:${worktree.hasAttachedPty ? 'yes' : 'no'} unread:${worktree.unread ? 'yes' : 'no'}\n${worktree.path}${worktree.preview ? `\npreview: ${worktree.preview}` : ''}`
|
||||
`${worktree.repo} ${worktree.branch} host=${worktree.hostId ?? 'unverifiable'} ${formatWorktreePsTerminalFields(worktree)} unread:${worktree.unread ? 'yes' : 'no'}\n${worktree.path}${worktree.preview ? `\npreview: ${worktree.preview}` : ''}`
|
||||
)
|
||||
.join('\n\n')
|
||||
const bodyWithScope = `${body}\n\n${scope}`
|
||||
|
||||
@@ -0,0 +1,47 @@
|
||||
import { describe, expect, it } from 'vitest'
|
||||
import {
|
||||
formatWorktreePsTerminalFields,
|
||||
projectWorktreePsTerminalVerdict
|
||||
} from './worktree-ps-terminal-verdict'
|
||||
|
||||
describe('worktree ps terminal verdict', () => {
|
||||
it('prints counts for a reachable host', () => {
|
||||
const row = { liveTerminalCount: 2, hasAttachedPty: true, unverifiableTerminalCount: 0 }
|
||||
expect(formatWorktreePsTerminalFields(row)).toBe('live:2 pty:yes')
|
||||
expect(projectWorktreePsTerminalVerdict(row)).toEqual({ ...row, terminalVerdict: 'live' })
|
||||
})
|
||||
|
||||
it('prints unverifiable, never zero or no, for an unreachable host', () => {
|
||||
const row = { liveTerminalCount: 0, hasAttachedPty: false, unverifiableTerminalCount: 1 }
|
||||
expect(formatWorktreePsTerminalFields(row)).toBe('live:unverifiable pty:unverifiable')
|
||||
expect(projectWorktreePsTerminalVerdict(row), 'count fields keep their JSON types').toEqual({
|
||||
...row,
|
||||
terminalVerdict: 'unverifiable'
|
||||
})
|
||||
})
|
||||
|
||||
it('keeps verified terminals alongside unverifiable ones', () => {
|
||||
const row = { liveTerminalCount: 1, hasAttachedPty: true, unverifiableTerminalCount: 2 }
|
||||
expect(formatWorktreePsTerminalFields(row)).toBe('live:1+2 unverifiable pty:yes')
|
||||
expect(
|
||||
formatWorktreePsTerminalFields({ ...row, hasAttachedPty: false }),
|
||||
'an unverifiable terminal may hold the pty'
|
||||
).toBe('live:1+2 unverifiable pty:unverifiable')
|
||||
expect(projectWorktreePsTerminalVerdict(row)).toEqual({ ...row, terminalVerdict: 'live' })
|
||||
})
|
||||
|
||||
it('reports none, not exited, for a row with no terminals', () => {
|
||||
const row = { liveTerminalCount: 0, hasAttachedPty: false, unverifiableTerminalCount: 0 }
|
||||
expect(formatWorktreePsTerminalFields(row)).toBe('live:0 pty:no')
|
||||
expect(projectWorktreePsTerminalVerdict(row)).toEqual({ ...row, terminalVerdict: 'none' })
|
||||
})
|
||||
|
||||
it('reports unverifiable for an idle row from a host that predates the count', () => {
|
||||
const row = { liveTerminalCount: 0, hasAttachedPty: false }
|
||||
expect(formatWorktreePsTerminalFields(row)).toBe('live:unverifiable pty:unverifiable')
|
||||
expect(projectWorktreePsTerminalVerdict(row)).toEqual({
|
||||
...row,
|
||||
terminalVerdict: 'unverifiable'
|
||||
})
|
||||
})
|
||||
})
|
||||
@@ -0,0 +1,42 @@
|
||||
import type { RuntimeWorktreePsSummary } from '../shared/runtime-types'
|
||||
|
||||
type TerminalCounts = Pick<
|
||||
RuntimeWorktreePsSummary,
|
||||
'liveTerminalCount' | 'hasAttachedPty' | 'unverifiableTerminalCount'
|
||||
>
|
||||
|
||||
/** `none`: the host reports no terminal for the row. Never `exited`, which needs proof of an exit. */
|
||||
type TerminalVerdict = 'live' | 'unverifiable' | 'none'
|
||||
|
||||
function terminalVerdict(row: TerminalCounts): TerminalVerdict {
|
||||
if (row.liveTerminalCount > 0) {
|
||||
return 'live'
|
||||
}
|
||||
// Absent count: a host that predates it, which cannot tell no terminals from lost contact.
|
||||
if (row.unverifiableTerminalCount === undefined || row.unverifiableTerminalCount > 0) {
|
||||
return 'unverifiable'
|
||||
}
|
||||
return 'none'
|
||||
}
|
||||
|
||||
/** `live:` and `pty:` words for one row; lost contact never reads as zero or no. */
|
||||
export function formatWorktreePsTerminalFields(row: TerminalCounts): string {
|
||||
if (terminalVerdict(row) === 'unverifiable') {
|
||||
return 'live:unverifiable pty:unverifiable'
|
||||
}
|
||||
const unverifiable = row.unverifiableTerminalCount ?? 0
|
||||
if (unverifiable === 0) {
|
||||
return `live:${row.liveTerminalCount} pty:${row.hasAttachedPty ? 'yes' : 'no'}`
|
||||
}
|
||||
return `live:${row.liveTerminalCount}+${unverifiable} unverifiable pty:${row.hasAttachedPty ? 'yes' : 'unverifiable'}`
|
||||
}
|
||||
|
||||
/**
|
||||
* JSON counterpart. The count fields keep their number/boolean types for existing scripts, so
|
||||
* `terminalVerdict` is what says a 0/false came from a host that could not be asked.
|
||||
*/
|
||||
export function projectWorktreePsTerminalVerdict<TRow extends TerminalCounts>(
|
||||
row: TRow
|
||||
): TRow & { terminalVerdict: TerminalVerdict } {
|
||||
return { ...row, terminalVerdict: terminalVerdict(row) }
|
||||
}
|
||||
@@ -0,0 +1,78 @@
|
||||
/**
|
||||
* What an idle pane says about a run. A ready shell prompt satisfies `tui-idle` too, so a prompt
|
||||
* typed at a shell whose agent is not installed ("command not found") would read as finished.
|
||||
* The agent's own status for the run's pane, reported since the run started, completes it as on
|
||||
* the desktop. Not every agent reports status (no hooks on the host, no recognised title), so for
|
||||
* those an idle pane still means done, but only after the agent had time to start and only when
|
||||
* no shell refused its command.
|
||||
*/
|
||||
import type { AutomationRun } from '../../shared/automations-types'
|
||||
|
||||
/** Covers agent spin-up over SSH before an idle pane without agent status is believed. */
|
||||
export const AGENT_START_GRACE_MS = 2 * 60 * 1000
|
||||
/** Only output after the prompt: a refusal further up predates this run. */
|
||||
const MISSING_COMMAND_TAIL_LINES = 6
|
||||
|
||||
// bash, zsh, dash/sh, fish, PowerShell and cmd.exe refusing a command that does not exist.
|
||||
const MISSING_COMMAND_PATTERNS = [
|
||||
/command not found/i,
|
||||
/^\S+: \d+: \S+: not found$/i,
|
||||
/unknown command/i,
|
||||
/is not recognized as (?:the name of a cmdlet|an internal or external command)/i
|
||||
]
|
||||
|
||||
export type AutomationRunAgentEvidence = {
|
||||
/** Agent status rows for a pane, from hooks, OSC and titles alike. */
|
||||
getAgentStatusRowsForPane(paneKey: string): readonly { receivedAt: number }[]
|
||||
/** The names the run's agent command may run under, to tell its refusal from its output. */
|
||||
agentCommandsForRun(run: AutomationRun): readonly string[]
|
||||
}
|
||||
|
||||
export type IdleRunVerdict =
|
||||
| { kind: 'completed' }
|
||||
| { kind: 'failed'; error: string }
|
||||
/** An idle shell that may not have started the agent yet. */
|
||||
| { kind: 'wait' }
|
||||
|
||||
/** A shell refusing one of the agent's own commands; an agent's tool output never matches. */
|
||||
export function findMissingCommandLine(
|
||||
tail: readonly string[],
|
||||
commands: readonly string[]
|
||||
): string | null {
|
||||
const names = commands.map((command) => command.split(/\s+/)[0]).filter(Boolean)
|
||||
const recent = tail
|
||||
.map((line) => line.trim())
|
||||
.filter(Boolean)
|
||||
.slice(-MISSING_COMMAND_TAIL_LINES)
|
||||
return (
|
||||
recent.find(
|
||||
(line) =>
|
||||
MISSING_COMMAND_PATTERNS.some((pattern) => pattern.test(line)) &&
|
||||
names.some((name) => line.includes(name))
|
||||
) ?? null
|
||||
)
|
||||
}
|
||||
|
||||
export function judgeIdleRun(
|
||||
evidence: AutomationRunAgentEvidence,
|
||||
run: AutomationRun,
|
||||
tail: readonly string[],
|
||||
runStartedAt: number,
|
||||
now: number
|
||||
): IdleRunVerdict {
|
||||
const paneKey = run.terminalPaneKey
|
||||
if (
|
||||
paneKey &&
|
||||
evidence.getAgentStatusRowsForPane(paneKey).some((row) => row.receivedAt >= runStartedAt)
|
||||
) {
|
||||
return { kind: 'completed' }
|
||||
}
|
||||
const missing = findMissingCommandLine(tail, evidence.agentCommandsForRun(run))
|
||||
if (missing) {
|
||||
return {
|
||||
kind: 'failed',
|
||||
error: `Automation agent did not start; this host could not run its command (${missing}).`
|
||||
}
|
||||
}
|
||||
return now - runStartedAt >= AGENT_START_GRACE_MS ? { kind: 'completed' } : { kind: 'wait' }
|
||||
}
|
||||
@@ -23,7 +23,7 @@ import type {
|
||||
ExternalAutomationTarget
|
||||
} from '../../shared/automations-types'
|
||||
import type { SshTarget } from '../../shared/ssh-types'
|
||||
import { isRuntimeOwnedSshTarget } from '../ssh/ssh-connection-store'
|
||||
import { isManagedOrcadSshTarget, isRuntimeOwnedSshTarget } from '../ssh/ssh-connection-store'
|
||||
|
||||
/** Current SSH registrations, hidden ones included so the guard can reject them itself. */
|
||||
export type DesktopSshTargetRegistry = {
|
||||
@@ -50,7 +50,7 @@ function resolveSshScope(
|
||||
throw externalAutomationTargetRemovedError()
|
||||
}
|
||||
// Why: checked before the generation compare so a hidden target reveals nothing about its registration.
|
||||
if (isRuntimeOwnedSshTarget(target)) {
|
||||
if (isRuntimeOwnedSshTarget(target) || isManagedOrcadSshTarget(target)) {
|
||||
throw new ExternalAutomationScopeError(EXTERNAL_AUTOMATION_SCOPE_CODES.targetHidden)
|
||||
}
|
||||
const current = sanitizeSshTargetGeneration(target.generation)
|
||||
|
||||
@@ -67,6 +67,8 @@ export async function runHeadlessAutomationDispatch(
|
||||
runId: run.id,
|
||||
status: 'dispatched',
|
||||
...launchRunTarget,
|
||||
// Kept here: a watched run's completion is written without it.
|
||||
...(precheckResult ? { precheckResult } : {}),
|
||||
error: null
|
||||
})
|
||||
// Observe the launched agent even while persistence is stalled or rejects its acknowledgement.
|
||||
|
||||
@@ -0,0 +1,228 @@
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
|
||||
import type { AutomationRun, AutomationRunStatus } from '../../shared/automations-types'
|
||||
import {
|
||||
createHeadlessRunTerminalRetention,
|
||||
RUN_TERMINAL_GRACE_MS,
|
||||
RUN_TERMINALS_KEPT_PER_AUTOMATION
|
||||
} from './headless-run-terminal-retention'
|
||||
|
||||
function makeRun(
|
||||
id: string,
|
||||
dispatchedAt: number,
|
||||
status: AutomationRunStatus = 'completed',
|
||||
automationId = 'nightly'
|
||||
): AutomationRun {
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: retention reads only these run fields.
|
||||
return {
|
||||
id,
|
||||
automationId,
|
||||
status,
|
||||
error: status === 'dispatch_failed' ? 'agent missing' : null,
|
||||
terminalPaneKey: `tab-${id}:1`,
|
||||
dispatchedAt,
|
||||
startedAt: dispatchedAt,
|
||||
createdAt: dispatchedAt
|
||||
} as AutomationRun
|
||||
}
|
||||
|
||||
function harness(runs: AutomationRun[]) {
|
||||
const closed: string[] = []
|
||||
const forgotten: AutomationRun[] = []
|
||||
const use = new Map<string, 'used' | 'unused' | 'unknown'>()
|
||||
// Runs whose shell is proven alone at its prompt; any other is unproven, as a live agent reads.
|
||||
const idleShell = new Set<string>()
|
||||
const dead = new Set<string>()
|
||||
const retention = createHeadlessRunTerminalRetention({
|
||||
listRuns: () => runs.filter((run) => !forgotten.some((gone) => gone.id === run.id)),
|
||||
terminalClientUse: (run) => use.get(run.id) ?? 'unused',
|
||||
runTerminalAlive: (run) => !dead.has(run.id),
|
||||
shellAloneAtPrompt: async (run) => idleShell.has(run.id),
|
||||
closeRunTerminal: async (run) => {
|
||||
closed.push(run.terminalPaneKey ?? '')
|
||||
return true
|
||||
},
|
||||
forgetRunTerminal: async (run) => {
|
||||
forgotten.push(run)
|
||||
}
|
||||
})
|
||||
return { retention, closed, forgotten, use, idleShell, dead }
|
||||
}
|
||||
|
||||
beforeEach(() => {
|
||||
vi.useFakeTimers()
|
||||
})
|
||||
|
||||
afterEach(() => {
|
||||
vi.useRealTimers()
|
||||
})
|
||||
|
||||
describe('headless run terminal retention', () => {
|
||||
it('closes finished run terminals past the grace period, keeping the newest few viewable', async () => {
|
||||
// Six hourly runs of one automation, all finished.
|
||||
const runs = [0, 1, 2, 3, 4, 5].map((hour) => makeRun(`r${hour}`, hour * 3_600_000))
|
||||
const h = harness(runs)
|
||||
|
||||
await h.retention.sweep()
|
||||
expect(h.closed).toEqual([])
|
||||
|
||||
vi.advanceTimersByTime(RUN_TERMINAL_GRACE_MS)
|
||||
await h.retention.sweep()
|
||||
expect(h.closed.toSorted()).toEqual(['tab-r0:1', 'tab-r1:1', 'tab-r2:1'])
|
||||
expect(6 - h.closed.length).toBe(RUN_TERMINALS_KEPT_PER_AUTOMATION)
|
||||
})
|
||||
|
||||
it('never closes a run that has not finished, however old', async () => {
|
||||
const runs = [
|
||||
makeRun('working', 0, 'dispatched'),
|
||||
makeRun('starting', 1, 'dispatching'),
|
||||
...[2, 3, 4, 5].map((n) => makeRun(`done${n}`, n))
|
||||
]
|
||||
const h = harness(runs)
|
||||
await h.retention.sweep()
|
||||
vi.advanceTimersByTime(RUN_TERMINAL_GRACE_MS * 10)
|
||||
await h.retention.sweep()
|
||||
expect(h.closed).toEqual(['tab-done2:1'])
|
||||
})
|
||||
|
||||
it.each([
|
||||
['a client typed into since dispatch', 'used'],
|
||||
['a client is attached to or viewing', 'used'],
|
||||
['this host cannot tell whether a client used', 'unknown']
|
||||
] as const)('never closes a terminal %s', async (_case, verdict) => {
|
||||
const runs = [0, 1, 2, 3, 4].map((n) => makeRun(`r${n}`, n))
|
||||
const h = harness(runs)
|
||||
h.use.set('r0', verdict)
|
||||
await h.retention.sweep()
|
||||
vi.advanceTimersByTime(RUN_TERMINAL_GRACE_MS * 10)
|
||||
await h.retention.sweep()
|
||||
expect(h.closed).toEqual(['tab-r1:1'])
|
||||
})
|
||||
|
||||
it.each(['dispatch_failed', 'skipped_precheck'] as const)(
|
||||
'keeps a %s run whose shell is not proven alone, since its agent may still be alive',
|
||||
async (status) => {
|
||||
const runs = [
|
||||
...[0, 1, 2, 3].map((n) => makeRun(`f${n}`, n, status)),
|
||||
...[4, 5, 6].map((n) => makeRun(`done${n}`, n))
|
||||
]
|
||||
const h = harness(runs)
|
||||
await h.retention.sweep()
|
||||
vi.advanceTimersByTime(RUN_TERMINAL_GRACE_MS * 10)
|
||||
await h.retention.sweep()
|
||||
expect(h.closed).toEqual([])
|
||||
}
|
||||
)
|
||||
|
||||
it('keeps the newest few per automation, not across all of them', async () => {
|
||||
const runs = [
|
||||
...[0, 1, 2].map((n) => makeRun(`a${n}`, n, 'completed', 'a')),
|
||||
...[0, 1, 2].map((n) => makeRun(`b${n}`, n, 'completed', 'b'))
|
||||
]
|
||||
const h = harness(runs)
|
||||
await h.retention.sweep()
|
||||
vi.advanceTimersByTime(RUN_TERMINAL_GRACE_MS)
|
||||
await h.retention.sweep()
|
||||
expect(h.closed).toEqual([])
|
||||
})
|
||||
|
||||
it('sweeps on its own once started, and stops with the service', async () => {
|
||||
const runs = [0, 1, 2, 3].map((n) => makeRun(`r${n}`, n))
|
||||
const h = harness(runs)
|
||||
h.retention.start()
|
||||
await vi.advanceTimersByTimeAsync(RUN_TERMINAL_GRACE_MS + 2 * 60_000)
|
||||
expect(h.closed).toEqual(['tab-r0:1'])
|
||||
|
||||
h.retention.stop()
|
||||
runs.push(makeRun('r4', 4), makeRun('r5', 5))
|
||||
await vi.advanceTimersByTimeAsync(RUN_TERMINAL_GRACE_MS * 2)
|
||||
expect(h.closed).toEqual(['tab-r0:1'])
|
||||
})
|
||||
|
||||
it('drains every completed, unused run terminal for an update, ignoring keep and grace', async () => {
|
||||
const runs = [
|
||||
...[0, 1, 2, 3, 4].map((n) => makeRun(`r${n}`, n)),
|
||||
makeRun('typed', 5),
|
||||
makeRun('adopted', 6),
|
||||
makeRun('failed', 7, 'dispatch_failed'),
|
||||
makeRun('working', 8, 'dispatched')
|
||||
]
|
||||
const h = harness(runs)
|
||||
h.use.set('typed', 'used')
|
||||
h.use.set('adopted', 'unknown')
|
||||
|
||||
// Fresh runs, inside the grace period: an update still releases them.
|
||||
expect(await h.retention.drain()).toBe(5)
|
||||
expect(h.closed.toSorted()).toEqual([
|
||||
'tab-r0:1',
|
||||
'tab-r1:1',
|
||||
'tab-r2:1',
|
||||
'tab-r3:1',
|
||||
'tab-r4:1'
|
||||
])
|
||||
})
|
||||
|
||||
it('closes a failed run whose agent command was not found, once its shell is idle', async () => {
|
||||
const runs = [
|
||||
makeRun('missing', 0, 'dispatch_failed'),
|
||||
...[1, 2, 3].map((n) => makeRun(`done${n}`, n))
|
||||
]
|
||||
const h = harness(runs)
|
||||
h.idleShell.add('missing')
|
||||
await h.retention.sweep()
|
||||
vi.advanceTimersByTime(RUN_TERMINAL_GRACE_MS)
|
||||
await h.retention.sweep()
|
||||
expect(h.closed).toEqual(['tab-missing:1'])
|
||||
expect(h.forgotten[0]).toMatchObject({ id: 'missing', status: 'dispatch_failed' })
|
||||
})
|
||||
|
||||
it('keeps a timed-out run whose agent still runs in its shell', async () => {
|
||||
const runs = [
|
||||
makeRun('timed-out', 0, 'dispatch_failed'),
|
||||
...[1, 2, 3].map((n) => makeRun(`done${n}`, n))
|
||||
]
|
||||
const h = harness(runs)
|
||||
await h.retention.sweep()
|
||||
vi.advanceTimersByTime(RUN_TERMINAL_GRACE_MS * 10)
|
||||
await h.retention.sweep()
|
||||
expect(h.closed).toEqual([])
|
||||
})
|
||||
|
||||
it('closes a still-dispatched run whose agent exited, leaving its shell idle', async () => {
|
||||
const runs = [
|
||||
makeRun('exited', 0, 'dispatched'),
|
||||
...[1, 2, 3].map((n) => makeRun(`done${n}`, n))
|
||||
]
|
||||
const h = harness(runs)
|
||||
h.idleShell.add('exited')
|
||||
await h.retention.sweep()
|
||||
vi.advanceTimersByTime(RUN_TERMINAL_GRACE_MS)
|
||||
await h.retention.sweep()
|
||||
expect(h.closed).toEqual(['tab-exited:1'])
|
||||
})
|
||||
|
||||
it('keeps the newest three live terminals, not counting ones already gone', async () => {
|
||||
const runs = [0, 1, 2, 3, 4].map((n) => makeRun(`r${n}`, n))
|
||||
const h = harness(runs)
|
||||
// The newest run's terminal was closed by hand: it must not take a keep slot.
|
||||
h.dead.add('r4')
|
||||
await h.retention.sweep()
|
||||
vi.advanceTimersByTime(RUN_TERMINAL_GRACE_MS)
|
||||
await h.retention.sweep()
|
||||
expect(h.closed).toEqual(['tab-r0:1'])
|
||||
expect(h.forgotten.map((run) => run.id).toSorted()).toEqual(['r0', 'r4'])
|
||||
})
|
||||
|
||||
it('drains idle failed and exited runs for an update, never a live agent', async () => {
|
||||
const runs = [
|
||||
makeRun('missing', 0, 'dispatch_failed'),
|
||||
makeRun('exited', 1, 'dispatched'),
|
||||
makeRun('timed-out', 2, 'dispatch_failed'),
|
||||
makeRun('working', 3, 'dispatched')
|
||||
]
|
||||
const h = harness(runs)
|
||||
h.idleShell.add('missing')
|
||||
h.idleShell.add('exited')
|
||||
expect(await h.retention.drain()).toBe(2)
|
||||
expect(h.closed.toSorted()).toEqual(['tab-exited:1', 'tab-missing:1'])
|
||||
})
|
||||
})
|
||||
@@ -0,0 +1,150 @@
|
||||
/**
|
||||
* Closing run terminals on a headless host. The desktop closes a run's terminal when the run
|
||||
* completes; on orcad nobody does, so schedules leave a shell and a PTY per run until the host
|
||||
* runs out. A run terminal stays open for a grace period and the newest few per automation stay
|
||||
* viewable; older ones are closed. A completed run's terminal is closed as is. A failed or
|
||||
* never-finishing run's is closed only once the shell is proven alone at its prompt, since a timed
|
||||
* out agent may still be alive. A terminal a client typed into or is viewing, or one whose use
|
||||
* this host cannot tell, is never closed: as on the desktop, that terminal is the user's.
|
||||
*/
|
||||
import {
|
||||
isFinalAutomationRunStatus,
|
||||
type AutomationRun,
|
||||
type AutomationRunStatus
|
||||
} from '../../shared/automations-types'
|
||||
|
||||
export const RUN_TERMINAL_GRACE_MS = 10 * 60_000
|
||||
export const RUN_TERMINALS_KEPT_PER_AUTOMATION = 3
|
||||
const SWEEP_INTERVAL_MS = 60_000
|
||||
|
||||
export type HeadlessRunTerminalRetentionDeps = {
|
||||
listRuns: () => readonly AutomationRun[]
|
||||
/**
|
||||
* Whether any client drove or is viewing the run's terminal. Like the desktop's take-over rule,
|
||||
* a used terminal is the user's now; `unknown` keeps it too.
|
||||
*/
|
||||
terminalClientUse: (run: AutomationRun) => 'used' | 'unused' | 'unknown'
|
||||
/** Whether the run's pane still holds the run's own PTY; a dead or replaced one is forgotten. */
|
||||
runTerminalAlive: (run: AutomationRun) => boolean
|
||||
/** Fresh proof that only the shell runs in the run's PTY, at its prompt; false when unproven. */
|
||||
shellAloneAtPrompt: (run: AutomationRun) => Promise<boolean>
|
||||
/**
|
||||
* Closes the run's own pane, leaving any pane a user split into that tab. False, closing
|
||||
* nothing, when the pane is gone or now holds another PTY (a restart put a new one there).
|
||||
*/
|
||||
closeRunTerminal: (run: AutomationRun) => Promise<boolean>
|
||||
/** Drops the closed terminal from the run, keeping its status, error and output. */
|
||||
forgetRunTerminal: (run: AutomationRun) => Promise<void>
|
||||
now?: () => number
|
||||
}
|
||||
|
||||
/** Still starting: its agent is being launched, so its terminal is never a candidate. */
|
||||
const STARTING: ReadonlySet<AutomationRunStatus> = new Set(['pending', 'dispatching'])
|
||||
|
||||
export function createHeadlessRunTerminalRetention(deps: HeadlessRunTerminalRetentionDeps): {
|
||||
sweep: () => Promise<void>
|
||||
/** Before an update restarts the server: no grace, no newest-N, every other rule still holds. */
|
||||
drain: () => Promise<number>
|
||||
start: () => void
|
||||
stop: () => void
|
||||
} {
|
||||
const now = deps.now ?? Date.now
|
||||
// When each run terminal was first seen as a candidate; the grace runs from there.
|
||||
const firstSeenAt = new Map<string, number>()
|
||||
let timer: ReturnType<typeof setInterval> | null = null
|
||||
let sweeping: Promise<void> | null = null
|
||||
|
||||
const forget = async (run: AutomationRun): Promise<void> => {
|
||||
await deps.forgetRunTerminal(run)
|
||||
firstSeenAt.delete(run.id)
|
||||
}
|
||||
|
||||
const mayClose = async (run: AutomationRun, graceMs: number): Promise<boolean> => {
|
||||
if (now() - (firstSeenAt.get(run.id) ?? now()) < graceMs) {
|
||||
return false
|
||||
}
|
||||
if (deps.terminalClientUse(run) !== 'unused') {
|
||||
return false
|
||||
}
|
||||
return run.status === 'completed' || (await deps.shellAloneAtPrompt(run))
|
||||
}
|
||||
|
||||
const sweepOnce = async (policy: { keep: number; graceMs: number }): Promise<number> => {
|
||||
let closedCount = 0
|
||||
const byAutomation = new Map<string, AutomationRun[]>()
|
||||
for (const run of deps.listRuns()) {
|
||||
if (!run.terminalPaneKey || STARTING.has(run.status)) {
|
||||
continue
|
||||
}
|
||||
if (!firstSeenAt.has(run.id)) {
|
||||
firstSeenAt.set(run.id, now())
|
||||
}
|
||||
byAutomation.set(run.automationId, [...(byAutomation.get(run.automationId) ?? []), run])
|
||||
}
|
||||
for (const runs of byAutomation.values()) {
|
||||
let kept = 0
|
||||
for (const run of runs.toSorted((a, b) => runRecency(b) - runRecency(a))) {
|
||||
try {
|
||||
// A terminal already gone holds no keep slot; only live ones stay viewable.
|
||||
if (!deps.runTerminalAlive(run)) {
|
||||
if (isFinalAutomationRunStatus(run.status)) {
|
||||
await forget(run)
|
||||
}
|
||||
continue
|
||||
}
|
||||
if (kept < policy.keep) {
|
||||
kept += 1
|
||||
continue
|
||||
}
|
||||
if (!(await mayClose(run, policy.graceMs))) {
|
||||
continue
|
||||
}
|
||||
if (await deps.closeRunTerminal(run)) {
|
||||
closedCount += 1
|
||||
}
|
||||
await forget(run)
|
||||
} catch (error) {
|
||||
console.error('[automations] could not close a run terminal:', error)
|
||||
}
|
||||
}
|
||||
}
|
||||
return closedCount
|
||||
}
|
||||
|
||||
const sweep = (): Promise<void> => {
|
||||
sweeping ??= sweepOnce({
|
||||
keep: RUN_TERMINALS_KEPT_PER_AUTOMATION,
|
||||
graceMs: RUN_TERMINAL_GRACE_MS
|
||||
})
|
||||
.then(() => {})
|
||||
.finally(() => {
|
||||
sweeping = null
|
||||
})
|
||||
return sweeping
|
||||
}
|
||||
|
||||
const drain = async (): Promise<number> => {
|
||||
// Waits out a periodic sweep so the two never close the same terminal twice.
|
||||
await sweeping?.catch(() => {})
|
||||
return sweepOnce({ keep: 0, graceMs: 0 })
|
||||
}
|
||||
|
||||
return {
|
||||
sweep,
|
||||
drain,
|
||||
start: () => {
|
||||
timer ??= setInterval(() => void sweep(), SWEEP_INTERVAL_MS)
|
||||
timer.unref?.()
|
||||
},
|
||||
stop: () => {
|
||||
if (timer) {
|
||||
clearInterval(timer)
|
||||
timer = null
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
function runRecency(run: AutomationRun): number {
|
||||
return run.dispatchedAt ?? run.startedAt ?? run.createdAt
|
||||
}
|
||||
@@ -12,7 +12,7 @@ const sshManagerState = vi.hoisted(() => ({
|
||||
}
|
||||
}))
|
||||
|
||||
vi.mock('../ipc/ssh', () => ({
|
||||
vi.mock('../ssh/ssh-target-registry', () => ({
|
||||
getSshConnectionManager: () => sshManagerState.manager
|
||||
}))
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@ import { spawn, type ChildProcess } from 'node:child_process'
|
||||
import type { ClientChannel } from 'ssh2'
|
||||
import type { AutomationPrecheck, AutomationPrecheckResult } from '../../shared/automations-types'
|
||||
import { MAX_AUTOMATION_PRECHECK_OUTPUT_CHARS } from '../../shared/automation-precheck'
|
||||
import { getSshConnectionManager } from '../ipc/ssh'
|
||||
import { getSshConnectionManager } from '../ssh/ssh-target-registry'
|
||||
import { shellEscape } from '../ssh/ssh-connection-utils'
|
||||
import { admitSelfInitiatedTreeKill } from '../own-chromium-tree-kill-guard'
|
||||
|
||||
|
||||
@@ -18,7 +18,7 @@ export type AutomationRunTerminalObserver = {
|
||||
resolveRunTerminal: (run: AutomationRun) => string | null
|
||||
observeCompletion: (
|
||||
handle: string,
|
||||
options: { signal: AbortSignal }
|
||||
options: { signal: AbortSignal; run?: AutomationRun }
|
||||
) => Promise<AutomationRunCompletionObservation>
|
||||
}
|
||||
|
||||
@@ -118,7 +118,10 @@ export class AutomationRunCompletionWatcher {
|
||||
): Promise<void> {
|
||||
let observation: AutomationRunCompletionObservation
|
||||
try {
|
||||
observation = await this.observer.observeCompletion(handle, { signal: controller.signal })
|
||||
observation = await this.observer.observeCompletion(handle, {
|
||||
signal: controller.signal,
|
||||
run
|
||||
})
|
||||
} catch (error) {
|
||||
if (controller.signal.aborted) {
|
||||
return
|
||||
|
||||
@@ -0,0 +1,125 @@
|
||||
import { afterEach, describe, expect, it, vi } from 'vitest'
|
||||
import type { AutomationRun } from '../../shared/automations-types'
|
||||
import { RUN_TERMINAL_GRACE_MS } from './headless-run-terminal-retention'
|
||||
|
||||
vi.mock('./service', () => ({
|
||||
AutomationService: class {
|
||||
start = vi.fn()
|
||||
stop = vi.fn()
|
||||
markDispatchResult = vi.fn(async () => ({}))
|
||||
}
|
||||
}))
|
||||
|
||||
import { createRuntimeAutomationService } from './runtime-automation-service'
|
||||
|
||||
function finishedRuns(count: number): AutomationRun[] {
|
||||
return Array.from(
|
||||
{ length: count },
|
||||
(_, n) =>
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: retention reads only these run fields.
|
||||
({
|
||||
id: `r${n}`,
|
||||
automationId: 'nightly',
|
||||
status: 'completed',
|
||||
error: null,
|
||||
terminalPaneKey: `tab-${n}:1`,
|
||||
terminalPtyId: `pty-${n}`,
|
||||
dispatchedAt: n,
|
||||
startedAt: n,
|
||||
createdAt: n
|
||||
}) as AutomationRun
|
||||
)
|
||||
}
|
||||
|
||||
function build(headless: boolean) {
|
||||
const runtime = {
|
||||
setAutomationService: vi.fn(),
|
||||
notifyAutomationsChanged: vi.fn(),
|
||||
getTerminalHandleForPaneKey: vi.fn((paneKey: string) => `handle:${paneKey}`),
|
||||
// Each run's pane still holds its own PTY unless a test restarts it.
|
||||
getTerminalPtyIdForHandle: vi.fn((handle: string) => `pty-${handle.split('tab-')[1]?.[0]}`),
|
||||
// pty-1's terminal has a client on it: typed into or being viewed.
|
||||
readTerminalClientUse: vi.fn((ptyId: string) => (ptyId === 'pty-1' ? 'used' : 'unused')),
|
||||
closeTerminal: vi.fn(async (_handle: string) => ({})),
|
||||
confirmTerminalShellAlone: vi.fn(async (_ptyId: string) => false),
|
||||
closeTerminalTab: vi.fn(async (_handle: string) => ({}))
|
||||
}
|
||||
const store = { listAutomationRuns: vi.fn(() => finishedRuns(5)), listAutomations: () => [] }
|
||||
const service = createRuntimeAutomationService({
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the mocked service and retention read only listAutomationRuns.
|
||||
store: store as never,
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: retention reads only the members faked above.
|
||||
runtime: runtime as never,
|
||||
headless
|
||||
})
|
||||
return { runtime, service, store }
|
||||
}
|
||||
|
||||
afterEach(() => {
|
||||
vi.useRealTimers()
|
||||
})
|
||||
|
||||
describe('headless automation service run terminal retention', () => {
|
||||
it('closes finished run terminals on orcad once the service runs, and stops with it', async () => {
|
||||
vi.useFakeTimers()
|
||||
const { runtime, service } = build(true)
|
||||
service.start()
|
||||
await vi.advanceTimersByTimeAsync(RUN_TERMINAL_GRACE_MS + 2 * 60_000)
|
||||
|
||||
// The oldest two finished runs are past the newest three; the one a client used stays open.
|
||||
// Only the run's own pane closes: a pane a user split into its tab survives.
|
||||
expect(runtime.closeTerminal.mock.calls.map(([handle]) => handle)).toEqual(['handle:tab-0:1'])
|
||||
expect(runtime.closeTerminalTab).not.toHaveBeenCalled()
|
||||
expect(service.markDispatchResult).toHaveBeenCalledWith(
|
||||
expect.objectContaining({ runId: 'r0', status: 'completed', terminalPaneKey: null })
|
||||
)
|
||||
|
||||
service.stop()
|
||||
runtime.closeTerminal.mockClear()
|
||||
await vi.advanceTimersByTimeAsync(RUN_TERMINAL_GRACE_MS * 2)
|
||||
expect(runtime.closeTerminal).not.toHaveBeenCalled()
|
||||
})
|
||||
|
||||
it('closes nothing in a pane a restart gave a new PTY, and still forgets the old one', async () => {
|
||||
vi.useFakeTimers()
|
||||
const { runtime, service } = build(true)
|
||||
// The oldest run's pane was restarted (exited pane, account switch): a user's PTY lives there.
|
||||
runtime.getTerminalPtyIdForHandle.mockImplementation((handle: string) =>
|
||||
handle === 'handle:tab-0:1' ? 'pty-user' : `pty-${handle.split('tab-')[1]?.[0]}`
|
||||
)
|
||||
service.start()
|
||||
await vi.advanceTimersByTimeAsync(RUN_TERMINAL_GRACE_MS + 2 * 60_000)
|
||||
|
||||
expect(runtime.closeTerminal).not.toHaveBeenCalled()
|
||||
expect(service.markDispatchResult).toHaveBeenCalledWith(
|
||||
expect.objectContaining({ runId: 'r0', terminalPaneKey: null, terminalPtyId: null })
|
||||
)
|
||||
service.stop()
|
||||
})
|
||||
|
||||
it('fails and closes a dispatched run whose agent exited, proven by its shell alone', async () => {
|
||||
vi.useFakeTimers()
|
||||
const { runtime, service, store } = build(true)
|
||||
store.listAutomationRuns.mockReturnValue([
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: retention reads only these run fields.
|
||||
{ ...finishedRuns(1)[0], status: 'dispatched' } as AutomationRun
|
||||
])
|
||||
runtime.confirmTerminalShellAlone.mockResolvedValue(true)
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the mocked service exposes the drain hook it was given.
|
||||
const drain = (service as unknown as { releaseFinishedRunTerminals: () => Promise<number> })
|
||||
.releaseFinishedRunTerminals
|
||||
await expect(drain()).resolves.toBe(1)
|
||||
expect(runtime.confirmTerminalShellAlone).toHaveBeenCalledWith('pty-0')
|
||||
expect(service.markDispatchResult).toHaveBeenCalledWith(
|
||||
expect.objectContaining({ runId: 'r0', status: 'dispatch_failed', terminalPtyId: null })
|
||||
)
|
||||
})
|
||||
|
||||
it('leaves run terminals to the renderer on the desktop', async () => {
|
||||
vi.useFakeTimers()
|
||||
const { runtime, service } = build(false)
|
||||
service.start()
|
||||
await vi.advanceTimersByTimeAsync(RUN_TERMINAL_GRACE_MS * 2)
|
||||
expect(runtime.closeTerminal).not.toHaveBeenCalled()
|
||||
})
|
||||
})
|
||||
@@ -0,0 +1,72 @@
|
||||
import { describe, expect, it, vi } from 'vitest'
|
||||
import type { HeadlessAutomationDispatcher } from './headless-dispatch'
|
||||
import type { AutomationRunTerminalObserver } from './run-completion-watcher'
|
||||
|
||||
const captured = vi.hoisted((): { dispatcher: unknown; observer: unknown } => ({
|
||||
dispatcher: null,
|
||||
observer: null
|
||||
}))
|
||||
|
||||
vi.mock('./service', () => ({
|
||||
AutomationService: class {
|
||||
start(): void {}
|
||||
stop(): void {}
|
||||
constructor(
|
||||
_store: unknown,
|
||||
opts: { headlessDispatcher?: unknown; terminalObserver?: unknown }
|
||||
) {
|
||||
captured.dispatcher = opts.headlessDispatcher
|
||||
captured.observer = opts.terminalObserver
|
||||
}
|
||||
}
|
||||
}))
|
||||
|
||||
import { createRuntimeAutomationService } from './runtime-automation-service'
|
||||
|
||||
describe('headless automation dispatch', () => {
|
||||
it('hands the launched run to its watcher instead of awaiting one tui-idle wait itself', async () => {
|
||||
const runtime = {
|
||||
setAutomationService: vi.fn(),
|
||||
notifyAutomationsChanged: vi.fn(),
|
||||
launchAgentTerminal: vi.fn(async () => ({
|
||||
handle: 'terminal-1',
|
||||
tabId: 'tab-1',
|
||||
paneKey: 'tab-1:pane-1',
|
||||
ptyId: 'pty-1',
|
||||
worktreeId: 'wt-1'
|
||||
})),
|
||||
showManagedWorktree: vi.fn(async () => ({ displayName: 'repo' })),
|
||||
waitForTerminal: vi.fn(),
|
||||
// Not resolvable by pane key yet: the launch's own handle must still reach the watcher.
|
||||
getTerminalHandleForPaneKey: vi.fn(() => null),
|
||||
getAgentStatusRowsForPane: vi.fn(() => [])
|
||||
}
|
||||
createRuntimeAutomationService({
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the mocked service never reads the store.
|
||||
store: {} as never,
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the dispatcher reads only the members faked above.
|
||||
runtime: runtime as never,
|
||||
headless: true
|
||||
})
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: captured from the mocked constructor above.
|
||||
const dispatcher = captured.dispatcher as HeadlessAutomationDispatcher
|
||||
const launch = await dispatcher({
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: a reuse-workspace automation; only these fields are read.
|
||||
automation: { workspaceMode: 'existing', workspaceId: 'wt-1', agentId: 'goose' } as never,
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: only the title is read.
|
||||
run: { title: 'Nightly' } as never,
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: unused for an existing workspace.
|
||||
target: {} as never
|
||||
})
|
||||
|
||||
expect(launch.completion).toBeUndefined()
|
||||
expect(launch.terminalPaneKey).toBe('tab-1:pane-1')
|
||||
expect(runtime.waitForTerminal).not.toHaveBeenCalled()
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: captured from the mocked constructor above.
|
||||
const observer = captured.observer as AutomationRunTerminalObserver
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: resolution reads only the pane key.
|
||||
expect(observer.resolveRunTerminal({ terminalPaneKey: 'tab-1:pane-1' } as never)).toBe(
|
||||
'terminal-1'
|
||||
)
|
||||
})
|
||||
})
|
||||
@@ -0,0 +1,188 @@
|
||||
/**
|
||||
* The AutomationService a runtime host executes schedules with. Shared by Electron startup and
|
||||
* orcad so both hosts dispatch through the same headless path.
|
||||
*/
|
||||
import type { ClaudeUsageStore } from '../claude-usage/store'
|
||||
import type { CodexUsageStore } from '../codex-usage/store'
|
||||
import type { Store } from '../persistence'
|
||||
import type { OrcaRuntimeService } from '../runtime/orca-runtime'
|
||||
import { AutomationService } from './service'
|
||||
import type { AutomationRun } from '../../shared/automations-types'
|
||||
import { createHeadlessRunTerminalRetention } from './headless-run-terminal-retention'
|
||||
import {
|
||||
getTuiAgentDetectCommands,
|
||||
isTuiAgent,
|
||||
TUI_AGENT_CONFIG
|
||||
} from '../../shared/tui-agent-config'
|
||||
import { buildHeadlessAutomationWorktreeCreateArgs } from './headless-workspace-create'
|
||||
import { createRuntimeAutomationRunTerminalObserver } from './runtime-terminal-run-observer'
|
||||
|
||||
const MAX_REMEMBERED_LAUNCHES = 256
|
||||
|
||||
export function createRuntimeAutomationService(input: {
|
||||
store: Store
|
||||
runtime: OrcaRuntimeService
|
||||
claudeUsage?: ClaudeUsageStore
|
||||
codexUsage?: CodexUsageStore
|
||||
/** A server process: it executes remote_host_service-owned schedules and dispatches headlessly. */
|
||||
headless: boolean
|
||||
}): AutomationService {
|
||||
const { store, runtime, claudeUsage, codexUsage } = input
|
||||
// The handle each headless launch returned, so its watcher never depends on a pane-key lookup.
|
||||
const launchedHandles = new Map<string, string>()
|
||||
const observer = createRuntimeAutomationRunTerminalObserver(runtime, {
|
||||
getAgentStatusRowsForPane: (paneKey) => runtime.getAgentStatusRowsForPane(paneKey),
|
||||
agentCommandsForRun: (run) =>
|
||||
automationAgentCommands(
|
||||
store.listAutomations().find((entry) => entry.id === run.automationId)?.agentId
|
||||
)
|
||||
})
|
||||
const service = new AutomationService(store, {
|
||||
claudeUsage,
|
||||
codexUsage,
|
||||
terminalObserver: {
|
||||
...observer,
|
||||
// This host's own launch handle first: it names the run's terminal without a lookup.
|
||||
resolveRunTerminal: (run) =>
|
||||
(run.terminalPaneKey ? launchedHandles.get(run.terminalPaneKey) : undefined) ??
|
||||
observer.resolveRunTerminal(run)
|
||||
},
|
||||
onAutomationsChanged: (payload) => runtime.notifyAutomationsChanged(payload),
|
||||
allowRemoteHostScheduling: input.headless,
|
||||
headlessDispatcher: input.headless
|
||||
? async ({ automation, run, target }) => {
|
||||
let terminalHandle: string
|
||||
let terminalSessionId: string | null = null
|
||||
let terminalPaneKey: string | null = null
|
||||
let terminalPtyId: string | null = null
|
||||
let workspaceId: string
|
||||
let workspaceDisplayName: string | null = null
|
||||
if (automation.workspaceMode === 'new_per_run') {
|
||||
const created = await runtime.createManagedWorktree(
|
||||
buildHeadlessAutomationWorktreeCreateArgs({ automation, run, repo: target.repo })
|
||||
)
|
||||
terminalHandle = created.startupTerminal?.handle ?? ''
|
||||
terminalSessionId = created.startupTerminal?.tabId ?? null
|
||||
terminalPaneKey = created.startupTerminal?.paneKey ?? null
|
||||
terminalPtyId = created.startupTerminal?.ptyId ?? null
|
||||
workspaceId = created.worktree.id
|
||||
workspaceDisplayName = created.worktree.displayName ?? null
|
||||
if (!terminalHandle) {
|
||||
throw new Error(
|
||||
created.warning ||
|
||||
'Automation workspace was created, but no agent terminal started.'
|
||||
)
|
||||
}
|
||||
} else {
|
||||
if (!automation.workspaceId) {
|
||||
throw new Error('The target workspace is no longer available.')
|
||||
}
|
||||
const terminal = await runtime.launchAgentTerminal(`id:${automation.workspaceId}`, {
|
||||
agent: automation.agentId,
|
||||
prompt: automation.prompt,
|
||||
title: run.title
|
||||
})
|
||||
terminalHandle = terminal.handle
|
||||
terminalSessionId = terminal.tabId ?? null
|
||||
terminalPaneKey = terminal.paneKey ?? null
|
||||
terminalPtyId = terminal.ptyId ?? null
|
||||
workspaceId = terminal.worktreeId
|
||||
const worktree = await runtime.showManagedWorktree(`id:${workspaceId}`)
|
||||
workspaceDisplayName = worktree.displayName ?? null
|
||||
}
|
||||
if (terminalPaneKey) {
|
||||
launchedHandles.set(terminalPaneKey, terminalHandle)
|
||||
if (launchedHandles.size > MAX_REMEMBERED_LAUNCHES) {
|
||||
launchedHandles.delete(launchedHandles.keys().next().value ?? '')
|
||||
}
|
||||
}
|
||||
// No completion: the run's watcher observes it, retrying wait timeouts until it settles.
|
||||
return {
|
||||
workspaceId,
|
||||
workspaceDisplayName,
|
||||
terminalSessionId,
|
||||
terminalPaneKey,
|
||||
terminalPtyId
|
||||
}
|
||||
}
|
||||
: undefined
|
||||
})
|
||||
runtime.setAutomationService(service)
|
||||
if (input.headless) {
|
||||
bindHeadlessRunTerminalRetention(service, store, runtime)
|
||||
}
|
||||
return service
|
||||
}
|
||||
|
||||
/** A headless host has no renderer to close finished run terminals; its service does it. */
|
||||
function bindHeadlessRunTerminalRetention(
|
||||
service: AutomationService,
|
||||
store: Store,
|
||||
runtime: OrcaRuntimeService
|
||||
): void {
|
||||
const retention = createHeadlessRunTerminalRetention({
|
||||
listRuns: () => store.listAutomationRuns(),
|
||||
terminalClientUse: (run) =>
|
||||
run.terminalPtyId ? runtime.readTerminalClientUse(run.terminalPtyId) : 'unknown',
|
||||
runTerminalAlive: (run) => runOwnHandle(runtime, run) !== null,
|
||||
shellAloneAtPrompt: (run) =>
|
||||
run.terminalPtyId
|
||||
? runtime.confirmTerminalShellAlone(run.terminalPtyId)
|
||||
: Promise.resolve(false),
|
||||
closeRunTerminal: async (run) => {
|
||||
// Only the run's own process: a restarted pane holds a new PTY a user may be working in.
|
||||
const handle = runOwnHandle(runtime, run)
|
||||
if (!handle) {
|
||||
return false
|
||||
}
|
||||
// The run's pane only: a pane a user split into the same tab is theirs.
|
||||
await runtime.closeTerminal(handle)
|
||||
return true
|
||||
},
|
||||
forgetRunTerminal: async (run) => {
|
||||
// A run still dispatched had its agent exit without a result, proven by its idle shell.
|
||||
const unfinished = run.status === 'dispatched'
|
||||
await service.markDispatchResult({
|
||||
runId: run.id,
|
||||
status: unfinished ? 'dispatch_failed' : run.status,
|
||||
error: unfinished
|
||||
? (run.error ?? 'The agent exited without reporting completion.')
|
||||
: run.error,
|
||||
terminalSessionId: null,
|
||||
terminalPaneKey: null,
|
||||
terminalPtyId: null
|
||||
})
|
||||
}
|
||||
})
|
||||
service.releaseFinishedRunTerminals = () => retention.drain()
|
||||
const start = service.start.bind(service)
|
||||
const stop = service.stop.bind(service)
|
||||
service.start = () => {
|
||||
start()
|
||||
retention.start()
|
||||
}
|
||||
service.stop = () => {
|
||||
retention.stop()
|
||||
stop()
|
||||
}
|
||||
}
|
||||
|
||||
function automationAgentCommands(agentId: string | undefined): string[] {
|
||||
if (!isTuiAgent(agentId)) {
|
||||
return []
|
||||
}
|
||||
const config = TUI_AGENT_CONFIG[agentId]
|
||||
return [...getTuiAgentDetectCommands(config), config.launchCmd]
|
||||
}
|
||||
|
||||
/** The run's pane handle while that pane still holds the run's own PTY; null otherwise. */
|
||||
function runOwnHandle(runtime: OrcaRuntimeService, run: AutomationRun): string | null {
|
||||
const handle = run.terminalPaneKey
|
||||
? runtime.getTerminalHandleForPaneKey(run.terminalPaneKey)
|
||||
: null
|
||||
return handle &&
|
||||
run.terminalPtyId &&
|
||||
runtime.getTerminalPtyIdForHandle(handle) === run.terminalPtyId
|
||||
? handle
|
||||
: null
|
||||
}
|
||||
@@ -0,0 +1,173 @@
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
|
||||
import type { AutomationRun } from '../../shared/automations-types'
|
||||
import { AGENT_START_GRACE_MS } from './automation-run-agent-evidence'
|
||||
import {
|
||||
createRuntimeAutomationRunTerminalObserver,
|
||||
type AutomationRunTerminalHost
|
||||
} from './runtime-terminal-run-observer'
|
||||
|
||||
const HANDLE = 'terminal-1'
|
||||
const PANE_KEY = 'tab-1:pane-1'
|
||||
const RUNTIME_TUI_IDLE_TIMEOUT_MS = 5 * 60 * 1000
|
||||
|
||||
/** A pane where idle is idle whatever painted it: a ready shell prompt satisfies tui-idle too. */
|
||||
function createPane(initial: { idle: boolean; tail: string[]; idleEdgeOnly?: boolean }) {
|
||||
let idle = initial.idle
|
||||
// The real runtime may settle tui-idle once per idle edge: an already-idle shell only times out.
|
||||
let edgeSpent = false
|
||||
let tail = initial.tail
|
||||
let rows: { receivedAt: number }[] = []
|
||||
const waiters = new Set<() => void>()
|
||||
const runtime: AutomationRunTerminalHost = {
|
||||
getTerminalHandleForPaneKey: () => HANDLE,
|
||||
readTerminal: async () => ({ tail }),
|
||||
waitForTerminal: (_handle, options) => {
|
||||
if (options?.signal?.aborted) {
|
||||
return Promise.reject(new Error('request_aborted'))
|
||||
}
|
||||
if (idle && !(initial.idleEdgeOnly && edgeSpent)) {
|
||||
edgeSpent = true
|
||||
return Promise.resolve({ satisfied: true })
|
||||
}
|
||||
return new Promise((resolve, reject) => {
|
||||
const timer = setTimeout(() => {
|
||||
waiters.delete(wake)
|
||||
reject(new Error('timeout'))
|
||||
}, options?.timeoutMs ?? RUNTIME_TUI_IDLE_TIMEOUT_MS)
|
||||
const wake = (): void => {
|
||||
clearTimeout(timer)
|
||||
resolve({ satisfied: true })
|
||||
}
|
||||
waiters.add(wake)
|
||||
options?.signal?.addEventListener('abort', () => {
|
||||
clearTimeout(timer)
|
||||
waiters.delete(wake)
|
||||
reject(new Error('request_aborted'))
|
||||
})
|
||||
})
|
||||
}
|
||||
}
|
||||
return {
|
||||
runtime,
|
||||
agentRows: () => rows,
|
||||
report: (receivedAt = Date.now()) => (rows = [{ receivedAt }]),
|
||||
setTail: (next: string[]) => (tail = next),
|
||||
becomeIdle: () => {
|
||||
idle = true
|
||||
for (const wake of waiters) {
|
||||
waiters.delete(wake)
|
||||
wake()
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
function observe(pane: ReturnType<typeof createPane>, controller = new AbortController()) {
|
||||
const observer = createRuntimeAutomationRunTerminalObserver(pane.runtime, {
|
||||
getAgentStatusRowsForPane: (paneKey) => (paneKey === PANE_KEY ? pane.agentRows() : []),
|
||||
agentCommandsForRun: () => ['goose']
|
||||
})
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the observer reads only these run fields.
|
||||
const run = {
|
||||
terminalPaneKey: PANE_KEY,
|
||||
startedAt: Date.now(),
|
||||
dispatchedAt: null
|
||||
} as AutomationRun
|
||||
const settled = vi.fn()
|
||||
const failed = vi.fn()
|
||||
const promise = observer
|
||||
.observeCompletion(HANDLE, { signal: controller.signal, run })
|
||||
.then(settled, failed)
|
||||
return { settled, failed, promise, controller }
|
||||
}
|
||||
|
||||
beforeEach(() => {
|
||||
vi.useFakeTimers()
|
||||
})
|
||||
|
||||
afterEach(() => {
|
||||
vi.useRealTimers()
|
||||
})
|
||||
|
||||
describe('observing a run with agent evidence', () => {
|
||||
it.each([
|
||||
['bash', 'bash: goose: command not found'],
|
||||
['zsh', 'zsh: command not found: goose'],
|
||||
['dash', 'sh: 1: goose: not found'],
|
||||
['fish', 'fish: Unknown command: goose'],
|
||||
['PowerShell', "goose: The term 'goose' is not recognized as the name of a cmdlet"]
|
||||
])('fails, never completes, a run whose agent %s cannot find', async (_shell, refusal) => {
|
||||
const run = observe(createPane({ idle: true, tail: ['$ goose run', refusal, '$'] }))
|
||||
await run.promise
|
||||
expect(run.settled).toHaveBeenCalledWith(
|
||||
expect.objectContaining({
|
||||
status: 'dispatch_failed',
|
||||
error: expect.stringContaining(refusal)
|
||||
})
|
||||
)
|
||||
})
|
||||
|
||||
it('completes once the agent itself reported for the run pane', async () => {
|
||||
const pane = createPane({ idle: false, tail: ['working'] })
|
||||
const run = observe(pane)
|
||||
await vi.advanceTimersByTimeAsync(1_000)
|
||||
pane.report()
|
||||
pane.becomeIdle()
|
||||
await run.promise
|
||||
expect(run.settled).toHaveBeenCalledWith(expect.objectContaining({ status: 'completed' }))
|
||||
})
|
||||
|
||||
it('keeps idle-means-done for an agent that never reports, after its start window', async () => {
|
||||
const pane = createPane({ idle: true, tail: ['summary written', '$'] })
|
||||
pane.report(1)
|
||||
const run = observe(pane)
|
||||
await vi.advanceTimersByTimeAsync(10_000)
|
||||
expect(run.settled).not.toHaveBeenCalled()
|
||||
await vi.advanceTimersByTimeAsync(AGENT_START_GRACE_MS)
|
||||
await run.promise
|
||||
expect(run.settled).toHaveBeenCalledWith(expect.objectContaining({ status: 'completed' }))
|
||||
})
|
||||
|
||||
it("never reads an agent's own tool output as its command missing", async () => {
|
||||
const tail = ['running tests', 'bash: pytest: command not found', 'fell back to unittest', '$']
|
||||
const run = observe(createPane({ idle: true, tail }))
|
||||
await vi.advanceTimersByTimeAsync(AGENT_START_GRACE_MS + 1_000)
|
||||
await run.promise
|
||||
expect(run.settled).toHaveBeenCalledWith(expect.objectContaining({ status: 'completed' }))
|
||||
})
|
||||
|
||||
it('completes an agent that exited before the window once the window passes, not at a wait timeout', async () => {
|
||||
// The stub ran, exited 0 and left an idle shell; no agent status, no new idle edge.
|
||||
const pane = createPane({ idle: true, tail: ['stub done', '$'], idleEdgeOnly: true })
|
||||
const run = observe(pane)
|
||||
await vi.advanceTimersByTimeAsync(AGENT_START_GRACE_MS - 1_000)
|
||||
expect(run.settled).not.toHaveBeenCalled()
|
||||
await vi.advanceTimersByTimeAsync(2_000)
|
||||
expect(run.settled).toHaveBeenCalledWith(expect.objectContaining({ status: 'completed' }))
|
||||
expect(run.failed).not.toHaveBeenCalled()
|
||||
})
|
||||
|
||||
it('keeps watching an agent past a tui-idle wait timeout and completes it later', async () => {
|
||||
const pane = createPane({ idle: false, tail: ['working'] })
|
||||
const run = observe(pane)
|
||||
await vi.advanceTimersByTimeAsync(RUNTIME_TUI_IDLE_TIMEOUT_MS * 2 + 1_000)
|
||||
expect(run.settled).not.toHaveBeenCalled()
|
||||
expect(run.failed).not.toHaveBeenCalled()
|
||||
|
||||
pane.report()
|
||||
pane.becomeIdle()
|
||||
await run.promise
|
||||
expect(run.settled).toHaveBeenCalledWith(expect.objectContaining({ status: 'completed' }))
|
||||
})
|
||||
|
||||
it('stops on cancellation without recording any result', async () => {
|
||||
const pane = createPane({ idle: true, tail: ['$'] })
|
||||
const run = observe(pane)
|
||||
await vi.advanceTimersByTimeAsync(5_000)
|
||||
run.controller.abort()
|
||||
await vi.advanceTimersByTimeAsync(AGENT_START_GRACE_MS)
|
||||
await run.promise
|
||||
expect(run.settled).not.toHaveBeenCalled()
|
||||
expect(run.failed).toHaveBeenCalledWith(expect.objectContaining({ message: 'request_aborted' }))
|
||||
})
|
||||
})
|
||||
@@ -3,10 +3,13 @@ import type {
|
||||
AutomationRunCompletionObservation,
|
||||
AutomationRunTerminalObserver
|
||||
} from './run-completion-watcher'
|
||||
import type { AutomationRunOutputSnapshot } from '../../shared/automations-types'
|
||||
import type { AutomationRun, AutomationRunOutputSnapshot } from '../../shared/automations-types'
|
||||
import { judgeIdleRun, type AutomationRunAgentEvidence } from './automation-run-agent-evidence'
|
||||
|
||||
const TERMINAL_SNAPSHOT_LIMIT = 2_000
|
||||
|
||||
/** Cadence for re-checking an idle pane whose agent has not reported yet. */
|
||||
const AGENT_EVIDENCE_POLL_INTERVAL_MS = 1_000
|
||||
/** Cadence for re-probing a pane that already satisfied tui-idle at dispatch. */
|
||||
const AGENT_START_POLL_INTERVAL_MS = 250
|
||||
/** Kept under the runtime's 2s tui-idle fallback poll so a probe waiter is torn
|
||||
@@ -34,6 +37,63 @@ export type AutomationRunTerminalHost = {
|
||||
readTerminal(handle: string, opts?: { limit?: number }): Promise<{ tail: string[] }>
|
||||
}
|
||||
|
||||
async function readTail(runtime: AutomationRunTerminalHost, handle: string): Promise<string[]> {
|
||||
try {
|
||||
return (await runtime.readTerminal(handle, { limit: TERMINAL_SNAPSHOT_LIMIT })).tail
|
||||
} catch {
|
||||
return []
|
||||
}
|
||||
}
|
||||
|
||||
function snapshotOf(tail: readonly string[]): AutomationRunOutputSnapshot | null {
|
||||
const snapshotBuffer = createHeadlessAutomationOutputSnapshotBuffer()
|
||||
snapshotBuffer.append(tail.join('\n'))
|
||||
return snapshotBuffer.snapshot()
|
||||
}
|
||||
|
||||
/** A satisfied wait judged by the run's agent evidence; null while the agent may still start. */
|
||||
async function judgeIdleTail(
|
||||
tail: readonly string[],
|
||||
evidence: AutomationRunAgentEvidence,
|
||||
run: AutomationRun,
|
||||
runStartedAt: number
|
||||
): Promise<AutomationRunCompletionObservation | null> {
|
||||
const verdict = judgeIdleRun(evidence, run, tail, runStartedAt, Date.now())
|
||||
if (verdict.kind === 'wait') {
|
||||
return null
|
||||
}
|
||||
return verdict.kind === 'completed'
|
||||
? { status: 'completed', outputSnapshot: snapshotOf(tail), error: null }
|
||||
: { status: 'dispatch_failed', outputSnapshot: snapshotOf(tail), error: verdict.error }
|
||||
}
|
||||
|
||||
/**
|
||||
* After a satisfied wait inside the agent-start window. An idle shell never produces a new idle
|
||||
* edge, so a fresh wait would only time out; instead the pane stays idle while its output is
|
||||
* unchanged, and that state is re-judged until the window passes or the agent reports.
|
||||
* Null once the pane changes: something is running, so the caller waits again.
|
||||
*/
|
||||
async function settleIdlePane(
|
||||
runtime: AutomationRunTerminalHost,
|
||||
handle: string,
|
||||
judged: { evidence: AutomationRunAgentEvidence; run: AutomationRun },
|
||||
runStartedAt: number,
|
||||
signal: AbortSignal
|
||||
): Promise<AutomationRunCompletionObservation | null> {
|
||||
const idleTail = (await readTail(runtime, handle)).join('\n')
|
||||
for (;;) {
|
||||
const tail = await readTail(runtime, handle)
|
||||
if (tail.join('\n') !== idleTail) {
|
||||
return null
|
||||
}
|
||||
const observation = await judgeIdleTail(tail, judged.evidence, judged.run, runStartedAt)
|
||||
if (observation) {
|
||||
return observation
|
||||
}
|
||||
await sleep(AGENT_EVIDENCE_POLL_INTERVAL_MS, signal)
|
||||
}
|
||||
}
|
||||
|
||||
function isTerminalWaitTimeout(error: unknown): boolean {
|
||||
return error instanceof Error && error.message === 'timeout'
|
||||
}
|
||||
@@ -146,19 +206,25 @@ async function buildUnobservedObservation(
|
||||
}
|
||||
|
||||
export function createRuntimeAutomationRunTerminalObserver(
|
||||
runtime: AutomationRunTerminalHost
|
||||
runtime: AutomationRunTerminalHost,
|
||||
/** Judges idle panes by the agent's own status; without it, an idle pane after the busy edge completes. */
|
||||
evidence?: AutomationRunAgentEvidence
|
||||
): AutomationRunTerminalObserver {
|
||||
return {
|
||||
resolveRunTerminal: (run) =>
|
||||
run.terminalPaneKey ? runtime.getTerminalHandleForPaneKey(run.terminalPaneKey) : null,
|
||||
observeCompletion: async (handle, { signal }) => {
|
||||
observeCompletion: async (handle, { signal, run }) => {
|
||||
const startedAt = Date.now()
|
||||
const judged = evidence && run?.terminalPaneKey ? { evidence, run } : null
|
||||
// The run's own start bounds which agent status counts and when idleness is believed.
|
||||
const runStartedAt = run?.startedAt ?? run?.dispatchedAt ?? startedAt
|
||||
// Why: tui-idle is level-triggered, so a reused pane still idle from the
|
||||
// PREVIOUS run satisfies it before this run's agent has typed a character.
|
||||
// Evidence that predates dispatch proves nothing about this run, so require
|
||||
// the pane to leave that state first — the busy edge the renderer's own
|
||||
// dispatch observer requires on reuse (requireWorkingAfterStart).
|
||||
if (await isTuiIdleSatisfiedNow(runtime, handle, signal)) {
|
||||
// Agent evidence already discounts a previous run's idleness, so it skips the busy edge.
|
||||
if (!judged && (await isTuiIdleSatisfiedNow(runtime, handle, signal))) {
|
||||
const started = await waitForAgentStart(
|
||||
runtime,
|
||||
handle,
|
||||
@@ -177,7 +243,13 @@ export function createRuntimeAutomationRunTerminalObserver(
|
||||
for (;;) {
|
||||
try {
|
||||
const wait = await runtime.waitForTerminal(handle, { condition: 'tui-idle', signal })
|
||||
return await buildObservation(runtime, handle, wait)
|
||||
if (!judged || !wait.satisfied) {
|
||||
return await buildObservation(runtime, handle, wait)
|
||||
}
|
||||
const observation = await settleIdlePane(runtime, handle, judged, runStartedAt, signal)
|
||||
if (observation) {
|
||||
return observation
|
||||
}
|
||||
} catch (error) {
|
||||
// Why: tui-idle waits expire on their own schedule; an agent still
|
||||
// working past that window is live, so re-arm rather than fail it.
|
||||
|
||||
@@ -52,6 +52,8 @@ export class AutomationService {
|
||||
private readonly codexUsage: CodexUsageStore | null
|
||||
private readonly allowRemoteHostScheduling: boolean
|
||||
private readonly headlessDispatcher: HeadlessAutomationDispatcher | null
|
||||
/** Set on a headless host: closes completed, unused run terminals now; resolves how many. */
|
||||
releaseFinishedRunTerminals: (() => Promise<number>) | null = null
|
||||
private readonly publish: PublishAutomationsChanged | null
|
||||
private readonly runs: AutomationRunWriter
|
||||
private readonly completionWatcher: AutomationRunCompletionWatcher | null
|
||||
|
||||
@@ -214,9 +214,9 @@ describe('a Codex send made after Codex answered an earlier one, before it opene
|
||||
const sending = rig.send('client-2')
|
||||
expect(await settledWithin(sending)).toBe('held')
|
||||
const methods = () =>
|
||||
rig.codex.connections[0]!.calls
|
||||
.map(({ method }) => method)
|
||||
.filter((method) => method !== 'model/list' && method !== 'config/read')
|
||||
rig.codex.connections[0]!.calls.map(({ method }) => method).filter(
|
||||
(method) => method !== 'model/list' && method !== 'config/read'
|
||||
)
|
||||
expect(methods()).toEqual(['thread/start', 'turn/start'])
|
||||
return { ...rig, sending, methods }
|
||||
}
|
||||
|
||||
@@ -0,0 +1,5 @@
|
||||
/** What a `ptySpawnHealth` reply proves on this platform. */
|
||||
export function ptySpawnHealthPlatformCoverage(): 'pty-spawn' | 'handshake' {
|
||||
// Why handshake on Windows: preflightPtySpawnHealth skips the spawn probe there.
|
||||
return process.platform === 'win32' ? 'handshake' : 'pty-spawn'
|
||||
}
|
||||
@@ -11,6 +11,7 @@ import { getDaemonPidPath, serializeDaemonPidFile } from './daemon-spawner'
|
||||
import type { SocketProbeOutcome } from './daemon-endpoint-probe'
|
||||
import {
|
||||
checkDaemonHealth,
|
||||
checkDaemonHealthWithCoverage,
|
||||
E2E_FORCE_DAEMON_HEALTH_UNREACHABLE_ENV,
|
||||
healthCheckDaemon
|
||||
} from './daemon-health'
|
||||
@@ -105,12 +106,61 @@ describe('daemon health', () => {
|
||||
try {
|
||||
await expect(checkDaemonHealth(socketPath, tokenPath)).resolves.toBe('healthy')
|
||||
await expect(healthCheckDaemon(socketPath, tokenPath)).resolves.toBe(true)
|
||||
expect(ptySpawnHealthCheck).toHaveBeenCalledTimes(2)
|
||||
await expect(checkDaemonHealthWithCoverage(socketPath, tokenPath)).resolves.toEqual({
|
||||
verdict: 'healthy',
|
||||
coverage: process.platform === 'win32' ? 'handshake' : 'pty-spawn'
|
||||
})
|
||||
expect(ptySpawnHealthCheck).toHaveBeenCalledTimes(3)
|
||||
} finally {
|
||||
await server.shutdown()
|
||||
}
|
||||
})
|
||||
|
||||
it('treats missing coverage from a legacy Windows daemon as handshake-only', async () => {
|
||||
writeFileSync(tokenPath, 'legacy-token')
|
||||
const server = createServer((socket) => {
|
||||
let pending = ''
|
||||
socket.on('data', (chunk) => {
|
||||
pending += chunk.toString()
|
||||
for (;;) {
|
||||
const newline = pending.indexOf('\n')
|
||||
if (newline === -1) {
|
||||
return
|
||||
}
|
||||
const message: unknown = JSON.parse(pending.slice(0, newline))
|
||||
const type =
|
||||
typeof message === 'object' && message !== null && 'type' in message
|
||||
? message.type
|
||||
: undefined
|
||||
pending = pending.slice(newline + 1)
|
||||
if (type === 'hello') {
|
||||
socket.write(`${JSON.stringify({ type: 'hello', ok: true })}\n`)
|
||||
} else if (type === 'ptySpawnHealth') {
|
||||
socket.write(
|
||||
`${JSON.stringify({ id: 'health-1', ok: true, payload: { healthy: true } })}\n`
|
||||
)
|
||||
}
|
||||
}
|
||||
})
|
||||
})
|
||||
await new Promise<void>((resolve, reject) => {
|
||||
server.once('error', reject)
|
||||
server.listen(socketPath, resolve)
|
||||
})
|
||||
|
||||
const platform = Object.getOwnPropertyDescriptor(process, 'platform')!
|
||||
Object.defineProperty(process, 'platform', { configurable: true, value: 'win32' })
|
||||
try {
|
||||
await expect(checkDaemonHealthWithCoverage(socketPath, tokenPath)).resolves.toEqual({
|
||||
verdict: 'healthy',
|
||||
coverage: 'handshake'
|
||||
})
|
||||
} finally {
|
||||
Object.defineProperty(process, 'platform', platform)
|
||||
await closeServer(server)
|
||||
}
|
||||
})
|
||||
|
||||
it('fails when a protocol-healthy daemon cannot spawn PTYs', async () => {
|
||||
const server = new DaemonServer({
|
||||
socketPath,
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
import { existsSync, readFileSync } from 'node:fs'
|
||||
import { connect, type Socket } from 'node:net'
|
||||
import { encodeNdjson } from './ndjson'
|
||||
import { ptySpawnHealthPlatformCoverage } from './daemon-health-identity'
|
||||
import {
|
||||
PROTOCOL_VERSION,
|
||||
type HelloMessage,
|
||||
@@ -21,15 +22,39 @@ export const E2E_FORCE_DAEMON_HEALTH_UNREACHABLE_ENV = 'ORCA_E2E_FORCE_DAEMON_HE
|
||||
// also covers a live-but-wedged daemon that simply missed the RPC budget.
|
||||
export type DaemonHealth = 'healthy' | 'unreachable' | 'rejected' | 'pty-spawn-unhealthy'
|
||||
|
||||
export function checkDaemonHealth(socketPath: string, tokenPath: string): Promise<DaemonHealth> {
|
||||
export type DaemonHealthCheck = {
|
||||
verdict: DaemonHealth
|
||||
coverage: 'pty-spawn' | 'handshake'
|
||||
}
|
||||
|
||||
function readPtySpawnHealthCoverage(
|
||||
payload: unknown,
|
||||
fallback: DaemonHealthCheck['coverage']
|
||||
): DaemonHealthCheck['coverage'] {
|
||||
if (typeof payload !== 'object' || payload === null) {
|
||||
return fallback
|
||||
}
|
||||
const coverage = 'coverage' in payload ? payload.coverage : undefined
|
||||
return coverage === 'pty-spawn' || coverage === 'handshake' ? coverage : fallback
|
||||
}
|
||||
|
||||
export function checkDaemonHealthWithCoverage(
|
||||
socketPath: string,
|
||||
tokenPath: string
|
||||
): Promise<DaemonHealthCheck> {
|
||||
return new Promise((resolve) => {
|
||||
// Older Windows daemons answered this RPC without spawning; an absent optional coverage
|
||||
// field must preserve that weaker meaning during adoption.
|
||||
const fallbackCoverage = ptySpawnHealthPlatformCoverage()
|
||||
const resolveVerdict = (verdict: DaemonHealth): void =>
|
||||
resolve({ verdict, coverage: fallbackCoverage })
|
||||
if (process.env[E2E_FORCE_DAEMON_HEALTH_UNREACHABLE_ENV] === '1') {
|
||||
resolve('unreachable')
|
||||
resolveVerdict('unreachable')
|
||||
return
|
||||
}
|
||||
|
||||
if (process.platform !== 'win32' && !existsSync(socketPath)) {
|
||||
resolve('unreachable')
|
||||
resolveVerdict('unreachable')
|
||||
return
|
||||
}
|
||||
|
||||
@@ -37,13 +62,13 @@ export function checkDaemonHealth(socketPath: string, tokenPath: string): Promis
|
||||
try {
|
||||
token = readFileSync(tokenPath, 'utf8').trim()
|
||||
} catch {
|
||||
resolve('unreachable')
|
||||
resolveVerdict('unreachable')
|
||||
return
|
||||
}
|
||||
|
||||
let settled = false
|
||||
let sock: Socket | null = null
|
||||
const settle = (result: DaemonHealth): void => {
|
||||
const settle = (result: DaemonHealthCheck): void => {
|
||||
if (settled) {
|
||||
return
|
||||
}
|
||||
@@ -58,7 +83,7 @@ export function checkDaemonHealth(socketPath: string, tokenPath: string): Promis
|
||||
sock?.off('connect', onConnect)
|
||||
sock?.off('data', onData)
|
||||
}
|
||||
const onError = (): void => settle('unreachable')
|
||||
const onError = (): void => settle({ verdict: 'unreachable', coverage: fallbackCoverage })
|
||||
const onConnect = (): void => {
|
||||
const hello: HelloMessage = {
|
||||
type: 'hello',
|
||||
@@ -89,13 +114,13 @@ export function checkDaemonHealth(socketPath: string, tokenPath: string): Promis
|
||||
try {
|
||||
message = JSON.parse(line) as Record<string, unknown>
|
||||
} catch {
|
||||
settle('rejected')
|
||||
settle({ verdict: 'rejected', coverage: fallbackCoverage })
|
||||
return
|
||||
}
|
||||
|
||||
if (message.type === 'hello') {
|
||||
if (!(message as HelloResponse).ok) {
|
||||
settle('rejected')
|
||||
settle({ verdict: 'rejected', coverage: fallbackCoverage })
|
||||
return
|
||||
}
|
||||
// Why: a protocol-live daemon with a stale cwd or node-pty helper
|
||||
@@ -106,12 +131,18 @@ export function checkDaemonHealth(socketPath: string, tokenPath: string): Promis
|
||||
}
|
||||
|
||||
if (message.id === 'health-1') {
|
||||
settle(message.ok === true ? 'healthy' : 'pty-spawn-unhealthy')
|
||||
settle({
|
||||
verdict: message.ok === true ? 'healthy' : 'pty-spawn-unhealthy',
|
||||
coverage: readPtySpawnHealthCoverage(message.payload, fallbackCoverage)
|
||||
})
|
||||
return
|
||||
}
|
||||
}
|
||||
}
|
||||
const timer = setTimeout(() => settle('unreachable'), HEALTH_CHECK_TIMEOUT_MS)
|
||||
const timer = setTimeout(
|
||||
() => settle({ verdict: 'unreachable', coverage: fallbackCoverage }),
|
||||
HEALTH_CHECK_TIMEOUT_MS
|
||||
)
|
||||
|
||||
sock = connect({ path: socketPath })
|
||||
sock.on('error', onError)
|
||||
@@ -122,6 +153,13 @@ export function checkDaemonHealth(socketPath: string, tokenPath: string): Promis
|
||||
})
|
||||
}
|
||||
|
||||
export async function checkDaemonHealth(
|
||||
socketPath: string,
|
||||
tokenPath: string
|
||||
): Promise<DaemonHealth> {
|
||||
return (await checkDaemonHealthWithCoverage(socketPath, tokenPath)).verdict
|
||||
}
|
||||
|
||||
export async function healthCheckDaemon(socketPath: string, tokenPath: string): Promise<boolean> {
|
||||
return (await checkDaemonHealth(socketPath, tokenPath)) === 'healthy'
|
||||
}
|
||||
|
||||
@@ -442,6 +442,35 @@ describe('current daemon lifecycle retirement', () => {
|
||||
adopted.dispose()
|
||||
})
|
||||
|
||||
it('atomically retires an idle daemon and permanently fences adapter spawns', async () => {
|
||||
await startServer()
|
||||
const adapter = new DaemonPtyAdapter({ socketPath, tokenPath })
|
||||
|
||||
await expect(adapter.requestIdleRetirement()).resolves.toEqual({ state: 'retiring' })
|
||||
await expect(
|
||||
adapter.spawn({ sessionId: 'late-after-decommission', cols: 80, rows: 24 })
|
||||
).rejects.toThrow('Terminal daemon is decommissioning')
|
||||
await waitFor(() => onIdleShutdown.mock.calls.length === 1)
|
||||
adapter.dispose()
|
||||
})
|
||||
|
||||
it('reopens adapter admission when the daemon refuses retirement for a live session', async () => {
|
||||
await startServer()
|
||||
const adapter = new DaemonPtyAdapter({ socketPath, tokenPath })
|
||||
await adapter.spawn({ sessionId: 'already-live', cols: 80, rows: 24 })
|
||||
|
||||
await expect(adapter.requestIdleRetirement()).resolves.toEqual({
|
||||
state: 'busy',
|
||||
liveSessions: 1,
|
||||
admissionReopened: true
|
||||
})
|
||||
await expect(
|
||||
adapter.spawn({ sessionId: 'allowed-after-refusal', cols: 80, rows: 24 })
|
||||
).resolves.toMatchObject({ id: 'allowed-after-refusal' })
|
||||
expect(onIdleShutdown).not.toHaveBeenCalled()
|
||||
adapter.dispose()
|
||||
})
|
||||
|
||||
it('does not let repeated authenticated control probes extend the startup deadline', async () => {
|
||||
await startServer()
|
||||
const healthControl = connect(socketPath)
|
||||
|
||||
@@ -8,6 +8,9 @@ export {
|
||||
getDaemonEndpointFacts,
|
||||
getDaemonProvider,
|
||||
listLiveDaemonPtyIds,
|
||||
listLiveDaemonSessions,
|
||||
requestIdleDaemonRetirement,
|
||||
releaseDaemonRetirementFence,
|
||||
readDaemonPidRecord,
|
||||
replaceDaemonProvider,
|
||||
shutdownDaemon,
|
||||
|
||||
@@ -103,6 +103,22 @@ describe('launchDaemonChild identity', () => {
|
||||
})
|
||||
})
|
||||
|
||||
describe('launchDaemonChild startup budget', () => {
|
||||
it('waits for readiness past the default 10 s when the caller allows it', async () => {
|
||||
vi.useFakeTimers()
|
||||
const child = fakeDaemonChild(4242)
|
||||
spawnDaemonChildProcessMock.mockReturnValue(child)
|
||||
try {
|
||||
const launch = launchDaemonChild({ ...LAUNCH_OPTIONS, startupTimeoutMs: 30_000 })
|
||||
await vi.advanceTimersByTimeAsync(20_000)
|
||||
child.emit('message', { type: 'ready', pid: 4242, startedAtMs: 1_000_000 })
|
||||
await expect(launch).resolves.toMatchObject({ identity: { pid: 4242 } })
|
||||
} finally {
|
||||
vi.useRealTimers()
|
||||
}
|
||||
})
|
||||
})
|
||||
|
||||
describe('launchDaemonChild durable-scope fallback', () => {
|
||||
it('retries once without cgroup isolation when the scoped attempt fails', async () => {
|
||||
isDurableDaemonScopeSupportedMock.mockReturnValue(true)
|
||||
|
||||
@@ -14,6 +14,8 @@ export type DaemonChildSpawnOptions = {
|
||||
pidPath: string
|
||||
launchNonce: string
|
||||
macosLoginSessionWatch: boolean
|
||||
/** Defaults to 10 s; a cold first exec on Windows can need longer. */
|
||||
startupTimeoutMs?: number
|
||||
}
|
||||
|
||||
function buildDaemonScriptArgs(options: DaemonChildSpawnOptions): string[] {
|
||||
|
||||
@@ -12,6 +12,7 @@ import { unlinkOwnedDaemonPidFile } from './daemon-spawner'
|
||||
const DAEMON_CHILD_TERMINATION_GRACE_MS = 5_000
|
||||
const DAEMON_CHILD_FORCE_EXIT_WAIT_MS = 1_000
|
||||
const STARTUP_STDERR_MAX_BYTES = 8192
|
||||
const DEFAULT_DAEMON_STARTUP_TIMEOUT_MS = 10_000
|
||||
|
||||
export class DaemonEndpointUnavailableError extends Error {
|
||||
constructor(
|
||||
@@ -187,7 +188,7 @@ async function launchDaemonChildAttempt(
|
||||
|
||||
timer = setTimeout(() => {
|
||||
void fail(new Error('Daemon startup timed out'))
|
||||
}, 10000)
|
||||
}, options.startupTimeoutMs ?? DEFAULT_DAEMON_STARTUP_TIMEOUT_MS)
|
||||
|
||||
child.on('message', onReadyMessage)
|
||||
child.on('error', onStartupError)
|
||||
|
||||
@@ -52,9 +52,11 @@ function createPreservedDaemonHandle(
|
||||
return handle
|
||||
}
|
||||
|
||||
export type DaemonLaunchPolicy = { macosLoginSessionWatch?: boolean; startupTimeoutMs?: number }
|
||||
|
||||
export function createOutOfProcessLauncher(
|
||||
runtimeDir: string,
|
||||
macosLoginSessionWatch = false
|
||||
{ macosLoginSessionWatch = false, startupTimeoutMs }: DaemonLaunchPolicy = {}
|
||||
): DaemonLauncher {
|
||||
return async (socketPath, tokenPath, suppliedPidPath, suppliedLaunchNonce) => {
|
||||
const entryPath = getDaemonEntryPath()
|
||||
@@ -132,7 +134,8 @@ export function createOutOfProcessLauncher(
|
||||
tokenPath,
|
||||
pidPath,
|
||||
launchNonce,
|
||||
macosLoginSessionWatch
|
||||
macosLoginSessionWatch,
|
||||
startupTimeoutMs
|
||||
})
|
||||
} catch (error) {
|
||||
if (!(error instanceof DaemonEndpointUnavailableError) || error.reason !== 'occupied') {
|
||||
|
||||
@@ -33,41 +33,40 @@ describe('daemon process inspection', () => {
|
||||
expect(runCommand).not.toHaveBeenCalled()
|
||||
})
|
||||
|
||||
it('asks PowerShell to report a failed CIM query instead of an absent process', async () => {
|
||||
const runCommand = vi.fn(
|
||||
async (_file: string, _args: string[], _timeoutMs: number) =>
|
||||
'{"status":"present","cmd":"daemon","start":1}'
|
||||
)
|
||||
it('reads the Windows process from the process table, not a PowerShell spawn', async () => {
|
||||
const runCommand = vi.fn()
|
||||
const readProcessTable = vi.fn(async () => [
|
||||
{ pid: 42, ppid: 1, name: 'node.exe', command: 'daemon', creationTimeMs: 7 }
|
||||
])
|
||||
|
||||
await queryWindowsProcess(42, { runCommand })
|
||||
|
||||
const script = runCommand.mock.calls[0]?.[1].at(-1) ?? ''
|
||||
expect(script).toContain("$ErrorActionPreference = 'Stop'")
|
||||
expect(script).toMatch(/catch \{[^}]*query_failed/)
|
||||
await expect(queryWindowsProcess(42, { runCommand, readProcessTable })).resolves.toEqual({
|
||||
status: 'present',
|
||||
commandLine: 'daemon',
|
||||
startedAtMs: 7
|
||||
})
|
||||
expect(runCommand).not.toHaveBeenCalled()
|
||||
})
|
||||
|
||||
it('keeps a failed CIM query indeterminate instead of proving the process gone', async () => {
|
||||
const runCommand = vi.fn(async () => '{"status":"query_failed"}')
|
||||
it('keeps an unreadable process table indeterminate instead of proving the process gone', async () => {
|
||||
const readProcessTable = vi.fn(async () => {
|
||||
throw new Error('windows process table is unreadable')
|
||||
})
|
||||
|
||||
await expect(queryWindowsProcess(42, { runCommand })).resolves.toEqual({
|
||||
await expect(queryWindowsProcess(42, { readProcessTable })).resolves.toEqual({
|
||||
status: 'unavailable'
|
||||
})
|
||||
})
|
||||
|
||||
it('never reads a probe result without a success marker as proof of absence', async () => {
|
||||
const runCommand = vi.fn(async () => '{"exists":false}')
|
||||
it('reports absence only from a table that was read and lacks the PID', async () => {
|
||||
const readProcessTable = vi.fn(async () => [
|
||||
{ pid: 7, ppid: 1, name: 'other.exe', command: '' }
|
||||
])
|
||||
|
||||
await expect(queryWindowsProcess(42, { runCommand })).resolves.toEqual({
|
||||
status: 'unavailable'
|
||||
await expect(queryWindowsProcess(42, { readProcessTable })).resolves.toEqual({
|
||||
status: 'missing'
|
||||
})
|
||||
})
|
||||
|
||||
it('reports absence only from a CIM query that ran and found nothing', async () => {
|
||||
const runCommand = vi.fn(async () => '{"status":"missing"}')
|
||||
|
||||
await expect(queryWindowsProcess(42, { runCommand })).resolves.toEqual({ status: 'missing' })
|
||||
})
|
||||
|
||||
it('reads the macOS start time through an async spawn', async () => {
|
||||
const runCommand = vi.fn(async () => 'Sat Jan 1 00:00:00 2028\n')
|
||||
|
||||
@@ -158,14 +157,14 @@ describe('daemon process inspection', () => {
|
||||
})
|
||||
|
||||
it.each([0, -1, 1.5, Number.MAX_SAFE_INTEGER + 1, Number.NaN])(
|
||||
'rejects unsafe Windows pid %s before command interpolation',
|
||||
'rejects unsafe Windows pid %s before reading the process table',
|
||||
async (pid) => {
|
||||
const runCommand = vi.fn()
|
||||
const readProcessTable = vi.fn(async () => [])
|
||||
|
||||
await expect(queryWindowsProcess(pid, { runCommand })).resolves.toEqual({
|
||||
await expect(queryWindowsProcess(pid, { readProcessTable })).resolves.toEqual({
|
||||
status: 'unavailable'
|
||||
})
|
||||
expect(runCommand).not.toHaveBeenCalled()
|
||||
expect(readProcessTable).not.toHaveBeenCalled()
|
||||
}
|
||||
)
|
||||
})
|
||||
|
||||
@@ -8,6 +8,11 @@ import type {
|
||||
ProcessSignalEvidence,
|
||||
WindowsProcessEvidence
|
||||
} from './daemon-incarnation-evidence-types'
|
||||
import {
|
||||
readWindowsProcessCreationTime,
|
||||
readWindowsProcessTableFresh,
|
||||
type WindowsProcessRow
|
||||
} from '../windows/windows-process-table'
|
||||
|
||||
const execFileAsync = promisify(execFile)
|
||||
|
||||
@@ -16,6 +21,7 @@ type InspectionCommandRunner = (file: string, args: string[], timeoutMs: number)
|
||||
export type DaemonProcessInspectionDependencies = {
|
||||
readTextFile?: (path: string) => Promise<string>
|
||||
runCommand?: InspectionCommandRunner
|
||||
readProcessTable?: () => Promise<WindowsProcessRow[]>
|
||||
}
|
||||
|
||||
export function inspectProcessSignal(pid: number): ProcessSignalEvidence {
|
||||
@@ -33,6 +39,12 @@ export function inspectProcessSignal(pid: number): ProcessSignalEvidence {
|
||||
}
|
||||
}
|
||||
|
||||
/** EPERM counts as alive: it proves some process holds the PID. */
|
||||
export function isProcessAlive(pid: number): boolean {
|
||||
const signal = inspectProcessSignal(pid)
|
||||
return signal === 'occupied' || signal === 'permission_denied'
|
||||
}
|
||||
|
||||
export function inspectProcessLiveness(pid: number): ProcessLivenessVerdict {
|
||||
const signal = inspectProcessSignal(pid)
|
||||
switch (signal) {
|
||||
@@ -94,9 +106,7 @@ export async function readProcessCommandLine(
|
||||
}
|
||||
}
|
||||
|
||||
// Why: Get-CimInstance errors (Winmgmt down, corrupt WMI repository, access denied) are
|
||||
// non-terminating and exit 0 with an empty $p, which is indistinguishable from "no such
|
||||
// process" — so the script reports query failure explicitly instead of asserting absence.
|
||||
// Only a table that was read and lacks the PID proves absence; the reader rejects a truncated one.
|
||||
export async function queryWindowsProcess(
|
||||
pid: number,
|
||||
dependencies: DaemonProcessInspectionDependencies = {}
|
||||
@@ -104,45 +114,21 @@ export async function queryWindowsProcess(
|
||||
if (!Number.isSafeInteger(pid) || pid <= 0) {
|
||||
return { status: 'unavailable' }
|
||||
}
|
||||
const runCommand = dependencies.runCommand ?? runInspectionCommand
|
||||
let rows: WindowsProcessRow[]
|
||||
try {
|
||||
const stdout = await runCommand(
|
||||
'powershell.exe',
|
||||
[
|
||||
'-NoProfile',
|
||||
'-NonInteractive',
|
||||
'-Command',
|
||||
`$ErrorActionPreference = 'Stop'; ` +
|
||||
`try { $p = Get-CimInstance Win32_Process -Filter "ProcessId = ${pid}" } ` +
|
||||
`catch { @{ status = 'query_failed' } | ConvertTo-Json -Compress; exit 0 }; ` +
|
||||
`if (!$p) { @{ status = 'missing' } | ConvertTo-Json -Compress; exit 0 }; ` +
|
||||
`$start = $null; if ($p.CreationDate) { ` +
|
||||
`$start = [long]([DateTimeOffset]$p.CreationDate).ToUnixTimeMilliseconds() }; ` +
|
||||
`@{ status = 'present'; cmd = $p.CommandLine; start = $start } | ConvertTo-Json -Compress`
|
||||
],
|
||||
3_000
|
||||
)
|
||||
const parsed = JSON.parse(stdout.trim()) as {
|
||||
status?: unknown
|
||||
cmd?: unknown
|
||||
start?: unknown
|
||||
}
|
||||
// Only a query that ran and found nothing proves absence; anything else stays indeterminate.
|
||||
if (parsed.status === 'missing') {
|
||||
return { status: 'missing' }
|
||||
}
|
||||
if (parsed.status !== 'present') {
|
||||
return { status: 'unavailable' }
|
||||
}
|
||||
return {
|
||||
status: 'present',
|
||||
commandLine: typeof parsed.cmd === 'string' && parsed.cmd ? parsed.cmd : null,
|
||||
startedAtMs:
|
||||
typeof parsed.start === 'number' && Number.isFinite(parsed.start) ? parsed.start : null
|
||||
}
|
||||
rows = await (dependencies.readProcessTable ?? readWindowsProcessTableFresh)()
|
||||
} catch {
|
||||
return { status: 'unavailable' }
|
||||
}
|
||||
const row = rows.find((candidate) => candidate.pid === pid)
|
||||
if (!row) {
|
||||
return { status: 'missing' }
|
||||
}
|
||||
return {
|
||||
status: 'present',
|
||||
commandLine: row.command || null,
|
||||
startedAtMs: row.creationTimeMs ?? readWindowsProcessCreationTime(pid)
|
||||
}
|
||||
}
|
||||
|
||||
// Why: the sync procfs helper in daemon-process-start-time spawns getconf per call; CLK_TCK is fixed for
|
||||
@@ -219,8 +205,6 @@ async function runInspectionCommand(
|
||||
args: string[],
|
||||
timeoutMs: number
|
||||
): Promise<string> {
|
||||
// powershell.exe is console-subsystem: without this it flashes a conhost and
|
||||
// steals foreground on every inspection (#10488).
|
||||
const { stdout } = await execFileAsync(file, args, {
|
||||
encoding: 'utf8',
|
||||
timeout: timeoutMs,
|
||||
@@ -229,6 +213,6 @@ async function runInspectionCommand(
|
||||
return stdout
|
||||
}
|
||||
|
||||
function hasErrorCode(error: unknown, code: string): boolean {
|
||||
export function hasErrorCode(error: unknown, code: string): boolean {
|
||||
return typeof error === 'object' && error !== null && 'code' in error && error.code === code
|
||||
}
|
||||
|
||||
@@ -0,0 +1,54 @@
|
||||
import { afterEach, expect, it, vi } from 'vitest'
|
||||
|
||||
vi.mock('../ipc/pty', () => ({ setLocalPtyProvider: vi.fn() }))
|
||||
|
||||
import { DaemonPtyRouter } from './daemon-pty-router'
|
||||
import { createAdapter } from './daemon-pty-router-test-fixture'
|
||||
import {
|
||||
disconnectDaemon,
|
||||
listLiveDaemonSessions,
|
||||
listLiveDaemonSessionsWithProtocol,
|
||||
replaceDaemonProvider,
|
||||
requestIdleDaemonRetirement
|
||||
} from './daemon-provider-state'
|
||||
import { PROTOCOL_VERSION } from './types'
|
||||
|
||||
afterEach(async () => {
|
||||
await disconnectDaemon()
|
||||
})
|
||||
|
||||
it('reads a census without an installed daemon as unverifiable, never as empty', async () => {
|
||||
await expect(listLiveDaemonSessions()).resolves.toBeNull()
|
||||
await expect(requestIdleDaemonRetirement()).resolves.toEqual({ state: 'unverifiable' })
|
||||
})
|
||||
|
||||
it('reads a census with an unanswered generation as unverifiable', async () => {
|
||||
const current = createAdapter('current', ['live-1'], undefined, PROTOCOL_VERSION)
|
||||
const legacy = createAdapter('legacy', [], undefined, PROTOCOL_VERSION)
|
||||
vi.mocked(legacy.listSessions).mockRejectedValue(new Error('daemon unreachable'))
|
||||
replaceDaemonProvider(new DaemonPtyRouter({ current, legacy: [legacy] }))
|
||||
|
||||
await expect(listLiveDaemonSessions()).resolves.toBeNull()
|
||||
})
|
||||
|
||||
it('labels each live session with the protocol of the generation that owns it', async () => {
|
||||
const current = createAdapter('current', ['live-1'], undefined, PROTOCOL_VERSION)
|
||||
const legacy = createAdapter('legacy', ['live-2'], undefined, PROTOCOL_VERSION - 1)
|
||||
replaceDaemonProvider(new DaemonPtyRouter({ current, legacy: [legacy] }))
|
||||
|
||||
await expect(listLiveDaemonSessionsWithProtocol()).resolves.toEqual([
|
||||
{ sessionId: 'live-1', isAlive: true, protocolVersion: PROTOCOL_VERSION },
|
||||
{ sessionId: 'live-2', isAlive: true, protocolVersion: PROTOCOL_VERSION - 1 }
|
||||
])
|
||||
})
|
||||
|
||||
it('lists every generation when each one answers', async () => {
|
||||
const current = createAdapter('current', ['live-1'], undefined, PROTOCOL_VERSION)
|
||||
const legacy = createAdapter('legacy', ['live-2'], undefined, PROTOCOL_VERSION)
|
||||
replaceDaemonProvider(new DaemonPtyRouter({ current, legacy: [legacy] }))
|
||||
|
||||
await expect(listLiveDaemonSessions()).resolves.toEqual([
|
||||
{ sessionId: 'live-1', isAlive: true },
|
||||
{ sessionId: 'live-2', isAlive: true }
|
||||
])
|
||||
})
|
||||
@@ -19,7 +19,8 @@ import {
|
||||
} from './daemon-launch-paths'
|
||||
import {
|
||||
attributeNextDaemonReplacement,
|
||||
createOutOfProcessLauncher
|
||||
createOutOfProcessLauncher,
|
||||
type DaemonLaunchPolicy
|
||||
} from './daemon-out-of-process-launcher'
|
||||
import type { DaemonProvider } from './daemon-provider-routing'
|
||||
import { installDaemonProvider } from './daemon-provider-state'
|
||||
@@ -46,7 +47,7 @@ function logDaemonMilestone(event: string, details: Record<string, unknown> = {}
|
||||
|
||||
export async function initDaemonPtyProvider(
|
||||
signal?: AbortSignal,
|
||||
options: { macosLoginSessionWatch?: boolean } = {}
|
||||
options: DaemonLaunchPolicy = {}
|
||||
): Promise<void> {
|
||||
logDaemonMilestone('daemon-init-start')
|
||||
// Why: e2e coverage for the startup PTY gate (#5232) needs a daemon init that deterministically outlasts the first-window timeout.
|
||||
@@ -58,7 +59,7 @@ export async function initDaemonPtyProvider(
|
||||
|
||||
const newSpawner = new DaemonSpawner({
|
||||
runtimeDir,
|
||||
launcher: createOutOfProcessLauncher(runtimeDir, options.macosLoginSessionWatch ?? false)
|
||||
launcher: createOutOfProcessLauncher(runtimeDir, options)
|
||||
})
|
||||
|
||||
// Why: assign the module-level spawner/adapter only after both succeed, so a failed ensureRunning() leaves no stale spawner.
|
||||
|
||||
@@ -11,12 +11,16 @@ import {
|
||||
getMacDaemonTccAttributionHealth,
|
||||
type MacDaemonTccAttributionHealth
|
||||
} from './daemon-tcc-attribution'
|
||||
import { PROTOCOL_VERSION } from './types'
|
||||
import { PROTOCOL_VERSION, type DaemonSessionInfo, type SessionInfo } from './types'
|
||||
import type { DaemonIdleRetirementResult } from './daemon-pty-runtime-state'
|
||||
|
||||
let spawner: DaemonSpawner | null = null
|
||||
let adapter: DaemonProvider | null = null
|
||||
|
||||
export function installDaemonProvider(newSpawner: DaemonSpawner, newAdapter: DaemonProvider): void {
|
||||
export function installDaemonProvider(
|
||||
newSpawner: DaemonSpawner | null,
|
||||
newAdapter: DaemonProvider
|
||||
): void {
|
||||
spawner = newSpawner
|
||||
replaceDaemonProvider(newAdapter)
|
||||
}
|
||||
@@ -95,26 +99,6 @@ export async function getCurrentDaemonMacTccAttributionHealth(): Promise<MacDaem
|
||||
)
|
||||
}
|
||||
|
||||
/** Returns null unless every daemon generation supplied an authoritative inventory. */
|
||||
export async function listLiveDaemonPtyIds(): Promise<string[] | null> {
|
||||
if (!adapter) {
|
||||
return null
|
||||
}
|
||||
const adapters =
|
||||
adapter instanceof DaemonPtyRouter || adapter instanceof DegradedDaemonPtyProvider
|
||||
? adapter.getAllAdapters()
|
||||
: [adapter]
|
||||
const inventories = await Promise.allSettled(
|
||||
adapters.map((daemonAdapter) => daemonAdapter.listProcesses())
|
||||
)
|
||||
if (inventories.some((inventory) => inventory.status === 'rejected')) {
|
||||
return null
|
||||
}
|
||||
return inventories.flatMap((inventory) =>
|
||||
inventory.status === 'fulfilled' ? inventory.value.map((process) => process.id) : []
|
||||
)
|
||||
}
|
||||
|
||||
// Why: keep the module-level adapter and ipc/pty.ts's localProvider in sync so app-quit can't dispose a stale reference.
|
||||
export function replaceDaemonProvider(newAdapter: DaemonProvider): void {
|
||||
adapter = newAdapter
|
||||
@@ -135,3 +119,82 @@ export async function shutdownDaemon(): Promise<void> {
|
||||
await spawner?.shutdown()
|
||||
spawner = null
|
||||
}
|
||||
|
||||
/** Returns null unless every daemon generation supplied an authoritative inventory. */
|
||||
export async function listLiveDaemonPtyIds(): Promise<string[] | null> {
|
||||
if (!adapter) {
|
||||
return null
|
||||
}
|
||||
const adapters =
|
||||
adapter instanceof DaemonPtyRouter || adapter instanceof DegradedDaemonPtyProvider
|
||||
? adapter.getAllAdapters()
|
||||
: [adapter]
|
||||
const inventories = await Promise.allSettled(
|
||||
adapters.map((daemonAdapter) => daemonAdapter.listProcesses())
|
||||
)
|
||||
if (inventories.some((inventory) => inventory.status === 'rejected')) {
|
||||
return null
|
||||
}
|
||||
return inventories.flatMap((inventory) =>
|
||||
inventory.status === 'fulfilled' ? inventory.value.map((process) => process.id) : []
|
||||
)
|
||||
}
|
||||
|
||||
/** Returns null unless every daemon generation supplied an authoritative session inventory. */
|
||||
export async function listLiveDaemonSessions(): Promise<SessionInfo[] | null> {
|
||||
const sessions = await listLiveDaemonSessionsWithProtocol()
|
||||
return sessions?.map(({ protocolVersion: _protocolVersion, ...session }) => session) ?? null
|
||||
}
|
||||
|
||||
/** Like listLiveDaemonSessions, with the protocol of the daemon generation owning each session. */
|
||||
export async function listLiveDaemonSessionsWithProtocol(): Promise<DaemonSessionInfo[] | null> {
|
||||
if (!adapter) {
|
||||
return null
|
||||
}
|
||||
const adapters =
|
||||
adapter instanceof DaemonPtyRouter || adapter instanceof DegradedDaemonPtyProvider
|
||||
? adapter.getAllAdapters()
|
||||
: [adapter]
|
||||
const inventories = await Promise.allSettled(
|
||||
adapters.map(async (daemonAdapter) =>
|
||||
(await daemonAdapter.listSessions()).map((session) => ({
|
||||
...session,
|
||||
protocolVersion: daemonAdapter.protocolVersion
|
||||
}))
|
||||
)
|
||||
)
|
||||
if (inventories.some((inventory) => inventory.status === 'rejected')) {
|
||||
return null
|
||||
}
|
||||
return inventories.flatMap((inventory) =>
|
||||
inventory.status === 'fulfilled' ? inventory.value : []
|
||||
)
|
||||
}
|
||||
|
||||
/** Terminals the degraded provider ran in-process; none outside degraded mode. */
|
||||
export async function countInProcessFallbackTerminals(): Promise<number> {
|
||||
return adapter instanceof DegradedDaemonPtyProvider
|
||||
? (await adapter.fallback.listProcesses()).length
|
||||
: 0
|
||||
}
|
||||
|
||||
/** Atomically fence new daemon terminals and retire only an idle, single-generation daemon. */
|
||||
export async function requestIdleDaemonRetirement(): Promise<DaemonIdleRetirementResult> {
|
||||
if (!adapter) {
|
||||
return { state: 'unverifiable' }
|
||||
}
|
||||
if (adapter instanceof DegradedDaemonPtyProvider) {
|
||||
return { state: 'unverifiable' }
|
||||
}
|
||||
if (adapter instanceof DaemonPtyRouter) {
|
||||
return adapter.requestIdleRetirement()
|
||||
}
|
||||
return adapter.requestIdleRetirement()
|
||||
}
|
||||
|
||||
/** Reopens terminal admission when an idle-retirement attempt did not retire the daemon. */
|
||||
export function releaseDaemonRetirementFence(): void {
|
||||
if (adapter && !(adapter instanceof DegradedDaemonPtyProvider)) {
|
||||
adapter.releaseIdleRetirementFence()
|
||||
}
|
||||
}
|
||||
|
||||
@@ -3,7 +3,12 @@ import { removeDaemonListener } from './daemon-listener-registry'
|
||||
import { emitPtyListeners } from './daemon-pty-listener-emission'
|
||||
import type { PtyIncarnationId } from '../../shared/pty-incarnation'
|
||||
import { DaemonPtySessionInventory } from './daemon-pty-session-inventory'
|
||||
import { CLEAN_DISCONNECT_PROTOCOL_VERSION } from './types'
|
||||
import {
|
||||
CLEAN_DISCONNECT_PROTOCOL_VERSION,
|
||||
type ListSessionsResult,
|
||||
type ShutdownIfIdleResult
|
||||
} from './types'
|
||||
import type { DaemonIdleRetirementResult } from './daemon-pty-runtime-state'
|
||||
import type { PtyBackgroundStreamEvent } from '../providers/types'
|
||||
|
||||
export abstract class DaemonPtyEventSubscriptions extends DaemonPtySessionInventory {
|
||||
@@ -85,6 +90,86 @@ export abstract class DaemonPtyEventSubscriptions extends DaemonPtySessionInvent
|
||||
this.recordAuthenticatedIdentity()
|
||||
}
|
||||
|
||||
async requestIdleRetirement(): Promise<DaemonIdleRetirementResult> {
|
||||
if (this.protocolVersion < CLEAN_DISCONNECT_PROTOCOL_VERSION) {
|
||||
return { state: 'unsupported' }
|
||||
}
|
||||
if (this.idleRetirementState === 'retiring') {
|
||||
return { state: 'retiring' }
|
||||
}
|
||||
if (this.idleRetirementPromise) {
|
||||
return this.idleRetirementPromise
|
||||
}
|
||||
if (
|
||||
this.disconnectOnlyPromise ||
|
||||
(this.respawnAdoptionClosed && this.idleRetirementState === 'open')
|
||||
) {
|
||||
return { state: 'unverifiable' }
|
||||
}
|
||||
this.idleRetirementAdmissionClosed = true
|
||||
this.respawnAdoptionClosed = true
|
||||
this.idleRetirementState = 'checking'
|
||||
const request = this.finishIdleRetirementRequest().finally(() => {
|
||||
if (this.idleRetirementPromise === request) {
|
||||
this.idleRetirementPromise = null
|
||||
}
|
||||
})
|
||||
this.idleRetirementPromise = request
|
||||
return request
|
||||
}
|
||||
|
||||
private async finishIdleRetirementRequest(): Promise<DaemonIdleRetirementResult> {
|
||||
try {
|
||||
await this.client.ensureConnected()
|
||||
} catch {
|
||||
// Nothing was asked of the daemon, so nothing can be retiring.
|
||||
this.reopenAfterRefusedIdleRetirement()
|
||||
return { state: 'unverifiable' }
|
||||
}
|
||||
try {
|
||||
const result = await this.client.request<ShutdownIfIdleResult>('shutdownIfIdle', undefined)
|
||||
if (result.retiring) {
|
||||
this.idleRetirementState = 'retiring'
|
||||
return { state: 'retiring' }
|
||||
}
|
||||
let liveSessions: number | null = null
|
||||
try {
|
||||
const inventory = await this.client.request<ListSessionsResult>('listSessions', undefined)
|
||||
liveSessions = inventory.sessions.filter((session) => session.isAlive).length
|
||||
} catch {
|
||||
liveSessions = null
|
||||
}
|
||||
this.reopenAfterRefusedIdleRetirement()
|
||||
return {
|
||||
state: 'busy',
|
||||
liveSessions,
|
||||
admissionReopened: true
|
||||
}
|
||||
} catch {
|
||||
// The daemon may have accepted before contact was lost; keep admission and respawn fenced.
|
||||
this.idleRetirementState = 'unverifiable'
|
||||
return { state: 'unverifiable' }
|
||||
}
|
||||
}
|
||||
|
||||
/** Reopens admission an idle-retirement attempt fenced without retiring the daemon. */
|
||||
releaseIdleRetirementFence(): void {
|
||||
// An unverifiable attempt may have been accepted, so it stays fenced like a retiring one.
|
||||
if (
|
||||
this.idleRetirementState !== 'retiring' &&
|
||||
this.idleRetirementState !== 'unverifiable' &&
|
||||
!this.idleRetirementPromise
|
||||
) {
|
||||
this.reopenAfterRefusedIdleRetirement()
|
||||
}
|
||||
}
|
||||
|
||||
private reopenAfterRefusedIdleRetirement(): void {
|
||||
this.idleRetirementState = 'open'
|
||||
this.idleRetirementAdmissionClosed = false
|
||||
this.respawnAdoptionClosed = false
|
||||
}
|
||||
|
||||
// Why: unlike dispose(), leave history files unclean (no endedAt) so the next launch treats them as crash-recoverable,
|
||||
// but still write a final checkpoint so a daemon crash while Orca is closed has recovery data.
|
||||
async disconnectOnly(): Promise<void> {
|
||||
|
||||
@@ -0,0 +1,168 @@
|
||||
import { vi } from 'vitest'
|
||||
import { settledWriteStub } from '../providers/settled-pty-write-stub'
|
||||
import type { DaemonPtyAdapter } from './daemon-pty-adapter'
|
||||
import type { PtyBackgroundStreamEvent, PtySpawnOptions, PtySpawnResult } from '../providers/types'
|
||||
import {
|
||||
AGENT_SESSION_CLAIM_DAEMON_PROTOCOL_VERSION,
|
||||
AGENT_SESSION_CREATE_OPERATION_DAEMON_PROTOCOL_VERSION,
|
||||
GIT_CREDENTIAL_GUARD_HOST_PROTOCOL_VERSION
|
||||
} from './types'
|
||||
import { SNAPSHOT_SERIALIZER_FIDELITY_DAEMON_PROTOCOL_VERSION } from './daemon-protocol-version'
|
||||
|
||||
type AdapterMock = DaemonPtyAdapter & {
|
||||
emitData: (id: string, data: string, sequenceChars?: number) => void
|
||||
emitBackground: (event: PtyBackgroundStreamEvent) => void
|
||||
emitExit: (id: string, code: number, incarnationId?: string) => void
|
||||
emitIdentityChange: () => void
|
||||
triggerWriteUnavailable: (id: string) => void
|
||||
}
|
||||
|
||||
export function createAdapter(
|
||||
label: string,
|
||||
sessions: string[] = [],
|
||||
reconcileResult?: { alive: string[]; killed: string[] },
|
||||
protocolVersion = GIT_CREDENTIAL_GUARD_HOST_PROTOCOL_VERSION
|
||||
): AdapterMock {
|
||||
const writes: { id: string; data: string }[] = []
|
||||
const dataListeners: ((payload: { id: string; data: string; sequenceChars?: number }) => void)[] =
|
||||
[]
|
||||
const backgroundListeners: ((payload: PtyBackgroundStreamEvent) => void)[] = []
|
||||
const writeUnavailableListeners: ((payload: { id: string }) => void)[] = []
|
||||
const exitListeners: ((payload: { id: string; code: number; incarnationId?: string }) => void)[] =
|
||||
[]
|
||||
const identityChangeListeners: (() => void)[] = []
|
||||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the router calls only the adapter members this mock defines.
|
||||
return {
|
||||
protocolVersion,
|
||||
supportsGitCredentialGuardHost: () =>
|
||||
protocolVersion >= GIT_CREDENTIAL_GUARD_HOST_PROTOCOL_VERSION,
|
||||
supportsAgentSessionClaims: () =>
|
||||
protocolVersion >= AGENT_SESSION_CLAIM_DAEMON_PROTOCOL_VERSION,
|
||||
supportsAgentSessionCreateOperations: () =>
|
||||
protocolVersion >= AGENT_SESSION_CREATE_OPERATION_DAEMON_PROTOCOL_VERSION,
|
||||
providesAgentSessionOwnerListings: () =>
|
||||
protocolVersion >= AGENT_SESSION_CLAIM_DAEMON_PROTOCOL_VERSION,
|
||||
canProvideAuthoritativeBufferSnapshot: () =>
|
||||
protocolVersion >= SNAPSHOT_SERIALIZER_FIDELITY_DAEMON_PROTOCOL_VERSION,
|
||||
spawn: vi.fn(async (opts: PtySpawnOptions): Promise<PtySpawnResult> => {
|
||||
const id = opts.sessionId ?? `${label}-new`
|
||||
sessions.push(id)
|
||||
return { id }
|
||||
}),
|
||||
listProcesses: vi.fn(async () =>
|
||||
sessions.map((id) => ({
|
||||
id,
|
||||
cwd: '',
|
||||
title: label
|
||||
}))
|
||||
),
|
||||
listSessions: vi.fn(async () => sessions.map((sessionId) => ({ sessionId, isAlive: true }))),
|
||||
requestIdleRetirement: vi.fn(async () => ({ state: 'retiring' as const })),
|
||||
releaseIdleRetirementFence: vi.fn(),
|
||||
hasPty: vi.fn((id: string) => sessions.includes(id)),
|
||||
probePtyLiveness: vi.fn(async (id: string) => sessions.includes(id)),
|
||||
write: vi.fn((id: string, data: string) => {
|
||||
writes.push({ id, data })
|
||||
}),
|
||||
writeWithSettlement: vi.fn(settledWriteStub()),
|
||||
resize: vi.fn(),
|
||||
setPtyBackgrounded: vi.fn(),
|
||||
getBufferSnapshot: vi.fn(async () => null),
|
||||
shutdown: vi.fn(async (id: string) => {
|
||||
const idx = sessions.indexOf(id)
|
||||
if (idx !== -1) {
|
||||
sessions.splice(idx, 1)
|
||||
}
|
||||
}),
|
||||
attach: vi.fn(async () => {}),
|
||||
sendSignal: vi.fn(async () => {}),
|
||||
getCwd: vi.fn(async () => ''),
|
||||
getInitialCwd: vi.fn(async () => ''),
|
||||
clearBuffer: vi.fn(async () => {}),
|
||||
acknowledgeDataEvent: vi.fn(),
|
||||
hasChildProcesses: vi.fn(async () => false),
|
||||
getForegroundProcess: vi.fn(async () => null),
|
||||
inspectProcess: vi.fn(async () => ({ foregroundProcess: null, hasChildProcesses: false })),
|
||||
confirmForegroundProcess: vi.fn(async () => `${label}-confirmed`),
|
||||
serialize: vi.fn(async () => '{}'),
|
||||
revive: vi.fn(async () => {}),
|
||||
getDefaultShell: vi.fn(async () => '/bin/zsh'),
|
||||
getProfiles: vi.fn(async () => []),
|
||||
onData: vi.fn(
|
||||
(callback: (payload: { id: string; data: string; sequenceChars?: number }) => void) => {
|
||||
dataListeners.push(callback)
|
||||
return () => {
|
||||
const idx = dataListeners.indexOf(callback)
|
||||
if (idx !== -1) {
|
||||
dataListeners.splice(idx, 1)
|
||||
}
|
||||
}
|
||||
}
|
||||
),
|
||||
onBackgroundStreamEvent: vi.fn((callback: (payload: PtyBackgroundStreamEvent) => void) => {
|
||||
backgroundListeners.push(callback)
|
||||
return () => {
|
||||
const idx = backgroundListeners.indexOf(callback)
|
||||
if (idx !== -1) {
|
||||
backgroundListeners.splice(idx, 1)
|
||||
}
|
||||
}
|
||||
}),
|
||||
onWriteUnavailable: vi.fn((callback: (payload: { id: string }) => void) => {
|
||||
writeUnavailableListeners.push(callback)
|
||||
return () => {
|
||||
const idx = writeUnavailableListeners.indexOf(callback)
|
||||
if (idx !== -1) {
|
||||
writeUnavailableListeners.splice(idx, 1)
|
||||
}
|
||||
}
|
||||
}),
|
||||
onExit: vi.fn(
|
||||
(callback: (payload: { id: string; code: number; incarnationId?: string }) => void) => {
|
||||
exitListeners.push(callback)
|
||||
return () => {
|
||||
const idx = exitListeners.indexOf(callback)
|
||||
if (idx !== -1) {
|
||||
exitListeners.splice(idx, 1)
|
||||
}
|
||||
}
|
||||
}
|
||||
),
|
||||
onDaemonIdentityChanged: vi.fn((callback: () => void) => {
|
||||
identityChangeListeners.push(callback)
|
||||
return () => {
|
||||
const idx = identityChangeListeners.indexOf(callback)
|
||||
if (idx !== -1) {
|
||||
identityChangeListeners.splice(idx, 1)
|
||||
}
|
||||
}
|
||||
}),
|
||||
ackColdRestore: vi.fn(),
|
||||
clearTombstone: vi.fn(),
|
||||
reconcileOnStartup: vi.fn(async () => reconcileResult ?? { alive: sessions, killed: [] }),
|
||||
dispose: vi.fn(),
|
||||
disconnectOnly: vi.fn(async () => {}),
|
||||
emitData: (id: string, data: string, sequenceChars?: number) => {
|
||||
for (const listener of dataListeners) {
|
||||
listener({ id, data, ...(sequenceChars === undefined ? {} : { sequenceChars }) })
|
||||
}
|
||||
},
|
||||
emitBackground: (event: PtyBackgroundStreamEvent) => {
|
||||
for (const listener of backgroundListeners) {
|
||||
listener(event)
|
||||
}
|
||||
},
|
||||
emitExit: (id: string, code: number, incarnationId?: string) => {
|
||||
for (const listener of exitListeners) {
|
||||
listener({ id, code, ...(incarnationId ? { incarnationId } : {}) })
|
||||
}
|
||||
},
|
||||
emitIdentityChange: () => identityChangeListeners.forEach((listener) => listener()),
|
||||
triggerWriteUnavailable: (id: string) => {
|
||||
for (const listener of writeUnavailableListeners) {
|
||||
listener({ id })
|
||||
}
|
||||
},
|
||||
_writes: writes
|
||||
} as unknown as AdapterMock
|
||||
}
|
||||
@@ -1,29 +1,20 @@
|
||||
import { createAdapter } from './daemon-pty-router-test-fixture'
|
||||
import { describe, expect, it, vi } from 'vitest'
|
||||
import { DaemonPtyRouter } from './daemon-pty-router'
|
||||
import { stubWriteSettlement } from '../providers/settled-pty-write-stub'
|
||||
import { SessionNotFoundError, TerminalSessionOwnerUnverifiedError } from './daemon-errors'
|
||||
import type { DaemonPtyAdapter } from './daemon-pty-adapter'
|
||||
import { settledWriteStub, stubWriteSettlement } from '../providers/settled-pty-write-stub'
|
||||
import type { PtyBackgroundStreamEvent, PtySpawnOptions, PtySpawnResult } from '../providers/types'
|
||||
import type { PtySpawnResult } from '../providers/types'
|
||||
import {
|
||||
AGENT_SESSION_CLAIM_DAEMON_PROTOCOL_VERSION,
|
||||
AGENT_SESSION_CREATE_OPERATION_DAEMON_PROTOCOL_VERSION,
|
||||
GIT_CREDENTIAL_GUARD_HOST_PROTOCOL_VERSION
|
||||
AGENT_SESSION_CREATE_OPERATION_DAEMON_PROTOCOL_VERSION
|
||||
} from './types'
|
||||
import {
|
||||
HISTORY_SEED_TRANSFER_PROTOCOL_VERSION,
|
||||
PROTOCOL_VERSION,
|
||||
SNAPSHOT_SERIALIZER_FIDELITY_DAEMON_PROTOCOL_VERSION,
|
||||
STABLE_PANE_ATTACH_ONLY_DAEMON_PROTOCOL_VERSION
|
||||
} from './daemon-protocol-version'
|
||||
|
||||
type AdapterMock = DaemonPtyAdapter & {
|
||||
emitData: (id: string, data: string, sequenceChars?: number) => void
|
||||
emitBackground: (event: PtyBackgroundStreamEvent) => void
|
||||
emitExit: (id: string, code: number, incarnationId?: string) => void
|
||||
emitIdentityChange: () => void
|
||||
triggerWriteUnavailable: (id: string) => void
|
||||
}
|
||||
|
||||
const LARGE_RECONCILE_SESSION_COUNT = 150_000
|
||||
|
||||
function buildSessionIds(prefix: string, count: number): string[] {
|
||||
@@ -34,152 +25,6 @@ function buildSessionIds(prefix: string, count: number): string[] {
|
||||
return ids
|
||||
}
|
||||
|
||||
function createAdapter(
|
||||
label: string,
|
||||
sessions: string[] = [],
|
||||
reconcileResult?: { alive: string[]; killed: string[] },
|
||||
protocolVersion = GIT_CREDENTIAL_GUARD_HOST_PROTOCOL_VERSION
|
||||
): AdapterMock {
|
||||
const writes: { id: string; data: string }[] = []
|
||||
const dataListeners: ((payload: { id: string; data: string; sequenceChars?: number }) => void)[] =
|
||||
[]
|
||||
const backgroundListeners: ((payload: PtyBackgroundStreamEvent) => void)[] = []
|
||||
const writeUnavailableListeners: ((payload: { id: string }) => void)[] = []
|
||||
const exitListeners: ((payload: { id: string; code: number; incarnationId?: string }) => void)[] =
|
||||
[]
|
||||
const identityChangeListeners: (() => void)[] = []
|
||||
return {
|
||||
protocolVersion,
|
||||
supportsGitCredentialGuardHost: () =>
|
||||
protocolVersion >= GIT_CREDENTIAL_GUARD_HOST_PROTOCOL_VERSION,
|
||||
supportsAgentSessionClaims: () =>
|
||||
protocolVersion >= AGENT_SESSION_CLAIM_DAEMON_PROTOCOL_VERSION,
|
||||
supportsAgentSessionCreateOperations: () =>
|
||||
protocolVersion >= AGENT_SESSION_CREATE_OPERATION_DAEMON_PROTOCOL_VERSION,
|
||||
providesAgentSessionOwnerListings: () =>
|
||||
protocolVersion >= AGENT_SESSION_CLAIM_DAEMON_PROTOCOL_VERSION,
|
||||
canProvideAuthoritativeBufferSnapshot: () =>
|
||||
protocolVersion >= SNAPSHOT_SERIALIZER_FIDELITY_DAEMON_PROTOCOL_VERSION,
|
||||
spawn: vi.fn(async (opts: PtySpawnOptions): Promise<PtySpawnResult> => {
|
||||
const id = opts.sessionId ?? `${label}-new`
|
||||
sessions.push(id)
|
||||
return { id }
|
||||
}),
|
||||
listProcesses: vi.fn(async () =>
|
||||
sessions.map((id) => ({
|
||||
id,
|
||||
cwd: '',
|
||||
title: label
|
||||
}))
|
||||
),
|
||||
hasPty: vi.fn((id: string) => sessions.includes(id)),
|
||||
probePtyLiveness: vi.fn(async (id: string) => sessions.includes(id)),
|
||||
write: vi.fn((id: string, data: string) => {
|
||||
writes.push({ id, data })
|
||||
}),
|
||||
writeWithSettlement: vi.fn(settledWriteStub()),
|
||||
resize: vi.fn(),
|
||||
setPtyBackgrounded: vi.fn(),
|
||||
getBufferSnapshot: vi.fn(async () => null),
|
||||
shutdown: vi.fn(async (id: string) => {
|
||||
const idx = sessions.indexOf(id)
|
||||
if (idx !== -1) {
|
||||
sessions.splice(idx, 1)
|
||||
}
|
||||
}),
|
||||
attach: vi.fn(async () => {}),
|
||||
sendSignal: vi.fn(async () => {}),
|
||||
getCwd: vi.fn(async () => ''),
|
||||
getInitialCwd: vi.fn(async () => ''),
|
||||
clearBuffer: vi.fn(async () => {}),
|
||||
acknowledgeDataEvent: vi.fn(),
|
||||
hasChildProcesses: vi.fn(async () => false),
|
||||
getForegroundProcess: vi.fn(async () => null),
|
||||
inspectProcess: vi.fn(async () => ({ foregroundProcess: null, hasChildProcesses: false })),
|
||||
confirmForegroundProcess: vi.fn(async () => `${label}-confirmed`),
|
||||
serialize: vi.fn(async () => '{}'),
|
||||
revive: vi.fn(async () => {}),
|
||||
getDefaultShell: vi.fn(async () => '/bin/zsh'),
|
||||
getProfiles: vi.fn(async () => []),
|
||||
onData: vi.fn(
|
||||
(callback: (payload: { id: string; data: string; sequenceChars?: number }) => void) => {
|
||||
dataListeners.push(callback)
|
||||
return () => {
|
||||
const idx = dataListeners.indexOf(callback)
|
||||
if (idx !== -1) {
|
||||
dataListeners.splice(idx, 1)
|
||||
}
|
||||
}
|
||||
}
|
||||
),
|
||||
onBackgroundStreamEvent: vi.fn((callback: (payload: PtyBackgroundStreamEvent) => void) => {
|
||||
backgroundListeners.push(callback)
|
||||
return () => {
|
||||
const idx = backgroundListeners.indexOf(callback)
|
||||
if (idx !== -1) {
|
||||
backgroundListeners.splice(idx, 1)
|
||||
}
|
||||
}
|
||||
}),
|
||||
onWriteUnavailable: vi.fn((callback: (payload: { id: string }) => void) => {
|
||||
writeUnavailableListeners.push(callback)
|
||||
return () => {
|
||||
const idx = writeUnavailableListeners.indexOf(callback)
|
||||
if (idx !== -1) {
|
||||
writeUnavailableListeners.splice(idx, 1)
|
||||
}
|
||||
}
|
||||
}),
|
||||
onExit: vi.fn(
|
||||
(callback: (payload: { id: string; code: number; incarnationId?: string }) => void) => {
|
||||
exitListeners.push(callback)
|
||||
return () => {
|
||||
const idx = exitListeners.indexOf(callback)
|
||||
if (idx !== -1) {
|
||||
exitListeners.splice(idx, 1)
|
||||
}
|
||||
}
|
||||
}
|
||||
),
|
||||
onDaemonIdentityChanged: vi.fn((callback: () => void) => {
|
||||
identityChangeListeners.push(callback)
|
||||
return () => {
|
||||
const idx = identityChangeListeners.indexOf(callback)
|
||||
if (idx !== -1) {
|
||||
identityChangeListeners.splice(idx, 1)
|
||||
}
|
||||
}
|
||||
}),
|
||||
ackColdRestore: vi.fn(),
|
||||
clearTombstone: vi.fn(),
|
||||
reconcileOnStartup: vi.fn(async () => reconcileResult ?? { alive: sessions, killed: [] }),
|
||||
dispose: vi.fn(),
|
||||
disconnectOnly: vi.fn(async () => {}),
|
||||
emitData: (id: string, data: string, sequenceChars?: number) => {
|
||||
for (const listener of dataListeners) {
|
||||
listener({ id, data, ...(sequenceChars === undefined ? {} : { sequenceChars }) })
|
||||
}
|
||||
},
|
||||
emitBackground: (event: PtyBackgroundStreamEvent) => {
|
||||
for (const listener of backgroundListeners) {
|
||||
listener(event)
|
||||
}
|
||||
},
|
||||
emitExit: (id: string, code: number, incarnationId?: string) => {
|
||||
for (const listener of exitListeners) {
|
||||
listener({ id, code, ...(incarnationId ? { incarnationId } : {}) })
|
||||
}
|
||||
},
|
||||
emitIdentityChange: () => identityChangeListeners.forEach((listener) => listener()),
|
||||
triggerWriteUnavailable: (id: string) => {
|
||||
for (const listener of writeUnavailableListeners) {
|
||||
listener({ id })
|
||||
}
|
||||
},
|
||||
_writes: writes
|
||||
} as unknown as AdapterMock
|
||||
}
|
||||
|
||||
it('forwards dead-endpoint write-unavailable signals from every routed adapter', () => {
|
||||
// Why revert-sensitive: main subscribes on the ROUTED provider, so if the router
|
||||
// does not forward this the STA-2373 fan-out never reaches the renderer and only
|
||||
@@ -247,6 +92,83 @@ it('forwards the owning legacy daemon sequence from attach', async () => {
|
||||
})
|
||||
|
||||
describe('DaemonPtyRouter', () => {
|
||||
describe('idle retirement', () => {
|
||||
it('retires every empty daemon generation and fences subsequent spawns', async () => {
|
||||
const current = createAdapter('current', [], undefined, PROTOCOL_VERSION)
|
||||
const legacy = createAdapter('legacy', [], undefined, PROTOCOL_VERSION)
|
||||
const router = new DaemonPtyRouter({ current, legacy: [legacy] })
|
||||
|
||||
await expect(router.requestIdleRetirement()).resolves.toEqual({ state: 'retiring' })
|
||||
expect(current.requestIdleRetirement).toHaveBeenCalledOnce()
|
||||
expect(legacy.requestIdleRetirement).toHaveBeenCalledOnce()
|
||||
await expect(router.spawn({ sessionId: 'late', cols: 80, rows: 24 })).rejects.toThrow(
|
||||
'Terminal daemon is decommissioning'
|
||||
)
|
||||
})
|
||||
|
||||
it('reports live inventory before retiring any generation and reopens admission', async () => {
|
||||
const current = createAdapter('current', [], undefined, PROTOCOL_VERSION)
|
||||
const legacy = createAdapter('legacy', ['legacy-live'], undefined, PROTOCOL_VERSION)
|
||||
const router = new DaemonPtyRouter({ current, legacy: [legacy] })
|
||||
|
||||
await expect(router.requestIdleRetirement()).resolves.toEqual({
|
||||
state: 'busy',
|
||||
liveSessions: 1,
|
||||
admissionReopened: true
|
||||
})
|
||||
expect(current.requestIdleRetirement).not.toHaveBeenCalled()
|
||||
expect(legacy.requestIdleRetirement).not.toHaveBeenCalled()
|
||||
await expect(
|
||||
router.spawn({ sessionId: 'after-refusal', cols: 80, rows: 24 })
|
||||
).resolves.toEqual({
|
||||
id: 'after-refusal'
|
||||
})
|
||||
})
|
||||
|
||||
it('does not partially retire when a generation predates clean idle shutdown', async () => {
|
||||
const current = createAdapter('current', [], undefined, PROTOCOL_VERSION)
|
||||
const legacy = createAdapter('legacy', [], undefined, 23)
|
||||
const router = new DaemonPtyRouter({ current, legacy: [legacy] })
|
||||
|
||||
await expect(router.requestIdleRetirement()).resolves.toEqual({ state: 'unsupported' })
|
||||
expect(current.requestIdleRetirement).not.toHaveBeenCalled()
|
||||
expect(legacy.requestIdleRetirement).not.toHaveBeenCalled()
|
||||
})
|
||||
|
||||
it('keeps admission fenced after a partial multi-generation retirement', async () => {
|
||||
const current = createAdapter('current', [], undefined, PROTOCOL_VERSION)
|
||||
const legacy = createAdapter('legacy', [], undefined, PROTOCOL_VERSION)
|
||||
vi.mocked(legacy.requestIdleRetirement).mockResolvedValueOnce({
|
||||
state: 'busy',
|
||||
liveSessions: 0
|
||||
})
|
||||
const router = new DaemonPtyRouter({ current, legacy: [legacy] })
|
||||
|
||||
await expect(router.requestIdleRetirement()).resolves.toEqual({ state: 'unverifiable' })
|
||||
await expect(router.spawn({ sessionId: 'unsafe', cols: 80, rows: 24 })).rejects.toThrow(
|
||||
'Terminal daemon is decommissioning'
|
||||
)
|
||||
})
|
||||
|
||||
it('does not certify reopened admission when another generation retired beside live sessions', async () => {
|
||||
const current = createAdapter('current', [], undefined, PROTOCOL_VERSION)
|
||||
const legacy = createAdapter('legacy', [], undefined, PROTOCOL_VERSION)
|
||||
vi.mocked(legacy.requestIdleRetirement).mockResolvedValueOnce({
|
||||
state: 'busy',
|
||||
liveSessions: 1,
|
||||
admissionReopened: true
|
||||
})
|
||||
const router = new DaemonPtyRouter({ current, legacy: [legacy] })
|
||||
await expect(router.requestIdleRetirement()).resolves.toEqual({
|
||||
state: 'busy',
|
||||
liveSessions: 1
|
||||
})
|
||||
await expect(
|
||||
router.spawn({ sessionId: 'unsafe-partial', cols: 80, rows: 24 })
|
||||
).rejects.toThrow('Terminal daemon is decommissioning')
|
||||
})
|
||||
})
|
||||
|
||||
it('reports separate conservative resume and fresh-create boundaries', () => {
|
||||
const current = createAdapter(
|
||||
'current',
|
||||
|
||||
@@ -1,3 +1,4 @@
|
||||
import { reconcileDaemonRouterSessions } from './daemon-router-session-reconciliation'
|
||||
import type { DaemonPtyAdapter } from './daemon-pty-adapter'
|
||||
import { DaemonPtyAdapterSubscriptionFanout } from './daemon-pty-adapter-subscription-fanout'
|
||||
import type {
|
||||
@@ -12,6 +13,8 @@ import type { PtyProcessInspection } from '../providers/pty-process-inspection'
|
||||
import { shouldHandoffDaemonHistory } from './daemon-history-handoff'
|
||||
import type { DaemonPtyRouterDataEvent, DaemonPtyRouterExitEvent } from './daemon-pty-router-events'
|
||||
import { DaemonSessionOwnerResolver } from './daemon-session-owner-resolution'
|
||||
import type { DaemonIdleRetirementResult } from './daemon-pty-runtime-state'
|
||||
import { DaemonRouterRetirement } from './daemon-router-retirement'
|
||||
import type { WriteSettlement } from '../../shared/pty-write-settlement'
|
||||
import type { TerminalOscColorQueryReplyColors } from '../../shared/terminal-osc-color-reply'
|
||||
|
||||
@@ -21,6 +24,7 @@ export class DaemonPtyRouter implements IPtyProvider {
|
||||
private sessionAdapters = new Map<string, DaemonPtyAdapter>()
|
||||
private readonly ownerResolver: DaemonSessionOwnerResolver<DaemonPtyAdapter>
|
||||
private readonly subscriptions: DaemonPtyAdapterSubscriptionFanout
|
||||
private readonly retirement = new DaemonRouterRetirement(() => this.allAdapters())
|
||||
|
||||
constructor(opts: { current: DaemonPtyAdapter; legacy: DaemonPtyAdapter[] }) {
|
||||
this.current = opts.current
|
||||
@@ -40,17 +44,34 @@ export class DaemonPtyRouter implements IPtyProvider {
|
||||
}
|
||||
|
||||
async spawn(opts: PtySpawnOptions): Promise<PtySpawnResult> {
|
||||
if (opts.attachOnly && opts.sessionId) {
|
||||
return await this.ownerResolver.spawnAttachOnly({ ...opts, sessionId: opts.sessionId })
|
||||
if (this.retirement.admissionClosed) {
|
||||
throw new Error('Terminal daemon is decommissioning')
|
||||
}
|
||||
const adapter = opts.sessionId ? this.sessionAdapters.get(opts.sessionId) : undefined
|
||||
const target = adapter ?? this.current
|
||||
const result = await target.spawn(opts)
|
||||
// Why: the adapter filters intentional recovery exits and canonical-ID races before publishing proof.
|
||||
if (!result.exitedBeforeSpawnReply) {
|
||||
this.ownerResolver.recordRoute(result.id, target, result.incarnationId)
|
||||
// Why counted: an idle-retirement census must not race a spawn it cannot yet see.
|
||||
this.retirement.spawnInFlight++
|
||||
try {
|
||||
if (opts.attachOnly && opts.sessionId) {
|
||||
return await this.ownerResolver.spawnAttachOnly({ ...opts, sessionId: opts.sessionId })
|
||||
}
|
||||
const adapter = opts.sessionId ? this.sessionAdapters.get(opts.sessionId) : undefined
|
||||
const target = adapter ?? this.current
|
||||
const result = await target.spawn(opts)
|
||||
// Why: the adapter filters intentional recovery exits and canonical-ID races before publishing proof.
|
||||
if (!result.exitedBeforeSpawnReply) {
|
||||
this.ownerResolver.recordRoute(result.id, target, result.incarnationId)
|
||||
}
|
||||
return result
|
||||
} finally {
|
||||
this.retirement.spawnInFlight--
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
requestIdleRetirement(): Promise<DaemonIdleRetirementResult> {
|
||||
return this.retirement.requestIdleRetirement()
|
||||
}
|
||||
|
||||
releaseIdleRetirementFence(): void {
|
||||
this.retirement.releaseFence()
|
||||
}
|
||||
|
||||
supportsGitCredentialGuardHost(sessionId?: string): boolean {
|
||||
@@ -257,38 +278,10 @@ export class DaemonPtyRouter implements IPtyProvider {
|
||||
this.adapterFor(sessionId).clearTombstone(sessionId)
|
||||
}
|
||||
|
||||
async reconcileOnStartup(validWorktreeIds: Set<string>): Promise<{
|
||||
alive: string[]
|
||||
killed: string[]
|
||||
}> {
|
||||
const alive: string[] = []
|
||||
const killed: string[] = []
|
||||
const aliveProviders = new Map<string, Set<DaemonPtyAdapter>>()
|
||||
for (const adapter of this.allAdapters()) {
|
||||
const result = await adapter.reconcileOnStartup(validWorktreeIds)
|
||||
// Why: daemon startup can reconcile many restored sessions; spreading
|
||||
// those arrays into push can exceed JavaScript's argument limit.
|
||||
for (const id of result.alive) {
|
||||
alive.push(id)
|
||||
}
|
||||
for (const id of result.killed) {
|
||||
killed.push(id)
|
||||
}
|
||||
for (const id of result.alive) {
|
||||
const providers = aliveProviders.get(id) ?? new Set<DaemonPtyAdapter>()
|
||||
providers.add(adapter)
|
||||
aliveProviders.set(id, providers)
|
||||
}
|
||||
}
|
||||
for (const id of new Set([...alive, ...killed])) {
|
||||
const providers = aliveProviders.get(id)
|
||||
if (providers?.size === 1) {
|
||||
this.ownerResolver.recordRoute(id, providers.values().next().value!)
|
||||
} else {
|
||||
this.ownerResolver.forgetRoute(id)
|
||||
}
|
||||
}
|
||||
return { alive, killed }
|
||||
async reconcileOnStartup(
|
||||
validWorktreeIds: Set<string>
|
||||
): Promise<{ alive: string[]; killed: string[] }> {
|
||||
return reconcileDaemonRouterSessions(this.allAdapters(), this.ownerResolver, validWorktreeIds)
|
||||
}
|
||||
|
||||
dispose(): void {
|
||||
|
||||
@@ -71,6 +71,12 @@ export type DaemonIdentityChangeEvent = {
|
||||
current: DaemonEndpointIdentity
|
||||
}
|
||||
|
||||
export type DaemonIdleRetirementResult =
|
||||
| { state: 'retiring' }
|
||||
| { state: 'busy'; liveSessions: number | null; admissionReopened?: true }
|
||||
| { state: 'unsupported' }
|
||||
| { state: 'unverifiable' }
|
||||
|
||||
export abstract class DaemonPtyRuntimeState {
|
||||
readonly protocolVersion: number
|
||||
protected socketPath: string
|
||||
@@ -92,6 +98,9 @@ export abstract class DaemonPtyRuntimeState {
|
||||
protected packagedAppVersion: string | null
|
||||
protected pendingRespawnAdoptionRelease: (() => void) | null = null
|
||||
protected respawnAdoptionClosed = false
|
||||
protected idleRetirementAdmissionClosed = false
|
||||
protected idleRetirementState: 'open' | 'checking' | 'retiring' | 'unverifiable' = 'open'
|
||||
protected idleRetirementPromise: Promise<DaemonIdleRetirementResult> | null = null
|
||||
protected respawnPromise: Promise<void> | null = null
|
||||
protected staleBundleReplacementPromise: Promise<void> | null = null
|
||||
protected writeRecoveryPromise: Promise<void> | null = null
|
||||
|
||||
@@ -25,7 +25,15 @@ import { injectHistoryEnv, injectWslFishHistoryEnv, logHistoryInjection } from '
|
||||
import { addWslEnvKeys } from '../wsl-env'
|
||||
|
||||
export abstract class DaemonPtySessionSpawn extends DaemonPtySpawnResult {
|
||||
// Checked again in doSpawn: retirement can close admission while spawn awaits.
|
||||
private assertSpawnAdmission(): void {
|
||||
if (this.idleRetirementAdmissionClosed) {
|
||||
throw new Error('Terminal daemon is decommissioning')
|
||||
}
|
||||
}
|
||||
|
||||
async spawn(opts: PtySpawnOptions): Promise<PtySpawnResult> {
|
||||
this.assertSpawnAdmission()
|
||||
const spawnOpts = this.withHistoryIsolation(opts)
|
||||
const sessionId = spawnOpts.sessionId ?? mintPtySessionId(spawnOpts.worktreeId)
|
||||
const operation: PendingDaemonSpawnOperation = {
|
||||
@@ -105,6 +113,7 @@ export abstract class DaemonPtySessionSpawn extends DaemonPtySpawnResult {
|
||||
operation: PendingDaemonSpawnOperation,
|
||||
historyRecovery: HistoryRecoveryContext
|
||||
): Promise<PtySpawnResult> {
|
||||
this.assertSpawnAdmission()
|
||||
if (
|
||||
opts.agentSessionEnsure &&
|
||||
this.protocolVersion < AGENT_SESSION_CLAIM_DAEMON_PROTOCOL_VERSION
|
||||
|
||||
@@ -10,6 +10,7 @@ import type { DaemonSessionBackgroundRouting } from './daemon-session-background
|
||||
import { recordDaemonStreamBacklogEvent } from './daemon-stream-backlog-probe'
|
||||
import type { DaemonStreamDataBatcher } from './daemon-stream-data-batcher'
|
||||
import type { DaemonTerminalAdmission } from './daemon-terminal-admission'
|
||||
import { ptySpawnHealthPlatformCoverage } from './daemon-health-identity'
|
||||
import type { TerminalHistorySeedTransferRegistry } from './terminal-history-seed-transfer-registry'
|
||||
import type { TerminalHost } from './terminal-host'
|
||||
import { SessionNotFoundError, type DaemonRequest } from './types'
|
||||
@@ -156,7 +157,7 @@ export class DaemonRequestRouter {
|
||||
return { health: await readCurrentProcessMacSystemResolverHealth() }
|
||||
case 'ptySpawnHealth':
|
||||
await this.options.ptySpawnHealthCheck()
|
||||
return { healthy: true }
|
||||
return { healthy: true, coverage: ptySpawnHealthPlatformCoverage() }
|
||||
case 'shutdown':
|
||||
return this.shutdown(clientId, request.id, request.payload.killSessions)
|
||||
}
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
import { expect, it, vi } from 'vitest'
|
||||
import { DaemonRouterRetirement } from './daemon-router-retirement'
|
||||
import { createAdapter } from './daemon-pty-router-test-fixture'
|
||||
import { PROTOCOL_VERSION } from './types'
|
||||
|
||||
it.each(['inventory', 'protocol', 'spawn', 'live'] as const)(
|
||||
'does not reopen admission on a %s retry after partial retirement',
|
||||
async (failure) => {
|
||||
const current = createAdapter('current', [], undefined, PROTOCOL_VERSION)
|
||||
const legacy = createAdapter('legacy', [], undefined, PROTOCOL_VERSION)
|
||||
let adapters = [current, legacy]
|
||||
const retirement = new DaemonRouterRetirement(() => adapters)
|
||||
vi.mocked(legacy.requestIdleRetirement).mockResolvedValueOnce({
|
||||
state: 'busy',
|
||||
liveSessions: 0
|
||||
})
|
||||
await expect(retirement.requestIdleRetirement()).resolves.toEqual({ state: 'unverifiable' })
|
||||
expect(retirement.admissionClosed).toBe(true)
|
||||
if (failure === 'inventory') {
|
||||
vi.mocked(current.listSessions).mockRejectedValueOnce(new Error('lost connection'))
|
||||
} else if (failure === 'protocol') {
|
||||
adapters = [createAdapter('old', [], undefined, 23)]
|
||||
} else if (failure === 'spawn') {
|
||||
retirement.spawnInFlight = 1
|
||||
} else {
|
||||
adapters = [createAdapter('live', ['existing'], undefined, PROTOCOL_VERSION)]
|
||||
}
|
||||
expect(await retirement.requestIdleRetirement()).not.toHaveProperty('admissionReopened')
|
||||
expect(retirement.admissionClosed).toBe(true)
|
||||
}
|
||||
)
|
||||
|
||||
it('does not reopen when every native result is busy without reopening proof', async () => {
|
||||
const current = createAdapter('current', [], undefined, PROTOCOL_VERSION)
|
||||
vi.mocked(current.requestIdleRetirement).mockResolvedValue({ state: 'busy', liveSessions: 0 })
|
||||
const retirement = new DaemonRouterRetirement(() => [current])
|
||||
await expect(retirement.requestIdleRetirement()).resolves.toEqual({ state: 'unverifiable' })
|
||||
expect(retirement.admissionClosed).toBe(true)
|
||||
})
|
||||
|
||||
it('keeps the fence through a lost native reply and failed retry inventory', async () => {
|
||||
const current = createAdapter('current', [], undefined, PROTOCOL_VERSION)
|
||||
vi.mocked(current.requestIdleRetirement).mockRejectedValueOnce(new Error('lost stop reply'))
|
||||
const retirement = new DaemonRouterRetirement(() => [current])
|
||||
await expect(retirement.requestIdleRetirement()).rejects.toThrow('lost stop reply')
|
||||
vi.mocked(current.listSessions).mockRejectedValueOnce(new Error('lost connection'))
|
||||
await expect(retirement.requestIdleRetirement()).resolves.toEqual({ state: 'unverifiable' })
|
||||
expect(retirement.admissionClosed).toBe(true)
|
||||
})
|
||||
|
||||
it('releases a fence left by an incomplete retirement, but never one that retired', async () => {
|
||||
const current = createAdapter('current', [], undefined, PROTOCOL_VERSION)
|
||||
vi.mocked(current.requestIdleRetirement).mockResolvedValueOnce({ state: 'busy', liveSessions: 0 })
|
||||
const retirement = new DaemonRouterRetirement(() => [current])
|
||||
await expect(retirement.requestIdleRetirement()).resolves.toEqual({ state: 'unverifiable' })
|
||||
expect(retirement.admissionClosed).toBe(true)
|
||||
retirement.releaseFence()
|
||||
expect(retirement.admissionClosed).toBe(false)
|
||||
expect(current.releaseIdleRetirementFence).toHaveBeenCalledOnce()
|
||||
|
||||
await expect(retirement.requestIdleRetirement()).resolves.toEqual({ state: 'retiring' })
|
||||
retirement.releaseFence()
|
||||
expect(retirement.admissionClosed).toBe(true)
|
||||
})
|
||||
@@ -0,0 +1,82 @@
|
||||
import type { DaemonPtyAdapter } from './daemon-pty-adapter'
|
||||
import type { DaemonIdleRetirementResult } from './daemon-pty-runtime-state'
|
||||
import { CLEAN_DISCONNECT_PROTOCOL_VERSION } from './types'
|
||||
|
||||
export class DaemonRouterRetirement {
|
||||
admissionClosed = false
|
||||
spawnInFlight = 0
|
||||
private retirementAttempted = false
|
||||
private retired = false
|
||||
private idleRetirementPromise: Promise<DaemonIdleRetirementResult> | null = null
|
||||
|
||||
constructor(private readonly allAdapters: () => DaemonPtyAdapter[]) {}
|
||||
|
||||
/** Reopens a fence left by an attempt that did not retire every generation. */
|
||||
releaseFence(): void {
|
||||
if (this.idleRetirementPromise || this.retired) {
|
||||
return
|
||||
}
|
||||
this.admissionClosed = false
|
||||
for (const adapter of this.allAdapters()) {
|
||||
adapter.releaseIdleRetirementFence()
|
||||
}
|
||||
}
|
||||
|
||||
async requestIdleRetirement(): Promise<DaemonIdleRetirementResult> {
|
||||
if (this.idleRetirementPromise) {
|
||||
return this.idleRetirementPromise
|
||||
}
|
||||
this.admissionClosed = true
|
||||
const request = this.finishIdleRetirementRequest().finally(() => {
|
||||
if (this.idleRetirementPromise === request) {
|
||||
this.idleRetirementPromise = null
|
||||
}
|
||||
})
|
||||
this.idleRetirementPromise = request
|
||||
return request
|
||||
}
|
||||
|
||||
private async finishIdleRetirementRequest(): Promise<DaemonIdleRetirementResult> {
|
||||
const adapters = this.allAdapters()
|
||||
if (this.spawnInFlight > 0) {
|
||||
this.admissionClosed = this.retirementAttempted
|
||||
return { state: 'busy', liveSessions: null }
|
||||
}
|
||||
if (adapters.some((adapter) => adapter.protocolVersion < CLEAN_DISCONNECT_PROTOCOL_VERSION)) {
|
||||
this.admissionClosed = this.retirementAttempted
|
||||
return { state: 'unsupported' }
|
||||
}
|
||||
const inventories = await Promise.allSettled(adapters.map((adapter) => adapter.listSessions()))
|
||||
if (inventories.some((inventory) => inventory.status === 'rejected')) {
|
||||
this.admissionClosed = this.retirementAttempted
|
||||
return { state: 'unverifiable' }
|
||||
}
|
||||
// Each adapter's inventory lists live sessions only.
|
||||
const liveSessions = inventories.reduce(
|
||||
(count, inventory) => count + (inventory.status === 'fulfilled' ? inventory.value.length : 0),
|
||||
0
|
||||
)
|
||||
if (liveSessions > 0) {
|
||||
this.admissionClosed = this.retirementAttempted
|
||||
return {
|
||||
state: 'busy',
|
||||
liveSessions,
|
||||
...(!this.retirementAttempted ? { admissionReopened: true as const } : {})
|
||||
}
|
||||
}
|
||||
this.retirementAttempted = true
|
||||
const results = await Promise.all(adapters.map((adapter) => adapter.requestIdleRetirement()))
|
||||
if (results.every((result) => result.state === 'retiring')) {
|
||||
this.retired = true
|
||||
return { state: 'retiring' }
|
||||
}
|
||||
const refusedLiveSessions = results.reduce(
|
||||
(count, result) => count + (result.state === 'busy' ? (result.liveSessions ?? 0) : 0),
|
||||
0
|
||||
)
|
||||
if (refusedLiveSessions > 0) {
|
||||
return { state: 'busy', liveSessions: refusedLiveSessions }
|
||||
}
|
||||
return { state: 'unverifiable' }
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,37 @@
|
||||
import type { DaemonPtyAdapter } from './daemon-pty-adapter'
|
||||
import type { DaemonSessionOwnerResolver } from './daemon-session-owner-resolution'
|
||||
|
||||
export async function reconcileDaemonRouterSessions(
|
||||
adapters: readonly DaemonPtyAdapter[],
|
||||
ownerResolver: DaemonSessionOwnerResolver<DaemonPtyAdapter>,
|
||||
validWorktreeIds: Set<string>
|
||||
): Promise<{ alive: string[]; killed: string[] }> {
|
||||
const alive: string[] = []
|
||||
const killed: string[] = []
|
||||
const aliveProviders = new Map<string, Set<DaemonPtyAdapter>>()
|
||||
for (const adapter of adapters) {
|
||||
const result = await adapter.reconcileOnStartup(validWorktreeIds)
|
||||
// Why: daemon startup can reconcile many restored sessions; spreading
|
||||
// those arrays into push can exceed JavaScript's argument limit.
|
||||
for (const id of result.alive) {
|
||||
alive.push(id)
|
||||
}
|
||||
for (const id of result.killed) {
|
||||
killed.push(id)
|
||||
}
|
||||
for (const id of result.alive) {
|
||||
const providers = aliveProviders.get(id) ?? new Set<DaemonPtyAdapter>()
|
||||
providers.add(adapter)
|
||||
aliveProviders.set(id, providers)
|
||||
}
|
||||
}
|
||||
for (const id of new Set([...alive, ...killed])) {
|
||||
const providers = aliveProviders.get(id)
|
||||
if (providers?.size === 1) {
|
||||
ownerResolver.recordRoute(id, providers.values().next().value!)
|
||||
} else {
|
||||
ownerResolver.forgetRoute(id)
|
||||
}
|
||||
}
|
||||
return { alive, killed }
|
||||
}
|
||||
@@ -685,3 +685,17 @@ describe('DegradedDaemonPtyProvider', () => {
|
||||
expect(fallback.listProcesses).toHaveBeenCalledTimes(3)
|
||||
})
|
||||
})
|
||||
|
||||
it('lists the in-process terminals a census must count, apart from the daemon inventory', async () => {
|
||||
const current = createDaemonAdapter('current')
|
||||
const fallback = createProvider('fallback')
|
||||
const provider = new DegradedDaemonPtyProvider({ current, legacy: [], fallback })
|
||||
|
||||
const fresh = await provider.spawn({ cols: 80, rows: 24 })
|
||||
|
||||
expect(fallback.spawn).toHaveBeenCalledOnce()
|
||||
expect(await provider.fallback.listProcesses()).toEqual([
|
||||
expect.objectContaining({ id: fresh.id })
|
||||
])
|
||||
expect(provider.getAllAdapters()).toEqual([current])
|
||||
})
|
||||
|
||||
@@ -26,7 +26,8 @@ export class DegradedDaemonPtyProvider implements IPtyProvider {
|
||||
|
||||
private current: DaemonPtyAdapter
|
||||
private legacy: DaemonPtyAdapter[]
|
||||
private fallback: IPtyProvider
|
||||
/** Runs terminals in this process, so they die with it. */
|
||||
readonly fallback: IPtyProvider
|
||||
private sessionProviders = new Map<string, IPtyProvider>()
|
||||
private freshSpawns: DegradedDaemonFreshSpawnRouter
|
||||
private ownerRecovery: DegradedDaemonOwnerRecovery
|
||||
|
||||
@@ -12,7 +12,7 @@ import { resolveSafePtyDefaultCwd } from '../../providers/pty-default-cwd'
|
||||
import { TerminalAttachCanceledError } from '../daemon-errors'
|
||||
import { DaemonProtocolError } from '../types'
|
||||
|
||||
const PTY_SPAWN_HEALTH_TIMEOUT_MS = 4_000
|
||||
export const PTY_SPAWN_HEALTH_TIMEOUT_MS = 4_000
|
||||
|
||||
async function loadNodePty(): Promise<typeof pty> {
|
||||
return import('node-pty')
|
||||
@@ -148,6 +148,13 @@ export async function preflightPtySpawn(args: {
|
||||
}
|
||||
}
|
||||
|
||||
export class PtySpawnHealthTimeoutError extends Error {
|
||||
constructor(timeoutMs: number) {
|
||||
super(`PTY spawn health check timed out after ${timeoutMs}ms`)
|
||||
this.name = 'PtySpawnHealthTimeoutError'
|
||||
}
|
||||
}
|
||||
|
||||
export function formatPtySpawnError(err: unknown, shellPath: string, spawnCwd: string): Error {
|
||||
const message = err instanceof Error ? err.message : String(err)
|
||||
const formatted = new DaemonProtocolError(
|
||||
@@ -159,7 +166,9 @@ export function formatPtySpawnError(err: unknown, shellPath: string, spawnCwd: s
|
||||
return formatted
|
||||
}
|
||||
|
||||
export async function runPtySpawnHealthProbe(): Promise<void> {
|
||||
export async function runPtySpawnHealthProbe(
|
||||
timeoutMs = PTY_SPAWN_HEALTH_TIMEOUT_MS
|
||||
): Promise<void> {
|
||||
const cwd = isExistingDirectory(process.env.ORCA_USER_DATA_PATH)
|
||||
? process.env.ORCA_USER_DATA_PATH
|
||||
: resolveSafePtyDefaultCwd()
|
||||
@@ -214,10 +223,8 @@ export async function runPtySpawnHealthProbe(): Promise<void> {
|
||||
}
|
||||
}
|
||||
const timer = setTimeout(() => {
|
||||
finish(new Error(`PTY spawn health check timed out after ${PTY_SPAWN_HEALTH_TIMEOUT_MS}ms`), {
|
||||
kill: true
|
||||
})
|
||||
}, PTY_SPAWN_HEALTH_TIMEOUT_MS)
|
||||
finish(new PtySpawnHealthTimeoutError(timeoutMs), { kill: true })
|
||||
}, timeoutMs)
|
||||
exitDisposable = proc.onExit(({ exitCode }) => {
|
||||
if (exitCode === 0) {
|
||||
finish()
|
||||
|
||||
@@ -29,7 +29,8 @@ async function syncDirectory(directory: string): Promise<void> {
|
||||
}
|
||||
}
|
||||
|
||||
function syncDirectorySync(directory: string): void {
|
||||
/** Sync variant of the best-effort directory fsync, for callers that publish by rename or link. */
|
||||
export function syncDirectoryDurablySync(directory: string): void {
|
||||
let fd: number | null = null
|
||||
try {
|
||||
fd = openSync(directory, 'r')
|
||||
@@ -50,7 +51,7 @@ function syncDirectorySync(directory: string): void {
|
||||
/** Rename an already-fsynced file and make the containing directory durable. */
|
||||
export function renameDurableSync(tmpPath: string, finalPath: string): void {
|
||||
renameFileWithWindowsRetry(tmpPath, finalPath)
|
||||
syncDirectorySync(dirname(finalPath))
|
||||
syncDirectoryDurablySync(dirname(finalPath))
|
||||
}
|
||||
|
||||
/** Publish an already-fsynced file without replacing a concurrently created destination. */
|
||||
@@ -58,7 +59,7 @@ export function publishFileDurableSync(tmpPath: string, finalPath: string): bool
|
||||
if (!publishFileWithoutOverwrite(tmpPath, finalPath)) {
|
||||
return false
|
||||
}
|
||||
syncDirectorySync(dirname(finalPath))
|
||||
syncDirectoryDurablySync(dirname(finalPath))
|
||||
rmSync(tmpPath)
|
||||
return true
|
||||
}
|
||||
@@ -209,12 +210,14 @@ export async function removeStaleDurableWriteTempFiles(
|
||||
export function writeFileDurableSync(
|
||||
tmpPath: string,
|
||||
finalPath: string,
|
||||
payload: string | Uint8Array
|
||||
payload: string | Uint8Array,
|
||||
/** Creation mode for a new file, e.g. 0o600 for state other users must not read. */
|
||||
mode?: number
|
||||
): void {
|
||||
let renamed = false
|
||||
try {
|
||||
// A Uint8Array payload is written verbatim; a string still defaults to UTF-8.
|
||||
writeFileSync(tmpPath, payload)
|
||||
writeFileSync(tmpPath, payload, mode === undefined ? undefined : { mode })
|
||||
const fd = openSync(tmpPath, 'r+')
|
||||
try {
|
||||
fsyncSync(fd)
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user