mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 00:02:38 +00:00
cd8d03bc06398b856f6db734aef11df4e2afe15a
211
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d9fbb4eecf | Keep SSH typing replies inside narrow split terminal panes (#24682) | ||
|
|
b49abdb1f4 |
fix: recover renderer launch failures in the running app (#24250)
* fix(recovery): back off a launch-failed renderer instead of tripping the crash breaker
A renderer that the OS refused to spawn (macOS exit 1003 = LAUNCH_RESULT_FAILURE; field
cause: per-user process limit, posix_spawn EAGAIN) burned the 3-reload crash-loop budget
in ~750ms and raised a "graphics driver" prompt, while the condition lasted minutes.
- launch-failed retries in place on a 250ms..60s backoff (~2 min), outside the breaker;
a loaded document resets it. Other crash reasons keep the breaker.
- Each launch failure records renderer_launch_failed_probe {spawnError} from a cheap
spawn probe, so bundles name EAGAIN/EACCES/ENOENT directly.
- The exhausted prompt says the process limit was hit (probe EAGAIN), drops the
graphics-driver wording, keeps Try Again as default, and offers no Restart:
app.relaunch also needs a free process slot and silently fails without one.
* fix(recovery): skip the launch probe on Windows and probe the prompt once
- Re-check quitting after the prompt's probe; don't re-probe on Copy Commands.
- recordRendererLaunchFailureProbe never rejects (breadcrumb write guarded).
- Windows: no spawn probe; a child per failed launch is the per-operation burst EDR scores.
* test: cover quitting and duplicate renderer launch failures
* test: use typed access in PTY delay regression fixture
* fix: scope extended launch retries to POSIX hosts
* test: cover launch probe behavior on native Windows
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
|
||
|
|
53930a161b |
Keep SSH typing replies visible during background pressure (#24629)
* test: align source-control fixtures with current store contracts * Bound E2E package setup and retain cancelled-job traces * Remove empty passing sentinels from opt-in socket tests * Make SSH typing pressure fixture readiness and replies observable |
||
|
|
3f37fcc423 | test: align source-control fixtures with current store contracts (#24571) | ||
|
|
ba9af21d75 | test: classify terminal driver input with the current PTY contract (#24560) | ||
|
|
c9a9b8d109 | test: restore delayed PTY writes in large-paste coverage (#24556) | ||
|
|
b666d07117 | test: update worktree setup and enforce cleanup results (#24552) | ||
|
|
1fbfb13e0f | test: isolate seeded Git repositories per Playwright worker (#24550) | ||
|
|
02790c53e8 |
Fix Codex status after Ctrl+C copy and side-chat navigation (#24339)
* fix: preserve Codex status on ambiguous Ctrl+C input * fix: confirm Codex turn cancellations from host rollout records * fix: keep ephemeral side hooks separate from the Codex main turn * perf: watch active Codex rollouts and skip unrelated records * fix: retain confirmed Codex cancellation across late relay events |
||
|
|
8c1670f2c4 |
test(e2e): fake Codex answers the --no-daemon --help probe without a spawn (#24440)
* test(e2e): fake Codex answers the --no-daemon --help probe without a spawn Since #23933 Orca runs `codex --help` before a path-named Codex launch. The fakes logged it as an agent spawn and held the probe for its 5 s timeout. Move the app-server refusal and --help answer into one shared FAKE_CODEX_LAUNCH_PROBES_SOURCE used by every fake Codex. * test(e2e): stop asserting the dispatch capability column a current worker no longer has #23994 stopped minting the per-dispatch capability, so capability_hash is null. The --help spawn failure used to stop this test before it got here. |
||
|
|
9afd1101ff |
fix(orchestration): stop minting and printing the dispatch capability (#23994)
* fix(orchestration): authorize worker reports without the dispatch capability Worker lifecycle reports and questions no longer depend on the per-dispatch capability token that lives only in the agent's conversation. The host now: - ignores capability_hash/capability_revoked_at for authorization on every row and checks the exact worker process instead (ask gains that check); - refuses a report whose calling terminal is provably another orchestration party (a Run coordinator or another Dispatch's worker), treating env that names no live pane here as absent; - applies one worker-state rule locally and remotely: a stop in flight refuses, while stop_unknown and start_unknown accept and settle. Minting and printing the flag are unchanged, so an older host and older preambles keep working. * fix(orchestration): stop minting the dispatch capability Dispatches no longer mint a per-Dispatch token, and preambles, the bundled skill guide and the ask resume hint stop printing --dispatch-capability. The consumer-generation bump and delivery fence that minting carried stay, now as setDispatchConsumer. Readers that inferred meaning from capability_hash read what they meant instead: worker-show's injected stage comes from the attached consumer, and a failed start copies custody identity only when no authority was ever attached. The CLI keeps accepting and forwarding the flag for older hosts. Cancelling a Task is recorded as failed with a reason; task-list now shows that reason and the guide and task-update notes document the recipe. * fix(orchestration): name the fenced party without implying which Dispatch it owns * refactor(orchestration): one worker report rule, fence only a different party - One module owns the unproven/settleable worker states and the refusal rule; local send records it, ask and remote throw it. A stale process is worker_identity_changed on every path. - The caller fence passes the worker's own terminal when its --from handle went stale. - Document the shared-tmux-server limit; drop the dead dispatch_capability_invalid rejection member; tests assert dispatch state, not the capability column. * refactor(orchestration): drop setDispatchConsumer and the dead capability retention - dispatch --inject no longer re-points the row createDispatchContext just wrote; worker-show reports every worker-less Dispatch as context_only, since Orca keeps no record of the paste. Tests re-point through a fixture. - failWorkerStart always records when the lifecycle closed; nothing authorizes on it. - Restore the ask resume hint's echo of a passed --dispatch-capability: an old host checks it before --resume. - Move the cancellation convention to #23983. * test(orchestration): drop capability-era assertions other tests already cover * refactor(orchestration): drop the host-side capability field and no-op test fixtures - RpcRequest and the SSH bridge stop carrying orchestrationCapability; the CLI's wire field stays for older hosts. - Fixtures pass identity to createRootDispatch instead of re-pointing to the same values; drop absence checks for a flag that can no longer be produced. * test(orchestration): cover a current process whose terminal moved to another pane * chore(orchestration): finish the capability cleanup in test stubs and skill wording * test(orchestration): drop needless response casts; mark the db stub cast safe |
||
|
|
c3183a4556 |
test(e2e): let the completed-worker fake Codex answer the --help probe (#24033)
#23900 probes codex --help before each launch; the fake counted it as a worker spawn, breaking two specs. |
||
|
|
21d4ae9448 |
feat(agents): pre-trust the folder wherever Orca starts an agent (#23744)
* feat(claude): pre-trust worktrees Orca creates Claude Code asks "Do you trust this folder?" on first launch in any folder it has not seen, which blocks unattended launches in worktrees Orca itself made. Orca now records where a worktree's content came from when it creates it, and before each Claude launch writes Claude's own folder-trust entry for that worktree's root (never the main checkout) when the new setting is on and the content is the user's repository. Forks, bare commits, folder workspaces and external checkouts keep Claude's prompt. The write takes Claude's lock, never creates or breaks the file, runs on the SSH host itself, and is revoked when the worktree is removed or the setting is turned off. Launches that already pass --dangerously-skip-permissions also skip the trust prompt for that one process only, via CLAUDE_CODE_SANDBOXED=1 on the command. * fix(claude): parse the relay trust request with a schema and ship its search keys * fix(claude): never write a WSL guest's trust into the Windows config A WSL worktree's Claude reads the guest's own config. Two paths still wrote its trust into the Windows host's ~/.claude.json instead: the Claude auth prep's fallback (runtime 'wsl' but the host config dir, when the WSL home cannot be resolved), which wrote a Linux-path key the removal revoke can never delete; and the Agent Teams leader, which passes no auth or distro and wrote a UNC key. Require the guest's own config dir, and treat any WSL worktree path as guest-only. * revert(claude): drop the skip-permissions trust shortcut Pre-trust stays limited to worktrees Orca creates from the user's own repository. The per-launch CLAUDE_CODE_SANDBOXED prefix skipped Claude's trust question in every folder for launches carrying the skip-permissions flag (Orca's default Claude args), including the user's own folders and fork PR worktrees, and the setting could not turn it off. Remove the prefix, the inherited-variable strip that existed only for it, and the Agent Teams leader-to-teammate propagation; restore the tests that pinned the prefixed launch string. * fix(claude): revoke SSH trust in the config file the grant used At spawn the relay resolves Claude's config from the launch env, which carries a CLAUDE_CONFIG_DIR set in Orca's Claude default env. claudeTrust.converge had only the relay's own process env, so removing the worktree or turning the setting off revoked in the default file and the grant outlived the worktree. Send the config-file keys with the request, as the local revoke already uses. * i18n(settings): translate the Claude worktree trust setting * fix(settings): say Claude trust applies when Orca starts Claude The description said Claude skips its trust prompt in any worktree Orca created. Trust is written only when Orca itself starts Claude there, so a `claude` typed by hand in a fresh worktree still asks. Say that, and bring the es/fr/ja/ko/zh translations in line with the new text. * fix(worktrees): treat a base on an Orca-added fork remote as fork content A worktree based on a named ref was always stamped as the repository's own content, so picking the fork remote Orca adds for a pull request (or a local branch tracking it) as the base made a fork's code eligible for Claude trust. At create time, read the repo's `remote.<name>.orca-created` markers and each branch's tracked remote in one `git config` call; a base on such a remote is stamped as a fork's content, and a read failure is not vouched for. Remotes the user added, such as `upstream`, stay first-party. * perf(claude): revoke worktree trust once per config file, not per worktree Turning "Trust worktrees Orca creates for Claude" off read and parsed the whole Claude config once per Orca worktree on the main process. Group the revocations by config file locally and by SSH connection, and make claudeTrust.converge take a batch of requests. * feat(settings): one agent-wide "trust the folder" setting in Settings > Agents Replace the Claude-only worktree trust toggle with a single setting, agentWorkspaceTrustEnabled (on by default; unreleased, so no migration). The row says what it does for every agent: agents Orca starts skip their "trust this folder?" prompt in that worktree or folder, turning it off stops new trust while existing trust stays, and while it is off unattended launches (orchestration workers, automations, the phone) stop at the agent's trust question until someone answers. Translations for es/fr/ja/ko/zh. Also restores the `awaitingUnnamed` chat catalog keys an earlier merge of main dropped from this branch. * feat(agent-trust): pre-trust the workspace for every preset agent at PTY spawn Every Orca-started agent PTY passes through one of the two spawn builders with its declared launchAgent, which survives setup-script wrapping. The builders now call one hook that, for a fresh launch (never a reattach or restored pane) with the setting on, applies the agent's trust preset to the worktree, folder workspace or main checkout it starts in. - One dispatcher, applyAgentWorkspaceTrust(preset, workspacePath, launch context), carries what a writer needs: the final spawn env, the Claude managed-account auth prep, the WSL distro and the SSH connection. - Claude joins the presets on both the claude and claude-agent-teams entries. Its writer stays grant-only in claude-folder-trust-file.ts: the file Claude reads (CLAUDE_CONFIG_DIR / custom-OAuth suffix / legacy .config.json / a WSL guest's own file), Claude's <file>.lock never broken and taken only when a write is due, atomic temp+rename keeping mode and symlinks, never creating the file, NFC + realpath keys. - SSH Claude launches forward the optional claudeFolderTrust spawn field so the relay grants with its own spawn env; old relays ignore it and Claude asks. Other presets keep the SFTP writer. A WSL launch never writes the Windows home: non-Claude presets skip it, Claude writes the guest's file or nothing. - Codex keeps the 20 s deadline its shared config lane needs; every other preset gets 1.5 s. A miss means the agent asks; trust bookkeeping never fails or blocks a launch. Removes the Claude-only machinery this replaces: the eligibility/host/ lifecycle/spawn modules, the persisted creation content-origin field and its classification, revoke-on-removal, the setting-off sweep, the claudeTrust.converge relay method and the Agent Teams leader special case (the leader pane now spawns through the hook with the claude preset). The agent config types move to tui-agent-config-types.ts so the config table stays under the line budget. * refactor(agent-trust): delete the pre-spawn trust writes the spawn hook replaces The spawn hook is now the only owner of agent folder trust, so remove every other writer: - the agentTrust:markTrusted IPC channel, its preload bridge and types, and all renderer callers (agent-trust-preflight and its callers in the background session, work-item direct launch, session continuation, worktree creation, folder workspace composer and session fork); - the main pre-spawn sites: the createdWithAgent preflight in worktree-remote.ts, markLocalWorktreeTrusted/markRemoteWorktreeTrusted and the runtime's markWorkspaceTrustedForAgent family with the markTrusted ports of the runtime create flows; - Codex's own launch-prep and resume-prep trust writes. Each of those launches reaches a spawn builder with launchAgent set, so the hook covers it. This also fixes a live gap: the worktree-remote.ts copy of the preset switch omitted Antigravity, so an agy agent started from a desktop worktree create still asked; the single dispatcher covers it. Trust is also written on the host the PTY actually spawns on, which removes the #11163 class of writing the wrong host's config. * test(agent-trust): type the spawn-builder trust fixtures and prove the spawn waits for trust The builder test passed untyped args (a string launchAgent) and cast its deps, which failed tc:node. It now builds both spawn states from a fully typed deps fixture and a typed restored pane, with no casts. Adds a case that holds the trust write pending and checks the builder does not finish until it settles, the ordering the deleted renderer and launch-prep tests used to cover. * fix(agent-trust): give SSH trust writes the 20 s deadline again The dispatcher gave every non-Codex preset a 1.5 s budget, including the SSH writers for Cursor, Copilot and Qoder, which make several round trips over the link. Before this PR those writes had 20 s (desktop) or no limit (runtime), so on a slow link an unattended SSH worker would now stop at the agent's trust question. SSH writes get the 20 s deadline back; local non-Codex writers keep the short budget, and Codex keeps 20 s. The relay's Claude grant keeps the short budget: it writes the relay host's own disk and does not cross the link. * fix(agent-trust): never pre-trust a home folder or a filesystem root A folder workspace can be the user's home folder or a disk root. Claude and Copilot let a trusted folder cover every folder under it, so pre-trusting one of those would silently trust everything on the machine for those agents. One check, isHomeOrFilesystemRoot, now refuses them for every preset: the dispatcher checks roots and this machine's homes (including the spawn env's HOME and a cached WSL guest home), the SSH writer checks the remote home it already resolves, and the relay checks its own home. The agent then asks, as it would without Orca. * refactor(agent-trust): drop the Codex launch plumbing that only carried trust The spawn hook replaced the trust writes in Codex launch prep and resume prep, which left the fields that fed them unread: CodexHomeLaunchContext.workspacePath and .launchAgent, the resume prep's workspacePath, and the structured Codex launch input's workspacePath (plus the extra target lookup that produced it). Remove them and their plumbing; unavailableManagedHomePath stays. Also removes test stubs of runtime trust methods this PR deleted, whose not-called assertions could no longer fail, and two comments that still described the old trust preflight. * chore(reliability-gates): point the trust gate at the spawn-time trust tests The agent-session trust gate still listed three test files this PR deleted (the renderer preflight, the Codex launch-prep deadline and the e2e trust completion suites), so check:reliability-gates, which runs in PR CI and in pnpm lint, failed on missing files. Its invariant also described the deleted IPC handler and pre-spawn writers. The gate now covers what replaced them: the spawn builders holding the spawn until trust settles, the fresh-launch and setting gates, the per-preset deadlines, and the home and root refusal. * fix(agent-trust): skip the relay Claude grant for a WSL shell On a Windows SSH host whose pane shell is wsl.exe, Claude runs inside the WSL guest and reads the guest's config. The relay still granted trust in the Windows host's own .claude.json, writing the Windows home for a WSL launch, which the local path never does. The relay now skips the grant there, so that Claude asks, as a local WSL launch does when Orca cannot reach the guest file. * perf(agent-trust): only agent launches wait on the trust hook Both spawn builders awaited the trust hook on every spawn, including plain shells, reattaches and agents without a preset. Awaiting even a resolved promise adds microtask ticks ahead of the pane-spawn reservation check, and this handler already keeps non-Codex spawns off an await because an extra tick reorders those reservation races. The hook now returns null when there is nothing to write, and the builders await only a real trust write. * test(agent-trust): keep the home and root cases off any real Claude config The home and root cases ran the real Claude writer with the test process's env, so a regression in the guard would have written trust for the home folder and / into whatever Claude config that env named. They now point CLAUDE_CONFIG_DIR at a folder that does not exist, and the writer never creates a config. * test(runtime): drop needless casts from the launch-host test The renamed launch-host test kept three `as never` casts on launch options that already match launchAgentTerminal's parameter type. The changed-lines casting gate reads the renamed file as new and failed on them. * test(agent-trust): type the Claude grant mock with the real writer's signature The mock took an unknown target, so installing the real writer as its implementation would not typecheck under strict function types. * fix(codex): drop the launch context the trust move left unread in local spawn env * fix(agent-trust): queue Claude grants per config file so a launch burst keeps them all Concurrent grants in one process retried Claude's file lock in lockstep, so each retry round admitted about one winner. Starting 12 Claude agents at once left 6 of them at the trust question with nothing logged. Grants for one config file now queue in-process; only Claude's own writes contend for the lock. The relay shares the writer, so bursts of SSH launches are covered too. * fix(agent-trust): never pre-trust a folder above a home either The guard refused only an exact home or a filesystem root. A folder workspace at /Users, /home or C:\Users was still pre-trusted, and Claude walks up parent folders for a non-git folder, so every non-git folder in the user's home became trusted. The guard now also refuses any folder that contains a home, on every host, and is renamed to say what it decides. * perf(agent-trust): skip the SSH round trips for Antigravity, which has no remote writer Every Antigravity launch over SSH now reaches the remote trust writer, which resolved the remote home and realpath'd the workspace over the link before writing nothing (the known remote gap). That delayed each launch by two SSH round trips, and up to the 20 s deadline on a stalled link. It now returns first. * fix(settings): keep the hidden folder trust row out of web-client settings search The paired web client hides the host-only "Trust the folder" row, but settings search still listed it, so searching "trust" opened the Agents pane with no matching row. Its search entry is now filtered the same way as Agent Awake. * fix(agent-trust): a failing breadth guard skips trust instead of failing the spawn * docs(qoder): New Tab now pre-trusts through the agent-wide spawn hook * fix(agent-trust): never pre-trust a home reached through a symlink Every trust writer stores the workspace's resolved path, but the breadth guard compared only the path as given. A folder workspace that is a symlink to the home folder (or a real home picked while HOME names a symlinked one, as on distros that link /home to /var/home) passed the guard, and Claude, Copilot and Cursor then trusted the home itself. The local dispatcher and the relay now compare given and resolved forms of both the workspace and each home. The SSH writer resolves the remote home alongside the workspace, in parallel, so it adds no round trip. Local non-Claude WSL launches still skip before any filesystem call. * fix(relay): a failing breadth guard skips Claude trust instead of failing the SSH spawn The relay ran its home/root guard and homedir() before its catch, so a throw there rejected the relay's terminal spawn. Same fix as the main dispatcher's: the whole grant, guard included, is best-effort. * fix(agent-trust): guard the path each writer stores, not the path Orca was asked to trust The breadth guard checked the launch's workspace while each writer stored a transformed path, so every new transformation opened a hole. Codex stores a linked worktree's main checkout: with a git repo rooted at the home, a Codex launch in one of its worktrees wrote trust for the whole home. One relay-safe host module now computes the stored path (Codex's main-checkout hop, then given and resolved forms of it and of each home), refuses a root, a home or a folder above one, and only then writes. Main uses it for local and WSL launches and the relay for Claude. An unknown home writes nothing, and the WSL home cache is keyed case-insensitively by distro. * fix(ssh): the relay writes every preset's trust on the SSH host itself Codex, Cursor, Copilot and Qoder trust over SSH was written from the desktop over SFTP: four or five round trips per launch, so it needed a 20 s deadline that outlasted the 8 s draft paste, the 10 s phone wait and the 15 s web-client create. It also skipped Claude's atomic rename, ignored CODEX_HOME, and stored the worktree where local Codex stores the main checkout. The unreleased `claudeFolderTrust` spawn field becomes `agentWorkspaceTrust`, sent for every preset. The relay derives the preset from the `launchAgent` it already receives and runs the same host writer main uses, on its own disk, within 1.5 s and with no extra round trip. Antigravity still returns early on the relay (its writer is unverified on SSH hosts), a WSL shell still skips, and any throw means the agent asks. Deleted: the SFTP preset writer, the remote Qoder writer, the SSH deadline clause and the desktop-side SSH root pre-check. * test(e2e): keep CLAUDE_CONFIG_DIR out of isolated Electron launches The spawn hook now writes Claude folder trust into the config CLAUDE_CONFIG_DIR names, so an e2e run started from a shell that sets it could add trust entries to the developer's real Claude config. Also drops a stale comment that still named Codex launch prep as the trust owner. * fix(agent-trust): guard Claude's resolve() form of the workspace too Claude's writer stores both resolve(path) and the realpath. The breadth guard compared only the given path and its realpath, so a workspace path that does not exist and climbs back with `..` (for example <home>/missing/..) passed the guard while Claude stored a key for the home itself. The guard now also compares resolve(path), so it sees every form a writer stores. * test(relay): pty.spawn writes agent trust before the agent's process starts Nothing exercised the relay handler's call into the trust writer, so removing that call, or no longer awaiting it, left every suite green while SSH launches silently stopped pre-trusting. The new case holds the trust call pending and checks the spawn waits for it, and that the call gets the request, the declared agent and the final spawn env. The reliability gate lists the new suite and records the resolve() form the breadth guard now compares. * fix(agent-trust): refuse a home only for agents that inherit trust from it The home and root refusal applied to every preset, so Codex, Cursor and Antigravity started asking in a home folder workspace, where they did not before. Only Claude, Copilot and Qoder let trust on a folder cover the folders below it; Codex matches its start folder or that folder's repo root, Antigravity the exact folder, and Cursor itself never inherits from a home, a folder above one or a shallow path. The refusal now reads a per-preset table in the host module, so the local and relay writers share the rule. * fix(agent-trust): trust Codex at the folder it starts in, as before Before this PR, Codex launch prep trusted the spawn's start folder. The spawn hook trusted only the workspace root and skipped terminals with no workspace, so Codex began asking in a floating terminal and in a subfolder of a non-git folder workspace: its lookup checks the start folder, then that folder's repo root, and a plain folder above it is neither. The hook now passes the resolved start folder for presets marked as keyed by it (Codex only), falling back to the workspace root. * fix(agent-trust): pre-trust a structured Codex chat's folder, as before Before this PR, creating a structured (native) Codex chat pre-wrote Codex trust for its folder through launch preparation. The PR removed that write and routed trust through the PTY spawn builders, which a structured chat never passes. Codex's app-server trusts the folder itself only when the chat's permissions can write it, so a read-only chat started running untrusted and ignored the project's .codex config. Creating the chat now calls the same dispatcher, behind the same setting, before launch prep. * fix(settings): plainer folder trust setting text |
||
|
|
2807332735 |
refactor(shared): bring constants.ts back under the max-lines limit (#23923)
* refactor(shared): move onboarding, notification and terminal platform defaults out of constants.ts * chore(lint): keep the shapedSidebar naming exemption on the file that now holds it * chore(i18n): regenerate the runtime catalog so it covers main's shipped keys |
||
|
|
e87772b3a4 |
test: retire a dormant worker through host-owned status (#23686)
* test(e2e): await renderer recovery after worker exit * test(e2e): publish worker recovery through authenticated hooks * test: keep retired background worker dormant before activation |
||
|
|
9b46c3f0f2 |
fix(terminal): let unselected Cmd+C reach apps that own their selection (#23597)
* fix(terminal): let unselected Cmd+C reach apps that negotiated kitty keyboard On macOS Orca swallowed an unselected Cmd+C as a no-op copy, so a full-screen TUI like Codex, which captures the mouse and keeps its highlight out of xterm's selection, never received its own copy chord. - isAppOwnedCopyChord (xterm-bypass-policy) is the one rule: macOS, no xterm selection, and non-zero kitty flags from the pane's mirror. The pane's xterm bypass and the dashboard popout's key handler both use it. - A selection copy stays claimed through its repeats and release, including custom copy bindings, so kitty event reporting cannot leak them to the PTY. - Plain shells keep sending nothing; Linux and Windows are unchanged. - The e2e kitty helpers move to helpers/terminal-kitty-keyboard.ts so the shortcut spec stays under the line limit. * fix(terminal): popout copy ownership reads xterm's selection like the pane A highlight of blank cells trims to empty text but is still a selection, so the popout must not hand that Cmd+C to a kitty app while the pane withholds it. * test(terminal): keep the held-copy binding test beside the copy dispatch tests The shortcut-policy suite is at its line limit. * test(terminal): stub a blank-cell selection without widening the preview harness type * style(terminal): tighten the app-owned copy comment and read the popout selection after the early return |
||
|
|
6c56c0a3dc |
fix(browser): a failed SSH route keeps its card while the host redials (#23465)
* fix(browser): a failed SSH route keeps its card while the host redials A browser route that already failed swapped its "SSH connection unavailable" card for "Connecting" on every dial of its host, and re-ran prepare once per dial cycle, so the card flickered and its buttons detached mid-click. Only a route that is still preparing now waits on a dialing host; a failed route keeps its card until the host actually connects, which re-derives it. Retry and Try anyway land on preparing together with the new attempt, so a press while the host dials waits for the connect instead of starting a prepare the effect immediately cancels. The escape-hatch e2e now makes the host truly unreachable before the disconnect; it passed before only because the card stayed latched over a host the terminal had already reconnected. * fix(browser): an unrouted SSH route waits for a dialing host too Only a failed or ready route is exempt from the host wait; a route that just became routed (or still shows another target's page) has no answer for this target, so it must not start a prepare that its own preparing write cancels. The escape-hatch spec now reads the settled failure from the renderer store: main's ssh:getState drops the entry on disconnect and on a failed connect, so its status is null there and never matches the failure pattern. |
||
|
|
29847641ab |
test(e2e): fix failures and improve stability (#23480)
- Narrow toolbar width and set window minimums for consistent testing - Add node_modules symlink to fixture for ESM import resolution - Exercise manual paging and fix button selector - Preserve repo filters in reveal workflow - Adjust timing strategy and increase test timeout |
||
|
|
4cafa50ec0 |
fix(windows): reuse shared PowerShell literal quoting at every hand-rolled escaper (#23083)
Co-authored-by: Orca Worker <orca-worker@localhost> |
||
|
|
82412dab8b |
Persist profile state in SQLite with background writes (#22612)
Migrate profile state to SQLite and move writes and backups into a background worker. Acknowledge terminal, SSH and automation changes only after durable saves. Preserve JSON import, recovery, rollback and compatibility exports. Validate migration, worker failures, maintenance, cross-profile moves and terminal lifetime races with unit, integration and end-to-end coverage. |
||
|
|
7d4413b3d7 |
fix(pdf): keep the search counter in sync with selected matches (#22420)
* fix: update the PDF search counter Co-authored-by: BM Cho <bm1016bm@gmail.com> Co-authored-by: makoto-developer <72484465+makoto-developer@users.noreply.github.com> * test(pdf): respect explicitly headful launch mode * test(pdf): wait for rendered folder search highlights * test(pdf): wait for rendered text before initial search --------- Co-authored-by: BM Cho <bm1016bm@gmail.com> Co-authored-by: makoto-developer <72484465+makoto-developer@users.noreply.github.com> |
||
|
|
80e0bee23b |
fix(floating-workspace): keep agent launches from moving the main window's tab (#22603)
* fix(floating-workspace): keep agent launches from moving the main window's tab
Launching an agent from the floating workspace's "+" menu switched the main
window off whatever chat or editor tab it was showing and onto its terminals.
The main window's selection is supposed to move only for the worktree it is
showing: browser and editor tab creation, splits, moves and drops all check
`activeWorktreeId === worktreeId` before touching it. Two places did not:
- `launchAgentInNewTab` called `setActiveTabType('terminal')` without a
worktree, which targets the active worktree whatever worktree the launch
landed in.
- terminal `createTab` wrote the global `activeTabId` for a tab in any
worktree.
Both now follow the store rule. The launch still selects its tab within its
own worktree, which is what the floating panel renders.
The floating titlebar button had side-stepped this with an `activate: false`
opt-out plus manual selection. That opt-out had no other caller and is removed;
the button now launches and focuses like every other entry point.
* fix(tabs): scope the remaining launch surface writes to the launch's worktree
Three more launch paths create a terminal tab and then call
`setActiveTabType('terminal')` without a worktree, which targets whatever
worktree is active when the call runs rather than the one the tab landed in:
- the paired-host agent launch, after the host's asynchronous create
- Session History resume, which can target a worktree the user is not viewing
and activates it only afterwards
- sleeping-agent resume, which the activation gate runs after asynchronous
readiness checks, by which time the user may have moved to another worktree
Each now names its worktree, like the local agent launch. The new tab still
lands selected when the user switches to that worktree.
* test(tabs): pin the paired-host launch scope in its existing web-runtime test
* test(tabs): type the left-worktree resume fixture instead of casting it
* refactor(tabs): require the worktree that setActiveTabType applies to
`setActiveTabType(type, worktreeId?)` quietly fell back to the active
worktree when the caller left the worktree out. A caller acting on a tab in
another worktree (the floating workspace, a background launch, a reveal that
lands after an async step) therefore retyped whatever the main window was
showing. The launch paths fixed earlier in this branch were instances of that;
54 other callers still relied on the fallback.
The worktree is now a required argument (nullable only for the no-active-
worktree case), so every caller states which worktree it means and a new
unscoped call fails to compile. Each call site passes the worktree of the tab
it acts on; where that is by construction the active worktree (shortcuts,
palette, tab strip), the result is unchanged. `activateTabAndFocusPane`
resolves the tab's owning worktree the same way `setActiveTab` does.
End-to-end helpers that drive the store directly pass the active worktree,
which keeps their previous behaviour.
* fix(floating-workspace): let the floating New Terminal activate its own tab
The floating "+" New Terminal created its tab with `activate: false` and then
selected it with `activateTab`, because creating an active tab used to write
the main window's selected tab even for another worktree. `createTab` now
activates a tab only within its own worktree's group unless that worktree is
the one on screen, so the workaround is no longer needed.
Creating the tab active also moves the floating workspace's remembered tab to
the new one; before, it stayed on the previously selected floating tab, which
auto-acknowledge reads to decide which floating agent the user is looking at.
* refactor(floating-workspace): route every floating New Terminal through one creator
The floating "+" New Terminal had stopped deferring activation, but Cmd+T with the
floating panel focused still went through a separate creator that created the tab
inactive and activated it by hand, which left the floating workspace's remembered tab
on the previous tab. Both now call createFloatingWorkspaceTerminalTab, which creates
the tab active in its own group and focuses it.
* docs(tabs): say why an unowned tab id keeps the on-screen worktree scope
|
||
|
|
dfff3915c4 |
fix(browser): scope back/forward/reload/zoom/grab shortcuts to the originating split (#22340)
* fix(browser): scope back/forward/reload/zoom/grab shortcuts to the originating split With two browser panes visible in a split, Back, Forward, Reload, Hard Reload, page zoom and Focus Address Bar fired in every visible pane. Main forwarded these guest chords without the page id, and each split's active pane subscribed. The renderer-side listeners for the same chords were also window-wide per pane, so a key pressed in the toolbar (or in a terminal in another split) reached every active browser pane. Guest-forwarded chords now carry the originating browserPageId; preload admits only well-formed payloads and each pane ignores ids that aren't its own. Toolbar-path listeners use the same focused-split scope Find already uses. The streamed remote pane's history chord moves onto that scoped hook. Cmd/Ctrl+C grab (STA-3319) gets the same scope and no longer arms while a text selection exists outside the browser pane, so copying from the native chat transcript works again. * refactor(browser): simplify split shortcut scoping per review Drop the preload payload admission (main and preload ship together), fold the three inline scope checks into browserChromeShortcutOwnsEvent, and replace the outside-overlay selection check with a plain live-selection rule so Cmd+C copies from surfaces that do not move split focus. * refactor(browser): share one zoom command type and tidy shortcut comments BrowserPageZoomEventDetail and BrowserPageZoomCommand were the same shape; keep one in shared/browser-page-zoom.ts and route guest and local zoom through a single handler. * refactor(browser): narrow the zoom event with instanceof instead of a cast * test(e2e): pin split-scoped browser shortcuts Two browser splits (and a terminal beside a browser) now prove that Back, Forward, Reload, Hard Reload, page zoom, Focus Address Bar, and the element grab chord act only on the split that sent them, from both the guest page and the browser toolbar. A native chat selection proves Cmd/Ctrl+C copies instead of arming grab. Split fixtures move to a shared helper so both specs reuse them. |
||
|
|
ea02d90704 |
fix(omp): answer startup Kitty queries before renderer handoff (#20620)
* fix(omp): answer startup Kitty queries before renderer handoff Forward actual renderer capability through local and remote spawn. Preserve source ranges and following keyboard mode pushes, and retain independent ConPTY color authority. Refs #17081. Secondary review: #17082. Co-authored-by: stevelliu <stevelliu@tencent.com> * test(omp): cover fragmented keyboard modes and ConPTY handoff * fix: preserve keyboard startup intent without terminal colors * fix: negotiate keyboard support for host-authoritative agent launches * fix: keep terminal creation within line budget * fix(omp): negotiate keyboard support for paired web launches * test: remove obsolete message type import after main integration * fix: validate paired launch results and retry incomplete SSH test snapshots * fix(omp): negotiate keyboard support for background paired launches --------- Co-authored-by: stevelliu <stevelliu@tencent.com> |
||
|
|
2a53293b11 |
test(e2e): stop a spec's parking-delay override from leaking into the rest of its worker (#21571)
* test(e2e): scope the parking-delay override to the spec that needs it
A Playwright worker imports many spec files into one Node process, and the app
fixtures launch Electron with a spread of that process's env. The split-
orientation spec set ORCA_E2E_TERMINAL_PARKING_DELAY_MS at module scope, so the
2s override outlived the file and reconfigured every app launched by every spec
that followed it in the same worker — measured directly: a probe spec sees a
30000ms cold-park delay on its own and 2000ms when that spec runs first.
Module scope is the part that cannot be undone. A write inside a test body can
save and restore, as four other specs here do; a write at import time runs
before any hook exists to restore it. test.use({ orcaAppExtraEnv }) reaches the
app launch without touching the worker every other spec shares.
The ratchet holds the module-scope writer count at zero.
* test(e2e): re-apply screen-reader mode while reading the accessibility tree
Separate from the env leak above, and unproven against the CI failure it
resembles: this is robustness, not a diagnosed fix.
screenReaderMode is an option on the xterm instance and the accessibility tree
belongs to that instance's DOM. The SSH cold-activation spec set it once,
imperatively, then waited on the node. A pane that parks and remounts, or
rebinds after a reconnect, comes back as a new instance with the option off, so
the one-shot mutation stops producing the node the wait is waiting for and the
wait reports "element(s) not found" rather than a content mismatch.
The helper re-applies the option inside the poll and returns null when the node
is absent, so a replaced instance is retried instead of being fatal. Five other
specs still use the one-shot pattern and are left alone.
|
||
|
|
2038376d8e |
fix(terminal): a park must not discard the only copy of a remote pane's scrollback (#21285)
* fix(terminal): keep a client copy of a parked remote pane's scrollback
A remote-runtime pty's bytes never transit the client's main process, so the pane's
xterm buffer is the only client-side copy. The ordinary cold-park unmounted that pane
without capturing it, licensed by TERMINAL_PAIRED_PARKING_RUNTIME_CAPABILITY — a static
build string that says nothing about whether the host retained this pty's buffer. On
reveal, a host that answers 'no-serializable-buffer' (or stays silent past the request
timeout) collapses to a null snapshot and the pane paints blank: tabs and splits survive,
the scrollback is gone.
Capture before every park, not only the retention-budget force-park, so the reveal has a
copy to replay when the host cannot answer. An unverifiable host answer is not proof the
pane was empty; keep the buffer, never discard it.
Adds ORCA_E2E_FORCE_REMOTE_TERMINAL_SNAPSHOT_UNAVAILABLE so an e2e can reproduce the
host-retains-nothing state, mirroring the existing forced-truncation lever.
* test(terminal): prove a parked remote pane survives a host that answers nothing
The oracle is a token the test types into the terminal before the park and the fixture
echoes back. Nothing replays stdin, so a respawned command cannot reproduce that line —
only the pre-park buffer can. An earlier argv marker passed vacuously for exactly that
reason.
The control ('host retains the buffer') is insensitive to the fix and fails if the harness
never parks, never reveals, or never echoed the token, so the regression case cannot be
green for a harness reason.
* refactor(terminal): validate the paired host terminal RPC shape instead of casting it
The merge-commit consistent-type-assertions gate flags every new `as`. Two were fixture
shapes that a type annotation states directly, and the third hid an unchecked RPC payload —
readCreatedTerminalTab now fails with the shape named rather than surfacing later as an
undefined surface id.
* fix(terminal): let a park capture survive an unhydrated repo catalog
Reading state.repos unguarded threw out of the cold-park effect whenever the catalog was
absent, which would break parking itself. Capture is best-effort evidence; an empty catalog
also fails open in shouldPreserveTerminalScrollbackBuffers, the safe direction for a park.
* docs(terminal): pin why the two unhydrated-catalog fallbacks point opposite ways
shouldPreserveTerminalScrollbackBuffers fails open toward 'remote' because a worktree wrongly
judged local parks with no copy at all. worktree-runtime-owner.ts resolves the same unhydrated
catalog to 'local', which is safe there and would be data loss here. A reader pattern-matching
'fail open' across the two gets one of them backwards.
* fix(terminal): keep a parked pane's scrollback across a reconnect merge
The direct-SSH pull replaces a replaced tab's layout wholesale, and a park capture does not
bump tab.generation — so a just-parked tab is not in locallyPreservedTabIds and the only
client-side copy of its remote scrollback went with the layout it replaced. That is the same
data loss this branch already fixes, one layer down, and it is the layer that decides whether
the fix survives the app update the user actually performed.
Carry the client's leaf-keyed scrollback into the host's layout, filtered to the host's own
root leaves. Structure stays the host's verbatim, so a split it added while we were away still
wins and a leaf it retired still drops its bytes. Local wins a conflict: neither copy is then
the only one, but remote-wins would overwrite the tail captured since the last upload and
propagate that backwards on the next replace-session patch.
Not a generation bump: the pane key is `${tab.id}-${tab.generation}`, so bumping would remount
the pane and destroy the very buffer the capture just serialized, lift the recovery-storm
ledger ceiling, and let a stale local ptyId win through preserveNewerLocalTerminalFields.
* fix(terminal): carry a parked pane's scrollback through the mirrored-layout rebuild
Found in review of this PR by rc-ssh-remoting. chooseRemoteTerminalLayout rebuilds a
mirrored tab's layout from the host's picture and never carried buffersByLeafId or
scrollbackRefsByLeafId forward, though it already receives existingLayout. The host
publishes no scrollback of its own, so ANY session-inventory frame landing between park and
reveal dropped the only client-side copy: the rebuild is bufferless, terminalLayoutEqual
compares buffers so the write is not bailed out, and apply-terminal-records assigns it
wholesale.
Measured before the fix: 336 bytes captured at park, 0 after one forced frame, blank pane on
reveal. After: 411 bytes survive the frame and the reveal repaints.
The e2e passed either way because no frame happened to land in its window, so it was not
covering the destroying event. It now forces one inside the park -> reveal window and asserts
the capture survives it.
An identical fix was written and reverted earlier in this branch as 'no measurable effect' —
that measurement ran on a harness deleting the client profile between launches, so nothing
downstream of persistence could register. It was never actually tested.
* feat(session): add a local-only home for ordinary-park scrollback
localOnlyScrollbackByTabId is a top-level session field, tabId -> leafId -> buffer, that never
rides the remote projection: exportRemoteWorkspaceSession is an explicit allowlist of named
top-level fields, so a new one is omitted for free, whereas anything added to
TerminalLayoutSnapshot is copied whole. It is also outside the two records the mirrored-tab apply
rewrites, so a host inventory frame cannot wipe it.
Registered in every exhaustive session registry ('tabKeyed'), hydrated and scoped like the layout
map, dropped with its tab on close/removal/purge/repo removal/mirrored retirement, copied on profile
transfer, emitted by the incremental patch builder, and capped by pruneLocalTerminalScrollbackBuffers
alongside the shared home — with a per-home test so an uncapped path cannot go unnoticed.
Known ceiling, not widened here: the field routes through the partition router that falls back to
'local' when the repo catalog is unknown at write time (#21295).
* fix(terminal): keep ordinary-park scrollback off the upload, and read both homes through one resolver
The ordinary cold park fires on every workspace hide. Its capture now splits: structure (root,
ptyIds, titles) stays in the shared layout, bytes go to localOnlyScrollbackByTabId. Force-park,
hibernate, sleep and shutdown keep writing buffersByLeafId, because that copy is what a second
desktop cold-restores from; a shared capture clears the local copy so the two homes never hold two
versions of one leaf.
resolveLeafScrollbackBuffers is the only read across the two homes (local wins a conflict: it is
the later write by construction). restoreTerminalPaneLayout no longer reads buffersByLeafId
directly, the capture's merge prior comes from the resolver, and the post-replay release covers
both homes.
Measured with the projection at 20 tabs x 2 panes at the per-leaf cap: the shared-layout shape
exports ~22 MiB per replace-session; the local-only shape exports the bufferless baseline.
* test(sync): pin that the mirrored rebuild carries the client scrollback refs
The carry-through added in
|
||
|
|
09073086a8 |
feat(terminal): inline images via @xterm/addon-image (perf-first) (#19512)
* feat(terminal): inline images via @xterm/addon-image, perf-first Add opt-in inline terminal images (SIXEL, iTerm2 IIP, Kitty graphics) through @xterm/addon-image, designed to keep idle terminals unaffected. Performance: - The addon (base64-inlined wasm decoders + protocol handlers) loads off the boot critical path via a deferred loader that mirrors the WebGL addon: primed after first paint only when the setting is on, read back synchronously at attach, with a 3-attempt cap so a transient failure never disables images for the session and a missing chunk never refetches per pane. renderer-boot-graph guards against eager import. - enableSizeReports:false so the addon never sets windowOptions and double-answers Orca's own CSI 14t/16t responder. - Perf-tuned decode/storage limits (storageLimit, sixel/iip/kitty size caps) in one place. Correctness: - Orca's DA1 handler wins over the addon's (last-registered-first), and the default DA1 response never advertised Sixel (;4), so DA1-detecting tools (chafa, img2sixel, viu, timg) never emitted it. The winning handler now appends ;4 while the setting is on, resolved per query so a live toggle changes the next DA1; idempotent against the ConPTY response that already lists it. - ORCA_IMAGE_PROTOCOL=kitty is exported to spawned shells (local, daemon, relay/SSH) and forwarded across the WSL boundary, so image-capable agents can pick an encoder. Unknown image sequences are swallowed by xterm when the addon is detached, so this never garbles output. - Settings toggle (default on) gates rendering and DA1 advertisement. Cross-checked against community PRs #7775, #11706, and #19201 at the end; credited below. Co-authored-by: s546126 <s546126@users.noreply.github.com> Co-authored-by: XRX193 <XRX193@users.noreply.github.com> Co-authored-by: lmsh7 <lmsh7@users.noreply.github.com> * fix(terminal): bound inline image memory and classify Kitty replies * fix(terminal): bound image decode and release image resources on cleanup * fix(terminal): address image addon review feedback * test(terminal): stub setPaneInlineImagesEnabled in appearance manager fakes * fix(terminal): evict unplaced kitty payloads before displayed images Byte-budget eviction dropped the oldest transmitted blob regardless of placement, so a new upload could erase a visible image while abandoned blobs still held budget. Unplaced payloads now go first and displayed ones only when that is not enough. The incoming image is always stored, so an oversized one overshoots the cap by one payload instead of being dropped after the protocol already acked OK. * fix(terminal): gate DA1 Sixel on real addon attachment; claim SSH image spec in CI - DA1 advertised Sixel from the setting alone, so a pane whose lazy addon chunk was still loading (or had failed all three attempts) told feature-detecting tools to emit DCS that nothing could render. Track the attached decoder per terminal and require it before setting the ;4 bit. - tests/e2e/terminal-inline-images-ssh.spec.ts was Docker-gated but claimed by no lane runner, so pr-e2e-gate-contract failed and the spec would have self-skipped green forever. - Reject non-positive PNG IHDR dimensions before decode: they are parsed with signed shifts, so a dimension >= 0x80000000 came back negative and slipped past the pixel-limit comparison. - One resolveTerminalInlineImagesEnabled() for the default-on setting; the four call sites mixed '?? true' with '!== false', which disagree on null. - One readInlineImageResources() walk of the addon internals instead of two copies that could drift against the patched dependency. - Isolate the deferred-attach drain per pane; make the zoom-invariance and backing-storage e2e assertions fail when the feature is dead. * refactor(terminal): one lazy xterm addon loader for webgl and image terminal-image-addon-loader was a structural clone of the webgl one — same memo, attempt cap, and .then(ok,err)-clears-memo recovery. Both now wrap createLazyXtermAddonLoader; each keeps its literal import() specifier so the bundler still splits the chunk (verified against a fresh build: addon-image stays out of the boot graph). * refactor(terminal): name openTerminal's addon flags; pin image addon limits Two adjacent optional booleans could be swapped without a type error once inline images added the second one. * docs(terminal): state the real per-pane image ceiling; drop test ordering dependency storageLimit:32 reads like the pane's budget but keys three pools — decoded pixels, retained encoded Kitty blobs, and pending WASM decoders — so the worst case is ~98 MB per pane with no cross-pane governor. Say so at the constant. pane-inline-images.test.ts's deferred case needed to run first; it now takes a fresh module instead, and the rest prime in beforeAll. Verified by running the file with that test moved last. * fix(terminal): satisfy rebased static analysis gate * fix(terminal): complete casting gate cleanup * fix(terminal): recover failed image addon loads * fix(terminal): bound image decoder allocations --------- Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local> Co-authored-by: s546126 <s546126@users.noreply.github.com> Co-authored-by: XRX193 <XRX193@users.noreply.github.com> Co-authored-by: lmsh7 <lmsh7@users.noreply.github.com> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
5c8540948d |
test(e2e): name the paired-client quit that preserves the profile (#21300)
Quit-without-deleting is closeElectronAppForE2E + cleanupE2EDaemons — dispose's first two steps without removeProfile. The composition is correct today but undiscoverable, and getting it wrong is silent and expensive in both directions. dispose() + reuseUserDataDir yields a FIRST RUN on an empty profile, so every persistence assertion after it reads empty and is indistinguishable from data loss. That produced a phantom data-loss report, live in two write-ups before a diagnostic listing zero session FILES (rather than zero buffers) contradicted it. Reaching for a bare app.close() to skip the deletion hangs instead: it lacks the timeout and force-kill fallback that closeElectronAppForE2E wraps around it, and burned a ten minute test deadline producing no reading at all. Test infrastructure only; no production code. Unblocks restart-persistence coverage for the paired topology. |
||
|
|
852ee907ee |
fix(e2e): stabilize flaky E2E tests against timing races (#20900)
* fix(e2e): stabilize flaky E2E tests against timing races - Paired terminal: use stable cold activation assertion instead of racy one-shot read; background tabs park eagerly. - Native chat: scope hydration assertions to transcript subtree to avoid false positives from UI chrome (worktree rows, tab titles). - Onboarding: inject verified status snapshot with max sequence to prevent hydration from downgrading host health during skip-to- project-setup. - Paired web: encode host health faults in snapshots with high sequence so real hydrations cannot outbid injected state. - Quick open: clear prior tooltips and increase hover timeouts to handle streaming result remounting. - Terminal attention: pass 'terminal-bell' to unread marker to match production contract (reads marker value, not presence). * fix one last test |
||
|
|
231e805b1e |
fix(lint): enable anti-slop/no-shape-in-symbol-names (#20785)
Flip `anti-slop/no-shape-in-symbol-names` from "off" to "error" and clear
every violation under src, config, tests and mobile.
What the rule bans
------------------
The case-insensitive substring "shape" in any JS/TS identifier: variables,
functions, parameters, types, type parameters, class members, private names,
object-literal keys and JSX identifiers. The one exemption is a statically
accessed member read owned by another value (`zodObject.shape` is fine), so
third-party APIs stay readable without a suppression.
"Shape" names a value's structure rather than its domain role. `UserShape`,
`validateArgShape` and `errorShape` all tell you the symbol is "an object
with some fields" -- which is already what a type says -- while saying
nothing about what the value is for or who owns it. The rule forces the
name to carry the domain instead.
Violations fixed
----------------
689 violations across 109 files at baseline (verified by re-running the
audit against the pre-change tree with the rule set to "error").
Fix pattern
-----------
Rename for the domain role, not the structure:
-type FieldShape = 'list' | 'map' | 'whole'
-const FIELD_SHAPES = { ... } satisfies Record<keyof Observation, FieldShape>
+type FieldEncoding = 'list' | 'map' | 'whole'
+const FIELD_ENCODINGS = { ... } satisfies Record<keyof Observation, FieldEncoding>
-function assertGitPushTargetShape(target: unknown): void
+function assertValidGitPushTarget(target: unknown): void
-function describeReadDirPathShape(p: string): ReadDirPathKind
+function classifyReadDirPath(p: string): ReadDirPathKind
Predicates became statements about the value (`isDeltaShapedProviderFrameKind`
-> `isDeltaProviderFrameKind`, `isDeleteShapedDiscardEntry` ->
`discardDeletesEntryFile`, `isSkillsCliAgentKeyShaped` ->
`isUsableSkillsCliAgentKey`). Type aliases dropped the suffix where the
remaining name was already unambiguous (`GhGraphqlErrorShape` ->
`GhGraphqlError`).
No wire-visible name was renamed: no IPC or RPC channel, stream opcode,
request/response param, persisted field, or i18n key. The `--shape=symlink|copy`
CLI flag read by .github/workflows/skill-update-roundtrip.yml is unchanged --
only the local variable holding it was renamed.
Exemptions
----------
They are file-scoped entries in config/oxlint-anti-slop.json, not inline
`oxlint-disable` comments. An inline directive naming an anti-slop rule reads
back as an UNUSED directive under the root lint scan, which does not load this
plugin -- the changed-code quality gate counts that warning, so the comment form
cannot be used for a rule that lives only in this config.
* src/renderer/src/components/browser-pane/annotate/**:
in the screenshot annotator a "shape" is the drawn geometry -- pen, arrow,
rect, ellipse, highlight. That is a genuine domain noun, and it pervades
every symbol in the module.
* repo-icon.tsx, repo-header-project-actions.tsx, mobile MobileRepoIcon.tsx:
lucide exports the icon component as `Shapes`. The name is theirs, and the
matching REPO_LUCIDE_ICONS key is the persisted icon name shared with the
desktop picker -- renaming it would orphan saved repo icons.
* src/shared/onboarding-state-types.ts, src/shared/constants.ts:
`shapedSidebar` is a persisted onboarding-checklist field and a telemetry
enum member; renaming it would orphan saved state.
* src/shared/rpc-contract/rpc-send-params.ts: matching zod's own literal `shape`
property is what selects the ZodObject branch of the conditional type.
No exemption was added merely to avoid a rename. Eight symbols initially
suppressed as "a cross-module refactor outside this change" were proven to have
zero non-TypeScript references repo-wide and renamed instead.
Zod's `ZodRawShape` needed no exemption at all: `Readonly<Record<string,
z.ZodType>>` is its definition, so repo-update-params.ts and
ui-update-value-tolerance-params.ts spell it out instead. Likewise
telemetry-event-classification.ts now reads `.shape` through an `in` narrowing,
which also retires two pre-existing type assertions; three more assertions the
rename had dragged onto changed lines (two `JSON.parse` sites, one node:sqlite
row read) became annotations and an explicit row mapping.
Verified
--------
* Audit reports zero violations; confirmed the rule genuinely fires by
planting a probe violation.
* node config/scripts/run-typecheck-projects-in-parallel.mjs exits 0.
* Vitest over src/shared, src/main/github/project-view, the annotate module,
the repo-icon components and the Chromium SameSite electron spec: all green.
* All 66 removed "shape" identifiers grepped repo-wide across every file type;
none survive.
* node config/scripts/generate-rpc-params-catalog.mjs --check exits 0.
* node --check on every changed .mjs; oxfmt clean on all changed files.
* `pnpm run check:code-quality:changed` reports 0 findings.
Not machine-verified: the 3 mobile/ files (its Vitest run cannot resolve
`expo/tsconfig.base.json` in this worktree), and the WSL- and Playwright-gated
specs. All are rename- or comment-only hunks, read in full.
|
||
|
|
7b53b5abd1 | test: replace fixed UI waits with observable readiness (#20369) | ||
|
|
729491597f | feat(desktop): measure relay regions and reconnect after idle cutover (#20106) | ||
|
|
20c56249d5 |
fix(terminal): keep a deliberately slept workspace cold until it is woken (#20075)
* fix(terminal): keep a deliberately slept workspace cold until it is woken Sleeping a workspace kills its PTYs but keeps its panes mounted and keeps each tab's session id as a wake hint. Any later remount of those panes (recovery, parking, portals) reattached that dead id, and the daemon's create-or-attach spawned a fresh shell, so slept workspaces revived on their own (#10205). The existing sleep-intent marker now outlives teardown and gates the deferred connect itself, so both the reattach and fresh-spawn arms stay cold. It is released by activating the workspace, by any PTY binding to one of its tabs (CLI, automation, client wake), and by purge. A queued startup still connects. Reproduces the community root cause from gatsby74 in #13343; the regression e2e remounts a slept hidden pane and fails on main. Co-authored-by: gatsby74 <gatsby74@users.noreply.github.com> Co-authored-by: mmarabel <mmarabel@users.noreply.github.com> * fix(terminal): let a slept pane wait for its wake instead of latching cold A pane whose connect ran while its workspace was slept used to mark itself connected and stop; nothing re-armed it, so a wake that produced a live PTY before the user clicked (CLI create, background agent resume, split panes) left panes stranded. The connect now waits on the sleep marker and resumes when the marker clears, and a torn-down pane drops its listener. Tabs created with a live PTY clear the marker too, the sleep flow marks each workspace only when its own teardown starts, and purge forgets the marker without waking anything. * fix(terminal): wake a waiting pane once, in its remounted generation Activation clears the sleep marker after the set() that bumps dead tabs' generations, and the waiting pane only resumes its connect when its tab generation is still current. Otherwise the stale pane and its remounted successor both reattached the same session id on a deliberate wake. * fix(terminal): resolve the waiting pane's tab by either id and re-arm after wake The wake listener looked the tab up by the pane's render id, which can be a unified id whose terminal tab lives under entityId, so the generation check declined forever for those panes. Mount, fresh spawn, and the wake listener now share one live resolver. The wait flag resets when the listener fires so a second sleep can hold the pane again, listener dispatch is guarded, folder activation clears after its own set(), and the sleep flow re-asserts the marker after each teardown while releasing a workspace the user activated meanwhile. * fix(terminal): ignore PTY binds that land inside the sleep teardown window A spawn resolving while shutdown was still awaiting the host bound a PTY and cleared the marker, waking every waiting pane mid-sleep; re-marking afterwards could not un-connect them. The sleep flow now scopes each teardown so binds in that window are not wakes. The e2e asserts a deliberate wake yields exactly one PTY, and the dispose test proves the listener is gone. --------- Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: mmarabel <mmarabel@users.noreply.github.com> |
||
|
|
fb9ba4b681 |
fix(editor): make markdown images inline so a paragraph stays schema-valid (#19746)
* fix(editor): make markdown images inline so a paragraph stays schema-valid Image was registered as a block node while paragraph is content:'inline*', but the markdown pipeline nests an inline image as a paragraph child. Schema.nodeFromJSON does not validate content, so the editor built a schema-invalid document that rendered fine and threw on the first step that reassembled the paragraph - i.e. on the user's next keystroke. Report 0e46c048 (1.4.198, macOS): RangeError "Invalid content for node paragraph" from checkContent via Node.replace, tearing down the editor.rich-markdown boundary. Register Image as inline and override paragraph's parseMarkdown so a lone image is not hoisted out of its paragraph. Also fixes the same crash class reachable through details/summary. Markdown output is byte-identical. * fix(editor): keep a fenced code block intact when an image is inserted into it Making the image node inline meant it could no longer be fitted into codeBlock (content:'text*', marks:''), so inserting one with the cursor inside a fence made ProseMirror close the block at the insertion point: the remaining code escaped as plain prose and the language attribute was lost, and autosave wrote that markdown to the user's file. The pre-fix block image split the fence into two intact blocks instead. Resolve the insert content against the target position: when an inline image cannot be fitted where the caret sits, wrap it in a paragraph so ProseMirror splits the block and both halves keep their ``` fencing and language. Prose insertion is unchanged. Every production insert path now shares that resolution - the toolbar picker, the slash command and the clipboard-screenshot paste through insertRichMarkdownImageFromPath, plus the GitHub/GitLab composer's image-URL insert - each with a regression test. Also guard the unchecked cast of Paragraph.config.parseMarkdown: a Tiptap upgrade that drops the field would otherwise turn every paragraph parse into a TypeError and take the whole editor down, instead of degrading to parseInline. Four of the new round-trip cases asserted only on getMarkdown(), which walks the document without running NodeType.checkContent and so emits byte-identical output from a schema-invalid document - they passed on the pre-fix code. roundTripMarkdown now runs doc.check(), the list-item and table-cell case performs a real edit, and the standalone-image case types beside the image. All twelve cases now fail on the merge-base. Adds an Electron e2e spec driving the real renderer: a paragraph image and a toggle-summary image each survive a keystroke, and Bold over a selection spanning the image keeps it. --------- Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
73d0521410 | Replace the sidebar create dropdown with two direct action buttons (#19653) | ||
|
|
12f53da542 |
Remove settled-worker automatic resume and hibernation fences (#19544)
* Remove settled-worker automatic resume and hibernation fences * test: retirement rollback case follows the no-fence policy Case 4 seeded and asserted automaticResumeBlockedBy, which this branch deletes. A rolled-back settled worker is now an ordinary done record that wake clears as passive evidence, same as any finished agent pane. * chore(i18n): regenerate the runtime-required catalog for the contrast floor strings * test(orchestration): give the stopping-worker guard fixtures a Run |
||
|
|
c056c6f9ac |
Unify sidebar create actions into single dropdown menu (#19375)
* Unify sidebar create actions into a single dropdown menu - Combine "New workspace" and "Add project" under a unified "Create" button - Remove layout logic that split these actions based on sidebar width - Normalize "Add Project" to "Add project" (lowercase) throughout the UI * Use null instead of 'Unassigned' for unassigned shortcut labels Add formatOptionalPrimaryShortcutLabel that returns null when a shortcut is unassigned, enabling simpler conditional rendering in dropdown menus. Remove associated translation strings. |
||
|
|
1848855515 |
test: enable software WebGL for Linux CI headful specs (#19001)
* test: enable CI WebGL and route GPU-dependent regressions * test: retain headful atlas cases in terminal rendering goldens * test: reuse golden command in project coverage assertions |
||
|
|
b3acef218a | test: verify imported projects through the virtualized sidebar (#19003) | ||
|
|
0c33f58e8a |
fix(ssh-relay): daemon owns the endpoint credential; a losing start never rotates it (#19052)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 19 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$962 | $\color{#cf222e}{\Huge{\mathbf{−}}}$136 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$826 |
| Prod | 18 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$295 | $\color{#cf222e}{\Huge{\mathbf{−}}}$116 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$179 |
<!-- /orca-pr-loc -->
## Symptom
Live 2026-09-05 (Orca 1.4.198 client, Ubuntu host): both relay processes `kill -STOP`ped for 20 s, then `-CONT`. The client redeployed while the host was frozen. Its fresh daemon lost the socket bind (`Socket path already in use`) but had **already rewritten** `relay-<id>.sock.credential`. The surviving daemon kept its in-memory credential, so every later `--connect` got `Endpoint credential mismatch; closing socket`, then `Grace started … timeoutMs=0 … ptys=1, clients=0` every ~20 s, forever. Only a manual `kill -TERM` cleared it. Receipts: `review-archive/orchestration-v3-pr16904/smoke-receipts-t012b/E16,E17,E18,E24`.
Three independent defects kept the wedge alive; each is fixed at its own seam.
## Fix
**1. The relay daemon owns credential publication (race-free under two concurrent starters).**
`relay-daemon.ts` binds the socket first, then publishes via the new `src/relay/relay-endpoint-credential-publication.ts`: adopt a valid pre-existing file (older clients still pre-write), else mint 32 random bytes and write temp+rename at 0600. A start that loses the bind exits inside `listen()` and never reaches the file. Why this option and not restore-on-loss or a client-side write: the only process that can *prove* ownership is the one whose `listen()` succeeded, and that proof is atomic with the bind. The client-side pre-write (`ssh-relay-endpoint-credential.ts`) and the launch-command `chmod 600`/`icacls` are removed on POSIX and Windows. The racing test also exposed that macOS reports a mid-bind collision as `EEXIST` rather than `EADDRINUSE`; `relay-socket-ownership.ts` now treats both as "held or stale".
**2. The client distinguishes "no daemon" from "daemon present but not answering", and never rewrites.**
A credential refusal is now typed on the wire: the daemon replies `orca-relay-handshake-credential-mismatch` (same frame type, no new opcode) and the bridge exits **43**; `waitForSentinel` maps it to `RelayCredentialMismatchError`, which the takeover treats as handshake-refusal evidence exactly like exit 42. A relay that holds the endpoint but **never refused** (the stalled-host shape: kernel backlog accepts the probe, handshake gets no answer) is now `RelayEndpointUnresponsiveError`, routed to the relay-lost backoff instead of the terminal Reset Relay path. Silence is not a decision (`docs/reference/ssh-execution-boundary.md`).
**2b. Deploy honours the verdict.** The 40 s live run exposed that the `--connect` catch block in `deployAndLaunchRelay` predates the incumbent probe and swallowed both verdicts as "probe failed, launch fresh", so a fresh daemon was still launched over the live one (it lost the bind by luck, which is exactly the collision in the incident). Held and Unresponsive now propagate; the session backs off on Unresponsive and surfaces Reset Relay on Held. Red-first in `ssh-relay-deploy-incumbent-verdict.test.ts`.
**3. The daemon cannot be wedged by a rotated file, because nothing can rotate it.**
The credential lives in the content-hashed relay dir, and after (1) the only writer is the daemon that owns the socket, so the "file changed under a live daemon" state the incident depended on is no longer reachable in-product. The credential is therefore fixed for the daemon's lifetime, as a plain secret should be. A hand-edited file is refused with the typed reply until restored (tested). Startup adoption of a pre-written file applies an owner-only + same-uid rule (review finding): anything else is replaced by a fresh mint. An earlier revision of this PR also re-read the file on mismatch and adopted it; that was removed as unreachable machinery that turned the credential into a per-handshake file-ownership check.
**3b. Fail closed between bind and publication.** A client that arrives after `listen()` resolves but before the credential is set is refused, not admitted as `unproved`. Nothing can be delivered in that window today; the guard makes the boundary structural instead of an event-loop ordering fact. Red-first in `relay-reconnect-listener-credential-gate.test.ts`.
**Wire compat.** New optional handshake reply only; an old `--connect` hits `Unknown handshake type` and exits 1 pre-sentinel, which it already treated as a generic failure. New daemon adopts an old client's pre-written file; new client still passes `--credential-file` so an old daemon reads it as before. Absence of exit 43 is never used as evidence.
**Also.** `terminal create` on a reconnecting SSH host now says what to do instead of a bare `No PTY provider for connection "<id>"` (prefix preserved; the renderer matches it).
## Tests (red first)
- `src/relay/subprocess.test.ts`: two `--detached` starts race one socket + credential file → exactly one reaches the sentinel, loser exits 1 with `Socket path already in use`, file valid + 0600, a `--connect` reading it reaches `relay.status` and reports the winner's pid. Red before (both starters died: daemon required a pre-existing file), green 6/6 after.
- `src/relay/relay-endpoint-credential-publication.test.ts`: mints after bind; adopts a pre-written 0600 file; replaces a pre-written 0644 file with a fresh mint; refuses a stale credential with exit 43 while still serving the real one, and keeps refusing a rewritten file until it is restored.
- `src/relay/relay-reconnect-listener-credential-gate.test.ts`: a client in the bind-to-publish window is refused and never attached; after publication the right credential is accepted and a wrong one refused; a daemon launched without a credential file is not gated. Red without the guard.
- `ssh-relay-deploy-incumbent-verdict.test.ts`: live-but-silent incumbent → `RelayEndpointUnresponsiveError`, refused → `RelayEndpointHeldError`, and in neither case is `--detached` launched; a failed `test -S` probe still launches fresh. Red 2/3 without the deploy change.
- `ssh-relay-deploy-helpers.test.ts` (exit 43), `ssh-relay-endpoint-takeover.test.ts` (refused → Held even with no `lsof`; silent → Unresponsive, nothing unlinked or signalled), `ssh-relay-session-terminal-error.test.ts` (Unresponsive → `onRelayLost`, not terminal). Deploy/namespace/native-deps tests updated to assert the client writes **no** credential.
## Live proof
New `tests/e2e/ssh-docker-relay-stall-credential.spec.ts` (claimed in `run-ssh-docker-e2e.mjs` and PR source routing), two cases: `kill -STOP` every relay pid in the container, send input during the freeze, hold **20 s** (the incident's duration, which races the mux liveness timeout) or **40 s** (past it for sure), `kill -CONT`; assert status back to `connected`, same pty, same daemon pid, same credential inode and content, relay.log did not shrink (a relaunch truncates it) and has zero `Endpoint credential mismatch` / `Socket path already in use` lines, in-stall input delivered at most once.
Run output (local, fixture image `orca-e2e-ssh-relay:3a864c665ba2cefd`, `ORCA_E2E_SSH_DOCKER=1 SKIP_BUILD=1 ORCA_E2E_FORWARD_APP_LOGS=1 … --project electron-headless --workers=1`, head `c2c20fd994`; re-run identically on the final head after the credential-lifetime change, 2 passed (1.7m), same annotations, and the bind-to-publish refusal never fired):
```
✓ keeps the same daemon and credential across a 20s relay freeze (38.3s)
relay-processes-stopped: 2 relay-processes-continued: 2
bridge-pids-before-after: 480 -> 480
socket-clients-accepted-before-after: 1 -> 1
in-stall-input-delivered: 1
✓ backs off and reattaches, never relaunching, across a 40s relay freeze (57.5s)
relay-processes-stopped: 2 relay-processes-continued: 4
bridge-pids-before-after: 480 -> 1202
socket-clients-accepted-before-after: 1 -> 3
in-stall-input-delivered: 1
2 passed (1.6m)
```
Client log in the 40 s case shows the new path end to end: `Relay channel lost … reconnect attempt 1/6` → `Socket probe result: "ALIVE"` → `Socket reconnect failed … Relay failed to start within 10s` → `Relay endpoint incumbent: … verdict=live evidence=accepted-connection holders=unenumerable` → `Failed to re-establish relay … A relay still owns … but did not answer the handshake … Orca will retry` → `reconnect attempt 2/6` → `Reconnected to existing relay via socket`. The 20 s case never left the frozen bridge (same bridge pid, one accept), so it exercises the "silence is not death" side of the same race. The 20 s case passed 6/6 across the session; the 40 s case was red on the prior head (`Socket path already in use` + `Startup failed: listen EADDRINUSE` in relay.log from the swallowed verdict) and is green after 2b. Before the fix the same injection produced a fresh daemon that rewrote the credential and a survivor refusing every client.
The `relay-processes-continued` count exceeds `stopped` in the 40 s case because the timed-out client's `--connect` bridge and the loser-side processes are parked behind the frozen listener when `CONT` runs; they exit on their own once it resumes.
## Gates
`pnpm test src/relay src/main/ssh` 332 files / 3884 tests pass · `pnpm typecheck:tsc:node` clean · `check:code-quality:changed` 0 findings · `check:react-doctor:changed` 0 findings · `pr-e2e-gate-contract.test.mjs` 42 pass · no lint disables or max-lines bumps added.
## Noted, not fixed here
- `terminal list` `orphaned:false` / `terminal close` `ptyKilled:true` for a pane whose relay is gone (`orca-runtime-stop-explicitly-closed-tab-ptys.ts`): different seam, `@ts-nocheck` characterization-covered file.
- On a host with no `lsof`, a stalled relay still cannot be enumerated as the holder; it is now retried rather than declared held, but a relay frozen past the backoff budget still ends in the existing "reconnect manually" banner.
|
||
|
|
06a607a1d7 |
feat(orchestration): make multi-agent workflows durable (#16904)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 225 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$21666 | $\color{#cf222e}{\Huge{\mathbf{−}}}$2820 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$18846 |
| Prod | 348 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$17107 | $\color{#cf222e}{\Huge{\mathbf{−}}}$4706 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$12401 |
<!-- /orca-pr-loc -->
## ELI5
Orca now treats orchestration like a durable control plane instead of inferring success from terminal keystrokes. Agents can tell whether a prompt was accepted or a turn started, replay an ambiguous request without sending twice, and recover coordinator mail after a crash. Completed workers can be inspected, released, or retained, and their panes no longer auto-resume as if the work were still running.
## What changed
- **Run receipts** from `run-create/use/current/show/list` are the row without routing plumbing (`home_database`, `coordinator_pane_key`) and without the duplicate `binding` object.
- **`terminal send` receipts are honest and idempotent.** `input_accepted` and `turn_started` are the only stages; `--wait-submit` observes without resending; `--retry-request <uuid>` replays the exact request against the same process incarnation. A transport timeout keeps the retry ID; only a different runtime answering strips it. Value-less or non-UUID `--retry-request` is rejected on the CLI and the SSH shim.
- **Mailbox delivery is committed before wakeup.** Pointer writes are staged in the DB before any PTY byte, replayed once after restart, and never emit a naked Enter. The watermark that parks concurrent deliveries is released with the DB reservation. Restart rescans pointer-pending and `dispatch:` mailboxes.
- **Lifecycle is a guarded transition graph** (`lifecycle-transition.ts`) with a table-driven test over every caller edge. Task reopen/overturn stays in the public contract. A PTY exit during `worker-stop` is the stop succeeding, not a failure.
- **Worker lifecycle CLI:** `worker-start` (`--spec` creates Task + attempt in one call), `worker-show`, `worker-read` (provider transcript first, bounded terminal fallback with a typed reason, local/WSL/SSH), `worker-stop`, `worker-abandon`, `worker-release`, `worker-retain`, `worker-list` (rowid-fenced pagination, fleet liveness, `attention`, literal `nextAction`).
- **Release is an explicit ownership table** (`decideWorkerTerminalRelease`): only an `owned` resource can be settled, the archive is mandatory where reachable, and an owner whose process is proven exited can always get out of `retained` via `archive_status: unavailable`. User-taken-over, external, and transferred panes stay retained.
- **Settled-worker resume fence** (folds in #17651): a settled dispatch whose pane is still open is fenced at settlement, on stop/abandon/exit, and at startup; lifted on release, retain, takeover, and pane reuse.
- **Liveness is `live` / `unverifiable` / `exited` only**, from execution-host evidence. Fleet projection reads the evidence clock, not the relay delivery clock. A host-certified exit outranks the worker's settled state. `unverifiable` never authorizes stop, abandon, retry, or release, in code or in the guide.
- **Federation:** structured reads negotiate by `method_not_found` so every shipped host keeps transcript-first output; exited remote workers are closed before being reported closed; epoch fencing holds across peer restart, downgrade, and pairing rotation; no per-second forced capability probe.
- **Schema v35:** repairs databases stamped v34 by the pre-fix branch (mailbox_handle default, index predicates), drops the write-only `lifecycle_transition_receipts` ledger and five never-read v31 identity columns.
- **Schema v36:** `dispatch:<id>` mailboxes get a real consumer generation on `dispatch_contexts` and `remote_dispatch_attachments`, bumped and fenced in the same transaction on every re-attach (manual inject, worker-start, federated attach). A stale worker whose Dispatch moved to another process now gets `consumer_fenced` instead of silently acking the new worker's Delivery. Run mailboxes already worked this way.
- **Schema v37:** `dispatch_contexts` records its creator (`creator_handle`, `creator_pane_key`), so a coordinator's context-only self-dispatch is bookkeeping rather than a nesting parent; before this, one self-dispatch made every later `worker-start` from that coordinator fail the depth cap. Pre-v37 rows keep counting (fails closed).
- **Dispatch-mailbox ownership is checked, not inferred.** A `check` from a process whose pane no longer holds the Dispatch, or whose last Attempt was abandoned/failed and moved to another terminal, gets `consumer_fenced` instead of an empty inbox that reads as "no mail yet". `--peek`/`--all` stay readable. A paneless caller still gets `stable_pane_required` with the rebind recovery.
- **Liveness certification is stricter:** a `process_exited` stage whose termination reason is `unknown` (a stop that was issued but never observed) projects `unverifiable`, not `exited`. Federated `worker-show` carries the execution host's verdict and host kind instead of a local guess. A live, ready worker with nothing pending has `nextAction: none` rather than pointing at the `worker-show` that produced it.
- **Wire:** `workerShow` keeps `dispatch.task_id` next to `taskId` for shipped CLIs. `ask --json` uses the standard `{ok, result}` envelope like every sibling verb.
- **Migration start-version detection** treats the two v32 recovery columns as versioned. Before this, every shipped database stamped below 32 resolved to the v6 floor and replayed the whole chain (the v23 backfill synthesized 68 phantom retained workers on a real v30 profile). Verified on a copy of a real 62 MB v30 profile: starts at 30, no row delta, integrity ok, 11 ms.
- **Skill guide** rewritten as a ≤200-line kernel plus seven references, to the outcome-first standard (Result / Done / Safe failure first, conditions not case lists, one done bar, references loaded at the point of use). The canonical loop uses `worker-start --spec`, names `worker-list` for completion accounting, documents `--retry-request` / `request-show` / `--wait-submit`, and requires positive evidence before any stall action. The other seven guides get the same treatment in #18724, split out so this PR stays orchestration-only.
- **`rpc/methods/orchestration-*`** (126 flat files) regrouped into `orchestration/{worker,federation,messaging,runs,gates}/`.
## Why
User reports showed the same boundary failures: false `agent_prompt_stalled` causing duplicate sends (#15180), coordinators unable to trust screen scrapes, cold-parked terminals receiving a pointer without the submit, settled workers accumulating as live tabs and auto-resuming after restart, and no way to tell a stalled worker from a working one.
## Linked issues
Fixes #15180. Fixes #17935 (orchestration skill description is 866 characters; a guard now caps every bundled skill at 1,024). Supersedes #17651 (fence folded in). Advances #16660, #16522, #14907, #13047.
## Review record
This PR was reviewed adversarially after revival: eight independent lenses (lifecycle, mailbox, send, worker, federation, transcript, complexity, live ergonomics), each required to prove findings with a failing test. That produced 16 proven blockers, all fixed with red-then-green regression tests, followed by two re-review rounds and a third fix wave that caught 3 regressions introduced by the fixes and 7 fixes that missed their target; all closed. A final pass (five lenses incl. a live built-runtime smoke, then a re-review of the fix wave) found and fixed seven more, chiefly the stale-worker mailbox steal, the self-dispatch depth wedge, and the unproven-exit certification. Three independent Codex (gpt-6-astra) passes followed: the first found nothing new, the second found and fixed 3 defects (task-status reachability, WSL-local host classification, peer-capability epoch), the third found and fixed 6 (production PTY controller never installed settled writes, ambiguous in-flight pointer failures allowed duplicate replay, SSH/relay deadlines cut off a valid `--wait-submit`, stop-vs-exit race during inspection, and two release-recovery paths for vanished or exited terminals). The full record (findings, proof tests, triage, declines with reasons) is archived outside the repo.
**Rework after the live smoke.** A first live cross-host run on the shipped adhoc build (this Mac, a paired Windows host on the same build, a paired Mac on 1.4.195, and an SSH host) found a P1: a running local worker read `unverifiable`/`missing_status` because the fleet snapshot rows lacked the terminal handle the matcher keyed on. A 59-row failure table over every bug fixed during review showed the same two classes recurring: a fact dropped in transit through optional fields, and two authorities for one fact. Two blind designs (Opus, Codex) converged on the same mechanisms, and the scoped tranches landed here with red-then-green seam tests from the real producer to the real consumer, faults injected only at the transport or hook-ingest boundary:
- **Settlement (data-loss class):** one three-valued `WriteSettlement` (`accepted | refused{reason} | unverifiable{reason, bytesHandedToTransport}`) from the SSH multiplexer through daemon client, providers, controller, to pointer staging. No boolean, no rejection-as-third-state. The two silent degrades that fabricated a handoff are deleted; a provider that cannot settle refuses before any effect. Pointer text and Enter share the contract; a partial flush is `unverifiable`, never `refused`.
- **Evidence identity (false-liveness class):** fleet agent-status evidence is a tagged union (`binding: worker | pane | unresolved{reason}`, `clock: observed | delivery`) minted once at ingest, so a hook row captured on one process incarnation can never bind to a later dispatch on the same pane. The matcher's `!worker.paneKey ||` defaults are gone. One host-scope parser replaces two.
- **Small pre-merge items:** `capability_unsupported` from an old peer is no longer relabelled `host_unavailable`; a producer census test asserts every agent-status consumer path projects a pane-only hook row as `live`.
Two ergonomics defects the second live run surfaced on a real database are fixed here too: a pre-v3 dispatch already marked `completed` projected as `outcome_unknown` / `requiresAction: true` forever (three copies of the outcome ladder disagreed on legacy rows; now one resolver, legacy `completed` reads `succeeded` with nothing to act on, legacy `failed` stays actionable on the failure), and an unscoped `worker-list` enumerated the entire database (now defaults to the Run bound to the calling terminal, `--run` overrides, and the receipt's additive `scope` field says which).
A third live round on the shipped adhoc build of `b082443e1f` (same four hosts) plus an unscripted run in the user's own prompt style (a plain Claude Code shell, `/orchestration`, three workers, zero errors, bound-Run default confirmed) found two more branch defects, fixed with red-then-green tests: a worker freshly started on a paired server projected `unverifiable`/`host_indeterminate` with `requiresAction` for ~3 minutes, including after its own `worker_done`, because the host's federation observation returned `missing_liveness_verdict` for any PTY the liveness register had not yet swept (the host now reads a connected pane it owns locally as `live`; disconnected or SSH-scoped panes stay `unverifiable`); and six pre-v3 completed rows still carried an `input` category because settling through the task-status path or `failDispatch` never closed the Dispatch's pending question threads (both paths close them now, and schema v38 closes threads already pending on settled rows). The guide's `worker-start` examples now show `--model sonnet`, since an omitted model inherits the launcher's default.
A Codex adversarial pass on the tranche diff found one real design hole (identity minted at read time instead of ingest, now closed) and two daemon settlement paths that threw instead of settling (fixed). Two `@ts-nocheck` runtime mixins on these paths were extracted into checked modules; the repo-wide `@ts-nocheck` count is unchanged at 171.
Deletions during review: ~1,900 lines (write-only ledger, unread columns, dead v1 archive path, test harnesses shipped in prod, duplicated liveness and state-machine copies, self-capability checks that were compile-time true).
## Testing
- `pnpm typecheck:tsc:node|cli|web` clean
- `pnpm run check:code-quality:changed` 0 findings; `check:react-doctor:changed` 0
- `pnpm verify:bundled-skill-guides`, `verify:skill-bundle-manifest`
- full `pnpm test` on the integrated head: 72,332 pass / 292 skipped; the only failures were three non-PR files (two zsh live-shell suites hit a node-pty spawn-helper ENOENT while a concurrent native rebuild ran, 44/44 in isolation; `release-checkout.unit.test.ts` is a known 30 s load timeout that passes in isolation on `origin/main` too).
- CI on
|
||
|
|
6aa0aaee6b |
test: isolate source-control generation repositories per scenario (#19105)
* test: isolate source control generation repositories per scenario * test: explain scenario repository fixture scope |
||
|
|
b44aaf20c6 |
test: reuse authoritative SSH connection readiness in localhost fixture (#19102)
* test: reuse authoritative SSH connection readiness in localhost fixture * test: retain localhost SSH setup diagnostics |
||
|
|
b459b8f16d |
test: repair nested SSH fixture after HUB restart (#19098)
* test: restore paired nested SSH fixture after HUB restart * test: cover failed re-pair selection and background window safety * test: use required braces in re-pair regression fixture * test: use current paired runtime identity after re-pairing |
||
|
|
9837adaa07 | test: reconnect after replacing same-ID runtime pairing (#19094) | ||
|
|
ec64df335e | test: reject unsupported app-server in WSL golden stub (#19062) | ||
|
|
85c7696427 | test: exercise supported ConPTY keyboard protocol reset (#19054) | ||
|
|
a37a0b50d1 |
test: await fresh inventory after headless terminal materialization (#19028)
* test: restore headless folder terminal materialization coverage * test: await a fresh terminal census after materialization |
||
|
|
8f97048d60 |
fix(tests): complete hidden SSH dialog exits during cleanup (#18993)
* test: await nested SSH dialog exit before further dismissal * test: wait for the dismissed SSH dialog identity * test: wait for picker Back to reveal the reused host form * validation: keep hidden E2E compositor frames active * test: extract hidden Electron compositor setup * test: complete hidden dialog exit animations without global throttling changes |
||
|
|
f811ee0740 |
Open open new link should not navigate away from current link (#18873)
* fix(browser): open modifier-click and middle-click links in background t Links opened with modifier keys (Cmd/Ctrl+click) and middle-click now open in background tabs, matching Chrome's behavior. Shift+middle-click continues to open in the foreground tab. The routing system now tracks separate foreground and background frame names, with an `activate` flag controlling whether the new tab is brought to focus. * fix(browser): don't navigate away when opening links in background tabs When opening links via context menu or other mechanisms that create background tabs, keep focus on the source tab rather than automatically switching to the newly opened tab. Set `activate: false` on tab creation to prevent unwanted navigation away from the current page. * fix(browser): silence popup notices for links opened in Orca tabs Links that open in new Orca tabs are immediately visible to the user and don't warrant a toast notification. Only external popup opens now show notifications, reducing unnecessary clutter while still alerting the user to unexpected external window opens. * Replace loading dots with animated spinner icons Replaces the small dot indicators with animated Loader2 icons that appear in place of the favicon while tabs are loading. Provides clearer, more prominent visual feedback during navigation. * test(browser-tab): verify target=_blank links don't navigate source tab Add a test case checking that plain main-frame target=_blank clicks open in a new tab without navigating the source tab away. Extract startBrowserLinkServer to a helper module and add the /blank-destination endpoint to support the new test case. * refactor(browser): localize clicked-link routing frame names Remove the global clickedLinkFrameNamesByGuestId state map and generate frame names locally within installGuestPopupPolicy, improving state encapsulation and simplifying cleanup logic. Functionality unchanged. * test(browser-tab): hold shift for middle-click gestures * test(browser-tab): drop duplicate shift-middle gesture * test(browser-favicon): verify spinner shown while favicon reloads Updated test expectations to reflect that the favicon component shows a loading spinner during reload instead of keeping the previous image mounted. * fix ci |