mirror of
https://github.com/stablyai/orca.git
synced 2026-10-04 16:02:08 +00:00
a0d36f5290a6728458e003f48efe271fc2f6b886
1016
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a0d36f5290 |
fix(agents): keep OMP identity and forward ask/approval events (#15713)
* fix(agents): keep OMP identity and forward ask/approval events (STA-4130) A live OMP pane was re-owned as Pi because the generic pi-compatible fallback always won, ask events blocked without a question payload, and OMP suppressed tool_approval_* unless an extension registered handlers. Mark Pi as the title-group fallback so a specific OMP identity is not downgraded, publish OMP ask input as the existing questions envelope, and forward tool_approval_requested/resolved onto blocked/working. STA-4130 Related to #14278 Co-authored-by: devatnull <59279509+devatnull@users.noreply.github.com> * fix(agents): keep launch Pi ownership over OMP wrapper frames (STA-4130) The pi-compatible fallback treated every generic Pi owner as inferred, so an explicit launch-Pi pane (and a launchless Pi pane with an OMP-shaped title) was re-owned as OMP. Launch provenance now stays authoritative; only an inferred status-frame owner yields to a specific sibling, and same-group titles no longer count as reuse. STA-4130 * fix(agents): drop Pi wrapper idle titles while OMP hook is active (STA-4130) Title-completion suppression compared pick-a-winner ownership, so a Pi ready frame looked like a different agent than a live OMP hook and fired a spurious task-complete notification. Reuse checks now use the title-identity group. STA-4130 * fix(agents): restore OMP approval forwarding after merge * fix(agents): restore title-owner API after merge * test(agents): update identity inventory ratchet --------- Co-authored-by: devatnull <59279509+devatnull@users.noreply.github.com> Co-authored-by: Merge Sim <sim@local> |
||
|
|
cc6b600e21 |
Fix orchestration CLI recovery, settled-Dispatch mail, and guide defects (#16919)
* Fix orchestration CLI recovery, settled-Dispatch mail, and guide defects Five reported orchestration CLI defects, verified individually before fixing. Two were real code defects, one was a docs error, one was correct as-is, and one was correct on both ends except for its recovery wording. - Mail addressed to a settled `dispatch:<id>` was accepted and silently dropped. Local sends bypassed the settlement check the federated branch already had, so the caller was told success for a delivery no worker would ever read. Reject with `dispatch_inactive` and name the Run mailbox to use instead. - A lost mutation response offered no read-only way to ask whether it took effect. `--retry-request` does dedupe correctly, but the recovery guidance emitted a query command only when the payload carried a dispatch id, which is exactly what a lost response lacks. Add read-only `orca orchestration request-show --request <id>` over the durable receipt ledger, and always emit a read-only step before the keyed retry. - The bundled `orca-cli` guide documented `check --unread --inject`, a flag the parser rejects. Correct it to `--format` and add a ratchet that runs every orchestration invocation in the bundled guides through the real CLI parser. - `check --json` is one stdout document and its keepalives are stderr-only; the reported `Extra data: line 2` came from merging the streams. Document the contract rather than changing the wire. - A rejected lifecycle message is loud on both ends already, but the rejection never named the flag that supplies the missing capability. Name it. * Harden orchestration mutation recovery guidance |
||
|
|
b8d5b0486e |
Fix opening HTTPS URLs from headless runtimes (#17467)
* fix(runtime): open headless browser URLs on paired client * fix(pty): tolerate runtimes without browser relay probe * fix(browser): require automation-capable client host * fix(browser): map client URL opener in sidecar --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
5bd66bac8b |
fix(cli): resolve a WSL worktree by the Linux path its own shell prints (#16628) (#17440)
On a Windows host the runtime stores a WSL worktree as the UNC path Windows sees, but a user inside the distro types the Linux spelling, so every `path:` selector missed: `worktree show`, `terminal list --worktree` and `worktree rm --worktree` all reported selector_not_found for a directory Orca manages. Translate once in the CLI, which is the only side that can prove which distro the typed path belongs to — from its own UNC cwd, never from WSL_DISTRO_NAME, which a Linux-native CLI also sets. The runtime's `path:` branch stays exact-spelling-only for the same reason: this resolver feeds delete, so a tail-only match would remove another distro's copy. |
||
|
|
1268fb56f1 |
fix(worktree): complete a create Git can confirm but cannot list (#17388)
* fix(worktree): complete a create Git can confirm but cannot list `worktree.create` verified against `listWorktrees`, which softens every git failure to `[]`. Any listing failure therefore failed a create whose worktree and branch `git worktree add` had already written, orphaning both, and reported only 'Worktree created but not found in listing' — the real cause reached the main-process console and never the user. Verify against the error-propagating listing instead, and when that fails or omits the row, rebuild the row by asking Git about the worktree itself. The direct read returns nothing unless Git resolves the path into this repo's object store with the expected branch checked out, so an unrelated or half-made checkout still fails the create. Fixes #16520 * fix(worktree): authorize a recovered create and reject an unreadable HEAD Review follow-ups on the create-verification fallback: - register the recovered worktree's own root, additively, so the create the user just made is not rejected by filesystem/git-status IPC - treat an unreadable HEAD as no recovery instead of a blank OID - keep the direct read's failure when the listing merely omitted the row - skip the symlink cases on Windows and reset the new harness mock * fix(worktree): bound the create-recovery disk read and keep WSL paths case-sensitive Readiness-scan follow-ups: - deadline the filesystem common-dir read; a .git on a hung mount left the whole create IPC pending where it used to fail after the Git deadline - offer no disk candidate for a bare repo instead of a fabricated <repo>/.git - compare POSIX common dirs case-sensitively, so two WSL repos differing only in case are not accepted as one object store on a Windows desktop - move toGitOutputSpace to shared/wsl-paths as toWslExecutionSpace, next to the parseWslUncPath callers that already open-code it * fix(worktree): share one budget for create verification and keep recovered roots Three follow-ups from review of the create-recovery path: - The recovery no longer starts a fresh 30s deadline after the listing already burned one, so worst-case create verification stays at ~30s instead of ~60s. A 5s floor keeps the direct read a chance to answer when the listing spent the whole budget. - rebuildAuthorizedRootsCache now carries a repo's previously registered roots forward when its listing throws. A rebuild running while Git is still broken could otherwise un-authorize the worktree a create just recovered. - Corrected the scan-cache doc comment: it claimed strict and lenient listings coalesce, but the cache key includes the runner name precisely to keep them apart, so a strict joiner can never inherit a lenient scan's softened []. Each change has a negative control: reverting the hunk fails exactly its own test and nothing else. * fix(worktree): keep a recovered worktree authorized across roots-cache rebuilds The previous approach registered a recovered create into the same per-repo set the rebuild recomputes from `git worktree list`. That set is derived from the very listing that failed, so a rebuild would re-deny the worktree — either by overlapping the registration, or by simply listing again and omitting the row. Carrying old roots forward on a thrown listing did not cover either case. Recovered roots now live in their own additive layer that rebuilds union in rather than replace. The layer is retired on evidence, not on a timer: - the listing can see the worktree again (Git recovered), or - the listing succeeded and the directory is gone (worktree removed). A repo whose listing threw is left untouched, because a dead mount fails both the listing and the stat, and treating that as "removed" would revoke the worktree in exactly the outage this layer exists for. The layer is capped so it cannot grow unbounded, and survives cache invalidation deliberately: repo mutations are frequent and would otherwise re-deny a recovered worktree. Three tests cover the healthy-rebuild-omits-the-row case, the in-flight rebuild race, and retirement once the listing sees it again. Removing the union fails exactly the two keep-tests and nothing else. * perf(worktree): only read the repo's .git from disk when Git's own answer disagrees The disk read is a second opinion on Git's reading of the common dir, but it ran unconditionally as part of the same Promise.all. A deadline bounds the IPC, not the syscall: Promise.race cannot cancel an in-flight fs operation, and a `.git` on a hung mount (dead NFS/SSHFS, stalled WSL 9p) pins a libuv threadpool thread that no timeout can reclaim. AbortSignal would not help either — fsPromises.stat takes no signal, and a blocked syscall is not interruptible from userland. So stop paying it on the happy path: read from disk only when Git's own reading did not already confirm the common dir. Same accept/reject outcome, but the threadpool exposure now requires both a failed listing and Git disagreeing about the repo, instead of every recovered create. * fix(worktree): compare the disk common-dir witness in Git's execution space Exercising the fix on a real Windows host against WSL Ubuntu-24.04 found the filesystem second opinion is inert there. Node reads `.git` in the caller's space and answers `\\wsl.localhost\<Distro>\home\...\.git`, while Git-in-the- distro answers `/home/...`. isSameCommonDirPath refuses to compare a POSIX path against a Windows one, and canonicalizeLocalPath cannot bridge them because realpath on a Linux path from a Windows process is ENOENT. So the candidate could never match, and the one case that depends on this witness alone — a symlinked repo root on the Git 2.25 fallback — declined a worktree Git had already confirmed. Run the disk result through toWslExecutionSpace, the same translation readRepoLocation already uses. This is a false reject, not a false accept: it made recovery give up, never adopt the wrong repo. Verified on awin; the modern --path-format=absolute branch was unaffected because Git answers both sides itself there. * fix(worktree): retire a recovered root only on proof, never on a stalled probe The prune ran an unbounded stat and read every failure as removal. Two consequences, both in the outage the recovered layer exists for: a hung mount stalled the rebuild that gates filesystem auth, and a transient EACCES/EIO revoked a live worktree. The listingFailed guard did not cover either, because listWorktrees softens Git failures to [] and never throws. Prune now retires on definitive ENOENT only, probes in parallel under a deadline, and treats a stall as inconclusive. The capacity bound refuses a new root instead of evicting an authorized one, so an over-cap create is merely unauthorized rather than a live worktree being revoked. |
||
|
|
cad9206839 |
fix(native-chat): preserve structured chat across rollback (#17439)
* fix(native-chat): preserve structured tabs across rollback * fix(native-chat): preserve rollback visibility state --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
b5a85890ac |
perf(git): bound git subprocess execution with an atomic admission scheduler (#16874)
* perf(git): bound git subprocess execution with an atomic admission scheduler Field traces (#16038, #11363) show Windows freeze storms driven by unbounded concurrent git children (12+ at once, 50-65s status convoys for 25+ minutes). Admit every main-process git child against atomic per-budget base+headroom counters (general / network / per-route), with reserved interactive capacity, ordering-only aging, close-bound permit release, a 120s fail-safe read timeout that feeds scheduler backoff, tier plumbing through every option carrier, and coalesced+jittered visibility pollers. Killswitch: ORCA_GIT_ADMISSION_DISABLED=1. Storm harness A/B: max concurrent children 65 -> 6, interactive p95 791ms -> 88ms; output-parity battery byte-identical with admission on vs off. * test(git): run the admission output-parity battery on every platform Parity needs real git, not the storm harness's PATH stub, so it must not share that file's POSIX gate - Windows is the platform where parity evidence matters. * fix(git): preserve interactive admission invariants * perf(git): keep admission queue drains linear * fix(git): close final admission gaps * perf(git): bound eligible route selection * fix(merge): remove unrelated stale snapshot changes * fix(git): preserve refresh lifecycle authority * test(git): align admission lifetime contracts * fix(git): harden admission across runtime paths * fix(git): restore freshness for bulk status reads * test(git): repoint delete-dialog source pins after admission plumbing The hydration effect now orders its targets through orderDeleteWorktreeStatusHydrationTargets and passes includeLineStats alongside the abort signal, so both literal anchors stopped matching. The invariants are unchanged and still pinned: dropping the signal, the main-worktree/folder filter, or getState-instead-of-subscribe each still reddens this test. * Fix git admission tier propagation and lock ordering Decode optional Git status tiers permissively and default runtime RPC status reads to the status lane while preserving renderer caller intent. Acquire the FETCH_HEAD mutex before atomic admission so same-repository fetch waiters hold no global or route permits. Preserve automatic pull-request refresh reasons, keep explicit hosted-review refreshes interactive, remove the dead candidate tier, and keep relay scheduling unchanged. Use tier-aware status lease keys because a shared lease cannot be safely promoted after its admission request is queued or granted. * test: align expectations with admission plumbing * refactor(child-process): move the process contract types to process-spec run-process.ts crossed its line cap after gaining the termination observer; the public types and defaults move out with re-exports so no caller changes. * chore: restore pnpm-lock.yaml to main (unintended local drift) --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
c3aceacc7b |
Fix PR unlink for auto-detected reviews (#16898)
* fix: make PR unlink hide auto-detected reviews * Type the empty-content test double against the real model The literal narrowed suppressedGitHubPR to number and typed the callback as Mock, so neither direction was comparable and tsconfig.tc.web.json failed on TS2352. Keeping the 'as' cast preserves checking of the fields the double does supply. * Add localization keys for the unlinked checks-panel state The unlinked title, relink action, and the remote-runtime upgrade notice introduced untranslated keys that static analysis requires in en.json. * Advertise PR suppression capability in the transport test The client capability list is pinned by websocket-transport.test.ts, and adding WORKTREE_GITHUB_PR_SUPPRESSION left the expected list stale. * Fix stale PR suppression in Checks * fix: harden PR unlink suppression state * refactor: extract PR unlink state handling * fix: show PR relink recovery in source control * fix: add unlinked PR localization * Clarify workspace-scoped PR unlinking --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
3ab9766e38 |
perf(worktree): prepare checkouts while the composer is open
Squashed merge of PR #17290. |
||
|
|
a8183884bd |
perf(wsl): place worktrees inside the distro when the project runs in WSL
Fix-forward for readiness review: align retirement placement with WSL mirrors and preserve Windows-side git-common watchers. |
||
|
|
ac02232015 |
perf: overlap independent worktree create preflight (#17386)
Readiness checklist passed; required CI and review checks are green. |
||
|
|
6677ae4e5e | test: correct 8 stale specs surfaced by the test-detected-bugs sweep (#17434) | ||
|
|
d9870c6c75 |
fix(browser): apply the app-wide HTTP proxy to embedded browser sessions (#15536)
* fix(browser): apply the app-wide HTTP proxy to embedded browser sessions The proxy setting was only ever written to `session.defaultSession`, but browser guests run on their own `persist:orca-*` partitions. Any host reachable only via the configured proxy failed to load in an embedded tab, landing on `chrome-error://chromewebdata/`, while the same setting worked everywhere else. Adds a per-session applier alongside the existing defaultSession path, keyed by a WeakMap so one session's applied config can't suppress another's, and applies it to every browser partition through the single installer they all pass through. Startup awaits an explicit sweep so the first guest navigation can't race the installer's fire-and-forget write, and a settings change re-sweeps so toggling the proxy takes effect without a restart. Env-var fallback and the system-proxy probe mirror the defaultSession behaviour, so a browser partition resolves the proxy the same way the rest of the app does. Fixes STA-4779 * fix(browser): await per-session proxy readiness * fix(proxy): preserve loopback and authenticate * fix(proxy): settle browser partition update races * fix(proxy): close partition policy races * fix(proxy): order settings and release removed sessions * test(browser): await partition proxy readiness * fix(proxy): cancel removed partition retries * refactor(proxy): keep OpenCode rate limits out of scope * fix(proxy): preserve sessionless host policy * fix(proxy): gate requests on policy readiness * fix(proxy): retire deleted browser sessions * fix(proxy): close retired browser guests * fix(proxy): retain retired session guards * fix(proxy): retain retired partition policies * fix(browser): retry transient proxy application failures * fix(browser): release deleted partition installer state * fix(proxy): retry delayed transient failures * fix(proxy): preserve route session authority after rebase * fix(proxy): clear retired session credentials * fix(proxy): retire failed browser profiles * fix(proxy): harden failed session cleanup --------- Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> |
||
|
|
3af2c665c0 |
fix(cli): name PowerShell when it strips quotes from JSON flags (#17351)
* fix(cli): name PowerShell when it strips quotes from JSON flags Windows PowerShell 5.1 does not escape inner quotes when building a native command line, so `--options '["a","b"]'` reaches orca.exe as `--options [a,b]`. The value is correct when printed and damaged by the time argv is parsed, so the resulting "invalid JSON" error blamed the user's input rather than the shell. #16743 recovered this for `--deps`, which is safe only because generated task IDs have a fixed 12-hex grammar. The same mangling hits `--options`, `--payload` and `--result`, and those are NOT safely recoverable: `["1","2"]` and `[1,2]` arrive at argv identically, so a general repair would silently turn strings into numbers. Detect instead. `getOptionalJsonFlag` rejects the damaged shape up front with an error that names the shell and shows the workaround. It fires only when the value is bracketed, quote-free, fails JSON.parse, AND consists entirely of bare tokens that quoting would rescue, so valid JSON is untouched. Also share the generated-id contract: `task-deps-flag` hardcoded /^task_[0-9a-f]{12}$/i, which silently diverges if `generateId`'s byte count changes. It now calls `isGeneratedId`, with a test pinning the two together. Verified on a Windows host. Measured argv, which the new test pins as a fixture: PS_VALUE=["task_b2a580db74d8","task_c3b691ec85e9"] ARGV=["--deps","[task_b2a580db74d8,task_c3b691ec85e9]"] Before: Invalid --options: must be a JSON array of strings After: --options arrived as [a,b], which is not valid JSON. Windows PowerShell 5.1 strips the inner quotes ... * fix(cli): scope JSON-flag detection to genuinely JSON flags Review found the detector wired to two flags that are not JSON: - `orchestration ask --options` is documented `<csv>` and the runtime splits it on commas, so `--options [a,b]` was a legitimate value being rejected. - `task-update --result` is stored verbatim and reused as dispatch failure text; existing tests pass free text, so a bracketed `[ok]` was being rejected. Both revert to `getOptionalStringFlag`. Only `gate-create --options` (`<json_array>`) and `send --payload` (`<json>`) are JSON-parsed and keep it. Three further review fixes: - Objects now require a `key:value` pair per entry. `{a,b}` and `{a:b,c}` were reported as quote-stripped although quoting them cannot produce valid JSON. - The raw value is no longer echoed. A `--payload` can carry secrets and this message reaches `--json` output; the flag name and guidance are enough. - The message hedges the shell attribution. Detection inspects only the value's shape, so it also fires when a macOS/Linux user forgets to quote, where PowerShell is not involved. Verified against a Windows host, all six cases: both JSON flags fire on the mangled shape and pass valid JSON through to the runtime; both non-JSON flags now reach the runtime again; and the secret in `{token:hunter2}` appears zero times in the error output. |
||
|
|
4bc2085271 | Revert "perf(rpc): compile Zod request schemas lazily" (#17368) | ||
|
|
7b86833120 |
perf(rpc): compile Zod request schemas lazily (#17353)
* perf(rpc): compile Zod request schemas lazily * test: align window reveal assertion |
||
|
|
252dbd60ea |
fix(terminal): restore lossy initial remote snapshots (#17113)
* fix(terminal): restore lossy initial remote snapshots * test(terminal): strengthen lossy snapshot causal oracle |
||
|
|
ae0f3675a1 |
fix(remote): focus host-delegated split panes (#16886)
* fix(remote): focus host-delegated split panes Return the authoritative leaf identity from terminal.split, record viewer-local focus intent behind the captured pairing revision, and replay the mirrored layout before focusing the exact pane. Preserve old-host fallback and prevent delayed split responses from stealing focus after the viewer moves away. Add deterministic runtime, renderer, concurrency, compatibility, and headed paired-Electron coverage for Cmd+D, header splits, and immediate PTY input routing. Fixes #16510 * fix(remote): preserve split focus across tab groups Resolve the initiating source tab and leaf from the remote PTY, while keeping the viewer's current focus as a separate anti-steal baseline. This lets context-menu/header splits from non-focused group tabs focus their result without allowing delayed responses to override a later navigation. * test(remote): drive split focus with key events * test(remote): use the platform split shortcut * fix(remote): fence concurrent split focus intent * fix(remote): harden split focus ordering * fix(remote): preserve split focus after runtime refactor * fix(remote): fence stale split focus gestures * test(remote): keep split focus regression within line budget |
||
|
|
07df4bf0be | fix(orchestration): recover stripped task deps | ||
|
|
47827b7539 | fix(add-project): use the entered group name when opening a folder (#16881) | ||
|
|
86d9c07a3c |
Split runtime RPC server responsibilities (#17270)
* Split speech session lifecycle * Split terminal output scheduler pipeline * Split mobile browser pane modules * Prune resolved max-lines suppressions * Split pane tree equalization logic * Extract mobile troubleshoot screen styles * Split external automation manager * Split main window service attachments * Split hosted review creation checks * Split automation dispatch event handling * Split settings navigation metadata * Split daemon initialization lifecycle * Split GitLab item dialog * Split relay dispatcher layers * Split mobile host screen * Retarget mobile view settings source test * Split runtime file client layers * Split ports panel layers * Split runtime environments pane layers * Split local PTY provider responsibilities * Split CDP bridge responsibilities * Split relay Git handler responsibilities * Track moved relay Git fetch audit * Split Linear item drawer responsibilities * Split telemetry event schema responsibilities * Split resource usage status responsibilities * Split remote terminal multiplexer responsibilities * Split Git worktree responsibilities * Split Codex hook service responsibilities * Keep mirrored hook trust type private * Split web runtime session responsibilities * Split GitHub project view read path * Split Claude runtime auth responsibilities * Split runtime RPC server responsibilities * Fix F3-speech for #17123 * Fix F1-cycle for #17131 * Fix F4-navtest for #17157 * Fix F2-allowlist for #17161 |
||
|
|
2dfaa676d8 | chore: update oxlint and oxfmt (#17150) | ||
|
|
8d3e32a2ff |
Fix setup-provisioned skills missing at agent startup (#17124)
* fix(setup): let repos gate agent startup * test(setup): update runner call expectations |
||
|
|
fd9125ea8c |
feat(native-chat): Codex structured native chat restructure (#16729)
* feat(native-chat): port structured Codex sessions from restructure-recovery Rebuilds the desktop structured native-chat implementation from brennanb2025/native-chat-restructure-recovery (tip 4e31c08db3) on top of current main as a single commit, scoped to the local Codex path. Ported: - Structured agent-session core: durable record store + single-writer lease, canonical journal, agent-session wire host/attach/eviction/subscribers, `agentSession.*` RPC surface (registered via ALL_RPC_METHODS; host-side mobile allowlist included for wire compat), pty write gate, transcript additions, and the Codex app-server adapter/launch resolution. - Renderer: NativeChatStructuredSession view/composer stack, structured launch path with the single-flight guard, local structured session tabs sync, activation gate + structured inventory (read-only `agentSession.handoffStatus` probe), agent-session tabs in the tab strip, AI-vault structured session activation, and the settings pane with the parent Experimental Chat UI toggle plus the nested "Use updated structured native chat" toggle. New sessions require both flags, agent codex, no prompt, and a local non-WSL, non-Windows-host execution host (structured-native-chat-availability). - Fixes 72c013cea6 (verified Codex launch recovery), 8ddbaf5e3d (defer native terminal view switching affordances), and 4e31c08db3 (release the launch gate after a visibility retry) with their regression tests, including the third-launch-after-retry guard case. - Cross-version agent-session wire test + CI lane, packaging entries (proper-lockfile, agent-tooling asar excludes), and the wire-compat doc section. Deliberately not ported: mobile/ changes, the Claude structured runtime (only the claude-transcript-branch-proof and claude-structured-owner-identity leaf modules remain, backing the kept TUI-recovery arms), the terminal↔chat adoption/handoff flow (`agentSession.adoptTerminal`/`requestHandoff`, the handoff request engine, TUI adoption machinery, orca-runtime adoption methods), renderer switching affordances and their dead leftovers, the hook/subagent-status refactor cluster, and unrelated branch changes. The crash-during-acquisition recovery path (restart handoff adjudication, restore/reverse re-acquire, lease schema handoff keys) is kept because every plain direct launch depends on it; a trimmed handoff coordinator exposes only status/restore/close. Branch edits that targeted files main has since split (ipc/pty.ts, worktrees.ts, rpc/methods/terminal.ts, useIpcEvents, pty-connection, store/slices/terminals.ts, runtime-types, web preload) were re-applied to the split modules, preserving main's newer logic (Windows CIM fallback, browser tab close rework, cold-restore resume flow, dispatcher threading). Known seam: the mobile clipboard image-provenance CONSUMER gate ships (agentSession.send refuses unproven mobile image refs with agent_session_image_untrusted) but the producer hunk in rpc/methods/clipboard.ts stays with the unported mobile cluster, so mobile image sends into structured chat fail closed until that side ports. * fix(native-chat): trust only authenticated local image uploads * fix(build): preserve Windows process-tree patch application * test(windows): include process creation time in addon fixture * fix(build): run windows-process-tree node-gyp from the physical package dir gyp expands the node-addon-api dependency by probing node, whose cwd resolves to the package's physical directory in the store, so the emitted target is a store-relative ../../../../node-addon-api@... hop. gyp then resolves that hop against the rebuild cwd; from the node_modules symlink/junction it escapes the store and configure fails with "node_addon_api.gyp not found" (run 32999886072). Rebuild from realpath(package dir) so both bases agree, matching how the package manager itself runs native install scripts. The regression test replays gyp's expansion+resolution against the planned cwd and fails without the fix. * fix(native-chat): keep chat tabs visible through terminal closes and empty-worktree launches Two proven blockers in the native Codex tab contract: closeTerminalTab pre-empted the canonical unified close. With one terminal left it deactivated the worktree on a terminal/editor/browser-only check, blanking a workspace that still held a renderable agent-session tab; with two or more it pre-picked a successor from terminal entities only, re-stamping the group active before closeUnifiedTab's MRU/neighbor repair could land on the chat tab. Successor choice now defers to the unified contract whenever the terminal has a unified row, and deactivation is gated on the unified renderable count (matching leaveWorktreeIfEmpty), with the legacy pre-pick kept only for terminals without a unified row. A structured session created on an empty worktree was published into the host's headless group while preserveLocalLayout froze the local layout, leaving the tab in store but permanently off screen. A preserveLocalLayout owner now always takes client-owned placement — repairing a rendered leaf whose group record is missing, or materializing a rendered group on a truly empty worktree — and applies the client-derived layout repair while still rejecting host-authored layout. Regression tests drive the real store through closeTerminalTab (git worktree and folder workspace) and the real snapshot applier for the empty-worktree adoption states; all fail without the fixes. * fix(native-chat): close stale turns and retry rejected sends * fix(native-chat): retire hosted rows on structured tab activation * fix(native-chat): preserve rpc defaults across main merge * chore: format remote wire compatibility guide * test(native-chat): cover retry after unconfirmed send * fix(native-chat): reload outbox on session switch * docs(settings): disclose structured chat platform limits * fix(native-chat): await Codex launch-home preparation * fix(codex): align child-process allowlist with async trust bridge * test(identity): update inventory for tab surface refactor * fix(windows): preserve process-tree CRLF patch sources * fix(native-chat): anchor an unmatched chat echo where it was sent (#16117) * fix(native-chat): anchor an unmatched chat echo where it was sent The reported symptom was old user messages replaying below every new turn, so the conversation read as scrambled. The cause was not that the echo failed to match a transcript row. Claude consumes a mid-turn send through a `queued_command` attachment and writes no `type:"user"` record for it, so some echoes can never match, and no amount of matching will change that. The cause was WHERE an unmatched echo rendered: buildMobileNativeChatTransientData appended every pending item after the entire transcript, so it re-read below each turn that landed afterwards. Render each echo directly after the transcript row it was sent against, using the baseline the send already captures. An unmatched echo is then at worst a duplicate in the right position rather than a scrambled one, and it stays visible. Echoes sharing an anchor keep send order; a send with no baseline, or one whose anchor folding dropped, still falls back to the tail. Deliberately NOT fixed by deleting the echo. Inferring from send ordering that an echo can never match, then removing it, loses the user's own text for a message the agent did receive, and it cannot fire in the common case anyway - measured drain groups are 1,017 of size 1 against 55 larger. It also escalates an existing gap: the count pass has no baseline-tail guard, unlike the glue pass, while `messages` is a 40-row window that head-trims, resets on reconnect and grows at the front on loadEarlier, so a false landing there would license deleting a DIFFERENT outstanding message. That count-pass gap is real and left for a separate change; anchoring makes its worst case a duplicate in place rather than a scrambled conversation. * fix(native-chat): preserve folded echo anchors * fix(native-chat): preserve forward-folded echo anchors * fix(native-chat): keep leading folded echoes in place * fix(workspace-cleanup): show git status for every row (#16690) * fix(native-chat): refuse structured chat on every Windows execution path canUseStructuredNativeChat only refused win32 when a project runtime resolved, so folder-workspace keys (and other keys with no project runtime) failed open into structured chat on Windows. Fail closed on win32 unconditionally after the host check, matching the settings copy: local macOS/Linux only; Windows/WSL/SSH stay on terminal chat. * fix(native-chat): restore runtime refusals behind the win32 gate |
||
|
|
6256f3d137 | Fix lost session-tab changes during initial census (#17064) | ||
|
|
df95f03101 |
Show ready and close actions for draft reviews (#16889)
* Fix draft review sidebar actions
* Drop unused React import in draft actions test
The automatic JSX runtime makes the default React import dead, and
tsconfig.tc.web.json failed the branch on TS6133.
* Add localization keys for draft review actions
The new Ready for review controls introduced five untranslated keys and
the static analysis job requires them present in en.json.
* Name the draft action for what it does
The button read 'Ready for review', which states a status rather than an
action, directly under a header already showing the PR state. The i18n
key (markReady), the in-flight label ('Marking ready...') and the success
toast ('marked ready for review') all already used the verb.
|
||
|
|
caef20fec8 |
fix(gitlab): paginate TaskPage issues beyond 50 (#13538)
Co-authored-by: Neil <neil@stably.ai> |
||
|
|
fc8c981103 | fix(browser-preview): enforce canonical runtime grants (#16975) | ||
|
|
52ade074a9 |
Display host on automation details and dialog (#16823)
* Show automation host in details and support moving between hosts - Rename AutomationCreateDestinationField to AutomationDestinationField to reflect dual use in create and edit modes - Add host display to automation detail view, showing storage authority - Allow editing automations to move them to different hosts within same authority; project list filters to available projects on chosen host - Update copy from create-only terminology to mode-agnostic wording * Display automation host and support cross-authority moves Users can now move automations to different storage authorities. The destination picker shows all available hosts, and selecting a new one displays a warning about the move. The save creates the automation on the destination and deletes it from the source; if deletion fails, both copies remain and the user is notified. * Remove cross-authority move support for automations An automation's authority (the Orca instance that stores and schedules it) cannot change; edits now only offer hosts within the same authority and move logic is removed entirely. This simplifies the destination picker and removes move-specific UI messaging. * Fix undefined selectedRowKey in automation host recovery Replace references to the undefined selectedRowKey variable with selectedRow?.key to properly access the row's key when recovering automation runs across hosts. * Support moving automations across execution authorities Allows users to move automations between different authorities (desktop ↔ runtime environments) during editing. A save to a different authority creates the automation on the destination and deletes the original with its run history. Includes clear messaging about the move operation, proper handling of workspace id resets, and graceful error handling when deletion fails. Supports destination-aware project and worktree fetching. * Reuse creationKey across move retries when schedule changes When retrying a failed automation move, dtstart is minted fresh each attempt, changing the payload. Previously, operationKey included the full payload, so retries would mint new creationKeys and risk duplicate automations on the destination if the initial create failed in transport. Now key only by the move (source + destination) to ensure stable creationKey across retries. Also fix workspace auto-selection to use authority-scoped worktrees instead of the merged cache, preventing unwanted restoration of source-host workspaces after switching authorities. * Rename `note` to `moveWarning` for automation host moves Clarifies that the field specifically warns when an automation would move to another host, replacing the plain storage line. |
||
|
|
a7b1da9e14 |
fix(ssh): a reattach may bind a pane but must never create one (#16751)
`reattachKnownPtys` treats every non-terminated lease as live and calls `persistPtyBinding`, which had no way to say "bind only". Two of its branches then rebuild UI the user is not asking for: - `pty-binding-persistence.ts:133-143` -- `if (args.incarnationId)` unconditionally deletes the pane's close tombstone. - `:145-160` -- on a tab-not-found it mints one via `createMinimalPersistedTerminalTab`. The in-code comment names its only intended caller: "pty:spawn can beat the debounced writer." Spawn. Reattach took the same branch. This is the mechanism canceled ticket STA-4268 described: "Leases have no pane incarnation and upsert only by target/PTY. Every nonterminal lease is reattached; frozen coordinates are passed to persistPtyBinding, which creates and flushes missing tabs and layout leaves." It was fixed in #13326, reverted by #14361, re-applied by #14384, and reverted again by #14395 (opened and merged nine seconds apart, empty commit body). `grep -rn "mayCreate" src/` returns nothing on main -- the mechanism is genuinely out of the tree. Four parts: 1. `mayCreate` (default true). When false, one pre-mutation check mirrors all four creating branches and returns `false` without mutating, so a refusal leaves nothing half-written. 2. The authority gate -- the part both prior attempts lacked, and probably why both were reverted. `mayCreate: false` alone refuses in two situations: "the user closed it" AND "the renderer has not published its layout yet". The second is routine on disconnect->reconnect and is almost certainly the #14361 tab-loss mechanism. The fence was not wrong; it was UNCONDITIONED. It is now passed only when `hasHostAuthoritativeTerminalMembership()` says the persisted membership speaks for this worktree, reusing the function already guarding the same question at `orca-runtime.ts:8608`. Losing a tab is worse than keeping a duplicate, so an unauthoritative session still gets the creating write. Authority is read from `local` because that is the partition the write lands in -- it is local's absence being interpreted. But a pane the `ssh:<target>` partition still holds is not gone, so it keeps its creating write; refusing there would strand a live pane behind a binding reattach can no longer reach. (SSH spawns bind into `ssh:<target>` while this reattach binds into `local` -- GH #12721/#12723, STA-3980. This does not fix that split; it refuses to judge from one side of it.) 3. `findTerminalTabIdForLeaf` -- bind resolves the tab from the live layout instead of the lease's frozen `tabId`. Only the leaf half of a pane key is remint-stable: `detachTerminalPaneToTab` moves a live pane into a new tab, so a stored tabId names the tab the pane left. Identity vs location. 4. Pane-keyed supersession -- retires siblings on `(targetId, worktreeId, leafId)` to `expired`, guarded by the durable binding, which reads `local` then `ssh:<target>` so it is correct whichever partition the binding landed in. `upsertSshRemotePtyLease` matched `(targetId, ptyId)` alone (`ssh-pty-lease-operations.ts:34-36`), so a new relay pty id on reattach minted a SECOND lease instead of updating the first, leaving the predecessor non-terminated with nothing to retire it. On refusal the lease goes `expired`, never `terminated` -- `expired` records that this shell has no surface to reach it through; `terminated` would assert an exit nothing here observed, which `docs/reference/ssh-execution-boundary.md` forbids. The remote process is left running. Deliberately NOT done: - A collision guard for `upsertSshRemotePtyLease`. Built, tested, and REMOVED -- its own test passed with the guard disabled, i.e. vacuous. Telling "same lease" from "recycled id on a different shell" needs a relay-start identity, which would be a twelfth per-tab identity concept; the codebase already carries eleven (784 refs) that SSH-v3 Phase 3 deletes. Left as an in-code NOTE. This handles lease DIVERGENCE, not COLLISION. - A port of #13324. Its own authors deleted its load fold and reverted its local-only reader in #13326 ("a headless-owned pane still gets a vote before its lease is retired"); porting it ships a state-destroying migration they removed. - `bindPaneShell` from #13325. Its purpose is making `isSupersededPtyId` live, and that fence does not exist in main. It would have been a refactor plus a silently-ignored `mayCreate` -- TS drops excess props through spreads (verified), which is why `mayCreate` is passed as a conditional spread here. Honesty about scope: this does NOT close the daily-tab-growth report. The deterministic e2e repro (`ssh-lost-kill-tab-resurrection.spec.ts`, later in this stack) is byte-identical before and after, and instrumentation shows why -- in that scenario `restoreReattachedPtyRuntime` is never called at all (`CREATING TAB` 9 hits, `reattach gate` 0, `BYPASS` 0). The fence is on a path that bug does not take. It is a real, separately-provable defect; it is not the headline fix, and must not be claimed as one. Evidence, A/B on this tree. Disabling the authority gate (`mayCreate = true`) and the supersession call by hand: **8 of 12 fail**, including "does not resurrect a tab whose closing pty.kill failed with a transport error" and "holds the live lease count flat across ten reconnects of one pane" (`[ 'pty-0', 'pty-1', 'pty-2', …(7) ]` vs `[ 'pty-9' ]`). The 4 that pass both ways are the over-refusal tripwires, which is the point of having them. Restored: 12 passed; 2,439 passed across `src/main/ssh`, `src/main/persistence` and `src/main/runtime/workspace-session`. |
||
|
|
2ac23c29be |
Add parent worktree selection when creating workspace (#15420)
* feat: select parent worktree for nesting when creating workspace Enable users to pick a parent workspace in the composer's Advanced drawer, organizing newly created worktrees hierarchically in the sidebar. The backend validates lineage relationships and gracefully retries without the parent if it becomes unavailable during creation. * feat: select parent worktree for nesting when creating workspace - Record app-picked parents as manual actions, not CLI-flag equivalents, ensuring the same user action carries consistent cleanup semantics across hosts - Gracefully retry without parent if the selection goes stale, warn the user instead of failing - Refactor create into modules: parent resolution, payload building, state merge * Allow selecting parent worktree when creating workspace - New parent worktree picker filtered by execution host and project - Preserve concurrent local writes to lineage by comparing per-record state instead of key-set membership * Allow selecting parent worktree when creating workspace - Rename "Parent workspace" label to "Parent worktree" - Filter candidates by execution host and project to prevent nesting across hosts - Wire parentWorktreeId through composer state and creation request pipeline - Update translations and add new copy for nesting-related messages |
||
|
|
2aec442f42 | Split Linear issue service by operation (#16770) | ||
|
|
9a0a2b1c31 |
fix(orchestration): settle worker release without web layout state (#16842)
* fix(orchestration): tolerate missing terminal layout partitions * fix(orchestration): handle legacy release without layouts |
||
|
|
7ee8b5e1a6 | Refactor lower max-lines modules (#16760) | ||
|
|
bb77a2b159 |
Add coverage for skill lock release and fix WebRTC test flakiness (#16846)
* test: add coverage for skill lock release and simplify WebRTC test - Add test for cleanupReleasedSkillInstallLock handling rmdir races - Improve error handling to cover all documented directory removal error codes - Simplify WebRTC egress test to use localhost addresses consistently * test: use network interface address for WebRTC egress probe - Discover the first non-internal IPv4 address instead of hardcoding localhost, allowing the test to work in CI and varied environments - Update proxy rules to use loopback designation for clarity - Bind UDP socket to all interfaces (0.0.0.0) to receive on the discovered address |
||
|
|
81ae98e10d |
fix(mobile): honor host worktree create retention (#16342)
* fix(mobile): honor host worktree create retention * fix(mobile): cover malformed worktree retention policy * fix(mobile): fail closed on malformed retention policy * fix(mobile): fail closed on missing dedupe ttl |
||
|
|
6ba6d58cd2 |
fix(orchestration): route @agent messages by resolved identity, not terminal title (#16237)
* feat(agent-status): add the pane agent identity resolver
Four ladders answer "which agent is in this pane" independently — the tab icon, the
open-tab/search occupant, the sidebar title rows, and the sidebar hook-row fallback — and they
disagree. Two consult the terminal title before the launch record, so a string Orca parsed
outranks a fact Orca owns.
resolvePaneAgentIdentity is the single ranked answer. Two rules, one of which is not an ordering:
1. Evidence is ranked by how directly it observes the process; a display title is last.
2. Each observation carries the runId of the agent run it describes. Evidence from a superseded
run is INELIGIBLE, not merely outranked.
Rule 2 is the part reordering could never supply. A completed hook naming A plus a title naming
B is either a bug (hook right, title stale) or a legitimate pane reclaim (title right) —
identical signals, opposite correct answers. Run ids make them different facts: in the bug both
belong to the current run; in the reclaim the hook belongs to a previous one. That pair ships as
a test asserting the two produce opposite answers from the same evidence.
Missing run ids are treated as eligible. Absence means "this peer does not publish them", not
"this is stale", so an old host's rows are never blanked. Sibling evidence is opt-in so
pane-scoped consumers cannot inherit another pane's agent.
No consumer imports this yet; each migrates separately with its own evidence.
Verified non-vacuous: reversing the authority order fails 10 of 18 assertions and removing the
run filter fails 3.
* fix(agent-status): close three resolver contract holes found in review
**Duplicate evidence of one source resolved by array order.** `eligible.find(...)` returned the
first match, so two live hooks naming different agents were settled by input position — the exact
property this resolver exists to remove. The original order-independence test only used DISTINCT
sources, so it never exercised it. Conflicting same-class evidence now returns null with
`ambiguousAt`, and does NOT fall through to a weaker source: letting a title answer whenever two
hooks disagree is worse than saying nothing.
**A bare numeric runId collided across authority restarts.** `incarnation` is a total order only
within one `authorityId` (agent-status-observation.ts states this), and the id is regenerated per
authority instance, so a restarted host counting from its own floor would report `1` and match an
unrelated live run 1. The run key now carries its authority, and evidence from a DIFFERENT
authority is treated as incomparable — kept, like an absent key — rather than as stale.
**Title stayed reachable by consumers that authorize writes.** Ranking it last makes misuse
unlikely; `minimumSource` makes it impossible. An action consumer passes `'launch'` and weaker
evidence is dropped before ranking, so routing or delivery cannot name a target from a parsed
string even by reordering its inputs. Display surfaces omit it and are unaffected.
Also restores the generic agent-vocabulary parameter, which lives on the routing branch and was
lost when this branch was rebased.
Each fix is mutation-verified: first-match restored fails 3, ignoring authority fails 1, dropping
the floor fails 2. The authority test was itself vacuous on the first attempt — both sides used
`incarnation: 1`, so a resolver ignoring authority still passed on the numeric compare. It now uses
differing incarnations.
The remaining review finding, that `process > launch` has no freshness bound, is NOT fixed here:
it needs an observation timestamp the evidence type does not yet carry. Recorded rather than
silently dropped.
* fix(orchestration): route @agent messages by resolved identity, not terminal title
`@claude` picked its recipients with `buildAgentNameRe('claude').test(title)`, so any pane whose
TITLE contained the word received Claude's messages. Terminal titles carry task text, and people
describe agent work in them, so this is the ordinary case rather than a contrived one: the
recorded title "Switch Claude and Codex off the load balancer… - grok" is a Grok pane that
received both @claude and @codex. Misdelivered instructions, not a cosmetic slip.
The cause is that `RuntimeTerminalSummary` carried no identity at all — `title` was the only
identity-ish field on it, so routing by title was the only option available. Fix the input:
- `RuntimeTerminalSummary.agentIdentity?: TuiAgent` — optional, host-resolved from launch and
foreground-process evidence the host owns, with the title ranked last and contributing only
when the evidence parser finds an unambiguous name. A title that merely mentions an agent
yields no evidence, which is the whole point.
- `resolvePublishedPaneAgentIdentity` in `src/shared` rather than inside the runtime class, so
the decision is testable without a runtime and so routing, delivery and the UI cannot drift.
- Groups match `agentIdentity`; the title matcher and its bespoke Cursor predicate are deleted.
Unknown fails closed. `agentIdentity` is absent when the host predates the field or had no
evidence beyond the title, and delivery is an action: not delivering is visible and recoverable
(the sender sees no recipients), while delivering to the wrong agent is neither. The optional
field is additive, so an old client simply ignores it (wire rule 1).
This is also the first real caller of the evidence parser and the identity resolver.
Tests: 27 in groups, 8 for the publisher, 3 RPC fan-out cases updated to the new contract. The
`@cursor`-must-not-match-"text cursor blink" hazard is now excluded structurally instead of by a
per-agent predicate.
Verified non-vacuous by mutation: swapping the process/title ranks fails 2 publisher assertions,
and reverting groups to title matching fails 15 of 27. One earlier mutation silently failed to
apply after formatting reflowed the block — the file was checked before trusting the result.
* perf(runtime): reuse terminal title during summary build
* fix(orchestration): refuse title evidence when publishing identity for routing
Rebuilt on current main so this carries the hardened parser from #16148 and the corrected
resolver from #16157 (authority-scoped run keys, no order-dependent duplicate resolution).
Applies the resolver's new `minimumSource` floor at the publisher. What this publishes authorizes
an action — routing decides which real agent pane receives a message — so ranking title last is
not enough; the floor removes it from consideration entirely, and no amount of reordering by a
caller can bring it back.
The trade, stated because it is a real capability loss: a hook-less agent over SSH that Orca did
not launch, and whose foreground process the host cannot read, is no longer addressable by @agent.
Accepted because a message delivered into the wrong agent's prompt is unrecoverable while an
undelivered one is visible — the sender sees zero recipients. Whether real panes actually carry
launch/foreground evidence is the open question, and is what live validation must answer.
* fix(pty): preserve agent identity on daemon reattach
* fix(runtime): retire stale pane agent identity
* chore: normalize runtime types formatting
* fix(agent-status): identify a pane from its own hook, not from how it was started
Two defects, one cause: identity was inferred from the outside instead of read from the agent.
**Hook evidence was never plumbed in.** The publisher considered `process`, `launch` and `title`
and contained zero hook references — while the resolver ranks `live-hook` first. The top rung of
the ladder was never connected.
That made identity depend on Orca having launched the agent. Most agents are started by typing
`claude` or `codex` at a shell, which leaves no launch record. On macOS the foreground process
still names them, so the gap was invisible. On WSL the Windows host reads the foreground process
as `wsl.exe` — the distro wrapper, not the agent inside it — so those panes had no signal at all
and became unaddressable by `@agent`.
A hook is the agent reporting itself, so it survives both: no launch record needed, and no
dependency on reading a process across the WSL boundary.
**`launch` outranked `completed-hook`.** Ranking is now by TENSE rather than by how authoritative
a source sounds:
present: live-hook > process
past: completed-hook > launch > sleeping-session > sibling > title
A launch record is an event, not a state — it stays true after the agent exits, which is why a
pane reused after closing its agent kept reading as the old one. A completed hook at least proves
the agent actually ran in that pane; a launch record only proves Orca tried to start one.
Neither rank was covered: all 392 existing tests passed unchanged after reordering. Mutation now
fails 2 on the old order and 4 with hook evidence removed.
Known remaining gap, deliberately not papered over: a hand-started WSL agent with no managed hooks
has no identity signal at all. Restoring a title guess there would reinstate the misdelivery this
PR exists to prevent.
* fix(orchestration): restore title as the last resort, not a forbidden source
An earlier revision passed `minimumSource: 'launch'` so routing could not see a title at any rank,
reasoning that a display string must never authorize a write. That conflated the evidence parser
with the raw substring match it replaced.
`buildAgentNameRe('claude').test(title)` was the misdelivery. `collectAgentTitleEvidence` returns
null on exactly those shapes: "Review the Claude session-history fix" on a Codex pane yields
nothing, and "Switch Claude and Codex off the load balancer… - grok" yields grok from its owner
suffix. Ranking title last is therefore sufficient; refusing it is not necessary.
Refusing it had a real cost. An agent a user starts by hand inside an Orca WSL terminal has no
launch record, no readable foreground process (the Windows host sees `wsl.exe`, not the agent in
the distro), and — until managed Codex hooks install there — no hook either. An unambiguous title
was the only thing left, and dropping it made that pane unaddressable by @agent where the previous
code could reach it. That is a regression, and most agents are started that way.
End-to-end coverage added at the routing layer with title allowed: @claude still does not reach a
Codex pane whose task text names Claude, @codex still does not reach a Grok pane whose task text
names Codex, and a pane identified only by an unambiguous title is reachable again.
* revert(agent-status): keep launch above completed-hook until run keys exist
Reverts the tense-based reorder from this branch. The reasoning behind it was sound as far as it
went — a launch record is a past event, not an observation, which is why a reused pane kept reading
as its previous agent — but it fixed one staleness by opening a worse one.
A completed hook is past tense too, and without an agent-run key it never expires at all. Ranking
it above `launch` lets a stale hook from a previous agent outrank the launch record Orca stamped
for the process running NOW. pane-agent-owner.ts already says this in its own comment: "Ranking
launch/live-hook above the completed/sleeping records keeps a genuine pane on its real agent and
stops a stale record from hijacking it."
The reorder belongs with authority-scoped run generation, which is what makes any past-tense
evidence expire. It is staged in the migration plan rather than shipped here.
What this branch keeps: hook evidence feeding pane identity (so an agent a user starts by hand is
identified from its own report rather than needing a launch record), and title restored as a
genuine last resort behind the evidence parser.
* fix(runtime): guard the pane key so terminal.list survives a non-UUID leaf
`makePaneKey` throws on a leaf id that is not a UUID. The hook-evidence lookup called it unguarded
inside `buildTerminalSummary`, so a single such leaf took down `terminal.list` for the whole list
rather than degrading that one pane — 136 tests across 5 files, and the native code-quality gate
tripped separately on a duplicate test title.
Both were mine, and both were caught by CI rather than by me: I ran the focused suites before
pushing instead of the affected directories.
* fix(runtime): declare published terminal agent identity
* fix(runtime): demote completed hook identity evidence
|
||
|
|
419e3b4496 |
Fix terminal reads that flatten composer drafts into output (#16711)
* fix(terminal): separate composer drafts from read output Rendered screen reads treated cursor-line suggestion overlays as PTY output. Detect composer-owned text from cell attributes and cursor context, remove it from tail, and expose it as structured draft metadata. * fix(terminal): handle wrapped composer overlays * fix(terminal): preserve draft wrapping and tail alignment * fix(terminal): preserve composer wrap boundaries * fix(terminal): preserve draft continuations with middle dots * fix(terminal): recognize configurable Codex status lines |
||
|
|
913509edeb |
fix(orchestration): prevent slow worker-start stalls (#16300)
* Extend orchestration agent submission timing budgets * fix(orchestration): preserve mutation recovery identity * fix(orchestration): preserve recovery executable identity * fix(orchestration): keep worker starts and recovery commands safe * test(orchestration): cover federated worker preflight * fix(orchestration): harden mutation recovery * fix(orchestration): redact dispatch recovery credentials * chore: preserve upstream skill dialog formatting * test(orchestration): stabilize agent prompt submit e2e * fix(orchestration): validate federated start receipts * perf(runtime): cache unchanged prompt verification tail * fix(orchestration): reject worker-start timer overflow * fix(orchestration): normalize worker-start timeout defaults * fix(orchestration): normalize worker-start readiness budgets * fix(orchestration): normalize federated readiness timeout * test(runtime): tolerate current-main degradation exports * chore: preserve current-main orcad formatting * chore: drop unrelated formatting carryover |
||
|
|
b241a68ae4 |
Fix worktree identity collisions across hosts (#16691)
* fix(workspaces): add collision-safe worktree identity * fix(workspaces): read worktree metadata per host and repair ambiguous identities The canonical identity store landed write-only: getWorktreeMetaForHost had no production callers while setWorktreeMetaForHost kept the legacy projection only for the first known owner, so a second host's edits persisted and were never read back. Wire the listing paths through host-qualified reads. An ambiguous alias was also unrecoverable — reads returned undefined and writes threw forever, and the throw escaped the detected-worktree loop, emptying the whole repo's sidebar. Fail open onto the most recently active instance instead. - collapse ambiguous aliases deterministically and persist the repair - reclaim identity rows in the metadata GC so they cannot outlive their locator or resurrect onto a worktree recreated at the same path - drop every host's rows when a locator is removed outright, not just the owner's - honour an explicit instanceId so the stale-lineage rotation guard still works - scope a rename to the moving host; other hosts keep their own locator - prefer the project host setup matching the repo's own execution host, so a repoId registered on two hosts no longer stamps the wrong one durably - reject an unencoded `|` in a host id, the invariant the alias delimiter needs - drop the never-populated hostGeneration from the canonical key * fix(workspaces): close remaining identity review gaps * fix(workspaces): close remaining review gaps * fix(workspaces): address review and CI regressions * test(workspaces): update host-qualified metadata expectations * fix(workspaces): preserve ambiguous identity records * fix(workspaces): snapshot metadata during listing * test(workspaces): mirror listing metadata snapshot in windows fixture * fix(workspaces): preserve identity routing for metadata writes * fix(workspaces): scope stale metadata cleanup by host * fix(workspaces): rekey identities on SSH readoption * fix(workspaces): fail closed for ambiguous board ids * perf(workspaces): snapshot metadata across catalog listing * fix(workspaces): retain neighboring manual order updates * test(workspaces): cover ambiguous board id index * fix(persistence): harden host-qualified worktree metadata * refactor(shared): split project host setup lookup * refactor(workspaces): simplify host-qualified metadata |
||
|
|
3558cf943f |
fix(codex): heal WSL hooks before typed launches (#16535)
* fix(codex): heal WSL hooks before typed launches * test(codex): keep launcher fixture type-safe on Windows * fix(build): list codex-home-wsl-env in the CLI typecheck project `managed-home-shell-preflight.ts` is already in the CLI project's include list and now imports `wslCodexRuntimeHomeForGuestHome` from `src/main/pty/codex-home-wsl-env.ts`, which the list did not cover — TS6307, so the CLI typecheck failed on every push. Added the single module rather than a `src/main/pty/**` glob: it is a 31-line leaf with no imports of its own, so it does not widen what the CLI bundle can reach. * fix(codex): converge the two WSL hook install lanes onto one writer Two independent readiness reviews agreed the Orca-terminal boundary holds, but Codex Sol found a P1 the other rated P2: the new just-in-time repair raced the existing relay installer and the two produced DIFFERENT hook and trust representations for the same managed home. Two unserialized writers emitting different formats is worse than the bug this PR fixes, because it fails intermittently rather than cleanly — a pane works or does not depending on which lane won. - Relay Codex installs now delegate to the runtime-home writer, so there is one canonical representation instead of two. Redirected scripts use the runtime path, the readable wrapper, and the prepended group. - `installForRuntimeHomeSerialized` puts every asynchronous WSL caller for a given home on one queue (`wslInstallQueues`), so concurrent panes cannot interleave writes. Also rewrites the stale pin test the new `-x` guard broke. It asserted the defect — "would run the impostor if the preflight carried an unqualified command name", expecting the hijack marker to exist. The guard is a security improvement, so the test now asserts the contract: an unqualified preflight is skipped and the marker is never written. Rewritten to the new behavior, not loosened or deleted. 818 tests pass across the affected suites; typecheck clean. The changed-file quality gate could not run locally — its pnpm engine-warning JSON parser fails under Node 26 — so CI covers it. The boundary both reviews verified is untouched: paired/relay/mobile clients stay hard-blocked from the RPC, params remain shape-locked to the managed home suffix with traversal rejection, nothing is written outside the managed home, and macOS/Linux stay inert. * fix(codex): serialize resolved WSL hook homes * fix(codex): recover managed WSL homes after restart * fix(wsl): translate Codex preflight through WSLENV * fix(cli): cover bounded WSL Codex repair * fix(codex): coalesce duplicate WSL hook repairs * fix(codex): verify reconstructed WSL homes |
||
|
|
cc384c5a3d |
fix(agent-hooks): post posix payloads as json (#11292)
* fix(agent-hooks): post posix payloads as json * fix(agent-hooks): mark header merged envelopes * docs(agent-hooks): describe header merge envelope * fix(agent-hooks): encode posix metadata headers * test(agent-hooks): update WSL JSON hook assertions * fix(agent-hooks): negotiate raw JSON transport * fix(agent-hooks): preserve packed metadata in POSIX shells * test(agent-hooks): include hook envelope in relay boundary inventory --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
5631aa00dd |
feat(orcad): items 2–7 — degradation, natives, daemon, ops, deploy (#16398)
* fix(ports): stop joining an undefined resourcesPath on a non-Electron host `resolveWorkerEntryPath` branched on `isPackaged` alone and joined `process.resourcesPath`. orcad reports `isPackaged` true — correctly, it is a production build, and ~15 consumers read it that way to gate HTTPS-only skill downloads and the real CLI name — but `process.resourcesPath` is Electron-only and `undefined` under plain Node. So the packaged branch threw `TypeError [ERR_INVALID_ARG_TYPE]: The "path" argument must be of type string` where a clean "worker unavailable" was the honest outcome. The type said `resourcesPath: string`, which is how it went unnoticed; it is now `string | undefined`, so the compiler carries the fact. A host with no Electron resources tree has no asar to look in, so it falls back to the module directory and lets the caller report a missing worker. Found by the item 1 agent while auditing the same `isPackaged` defect class in the watcher. Verified in both directions: reverting the guard reproduces the TypeError. * feat(orcad): prove node-pty loads before anything requires it Of the two ways node-pty fails, only one is catchable. A missing module throws MODULE_NOT_FOUND. A module built against the wrong libc or Node ABI is refused by the dynamic loader, and in the worst case takes the process down before any handler exists — that is #9902, which crashed the desktop app on Ubuntu 20.04 before a window appeared. There was no libc or ABI precondition anywhere in the tree. So orcad now proves the load in a CHILD process, from main.ts, before anything requires node-pty. Whatever the child does — throw, abort, die on a signal — is data rather than our own death, and the operator gets a sentence naming the host's libc, Node ABI and prebuild slot plus the command to run. Proven-unloadable exits 78 (EX_CONFIG), so a supervisor does not restart an unequippable host forever. A probe that never answered is unverifiable, not blocked: refusing to boot on an inconclusive signal would take down hosts that work. The child dlopens the file node-pty would have chosen, before requiring the package. node-pty's loader walks several directories and rethrows only the LAST error, so a refused binary reads as "Cannot find module ./prebuilds/..." — which sends the operator to install a module that is already there. It also reports through stdout: node echoes the whole -e source above a stack trace, and matching tokens against stderr made the probe's own source text answer for the verdict. Verdicts reach clients as a terminal_unavailable degradation alongside the existing browser_unavailable one, through the same cause-registry shape. degradations[].code is now an open vocabulary; clients already render only `message`. Prebuilds are compiled from PATCHED sources — the patch IS the glibc-floor fix, so an upstream tarball reproduces #9902 — into linux-{x64,arm64}-{glibc,musl} and darwin-{x64,arm64} slots. libc is in the slot name because node-pty's loader falls back to prebuilds/<platform>-<arch> and cannot tell glibc from musl. orcad installs the matching slot at boot, so a host with no compiler serves terminals. The relay's five pure toolchain-diagnosis functions moved to a transport-free module so the Node bundle can reuse them without dragging ssh2 in behind them; the relay keeps its API by re-export. macOS gets `xcode-select --install` rather than the cross-distro apt/dnf/pacman/apk menu, every line of which is wrong there. * test(orcad): pin the node-pty precondition to ground truth, not a prepared host CI's test shard runs `vitest` directly, so `ensure-native-runtime --runtime=node` never prepares node-pty for the Node ABI — `degraded` is the correct verdict there, and asserting 'ok' encoded an environment the shard does not have. Asserting whatever it returned would be vacuous, so the expectation is now derived from an independent require() of node-pty. Verified it still bites: forcing the precondition to always report 'ok' fails the suite. * feat(orcad): run the terminal daemon, and the ops contract around it orcad declared `canRecoverPersistentLocalPtys: () => false` because it did not run the terminal daemon, so every restart, update and rollback SIGKILLed every running terminal — on the host whose selling point is that work survives the client going away. That is the one property `ssh-execution-boundary.md` recommends the peer model for. Item 4 — the daemon: - Port the launch path off electron: `daemon-init.ts`, `daemon-host-relocation.ts` and `observability/logs-directory.ts` now read the `AppEnvironment` port. Relocation additionally asks whether the app root is an asar archive rather than whether the build is packaged, so a Node host answering `isPackaged() === true` no longer walks into an Electron-only NSIS-escape path (same precedent as `parcel-watcher-entry-path.ts`). - `build-orcad.mjs` emits `daemon-entry.js` beside `orcad.js`, scans the forked children's metafiles for electron/node:sqlite, and load-checks the child under plain Node. - orcad spawns and adopts the daemon; shutdown disconnects and never kills it. `canRecoverPersistentLocalPtys` now reads the live provider and is false under degraded routing, where fresh terminals would die with the process. Item 3 — the ops contract (docs/reference/orcad-operations.md): - Bind policy: `--bind`, default loopback, pinned so neither `orca serve`'s wide default nor the connected-device widen can override it, and so a paired client cannot rebind the listener from outside. - Instance lock on the data root before profile load, scoped to the runtime role so it never refuses a restart that a live daemon makes worthwhile. - Supervision: exit codes a supervisor can act on (78 = do not retry), second-signal escalation, a shutdown deadline, and crash-loop containment on daemon respawn. - Health in the readiness payload: build hash, Node ABI, and a PTY self-test that spans both processes — the daemon spawns a real PTY in its own process and the verdict crosses its socket. Both bundle load-checks now assert on exit codes: these bundles are minified onto one line, so Node's uncaught-exception report echoes every string literal in the bundle and the previous message match passed against a bundle that never loaded. * feat(orcad): deploy, activate and roll back a versioned orcad install Plan items 6 and 7 from docs/design/shipping-orcad.html. Install reuses the relay's transaction verbatim — per-version lock, staged SFTP write, .install-complete sentinel, stale-lock recovery — under a parameterized namespace, so orcad-<v>/ sits beside relay-<v>/ permanently (§06). Parameterizing GC is the trap that creates: each model now collects only its own directories, enforced twice (prefix-scoped remote listing plus a local ownership re-check), and a client picks its model from how the host is registered, never from what it finds on disk. Activation is separate from installation, because a versioned directory selects nothing. A candidate is launched, publishes orca_server_ready, and only becomes active if its cross-process health payload passes: right build hash, listening, daemon live, PTY self-test green. A rejected candidate is stopped and the incumbent restarted, so a careful deploy cannot cause the outage it was being careful about. Update and rollback are shaped by the daemon. An update restarts orcad, the daemon outlives it, and the surviving daemon was forked from the outgoing bundle — so live terminals defer the update rather than proceed, and GC pins the active version, the rollback target and the live daemon's bundle. Orca's persisted state carries no schema version, so rollback restores a pre-activation snapshot rather than trusting backward-readability; the point past which it is unsafe is the first terminal created after activation, which the snapshot cannot describe and the surviving daemon still owns. Running the generated shell for real found two bugs the text assertions missed: tar members re-quoted inside a shell variable captured nothing, and kill -0 reports a zombie as alive. * test(orcad): assert the precondition is self-consistent, not environment-shaped The real-host case cannot predict a status: CI's shard runs vitest directly, so node-pty is never built for the Node ABI and 'degraded' is correct there, while a prepared checkout gives 'ok'. The previous attempt used require('node-pty') as ground truth, which resolves the JS wrapper while the native binding loads lazily — it proved strictly less than the precondition checks, and failed CI for exactly that reason. What is invariant on a host with node-pty installed: never 'blocked', and never a degraded verdict carrying an unestablished reason. The injected-input tests keep the logic coverage. * fix(orcad): drop an eslint-disable the rule no longer needs * test(orcad): separate slot placement from the load verdict Both remaining CI failures were the same shape: tests reaching into node_modules for a pty.node that only exists after `ensure-native-runtime --runtime=node`, which CI's shard never runs because it invokes vitest directly. Slot *placement* is the logic worth checking on every host, so it now uses a synthetic payload and asserts the verdict stays honest about not loading. The three assertions that genuinely need a Node-ABI binding are gated on it existing. Verified: breaking slot installation fails both placement tests; with the real pty.node hidden the file is 17 passed / 3 skipped instead of ENOENT. * test(orcad): gate the load-dependent cases on a real load, not on the file existing CI ships a pty.node built for Electron's ABI, so existsSync was true while require still failed — the gate ran exactly the tests that host can never satisfy. It now probes the binding in a child process, so a bad one cannot take the runner down. The self-consistency assertion also allowed too little: 'blocked' is the honest verdict for a corrupt binding, alongside 'ok' on a prepared host and 'degraded' on an unprepared one. What stays invariant is that anything other than 'ok' names an established cause, so a terminal is never declined for a reason nobody worked out. Verified against all three host states: prepared (19 passed), unprepared, and a corrupt binding (17 passed / 3 skipped, no failures). * test(orcad): gate on the whole premise — binding AND spawn-helper CI has a loadable pty.node but no spawn-helper, and a slot without the helper is legitimately 'degraded'. So the previous gate let a test run whose premise ('a complete slot yields ok') that host cannot satisfy. Verified in both states: with the helper present 19 pass; with it removed the load-dependent cases skip (17 passed / 3 skipped) instead of failing. * fix(orcad): preserve degradation types after rebase |
||
|
|
de6fe8b7ea |
fix(worktrees): resolve id: worktree selectors by path equivalence (#16243) (#16494)
* test(worktrees): cover id: selector path-spelling parity with path: (#16243) The renderer can only address a workspace by id (toRuntimeWorktreeSelector always emits id:<repoId>::<path>), and the runtime matches that id byte for byte while a path: selector has always compared through normalizeRuntimePathForComparison. A stored id that spells its path differently from `git worktree list` therefore resolves for the CLI and answers selector_not_found for the UI, which reads that as a stale local mirror, calls forgetLocal, reports success, and lets the row return on the next catalog refresh: a silent delete. These tests fail on both resolution sites -- the fleet `id:` branch of resolveWorktreeSelector and the scoped resolveScopedWorktreeIdRow a host-qualified removal takes -- and pin what must stay closed: an exact repo id (STA-4343), host qualification, dot segments neither selector canonicalizes, and a refusal rather than a guess when two rows spell one path. 13 failing, 35 passing. * fix(worktrees): resolve id: worktree selectors by path equivalence (#16243) worktreeIdComparisonKey names one repo, one filesystem location, and one folder-workspace instance, folding exactly the path spellings normalizeRuntimePathForComparison already folds for a path: selector -- and nothing more, so dot segments stay unresolved for both shapes. Both id: resolution sites consult it only after an exact match finds nothing: the fleet branch of resolveWorktreeSelector and resolveScopedWorktreeIdRow, which a host-qualified removal takes. runtimeWorktreeIdsEqual now derives from the same key so the runtime has one normalizer rather than a parallel one. Not a pure refactor at that last site: runtimeWorktreeIdsEqual used to normalize-compare ids that parse but carry an empty repoId or an empty path ('::/p', or 'repo::' against 'repo::/'), and worktreeIdComparisonKey returns null for those, so across its call sites (PTY identity, refresh, mutation queue) such ids now compare byte-exact instead. That narrows matching rather than widening it, no real worktree carries such an id, and it is the behavior #15616 guarantees for malformed ids -- but it is a behavior delta, not just a tidy-up. Perf (#14399): the exact match is still tried first and still wins outright, so a resolvable id costs exactly what it did before. Neither site adds a scan -- the fleet branch re-filters the array it had already listed, the scoped lookup re-filters the single owning repo's projected rows -- so an explicit id still never scans every repo. Fail-closed behavior is unchanged: the repo id compares exactly (STA-4343), host qualification is untouched, the folder-workspace instance suffix stays part of the path, and a scoped lookup with two equivalent rows refuses instead of guessing. The bare unprefixed selector branch keeps byte-exact id matching, since only the id: shape reaches a renderer caller. Shares src/shared/worktree/id.ts with the open #15616, which introduces worktreeIdComparisonKey for the same divergence in lineage pruning and authoritative-scan purging; this adopts that helper rather than adding a second one. Complementary to the open #16295, which makes the miss visible; this removes the miss. * chore(worktrees): satisfy oxfmt and oxlint on #16243 tests oxfmt --check flagged both new test files and oxlint's unicorn/no-useless-fallback-in-spread flagged the store mock; the full lint and format gates now match the pre-change baseline. * test(worktrees): pin Windows spellings and malformed-id exactness (#16243) Review found two axes the first pass left unproven at the two id: resolution sites. Both are the invariants the open #15616 guarantees for the shared worktreeIdComparisonKey it introduces for #15598, so violating either here would break a contract a sibling PR depends on. Windows: #15598's whole defect is that one checkout is recorded under both `D:\Agentic\game2` and `D:/Agentic/game2`. The fleet branch, the scoped removal lookup, and the key itself now each resolve the backslash spelling against the forward-slash spelling git reports, and fold drive-letter case -- while a backslash inside a POSIX path stays a filename character and a POSIX root stays case-sensitive, exactly as normalizeRuntimePathForComparison already decides for a path: selector. Malformed ids keep exact matching at both sites: an id with no repo boundary or an empty path still refuses, and the scoped lookup still refuses it without scanning. Four of these fail without the production change (three fleet/removal Windows cases and the scoped one); the malformed-id and POSIX-backslash cases are invariant guards that hold either way. Verified: 58 passed in the three files; 18 fail with the production hunks reverted; orca-runtime.test.ts and worktree-teardown-unstopped-pty.test.ts green (1270 passed | 1 skipped); pnpm tc:node clean. * test(worktrees): pin Windows id: spelling folds and fleet ambiguity refusal (#16243) The Windows backslash spelling now rides the ID_SPELLINGS rows, so it is driven through both id: sites -- resolveWorktreeSelector and the scoped removal target -- and compared against what the same workspace's path: selector resolves, rather than only through worktreeIdComparisonKey. That is the spelling #15598/#15616 found in the wild and the one the owner's Windows client produces. The fleet path's ambiguity refusal had no test: two same-repo rows spelling one directory, an id: matching neither exactly, must reject selector_ambiguous. It is the fail-closed guard on a delete-capable resolver, and the property a later refactor is most likely to turn into a silent pick. Also records two limits at the source instead of leaving them to be rediscovered: a UNC or WSL root never folds into a drive-letter location (while Windows' two WSL UNC aliases do name one location), and a folder-workspace id keeps a trailing slash placed before the ::workspace:<uuid> suffix, so that spelling stays exact-match-only. Neither behavior changes here. The file docblock overclaimed parity. path: collapses duplicate same-host registrations to the first row while a folded id: refuses them; the contract this file pins is path-spelling parity, not dedup parity, and the divergence is deliberate because this resolver also serves delete. Non-vacuity, verified by temporarily reverting the production hunks: neutralizing both id: fallbacks turns 10 of these tests red, including both new Windows rows and the ambiguity refusal (it degrades to selector_not_found). Making the fleet fallback pick the first folded match instead of collecting all of them turns the ambiguity test red on its own. The remaining cases -- malformed ids, dot segments, the POSIX backslash, the folder-workspace slash -- pass against the pre-fix code too: they guard against future widening rather than proving this fix. Drops the two Windows cases the ID_SPELLINGS row subsumes. The drive-letter case test asserted only the Windows half its name promised; it now also pins that a POSIX root does NOT fold case, since an unconditional lowercase would merge /data/Foo with /data/foo on the platform CI runs on. Fixture paths use the upstream-attested /srv/projects prefix (and a neutral plugin-host leaf) instead of a local install root; the spelling variations the tests exist to pin -- doubled separator, dot segment, trailing slash, uppercase POSIX, cafe NFC/NFD, and the Windows D: rows -- are unchanged in form. * docs(worktrees): trim the id: selector test header and document the comparison key (#16243) |
||
|
|
0f522c35e5 | fix(remote): gate empty session inventory on host authority (#16546) | ||
|
|
d60a3c900b | Reset stale terminal modes after dead TUI replay (#16379) | ||
|
|
07f2e14c08 |
refactor(codex): make WSL account surfaces direct-home aware (#16499)
* refactor(codex): make WSL account surfaces direct-home aware * fix(wsl): keep Codex relay hooks on managed runtime home * test(wsl): assert relay hooks use managed Codex home |
||
|
|
26721bd632 |
fix(codex): stop blocking the main thread on trust grants (#16441) (#16594)
* fix(codex): stop blocking the main thread on trust grants (#16441) Codex hook trust was granted by blocking the Electron main thread on `spawnSync` of a bundled ELECTRON_RUN_AS_NODE entry for the whole app-server deadline: 15s native, 35s WSL, ~45s on the real-home path (rebase inspect + repair + grant). Cold start and every Codex pane launch showed "Not Responding"; the reported event-loop gap was 15,049 ms. The subprocess only ever existed to donate an event loop to a deliberately blocked parent — `runCodexHookTrustGrantSession` was already the real async implementation. Make the callers async and the fork is unnecessary, so the bridge, the forked entry and its envelope are deleted along with their build/knip/tsconfig registrations. The CLI `agent hooks prepare-codex` handler is already async, so it awaits the in-process session and saves a process spawn per managed-home shell. `resolveCodexTrustGrantHost` is async too; the WSL identity probe moves from `execFileSync` to `runProcess`, dropping that file from the child-process import allowlist. Status reads keep a synchronous native-only stamp path. Two invariants that held only because the lane blocked: - Overlapping capability probes were impossible by construction. `GitCapabilityCache`'s dedupe engine is extracted to a shared `CapabilityProbeCache` and `CodexAppServerCapabilityCache` now inherits it, so concurrent launches against a cold host share one app-server session instead of one each. - Two grants on one `config.toml` could not interleave capture and restore. A reentrant per-file lane now serializes the whole install sequence (managed, WSL runtime, real-home ensure, legacy sweep) and the grant and rebase inside it. Cold-start work moves off the critical path: retained-home reconciliation (N sequential sessions) is fire-and-forget behind the daemon provider, and the startup real-home ensure chains into managed hook reconciliation instead of blocking app init. Every preserved semantic is unchanged: never throws, the ORCA_DISABLE_CODEX_TRUST_RPC kill switch, ledger hits, backfill-pending and cooldown fallbacks, config rollback on every failure path, pre-grant self-computed trust removal, the verify-failure taxonomy, diagnostics and telemetry. * fix(codex): widen the trust-config lane to every config.toml writer Review follow-ups on #16441's async trust grant: - `markCodexProjectTrusted` now runs inside the runtime+system config.toml lanes, so a project-trust write can no longer land inside a hook grant's capture->restore window and be silently reverted. Its callers await it. - `install`/`refreshRuntimeUserHooks`/`remove` hold the system config.toml lane as well as the runtime one — they promote approvals into ~/.codex/config.toml and mirror it back. Lock order is runtime-before-system everywhere. - The real-home ensure chain resumes after a rejection instead of returning the same rejected promise to every later pane launch, and resolving the real home is now inside the module's never-throws boundary. - `buildSpawnEnv` awaits inside a cancelable pending-spawn registration, so shutdown during the (now long) env build stops the PTY from launching. `prepareLocalPtySpawn` generalizes into `awaitCancelableLocalPtySpawn`. - CapabilityProbeCache drops the test-only `nowMs` passthrough; its probe backstop comment now describes what it actually guards. - Preflight is a plain async function; the trust dispatch in orca-runtime collapses into one `markWorkspaceTrustedForAgent`. * test(codex): exercise the trust-config lane under real concurrency The async grant makes two pane launches overlap for the first time. These drive the real modules end to end on real files: a rollback swallowing a sibling's grant, a markCodexProjectTrusted write landing inside a capture -> restore window, shared capability-probe dedupe on a cold host, the host-scoped transient cooldown, and reentrancy from inside an installer. Each was verified to fail against a deliberately broken implementation (lane removed, dedupe disabled, cooldown made global, reentrancy pass- through disabled). * test(codex): stop hook-service suites spawning the developer's real codex The forked grant bundle never existed under vitest, so the RPC lane was unreachable in tests on main. Running it in-process makes these suites spawn a real `codex app-server` when one is installed: 38 spawns and two failures in hook-service-runtime-trust-repair on a machine with codex, green in CI where there is none. Stand in for the missing binary so both environments exercise the same fallback lane. * docs(codex): scope the trust-RPC kill switch comment to what it actually gates The comment read as though the flag forces the fallback lane everywhere. It gates the managed grant only: the real-home rebase still runs its own inspect/repair app-server sessions when Orca's insertion shifts a user's hook positions, and never reads the flag. Verified by exercise, not by reading — with the flag set, both inspect-user-hook-trust and repair-user-hook-trust still ran. Pre-existing: main has no check there either, it just blocked the main thread while doing it. Widening the flag to cover the rebase is a follow-up; this only stops the comment promising something the constant does not do. |
||
|
|
68d5b9206e |
fix(agent-prompt): stop reporting delivered prompts as stalled (#16095) (#16590)
* fix(agent-prompt): stop reporting delivered prompts as stalled (#16095)
Enter is written before verification runs, so `agent_prompt_stalled` can only
ever mean "turn start not observed" — never "prompt not delivered". Three of the
verifier's blind spots made that misreading routine, and the coordinator then
treated it as non-delivery and pasted the whole preamble a second time into a
worker already running it.
- Accept a hook-reported `working` recorded after the baseline. Hook rows reach
the runtime through getAgentStatusSnapshot with no window involved, unlike the
synthetic-title route that feeds workingSequence (suppressed for codex, absent
for kimi, and gated on window visibility for everyone else).
- Accept pane output after Enter when the agent was already working: a
`->working` edge is unreachable there, so the old predicate could never be
satisfied by a follow-up prompt. An idle agent still owes a real turn start,
so a swallowed Enter stays detectable.
- Give codex/kimi panes a longer effect window; their only turn-start proof is
an out-of-process hook round-trip, not a TUI repaint.
- Coordinator dispatch no longer fails (and re-dispatches) a task whose prompt
stalled; the dispatch stays active with its capability intact so the worker's
own report settles it.
* fix(orchestration): let a worker's own report correct an unobserved prompt (#16095)
Follow-up to
|
||
|
|
08a447dfb2 |
fix(terminal): size the pre-Enter wait to what the host actually ingests (#15925) (#16586)
* fix(terminal): size the pre-Enter wait to what the host actually ingests The Windows agent-prompt submit delay was a flat 1_500 ms frozen from the client's process.platform at import. Measured on two real Win11 hosts, ConPTY ingests a bracketed paste linearly at ~0.009-0.010 ms/byte, so the constant was both far too long for a 2-8 KB prompt (14-89 ms of real cost) and too short past ~145 KB — at 160 KB one host took 1_499 ms, meaning Enter landed mid-paste, exactly the corruption the delay exists to prevent, up to the 16 MB input ceiling. Replace it with getTerminalPasteIngestMs(platform, byteLength) and derive every pre-Enter wait from it: - open-loop fallback = 500 ms settle + ingest bound, uncapped - claude/codex render gate cannot start its quiet window before the ingest bound elapses (an agent that repaints mid-ingest could otherwise satisfy marker-then-quiet while ConPTY was still feeding the paste), and its 8 s hard cap now sits on top of the ingest bound instead of standing in for it - the plain terminal.send suffix path, which had an undocumented flat 500 ms The rate follows the host that owns the pty transport, not the client: a WSL pane is spawned as wsl.exe behind the Windows pseudoconsole so it still pays ConPTY, while an SSH pane follows the relay's reported remotePlatform. Also swap the inter-chunk setTimeout(0) for setImmediate. It cost a full ~15 ms Windows timer tick per 16 KiB chunk (~0.95 s/MB) while pacing ~1.07 MB/s — 11x above ConPTY's drain rate — so it never provided backpressure; the event-loop yield it did provide is preserved. * fix(terminal): stop double-charging paste ingest in the render gate The render gate's hard cap is armed twice -- once at arm() and again when the show-cursor marker arrives -- but it re-added the whole ingest window each time while the ingest clock itself runs once from gate construction. A marker seen mid-ingest pushed the cap out by a second full ingest term (~34 s instead of ~24 s for a 1 MB prompt on ConPTY). Capture the ingest deadline absolutely and arm with what is left of it. Also thread the request AbortSignal through terminal.send so the now payload-scaled suffix wait can be cancelled: at 16 MB it runs ~262 s, well past the CLI's 60 s request budget, and previously nothing stopped the eventual Enter. Cleanups: a pty record's connectionId is only ever an SSH target id, so the wsl: relay-id guard in getPtyWriteHostPlatform was dead; and hoisting action.text removes both non-null assertions in writeTerminalAction. |