mirror of
https://github.com/stablyai/orca.git
synced 2026-09-22 08:02:28 +00:00
e2b70a5eba68416972f94de737f39fcd60fdeec7
324
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e84fb46eaf |
test: add golden E2E tests for source control workflows (#14260)
* test: add golden E2E tests for source control workflows - Tests core source control interactions: file edit/save, commit staging, and diff viewing - Integrated into CI/CD pipelines for Linux, macOS, and Windows - Includes helper utilities for test setup and worktree management * test(e2e): verify golden commit author and fix test flakiness - Configure git author name/email at worktree level during setup - Verify commits are made with correct author details in assertions - Add explicit timeouts to file visibility waits and git status polling - Fix test ordering to seed edits after source control is open - Simplify git status refresh logic to rely on automatic updates * Add rollback to createGoldenWorktree on setup failure Cleanup callbacks only register after setup succeeds. When a config command fails, the half-built worktree and branch leak into later test runs, causing flakiness. Now we roll back immediately and re-throw the setup error. * test(e2e): match explorer rows after the git status badge appears The golden file-save spec used an exact /^README.md$/ filter. After save, the explorer row text becomes "README.md M", so reopen clicked nothing. * test: strengthen golden worktree setup verification - Track working directory in git call inspection to verify correct execution context - Verify user.name/email config applies to worktree-specific settings, not repo - Add exhaustive setup call sequence assertions to catch setup/rollback leaks |
||
|
|
e7b85266f5 |
Add golden E2E tests for fresh terminal and shell commands (#14302)
* test(e2e): add golden tests for fresh terminal and shell commands Adds regression tests for terminal initialization in fresh profiles and shell command execution to the golden test suite, integrated across Linux, macOS, and Windows CI. * test(e2e): bracket shell command output between markers The echoed command can wrap or be clipped by the buffer tail. Bracket output between begin and end markers to reliably identify real output, and strip ANSI escape sequences that interfere with parsing. |
||
|
|
dd63d35d09 |
fix(ci): skip missing golden scripts on older release tags (#14267)
Cut Release is dispatched from main but checks out the tagged tree. Cherry-pick tags such as v1.4.182-rc.1 do not define test:e2e:windows-fresh-startup-golden, so the Windows golden job failed with ERR_PNPM_NO_SCRIPT. Run tag-optional goldens with --if-present. |
||
|
|
7c93aed6dc |
Fix fsync of read-only files on POSIX (#14235)
* fix(files): fsync read-only files on POSIX * test(e2e): add golden E2E tests for POSIX profile index fsync Validates that profile index files are properly persisted on POSIX systems, including with restrictive umask settings. These are release-blocking golden tests for Linux and macOS. * test(terminal): wait for fish child ownership before stdin write Fish 4.8 withdraws DECSET 2031 before spawning the child, so the shell-contracts harness could send hello into an intermediate prompt and hang waiting for CHILD-READ. Wait for the child's CHILD-READY marker and answer split DA1/CPR/OSC queries across chunk boundaries. * test(e2e): verify profile index persists to disk with restrictive umask Strengthen the POSIX fsync test to verify the rebuilt index is actually written to disk and has correct permissions under a restrictive umask, not just cached in memory. |
||
|
|
9a10561258 | fix(terminal): retain SSH startup delivery through reconnect (#14161) | ||
|
|
de729d6067 | Fix paired web creation failure handling (STA-4024, STA-4025, STA-4063) (#14100) | ||
|
|
b115f8d256 |
Add Windows golden E2E test for fresh-startup regression
Windows terminal rendering golden is flaky on CI runners. Re-enable Windows in the golden E2E gate with a scoped test for the fresh-profile startup regression from #14130. Terminal rendering continues on Linux and macOS; Windows runs fresh-startup only. |
||
|
|
e8044b1b30 |
fix(windows): restore fresh-profile startup after durable fsync (#14173)
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> Co-authored-by: DHTheOne <238933622+DHTheOne@users.noreply.github.com> Co-authored-by: Anton Tupitsyn <70199858+PLUTONYY@users.noreply.github.com> Co-authored-by: 7loop <48346764+7loop@users.noreply.github.com> |
||
|
|
90b8554fc9 | fix(terminal): recover readiness after startup exec (#14027) | ||
|
|
5ea7df1a5b |
fix(terminal): make DECSET 2031 subscriptions silent (#13904)
fish arms `CSI ?2031h` before painting each prompt and withdraws it when it hands the tty to a child — a ~1ms window. Orca answered that subscribe with `CSI ?997;Nn` across a 1-3ms renderer hop, so the reply landed after the withdrawal and was read as stdin by the next child, corrupting `brew`/`npx` `[y/N]` prompts. The reply is not stale by Orca's own view when written (measured staleReplies: 0), so no suppress-the-stale-reply scheme can close this — the information needed to suppress does not exist yet. Nothing asked for the reply either. The Contour spec says a terminal "should only send out the DSR when the palette has been updated"; Ghostty (Termio.zig:729 — force=true reachable only from the ?996n DSR), iTerm2 (VT100Terminal.m:995 — flag only) and xterm.js (InputHandler.ts:2035 — flag only) all emit nothing on the DECSET. So stop entering the race: record the subscription, answer nothing. Of 17 real programs measured under a pty, only fish, tmux, claude and opencode subscribe; none block on a reply, and answering produces one redundant palette re-query and zero rendering difference. tmux is the only one that sends `?996n`, which Orca still answers. - Subscribes are record-only at all four emitters (live scan, hidden-gate fact, parked byte watcher, parked responder — the last is deleted, it only replied). - `?996n` answers, the subscription registry, and the theme-flip push are unchanged. `paneLastThemeMode` is still seeded at subscribe so the next appearance re-apply is not read as a flip. - Replay grammar carries `?2031l` alongside `?2031h`, so a late-attaching remote client no longer registers a subscription the TUI already retired. Also closes fish-integration gaps found alongside: `unset` (which fish lacks) becomes `set -e` on paths parsed by the client's login shell, `config.fish` is parsed for agent-home detection, and bracketed-paste startup delivery is made consistent across local/daemon/relay. Regression test drives real fish 4.7.1 under node-pty and asserts on what the child process reads; it fails against pre-fix code with the exact payload from the issue. CI installs fish 4 and fails loudly rather than skipping. Closes #9993 Co-authored-by: Orca <help@stably.ai> |
||
|
|
d6e1d84235 |
fix(wsl): forward native CLI arguments losslessly (#12582)
* fix(cli): preserve WSL --deps quotes and parse task ids strictly PowerShell 5.1 native splat was stripping ASCII double quotes on the WSL bridge, so non-empty JSON --deps arrays failed while [] still worked. Pre-escape quotes before launching orca.exe, and recover quote-stripped task-id arrays while rejecting non-task-id and malformed CSV input (#12188). * fix(orchestration): narrow WSL deps recovery * fix(wsl): forward native CLI arguments losslessly * ci(windows): exercise WSL PowerShell argv boundary --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
4c69e45552 |
Strengthen plain-node-entry-guard with entry name validation (#12761)
* Strengthen plain-node-entry-guard with entry name validation - Add buildStart hook to validate guarded entry names exist in rollup inputs, preventing stale names from silently stopping guards - Extend electron require detection to subpaths (electron/main, etc) - Improve smoke test signal/exit handling and use constants - Add comprehensive tests for entry validation and new behaviors * Add SIGKILL escalation to plain-node-entry-guard timeout Switch from spawnSync to async spawn to properly handle daemons that trap SIGTERM. spawnSync's timeout only sends the signal and waits, so a daemon that ignores SIGTERM causes the build to hang. The new runDaemonEntry function escalates to SIGKILL after a grace period to enforce the deadline. Configurable timeouts and grace periods via SmokeTimings type; closeBundle hook becomes async to support the change. |
||
|
|
119b1e1c53 |
fix(release): reject non-release publications (#13591)
* fix(release): quarantine unauthorized publications * fix(release): remove unauthorized publication tags * fix(release): require canonical version tags * refactor(release): narrow policy to one workflow * fix(release): reassert latest stable release |
||
|
|
418fbb3192 | fix(mobile-ios): gate release on TestFlight distribution (#13419) | ||
|
|
3b1017c4fb |
Add nightly cut (#13410)
* Add daily macOS dev build release channel Publish once-daily signed macOS builds from main at a dedicated cadence, separate from hourly (too noisy) and release branches (too infrequent). Builds are notarized and installable via the updater, but unvetted — published to stablyai/orca-daily rather than the main repo to avoid evicting stable/RC entries from the releases feed. * fix lint * fix commit * Add third token mint to daily macOS build workflow The upload step's 2x45m retry budget can outlive the one-hour token, so a third is minted after it for verify and cleanup operations. Release notes are moved to a file to ensure consistency between draft creation and publish. Daily channel description updated with specific UTC release time. |
||
|
|
6858e072cf |
fix(terminal): agent pane auto-launch lost under fish + Starship (STA-3417) (#12840)
* fix(terminal): extend the shell-ready startup barrier to fish (STA-3417) Fish never emitted the OSC 777 shell-ready marker, so agent launch commands were written into the PTY while fish/Starship were still initializing: the daemon path wrote them synchronously at session create and the local path blind-wrote ~30ms after the first output byte. The command was echoed by the kernel but never executed. - shell-templates: shared fish --init-command that emits the marker once on the first fish_prompt event (the earliest point fish's own reader owns the PTY, mirroring zsh's zle-line-init marker) - daemon shell-ready: fish joins the startup barrier so the launch command queues until the marker (timeout fallback unchanged) - local-pty-shell-ready: fish launch config gains the marker wrapper - codex-startup-delivery/tui-agent-startup: omp/pi/opencode plans now request shell-ready delivery (codex parity) so the SSH renderer path also waits for the prompt; plain payload-free codex stays on the markerless fast path * fix(terminal): answer DA1 past the shell-ready barrier The barrier queues all inbound input until the ready marker, including the renderer's DA1 reply. A shell that withholds its first prompt until DA1 is answered — fish waits 10s — therefore never emits the marker that would release the reply it is waiting for. Measured: 10.37s to launch an agent, versus 0.35s once the reply lands. Answer DA1 from the daemon while the barrier holds, writing straight to the subprocess so the reply bypasses the queue, and consume the query so the renderer's xterm cannot also reply. Released on ready, timeout, or dispose, handing DA1 back to the renderer for steady state. Consolidates the identical DA1 handler the ConPTY override already used. * fix(terminal): prevent duplicate startup DA1 replies |
||
|
|
850342a3e0 |
fix(ci): run the root-directory guard on stock macOS bash 3.2 (#12879)
* fix(ci): run the root-directory guard on stock macOS bash 3.2 The guard script builds its base-tree lookup with `declare -A`, which needs bash 4+. Its test spawns plain `bash` from PATH, and stock macOS has shipped /bin/bash 3.2 since 2007, so on any Mac without a Homebrew bash the script exits 2 before asserting anything and the default `pnpm test` suite fails 3 of the guard's 4 cases. Machines with a Homebrew bash on PATH never see it, which is why it went unnoticed. Replace the associative array with a plain-array linear scan. Root directories number in the dozens, so the O(n^2) membership check is negligible, and the NUL-delimited reads that protect unusual filenames stay as they were. The empty-array expansion is guarded for `set -u` under bash 3.2. All four guard tests now pass with /bin/bash 3.2; behavior under CI's bash 5 is unchanged. * fix(ci): run the root-directory guard under node instead of bash The guard is the only check in the repo written in shell, and it used `declare -A`, which stock macOS `/bin/bash` 3.2 does not have — so the guard's own test suite failed 3 of 4 cases on any Mac without a Homebrew bash. CI never noticed because runners ship bash 5. Porting it to node removes the interpreter-version variable instead of working around one construct: node is what the sibling script in this directory already uses, it is the runtime that runs the test, and the NUL-delimited read is the same shape as check-changed-code-quality.mjs. It also drops a latent false pass — a failing `git ls-tree` inside the shell's `< <(...)` was not caught by `pipefail`, so the read loop saw nothing and the guard reported success. `execFileSync` throws instead, which is why the two `git rev-parse --verify` probes are no longer needed. Output and exit codes are otherwise unchanged; the usage line now prints node's script path where the shell printed `$0`. Tests pin each guarantee and fail when it is reverted: NUL-delimited reads so odd paths are reported unmangled, exit 2 on bad usage, and git's own 128 with no node stack trace when a sha does not resolve. * fix(ci): keep root entry bytes intact and fence guard output git pathnames are arbitrary bytes, but the guard read ls-tree with encoding 'utf8', so every invalid sequence collapsed to U+FFFD. That mangled the reported name and, because the replacement is not injective, let two different entries compare equal — a genuinely new root entry could be waved through as pre-existing. Read the bytes as latin1 and write them back unchanged. The blocked-entry list is also attacker-controlled and went straight to stdout. The runner trims leading whitespace before matching '::', so an indented entry name still parses as a workflow command, and a pathname may embed a newline. Wrap the list in ::stop-commands:: with a random resume token so only the guard's own annotation is acted on. --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
17cfc968cf |
Revert the terminal IME composition-ownership change (#13282)
* Revert "test(ime): restore coverage the composition-ownership change removed (#13168)" This reverts commit |
||
|
|
05c30166f4 |
fix(ci): stop hourly prune from deleting the just-published release (#13266)
* fix(ci): stop hourly prune from deleting the just-published release Hourly prune sorted non-draft releases by createdAt, but nearly every orca-hourly release shares one createdAt from bulk import. At the retain cap, stable sort + reverse put the newest release past the window and immediately deleted it with --cleanup-tag. Sort by publishedAt (tagName as tie-break) and hard-skip the tag this run just published so prune cannot self-delete. * fix(ci): harden hourly prune protect without retain+1 drift Review found that skipping only in the delete loop under-prunes when sort is wrong, and excluding TAG before the retain slice would permanently keep retain+1 releases. Force this run's tag to the front of the sorted list before slicing so it always has a retain seat and oldest builds still prune. Also gate prune on publish_live success and warn if TAG still appears stale. |
||
|
|
17b3dff3c4 |
refactor(terminal): return IME composition ownership to xterm (#13128)
* fix(terminal): return IME composition ownership to xterm * fix(mobile): derive terminal input from native replacement ranges * test(mobile): record iOS Japanese IME traces * fix(mobile): preserve native IME replacement ranges * fix(xterm): flush queued application input after IME commit * test(terminal): pin Korean intermediate commit * test: pin Windows IME shortcut ownership * test: replay IBus number candidate commit * fix: preserve native macOS input-method punctuation * refactor(terminal): remove stale mac focus override * fix(mobile): preserve soft keyboard deletion ranges * fix: keep IME-owned palette chords in renderer * fix: stop carried IME shortcuts at renderer owner * fix: preserve carried IME shortcut dispatch * fix: narrow main-owned shortcut actions * test(mobile): pin Japanese IME replacement traces * test(terminal): retain paired native IME trace * fix(chat): preserve browser IME composition ownership * fix(chat): retain macOS IME confirm gesture * fix(chat): expire unmatched IME confirm carry * fix(chat): isolate IME confirmation expiry * fix(chat): retain active IME confirmation * refactor(terminal): remove dead composition handler * feat(ime): add shared Enter-ownership seams for CJK composition The confirming Enter of a CJK composition arrives as two keydowns and the orderings differ by platform: Windows/Linux redispatch the unmarked Enter/13 before keyup, macOS delivers keyup first. A guard reading only isComposing or keyCode 229 misses the redispatch, so surfaces submitted on a confirm. Adds useImeEnterGestureOwnership (carry token, next-frame expiry), a shared ImeEnterGuardedForm for native implicit submission, and the cmdk seam covering 18 CommandInput surfaces at one site. A chorded Enter arms the carry but is never swallowed — the reverse would eat a user's deliberate Cmd/Ctrl+Enter. Both failure modes are pinned by ime-enter-gesture-ownership-contract.test.ts. Co-authored-by: Orca <help@stably.ai> * refactor(terminal): consolidate native input listeners and parked-screen owner Extracts the shared native-input listener installer and renames the parked-screen detector for what it actually does, replacing per-call-site duplication. The listener installer keeps a forgetOptionKeyLocationOnBlur flag so per-window semantics are preserved rather than flattened. Net deletion; no behaviour change intended. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin recorded IME shapes as regression tests Nine regression tests built from hashed affected-platform captures, each with a paired ordinary negative and a discriminating mutation verified to take the file from all-passing to exactly one failure. Covers the Windows MS-Korean Shift family (#12179, #11878, #12151, #11946, #12152) and the Korean TUI line-break rows (STA-3237, STA-3222, STA-3129). STA-3237 pins the empirical 3-Shift / 2-active-composition / 2-newline ratio the device run established — the third Shift produces nothing because Space has already committed. That ratio is not derivable from a static capture. Co-authored-by: Orca <help@stably.ai> * fix(ime): guard Enter-commit surfaces against CJK confirm Applies the Enter-ownership guards across the surfaces whose Enter commits something: publishes, clones, pairs, installs, posts, or persists. Tiered deliberately rather than uniformly. Irreversible and remote-effect sites take the carry token, which also blocks the unmarked redispatch. Locally reversible sites take the oracle check with a one-line comment naming the residual, because a spurious commit there costs one undo. Three numeric fields are left unguarded with the reason in-code: Chromium blanks number inputs at compositionstart, so a confirm-Enter only ever reaches an empty-draft reset. Measured with a CDP probe rather than assumed — a guard that cannot fire is noise. Co-authored-by: Orca <help@stably.ai> * test(ime): teeth-check the Enter guards on every guarded surface One suite per guarded surface, each verified by deleting the guard and confirming the test fails. A green guard test without that check is unverified, not verified. Two shapes pass vacuously in happy-dom and are avoided here: native implicit form submission never fires, and blur() is inert on an unfocused element. Both made "the commit did not happen" assertions pass with the guard removed, so the suites assert the guard's contract directly instead. Co-authored-by: Orca <help@stably.ai> * fix(mobile): keep iOS Korean commits whole through the live-input path iOS Korean reports isComposing: false on every event, so it bypasses the composition guard entirely. The strict owner rejected UIKit's transformed post-change field and sent only the leading jamo — the reported symptom. Prefers the authoritative same-event field text over the predicted text when the supplied operation cannot produce it. Generic: no Korean special-case, no locale classifier, no normalization. Adds the RN-target-keyed submit carry alongside it. Co-authored-by: Orca <help@stably.ai> * test(e2e): make IME capture harnesses fail loudly instead of silently Four instruments recorded silence as success, so a void run scored as a clean one: - readTerminalImeBoundaryTrace returned an empty trace when the probe never installed, making every "nothing leaked" negative pass vacuously - summarizeLatencies([]) returned a perfect zero distribution that passed all three latency thresholds - the macOS Vietnamese spec pinned an input-source ID that does not exist, and failed as though the operator had chosen the wrong source - the expectedLineCount=1 prefix property was undocumented and one edit from silently downgrading a PTY assertion Input sources now resolve by enumeration and name the near-matches on failure. Co-authored-by: Orca <help@stably.ai> * test(terminal): cover Cangjie cancellation and fix a cross-namespace assertion Adds #11951's recorded Cangjie cancel shape to the existing cancellation suite, which covered Pinyin and Sogou but not Cangjie. One keystroke then Backspace arriving as deleteContentBackward with data: null, so the stale preedit is the only thing a fallback could replay. Verified against the historical pre-6cd944c62b3 bundle: the positive fails with ['尸'] where [] is expected, while the ordinary negative stays green. Also fixes the Vietnamese spec, which asserted a TIS-space input-source ID against getKeyboardInputSourceId(). Those two Orca APIs report the same source in different namespaces — TIS nests it under VietnameseIM, the app API does not. The resolver stays as an installation precondition; the assertion matches the leaf. Co-authored-by: Orca <help@stably.ai> * test(e2e): add a real-IME macOS arm for the Korean chord commit The existing korean-ime-terminal-shift-enter-commit spec synthesizes composition over CDP: Input.imeSetComposition sets the preedit directly and Input.insertText performs the commit. Asserting the IME produced events you injected yourself is circular, so that spec cannot certify real-IME behaviour. This arm selects 2-Set Korean via TIS, reads it back live, and injects through System Events key codes, so the OS owns the preedit, the commit instant, and isComposing. PTY byte expectations are preserved verbatim. Covers 2 of the original 4 cases by design. The other two are the Windows/Linux redispatch-before-keyup ordering, which macOS cannot produce and which cannot be selected -- the OS decides it. Reintroducing synthesis to "restore coverage" would reintroduce the circularity. Co-authored-by: Orca <help@stably.ai> * test(e2e): assert the macOS chord arm at the PTY boundary, not the renderer The byte expectations were transcribed from korean-ime-terminal-shift-enter-commit :364/:383, which assert against onData -- a renderer boundary where the terminator is CR. This spec reads the PTY child, where the tty has already converted CR to LF. Names both forms per row rather than swapping the constant, so the conversion reads as evidence that the capture reached past the renderer, as #11936 and #11951 record. Ctrl+Enter's CSI-u sequence is unaffected and is identical at both boundaries. Co-authored-by: Orca <help@stably.ai> * test(e2e): measure composer-to-onData latency and stop dropping IME keystrokes Two defects in the echo latency probe. It hooked onWriteParsed and onRender but never onData, so it measured key->parse->render echo rather than the composer-vs-onData delta the latency rows need. Adds a third hook feeding its own sample set. And `event.key.length !== 1` silently dropped IME keystrokes: Pinyin and Cangjie keydowns arrive as key:'Process' (length 7). Replayed over the captured corpus, the old filter accepted 580 of 4137 Chinese IME keydowns -- it was discarding 80% of them. The new filter matches the shape the owner itself branches on. Attribution charges each onData to the latest keydown rather than a FIFO head, because composing jamo emit no onData at all and a queue would credit a whole composition to its first keystroke. The consumer now asserts sample count before any percentile, so a zero-sample run cannot render as a flawless distribution. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin the WSL shifted-jamo newline shape for #11919 In Korean 2-set, Shift types ordinary letters -- the double consonants and the compound vowels. Each such keystroke reaches Chromium as key='Process', keyCode=229, shiftKey=true. The v1.4.163 classifier matched exactly that pattern with no code guard, so it called those keystrokes Enter, rewrote them to a synthetic Shift+Enter, and injected a newline into the middle of the word -- with no Enter key pressed. That is why the reporters said "no modifier key pressed": they had not chorded Shift+Enter, but they had pressed Shift, to type the double consonant. Asserts the row's own recorded capture: 40 immediate keydowns, exactly 3 of them Shift-carrying inside a single syllable, and an onData stream with one newline per Enter press and none mid-word. Two ordinary negatives keep it from being a blanket mute -- the same session's non-IME keydowns still reach shortcut policy, and an ordinary Shift+Enter still resolves through the real policy. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin the composition commit lag that made Korean type one behind macOS Korean 2-Set commits syllable N only when the first jamo of N+1 arrives, so compositionend and compositionstart land in the same task. A composition-start handler cancelled the pending finalizer that was the only path to triggerDataEvent and ended the session without emitting bytes, so every committed syllable reached onData exactly one syllable late and the backlog cleared only at a Space or Enter. Types continuously with no Enter and no Space -- either would flush the backlog and hide it -- and samples onData at every syllable boundary. Paired with a length-matched ASCII arm that stays green throughout, so the positive is a fact about composition rather than about timing in general. Bisected to a single call site across five builds: pristine, 1.4.155 and 1.4.162 pass, 1.4.163 fails, removing the one call repairs it, restoring it fails identically. That window is exactly the reporter's "started immediately after updating". Co-authored-by: Orca <help@stably.ai> * test(mobile): cover the send-queue abort that silently drops queued keystrokes One failed send in use-terminal-live-input-commit aborts every keystroke queued behind it, with the error swallowed by .catch(() => false). The existing test resolves(true) on every send, so the failure branch was uncovered. Four arms: the abort itself, an ordinary negative on the healthy path, a throwing sender, and a liveness control proving the queue recovers once the chain settles. Deleting the abort takes 4 passed to 3 failed, with the ordinary negative correctly surviving. Scope is stated in the docblock: this is a transport send-queue abort, reachable only via a real disconnect or RPC error. REQUEST_TIMEOUT_MS is 30s, so latency alone cannot reach the branch — consistent with #7094's symptom class, not proven to be its cause. * test(terminal): pin that daemon snapshot/restore cannot disturb a composition Two independent reporters attributed broken Korean composition to the always-on PTY daemon repainting terminal state over the preedit. The attribution is wrong on ancestry — the daemon shipped three months before the version both call good — but the boundary was never actually tested. Runs the real applyMainBufferSnapshot choreography against a live composition, including the full 2J/3J/H wipe plus the resize and alt-screen branches. textarea.value, selectionStart/End, compositionView.textContent and .active all survive byte-identical, and interleaving a restore between every jamo of 문제 still commits 문제 at onData. Also pins that the uncommitted preedit is absent from the captured snapshot: it lives in the textarea, never the buffer, so a restore has nothing stale to echo back. Injecting one textarea.value = '' into the restore fails exactly the three restore-boundary tests. * test(terminal): pin that Cmd tears down a composition where Ctrl and Shift do not xterm's composition keydown exempts only keyCode 16/17/18 (Shift/Ctrl/Alt) plus 20/229. macOS Meta — 91/93/224 — is absent, so a Cmd press mid-composition takes _finalizeComposition(false): the overlay goes dark and never recovers, because compositionstart is not re-fired. The user composes the rest of the word blind. Linux and Windows users press Ctrl and are exempt. xterm already has a Meta-aware modifier predicate in wasModifierKeyOnlyEvent, so this is an internal inconsistency rather than a deliberate choice. Owns no reported row and is version-neutral: 5/5 on both 1.4.162 and 1.4.163. The branch is unexercised in all 328 recorded traces, so this is a hazard pin, not a regression guard. Only the teardown is asserted; the likely duplicated commit needs a compositionend the IME kept alive across the Cmd, which no capture contains. Deleting the exemption fails exactly the three paired negatives; adding Meta to it fails exactly the two Cmd arms. * test(native-chat): characterize preedit loss when a question card replaces the composer An AskUserQuestion card fully replaces the composer by design, but the in-flight composition goes with it: the composer unmounts before compositionend reaches it, so the preedit is never committed to the draft. The committed text survives only because the draft is cached and restored via defaultValue. Node identity changes, value 'abc' is preserved, the 가 is gone. Drives the real NativeChatView -> SessionGate -> InteractiveCard -> questionActive swap -> Composer -> ComposerField, flipped by writing the same store field an AskUserQuestion hook event writes. Flipping questionActive to false fails exactly this test and nothing else across 639 native-chat tests, so the path was entirely unguarded. CHARACTERIZATION TEST: it asserts the loss. Fixing the defect — committing the preedit before the swap, or keeping the composer mounted — will make this file fail. Update the expectations to the new contract rather than working around them. Owns no reported row. #12118/STA-3219 flicker is keyed to token counters, which provably do not remount, and a question card arrives once per question. * test(terminal): pin the duplicated commit when Meta interrupts a composition _finalizeComposition(false) sends textarea.value.substring(start, end) but cannot clear the IME-owned textarea, so a later compositionend re-sends the same range. Meta reaches that path because CompositionHelper exempts only Shift/Ctrl/Alt; xterm's own wasModifierKeyOnlyEvent covers Meta four ways, so the omission is an internal inconsistency rather than a choice. Companion to the modifier-exemption guard, which deliberately pins only the overlay teardown. This pins the data consequence. HAZARD PIN: owns no reported row. The trigger is unverified on hardware — no capture in the corpus contains a Meta-during-composition gesture, and whether macOS keeps the composition alive across it is unmeasured. The duplication follows from the code given that sequence; whether users reach the sequence is the open half. An earlier premise that Space (keyCode 32) reaches this path was refuted by a corpus scan: 0 of 731 evidence files carry a keyCode-32 Space while composing, against 171 at 229, and 229 returns early. * test(terminal): characterize the syllable lost when the textarea blurs mid-composition CoreBrowserTerminal._handleTextAreaBlur clears the helper textarea unconditionally — "Text can safely be removed on blur" — while CompositionHelper._finalizeComposition reads the committed text back out of that same value from a deferred timeout. By the time it runs the value is empty, the substring is '', and triggerDataEvent never sees the syllable. xterm checks composition state in _syncTextArea and omits the same check here. Six cases. Blurring mid-composition loses the syllable in every ordering, including compositionend-before-blur, which is Chromium's real order — so it is not an ordering artifact. A bare textarea.blur() with no Orca code loses it too, which places the owner upstream: Orca's unguarded release on outside pointerdown is one trigger, not the cause. Committing 한 then blurring mid-가 yields ['한'] where ['한','가'] is correct: one syllable gone, surrounding text intact. Teeth checked by inverting — adding an Orca-side composition guard flips exactly the three cases that route through the release path and leaves the bare-blur and no-blur cases green, which is the scope split: a fix in regular-terminal-focus-ownership alone would not close this. HAZARD PIN, but unlike the others this one has a real production injector — clicking outside the terminal mid-composition. Owns no reported row. The shape matches #9738's report; the injector does not, and a shape match with a mismatched injector is not an owner. * test(terminal): say which arm the STA-3237 fixture came from The recorded keydowns are wave 4's A-shift-unmarked-only — the arm that emits no PTY bytes. Nothing in the file said so, so two readers concluded the row's events fail the owner's predicate and that STA-3237 and STA-3222 were different defects. They share an owner; the arm that fires is Process/229+Shift, absent from this bubble-phase trace because the owner claims it in the capture phase. Also corrects "code-blind": the v1.4.163 policy emits \x1b\r only for a shift-only key:'Enter', and a jamo keydown reaches that branch solely via the isTerminalImeProcessEnter rewrite. The mock is deliberately wider so the ownership guard stays under test if that rewrite moves. Comments only — no assertion, fixture value, or mock behaviour changed. * test(e2e): track the input-source selector the macOS specs shell out to Five tracked macOS IME specs ran `swift .tmp/select-input-source.swift`, a file that is gitignored and existed only on one machine. Anyone else checking out the repo — or the same machine after .tmp is cleaned — could not run them, and they are the capture drivers for the macOS rows that are blocked waiting for exactly those runs. Moves it to tests/e2e/ beside its callers. The chord spec now resolves it from __dirname rather than reaching two levels up into .tmp. * test(terminal): pin the CJK repaint decision against the reporter's own output #12164 comment 1 and #5921 report agent output with double-width glyphs rendering duplicated character-by-character while ASCII in the same line stays clean. No IME, no composition, no keystroke — the user never types the CJK. Segmenting all three verbatim samples into maximal same-risk-class runs gives 33 runs and zero violations of "this run is corrupted iff the production detector flags it": 17 wide runs all corrupted, 16 narrow runs all byte-identical. The paired negative is co-located in the same line rather than in a separate run — the reporter supplied it without knowing. Doubling is asserted as present, not uniform: 자바스크립트 and 시스템 each leave a jamo undoubled, which is a repaint-region boundary artifact rather than a per-character transform. The discriminating arm is in the test rather than a source mutation: |
||
|
|
cf16eac7f6 |
fix(agent-hooks): keep Node 18 relay companion loadable (#13135)
Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com> |
||
|
|
20aeb0cb99 |
Prepare mobile 0.0.42 and fix TestFlight CI hang (#12961)
* Bump mobile app.json to 0.0.42 * fix(mobile-ios): stop TestFlight CI from waiting on ASC processing 0.0.42 builds 1–2 uploaded successfully then hung for hours polling processing with no Ready build and no Apple email. Exit after upload and cap the job at 90m so the next cut does not repeat that hang. * fix(mobile-ios): fully skip Pilot wait (no changelog) Pilot only returns immediately after upload when changelog is nil; passing notes re-enters the ASC build-list poll. |
||
|
|
7da9368b78 |
fix(terminal): fence detached daemon endpoint ownership (#12709)
* fix(terminal): fence daemon endpoint ownership * fix(terminal): clean failed daemon PID claims * fix(terminal): close daemon ownership review gaps * test(daemon): release startup IPC in boot smoke * test(daemon): mirror production stdio in boot smoke * fix(daemon): exit after rpc shutdown cleanup * fix(terminal): make the socket name the daemon endpoint authority The reported failure was a live daemon hosting PTYs that nothing could reach: terminals acknowledged input and never ran it, listings diverged from reality, and restarting the app never helped because the detached helper survived. The ownership fence added for it could not fire in the sequence that produces the split brain. libuv unlinks the pathname a server bound to when that server closes, with no ownership check. A daemon that lost its endpoint name therefore deleted whichever socket then sat at that path — including a live replacement's — stranding a daemon that still hosted every session. Bind a private same-directory name and hard-link it into place instead: libuv can only ever unlink our own bind name, the exclusive link is a kernel-enforced endpoint claim, and the canonical name is removed only under an inode ownership check. The bind name replaces the basename rather than extending it, so it cannot overflow sun_path. killStaleDaemon removed the PID record unconditionally immediately before every fork, so the exclusive PID claim was always uncontested at bind time. It also unlinked a live daemon's endpoint whenever a connect probe merely timed out, and treated a `ps` timeout as proof of PID recycling. Now only positive evidence of a dead endpoint authorizes reclaiming it, SIGKILL is confirmed rather than assumed, and a daemon that cannot be proven stopped keeps its record and endpoint while the launcher refuses to fork beside it. A daemon whose endpoint was taken over now retires itself, draining rather than killing, so an unreachable orphan stops being permanent. A repaired PID record re-derives entryPath, appVersion and the Linux incarnation markers from the authenticated owner instead of dropping them; without appVersion a healthy daemon read as a permanently stale bundle and, on Windows, went unpinned against daemon-host pruning. Repair failure now fails open — abandoning a healthy daemon over a pid file write cost every persistent terminal on the machine. Also: treat only ENOENT as an unclaimed record so a Windows file lock is not reported as an ownership conflict; settle start() before close() so an accepted connection cannot defer it forever; sweep abandoned claim and bind names; and type the endpoint-identity seam so a rename cannot silently disable the fence. Adds a real-process handover smoke that reproduces the failure with two daemons racing one endpoint, and wires it into the native-smoke job. * fix(daemon): retire only on proven endpoint ownership loss The ownership watchdog read a null identity for any stat failure, so a transient EACCES or EIO on the runtime directory would retire a daemon that was still serving every terminal on the machine. Distinguish "the entry is gone" from "the probe failed" and act only on the former. Also require the loss to persist across two polls: a replacement publishes by unlink-then-link, and a single observation can land in that gap. * fix(daemon): source repaired ownership metadata from the authenticated hello Adversarial review found three defects in the previous two commits. Re-deriving entryPath from the owner's command line truncated it at the first space. A command line is a single space-joined string, so `C:\Program Files\Orca\...` and `/Applications/Orca 2.app/...` came back as `"C:\Program` and `/Applications/Orca`. getDaemonLaunchIdentity treats a present entryPath as authoritative, so a healthy daemon read as `different_app_path` and was killed and re-forked — worse than the missing-metadata case the derivation was added to fix. Carry entryPath and appVersion as optional fields on the daemon hello identity instead: the daemon already has both from its own argv, and per docs/reference/remote-wire-compatibility.md a new optional field is safe because every reader falls back when it is absent. This also removes a synchronous `ps` spawn from the Electron main thread during startup. `start()` rolled back the PID record even when it never published one. Losing the endpoint link now runs that path, and the ownership-checked unlink briefly renames the incumbent's record aside — enough to strand a live daemon's ownership. Roll back only what we actually wrote. publishDaemonSocketPath read its identity from the canonical name after linking, so a concurrent unlink returned null: no ownership watchdog and no endpoint cleanup on any shutdown path. Read it from the bound name before linking, which shares the inode. Refusing to fork beside an unconfirmed daemon left the user with no daemon at all and no in-app recovery, since restart re-entered the same fence. We have just proved something answers the endpoint, so adopt it in degraded mode: live sessions keep working, fresh terminals run locally. SIGTERM is also individually guarded now — an EPERM fell into the blanket catch and reported "nothing alive", authorizing the very duplicate this fence exists to prevent. Also reset the ownership-loss streak on an inconclusive probe so the confirmations are consecutive, and sweep scratch names before the launch so a failed launch still reclaims them. |
||
|
|
fde816e4ee | move folders (#12758) | ||
|
|
06780260c0 |
test(remote-runtime): run an old client and an old server against current code (#12682)
Mixed versions are the normal state of the remote-server feature: users update clients and servers independently. Until now nothing tested that. Every cross-version claim was made by code reading plus unit tests with hand-written old/new shapes — enough to catch design problems, not enough to catch a real skew regression. This runs the REAL protocol implementations from two builds against each other in one process: the actual host methods and RPC dispatcher on one side, the actual renderer multiplexer on the other, with a transport that reproduces the production asymmetry — each side decodes with its OWN codec and drops frames whose opcode it does not know. A frame survives only if the RECEIVING build understands it, which is what makes this level sufficient without launching two apps. The old side is a genuine checkout extracted from the release tag; the extracted client was confirmed to lack a symbol that exists only on main. Journey: subscribe, first snapshot, input reaching the process, live output, hide/reveal snapshot, transport drop, resubscribe, input landing again — across old->new, new->old, and a current/current control. Every step ends on an observed-state barrier; no sleeps. The oracle asserts the recorded step list, the exact 16-frame named sequence, negotiated capabilities, the exact input the host wrote to the PTY, rendered content, and zero decoder-rejected frames. A host method the stub lacks is recorded by name and asserted empty, so a harness gap cannot masquerade as a wire break. Detection is proven per violation shape, and it attributes each to the correct side: an unnegotiated opcode goes red only where a decoder would reject it, a removed published field goes red only where an old client consumes it, and a legal additive field stays green in all three pairings so the harness will not cry wolf on safe changes. It also documents the three compatibility rules in docs/reference/remote-wire-compatibility.md, linked from AGENTS.md, since they previously existed only as folklore — notably that "decoders reject unknown opcodes" is true for the desktop decoder but NOT for mobile, which silently drops them. Deliberately scoped: terminal stream only. The session-tab sync channel is not covered, nor agent-session publications, file/Git RPCs, mobile E2EE framing, or the relay transport. Two version points, so a regression introduced and reverted between them is invisible. CI selection was verified rather than assumed — `vitest list` confirms 0 matches under the shard's exclude and 4 under the dedicated job — because a lane silently running zero tests is precisely how a host-side defect escaped CI earlier in this series. Closes STA-3469. |
||
|
|
fb27702100 |
feat(updater): restart hourly build numbers per version, restyle the timestamp (#12587)
The number answers "which build of 1.4.163 is this", so carrying it across versions made it meaningless — 1.4.164 opened at 38 for no reason a reader could see. It now counts titles matching the base version being built, so a version bump restarts the series at 01. Deriving it moves from workflow jq into the script, because the number depends on the base version and only the script knows which base the published tags resolved to. Timestamps go from `07-31 13:54` to `Jul 31, 1:54PM`, still Pacific. Co-authored-by: Orca <help@stably.ai> |
||
|
|
c4ae923baf |
fix(ci): vet the adhoc build ref before running it with release secrets (#12161)
The adhoc workflow checked out any requested ref and ran its scripts and electron-builder config with MAC_CERTS, the notary password, and the adhoc publisher token in reach — including refs/pull/* fork code a maintainer could dispatch in one innocuous-looking click. Vet the ref before checkout: PR refs are refused, branches/tags resolve in a bare tree:0 scratch fetch, raw SHAs must be reachable from a repo branch or tag (a partial clone lazily serves PR-only commits by SHA, so name resolution alone is not a trust test), and checkout pins the vetted SHA so a race push cannot swap the commit. Also reference an adhoc-mac-build environment so the secrets can later be fenced off from stale workflow copies via repo settings. |
||
|
|
8ab7d8a110 |
fix(updater): base dev builds on published tags, not main's package.json (#12376)
main's version only moves on `release:` commits, and stable patches are cut from release branches that never merge back. On 2026-08-03 main read 1.4.165-rc.0 for twenty hours while 1.4.165, 1.4.166 and 1.4.167 all shipped, so every hourly built in that window was stamped 1.4.165-hourly.* while carrying code newer than 1.4.167 — and sorted below the stable its user was already running. Resolve the base from the main repo's published tags instead, taking the patch above the highest shipped stable. package.json stays a floor for the case where main leads the tags. Co-authored-by: Orca <help@stably.ai> |
||
|
|
484273844a |
feat(updater): add an adhoc release channel for branch builds (#12051)
* feat(updater): add an adhoc release channel for branch builds Hourly covers main. This covers everything that is not main yet: a dispatchable macOS build of an unlanded branch, published to stablyai/orca-adhoc, so the team can run an experimental feature for a few days instead of reasoning about it from a diff. Adhoc sits at the bottom of the version order — 'adhoc' < 'hourly' < 'rc' < stable — so no routine check can walk anyone onto somebody's branch; only an explicit pinned jump reaches one. It gets its own repo rather than sharing orca-hourly's, because a branch build must not appear in the list a developer riding main is looking at. Signed and notarized exactly like hourly, for the same reason: macOS anchors a notarized app's TCC grants on identifier + team, so an unnotarized build reads as a new client and silently loses file access under Documents/Desktop/Downloads. Tags stamp to the second rather than the minute. Hourly runs under a concurrency group and cannot overlap itself; adhoc builds are dispatched on demand, so two people cutting from different branches inside one minute is ordinary — and a minute-resolution tag would collide and fail the second build after its whole pack-and-notarize run. Channel-specific behaviour now derives from one DEDICATED_REPO_CHANNELS list: repo mapping, macOS-only support, and UpdateSource. The RPC schema that validates releaseChannelOverride was a hand-copied enum missing the new channel, which would have rejected the override on its way to the main process; it reads the predicate now. * fix(updater): merge the duplicated shared/types import Co-authored-by: Orca <help@stably.ai> * fix(ci): default the adhoc build ref to the dispatch branch The Actions UI puts its own "Use workflow from" branch picker directly above the ref field, and picking a branch there is what most people read as "build this". Making the field optional means the obvious action is also the correct one; naming a branch explicitly still wins, so main's copy of the workflow runs rather than a stale one on an old branch. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
2b44e9ed9e |
fix(updater): notarize hourly macOS builds so TCC grants survive updates (#12007)
macOS anchors a notarized Developer ID app's TCC grants on identifier + team, which is cdhash-independent and so survives an in-place update. Without a notarization ticket there is no such stable identity, so every hourly reads as a different client: the grant row stays but stops matching, and file access under Documents/Desktop/Downloads fails with EPERM and no re-prompt. `tccutil reset` fixes it until the next build — and orca-hourly has shipped as many as 14 builds in a day. Skipping notarization was chosen because Squirrel.Mac validates the replacement bundle's signature, not its notarization. That is true, but it is the wrong requirement; the in-place swap was never the problem. Budgets grow to absorb the notary round trip (publish 2x45, job 150), and the App token is re-minted after the build so its one-hour life starts at the first call that uses it rather than during `pnpm install`. |
||
|
|
edb5607e28 |
ci: block new root-level entries (#11903)
* ci: guard repository root additions * fix: clear existing type-aware lint warnings |
||
|
|
ad1e58d966 |
chore: declutter top-level repo layout (#11890)
Remove one-off incident docs and committed test-results noise, move dev/repro/bench tools under tests/tools, and relocate i18next config into config/ so the GitHub root scrolls to the description faster. |
||
|
|
676964b099 | ci: run only changed e2e specs on pull requests (#11834) | ||
|
|
cd2b62ed14 |
feat(updater): name hourly releases by version, build number, time, and sha (#11817)
* feat(updater): name hourly releases by version, build number, time, and sha Hourly releases were titled with their raw tag (`v1.4.163-hourly.202607312054`), which reads as one opaque digit run and does not say which commit it came from. Title them `1.4.163 • 01 • 07-31 13:54 • e698241` instead, and show that same string in the in-app build picker by having the picker render the release's stored name rather than deriving its own label. Composing it in one place means the two surfaces cannot drift. The build number is monotonic across the channel. It is read as the highest number already in use rather than as a count of releases: the prune step trims to 72, so a count would roll backwards after three days and reissue numbers. Drafts count toward it — unlike in the freshness check, which asks whether a commit shipped, this asks whether a number is free, and a stranded draft still holds one. Times are Pacific while the tag's stamp stays UTC. The stamp is a sort key and a local one would repeat an hour at every DST fall-back, making two distinct builds compare equal; the title is only ever read. * fix(updater): fail the hourly build when the release name is missing The workflow checks out `ref: main`, but a workflow_dispatch runs the workflow file from whatever branch was dispatched. A branch that edits this step while main still carries the old script produces an empty name and an untitled release — silent, and only visible once someone opens the releases page. Verified by hitting exactly that on run 30665586904. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
79251d7a98 |
[P2] fix(release,settings): restore signing preflight portability, bootstrap diagnostics, and skill re-check (#11692)
* fix(release): restore the SignPath composite action when cutting from an older ref Co-authored-by: Orca <help@stably.ai> * fix(startup): record a durable diagnostic before the bootstrap fatal-exit guard exits Co-authored-by: Orca <help@stably.ai> * fix(settings): make agent-skill Re-check rescan skill freshness Co-authored-by: Orca <help@stably.ai> * fix(startup): keep the bootstrap fatal diagnostic when the log override is unwritable Create the parent directory an overridden ORCA_BOOTSTRAP_FATAL_LOG names and fall back to the default location when that path still cannot be opened, so a missing parent no longer costs the only account of the failure. Also pins the Re-check freshness rescan to the completed install scan rather than the click. Co-authored-by: Orca <help@stably.ai> * refactor(settings): move the post-recheck surface sync out of the panel Co-authored-by: Orca <help@stably.ai> * fix(startup): retain diagnostics without node fs * fix(skills): keep freshness scoped to the local runtime * fix(settings): register freshness status translations * fix(settings): scope and sequence skill freshness refreshes * fix(settings): refresh freshness across runtime transitions --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
e467b3ff7b |
fix(remote): stabilize shared control and terminal parking (#11656)
* fix(remote): stabilize shared control and terminal parking * fix(remote): harden parking review edge cases * fix(terminal): restore parked local floating buffer * fix(ci): drop superseded paired parking evidence * fix(terminal): preserve floating park watchers * fix(ci): include web client in paired e2e artifact * fix(ci): reuse renderer build for paired e2e |
||
|
|
f998f7ec62 |
feat(updater): add hourly dev channel and build switching (#11250)
* feat(updater): add hourly dev channel and build switching Adds an hourly macOS build channel plus a dev-only surface for switching update channels and jumping to any published build, including older ones. Hourly builds publish to a separate stablyai/orca-hourly repo. The routine update path resolves tags from the main repo's releases atom feed, which exposes only its 10 newest entries — 24 hourly tags a day would evict every stable/RC entry there and leave real users with nothing to update to. Hourly artifacts carry the release bundle id and Developer ID signature so Squirrel.Mac can swap them in place; only notarization is skipped, which in-place updates never check. Version tails are stripped to the base (1.4.160-hourly.<stamp>, not 1.4.160-rc.3-hourly.<stamp>) so hourlies sort below both rc.N and stable and are reachable only by an explicit pinned jump, never by an ordinary check. The picker is revealed by Option-clicking the Updates header, matching the Help menu's existing hidden admin affordance. Pinned jumps set allowDowngrade and release the feed on every settle path so a jump can never leave background checks permanently deferred. * chore(hourly): create orca-hourly and add token provisioning script Adds setup-hourly-release-token.sh, which provisions HOURLY_RELEASE_TOKEN without the value ever reaching stdout, argv, or shell history: it is read with `read -rs`, passed to gh through GH_TOKEN in the environment rather than as an argument (argv is world-readable via ps), piped into `gh secret set` on stdin, and scrubbed by an EXIT trap. Verification creates and deletes a draft release in orca-hourly to prove Contents:write for real rather than trusting the permission checkbox. Drafts are absent from the releases atom feed, so the probe cannot disturb users. Refuses to run without a controlling terminal instead of falling through having set nothing, and refuses to run under xtrace, which would echo the token on every expansion. * fix(updater): address review feedback on the hourly channel Renderer: - Guard listBuilds against out-of-order responses. activeChannel flips once getVersion resolves, and rapid channel clicks stack requests, so a slower earlier load could land last and fill the list with builds from a channel the picker was no longer showing. - Selecting the running build's own channel now clears the override instead of pinning it. There was previously no way back to "follow this build's channel", so merely opening the panel left background checks pinned. - Validate releaseChannelOverride on hydration, matching every other enum-like field in that function. Main: - Exclude pinned jumps from recordCompletedUpdateCheck() in update-available. A dev browsing the picker was persisting lastUpdateCheckAt and suppressing the next real background check for a full day. - parseHourlyVersionStamp now anchors on the whole version and round-trips the parsed fields. It accepted garbage prefixes, and Date.UTC rolled impossible dates forward, so ...hourly.202602300000 rendered as March 2. Workflow: - Publish into a draft and flip it live only after the manifest check. The window between creating the release and verifying its assets previously exposed a tag the picker would offer and the download would 404 on; a draft is invisible to listReleaseBuilds, so a job that dies in that window — including a hard kill by the job timeout, which runs no cleanup step — leaves nothing user-visible behind. - Add a failure handler that discards the draft, gated on the publish step not having succeeded so a later prune failure cannot delete a live release. - Align retry budgets with the job timeout (was 60min against a worst case of ~185min, so a mid-retry kill skipped the cleanup that step exists for). - Exclude drafts from the freshness and retention queries. - persist-credentials: false; the job only reads this repo and never pushes. * refactor(hourly): authenticate with a GitHub App instead of a PAT A fine-grained PAT expires, and the hourly build would then fail silently on a schedule nobody watches. A GitHub App's private key has no expiry, so this is set up once. It is also owned by the org rather than by the person who created it, so the credential survives that person leaving. The workflow mints a short-lived installation token via actions/create-github-app-token and passes it as GH_TOKEN. Installation tokens live one hour, which is ample: this job runs no tests, no notarization, and no Windows signing, so it is pack + upload. The retry budgets and job timeout are re-sized to that reality rather than copied from the release pipeline, whose 3x45 publish budget exists for notarization and SignPath. setup-hourly-release-token.sh now provisions HOURLY_RELEASE_APP_ID and HOURLY_RELEASE_APP_PRIVATE_KEY. The key is redirected from a file straight into `gh secret set` on stdin, so its contents never enter a shell variable, argv, or the terminal. * fix(hourly): make the xtrace guard fire and cover cancelled runs The xtrace guard disabled tracing before testing for it, so `[[ -o xtrace ]]` read the state the previous line had just cleared and never fired. `bash -x` ran straight through, tracing exactly the key handling the guard exists to prevent. Test first, then disable. The draft cleanup only ran on failure(), but a run stopped from the Actions UI is cancelled(), not failed — a manual cancel mid-publish stranded the draft. Cover both. |
||
|
|
cc078a5021 |
perf(main): move hang watchdog into a worker thread (#11488)
* perf(main): add watchdog boundary memory benchmark Add a repeatable Electron 43 RSS harness that measures the production-built watchdog entry across the child-process and worker-thread boundaries. Record per-trial samples, the median, revision, runtime, and settling procedure for reproducible PR evidence. * perf(main): move hang watchdog into a worker thread Keep main-thread hang detection independent of the blocked Electron event loop without paying for a second ELECTRON_RUN_AS_NODE process. Preserve the marker and telemetry contract while moving timing configuration and heartbeats onto a bundled worker entry. * test(main): smoke packaged hang watchdog worker * fix(main): make packaged watchdog smoke able to fail The smoke reported failure only through process.exitCode, but its finally block quit Electron gracefully, and Electron takes its status from the browser exit code. Every failure mode — entry missing from app.asar, worker error, marker timeout, non-zero worker exit — exited 0 with the diagnostic discarded on stderr, so the required PR check could never go red. Propagate a real status via app.exit, assert the success line in stdout, and surface stderr. Verified against a packaged tree with the entry removed: exit 0 before, exit 1 after. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
14de3fa14d |
fix(computer): reap mac helper after client loss (#11493)
* perf(computer): add mac helper owner-loss benchmark Measure the release helper's resident memory before and after its owner-session deadline. Record exact revisions, per-trial RSS, retained state, and clean-exit latency so lifecycle reclamation is reproducible. * fix(computer): reap mac helper after client loss Bind the detached macOS helper lifetime to authenticated socket ownership. Reap the helper after its final authenticated client disconnects, and add a startup deadline for sessions that never authenticate. * test(computer): harden owner benchmark cleanup * test(computer): make owner benchmark cleanup failure-safe * test(computer): close remaining owner cleanup races |
||
|
|
8f7692aa12 |
Fix packaged skills CLI runtime ownership (#11627)
* fix(cli): make packaged skills runtime self-contained * fix(cli): address packaged skills review feedback * ci(cli): smoke packaged skills on Windows --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
a906f98baf |
fix(release): survive PSGallery outages in the Windows signing preflight
The Windows release job hard-failed in run 30125672117: every SignPath module install attempt got 403 Forbidden from the gallery's OData API, which is behind Azure Front Door and was also serving 502/504 at the time. That step was the only hard-fail in an otherwise fail-open signing chain, so a gallery incident blocked the whole release. The gallery CDN that serves the nupkg is a separate origin and stayed healthy throughout, so fall back to a pinned version fetched from it after the normal install path is exhausted. The fallback verifies a SHA-256 pin, since that route skips the gallery's own package validation. Extracted to a composite action so the release job and the signing rehearsal cannot drift apart. |
||
|
|
b370dc0900 | ci(release-cut): include source ref/commit and cutter in SignPath Slack (#10524) | ||
|
|
d0f341ad69 |
fix(computer-use): make modifier clicks interruption-safe (#11451)
* fix(computer-use): make modifier clicks interruption-safe * fix(computer-use): pace modified Windows multiclicks * fix(computer-use): address modifier safety review |
||
|
|
5e00a30e4e |
Decouple feature copy from locale parity (#8512)
* Decouple feature copy from locale parity * Fix undeclared dynamic localization key check * Fix localization code owner |
||
|
|
b339fe0346 |
Fix Node 26 test gate and happy-dom storage (#11434)
* ci: test PR shards on Node 26 * test: isolate happy-dom storage from Node globals |
||
|
|
fe6f929c6e |
fix(terminal): reconcile cross-platform IME composition lifecycle (#11293)
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: JeongUk Park <jeongph.dev@gmail.com> |
||
|
|
dde72f85de |
fix(windows): separate updater from orchestration migration (#11405)
* fix(windows): separate updater from orchestration migration * fix(terminal): attest adopted reveal identity --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
363e478909 |
fix(orchestration): preserve active workers across updates (#11271)
* fix(orchestration): preserve active workers across updates * test(ssh): model absent legacy adoption * test(orchestration): align compatibility contracts * fix(windows): escape updater PowerShell booleans * fix(windows): restore stock uninstall process check * fix(orchestration): keep recovery off renderer startup barrier * fix(orchestration): harden legacy recovery migration * fix(orchestration): close recovery review gaps * fix(orchestration): complete legacy worker cutover recovery * fix(orchestration): preserve legacy workers across updates --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
a7c8b8e071 |
fix(terminal): bound SSH & remote hidden-worktree terminal retention (C1) (#10625)
* fix(terminal): park SSH worktrees like local ones (C1 retention, slice A) SSH ptys were blanket-excluded from hidden-view parking, so a hidden SSH worktree retained every pane forever (C1: renderer heap climbs to the V8 ceiling). SSH bytes transit local main — fact-mode watchers already cover them, and main keeps a headless model served over pty:getMainBufferSnapshot that the SSH reattach path never consulted. - isParkRestorableTerminalPty: snapshot-backed OR (SSH + policy); threaded through both park verdicts, both selectors, watcher coverage, and the watcher start guard. Remote-runtime/fail-open/foreign/null unchanged. - Parked-SSH reveal paints from main's headless model (dimension-matched, ~5k rows) and degrades to the relay 100KiB replay unless the snapshot is a non-empty source==='headless' payload — never a blank/stale paint. - Kill switch: settings.terminalSshViewParking (default on). DESIGN.md records the approved plan and the H1 magnitude non-claim. Co-authored-by: Orca <help@stably.ai> * fix(terminal): bound hidden-worktree retention with a force-park budget (C1, slice B) Un-parkable worktrees (remote-runtime ptys, uncoverable tabs, SSH with the slice-A switch off) had unlimited retention: the parking cap/TTL only ever saw eligibility-passing worktrees, so one bad tab pinned a whole worktree's panes forever. Retention is now memory-bounded, not eligibility-bounded. - terminal-hidden-worktree-retention.ts: retention budget (12 hidden / 45min TTL, sized from the measured 2.5-19MB per-pane V8 cost, DESIGN.md §2) over hidden worktrees ordinary parking can never evict; reuses the hot-retain ranking so last-active exemption, deterministic ties, and deadline-driven rechecks hold. Fail-open/foreign-pty tabs are eviction-exempt (a remount would fresh-spawn and orphan the live shell). - Terminal.tsx: force-parked ids join the parked set AFTER the coverage veto (darkness for uncoverable tabs is the accepted cost); buffers captured via the sleep-flow registry before the unmount render; retention TTL added to the recheck deadlines for budget candidates only. - Verdict stays out of its own effect deps; policy test asserts idempotence and time-monotone membership (flip-loop dwell regression). - Kill switch: settings.terminalHiddenWorktreeRetentionBudget (default on). Co-authored-by: Orca <help@stably.ai> * fix(terminal): demote hidden scrollback for eviction-exempt worktrees (C1, slice C) The retention budget (slice B) must exempt worktrees holding fail-open or foreign-worktree ptys — a remount would fresh-spawn and orphan the live shell — which would leave that class unbounded again. Instead, past the same 45min retention TTL their hidden panes drop to the minimum scrollback tier (measured: ~19MB -> ~1.3MB V8 heap per 50k-row pane; trimmed history is gone by design, reveal restores the configured cap for future output). - terminal-hidden-scrollback-demotion.ts: module-state verdict registry (parked-watcher pattern) with content-equality notify damping; applied in the existing scrollback-rows effect in use-terminal-pane-lifecycle. - selectScrollbackDemotedTerminalWorktrees: pure, TTL-gated, time-monotone. - Retention TTL wakeups now also cover exempt worktrees so demotion fires. - Kill switch: settings.terminalHiddenScrollbackDemotion (default on). Co-authored-by: Orca <help@stably.ai> * fix(terminal): paint the SSH model snapshot inline, not via nested coordinator (C1 slice A fix) applyMainBufferSnapshot runs its own structuralReplayCoordinator.run; calling it from applyReattachPayload (already inside the coordinator when a relay replay exists) deadlocks on the coordinator's tail chain. The model paint now mirrors the daemon-snapshot branch inline (folded scrollback + rehydrate + screen, dimension-matched, escape tail last) and arms the restored-snapshot seq baseline so deferred/live chunks the snapshot covers dedupe instead of double-painting. Also falls through (no early return) so reattachPayloadApplied still latches. Adds the folder-workspace id parity unit case. Co-authored-by: Orca <help@stably.ai> * test(terminal): SSH park+reveal e2e round-trip + as-built design notes (C1) Docker-gated (ORCA_E2E_SSH_DOCKER=1) spec: SSH tab parks behind a decoy and reveal restores marker content at multi-viewport scrollback depth. DESIGN.md records the as-built deltas (inline paint, force-park shape, last-active floor) and the residuals so follow-ups aren't lost. Co-authored-by: Orca <help@stably.ai> * fix(terminal): paint SSH reveal from main's model even when the relay replay is empty (C1 review #1) A relay restart empties the replay buffer; the reveal previously painted nothing even when main's headless model held the session. The reattach now prefetches the model snapshot when no structural replay exists (SSH-shaped ptys only) and paints it inside the coordinator; emptiness is judged on the composed payload (scrollbackAnsi + data + pendingEscapeTailAnsi) so an alt-screen snapshot with an empty screen frame still paints. Co-authored-by: Orca <help@stably.ai> * fix(terminal): decouple scrollback demotion (slice C) from the retention-budget switch (C1 review #2) Per the approved contract each slice reverts behind its own switch: slice C now requires only the master terminalHiddenViewParking plus its own terminalHiddenScrollbackDemotion flag. The TTL wakeup timer fires for demotion candidates even with the budget switch off. No DEFAULT_SETTINGS entries exist for sibling flags (defaults are the '!== false' optional pattern), so no explicit defaults are added. Co-authored-by: Orca <help@stably.ai> * fix(terminal): scope eviction exemption to the tab, not the worktree (C1 review #3) One eviction-exempt tab (fail-open/foreign pty) previously vetoed force-park for its whole worktree, pinning co-located remote-runtime tabs forever. The worktree now force-parks while exempt tabs keep their mounted panes via a per-tab exclusion mirroring the Activity-portal pattern (legacy watcher sync, legacy render, and the overlay cold-parking hook). Ordinary parking is untouched — a worktree with an exempt tab still cannot ordinary-park. Slice C now also demotes exempt tabs' panes as soon as their worktree force-parks under the count budget (they are the only panes left mounted). Co-authored-by: Orca <help@stably.ai> * fix(terminal): demote un-parkable worktrees the force-park lever spared (C1 review #4) The last-active exemption means a single hidden un-parkable worktree never force-parks — and slice C previously only targeted exempt-tab worktrees, so its panes held full scrollback forever. Demotion now also covers un-parkable non-exempt worktrees past the retention TTL that are absent from the force-parked set (last-active spared, or slice B switched off). Membership stays time-monotone for fixed inputs; covered by new idempotence/monotone selector tests. Co-authored-by: Orca <help@stably.ai> * fix(terminal): keep the hidden clock running through transient background-measure windows (C1 review #5) Whole-worktree background mounts (browser-automation bootstrap lease, mobile mounts, agent wakes) open a ~3s self-clearing measure window that previously deleted hiddenSince — every remount restarted the 30s hysteresis and the 45min retention TTL, so a periodically re-mounted force-parked worktree never re-parked. The measure window still pauses parking/eviction verdicts (all selectors skip measuring candidates); only the clock survives, so the prior verdict resumes as soon as the window closes. Visible and portal-holding worktrees still reset the clock. Co-authored-by: Orca <help@stably.ai> * test(terminal): make the SSH park+reveal depth assertion prove the model paint (C1 review #6a) Pad the session with ~180KB of output after the numbered markers so the earliest marker falls outside the relay's 100KiB rolling replay buffer while staying inside main's ~5k-row headless model; asserting marker_1 after reveal now proves the headless-model paint rather than passing under the relay fallback. Co-authored-by: Orca <help@stably.ai> * docs(terminal): rewrite DESIGN.md as the single as-built C1 contract (review #7) One contract matching the code: status IMPLEMENTED around force-park (not the unmount proposal), real kill-switch names with coupling + revert matrices, the true retention-floor formula with measured per-pane and demotion numbers, an explicit when-OOM-is-still-possible paragraph naming the H2 pendingSideEffects residual, the applyMainBufferSnapshot deadlock constraint inside the slice-A section, stable-signal phrasing instead of a capability latch, fail-open AND foreign-worktree exemption class, verified cites, and a planned/landed/follow-up test matrix. Co-authored-by: Orca <help@stably.ai> * fix(terminal): resolve the eviction exemption per pane, not per tab (C1 review #8) isEvictionExemptTerminalTab read only tab.ptyId — the FIRST leaf's pty — while the coverage veto that makes a worktree a retention candidate walks every pane. A split tab whose second leaf held an unrestorable pty therefore failed coverage (→ force-park target) yet looked exempt-free, so force-park unmounted it and orphaned the live shell. The exemption now resolves panes through the same resolveParkedTerminalPaneCandidates, keeping tab.ptyId in the union for the no-layout/no-capture case. Also from the same review round: - force-park's capture passes includeLocalBuffers:false like every other shutdownBufferCaptures caller; it was serializing up to 512KB/pane of scrollback into the store inside a fix meant to bound renderer heap. - Terminal.tsx unmount resets the scrollback-demotion registry — module state with no reset path, read by a pane effect that runs before the host effect that would clear it, so a stale verdict trimmed restore replays. - memoize watcher coverage per tab within the parking pass; the retention candidates re-asked it for every mounted worktree, not just the parked few. * docs(terminal): drop DESIGN.md — the as-built C1 contract moves to the PR body Co-authored-by: Orca <help@stably.ai> * fix(terminal): cap the deferred PTY side-effect queue (C1 residual H2) pendingSideEffects grew without bound under background timer throttling (~64 drained/s vs hundreds queued/s overnight). Cap at 512 entries with oldest-first eviction: titles drop (last-wins), a pending bell latches onto the next survivor, agent-status payloads collapse onto the survivor keeping the newest 16 (last-wins store state, KB-scale strings). Co-authored-by: Orca <help@stably.ai> * fix(terminal): carry command-lifecycle facts through parked watchers (C1 follow-up) Parked fact-mode watchers omitted onCommandFinished/onCommandCode*, so OSC 133;D and Command Code scrape signals went dark while parked. New parked-terminal-command-status.ts ports the store-level subset: git-UI nudge on every command finish, same-turn status-row drop for SSH PTYs (exact mounted-path parity — the foreground tracker refuses SSH ids), and the Command Code working seed / 1500ms done settle. Byte mode scans the same shared parsers for authority-off parity. Local-PTY status drops stay with the mounted pane: they need pty-connection's process-confirm ladder to tell a leaked nested-shell 133;D from a real agent exit. Co-authored-by: Orca <help@stably.ai> * test(terminal): retention-budget force-park e2e with a retentionLimit override (C1 6b) ORCA_E2E_TERMINAL_RETENTION_LIMIT flows preload → e2e-config → getTerminalParkingPolicyOverrides (exposeStore-gated, positive-integer only) so a spec can shrink the force-park budget to 1. The Docker-gated spec opens two remote worktrees on one relay target (second pre-seeded remote repo), disables terminalSshViewParking to make both un-parkable, hides both behind the local context, and proves the older one force-parks while the last-active exemption spares the newest; re-activating the evicted worktree restores the marker tail via relay replay. Co-authored-by: Orca <help@stably.ai> * test(terminal): retention-budget e2e via same-repo remote worktrees (passes docker lane) The first draft added a second remote repo mid-session, whose pane pty spawn misroutes to the local daemon with the remote cwd (pre-existing multi-repo issue, reproducible without any retention override — a seeded local repo plus one remote repo shows the same misroute). The spec now budgets across three worktrees of the ONE connected repo, created through the product createWorktree path (an external git-worktree-add only lands as a detected worktree needing adoption) and polled through the relay's transient post-connect reconnect window. Verified green on the local Docker lane in 20.8s. Co-authored-by: Orca <help@stably.ai> * fix(terminal): prevent remount thrashing during post-measure cool-down ( Implements the C1 retention contract: preserve worktree `hiddenSinceMs` through a background-measure window (so TTL/ranking stay honest), but re-park waits for a full `coldParkDelayMs` cool-down after the measure ends. Without the cool-down, every ~3s measure lease on a past-deadline worktree thrashes remount/reattach. Core changes: - Terminal.tsx: add measure clock (measuringTerminalWorktreeIdsRef) and post-measure cool-down tracking (terminalWorktreeParkCooldownUntilRef); gate parking candidates until cool-down expires. - Extract snapshot replay choreography to shared terminal-snapshot-replay-paint.ts (used by SSH reattach + daemon restore paths). - Add SSH model snapshot timeout (750ms) with fallback to relay replay. - Move cold-park recheck deadline logic to terminal-cold-park-recheck-deadlines.ts; add cool-down deadline to scheduling. - useTerminalTabColdParking: implement matching measure-clock contract with per-tab cool-down gate to keep tab deadlines synced with worktree retention clock. - Add resolveTerminalMountScrollbackRows() to demote new xterms under demoted worktrees (pane births during demotion must take the demoted tier at create). - Add kill switches: terminalSshViewParking, terminalHiddenWorktreeRetentionBudget, terminalHiddenScrollbackDemotion. * fix(terminal): detect Command Code completion in parked mid-turn panes Seed the byte watcher with in-flight turn state from agent status: the watcher is recreated per park cycle with no startup command to arm it, and the banner scrolled away before parking. Also memoize eviction-exempt checks and use SSH PTY ID builder in tests. * fix(terminal): flush pending command-code settles on reveal remount When a parked pane reveals mid-Command Code turn, the new detector cannot re-observe the already-passed idle composer. Cancelling the settle leaves the row stranded at 'working', so dispose now flushes the pending settle instead. Extract readInFlightCommandCodeTurn to shared space and seed detectors with in-flight turns so remounts complete mid-flight commands. Also memoize SSH model probes to prevent double timeouts on reattach. * fix(terminal): remove scrollback demotion (C1 slice C) The scrollback demotion feature for eviction-exempt hidden worktrees is no longer needed. Retention budget limits are now sufficient without this additional bound. Remove the terminal-hidden-scrollback-demotion module, the selectScrollbackDemotedTerminalWorktrees function, and related per-pane demotion logic. * test(terminal): assert bounded probe during stalled reveal Add assertion to verify that a stalled reveal operation makes exactly one `getMainBufferSnapshot` call, ensuring retry logic doesn't introduce redundant probes that would extend the timeout window before relay fallback. * fix(terminal): implement C1 retention budget for hidden parked worktrees Addresses OOM regressions in hidden parked terminals by force-evicting worktrees past a retention budget: at most 12 mounted while hidden, none past 45 minutes (absolute, not exempted by last-active). Eviction is least-recently-hidden-first. Exempt tabs (unrestorable local PTYs) keep their panes to avoid orphaning shells; worktrees are force-parked even if they contain exempts, and their buffers released elsewhere. SSH/remote worktrees serialize buffers pre-eviction for reveal; local worktrees keep daemon snapshots. Command Code's done-settle window is transferred across park/reveal boundaries so the row cannot strand at 'working'. Model probe on SSH reattach is scoped to park-reveal only, not ordinary reconnects. Includes new E2E suite proving the budget actually releases memory. * memoize eviction-exempt terminal tabs to avoid redundant store reads Each tab's exemption check re-reads the store and walks the layout tree. Introduce selectEvictionExemptTerminalTabIds() to resolve all exempt tabs for a worktree in a single pass, then memoize the result in Terminal.tsx and useTerminalTabColdParking. This prevents O(n) store reads when checking exemptions across multiple tabs and ensures the set remains stable across unrelated re-renders. * refactor: reformat hidden-worktree retention comments Reflow to 80-character lines and remove internal ticket references (C1, C1 slice C). * fix(lint): split overlay slot and eviction-exempt tabs under max-lines Static analysis failed because TerminalPaneOverlayLayer (401) and terminal-parked-tab-watchers (304) exceeded oxlint max-lines. Extract the slot component and eviction-exempt helpers into dedicated modules. * test(terminal): stabilize retention budget e2e control arm Stage un-parkable remote pty ids only after both worktrees are hidden, and keep re-staging during the control-arm poll so a late updateTabPtyId cannot flip the decoy back to park-restorable and ordinary-park it before budget engages. * test(terminal): pin retention e2e decoy to a mounted pane snapshot Use the active pane-identity snapshot for the decoy tab instead of all worktree tabs, and re-assert un-parkable ids after the control-arm hold so a deferred/empty tab id cannot fail the budget-off mounted-count check. * fix: memoize terminal eviction exemptions on layout leaf PTYs Splits add leaf panes to the layout store without changing the tabs array. A memo keyed only on tabs misses this change, leaving new panes unexempted for unmount. Include layout leaf PTYs in the exemption memo key so it recalculates when splits occur or PTYs are re-minted. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
a40183389b |
feat: bound direct SSH reconnect fan-out and recovery (#11003)
* docs: design for direct SSH reconnect fan-out Capture the implementation-ready plan for host-qualified, epoch-fenced SSH reconnect recovery after two rounds of multi-model LLM counsel review. * docs: reconcile SSH reconnect fan-out design * docs: close reconnect design consistency gaps * feat: implement bounded direct SSH reconnect recovery * fix: bound direct SSH retry settlement * fix: harden direct SSH reconnect authority * fix: preserve split SSH retry ownership * fix: preserve SSH split continuation authority * docs: record final SSH reconnect validation * fix: preserve SSH authority through retained and detached state * fix: retain SSH authority across delayed split mounts * fix: close SSH authority recovery gaps * fix: fence stale SSH transport replacement * fix: serialize SSH target teardown * fix: settle SSH teardown failures before reconnect * fix: retire failed SSH reset sessions * test: reconcile current main E2E contracts * fix: close direct SSH reconnect review gaps * fix: fence stale SSH reconnect side effects * fix: close final SSH reconnect lifecycle gaps * test: stabilize current-main reliability gates * test: prove plugin navigation containment * test: make plugin navigation oracle authoritative * test: make plugin navigation oracle deterministic * ci: allow sharded e2e suite to finish * test: wait for runtime pane publication * test: classify pane readiness by error code * test: select close persistence terminal by tab identity * docs: mark reconnect implementation validated --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |