mirror of
https://github.com/stablyai/orca.git
synced 2026-09-23 16:02:24 +00:00
0037f7b3a4f3c19aeb9797b46a930ec83bb727a8
227
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0037f7b3a4 |
docs(terminal): the input quarantine is load-bearing, not superseded
G6 lists "no superseded quarantine remains reachable" and this module was assumed to be one. Disabling its single call site reproduces the hazard it exists for — `cho hi; rm -rf x` reaching the shell — so deleting it without a replacement re-opens command execution. The replacement was costed by building it rather than estimated: +26 production LOC to thread the incarnation, ~+33 complete, and the cross-remount state it needs outlives the destroyed pane so it becomes a module about the size of the one deleted. Floor is roughly +140 to delete 88, and it would add a second identity comparison to a gate already failing for having more than one. The decisive part is that the route is not uniformly available: remote runtime results carry no incarnation, old hosts cannot be made to publish one, and mixed versions are the normal state. A paired client reads unknown, which this program's own rule says is not proof — so either every remote reattach surfaces unresolved, or a fallback is needed and the only correct fallback is this module. Whether to amend the clause or accept something weaker on remote hosts is a user decision, so the clause verdict is left as failing rather than quietly reclassified. Co-authored-by: Orca <help@stably.ai> |
||
|
|
a02fa82924 |
docs(terminal): reconcile G6 with the recorded decision and assess its clauses
G6's body still demanded strictly-negative production LOC after the user relaxed it to minimise-and-justify, so the gate had two conflicting pass conditions and no single truth value. Its body now points at that decision. Assessed the remaining clauses against the branch rather than assuming. Two fail structurally: more than one identity comparison and mutation admission path still exist, and `terminal-input-quarantine.ts` is still reachable from two production files. Records why the quarantine is not subsumed by the superseded-PTY fence, which I had assumed and checked. The fence refuses writes aimed at a stale ptyId; the quarantine guards the user's next keystrokes landing on the successor under its current, correct id — a case the fence never sees. Removing it needs the recovery path to surface a different shell as unresolved, not a deletion. Co-authored-by: Orca <help@stably.ai> |
||
|
|
5ed7114c6a |
docs(terminal): record why the duplicate-resume fix was not built
I recommended adding a typed end-reason so a user quit stops looking like a resume candidate, then went to implement it and stopped. `SleepingAgentSessionRecord` already carries three fields that each exist to stop something resuming that should not have — `origin`, `restoreOnTabOpenOnly`, and `automaticResumeBlockedBy` — each traceable to its own incident, consulted at 22 non-test sites. A fourth predicate, however well typed, is the fifth containment cycle. The designs without this bug do not have a better flag; they resume only on an explicit action, into a new terminal id, and make two agents in one terminal unrepresentable in the schema. The first of those is a product decision about whether automatic resume stays a feature, so it is the user's call rather than mine. Co-authored-by: Orca <help@stably.ai> |
||
|
|
6b6a3b666a |
test(ssh): add MaxSessions, lazy-discovery and paired-skew oracles
Three journeys attempted; none promoted, and the reasons are recorded in the ledger rather than rounded up. MaxSessions=1 against real OpenSSH, with the cap read back from `sshd -T` rather than assumed, and remote pids read on the container two independent ways that must agree, each carrying its kernel start time. Two disjoint mutations discriminate — one reddens only the reconnect clause, the other only the two restart clauses. But the disconnect clause is a forward guard: four separate guard removals left it green, so nothing shipped is load-bearing for it. Lazy discovery samples sshd's own accept log and live session census across a 22s window with the in-use host as a positive control. No mutation reddens its third clause alone — the real cross-host lease scoping is load-bearing, but removing it breaks the sibling host during setup, so the failure carries no clause information. The paired-runtime skew spec pairs two real processes at different versions and refuses to run rather than degrade into a same-version pairing that would look green and prove nothing. No production code changes. Co-authored-by: Orca <help@stably.ai> |
||
|
|
d0f747b391 |
docs(terminal): correct the WSL provider-suite diagnosis
The WSL run blamed bash 5.3.9 for the unrelated `local-pty-shell-ready` failure. macOS runs the same bash version and passes 67/67, so the version is not the cause — the trigger is environmental to that distro, and the underlying defect is that the spec asserts an absolute count of OSC markers it does not own. Co-authored-by: Orca <help@stably.ai> |
||
|
|
06cd858dc0 |
docs(terminal): promote Journey 2 to proven on macOS, Linux and physical WSL
The oracle runs on every environment the journey names, and is clause-selective on all three: reverting three-valued `hasPty` reddens only the unknown-not-dead clause, and widening the sole-provider fallback reddens only the stale-generation clause. Selectivity in WSL was established rather than assumed. The spec runs serially, so a red first test reports the others as "did not run" — they were re-run alone under the same mutation and stayed green. Also records that an Orca WSL-mode terminal now starts on that host at all, which it could not before: the distro had no provisioned default Unix user, so every interactive launch blocked on first-run setup. One diagnosis from the WSL run is corrected here rather than repeated: the unrelated `local-pty-shell-ready` failure was attributed to bash 5.3.9, but macOS runs the same bash version and passes 67/67. The trigger is environmental to that distro, and the underlying defect is that the spec pins an absolute count of OSC markers it does not own. Co-authored-by: Orca <help@stably.ai> |
||
|
|
d0d942bb1c |
docs(terminal): record journey evidence that falls short of promotion
Four journeys now have discriminating oracles but none meets its full stated scope, and each shortfall is named rather than rounded up. Journey 2 is one WSL run from promotion. Journey 12's tests are in-process, so they do not close the live-skew gap the original ledger named. Journey 4's cross-host clause cannot be proven by mutation at all — a mux is per target, so its dispose cannot cross hosts, and the cross-host test stayed green under the mutation that reddens siblings. Journey 13 measured one dimension of ten, on lifted predicates rather than through real IPC. Co-authored-by: Orca <help@stably.ai> |
||
|
|
4622f33453 |
docs(terminal): promote Journey 1 to proven on all three platforms
The oracle now runs natively on macOS, Linux and Windows, and its discrimination was watched on each: a mutation reddens it, a restore greens it. On Linux and Windows both mutations were run, and the second reddens only the stale-operation test — so the journey's two clauses are proved independently rather than jointly. Windows is the new evidence. The PowerShell branches added blind at ebffb85a848 executed correctly on their first run: `$PID` expanded to real integers, which also proves the pane shell there is PowerShell-family rather than Git Bash, and `Get-Process StartTime` returned kernel start times 5.4s apart — so a recycled pid could not have passed as a survivor. First journey promoted in this program. The other twelve are unchanged, and the residual limit on "every stale exact operation" is recorded rather than glossed. Co-authored-by: Orca <help@stably.ai> |
||
|
|
27ad530ea6 |
docs(terminal): record the fence's real gap and what peer designs taught
Marks the client-constructed binding proposal as rejected with the three false claims that sank it, and records what shipped instead. States the shipped fence's actual limitation rather than leaving it implied: it compares a binding, not an incarnation, so a respawn under a reused ptyId passes. The obvious remedy is wrong here — the agent-create id is deterministic by design so a replayed create stays idempotent, and randomising it would trade this narrow gap for a duplicate-spawn bug. Also records the ranked lessons from four comparable agent IDEs, chiefly that a typed end-reason at end time is what stops a user quit from looking like a resume candidate. Co-authored-by: Orca <help@stably.ai> |
||
|
|
23fb083ea6 |
docs(terminal): propose one authoritative binding identity
Every defect this program has touched is the same defect: identity compared with the wrong key, or not compared at all. Lease keyed without the pane, reattach using a creating write, folder-workspace ids compared with the instance suffix stripped, local mutating IPC carrying only an id, a live shell classified as expired, liveness unable to say unknown. Proposal: one branded binding type built from fields that already exist and are already persisted, constructible only from an authoritative source, carried by mutating operations, compared by one shared function. Makes a wrong-key comparison a type error rather than the next incident. Under adversarial review, including against the open issue corpus. Not accepted. Co-authored-by: Orca <help@stably.ai> |
||
|
|
29e1f0f889 |
docs(terminal): record the user decision relaxing G6
G6 becomes minimise-and-justify rather than strictly net-negative. The deletion budget the plan assumed does not exist: an entrypoint-rooted import graph found 51 of 53 candidate files reachable and instantiated on live paths, leaving 263 deletable LOC against roughly +1,021 to offset. Correctness may still not be traded for line count. Co-authored-by: Orca <help@stably.ai> |
||
|
|
8ec025e4e4 |
test(ssh): census both durable session partitions on reconnect
Adds a second reconnect scenario and a helper that reads pane records from the local partition as well as the ssh host partition. That split matters: the reattach binding call passes no hostId, so a grafted pane lands in the LOCAL partition and an oracle reading only the host partition passes whether or not the guard is present. Both tests remain forward guards. The second one was reported as discriminating and did not reproduce: with `mayCreate: false` removed from the call site and the app rebuilt, both still passed. Its induction races `pty:kill` against a severed transport, so when the kill lands the lease is cleaned up and there is nothing left to graft. The handoff README is corrected to say so rather than claim a journey. Co-authored-by: Orca <help@stably.ai> |
||
|
|
72f4384a2d |
docs(terminal): track the terminal-session correctness handoff package
The package was untracked under a gitignored `docs/**`, with the un-ignore rules living only in an uncommitted .gitignore edit — a single `git clean -xdf` would have destroyed the authoritative plan. The 814-path construction snapshot is now pushed as `nwparker/react185-authority-snapshot` too; it had no remote ref. Co-authored-by: Orca <help@stably.ai> |
||
|
|
8480d8f26d |
docs(terminal): record that a guard must be pinned at its call site
A refusal that exists and is never passed is indistinguishable from no refusal, and store-level tests cannot tell the difference — they call the store directly. Learned from `mayCreate`, which was correct and had no production caller for several commits. Co-authored-by: Orca <help@stably.ai> |
||
|
|
0960a4c911 |
docs(terminal): record what makes a retention bound safe
Shortening a grace period is the wrong lever. Measuring process time and gating reclamation on an independent observation are what make one safe, and they are what deployed systems actually do. Also records that lifecycle belongs in the attach reply rather than a delivered event — that is what removes the need for a durable per-consumer cursor to guarantee an exit is never lost. Co-authored-by: Orca <help@stably.ai> |
||
|
|
9177022505 |
docs(terminal): record the terminal session behavior contract
Properties stated as observable behavior rather than mechanism, so an oracle written against them survives a change of implementation. Records the weaker, correct form of the timer rule — a timer may never be the sole cause of a destructive action — because recovery budgets and scratch-file age gates are correct code that an absolute ban would condemn. Also notes which mechanisms are deliberately not required, so each has to earn its place rather than arrive with an architecture. Co-authored-by: Orca <help@stably.ai> |
||
|
|
17cfc968cf |
Revert the terminal IME composition-ownership change (#13282)
* Revert "test(ime): restore coverage the composition-ownership change removed (#13168)" This reverts commit |
||
|
|
17b3dff3c4 |
refactor(terminal): return IME composition ownership to xterm (#13128)
* fix(terminal): return IME composition ownership to xterm * fix(mobile): derive terminal input from native replacement ranges * test(mobile): record iOS Japanese IME traces * fix(mobile): preserve native IME replacement ranges * fix(xterm): flush queued application input after IME commit * test(terminal): pin Korean intermediate commit * test: pin Windows IME shortcut ownership * test: replay IBus number candidate commit * fix: preserve native macOS input-method punctuation * refactor(terminal): remove stale mac focus override * fix(mobile): preserve soft keyboard deletion ranges * fix: keep IME-owned palette chords in renderer * fix: stop carried IME shortcuts at renderer owner * fix: preserve carried IME shortcut dispatch * fix: narrow main-owned shortcut actions * test(mobile): pin Japanese IME replacement traces * test(terminal): retain paired native IME trace * fix(chat): preserve browser IME composition ownership * fix(chat): retain macOS IME confirm gesture * fix(chat): expire unmatched IME confirm carry * fix(chat): isolate IME confirmation expiry * fix(chat): retain active IME confirmation * refactor(terminal): remove dead composition handler * feat(ime): add shared Enter-ownership seams for CJK composition The confirming Enter of a CJK composition arrives as two keydowns and the orderings differ by platform: Windows/Linux redispatch the unmarked Enter/13 before keyup, macOS delivers keyup first. A guard reading only isComposing or keyCode 229 misses the redispatch, so surfaces submitted on a confirm. Adds useImeEnterGestureOwnership (carry token, next-frame expiry), a shared ImeEnterGuardedForm for native implicit submission, and the cmdk seam covering 18 CommandInput surfaces at one site. A chorded Enter arms the carry but is never swallowed — the reverse would eat a user's deliberate Cmd/Ctrl+Enter. Both failure modes are pinned by ime-enter-gesture-ownership-contract.test.ts. Co-authored-by: Orca <help@stably.ai> * refactor(terminal): consolidate native input listeners and parked-screen owner Extracts the shared native-input listener installer and renames the parked-screen detector for what it actually does, replacing per-call-site duplication. The listener installer keeps a forgetOptionKeyLocationOnBlur flag so per-window semantics are preserved rather than flattened. Net deletion; no behaviour change intended. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin recorded IME shapes as regression tests Nine regression tests built from hashed affected-platform captures, each with a paired ordinary negative and a discriminating mutation verified to take the file from all-passing to exactly one failure. Covers the Windows MS-Korean Shift family (#12179, #11878, #12151, #11946, #12152) and the Korean TUI line-break rows (STA-3237, STA-3222, STA-3129). STA-3237 pins the empirical 3-Shift / 2-active-composition / 2-newline ratio the device run established — the third Shift produces nothing because Space has already committed. That ratio is not derivable from a static capture. Co-authored-by: Orca <help@stably.ai> * fix(ime): guard Enter-commit surfaces against CJK confirm Applies the Enter-ownership guards across the surfaces whose Enter commits something: publishes, clones, pairs, installs, posts, or persists. Tiered deliberately rather than uniformly. Irreversible and remote-effect sites take the carry token, which also blocks the unmarked redispatch. Locally reversible sites take the oracle check with a one-line comment naming the residual, because a spurious commit there costs one undo. Three numeric fields are left unguarded with the reason in-code: Chromium blanks number inputs at compositionstart, so a confirm-Enter only ever reaches an empty-draft reset. Measured with a CDP probe rather than assumed — a guard that cannot fire is noise. Co-authored-by: Orca <help@stably.ai> * test(ime): teeth-check the Enter guards on every guarded surface One suite per guarded surface, each verified by deleting the guard and confirming the test fails. A green guard test without that check is unverified, not verified. Two shapes pass vacuously in happy-dom and are avoided here: native implicit form submission never fires, and blur() is inert on an unfocused element. Both made "the commit did not happen" assertions pass with the guard removed, so the suites assert the guard's contract directly instead. Co-authored-by: Orca <help@stably.ai> * fix(mobile): keep iOS Korean commits whole through the live-input path iOS Korean reports isComposing: false on every event, so it bypasses the composition guard entirely. The strict owner rejected UIKit's transformed post-change field and sent only the leading jamo — the reported symptom. Prefers the authoritative same-event field text over the predicted text when the supplied operation cannot produce it. Generic: no Korean special-case, no locale classifier, no normalization. Adds the RN-target-keyed submit carry alongside it. Co-authored-by: Orca <help@stably.ai> * test(e2e): make IME capture harnesses fail loudly instead of silently Four instruments recorded silence as success, so a void run scored as a clean one: - readTerminalImeBoundaryTrace returned an empty trace when the probe never installed, making every "nothing leaked" negative pass vacuously - summarizeLatencies([]) returned a perfect zero distribution that passed all three latency thresholds - the macOS Vietnamese spec pinned an input-source ID that does not exist, and failed as though the operator had chosen the wrong source - the expectedLineCount=1 prefix property was undocumented and one edit from silently downgrading a PTY assertion Input sources now resolve by enumeration and name the near-matches on failure. Co-authored-by: Orca <help@stably.ai> * test(terminal): cover Cangjie cancellation and fix a cross-namespace assertion Adds #11951's recorded Cangjie cancel shape to the existing cancellation suite, which covered Pinyin and Sogou but not Cangjie. One keystroke then Backspace arriving as deleteContentBackward with data: null, so the stale preedit is the only thing a fallback could replay. Verified against the historical pre-6cd944c62b3 bundle: the positive fails with ['尸'] where [] is expected, while the ordinary negative stays green. Also fixes the Vietnamese spec, which asserted a TIS-space input-source ID against getKeyboardInputSourceId(). Those two Orca APIs report the same source in different namespaces — TIS nests it under VietnameseIM, the app API does not. The resolver stays as an installation precondition; the assertion matches the leaf. Co-authored-by: Orca <help@stably.ai> * test(e2e): add a real-IME macOS arm for the Korean chord commit The existing korean-ime-terminal-shift-enter-commit spec synthesizes composition over CDP: Input.imeSetComposition sets the preedit directly and Input.insertText performs the commit. Asserting the IME produced events you injected yourself is circular, so that spec cannot certify real-IME behaviour. This arm selects 2-Set Korean via TIS, reads it back live, and injects through System Events key codes, so the OS owns the preedit, the commit instant, and isComposing. PTY byte expectations are preserved verbatim. Covers 2 of the original 4 cases by design. The other two are the Windows/Linux redispatch-before-keyup ordering, which macOS cannot produce and which cannot be selected -- the OS decides it. Reintroducing synthesis to "restore coverage" would reintroduce the circularity. Co-authored-by: Orca <help@stably.ai> * test(e2e): assert the macOS chord arm at the PTY boundary, not the renderer The byte expectations were transcribed from korean-ime-terminal-shift-enter-commit :364/:383, which assert against onData -- a renderer boundary where the terminator is CR. This spec reads the PTY child, where the tty has already converted CR to LF. Names both forms per row rather than swapping the constant, so the conversion reads as evidence that the capture reached past the renderer, as #11936 and #11951 record. Ctrl+Enter's CSI-u sequence is unaffected and is identical at both boundaries. Co-authored-by: Orca <help@stably.ai> * test(e2e): measure composer-to-onData latency and stop dropping IME keystrokes Two defects in the echo latency probe. It hooked onWriteParsed and onRender but never onData, so it measured key->parse->render echo rather than the composer-vs-onData delta the latency rows need. Adds a third hook feeding its own sample set. And `event.key.length !== 1` silently dropped IME keystrokes: Pinyin and Cangjie keydowns arrive as key:'Process' (length 7). Replayed over the captured corpus, the old filter accepted 580 of 4137 Chinese IME keydowns -- it was discarding 80% of them. The new filter matches the shape the owner itself branches on. Attribution charges each onData to the latest keydown rather than a FIFO head, because composing jamo emit no onData at all and a queue would credit a whole composition to its first keystroke. The consumer now asserts sample count before any percentile, so a zero-sample run cannot render as a flawless distribution. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin the WSL shifted-jamo newline shape for #11919 In Korean 2-set, Shift types ordinary letters -- the double consonants and the compound vowels. Each such keystroke reaches Chromium as key='Process', keyCode=229, shiftKey=true. The v1.4.163 classifier matched exactly that pattern with no code guard, so it called those keystrokes Enter, rewrote them to a synthetic Shift+Enter, and injected a newline into the middle of the word -- with no Enter key pressed. That is why the reporters said "no modifier key pressed": they had not chorded Shift+Enter, but they had pressed Shift, to type the double consonant. Asserts the row's own recorded capture: 40 immediate keydowns, exactly 3 of them Shift-carrying inside a single syllable, and an onData stream with one newline per Enter press and none mid-word. Two ordinary negatives keep it from being a blanket mute -- the same session's non-IME keydowns still reach shortcut policy, and an ordinary Shift+Enter still resolves through the real policy. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin the composition commit lag that made Korean type one behind macOS Korean 2-Set commits syllable N only when the first jamo of N+1 arrives, so compositionend and compositionstart land in the same task. A composition-start handler cancelled the pending finalizer that was the only path to triggerDataEvent and ended the session without emitting bytes, so every committed syllable reached onData exactly one syllable late and the backlog cleared only at a Space or Enter. Types continuously with no Enter and no Space -- either would flush the backlog and hide it -- and samples onData at every syllable boundary. Paired with a length-matched ASCII arm that stays green throughout, so the positive is a fact about composition rather than about timing in general. Bisected to a single call site across five builds: pristine, 1.4.155 and 1.4.162 pass, 1.4.163 fails, removing the one call repairs it, restoring it fails identically. That window is exactly the reporter's "started immediately after updating". Co-authored-by: Orca <help@stably.ai> * test(mobile): cover the send-queue abort that silently drops queued keystrokes One failed send in use-terminal-live-input-commit aborts every keystroke queued behind it, with the error swallowed by .catch(() => false). The existing test resolves(true) on every send, so the failure branch was uncovered. Four arms: the abort itself, an ordinary negative on the healthy path, a throwing sender, and a liveness control proving the queue recovers once the chain settles. Deleting the abort takes 4 passed to 3 failed, with the ordinary negative correctly surviving. Scope is stated in the docblock: this is a transport send-queue abort, reachable only via a real disconnect or RPC error. REQUEST_TIMEOUT_MS is 30s, so latency alone cannot reach the branch — consistent with #7094's symptom class, not proven to be its cause. * test(terminal): pin that daemon snapshot/restore cannot disturb a composition Two independent reporters attributed broken Korean composition to the always-on PTY daemon repainting terminal state over the preedit. The attribution is wrong on ancestry — the daemon shipped three months before the version both call good — but the boundary was never actually tested. Runs the real applyMainBufferSnapshot choreography against a live composition, including the full 2J/3J/H wipe plus the resize and alt-screen branches. textarea.value, selectionStart/End, compositionView.textContent and .active all survive byte-identical, and interleaving a restore between every jamo of 문제 still commits 문제 at onData. Also pins that the uncommitted preedit is absent from the captured snapshot: it lives in the textarea, never the buffer, so a restore has nothing stale to echo back. Injecting one textarea.value = '' into the restore fails exactly the three restore-boundary tests. * test(terminal): pin that Cmd tears down a composition where Ctrl and Shift do not xterm's composition keydown exempts only keyCode 16/17/18 (Shift/Ctrl/Alt) plus 20/229. macOS Meta — 91/93/224 — is absent, so a Cmd press mid-composition takes _finalizeComposition(false): the overlay goes dark and never recovers, because compositionstart is not re-fired. The user composes the rest of the word blind. Linux and Windows users press Ctrl and are exempt. xterm already has a Meta-aware modifier predicate in wasModifierKeyOnlyEvent, so this is an internal inconsistency rather than a deliberate choice. Owns no reported row and is version-neutral: 5/5 on both 1.4.162 and 1.4.163. The branch is unexercised in all 328 recorded traces, so this is a hazard pin, not a regression guard. Only the teardown is asserted; the likely duplicated commit needs a compositionend the IME kept alive across the Cmd, which no capture contains. Deleting the exemption fails exactly the three paired negatives; adding Meta to it fails exactly the two Cmd arms. * test(native-chat): characterize preedit loss when a question card replaces the composer An AskUserQuestion card fully replaces the composer by design, but the in-flight composition goes with it: the composer unmounts before compositionend reaches it, so the preedit is never committed to the draft. The committed text survives only because the draft is cached and restored via defaultValue. Node identity changes, value 'abc' is preserved, the 가 is gone. Drives the real NativeChatView -> SessionGate -> InteractiveCard -> questionActive swap -> Composer -> ComposerField, flipped by writing the same store field an AskUserQuestion hook event writes. Flipping questionActive to false fails exactly this test and nothing else across 639 native-chat tests, so the path was entirely unguarded. CHARACTERIZATION TEST: it asserts the loss. Fixing the defect — committing the preedit before the swap, or keeping the composer mounted — will make this file fail. Update the expectations to the new contract rather than working around them. Owns no reported row. #12118/STA-3219 flicker is keyed to token counters, which provably do not remount, and a question card arrives once per question. * test(terminal): pin the duplicated commit when Meta interrupts a composition _finalizeComposition(false) sends textarea.value.substring(start, end) but cannot clear the IME-owned textarea, so a later compositionend re-sends the same range. Meta reaches that path because CompositionHelper exempts only Shift/Ctrl/Alt; xterm's own wasModifierKeyOnlyEvent covers Meta four ways, so the omission is an internal inconsistency rather than a choice. Companion to the modifier-exemption guard, which deliberately pins only the overlay teardown. This pins the data consequence. HAZARD PIN: owns no reported row. The trigger is unverified on hardware — no capture in the corpus contains a Meta-during-composition gesture, and whether macOS keeps the composition alive across it is unmeasured. The duplication follows from the code given that sequence; whether users reach the sequence is the open half. An earlier premise that Space (keyCode 32) reaches this path was refuted by a corpus scan: 0 of 731 evidence files carry a keyCode-32 Space while composing, against 171 at 229, and 229 returns early. * test(terminal): characterize the syllable lost when the textarea blurs mid-composition CoreBrowserTerminal._handleTextAreaBlur clears the helper textarea unconditionally — "Text can safely be removed on blur" — while CompositionHelper._finalizeComposition reads the committed text back out of that same value from a deferred timeout. By the time it runs the value is empty, the substring is '', and triggerDataEvent never sees the syllable. xterm checks composition state in _syncTextArea and omits the same check here. Six cases. Blurring mid-composition loses the syllable in every ordering, including compositionend-before-blur, which is Chromium's real order — so it is not an ordering artifact. A bare textarea.blur() with no Orca code loses it too, which places the owner upstream: Orca's unguarded release on outside pointerdown is one trigger, not the cause. Committing 한 then blurring mid-가 yields ['한'] where ['한','가'] is correct: one syllable gone, surrounding text intact. Teeth checked by inverting — adding an Orca-side composition guard flips exactly the three cases that route through the release path and leaves the bare-blur and no-blur cases green, which is the scope split: a fix in regular-terminal-focus-ownership alone would not close this. HAZARD PIN, but unlike the others this one has a real production injector — clicking outside the terminal mid-composition. Owns no reported row. The shape matches #9738's report; the injector does not, and a shape match with a mismatched injector is not an owner. * test(terminal): say which arm the STA-3237 fixture came from The recorded keydowns are wave 4's A-shift-unmarked-only — the arm that emits no PTY bytes. Nothing in the file said so, so two readers concluded the row's events fail the owner's predicate and that STA-3237 and STA-3222 were different defects. They share an owner; the arm that fires is Process/229+Shift, absent from this bubble-phase trace because the owner claims it in the capture phase. Also corrects "code-blind": the v1.4.163 policy emits \x1b\r only for a shift-only key:'Enter', and a jamo keydown reaches that branch solely via the isTerminalImeProcessEnter rewrite. The mock is deliberately wider so the ownership guard stays under test if that rewrite moves. Comments only — no assertion, fixture value, or mock behaviour changed. * test(e2e): track the input-source selector the macOS specs shell out to Five tracked macOS IME specs ran `swift .tmp/select-input-source.swift`, a file that is gitignored and existed only on one machine. Anyone else checking out the repo — or the same machine after .tmp is cleaned — could not run them, and they are the capture drivers for the macOS rows that are blocked waiting for exactly those runs. Moves it to tests/e2e/ beside its callers. The chord spec now resolves it from __dirname rather than reaching two levels up into .tmp. * test(terminal): pin the CJK repaint decision against the reporter's own output #12164 comment 1 and #5921 report agent output with double-width glyphs rendering duplicated character-by-character while ASCII in the same line stays clean. No IME, no composition, no keystroke — the user never types the CJK. Segmenting all three verbatim samples into maximal same-risk-class runs gives 33 runs and zero violations of "this run is corrupted iff the production detector flags it": 17 wide runs all corrupted, 16 narrow runs all byte-identical. The paired negative is co-located in the same line rather than in a separate run — the reporter supplied it without knowing. Doubling is asserted as present, not uniform: 자바스크립트 and 시스템 each leave a jamo undoubled, which is a repaint-region boundary artifact rather than a per-character transform. The discriminating arm is in the test rather than a source mutation: |
||
|
|
06780260c0 |
test(remote-runtime): run an old client and an old server against current code (#12682)
Mixed versions are the normal state of the remote-server feature: users update clients and servers independently. Until now nothing tested that. Every cross-version claim was made by code reading plus unit tests with hand-written old/new shapes — enough to catch design problems, not enough to catch a real skew regression. This runs the REAL protocol implementations from two builds against each other in one process: the actual host methods and RPC dispatcher on one side, the actual renderer multiplexer on the other, with a transport that reproduces the production asymmetry — each side decodes with its OWN codec and drops frames whose opcode it does not know. A frame survives only if the RECEIVING build understands it, which is what makes this level sufficient without launching two apps. The old side is a genuine checkout extracted from the release tag; the extracted client was confirmed to lack a symbol that exists only on main. Journey: subscribe, first snapshot, input reaching the process, live output, hide/reveal snapshot, transport drop, resubscribe, input landing again — across old->new, new->old, and a current/current control. Every step ends on an observed-state barrier; no sleeps. The oracle asserts the recorded step list, the exact 16-frame named sequence, negotiated capabilities, the exact input the host wrote to the PTY, rendered content, and zero decoder-rejected frames. A host method the stub lacks is recorded by name and asserted empty, so a harness gap cannot masquerade as a wire break. Detection is proven per violation shape, and it attributes each to the correct side: an unnegotiated opcode goes red only where a decoder would reject it, a removed published field goes red only where an old client consumes it, and a legal additive field stays green in all three pairings so the harness will not cry wolf on safe changes. It also documents the three compatibility rules in docs/reference/remote-wire-compatibility.md, linked from AGENTS.md, since they previously existed only as folklore — notably that "decoders reject unknown opcodes" is true for the desktop decoder but NOT for mobile, which silently drops them. Deliberately scoped: terminal stream only. The session-tab sync channel is not covered, nor agent-session publications, file/Git RPCs, mobile E2EE framing, or the relay transport. Two version points, so a regression introduced and reverted between them is invisible. CI selection was verified rather than assumed — `vitest list` confirms 0 matches under the shard's exclude and 4 under the dedicated job — because a lane silently running zero tests is precisely how a host-side defect escaped CI earlier in this series. Closes STA-3469. |
||
|
|
d8e5944b60 |
Stop a duplicate headless orca serve from crash-looping and exhausting AppImage FUSE mounts (#12212)
* fix(startup): stop a duplicate headless serve from crash-looping and leaking AppImage mounts A second Orca launch that loses the single-instance lock called app.quit() before `ready`. That quit is deferred, so the doomed process kept booting into Chromium's Linux display initialization, failed with "Missing X server or $DISPLAY", and died with SIGSEGV. systemd read that as a crash and restarted it forever; each restart re-mounted the AppImage and left the squashfuse mount behind, until the host hit the 1000-mount FUSE ceiling and every later launch failed. The lock-losing launch now calls app.exit(3), which terminates synchronously before any display init. Exit code 3 is a stable "another process already owns this userData profile" contract, and the documented systemd unit uses RestartPreventExitStatus=3 plus a real StartLimitIntervalSec/StartLimitBurst window so a permanently failing launch can no longer retry unbounded. Second-instance argv is now forwarded to the owner, and a duplicate `orca serve` no longer asks the live headless server to open a desktop window. Desktop activation for ordinary launches and macOS dock re-activation is unchanged. Closes #11935 * docs(headless): clear the start limit before the scripted service starts StartLimitIntervalSec=300/StartLimitBurst=5 rate-limits operator starts too, so after a crash-loop trips the burst systemd refuses a plain `systemctl start` for the rest of the window. The Upgrade and Roll back scripts run under `set -euo pipefail`, so that refusal aborted the rollback mid-flight and left the server down on the exact recovery path the doc prescribes. Both scripts (and their EXIT-trap recoveries) now run `systemctl reset-failed` first, the unit reference explains the interaction, and the crash-loop bullet points at it for manual starts. Co-authored-by: Orca <help@stably.ai> * test(startup): reproduce the #11935 duplicate-serve crash loop under real Electron The committed coverage for #11935 was source-text greps, so nothing gated the mechanism the fix rests on: pre-`ready` `app.quit()` is deferred, which is why the lock-losing headless `orca serve` kept booting into Linux display init. This runs two real Electron processes against one disposable profile. The duplicate executes the lock-loss gate's own `app.*` statement, lifted out of `src/main/index.ts`, so reverting to `app.quit()` fails the test. It also feeds the owner's real forwarded argv through `shouldActivateDesktopForSecondInstance`. Also record why the activation predicate matches `--serve` and not the `serve` subcommand: an AppImage launched as `orca serve` exits at the CLI redirect before requesting the lock. * test(startup): wait for the owner process to exit before removing its profile Windows holds the profile's handles for a beat after SIGKILL, so an immediate rmSync can fail with EBUSY/EPERM. Co-authored-by: Orca <help@stably.ai> * test(startup): pass the fixture marker path by env, not argv Chromium reorders argv and the duplicate's argv is itself under test, so a trailing positional was the wrong channel for it. Co-authored-by: Orca <help@stably.ai> * test(startup): only the activation case waits on the owner notification The exit-contract cases assert on the duplicate's own already-terminated process, so they should not block on cross-process delivery. Co-authored-by: Orca <help@stably.ai> * test(startup): drop the staged lock race, keep the real-Electron gate contract CI proved the two-process form cannot work on a display-less Linux runner: Chromium's ProcessSingleton needs the browser IO thread, which needs `ready`, which needs a display. The pre-`ready` owner looked stale and the duplicate took the lock (`expected [ 'DUPLICATE_WON_LOCK' ] to include 'DUPLICATE_LOST_LOCK'`). Lock acquisition and argv forwarding are already covered in single-instance-lock.test.ts. What only a real process can settle is what the loser does next, so that is all this file now runs -- display-independent. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
ed7849eb7b |
fix(worktrees): stop silently switching existing Windows setup scripts to Git Bash (#12406)
* fix(worktrees): stop silently switching existing Windows setup scripts to Git Bash #6967 derived the Windows setup-runner shell from `terminalWindowsShell`. On upgrade, any Windows user whose terminal preference resolved to Git Bash had their existing `orca.yaml` setup script (and issue command) handed to bash instead of cmd.exe. Scripts authored against the cmd runner — `copy`, `xcopy`, `set VAR=value`, `if errorlevel 1`, `%VAR%`, backslash paths — broke with no migration and no warning, and the failure looked like Orca broke the project. The conflation is also wrong in the steady state: a terminal preference is per-user, so two people on the same repo got different interpreters for the same orca.yaml and no project could write a setup script that worked for all of its Windows contributors. The interpreter is now a property of the script, declared the standard way: a leading `#!` line. Native Windows keeps the historical `.cmd` runner unless the script declares a POSIX shell, so no existing script changes behavior. `resolveSetupRunnerShell` keeps its role as the feasibility gate — a bash runner still requires the terminal to resolve to Git Bash, since the launch command is typed into that shell and uses MSYS `/c/...` paths. `buildWindowsRunnerScript` now drops a leading `#!` line rather than `call`ing it, so a declared-bash script that falls back to cmd (Git Bash missing) fails on a real setup line instead of aborting on errorlevel at line one. WSL worktrees, POSIX platforms, and SSH hosts are untouched. * fix(worktrees): keep the cmd setup runner launchable from a Git Bash pane Adversarial review of this PR found that pinning the runner format per script reopened issue #6896 one layer down. - `WorktreeSetupLaunch.shell` had been redefined to mean "the format the runner file was written in". `resolveSetupRunnerCommand` consumes it as "the shell that types the launch command", so a Git Bash terminal with a batch setup script produced `cmd.exe /c "C:\...\setup-runner.cmd"` typed into a bash pane, where MSYS rewrites the `/c` switch into a drive path: cmd opens interactively and setup never runs. `shell` is the terminal's family again; the runner file's .cmd/.sh extension carries the format, and a batch runner launched from a POSIX pane reuses the existing PowerShell ProcessStartInfo launcher. - The cmd runner dropped a leading `#!` line and ran the rest as batch, so a bash script reaching cmd (PowerShell/cmd terminal, or any SSH-to-Windows host) got its interpreter-agnostic prefix executed before failing mid-way. It now prints why and exits 1 without running anything. - A `#!` line's option flags were discarded: `#!/usr/bin/env -S bash -euo pipefail` lost pipefail because the runner is launched as `bash <path>`. The generated posix runner now replays declared flags via `set` and drops the duplicate interpreter line. - Docs cover the per-user setup command in repository hook settings, which goes through the same `#!` rule, and describe what the `#!` line does and does not select. Tests: composed launch command for a POSIX pane + cmd runner (hooks, shared runner command, setup sequencing gate, observed-setup signal), the cmd runner's shebang refusal, and shebang flag replay. Each fails with the source reverted. * fix(worktrees): replay only real `set` flags and keep the gate in the pane's shell Two round-2 review findings: - `#!/bin/bash -l` replayed `set -l`, which exits 2 and aborted the runner under its own `set -e` before a single setup line ran (all platforms). Only the flags `set` documents are replayed now; a bare `-o` with no option name is dropped instead of dumping the shell-option table. - The wait-for-setup gate picked its language from the runner file, so a batch runner launched from a Git Bash pane got the PowerShell gate while the agent startup command was already POSIX-quoted — `Invoke-Expression` cannot parse `'\''`. The gate now follows the pane; the runner still launches through the ProcessStartInfo launcher, never through bash. --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
96c954f3be |
chore: remove force-added design docs from docs/ (#11891)
Keep only the durable docs already allowlisted for tracking (STYLEGUIDE, assets, localized readme, and reference compatibility guides). Drop feature design notes, plans, and repro artifacts that were force-added past the existing docs ignore rules. |
||
|
|
676ef7fab8 |
feat(cli): add orca skills install and orca skills update for headless skill setup (#9201)
Adds `orca skills install` and `orca skills update` so skills can be set up without the GUI — SSH hosts, containers, CI. Previously `orca skills` had only `list` and `get`, so there was no headless path. **Agent targeting is scoped explicitly rather than delegated to detection.** The `skills` CLI decides which agents to install into, and with `-y` and zero detected agents it takes `targetAgents = validAgents` — all ~75. That is not a corner case for a headless CLI: a fresh SSH box or container with no agent installed is the normal starting state. Measured on a bare host, the unscoped command created **52 top-level agent directories and 54 junctions** (one real payload in `~/.agents/skills`, the rest links) on Windows, and 52/53 on macOS. The CLI now passes `--agent` derived from Orca's own detection, mapped to the `skills` key namespace, plus `universal`. Supplying `--agent` makes `runAdd` use it directly and never call `detectInstalledAgents()`, so the fan-out branch is unreachable. On a bare host it now refuses with `No coding agent detected on this host` and exit 1, creating nothing. Same command with scoping: **1 directory, 0 junctions.** `universal` alone would under-install — Claude Code is not in that set, and 19 of 28 mapped keys write agent-private homes `universal` never touches. `--agent '*'` is the bug itself. The mapping is hedged three ways: `null` for any agent whose key could not be confirmed, `satisfies Record<TuiAgent, …>` so a new Orca agent is a compile error, and a test pinning every mapped key against the CLI's own valid list. Fixed during review — two holes that each restored the full fan-out through a different door: - `--agent ','` trimmed to nothing, which skipped the refusal *and* emitted no `--agent`. - `--agent -y` passed an emptiness check, and the vendor CLI silently drops `-`-leading values, re-emptying its list. The real invariant is argument *shape*, not emptiness, and it is now enforced at the choke point in `buildAgentFeatureSkillInstallArgs`, so no caller can emit `-y` without a usable target. `*` remains allowed — asking for every agent explicitly is a choice, not an accident. Verified with 51 hostile inputs through the built binary, each recorded argv replayed through the vendor's own parser. Also fixed: the `ORCA_CLI_CWD` refusal now runs before target resolution (it was quoting the wrong host's agent list), and `--dry-run` is refused in a forwarded shell rather than printing a command naming the wrong machine. Validated on a real Windows host across PowerShell 7, PowerShell 5.1, cmd.exe and Git Bash: `.cmd` shims route through `cmd.exe` and `.exe` shims spawn directly (proved with instrumented shims, not inferred), the ENOENT path produces an actionable error rather than a silent failure, and `skills update` genuinely restores a corrupted skill byte-for-byte. Known, not addressed here — both upstream behaviours this only forwards: a partial install failure exits 0, and "no installed skills found" exits 0. Both are invisible to the headless callers this feature exists for. Co-authored-by: scastanoh21 <scastanoh21@gmail.com> |
||
|
|
dde72f85de |
fix(windows): separate updater from orchestration migration (#11405)
* fix(windows): separate updater from orchestration migration * fix(terminal): attest adopted reveal identity --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
363e478909 |
fix(orchestration): preserve active workers across updates (#11271)
* fix(orchestration): preserve active workers across updates * test(ssh): model absent legacy adoption * test(orchestration): align compatibility contracts * fix(windows): escape updater PowerShell booleans * fix(windows): restore stock uninstall process check * fix(orchestration): keep recovery off renderer startup barrier * fix(orchestration): harden legacy recovery migration * fix(orchestration): close recovery review gaps * fix(orchestration): complete legacy worker cutover recovery * fix(orchestration): preserve legacy workers across updates --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
3a80fbe162 |
Revert terminal rendering changes from #10692, #10794, #10871, and #10907 (#11338)
* Revert "fix(terminal): avoid flash while restoring parked terminals (#10871)" This reverts commit |
||
|
|
a40183389b |
feat: bound direct SSH reconnect fan-out and recovery (#11003)
* docs: design for direct SSH reconnect fan-out Capture the implementation-ready plan for host-qualified, epoch-fenced SSH reconnect recovery after two rounds of multi-model LLM counsel review. * docs: reconcile SSH reconnect fan-out design * docs: close reconnect design consistency gaps * feat: implement bounded direct SSH reconnect recovery * fix: bound direct SSH retry settlement * fix: harden direct SSH reconnect authority * fix: preserve split SSH retry ownership * fix: preserve SSH split continuation authority * docs: record final SSH reconnect validation * fix: preserve SSH authority through retained and detached state * fix: retain SSH authority across delayed split mounts * fix: close SSH authority recovery gaps * fix: fence stale SSH transport replacement * fix: serialize SSH target teardown * fix: settle SSH teardown failures before reconnect * fix: retire failed SSH reset sessions * test: reconcile current main E2E contracts * fix: close direct SSH reconnect review gaps * fix: fence stale SSH reconnect side effects * fix: close final SSH reconnect lifecycle gaps * test: stabilize current-main reliability gates * test: prove plugin navigation containment * test: make plugin navigation oracle authoritative * test: make plugin navigation oracle deterministic * ci: allow sharded e2e suite to finish * test: wait for runtime pane publication * test: classify pane readiness by error code * test: select close persistence terminal by tab identity * docs: mark reconnect implementation validated --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
97cb32c1cc |
fix(terminal): release an abandoned synchronized-output frame on reveal (STA-2694) (#10907)
* fix(terminal): release an abandoned synchronized-output frame on reveal Alt-screen agent TUIs (OpenCode/OpenTUI, Codex, grok) bracket every repaint in `?2026h … ?2026l`. Hiding a pane mid-bracket — which a worktree switch or cold-park lands on routinely, since these brackets are written many times a second — leaves xterm's `decPrivateModes.synchronizedOutput` latched. RenderService.refreshRows checks that latch *before* rendering, so while it holds, every repaint Orca owns is a no-op: the forced render-pause repaint, the plain `refresh()` fallback, and the shared glyph-atlas rebuild all render zero rows while the xterm buffer is perfectly correct. Release the latch at the two reveal repaint entry points so those repaints actually paint. Also adds an OpenCode-shaped alt-screen e2e fixture and spec. The existing inline-TUI convergence spec covers the normal-buffer shape (live block glued to the bottom, history scrolling into scrollback); this covers the full-screen alternate-buffer shape, where nothing scrolls and so no row ever self-heals through the scroll path. Scope note: xterm arms a 1s watchdog that clears this latch on its own, so this closes a bounded window rather than the whole STA-2694 report. The e2e spec passes with and without the production change for that reason; the unit tests are what pin the behavior. Refs STA-2694. * fix(terminal): clear the render model on the plain-refocus repaint path `schedulePaneRevealPresent` — the atlas-preserving path a plain window refocus takes — only called `terminal.refresh()`. xterm's renderers are diff-based: `_updateModel` early-continues on any cell whose code/fg/bg/ext still match the cached model, so a refresh repaints nothing for a pane whose buffer never changed. When an occluded window loses its canvas contents while that model stays populated, the refresh skips exactly the cells that went stale and the pane keeps compositing pre-hide pixels — until a window resize reallocates the model, which is the repair users find by hand. Clear the model first (`RenderService.clear()` → renderer `clear()` → `_clearModel(true)`) so the refresh becomes a guaranteed full repaint. That drops cached cells and glyph vertices but NOT the texture atlas, which is shared by every same-config terminal and whose mid-stream wipe re-arms xterm's page-merge garble race (xterm.js #4480) — the reason this path is atlas-preserving in the first place. Also covers the DOM-renderer fallback in `resetWebglTextureAtlas`: `clearTextureAtlas()` is what invalidated the model on the WebGL path, so a pane without an addon had nothing invalidate it and hit the same skip. Scope note: the e2e spec guards buffer/geometry convergence across the hide/reveal boundaries and adds idle-agent and headful desktop-hide cases, but it cannot observe a stale canvas — both oracles built for that (canvas-vs-buffer ink sampling, screenshot-vs-forced-repaint) were proven blind by injecting the defect, and the spec header documents why. The unit tests pin the ordering and the atlas-preservation invariant. Refs STA-2694. Co-authored-by: Orca <help@stably.ai> * docs(terminal): hand off the STA-2694 reveal-artifact investigation Records both fixed defects with their xterm mechanisms, the reveal/wake call graph, why every e2e oracle for a stale canvas was proven blind, how to arm the in-app render-desync sentinel on real hardware, and the one unverified lead (dimension staleness) that would explain why a window resize specifically is the repair users find. Refs STA-2694. Co-authored-by: Orca <help@stably.ai> * Revert "fix(terminal): clear the render model on the plain-refocus repaint path" This reverts commit |
||
|
|
6b16c20796 | fix(memory): clarify Resource Manager accounting (#10821) | ||
|
|
a0944cc129 |
fix(linux): restore Ubuntu 20.04 launch — pin node-pty glibc symbols + add glibc/libstdc++ packaging gate (#9902) (#10019)
* fix(linux): restore Ubuntu 20.04 launch by pinning node-pty glibc symbols (#9902) The bundled node-pty pty.node is compiled from source in release CI on ubuntu-latest (glibc 2.39). glibc's 2.32-2.34 libpthread/libutil merge relocated openpty/forkpty (GLIBC_2.34) and pthread_sigmask (GLIBC_2.32) into libc under new symbol versions, so the from-source build bound to versions absent on Ubuntu 20.04 (glibc 2.31). The main process imports node-pty at startup, so the app crashed on launch. pty.node is the sole blocker (Electron needs GLIBC_2.25; other native modules <= 2.17). - Patch node-pty: a .symver shim pins the 3 symbols to their pre-merge version (GLIBC_2.2.5 x64 / GLIBC_2.17 arm64), and Linux-only ldflags force libutil.so.1/libpthread.so.0 back into DT_NEEDED. Guarded to Linux; macOS/Windows untouched. - Add a packaging gate (verify-linux-glibc-floor.cjs, afterPack): reads each bundled native binary's objdump -p version needs and fails the Linux build if any strong GLIBC_/GLIBCXX_/CXXABI_ node exceeds stock Ubuntu 20.04 (glibc 2.31 / GLIBCXX_3.4.28 / CXXABI_1.3.12). Catches GLIBC_ABI_DT_RELR, rejects GLIBC_PRIVATE, skips weak needs, fail-closed. - Docs + tests; the lazy sherpa-onnx speech prebuilt (GLIBCXX_3.4.29, never loaded at launch) is a documented libstdc++-floor exemption. * fix(linux): assert DT_NEEDED provider deps in the glibc-floor gate Harden the packaging gate (flagged in adversarial re-eval): the version-floor check alone can false-pass if the patch's forced `-l:libutil.so.1` ever silently drops — the pinned openpty@GLIBC_2.2.5 still resolves from libc's compat alias at build time, but fails to load on Ubuntu 20.04 where openpty/forkpty live only in libutil. The gate now also asserts that any binary importing openpty/forkpty keeps libutil.so.1 in DT_NEEDED. Validated on a real symver-pinned .so with libutil dropped (now fails) vs. present (passes). Documents the recommended real-host smoke-test follow-up. |
||
|
|
b232df732b | fix(terminal): make remote agent sessions host-authoritative (#9687) | ||
|
|
f9f3cd2fbe | fix(terminal): prevent reconnect from killing live daemon sessions (#9804) | ||
|
|
34c160442f | Fix headless Linux serve pairing readiness (#9785) | ||
|
|
cc44acaaa3 | Fix Windows ConPTY OSC color reply leaks at the PTY owner (#9651) | ||
|
|
aad34cbb32 |
docs(headless-server): add upgrade SOP for orca serve on Linux (#9575)
* docs(headless-server): add upgrade SOP for orca serve on Linux The headless Linux guide covered install/run/systemd but had no upgrade section, leaving operators to guess how to move to a new AppImage without losing state. Add an "Upgrade" section documenting the manual SOP (serve mode never auto-updates) and one troubleshooting bullet: - State lives under the service user's ~/.config (orca + Orca dirs), independent of /opt/orca, and orca-data.json is forward-migrated on load, so a forward upgrade is safe. - Replace the binary with an atomic same-filesystem rename (download to .new, verify, mv) — never curl -o over the FUSE-mounted live binary. - Back up the whole .config before upgrading, because rollback is NOT binary-only safe: an older build strips newer orca-data.json fields it doesn't recognize, and the .bak.* ring is corruption-recovery, not a pre-upgrade copy. - Note there is no headless version command; track the release tag instead. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(headless-server): harden the orca serve upgrade/rollback runbook Address CodeRabbit review on #9575: - Fail closed: run the upgrade block under `set -euo pipefail`, remove any stale `.new` file before download, and gate the atomic `mv` on an explicit ELF check so a failed/partial/non-ELF download can never be promoted. - Keep /opt/orca/VERSION tied to the installed binary: a single `TAG` variable drives both the download URL and the recorded VERSION, saved as VERSION.prev on upgrade and restored on rollback so the audit file never drifts. - Crash-loop troubleshooting now points to Roll back first (restores the pre-upgrade orca-data.json) instead of re-running Upgrade. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(headless): harden server upgrade SOP --------- Co-authored-by: fanyunqian.1 <fanyunqian.1@bytedance.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
c0f0810dd9 |
Fix native Windows PTY startup query handling (#9500)
* Fix native Windows PTY startup query handling * Fix daemon boot smoke protocol lookup * Fix Windows daemon repro protocol lookup |
||
|
|
7adda25b0a |
fix(daemon): retire empty current-generation daemons (#9277)
* fix(daemon): retire empty current-generation daemons Co-authored-by: Orca <help@stably.ai> * fix(daemon): retire empty daemons on disconnect Co-authored-by: Orca <help@stably.ai> * test(daemon): authenticate Windows lifecycle harness Co-authored-by: Orca <help@stably.ai> * test(daemon): assert remaining shutdown budget Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
1def694e80 |
Add native chat skill picker with host-aware discovery (#9480)
* Add native chat skill and command picker with host-aware discovery Adds a unified, keyboard-first skill and command picker to native chat that: - Uses agent-native invocation syntax (slash for Claude/OpenClaude/Grok, dollar for Codex) - Discovers skills only on the pane's execution host (local, WSL, SSH-unavailable, or runtime) - Groups or separates commands and skills per agent configuration - Deduplicates by canonical path but preserves visibility through all contributing roots - Handles IME composition, loading states, and errors without claiming PTY-level control - Records picker telemetry (open, item accepted, send classification, discovery outcomes) - Extends shared agent profiles to define per-agent skill grammars and source ownership * Remove obsolete reference and design documentation Clean up stale design specs, implementation plans, and investigation notes from docs/reference/. These documents predate the current implementation and are no longer actively maintained or referenced by the codebase. * Extract shared skill discovery utilities and add skill invocation envelo - Move skill comparison and source classification to shared module for native/WSL reuse - Extract display text sanitization to prevent control/zero-width character spoofing - Add native-chat command envelope parser and surfacer for skill invocations - Extend discovery timeout backstop to account for WSL metadata read sequence * Localize skill picker UI for Spanish, Japanese, Korean, Chinese Translate skill picker UI strings including commands, skills, loading states, error messages, and scope labels for the new skill picker feature across four language locales. * Fix skill picker bugs and improve code robustness - Fix i18n plural handling: rename `count` to `sourceCount` to prevent unintended plural-key resolution in localized strings - Fix skill discovery array mutations: copy `root.providers` to prevent bugs during dedup merge - Fix image attachments being silently dropped when message text starts with /skill or agent prefix - Extract `quoteBashString` utility for WSL command code reuse across builders - Add line-separator safety characters (0x2028/0x2029) to skill display filter - Remove stale doc reference links and clarify inline comments * Add reference docs for git compatibility and headless Linux server setup Track previously untracked operational guides in `docs/reference/` that explain Git binary compatibility requirements across host types and how to run `orca serve` on headless Linux. Update AGENTS.md and README.md to link to these references. |
||
|
|
6e91ca6c0e |
fix(browser): local HTTPS Try HTTPS + cert proceed (#8454) (#9104)
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> |
||
|
|
319ae4e9ea |
fix(terminal): make whole-tab close durable (#8958)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
6e2a4a824d |
fix(worktrees): stop surfacing prunable git worktrees as live workspaces (#8409)
* fix(worktrees): stop surfacing prunable git worktrees as live workspaces A worktree still registered in git but whose directory was deleted (git's `prunable` state) was enumerated as a normal workspace, producing repeated pty:spawn DaemonProtocolError / fs:readDir ENOENT loops and a blank pane. - Parse the `prunable` porcelain field (Git >= 2.36) in both the main and relay worktree-list parsers. - For Git < 2.36 (no `prunable` field), probe each linked worktree path for existence on the fallback line-block path, skipping locked registrations to mirror git's own prunable rules. - Omit prunable worktrees from the detected-workspace enumeration only; removal/cleanup flows keep seeing them. - Extend the real-binary compatibility contract with the 2.36 `prunable` boundary. Fixes #8389 Claude-Session: https://claude.ai/code/session_018Rg1Bpq4GGwmz613hq6RSD * fix(worktrees): pin the prunable/locked porcelain annotations to their real Git 2.31 boundary The prunable and locked annotations landed in Git 2.31, five releases before `worktree list -z` (2.36); only -z defines the capability fallback boundary. Correct the compatibility contract so a future matrix entry in the 2.31-2.35 range passes, and reword the fallback comments: on 2.31-2.35 the annotations still parse and the existence probe is a backstop; only Git <2.31 relies on it outright. * fix(worktrees): omit prunable registrations from the Space scan A prunable registration has no directory to size or reclaim, so Space rendered it as a dead "Missing" row whose checkbox stayed disabled with no prune/remove affordance (reported on macOS after a reboot cleared /private/tmp under 16 registrations). Skip prunable entries in the scan, matching the workspace enumeration; removal flows list worktrees separately and still see them. --------- Co-authored-by: kaynan <kaynan.camargo@terceiro-sky.com.br> Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
40d0159926 |
fix(terminal): kill agent descendant processes on session teardown (STA-1800) (#8706)
* fix(terminal): kill agent descendant processes on session teardown (STA-1800) Agent CLIs spawn tool children in detached process groups that PTY SIGHUP can never reach. Killing an agent session (tab close, retire, sleep) left those children running as orphans — eight orphaned git processes burned ~8 cores for up to 11.5h under the agents-running keep-awake and drained a battery to 8%. New pty-descendant-termination module: snapshot the ppid tree BEFORE signalling (a dead root's descendants reparent to pid 1 and become unfindable), SIGTERM the root group and every descendant, then after a 2s grace SIGKILL survivors gated on a pid+start-time identity re-check so a recycled pid is never signalled. Snapshot is bounded and never rejects; failures degrade to today's shell-only kill. Wired for agent sessions only (plain terminals keep nohup semantics) at all three POSIX kill sites: local provider shutdown, daemon TerminalHost immediate kill (the pty:kill path — force-kill bypassed Session.kill entirely), and daemon Session graceful kill. Verified live in the built app: an agent pane with a detached-pgid child; the child survived on the unwired build (three control runs) and dies within ~5s with the fix. Windows ConPTY and SSH-hosted PTYs keep the previous foreground-tree contract (documented follow-ups). * fix(terminal): harden descendant teardown * fix(terminal): require fresh process snapshots * fix(terminal): close descendant teardown races * refactor(terminal): preserve teardown line budget * fix(terminal): keep descendant teardown fresh and identity-safe * docs(reliability): record integrated descendant E2E * fix(terminal): bound descendant teardown work * fix(terminal): share descendant snapshot indexes * docs(reliability): record descendant review evidence * revert: remove speculative descendant hardening |
||
|
|
0302ae86b8 |
feat(ssh): support Kerberos/GSSAPI hosts via the system OpenSSH transport (#7507)
* feat(ssh): support Kerberos/GSSAPI hosts via the system OpenSSH transport ssh2 has no gssapi-with-mic support, and adding it would mean forking its protocol layer plus packaging the kerberos native module for three platforms. Instead, route GSSAPI hosts through the existing system-OpenSSH transport, which delegates Kerberos (tickets, SSPI on Windows) to the platform ssh binary. Two tiers, because RHEL-family distros enable GSSAPIAuthentication globally in /etc/ssh/ssh_config and ssh -G therefore reports it for every host: - Targets whose ~/.ssh/config Host block explicitly sets GSSAPIAuthentication yes (imported as target.gssapiAuthentication) try system ssh first, falling through to ssh2 so key auth and credential prompts still work when no ticket is available. - When ssh2 exhausts key/agent auth and the ssh -G-resolved config enables GSSAPI, retry over system ssh before prompting for credentials, so Kerberos-only hosts on distro-default configs connect without a password prompt. Hosts where keys work never leave the ssh2 path. Manual targets flagged for GSSAPI pass -o GSSAPIAuthentication=yes explicitly since they bypass ssh_config. Both tiers work headless (no credential callbacks required). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ssh): harden GSSAPI transport selection (review fixes for PR #7507) Review fixes on top of the Kerberos/GSSAPI feature branch (s546126/kerberos-ssh): - HIGH: reset useSystemSshTransport on the ssh2 fall-through. doSystemSshProbe sets the flag before spawnSystemSshCommand, which throws synchronously when no system ssh binary is on PATH (outside the probe try/catch). The proactive fall-through previously reset only 2 of 3 transport fields, so exec/sftp kept routing through the failed transport - breaking GSSAPI on Windows-with-Git-ssh and headless Linux. - MEDIUM: throw a cancellation error (not the stale ssh2 authError) when a disconnect supersedes the reactive probe mid-flight, and guard connect()'s catch on disposed, so a deliberate disconnect is not overwritten with auth-failed. - MEDIUM: skip the encrypted-key passphrase prompt when the GSSAPI fallback applies, so a Kerberos ticket is tried before prompting; the general prompt still fires if the probe fails. Adds 3 mutation-verified regression tests and hardens two existing tests to assert the probe actually ran. Not connected to any PR remote. Co-authored-by: Orca <help@stably.ai> * fix(ssh): isolate GSSAPI system transport Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: s546126 <268420947+s546126@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
a8a8040589 |
fix(agent-hooks): make Windows cmd hook launcher directly spawnable (#8430 regression) (#8737)
Codex, Antigravity, and Devin launch their agent-hook `command` as a program (argv[0]), not through cmd.exe. PR #8430 changed wrapWindowsCmdHookCommand to emit an `if exist "path\." (drain) else if exist "path" (call "path") else (drain)` compound whose argv[0] is the cmd builtin `if` — unspawnable — so every Codex/Antigravity/Devin hook (SessionStart, UserPromptSubmit, Stop, ...) failed with "hook exited with code 1" on Windows starting in v1.4.138. Revert the cmd-safe fast path to the bare, directly-spawnable .cmd path (the proven pre-#8430 form). A cmd-builtin drain and direct-spawnability are mutually exclusive, and nesting cmd.exe /d /c breaks large-payload draining; the missing-script stdin drain stays on the encoded-PowerShell fallback (used for spaced/non-ASCII paths). Upgrades self-heal on first launch: startup install() unconditionally rewrites the command and Codex trust entry, sweeping the old compound form. Add a platform-independent regression guard (launcher must resolve to a real file, never a cmd-builtin fragment), update the lifecycle test + docs, and fix a stale Devin comment. |
||
|
|
36cd8a3347 |
fix(terminal): retire sessions when tabs close (#8628)
* fix(terminal): retire sessions when tabs close * fix(terminal): close remaining session lifecycle gaps * fix(terminal): close review-discovered lifecycle gaps * fix(terminal): revalidate bulk session retirement * fix(terminal): harden retirement review edges * fix(agent): reverify restored pane authority * test: make terminal retirement gate portable on Windows * test: use POSIX join in Linux PATH assertion * test: keep simulated Linux PATH host-consistent |
||
|
|
c3ab805d12 |
fix(agent-hooks): drain stdin before hook script early exits so agents never hit EPIPE (#8430)
* Fix hook scripts to drain stdin before any early-exit path Generated agent hook scripts and missing-script launchers could exit successfully before consuming the payload written to their stdin, leaving the writer with a broken pipe (EPIPE/ERROR_BROKEN_PIPE) once the reader closed early. Capture stdin (or drain it via a shared epilogue/fast-path guard) before any whole-script success exit across all POSIX, batch, PowerShell, and Git Bash launcher variants, and add a cross-agent lifecycle test suite plus a live Electron verification script to guard the contract going forward. * Harden hook scripts against unreadable managed scripts and add a Claude/ - Extend the POSIX launcher guard to also require `[ -r ]`, not just `-f`/`-x`, so an executable-but-unreadable managed script still drains stdin instead of erroring or silently misbehaving. - Add a verifier case (`verifyClaudeDevinSkip`) that spins up a local HTTP server and confirms the Claude hook never forwards a request that Devin already imported, catching accidental double-forwarding. - Update installer-utils tests and stdin-lifecycle docs to match the new readable-file guard and the added verification case. * Fix hook-launcher verification to derive script paths from the installed Extract the quoted path from the launcher's `if [ -f '...'` clause instead of reconstructing it via join(home, ...), so missing/failing-script test cases can't silently fall through to the real script if the install layout changes. --------- Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> |
||
|
|
7b12d38b17 | perf(terminal): tune cold-park keep-warm so common rotation never remounts (#8262) | ||
|
|
25ecf2eea2 |
fix: reconcile SSH repo rows after host re-add (#8201)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
fd6805a299 |
Fix mobile terminal query reply authority (#8227)
* Fix mobile terminal query reply authority * fix(terminal): harden mobile query reply handoffs * fix(terminal): exclude passive mobile query responders * fix(terminal): gate mobile query replies on host capability Older hosts strip terminal.send's inputKind (zod drops unknown keys), so a forwarded xterm reply would land as ordinary floor-taking shell input. Hosts now advertise terminal.query-reply-input.v1 via status.get and mobile drops replies unless the host advertises it (pre-fix behavior). Also documents the bounded desktop-to-mobile handoff double-reply residual. Co-authored-by: Orca <help@stably.ai> * fix(terminal): advance snapshot seq across recovery snapshots The pending-overflow recovery loop trims buffered output against recovery.seq while query replay and boundary strips kept using the initial snapshot seq. Unreachable under today's control flow (no await separates the initial-overflow consume from the loop), but the stale seq would silently drop covered query replies if that ordering ever changes. Track the seq that actually covered the buffered chunks. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
533992bdda |
fix(git): cache unsupported capabilities per host (#8109)
* fix(git): cache unsupported capabilities per host Old Git worktree, ref-search, and merge-tree fallbacks retried unsupported flags on recurring operations, flooding subprocess traces. Centralize capability probing per native, WSL, and SSH execution host, coalesce concurrent probes, and retry periodically for in-place Git upgrades. * fix(git): recognize real old-Git merge-tree rejection * test(git): enforce real binary compatibility matrix * fix(ci): preserve Git compatibility test ownership * fix(git): retain supported capability state |