mirror of
https://github.com/stablyai/orca.git
synced 2026-09-23 16:02:24 +00:00
c19f1b386b49837d7fdc8d3c8a92a47eeaad55dd
473
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
17cfc968cf |
Revert the terminal IME composition-ownership change (#13282)
* Revert "test(ime): restore coverage the composition-ownership change removed (#13168)" This reverts commit |
||
|
|
25a8c517e1 |
test(ime): restore coverage the composition-ownership change removed (#13168)
* test(terminal): pin the recorded Korean commit-before-newline order (STA-3132) Recorded first-party on Windows 11 + Microsoft Korean (HKL 0412) against the defect-era v1.4.164 build, with bytes read on the far side of the PTY: the terminal received ea b0 80 0d, the syllable strictly before the CR. The capture did not reproduce the suspected deferred-newline inversion. That route needed a session end carrying dataPendingReconciliation, which plain compose-then-Enter cannot produce because the IME finalizes first and the newline is never held; back-to-back arms at 25/60/120 ms did not reach it either. The test therefore pins the ordering rather than discriminating a fix. Co-authored-by: Orca <help@stably.ai> * test(terminal): restore Hangul back-to-back flush coverage deleted with the composition layer #12278 fixed a Hangul syllable that was not flushed before the next composition began — the force-end path, and the one that leaves stale glyphs behind. Returning composition ownership to xterm deleted both that patch and its test, so nothing guarded the behavior any more. Replays the recorded back-to-back arms (25/60/120 ms, read as 가\r나 at the PTY) against a real xterm Terminal. It passes on main: stock xterm flushes the committed syllable natively, so the removal was safe rather than a silent regression. Co-authored-by: Orca <help@stably.ai> * test(mobile): pin accessory-byte ordering behind a Hangul commit Returning composition ownership to xterm deleted the accessory-input commit tests along with the hook they targeted, but the guarantee they protected is user-visible and still applies: an accessory-bar keystroke must not overtake the syllable being committed, and must be suppressed when that commit fails. Drives the current hook with an Android composing-region trace rather than reconstructing the deleted coordinator. Co-authored-by: Orca <help@stably.ai> * test(terminal): replay recorded IBus and fcitx5 Hangul traces offline Commits interleaved with ASCII (한abc글) are the Linux IME gesture users report on, and its failure modes are a lost syllable and a doubled one. That gesture was only covered by tests/e2e/terminal-linux-ime-native.spec.ts, which needs a Linux host running a real input framework. Fixtures are the recorded captures from the sealed linux-final evidence run, replayed against a real xterm Terminal: exact onData, exactly-once counts across five repetitions, and the PTY bytes the recorded run actually received. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
17b3dff3c4 |
refactor(terminal): return IME composition ownership to xterm (#13128)
* fix(terminal): return IME composition ownership to xterm * fix(mobile): derive terminal input from native replacement ranges * test(mobile): record iOS Japanese IME traces * fix(mobile): preserve native IME replacement ranges * fix(xterm): flush queued application input after IME commit * test(terminal): pin Korean intermediate commit * test: pin Windows IME shortcut ownership * test: replay IBus number candidate commit * fix: preserve native macOS input-method punctuation * refactor(terminal): remove stale mac focus override * fix(mobile): preserve soft keyboard deletion ranges * fix: keep IME-owned palette chords in renderer * fix: stop carried IME shortcuts at renderer owner * fix: preserve carried IME shortcut dispatch * fix: narrow main-owned shortcut actions * test(mobile): pin Japanese IME replacement traces * test(terminal): retain paired native IME trace * fix(chat): preserve browser IME composition ownership * fix(chat): retain macOS IME confirm gesture * fix(chat): expire unmatched IME confirm carry * fix(chat): isolate IME confirmation expiry * fix(chat): retain active IME confirmation * refactor(terminal): remove dead composition handler * feat(ime): add shared Enter-ownership seams for CJK composition The confirming Enter of a CJK composition arrives as two keydowns and the orderings differ by platform: Windows/Linux redispatch the unmarked Enter/13 before keyup, macOS delivers keyup first. A guard reading only isComposing or keyCode 229 misses the redispatch, so surfaces submitted on a confirm. Adds useImeEnterGestureOwnership (carry token, next-frame expiry), a shared ImeEnterGuardedForm for native implicit submission, and the cmdk seam covering 18 CommandInput surfaces at one site. A chorded Enter arms the carry but is never swallowed — the reverse would eat a user's deliberate Cmd/Ctrl+Enter. Both failure modes are pinned by ime-enter-gesture-ownership-contract.test.ts. Co-authored-by: Orca <help@stably.ai> * refactor(terminal): consolidate native input listeners and parked-screen owner Extracts the shared native-input listener installer and renames the parked-screen detector for what it actually does, replacing per-call-site duplication. The listener installer keeps a forgetOptionKeyLocationOnBlur flag so per-window semantics are preserved rather than flattened. Net deletion; no behaviour change intended. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin recorded IME shapes as regression tests Nine regression tests built from hashed affected-platform captures, each with a paired ordinary negative and a discriminating mutation verified to take the file from all-passing to exactly one failure. Covers the Windows MS-Korean Shift family (#12179, #11878, #12151, #11946, #12152) and the Korean TUI line-break rows (STA-3237, STA-3222, STA-3129). STA-3237 pins the empirical 3-Shift / 2-active-composition / 2-newline ratio the device run established — the third Shift produces nothing because Space has already committed. That ratio is not derivable from a static capture. Co-authored-by: Orca <help@stably.ai> * fix(ime): guard Enter-commit surfaces against CJK confirm Applies the Enter-ownership guards across the surfaces whose Enter commits something: publishes, clones, pairs, installs, posts, or persists. Tiered deliberately rather than uniformly. Irreversible and remote-effect sites take the carry token, which also blocks the unmarked redispatch. Locally reversible sites take the oracle check with a one-line comment naming the residual, because a spurious commit there costs one undo. Three numeric fields are left unguarded with the reason in-code: Chromium blanks number inputs at compositionstart, so a confirm-Enter only ever reaches an empty-draft reset. Measured with a CDP probe rather than assumed — a guard that cannot fire is noise. Co-authored-by: Orca <help@stably.ai> * test(ime): teeth-check the Enter guards on every guarded surface One suite per guarded surface, each verified by deleting the guard and confirming the test fails. A green guard test without that check is unverified, not verified. Two shapes pass vacuously in happy-dom and are avoided here: native implicit form submission never fires, and blur() is inert on an unfocused element. Both made "the commit did not happen" assertions pass with the guard removed, so the suites assert the guard's contract directly instead. Co-authored-by: Orca <help@stably.ai> * fix(mobile): keep iOS Korean commits whole through the live-input path iOS Korean reports isComposing: false on every event, so it bypasses the composition guard entirely. The strict owner rejected UIKit's transformed post-change field and sent only the leading jamo — the reported symptom. Prefers the authoritative same-event field text over the predicted text when the supplied operation cannot produce it. Generic: no Korean special-case, no locale classifier, no normalization. Adds the RN-target-keyed submit carry alongside it. Co-authored-by: Orca <help@stably.ai> * test(e2e): make IME capture harnesses fail loudly instead of silently Four instruments recorded silence as success, so a void run scored as a clean one: - readTerminalImeBoundaryTrace returned an empty trace when the probe never installed, making every "nothing leaked" negative pass vacuously - summarizeLatencies([]) returned a perfect zero distribution that passed all three latency thresholds - the macOS Vietnamese spec pinned an input-source ID that does not exist, and failed as though the operator had chosen the wrong source - the expectedLineCount=1 prefix property was undocumented and one edit from silently downgrading a PTY assertion Input sources now resolve by enumeration and name the near-matches on failure. Co-authored-by: Orca <help@stably.ai> * test(terminal): cover Cangjie cancellation and fix a cross-namespace assertion Adds #11951's recorded Cangjie cancel shape to the existing cancellation suite, which covered Pinyin and Sogou but not Cangjie. One keystroke then Backspace arriving as deleteContentBackward with data: null, so the stale preedit is the only thing a fallback could replay. Verified against the historical pre-6cd944c62b3 bundle: the positive fails with ['尸'] where [] is expected, while the ordinary negative stays green. Also fixes the Vietnamese spec, which asserted a TIS-space input-source ID against getKeyboardInputSourceId(). Those two Orca APIs report the same source in different namespaces — TIS nests it under VietnameseIM, the app API does not. The resolver stays as an installation precondition; the assertion matches the leaf. Co-authored-by: Orca <help@stably.ai> * test(e2e): add a real-IME macOS arm for the Korean chord commit The existing korean-ime-terminal-shift-enter-commit spec synthesizes composition over CDP: Input.imeSetComposition sets the preedit directly and Input.insertText performs the commit. Asserting the IME produced events you injected yourself is circular, so that spec cannot certify real-IME behaviour. This arm selects 2-Set Korean via TIS, reads it back live, and injects through System Events key codes, so the OS owns the preedit, the commit instant, and isComposing. PTY byte expectations are preserved verbatim. Covers 2 of the original 4 cases by design. The other two are the Windows/Linux redispatch-before-keyup ordering, which macOS cannot produce and which cannot be selected -- the OS decides it. Reintroducing synthesis to "restore coverage" would reintroduce the circularity. Co-authored-by: Orca <help@stably.ai> * test(e2e): assert the macOS chord arm at the PTY boundary, not the renderer The byte expectations were transcribed from korean-ime-terminal-shift-enter-commit :364/:383, which assert against onData -- a renderer boundary where the terminator is CR. This spec reads the PTY child, where the tty has already converted CR to LF. Names both forms per row rather than swapping the constant, so the conversion reads as evidence that the capture reached past the renderer, as #11936 and #11951 record. Ctrl+Enter's CSI-u sequence is unaffected and is identical at both boundaries. Co-authored-by: Orca <help@stably.ai> * test(e2e): measure composer-to-onData latency and stop dropping IME keystrokes Two defects in the echo latency probe. It hooked onWriteParsed and onRender but never onData, so it measured key->parse->render echo rather than the composer-vs-onData delta the latency rows need. Adds a third hook feeding its own sample set. And `event.key.length !== 1` silently dropped IME keystrokes: Pinyin and Cangjie keydowns arrive as key:'Process' (length 7). Replayed over the captured corpus, the old filter accepted 580 of 4137 Chinese IME keydowns -- it was discarding 80% of them. The new filter matches the shape the owner itself branches on. Attribution charges each onData to the latest keydown rather than a FIFO head, because composing jamo emit no onData at all and a queue would credit a whole composition to its first keystroke. The consumer now asserts sample count before any percentile, so a zero-sample run cannot render as a flawless distribution. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin the WSL shifted-jamo newline shape for #11919 In Korean 2-set, Shift types ordinary letters -- the double consonants and the compound vowels. Each such keystroke reaches Chromium as key='Process', keyCode=229, shiftKey=true. The v1.4.163 classifier matched exactly that pattern with no code guard, so it called those keystrokes Enter, rewrote them to a synthetic Shift+Enter, and injected a newline into the middle of the word -- with no Enter key pressed. That is why the reporters said "no modifier key pressed": they had not chorded Shift+Enter, but they had pressed Shift, to type the double consonant. Asserts the row's own recorded capture: 40 immediate keydowns, exactly 3 of them Shift-carrying inside a single syllable, and an onData stream with one newline per Enter press and none mid-word. Two ordinary negatives keep it from being a blanket mute -- the same session's non-IME keydowns still reach shortcut policy, and an ordinary Shift+Enter still resolves through the real policy. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin the composition commit lag that made Korean type one behind macOS Korean 2-Set commits syllable N only when the first jamo of N+1 arrives, so compositionend and compositionstart land in the same task. A composition-start handler cancelled the pending finalizer that was the only path to triggerDataEvent and ended the session without emitting bytes, so every committed syllable reached onData exactly one syllable late and the backlog cleared only at a Space or Enter. Types continuously with no Enter and no Space -- either would flush the backlog and hide it -- and samples onData at every syllable boundary. Paired with a length-matched ASCII arm that stays green throughout, so the positive is a fact about composition rather than about timing in general. Bisected to a single call site across five builds: pristine, 1.4.155 and 1.4.162 pass, 1.4.163 fails, removing the one call repairs it, restoring it fails identically. That window is exactly the reporter's "started immediately after updating". Co-authored-by: Orca <help@stably.ai> * test(mobile): cover the send-queue abort that silently drops queued keystrokes One failed send in use-terminal-live-input-commit aborts every keystroke queued behind it, with the error swallowed by .catch(() => false). The existing test resolves(true) on every send, so the failure branch was uncovered. Four arms: the abort itself, an ordinary negative on the healthy path, a throwing sender, and a liveness control proving the queue recovers once the chain settles. Deleting the abort takes 4 passed to 3 failed, with the ordinary negative correctly surviving. Scope is stated in the docblock: this is a transport send-queue abort, reachable only via a real disconnect or RPC error. REQUEST_TIMEOUT_MS is 30s, so latency alone cannot reach the branch — consistent with #7094's symptom class, not proven to be its cause. * test(terminal): pin that daemon snapshot/restore cannot disturb a composition Two independent reporters attributed broken Korean composition to the always-on PTY daemon repainting terminal state over the preedit. The attribution is wrong on ancestry — the daemon shipped three months before the version both call good — but the boundary was never actually tested. Runs the real applyMainBufferSnapshot choreography against a live composition, including the full 2J/3J/H wipe plus the resize and alt-screen branches. textarea.value, selectionStart/End, compositionView.textContent and .active all survive byte-identical, and interleaving a restore between every jamo of 문제 still commits 문제 at onData. Also pins that the uncommitted preedit is absent from the captured snapshot: it lives in the textarea, never the buffer, so a restore has nothing stale to echo back. Injecting one textarea.value = '' into the restore fails exactly the three restore-boundary tests. * test(terminal): pin that Cmd tears down a composition where Ctrl and Shift do not xterm's composition keydown exempts only keyCode 16/17/18 (Shift/Ctrl/Alt) plus 20/229. macOS Meta — 91/93/224 — is absent, so a Cmd press mid-composition takes _finalizeComposition(false): the overlay goes dark and never recovers, because compositionstart is not re-fired. The user composes the rest of the word blind. Linux and Windows users press Ctrl and are exempt. xterm already has a Meta-aware modifier predicate in wasModifierKeyOnlyEvent, so this is an internal inconsistency rather than a deliberate choice. Owns no reported row and is version-neutral: 5/5 on both 1.4.162 and 1.4.163. The branch is unexercised in all 328 recorded traces, so this is a hazard pin, not a regression guard. Only the teardown is asserted; the likely duplicated commit needs a compositionend the IME kept alive across the Cmd, which no capture contains. Deleting the exemption fails exactly the three paired negatives; adding Meta to it fails exactly the two Cmd arms. * test(native-chat): characterize preedit loss when a question card replaces the composer An AskUserQuestion card fully replaces the composer by design, but the in-flight composition goes with it: the composer unmounts before compositionend reaches it, so the preedit is never committed to the draft. The committed text survives only because the draft is cached and restored via defaultValue. Node identity changes, value 'abc' is preserved, the 가 is gone. Drives the real NativeChatView -> SessionGate -> InteractiveCard -> questionActive swap -> Composer -> ComposerField, flipped by writing the same store field an AskUserQuestion hook event writes. Flipping questionActive to false fails exactly this test and nothing else across 639 native-chat tests, so the path was entirely unguarded. CHARACTERIZATION TEST: it asserts the loss. Fixing the defect — committing the preedit before the swap, or keeping the composer mounted — will make this file fail. Update the expectations to the new contract rather than working around them. Owns no reported row. #12118/STA-3219 flicker is keyed to token counters, which provably do not remount, and a question card arrives once per question. * test(terminal): pin the duplicated commit when Meta interrupts a composition _finalizeComposition(false) sends textarea.value.substring(start, end) but cannot clear the IME-owned textarea, so a later compositionend re-sends the same range. Meta reaches that path because CompositionHelper exempts only Shift/Ctrl/Alt; xterm's own wasModifierKeyOnlyEvent covers Meta four ways, so the omission is an internal inconsistency rather than a choice. Companion to the modifier-exemption guard, which deliberately pins only the overlay teardown. This pins the data consequence. HAZARD PIN: owns no reported row. The trigger is unverified on hardware — no capture in the corpus contains a Meta-during-composition gesture, and whether macOS keeps the composition alive across it is unmeasured. The duplication follows from the code given that sequence; whether users reach the sequence is the open half. An earlier premise that Space (keyCode 32) reaches this path was refuted by a corpus scan: 0 of 731 evidence files carry a keyCode-32 Space while composing, against 171 at 229, and 229 returns early. * test(terminal): characterize the syllable lost when the textarea blurs mid-composition CoreBrowserTerminal._handleTextAreaBlur clears the helper textarea unconditionally — "Text can safely be removed on blur" — while CompositionHelper._finalizeComposition reads the committed text back out of that same value from a deferred timeout. By the time it runs the value is empty, the substring is '', and triggerDataEvent never sees the syllable. xterm checks composition state in _syncTextArea and omits the same check here. Six cases. Blurring mid-composition loses the syllable in every ordering, including compositionend-before-blur, which is Chromium's real order — so it is not an ordering artifact. A bare textarea.blur() with no Orca code loses it too, which places the owner upstream: Orca's unguarded release on outside pointerdown is one trigger, not the cause. Committing 한 then blurring mid-가 yields ['한'] where ['한','가'] is correct: one syllable gone, surrounding text intact. Teeth checked by inverting — adding an Orca-side composition guard flips exactly the three cases that route through the release path and leaves the bare-blur and no-blur cases green, which is the scope split: a fix in regular-terminal-focus-ownership alone would not close this. HAZARD PIN, but unlike the others this one has a real production injector — clicking outside the terminal mid-composition. Owns no reported row. The shape matches #9738's report; the injector does not, and a shape match with a mismatched injector is not an owner. * test(terminal): say which arm the STA-3237 fixture came from The recorded keydowns are wave 4's A-shift-unmarked-only — the arm that emits no PTY bytes. Nothing in the file said so, so two readers concluded the row's events fail the owner's predicate and that STA-3237 and STA-3222 were different defects. They share an owner; the arm that fires is Process/229+Shift, absent from this bubble-phase trace because the owner claims it in the capture phase. Also corrects "code-blind": the v1.4.163 policy emits \x1b\r only for a shift-only key:'Enter', and a jamo keydown reaches that branch solely via the isTerminalImeProcessEnter rewrite. The mock is deliberately wider so the ownership guard stays under test if that rewrite moves. Comments only — no assertion, fixture value, or mock behaviour changed. * test(e2e): track the input-source selector the macOS specs shell out to Five tracked macOS IME specs ran `swift .tmp/select-input-source.swift`, a file that is gitignored and existed only on one machine. Anyone else checking out the repo — or the same machine after .tmp is cleaned — could not run them, and they are the capture drivers for the macOS rows that are blocked waiting for exactly those runs. Moves it to tests/e2e/ beside its callers. The chord spec now resolves it from __dirname rather than reaching two levels up into .tmp. * test(terminal): pin the CJK repaint decision against the reporter's own output #12164 comment 1 and #5921 report agent output with double-width glyphs rendering duplicated character-by-character while ASCII in the same line stays clean. No IME, no composition, no keystroke — the user never types the CJK. Segmenting all three verbatim samples into maximal same-risk-class runs gives 33 runs and zero violations of "this run is corrupted iff the production detector flags it": 17 wide runs all corrupted, 16 narrow runs all byte-identical. The paired negative is co-located in the same line rather than in a separate run — the reporter supplied it without knowing. Doubling is asserted as present, not uniform: 자바스크립트 and 시스템 each leave a jamo undoubled, which is a repaint-region boundary artifact rather than a per-character transform. The discriminating arm is in the test rather than a source mutation: |
||
|
|
a7e31e5e10 |
chore(mobile): apply the supply-chain release-age gate to mobile
mobile/ had no .npmrc, so unlike the root workspace it would resolve packages published moments ago. #13113 surfaced this concretely: it pulled nanoid 3.3.18 and postcss 8.5.26 at 0.4 and 1.7 days old, both newer than anything the root gate would have allowed. Copies the root's minimum-release-age=4320 (3 days). Deliberately not shamefully-hoist -- that one is Electron-specific and would change how mobile hoists. Re-resolves nanoid to 3.3.17 and postcss to 8.5.25 in the same commit because the gate is otherwise unusable: pnpm install fails with ERR_PNPM_NO_MATCHING_VERSION on the locked nanoid 3.3.18. Both picks stay above their advisory floors (CVE-2026-67213 needs >=3.3.17, CVE-2026-69153 needs >=8.5.23), so this is not a security regression. Co-authored-by: Orca <help@stably.ai> |
||
|
|
757b785e43 |
fix(deps): resolve Dependabot security alerts across root and mobile (#13113)
Clears 47 of 49 open Dependabot alerts across the root and mobile lockfiles. The 2 remaining (image-size) have no patched upstream release. Direct bumps: pdfjs-dist 5.7.284 -> 6.2.108 (CVE-2026-16633), mermaid 11.16.0 -> 11.16.1 (root + mobile), dompurify 3.4.12 -> 3.4.13. In-range re-resolves: brace-expansion, fast-uri, hono, ip-address, js-yaml 4.3.1/3.15.1, nanoid, postcss, tar, undici 6.28.0/7.29.0. Drops the @modelcontextprotocol/sdk>@hono/node-server override by bumping shadcn's transitive SDK to 1.30.0, which widens its range to ^1.19.9 || ^2.0.5 so @hono/node-server resolves to a patched 2.1.0 on its own. The other two overrides must stay: monaco-editor hard-pins dompurify 3.2.7 and xcode wants uuid ^7.0.3, both vulnerable. pdf.js 6 removed PDFDocumentProxy.destroy(); PdfViewer now tears the document down via the loading task it was already destroying. Supersedes #13074, #13090, #12960, #12952. Co-authored-by: mondaychen <monday.chen@gmail.com> Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
094d6821ef |
feat(native-chat): add model and effort pickers for grok (#12780)
* feat(native-chat): add model and effort pickers for grok Grok had no session-option catalog, so the native chat composer showed no pills and every launch ran the CLI's own defaults with no way to change them. Adds a `GROK_SESSION_OPTION_CATALOG` (model via `-m`/`/model`, reasoning effort via `--reasoning-effort`/`/effort`) and the discovery plumbing behind it. Grok's selectable ids depend on the signed-in account and on `[model.*]` config, so the seed carries only `grok-4.5` and a runtime `grok models` probe supplies the rest as authoritative — a retired id must be droppable, since launching one is a fatal exit rather than a warning. Because `grok models` publishes `Default model:` and marks the row `(default)`, the picker can name the model a fresh session is actually running: `defaultModelIsCliDefault` plus an untracked record means no `-m` was ever emitted, so the CLI is on its own default. That default scopes the effort row but is never written to persisted settings — that field is what authorizes `-m` on every later launch, and adopting a model the user never picked would pin today's default forever, fatally so on an account without it. `grok --help` publishes no default for `--reasoning-effort`, so the effort value stays unnamed until something sets it. Known gap: that refusal to persist is also a limit. An option set while on the CLI default is dispatched and honored in-session, but reaches no later launch — it persists under the default's id with `model` left unset, and both `resolveNativeChatSessionOptionDefaults` and `resolveAgentSessionOptionLaunch` bail without that key. Picking a model explicitly persists normally. Closing this means teaching both to resolve options from the default model while still refusing to emit `-m`, which is the launch-args path and wants its own review. Known gap: the picker infers "no `-m` was emitted" from its own in-memory record, so a model reaching argv from outside it — the user's own `agentDefaultArgs`, or a renderer reload that drops the record while the flagged PTY lives on — leaves the pill claiming the CLI default while another model runs. No wrong model is persisted. Extracts `hasFlag` and `labelFromModelId`, and splits the model-probe spec out of the commit-message registry so discovery no longer implies an agent can write commit messages. Co-authored-by: Orca <help@stably.ai> * docs(native-chat): note the invariant keeping modelIsCliDefault agent-safe The flag is computed without checking the catalog, so it reads as unsafe for the four agents with no CLI default. It is safe only because `persist` bails unless `modelId` is truthy, which for those agents implies a tracked model. Widening that guard would silently change persistence for every agent. Co-authored-by: Orca <help@stably.ai> * fix: retire persisted models on mount and handle -- terminator - When a pane mounts after model discovery has already settled, it now checks the cache and retires persisted models that are no longer available. - CLI flag detection now respects the `--` option terminator, treating everything after it as positional arguments rather than flags. * Fix: persist grok session options under probe-confirmed defaults Options set under the CLI default were silently lost on restart. Distinguish seed guesses from probe-confirmed defaults by renaming `modelIsCliDefault` to `modelIsUnverifiedDefault`. Once confirmed, adopt the default as a persisted flag so options survive restarts. * fix(native-chat): close the retired-model fatal-launch paths from counsel review Counsel report C1/C2 (High), C3, P1, C4: - Untrack a session model an authoritative discovery dropped and gate every persist path, so option writes can never re-adopt a retired id (C1). - Resolve launch defaults through the enrichment cache: a persisted model missing from every settled probe no longer becomes a fatal `-m` (C2). - Serialize retirement and picks on one settings write queue that re-reads live state at apply time (C3). - Stabilize onSwitchToTerminal so the session-option surface is not rebuilt every TerminalPane render (P1), and cap the enrichment host map (C4). Co-authored-by: Orca <help@stably.ai> * Store agent in enrichment entry and extract token utilities Refactor enrichment to store the agent field directly instead of parsing it from a composite key, and extract CLI flag token filtering into a shared utility. Use a dedicated function for tracked model ID lookup. Improves code reuse and reduces parsing overhead. * Rename modelIsUnverifiedDefault to adoptModelAsLaunchDefault Move the model adoption gate into the core session-options module, where probe confirmation and discovered-model status are known. This ensures adoption decisions are gate-checked before persisting to avoid fatal launch flags, and simplifies the picker surface by moving the logic to where it belongs. * Keep model probe evidence by agent, not host Store probed model IDs in agent-keyed cache independent of host cache, so evidence persists across host eviction. Prevents retired models from being treated as valid when host cache entries are evicted. * Store agent in enrichment entries instead of separate proof-evidence map Model probe evidence is now tied to enrichment entries rather than maintained in a separate per-agent map, eliminating the need for eviction logic that could disconnect proof from entries. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
20aeb0cb99 |
Prepare mobile 0.0.42 and fix TestFlight CI hang (#12961)
* Bump mobile app.json to 0.0.42 * fix(mobile-ios): stop TestFlight CI from waiting on ASC processing 0.0.42 builds 1–2 uploaded successfully then hung for hours polling processing with no Ready build and no Apple email. Exit after upload and cap the job at 90m so the next cut does not repeat that hang. * fix(mobile-ios): fully skip Pilot wait (no changelog) Pilot only returns immediately after upload when changelog is nil; passing notes re-enters the ASC build-list poll. |
||
|
|
02a1251c2d |
fix(native-chat): classify diff lines whose content begins with -- or ++ (#12459)
* fix(native-chat): stop diff colouring from misreading -- / ++ content lines as file headers diffFromText skipped every line starting with --- / +++ as a file header, so a deleted SQL/Lua '-- comment' (git emits '---<content>') or an added '++flag' fell through to gray context with its marker still attached — and when it was the only change, the two-marker gate dropped the coloured diff entirely. Detect real headers structurally instead: an adjacent '--- <old>' / '+++ <new>' pair outside any hunk. A hunk header or 'diff --git' line now also proves the text is a diff, so a genuine single-line change renders while prose keeps the guard. Co-authored-by: Orca <help@stably.ai> * test(native-chat): adopt #12335 diff-collision vectors and add mobile parity Pulls in @YuriNachos's test vectors from #12335 (header-less --- deletion, an adjacent --x/++y content pair, mobile re-export parity) and adds the spaced -- / ++ pair inside a hunk, which the pair-only rule in that PR misreads. Co-authored-by: Orca <help@stably.ai> * fix(native-chat): keep bare --- / +++ rules out of the diff marker count Dropping the `---`/`+++` prefix exclusions made a bare `---` — a Markdown thematic break or YAML document separator — classify as a deletion. Tool results routinely carry those, so `---\na: 1\n---\nb: 2` went from correctly rejected to rendering as a red diff. A bare rule is never a file header (those need a path after the marker) and is only content inside a hunk, so treat it as meta when outside one. Fold the separate `isStructuredDiff` scan into the same pre-pass and skip non-marker lines early, so the added guard costs no extra traversal: 5.1 -> 4.3 us per 120-line prose result, diff path unchanged. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
9e4e6ddae5 |
feat(native-chat): render omp transcripts (#11523)
* feat(native-chat): render omp transcripts
omp already ships as a first-class launchable agent with session_id resume, but
its transcripts had no decoder, so native chat could not render it — the agent
runs and the conversation stays a raw terminal. This adds the decoder and wires
it through the same path Claude, Codex and Grok use.
omp writes one envelope per line, `{ type, id, parentId, timestamp, … }`, where
conversation turns are `type: 'message'` and the rest is session bookkeeping.
Reasoning arrives as a `thinking` content block inside the assistant turn, so
the mapping follows Claude rather than Codex: thinking becomes a text block on
an assistant message, where Codex and Grok emit a separate reasoning role only
because their transcripts carry dedicated reasoning records.
- toolCall -> tool-call, arguments passed through as the object omp writes
- toolResult -> tool role, isError preserved
- developer -> system, matching the Codex non-user/non-assistant fallback
- blob-handle images drop, as the Claude mapper drops an image record with
neither path nor url
- bookkeeping and unrecognized types skip rather than throw
Session files are `<ISO timestamp>_<session id>.jsonl` under a per-cwd directory,
so the resolver matches the id as a base-name suffix the way Codex rollout files
are matched, and honors OMP_CODING_AGENT_DIR through normalizeAgentSessionsDir
so it stays consistent with the AI Vault scanner.
omp records no interruption or abort event, so unlike Claude and Codex there is
no NATIVE_CHAT_INTERRUPTED_STATUS_TEXT path.
Verified against 94,603 lines of real omp transcripts across four sessions:
50,546 records decoded, zero malformed, zero thrown.
* fix(native-chat): complete omp record coverage and gate remote transcripts
Review fixes on the omp transcript decoder.
omp writes several record types with no `content` field, so they decoded
to zero blocks and disappeared from the chat view entirely:
- `bashExecution` / `pythonExecution`: TUI `!command` runs, now a tool turn
- `fileMention`: `@path` attachments, listed by path (never `files[].content`,
which is an auto-read dump)
- `custom_message` and legacy `custom` / `hookMessage` rows, gated on
`display` the way omp's own renderer gates them
Also:
- `stopReason: 'aborted'` turns now surface as the interrupted row, matching
the Claude and Codex decoders. An abort carrying partial content keeps it.
- A cancelled command cell now reads as errored. Every omp cancel path emits
`exitCode: undefined`, which JSON drops, so an `exitCode !== 0` check read a
cancelled run as a clean success.
- omp joins Grok in requiring a locally readable transcript. Its hook reports
no transcript path, so under Model-A SSH the chat view opened against a disk
this process cannot read and never loaded. Applies on mobile too, which
shares the same allowlist.
- The session-file walk prunes omp's per-session subagent artifact
directories, matching the AI Vault scanner. It was returning a subagent
transcript instead of the parent session, and cost a full recursive readdir
on every resolve.
* style(native-chat): apply oxfmt to the omp review fixes
Mobile CI gates `oxfmt --check`; the two root files were unformatted too,
just ungated there. Line wrapping only, no behavior change.
---------
Co-authored-by: plotarmordev <299844489+plotarmordev@users.noreply.github.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
|
||
|
|
e68831f32c |
fix(github-project): index fork upstream slugs for project row matching (#12822)
* fix(github-project): index fork upstream slugs for project row matching Project cards often reference the public upstream repo while the open clone's origin is a personal fork. Map the parent slug to the same Repo so selected-repo filters no longer hide every board row. Preserves origin-based getRepoSlug identity for non-project callers. Fixes #12647 * fix(github-project): match project rows against fork upstream slugs Resolve the referenced call to a nonexistent `resolveRepoUpstreamSlug` and match the persisted `repo.upstream` parent instead of issuing an extra `github.repoUpstream` RPC per repo on every index build — that lookup shells out to `gh repo view` for non-forks, so it would have gated the Projects tab on N network calls. `repo.upstream` is already resolved at repo-add time and backfilled at startup, so the fix costs no IPC. Origin matches take precedence over upstream ones so an open clone of the upstream repo itself is never made ambiguous by someone's fork of it. Also covers the two surfaces the origin-only match broke alongside the desktop table: mobile's project row matcher and the store-slice row-mutation routing. * fix(github-project): scope fork upstream matching by host and selection Round-1 review fixes on top of the upstream-slug index: - Apply origin-over-upstream precedence among *selected* repos instead of globally. An open-but-unselected clone of the upstream repo was shadowing the selected fork, so #12647 still reproduced for anyone holding both — and repo selection collapses to one repo per project key, which is exactly that case. - Scope a fork's upstream identity key to the fork's own origin host. Persistence strips upstream.host, so GHES forks never matched their own rows and a GHES fork's parent could bind a same-named github.com row. * fix(github-project): skip the fork alias when its own origin is unresolved Round-2 review fix. `githubHostFromIdentityKey` cannot tell "origin resolved to github.com" from "origin did not resolve" — both yield no host. A GHES fork whose slug resolution had failed (auth lapse, unreachable runtime) therefore landed in the github.com namespace, so an unrelated public Project row matched it and Start work opened the wrong clone on the wrong server. Require a resolved origin before indexing the upstream alias: it is the only host evidence there is, and a repo with an unresolved origin was already absent from the origin index, so nothing is lost that origin matching had. * fix(repos): persist the fork upstream host instead of dropping it `sanitizeRepoUpstream` kept only `{owner, repo}`, so a fork's parent lost the server it lives on every time the record round-tripped through disk. That forced the Project row matcher to re-infer the host from `origin`. The inference is right for an API-resolved fork parent — `getRepoUpstream` stamps `origin.host` there precisely because "a fork parent lives on the same server as the fork". It is wrong for the other branch: a local `upstream` remote carries its own host, so a github.com clone with a GHES `upstream` remote was indexed into the github.com namespace, where an unrelated same-owner/name public repo could claim it and Start work would open the wrong clone. Keeping the host removes the guess. Absent stays absent, so records written before this hydrate unchanged and the origin-derived fallback still covers them. Also fixes the avatar for rehydrated GHES forks, which resolved against github.com for the same reason. * docs(github-project): correct upstream host fallback comment Persistence now keeps non-empty upstream.host; originIdentityKey remains the host fallback for older records without one (CodeRabbit nit). * fix(github-project): own slug-index retry timer cleanup Move the failure-retry setTimeout into its own effect so cleanup always clears it. Scheduling from the async buildIndex then-handler failed the react-doctor effect-needs-cleanup gate in static analysis. * test(github-project): guard the slug-index retry timer, fix the mobile twin comment Two follow-ups on |
||
|
|
8ddf575fe6 |
Revert "Remove source control group order preference (#12785)" (#12955)
This reverts commit
|
||
|
|
f6d0bde6fb |
fix(mobile): choose host for new workspace (#11647)
* fix(mobile): choose host for new workspace * fix(mobile): close stale workspace host picker * fix(mobile): disambiguate workspace host choices * fix(mobile): keep host endpoint paths private * fix(mobile): redact invalid host endpoints * fix(mobile): handle opaque host endpoints * fix(mobile): announce host picker options * fix(mobile): harden workspace host picker * fix(mobile): preserve host through workspace creation |
||
|
|
6182ca4b04 |
fix(mobile): expose host card actions (#11648)
* fix(mobile): expose host card actions * fix(mobile): preserve host status accessibility * fix(mobile): make host path labels speakable * fix(mobile): preserve host card accessibility details * fix(mobile): remove redundant host card chevron |
||
|
|
89f00658d2 |
fix(mobile): bound home host auto-connect fanout (#11642)
* fix(mobile): bound home host auto-connect fanout * fix(mobile): clarify bounded host connection state * fix(mobile): close unused settings host clients * fix(mobile): close released host clients * fix(mobile): sync host release policy in effect * fix(mobile): preserve focused host client ownership * test(mobile): cover manual Home client release * Fix React Doctor array type check |
||
|
|
b521837481 |
fix(mobile): keep hosts visible when credentials are unavailable (#12775)
* fix(mobile): keep hosts visible when credentials are unavailable * fix(mobile): guard unavailable host recovery * fix(mobile): protect replacement credential writes * fix(mobile): retain superseded cleanup intents * fix(mobile): preserve credential cleanup authority * fix(mobile): make host cleanup crash-safe |
||
|
|
630f13bfe6 | refactor(mobile): separate native chat controller contract (#12826) | ||
|
|
79896cb9a6 |
fix(native chat): retain transcript while reconnecting (STA-3333) (#12495)
* fix(mobile): keep the cached transcript visible while reconnecting A manual retry closes the client and opens a fresh one, so the chat session hook saw a new client under an unchanged identity, dropped its settled read, and handed out an empty list — the transcript collapsed to a full-screen spinner until the swapped client's snapshot landed. Hold the last settled list per identity (captured post-commit) and keep rendering it while the re-read is in flight. `transcriptLoading` still gates consumers that decide from an empty transcript, so the launch-draft seed is unaffected. The held list is keyed by a new `sourceIdentity` (host/workspace) in addition to agent/session/transcript, so it can never serve another source's messages. Refs STA-3333. * test(mobile): assert the whole reconnect window, not just its first frame The re-subscribe lands a commit after the first render of the swap, so a regression that cleared the held list there left frame 0 green and still blanked the transcript. Verified: clearing the cache in the subscribe cleanup now fails this test, where before only the view-toggle test caught it. * fix(mobile): don't derive a tappable ask card from the held transcript The cache this PR adds keeps the previous list rendered while a swapped client re-reads. useMobileNativeChatPrompts was the one consumer reading `messages` without honouring `transcriptLoading`, so an ask answered on the terminal resurrected as a live, tappable card during that window. Gating on `transcriptLoading` is exactly base behaviour: `setRead` only ever stores 'ready'/'error', so status==='loading' implied an empty list before this PR. The live `askFromStatus` path is untouched. * chore: keep merge formatting scoped |
||
|
|
6df8997c3c |
fix(file-explorer): sort numbered file names naturally across every listing surface (#11576)
* fix(file-explorer): sort numbered file names naturally The File Explorer compared names with bare localeCompare, so numbered files listed 100, 200 before 99. Hoist the numeric collator Source Control file rows already use (#10850) into src/shared and apply it to the local and runtime directory listings, the name-filtered view, and Source Control directory nodes, which were inconsistent with the file rows one line below (#11426). * fix(file-explorer): natural sort on SSH funnels, relay, and pickers Adversarial-review round 1 rework: - Both readDir funnels short-circuited to the SSH filesystem provider before the patched sort, so SSH workspaces kept lexicographic order; re-sort locally after the provider returns (the remote relay may be an older build), and fix the relay's own comparator for relay-native consumers. - sortDirEntries (shared, unit-tested) owns the directories-first + natural-order listing contract used by every funnel. - compareFileNames breaks numeric-collation ties ('2' vs '02') by code units so sibling order stays total instead of readdir order, and pins the collator locale to 'en' so every host produces one order. - The SSH folder browser and runtime server dir picker now match the Explorer they browse into. - Ordering pinned by tests at the relay, source-control tree, and shared helper. * fix(mobile): natural sort in the mobile file explorer Mobile re-sorted host readDir results with bare localeCompare, undoing the host funnel's natural order (round-2 review). Reuse the shared comparator and pin the order in the mobile suite. * fix(file-explorer): natural sort at the renderer choke point and remaining ties Round-3 review: the remote-runtime RPC and paired-web routes return the host's order verbatim, so re-sort in readFileExplorerDirectory where every desktop route converges; pin the SSH funnel with a handler-level test; and route Source Control path compares through compareFileNames so numeric-collation ties share one total order with the Explorer. * docs(file-name-sort): state the real perf baseline in the hoist comment * refactor(source-control): drop the dead collator export; pin the test oracle locale * fix(file-listings): cover remaining natural-sort surfaces |
||
|
|
ae1ed5e886 |
Remove source control group order preference (#12785)
* Reorder source control to show staged changes first by default Stages are closest to the commit action and most relevant to the commit workflow. Merges untracked files into Changes visually while preserving their Git area. Removes the untracked-first preset and includes migration logic for existing user settings. * Drop source control group order user preference Remove the sourceControlGroupOrder setting and related UI, migrations, and persistence logic. The source control view now always displays sections in the order: staged changes, unstaged changes, untracked files. * Reorder source control to show changes before staged Aligns with the edit-stage-commit workflow by showing unstaged changes (active edits) before staged changes (queued for commit). |
||
|
|
2c2a3266a6 |
Change question card submit button label to 'Submit' (#12782)
- Replace 'Send answer' with 'Submit' for clarity and consistency - Update all locale translations (en, es, ja, ko, zh) - Remove fixed button width and add whitespace-nowrap for flexible sizing - Update component and test references |
||
|
|
de152503b2 | chore(mobile): prepare 0.0.41 releases (#12781) | ||
|
|
73cd4c3f46 | fix(mobile): bound live terminal input latency (#12763) | ||
|
|
be2f9eddd3 | fix(mobile): abandon pairing journals that can no longer reconcile (#12773) | ||
|
|
7f4570c9a6 |
fix(mobile): activate Source Control diff tabs on phones (#12770)
* fix(mobile): activate source-control diff tabs on phones * fix(mobile): reveal legacy source control file tabs |
||
|
|
885afb55a9 |
fix(mobile): open host editor from root navigation (#12766)
* fix(mobile): open host editor from root navigation * test(mobile): update task navigation router contract |
||
|
|
d4dfc35ac4 |
fix(mobile): preserve multi-image chat attachments (#12639)
* fix(mobile): preserve multi-image chat attachments * fix(mobile): use preferred array syntax * fix(mobile): harden multi-image attachment flow * fix(mobile): retain first-send image previews |
||
|
|
23238aee0b |
fix(mobile): stop native-chat send button flicker (#12764)
* fix(mobile): stop native-chat send button flicker * fix(mobile): keep composer lock rendering pure |
||
|
|
86b878cfd6 | fix(mobile): parse classified PR lookup outcomes (#12659) | ||
|
|
5d2ad3597a |
fix(native-chat): add direct Codex model selection (#12657)
* fix(native-chat): select Codex models directly * fix(native-chat): confirm agent exits before switching views * fix(runtime): handle unavailable foreground probes |
||
|
|
a766ee4bcd |
fix(runtime): refuse to silently wake a deliberately slept pane (STA-3465) (#12672)
`activateMobileSessionTab` gated only on `publicTab.status !== 'ready'`. A deliberately slept pane publishes as `pending-handle` indefinitely — indistinguishable at that call site from a pane awaiting reconnect — so the reconnect probe added by #11542 respawned it with a re-resolved agent launch, waking something the user had deliberately put to sleep. The first attempt refused activation for any pane with a `worktree-sleep` record, applied to every path. Independent review found that broke the documented wake gesture: opening the tab IS how those panes are meant to cold-restore (`wake-sleeping-agents-in-background.ts`: "Those panes cold-restore --resume when their own tab is opened"). A mobile tap sends the byte-identical call the reproduction test used, and in three of four topologies no wake clears the record first — so the tap became a permanent no-op with no feedback. This carries intent explicitly instead of inferring it. A new shared `TabActivationIntent` ('user' | 'automatic') rides the existing ActivateTab schema as an optional additive field; `isAutomaticTabActivation` returns true only for an explicit 'automatic', so an absent value is permissive BY CONSTRUCTION in one place — an older client that does not send it keeps today's behavior rather than silently losing its wake gesture. The field is required on the mobile helper's params, so no call site can be added without declaring who asked. Every user path (mobile tab switches, paired tab clicks, shortcuts, palette, the pane's own open) is labelled 'user'. The only automatic sender in the codebase is `waitForResubscribeHostSessionHandle`, the #11542 reconnect probe. Verified per topology: user activation materializes a parked pane under headless serve, a paired runtime client, a completed agent with restoreOnTabOpenOnly, and a running agent whose wake cleared the record. The automatic probe is refused without retiring the surface, and #11542's reconnect tests stay green. Also fixes a test fixture that made a real bug untestable: the store stub ignored the host id, so mutating the partition lookup to 'local' left the suite green. Correcting it exposed three existing SSH reattach tests that had been relying on that looseness — their workspace session sat in the local partition while their repo was SSH-hosted, a store production would never read. Production was always right; the tests described an impossible world. Fixes STA-3465. |
||
|
|
738f640428 |
fix(mobile): relay UX overhaul — steady status colors, visible relay dials, coordinated deep links (F1-F10)
* fix(mobile): keep healthy relays green through focus and network nudges (F1+F2) Focus/app-resume nudges probe the active relay instead of suspending it; network-change nudges replace it make-before-break, suspending only after a failed dial. Mount, Retry, and host-swap windows read 'connecting' instead of 'disconnected'; the host list keeps last-known worktrees for every not-connected state and spins instead of rendering nothing. * feat(mobile): surface the pairing relay path in the pairing log (F3) The relay candidate was silent during pairing: dialing, E2EE handshake, director recovery, and the winning path now emit redacted phase lines through the same connectOptions.onLog the direct path already used. * docs(mobile): relay UX investigation findings and F0-F10 fix plan * feat(mobile): name and narrate relay dials while they happen (F5) migrateTo forwards the dialing session's connecting/handshaking/reconnecting phases whenever the client is suspended or disconnected — never downgrading a live session — and exposes getPendingPath so the host card can say '· Orca Relay' during the dial instead of only after it. * feat(mobile): race a relay dial when the direct dial stalls (F6) A 2.5s grace timer starts relay recovery while an unauthenticated direct dial is still inside its 12s connect window; the race gets one attempt through the existing mutex/cooldown machinery, cancels when direct authenticates, and never arms for hosts without a relay endpoint. * fix(mobile): overlay the protocol gate instead of unmounting the host stack (F9) A pending status.get used to swap the mounted HostStack for a spinner at the moment the socket connected, destroying in-flight nested navigation. Once children have rendered for a host they stay mounted under an opaque touch-blocking overlay; first visits and blocked verdicts keep the old behavior. * fix(mobile): keep loaded data through transient connection blips (F10) Git history no longer blanks on reconnect (and commit files refetch instead of caching an offline empty answer), the repo picker keeps its last-good list when an in-flight repo.list rejects, the diff review's ready-state preservation actually runs, and proven host capabilities survive a drop flagged unverified instead of being wiped. * feat(mobile): coordinate every home deep push and bounce dead resume targets (F4+F7+F8) Notification taps, the Accounts card, and host-edit now use the shared mount-then-replace transition (with a focused-route walker so root-layout scope works); the Resume card renders from the snapshot in a disabled state so its late arrival can't shift Tasks under the thumb; resume targets are validated against proven catalog data, and a session route whose worktree the host proves missing bounces to the host index with a notice banner instead of stranding on a dead screen. * test(mobile): cover the resume-target and notice policies (F7) Key notice dismissal by code so closing one banner cannot swallow a later, different one, and move the visibility rule into host-route-notice.ts where it is testable without a screen. Adds the missing units for F7's decision points: isResumeTargetConfirmedMissing (unproven catalog is silence, synthetic routes exempt), the validating last-visited reader, and the notice visibility rule. * fix(mobile): review-pass hardening for the gate overlay and diff preservation Adversarial review findings: the reader's hunk position now survives a connection blip (reset only on item change), the covered stack is hidden from TalkBack while the gate overlay is up, and the overlay's hit-test comment is scoped honestly to in-tree views (native-Modal drawers present above it — follow-up). * fix(mobile): CI + CodeRabbit review fixes for #12609 Move the findings doc under docs/ (root directory guard), drop two unused eslint-disable directives, and address review findings: an unproven snapshot seed can no longer downgrade a proven worktree catalog; a locally-aborted relay dial skips the director fallback; post-migration bookkeeping failures log instead of masquerading as dial failures (which could suspend the healthy session); the auth wait arms its timeout before subscribing; forwarded dial phases stop at close(); the legacy selector_not_found fallback requires runtime_error; the diff-loading effect depends on the fields it reads; and host-edit auto-cancellation is now pinned by a test. * fix(mobile): second review round — queued replacements, race fence, confirmed bounces A network-change replacement now survives the recovery mutex and cooldowns as a queued intent instead of being dropped or suspending a healthy session — only a failed dial or a dead probe tears one down. The happy-eyeballs migration withdraws when direct authenticated during the relay dial (first-authenticated-wins). A worktree bounce requires two consecutive host-proven misses, since a transient desktop repo-scan rejection answers selector_not_found for a live worktree. Background network flaps no longer wake a billed relay splice, the lifecycle foreground flag stays in sync, a screen unmount cancels only its own pending host-stack transition, and diff review keeps the loaded review when its reconnect refresh rejects. Extracted mobile-endpoint-nudge-router.ts and the establisher's dialEligible pass, and split the supervisor nudge tests, to stay under max-lines. * fix(mobile): satisfy the React Doctor changed-code gate Render-phase ref writes move into effects: the protocol gate's resolved/mounted latches now record committed outcomes only (a discarded children render can no longer count as mounted), and the bounce hook syncs its callback ref in an effect. Array<T> annotations become T[] in the extracted modules. * fix(mobile): keep the loaded diff when the reconnect refetch rejects (F10) The diff-loading hook's catch was the one path still erasing a ready diff — the same keepLoadedDiff guard its disconnect and loading branches already use, now pinned by a reject-after-ready test. * fix(mobile): process foreground revival nudges --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
c511e51442 |
fix(mobile): label native-chat tool rows with a clean, expandable input summary (STA-3333) (#12498)
* fix(mobile): label tool rows with a clean summary, expand full input (STA-3333)
Mobile tool rows showed the raw input JSON (`{"file_path":…}`) as the row
label, and the expanded detail just repeated that same truncated string.
- `describeToolInput` labels a row with the target file path, else the
primary argument (command/cmd/query/pattern/url/description), else the
bounded JSON preview.
- Codex delivers tool arguments as a JSON string; normalize those into the
object shape the helpers already understand, so labels, file links,
run summaries and the expanded detail all work for Codex calls too.
- The expanded detail now renders the fully formatted input, capped at
MAX_TOOL_RESULT_CHARS like desktop's tool detail (and like the result
body), and a structured input makes the row expandable.
* fix(mobile): name search rows by their term and keep the filename in path labels (STA-3333)
Review follow-ups to the tool-row summary, all in the shared helper:
- A Grep/Glob row labelled itself with the directory it scanned and dropped
the pattern entirely, because `toolFilePath` treats `path` as a file target.
That path is a scan root, so it also rendered a tap-to-open link that asked
the app to open a folder. `toolFilePath` now ignores the generic `path` key
for search-shaped input, which lets the pattern win the label and drops the
bogus link; an explicit `file_path` still wins.
- An overlong path was truncated from the head, cutting off the basename —
the one part that tells two rows apart. Trim from the front instead, so
the label reads `…/session/MobileNativeChatMessage.tsx`.
- The primary-argument chain used `??`, so a present-but-blank key selected
itself and swallowed the keys ranked after it, dropping the label all the
way back to raw JSON. Take the first key that actually yields a label.
Refs STA-3333.
* fix(mobile): don't offer an expander whose detail repeats the row (STA-3333)
An empty tool input formats back to the row label verbatim, so `{}` and `[]`
advertised an expander and then re-showed the label — the same repeat-the-JSON
problem this change set out to remove. Gate `isStructuredToolInput` on the
collection actually having contents; the lazy detail path is untouched.
Also pins the overlong-path test to the path itself: asserting only length<=80
plus a `…` passed just as well with path labelling deleted.
* fix(mobile): gate the tool detail panel on having detail (STA-3333)
The Tools toggle opens every row at once, bypassing the row's tap guard,
so a row with nothing to expand rendered its own label again underneath
itself — and the tap that would dismiss it is a no-op. Matches desktop.
* fix(mobile): keep a blank tool argument out of the run header (STA-3333)
Skipping a present-but-blank primary key let `briefToolArg` fall through
to the raw JSON preview, so a run header read `Bash {"command":""}` where
it used to read `Bash`. Also state the search-path trade-off honestly:
suppressing the link costs a file-scoped search its tap target.
* fix(mobile): only treat a blank primary key as a missing argument (STA-3333)
The previous guard tested key presence, so a populated but non-string
argument — a mixed argv like ['kill','-9',pid], or a structured query —
dropped out of the run header instead of falling back to the preview.
* test(mobile): pin the tool-row chevron to the detail panel (STA-3333)
The panel gate was covered but the chevron beside it was not: swapping
`showDetail` back to `expanded` on the icon alone left all 909 mobile
tests green, so the affordance lie this branch fixes could return
unnoticed — a down-chevron over no panel, on a row whose tap is guarded
off.
Asserts both icon counts on the fixture that test already renders. The
two halves now die for distinct reasons: the panel gate on the duplicate
label text, the chevron on the icon count.
* test(shared): pin the blank-search-key guard in the tool label (STA-3333)
Dropping `.trim()` from summarizePrimaryToolArg left all 32 tests green,
yet it leaks through isSearchToolInput: a whitespace-only `query` starts
counting as a search term, which suppresses `path`. One character takes
the row's label, its tap-to-open link and its run-header argument at
once, and puts the raw JSON label back — the bug this branch removes.
Asserts all three outputs on that shape. Kills only that mutant; the
isSearchToolInput mutant still dies on the existing search test.
* fix(native-chat): share tool input display semantics (STA-3333)
Build the tool row label, file target, detail eligibility and bounded detail from one normalized input model. Mobile no longer reparses JSON-string input across independent helpers or repeats an already-complete plain label, and desktop now uses the same clean row summary instead of retaining raw JSON.\n\nKeep full detail formatting lazy for collapsed rows and share the 4000-character detail cap across both renderers. Tests pin desktop adoption, mobile disclosure parity, one-pass JSON parsing and the shared bound.
|
||
|
|
c3ddc0d5df |
fix(mobile): keep native chat ask dismissals tab-scoped and gated (STA-3333) (#12497)
* fix(mobile): keep native chat ask dismissals tab-scoped and gated
Dismissal state lived in the chat view subtree, which unmounts on a
chat<->terminal toggle, so an answered ask card came back on return. It
also had no tab scope and no waiting/blocked gate.
- move dismissal into the controller, keyed per session tab
- gate ask cards on waiting/blocked like the permission path already is,
and retire a dismissal off the ungated detected prompt so a working/done
status can't be mistaken for the prompt clearing
- ignore a dismissal that settles after its prompt cleared or was replaced
Refs STA-3333.
* fix(mobile): keep an ask dismissal through the transcript re-subscribe
A view toggle or tab switch re-subscribes the native-chat transcript, and
useMobileNativeChatSession withholds `messages` until that read settles. A
transcript-derived ask therefore reads as null while the chat surface is
already visible, so the reset effect took it as "the agent moved on" and
retired a live dismissal — the answered card came back, which is the bug
the off-chat guard was meant to close.
Treat an unobserved null as unobserved: `observing` now also requires the
read to have settled. A prompt that is already detected stays observable on
its own, so a status-derived ask still registers on first paint and an
answer taken during that first load is still accepted.
* fix(mobile): keep the transcript-derived ask outside the paused gate
A hook row idle past AGENT_STATUS_STALE_AFTER_MS (30m) projects to `done`
with no interactivePrompt, so the transcript fallback is the only source
left for a still-pending question. Gating it behind waiting/blocked made
that question unanswerable from mobile. Only the sticky status payload
needs the gate; `extractPendingAsk` clears itself on the tool result.
Also pins the load-window clause in the ask-observability guard, which
was behaviourally load-bearing but killed no test.
* fix(mobile): treat a never-read transcript as unobserved, not as "no ask"
The ask-observability guard only excused `transcriptLoading`, which is true
for an in-flight read alone. useMobileNativeChatSession also withholds
`messages` when the client is gone ('idle') or the tab has not reported a
provider session yet ('waiting-session') — both leave the flag false over an
empty list that was never read. The derived prompt then read as null, the
reset effect took that as "the agent moved on", and a live dismissal was
retired; when the read landed with the question still pending the answered
card came back — the resurfacing bug this guard exists to close.
Gate on the read having actually settled instead. 'error' still counts: it
keeps the last successful read in `messages`, so a prompt that clears under
it is real evidence, unlike a list that was never populated.
Also locks three guards that killed no test: the sticky-status suppression
of the transcript fallback (which is what makes the new paused gate hold in
the post-answer window), the reset effect's identity bail-out, and showAsk's
empty-prompt case. The transcript stand-in now derives `transcriptLoading`
from `status` the way the real hook couples them, so these tests can only
express states the session hook can reach.
Refs STA-3333.
* test(mobile): pin the ask dismissal's tab scope and ungated retirement input
Both wirings were unpinned: swapping `scopeKey` to a constant or feeding the
gated `ask` in as `detectedAsk` left the whole mobile suite green.
* fix(mobile): require a landed read before an errored transcript retires a dismissal
`status === 'error'` was treated as settled on the claim that an error keeps
the last successful read in `messages`. That only holds for an error that lands
on top of an earlier read. The host forwards an initial-drain failure as an
error frame carrying an EMPTY list (transcript-watch-error.test.ts), the mobile
frame applier checks `frame.error` before the messages array so those rows are
discarded, and the session hook's error path never calls `setMessages` — so a
first-read error leaves `messages` at the `[]` the identity-change effect wrote.
That frame is also not terminal: the watcher keeps `initialDrain` true and a
real snapshot follows once the read recovers. So a re-subscribe whose first
read errors made the never-populated list read as "no ask", retired the live
dismissal, and the recovered snapshot brought the answered card back over the
composer — the exact resurfacing this guard exists to close, and most likely on
remote/SSH transcript reads.
Require rows for the error case. Rows can only be present once a read landed,
so the predicate is never wrong in the resurfacing direction; it only declines
to retire a dismissal when the transcript was never observed.
Also drop the dismiss hook's `detectedAsk = ask` default and make both prompts
required. That default silently fed the gated prompt in as the detected one,
which is the pre-fix behavior: a paused-out card would read as "prompt gone"
and retire the dismissal. tsc now enforces the ungated payload at every call
site instead of leaving a trap for the next caller.
* fix(mobile): scope the ask dismissal to the provider session, not the tab
A restart, /clear, or resume swaps the provider session inside one tab. The
next session's first question is often byte-identical, so a tab-keyed dismissal
hid the live card and left the turn blocked with nothing to act on.
* chore: restore upstream formatting
|
||
|
|
38a892c980 |
feat(mobile): native-chat model/session-option picker + shared slash catalog (STA-3332) (#12366)
* feat(mobile): native-chat model/session-option picker + shared slash catalog (STA-3332) Piece A — shared slash catalog + send classification: - Mobile composer now serves getVerifiedNativeChatCommands from the shared catalog (agent-aware, with description rows) instead of a hardcoded provider-agnostic list that advertised commands Claude does not have. - classifyNativeChatSend moves to src/shared/native-chat-slash-commands.ts (renderer re-exports keep desktop import paths stable); mobile's send seam now gates optimistic echoes on it, so slash sends no longer create a 'Queued' bubble that no transcript echo can ever retire, and the ack-lost hold only arms for chat sends. Piece B — mobile model/session-option pickers: - New per-tab session-option tracking (state/commands/labels modules) ported from the desktop live flow, reading the shared agent-session-option catalog for Claude AND Codex. - Composer pill row (model + options) opening an inline choice card in the proven Ask-card pattern; applies use catalog modelApply semantics (/model <value> via the existing send path), Codex-style agent-picker entries dispatch the picker command and flip the tab to the terminal view. - Current model seeds from the hook-reported provider model when derivable; typed /model-style commands update tracked state (recordOutgoingCommand parity); dispatched values render as sent-not-confirmed. * fix(mobile): keep session option sends scoped * fix(mobile): synchronize native chat refs after commit * refactor: share native chat session option logic * fix(mobile): keep the live tab's session-option record from eviction `getScopedRecord` returned an existing record without re-inserting it, so the per-tab record map evicted by insertion order rather than recency. A long-lived active tab is the oldest key, so crossing the 32-scope cap silently dropped its tracked model and reset the pill to "Model". Desktop's scope cache does delete-then-set for exactly this reason. Also moves the shared session-option tests to src/shared so the root suite runs them (they only exercised src/shared logic the Electron renderer consumes, but sat under mobile/ where only mobile's vitest project sees them), and restores two "why" comments dropped while extracting the shared modules. * fix(mobile): stop a stale session-start report reverting a model pick Re-entering a chat tab re-delivers the same `agentStatus.model`, and the reported-model effect re-applied it unconditionally — so picking a model, moving to another tab, and coming back reverted the pill to the model the agent reported at session start, which cannot have observed the `/model` sent after it. The status stream reconnecting had the same effect. A report is now only treated as evidence when the matched catalog id CHANGES for that scope; a genuinely new report still supersedes a local pick. Mobile has no screen read to confirm a switch against, so the repeat is all we can key off. * fix(mobile): close four session-option picker defects found in review D1 — a picker apply could interleave with a composer send. The composer already blocks a text send while an apply is dispatching, but not the reverse: the host spaces a send's body and its Enter ~500ms apart, so an apply tapped inside that window was submitted as part of the user's prompt, and the pill then claimed a model change that never ran as a command. The pickers render inside the composer, so they now take its in-flight state directly — the same guard, mirrored. D2 — an option was filed under the wrong model. `setTrackedSessionOption` resolves the owning model when it commits, not when the command was built, and the report effect mutates the same record off-queue. A report landing mid-dispatch therefore recorded `/effort low` against the model it switched TO. Ports desktop's supersession guard, which skips the commit when the baseline moved. D3 — a command template's prefix also matches prose that starts with it, so "/model is a weird word" tracked that prose as the current model, rendered it as the pill label, and matched no catalog model, dropping every per-model option. Parsed values are now canonicalized against the catalog; a typed value containing whitespace is treated as a prompt rather than a command. Perf — `/` on a Codex tab returned all 45 commands into a non-virtualized ScrollView showing ~5, re-reconciled on every streaming tick above the transcript. Capped at 12. Also splits the row primitives out of MobileNativeChatSessionOptionPickers.tsx, which the D1 guard pushed to 402 effective lines against a 400 cap. * refactor: share the session-option display ordering CATEGORY_ORDER and the non-model sort were byte-identical in NativeChatSessionOptionPickers.tsx and mobile's labels module — pure logic with no i18n in it, so there was no reason for two copies that can drift. Both now call sortNativeChatSessionOptions from the shared snapshot module. * refactor(mobile): align model picker layout * style(mobile): round native chat composer * fix(mobile): inset rounded chat composer |
||
|
|
5ed45739e9 |
fix(runtime): make sibling-workspace terminal-path resolution an explicit client opt-in (#12616)
files.resolveTerminalPath began returning a foreign worktree id + relativePath for absolute paths owned by a sibling workspace, with no protocol or capability gate. Mobile 0.0.36 in the field ignores resolved.worktree and reuses its own worktree id for the follow-up files.open, so a tap on a sibling-worktree path opened the WRONG worktree's copy of that file (on 1.4.168 the tap was a safe no-op). Gate the sibling-workspace lookup behind a new optional crossWorkspace request field: clients that honor resolved.worktree opt in; everything else keeps the pre-sibling-resolution contract. Old servers strip the unknown field (zod), so every version pairing degrades to the safe legacy behavior. Optional-field addition, so no RUNTIME_PROTOCOL_VERSION bump per protocol-version.ts rules. The terminal-path RPC tests move to files-terminal-path-resolution.test.ts because files.test.ts sits at the max-lines cap. Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
39c3c58d55 |
perf(runtime): gate terminal.list visual layouts (#12450)
* perf(runtime): gate terminal.list visual layouts and stop the false writable claim visualLayouts is ~31% of a large terminal.list payload (44,208 B of 137,412 B on a live 134-terminal remote runtime) and has exactly one consumer: the human-readable CLI formatter. Gate it behind an includeVisualLayouts request param that defaults to included, so pre-flag clients are unaffected, and have every --json/internal caller opt out. Also drop the record-backed builder's writable, which was a verbatim copy of connected. terminal.show now states writability explicitly as exactly what terminal.send's PTY gate enforces. * test(runtime): type the payload-size fixture arrays for tsc * fix(runtime): preserve terminal list compatibility * test(runtime): guard terminal list optimization * fix(cli): preserve agent access to terminal layouts |
||
|
|
7948e46db8 |
fix(mobile): open the Resume workspace through a mounted host stack (#12001)
* fix(mobile): open the Resume workspace through a mounted host stack Tapping Resume on Home landed on a blank host screen instead of the session. A cold push straight into the nested /h/[hostId] navigator resolves to the host index route without the dynamic id, so HostProtocolGate mounts with hostId undefined and never connects. Home already worked around this for the host editor (#11635) and Tasks (#11853) by mounting /h/[hostId] first and replacing it once the stack is committed. Extract that mechanism into host-stack-navigation so Resume uses the same transition instead of a direct push. The previous Resume fix (#11876) swapped the manual href for a typed dynamic href, but expo-router's encodeParam already applies encodeURIComponent to dynamic segments, so it resolved to the same URL the manual string produced and left the cold-navigator path unchanged. Claude-Session: https://claude.ai/code/session_01RMoaxp7MLg2ydP28KFLX7B * fix(mobile): harden the host-stack transition after bot review - match a host route committed as the encoded segment it was pushed as, so an id containing `/`, `#`, or `%` still triggers the REPLACE - share one pending transition across the Home entry points; per-hook refs let a Tasks tap and a Resume tap arm two pushes that could not cancel each other - assert the source markers before slicing in the Resume wiring test Claude-Session: https://claude.ai/code/session_01RMoaxp7MLg2ydP28KFLX7B * test(mobile): lock the host-stack transition state machine Assert the replace waits for the host mount (zero dispatches before it, exactly one after), and cover cancel/retarget — the paths the shared pending transition relies on. * test(mobile): model listener removal in the navigation harness A no-op unsubscribe let setState keep calling a canceled listener, so the teardown assertions only exercised the active guard. Dropping the unsubscribe call from dispose now fails the suite. --------- Co-authored-by: kaynan <kaynan.camargo@terceiro-sky.com.br> |
||
|
|
535b5a594d |
fix(mobile): keep paged chat history coherent across reconnect replays (STA-3333) (#12494)
* fix(mobile): keep paged chat history coherent across reconnect replays The transport replays nativeChat.subscribe with its original params after an in-place reconnect, and the session hook treated every snapshot as a fresh base — so a socket blip truncated paged-in history back to the initial 40. A replay snapshot that extends a contiguous retained tail now merges in by id; a disjoint replay (long outage, compaction while away) still replaces, since stitching would leave a silent gap. Only a genuinely replaced window resets the grown read limit and paging cursor, and any snapshot invalidates an in-flight older-page request so a stale cursor result cannot land on the new window. Refs STA-3333. * test(mobile): pin the replay contiguity-scan rejection branches The scan's three rejection rules were unpinned: deleting the `sawNewMessage` guard, the ordering check, or the retained-tail anchor each left the whole suite green. Cover the interleaved-new-row, reordered-id, and short-of-tail cases, and the replaced-window hasMore fallback. Each new test is mutation-proven to kill exactly one mutant. * test(mobile): pin replay paging-metadata and base-snapshot authority Two more branches of the replay logic were unpinned. Adopting a replay's `beforeOffset` when it starts partway into paged-in history would make the next loadEarlier re-fetch on-screen rows and prepend duplicates; treating a post-replacement snapshot as a replay would retain a row the authoritative window dropped. Both mutants now fail exactly one test. * test(mobile): pin the replay removal boundary and base-snapshot bookkeeping Two branches introduced by this PR survived the suite unpinned: - `firstIndex > 0` was only pinned one-directionally. Weakening it to `firstIndex > 1` kept all 28 tests green, so an off-by-one would silently retain one row the host had already dropped. - `snapshotSeenRef` is set only for snapshot frames. Setting it unconditionally is invisible in normal flows, where the first frame is the snapshot, but demotes the real base snapshot to a replay when a live append lands first. Each new test kills exactly one of those mutants and nothing else. No source change. * test(mobile): pin the older-page fence against a cursor-re-cutting replay The snapshot arm of `if (applied.windowReplaced || frame.type === 'snapshot')` was unpinned: deleting it kept the suite green. It is load-bearing. A replay that merges cleanly can still carry a new `beforeOffset`, which the hook adopts via `replayStillStartsAtOldest`. The page already in flight was addressed with the old offset, so without the fence it lands and writes its own stale cursor back over the fresh one, leaving the next `loadEarlier` addressed from a byte offset that no longer describes the file. Row order alone stays correct, which is why the ordering-only reasoning missed this. Sole failure under the mutation. No source change. |
||
|
|
9ee359550b |
fix(mobile): make native-chat file links and path citations tappable (STA-3331) (#12364)
* fix(mobile): make native-chat file links and path citations tappable (STA-3331)
- Linkify POSIX absolute paths in chat prose (leading-/ regex alternative;
URL guard now keys off the char before the matched slash)
- Parse agent-style path:line(:col) citations in prose, code spans, and the
open flow; line/column ride into the mobile file preview route
- Route non-web markdown hrefs (file: URIs, relative/absolute paths) to the
file opener instead of silently dropping them; unknown schemes stay dead
- Resolve chat paths against the worktree root, not the terminal's live cwd
- Reuse the terminal tap-to-open flow for chat taps (haptic, preview route,
tab activation with retries) via a shared identity-stable hook, and toast
on misses instead of silent no-ops
- Keep snake_case paths whole (intraword underscores are literal text),
scan bold/italic/strike spans for paths, split trailing punctuation off
autolinks, and let taps land while the composer keyboard is up
* fix(mobile): harden chat file tap handling
* refactor(chat): share native chat href routing
* fix(mobile): detect files directly under path roots
* fix(mobile): keep inline tokens and dunder paths intact around emphasis
Review follow-ups on the chat file-link work:
- A rejected intraword `_` token left the scan index past its closing
underscore, so every inline token between two snake_case words was
swallowed and rendered as literal source — including markdown links,
which became untappable. Rescan from just past the opening delimiter.
- Treat a path separator as an intraword flank so `src/__init__.py` and
`a/__tests__/x.ts` stay whole; previously they rendered as bold plus a
remnant that the new absolute-root pattern turned into a tap on `/x.ts`.
- Bound the `:line(:col)` tail so `src/app.ts:1e3` and `:80%` no longer
parse a line number, while a cited range still opens its first line.
- Route chat tap failures through the composer banner (toast fallback):
chat taps happen with the keyboard up, which covers the toast.
- Drop the tap-handler mirror's dep list; the call site rebuilds its
accessors every render, so it could never skip on a route that
rerenders per keystroke.
* Revert "fix(mobile): keep inline tokens and dunder paths intact around emphasis"
This reverts commit
|
||
|
|
d52df52eea | fix(mobile): open editor for disconnected hosts (#12575) | ||
|
|
d27d69fbff |
fix(mobile): serialize native chat PTY writes and fence superseded ask keystrokes (STA-3333) (#12502)
* fix(mobile): serialize native chat PTY writes
Two composed native-chat write sequences into one PTY interleaved their
bytes: the per-terminal send-in-flight guard lived inside the image
attachments hook, so ask answers, permission choices, and question
answers wrote straight past it. Move the guard to a shared module-scope
write lock (the terminal outlives any one screen) and take it on all
four paths.
Ask answers additionally queue behind the prior chain's RPC rather than
racing it, and a superseding answer is fenced once a key has actually
landed: an accepted or ambiguously-delivered keystroke already moved the
remote selector, so a replacement's from-scratch key plan would answer
the wrong question. A superseding answer inherits the cancelled chain's
hold (refcounted, last chain out releases) so changing your mind
mid-answer still works, and Stop/cancel stay unguarded so an interrupt
can never deadlock against the send it cancels.
Refs STA-3333.
* fix(mobile): report a fenced native-chat answer instead of dropping it
The fence added for superseding Ask answers returned false with no
onSendError, and the card re-enables on a false result — so a queued
answer vanished with no banner and no toast, looking exactly like a dead
button. Every other bail in answerAsk reports. Keep the generation-
mismatch branch silent: a newer answer owns the error surface.
Also covers the hold-count release, which had no test at all: replacing
it with an unconditional release left all 58 tests green while silently
reopening the terminal mid-sequence — the exact interleave this PR fixes.
* fix(mobile): stop a superseded answer from clearing the fence banner
A chain that finished its key plan reported `true` even after a newer answer
superseded it. The route sends through useNativeChatAcceptedAction, whose
accepted callback retires the send-error banner — and that callback runs after
the successor's fence report, because finishTurn() fires in `finally`, before
the chain's own promise settles. So the successful predecessor deterministically
wiped the fence message the successor had just raised: every healthy write took
that branch, which made the previous commit's report vacuous exactly where it
mattered.
A superseded chain now reports no success, matching every other supersession
checkpoint in this hook.
* fix(mobile): fence on the turn slot, not the generation counter
The supersession guards added in
|
||
|
|
9759bd2e76 | fix(mobile): center native chat checkmarks (#12565) | ||
|
|
e4aadcceff |
fix(mobile): keep repeated-prefix native chat replies streaming (STA-3333) (#12501)
* fix(mobile): keep repeated-prefix chat replies streaming Text alone can't tell "the transcript caught up with this stream" from "a new reply repeats the previous turn's prefix", so the old suppress-on- prefix rule swallowed genuine repeated replies. A stateful gate remembers which transcript tail predates the current stream segment and hides the bubble only when that tail moved during the segment, scoped to the active host/workspace/tab/session so a swapped chat can't inherit a baseline. Refs STA-3333. * fix(mobile): keep the streaming gate alive across chat/terminal toggles The gate lived in MobileNativeChatView, but MobileNativeChatOverlay returns null whenever the user peeks at the terminal — that unmounts the view and throws the baseline away, so the repeated-prefix reply was swallowed again on the way back. Move the gate (and the fold memo it reads) up to the overlay, which stays mounted across those toggles. While hidden the transcript is empty and the throttled stream reports no text, which the gate would have read as "idle" and re-anchored on. Pass the agent's working state so a textless tick inside a live segment holds the baseline instead. The scope key is now keyed off the tab rather than the view-gated chat resolution, so it survives the toggle too; streamIdentity keeps its exact previous value because the delayed-send guards compare against it. Also drops a dead disjunct in the caught-up test: a null baseline is already unequal to every real tail id. * test(mobile): model the real re-show ordering in the streaming-gate tests The overlay regression test replayed the transcript before the stream text on the way back from the terminal view. That ordering is backwards: the session withholds `messages` until a fresh read settles (an RPC round trip) while the throttled stream text returns in ~50ms — and with the transcript already back, a gate that got discarded on the toggle still passes. Replay the real order, which pins the gate's lifetime as intended. Swaps the hidden-gap duplicate case for the in-view one (a tool frame clears the assistant text mid-turn), which is where the hold actually earns its keep; the hidden-gap direction stays covered at the gate level. * fix(mobile): stop the streaming gate adopting a reply as its own history A textless status tick was re-anchoring the gate's pre-stream baseline, so two paths still rendered wrong: - The reply's transcript push beats its throttled status text whenever the pane stays `working` past the turn (a live subagent or background task). The tick in between adopted the just-landed reply as history, and the status text that followed rendered it a second time — a duplicate bubble, and a regression against main's suppress-on-prefix rule. - Peeking at the terminal between turns empties the transcript. That empty tail was adopted as the baseline, so the next repeated-prefix reply was swallowed again — the bug this PR exists to fix. Only a tick that carries a real tail and sits outside a live turn anchors now, with an exception for a gate that has never anchored: mounted mid-turn, the first real tail it sees is the best history it will ever get. Also drop `buildMobileNativeChatData`, a test-only builder this PR had wired the new gate into; its green test asserted the exact suppression this PR removes. Its fold/pending/image coverage moves to the builder the view calls. * test(mobile): pin the textless anchor's text reset Mutation testing found the `prevText` reset on an anchoring textless tick unpinned: keeping the previous turn's text there reads the next turn's opener as a new segment, re-anchors onto the reply that just landed, and renders it a second time — the same duplicate-bubble class already fixed twice on this branch. |
||
|
|
96e31e7bf8 |
Remove a paired computer's deleted projects from every connected device (#12215)
* fix(repos): remove a paired computer's deleted projects from every connected device A project deleted on a paired Orca host stayed in every connected client's sidebar and could not be removed there. Two independent defects: 1. Host-local repo IPC mutations only sent `repos:changed` to the host's own renderer (src/main/ipc/repos.ts:2711). The runtime client-event stream was fed only by mutations arriving over runtime RPC, and clients refetch a remote catalog only on a `reposChanged` event -- there is no polling on desktop -- so the deleted rows persisted indefinitely. The shared `notifyReposChanged` helper now also calls the new `OrcaRuntimeService.notifyReposChangedForRemoteClients()` (src/main/runtime/orca-runtime.ts:5175), mirroring the existing `notifyWorktreesChangedForRemoteClients` precedent. This covers every repo, project-group and folder-workspace IPC mutation, so renames, colors, reorders and adds propagate too. 2. Deleting the ghost row on the client routed `repo.rm` to the owner, which answered `repo_not_found`. `removeProject` wrapped its whole body in one try/catch, so the rejection aborted the local purge before the `set()` (src/renderer/src/store/slices/repos.ts:3466) and the delete button appeared to do nothing. Only `repo_not_found` is now tolerated; any other failure still keeps the row, and an opt-in `errorFeedback: 'toast'` makes it visible at the three single-project user-initiated entry points. Bulk and background callers keep today's silence plus their own aggregate reporting. Closes #11994 Co-authored-by: Orca <help@stably.ai> * fix(repos): revert inert RepositoryPane removeProject arg The settings pane's only render site drops the argument; the toast is already delivered by removeSettingsProjectFromAllHosts. Co-authored-by: Orca <help@stably.ai> * fix(repos): scope duplicate-repo-id deletes to the owning execution host Cover the cross-host collisions #11994's broadcast now fans out to every paired device. Same-name projects on different hosts were already isolated (per-host UUIDs, host-scoped catalog merge and purge) and are pinned by regression tests. Two same-repo-id paths were not: `repo.rm` with a `path:`/`name:` selector and `deleteProjectHostSetup` both resolved one row and then deleted by bare id, taking the sibling host's registration with it. Co-authored-by: Orca <help@stably.ai> * test(mobile): align the poll-interval rationale with the new reposChanged emission Co-authored-by: Orca <help@stably.ai> * fix(repos): resolve deleteProjectHostSetup's repo row only on the setup's own host The sibling-host fallback could only ever pick a row on a host the caller did not name; with no exact match the setup is stale and the existing path already drops just the setup. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
fb1259a09d |
fix(mobile): keep cached workspace counts across a transient RPC failure (#12408)
* fix(mobile): keep cached workspace counts across a transient RPC failure The Home host card showed "12 worktrees · 2 active" until any worktree.ps failed — a backgrounded app, a Wi-Fi→cellular handoff, or a sleep/resume that kills the socket mid-request. Two things then went wrong: - render dropped the counts: `markHomeWorktreeCatalogUnavailable` kept the proven numbers in state, but the card only rendered them when `catalogUnavailable` was unset, so the line collapsed to "Worktree list unavailable" even though the last successful counts were right there. - nothing re-drove the fetch: the per-host wiring latched a `statsFetched` boolean on the first connect, and the logical client survives socket drops, so its reconnect never re-read the catalog. The card stayed wrong until the user navigated away and back. Keep the proven counts and flag them stale (`staleCounts`), rendered as "Last known: 12 worktrees · 2 active"; a host whose catalog never loaded still reads "Worktree list unavailable" (STA-3123). Replace the one-shot latch with createHostConnectRefetchGate, which fires on each transition INTO 'connected' — one refetch per reconnect, no polling timer — mirroring useWorktreeResync on the host screen. fetchHomeHostWorktreeInfo moves out of app/index.tsx so its rejection path is covered by tests. * fix(mobile): bound "Last known" counts and survive a path cutover Review found two ways the home host card's stale-count fix misbehaves. 1. A migrateTo cutover (relay->direct probe, forced replacement) rejects in-flight requests with LogicalClientCutoverError and republishes 'connected' from 'connected', so the connect gate never re-arms and the card latched on "Last known: ..." with nothing left to clear it. worktree.ps now re-issues on the authenticated replacement, bounded, like runtime-capability-probe and worktree-create-retry already do. 2. "Last known: N worktrees" had no age bound. The home snapshot is persisted, so a cold start whose first worktree.ps failed rendered counts proven days ago exactly like counts proven seconds ago - the case STA-3123 deliberately rendered as "Worktree list unavailable". Counts now carry countsProvenAt and expire out of the "last known" wording after 10 minutes; counts persisted by an older build count as expired. Also, per review: the card derives its own worktree line from HostWorktreeInfo, so a caller can no longer re-gate the counts away (that was the original defect), and the derivation is covered by a render test - mobile/vitest.config.ts never collected *.test.tsx, so component tests were silently dead. Home stats are keyed by host and summed instead of letting whichever desktop replied last overwrite the shared header row, which the per-reconnect refetch made churn on flaky links. * fix(mobile): age bounds liveness, not the counts; scope the header total to paired hosts Round-2 review follow-up. Age bound was anchored on proof time inside the failure branch only, so a session connected past the window that then hit one failed refresh rendered the pre-fix "Worktree list unavailable" — the exact case this PR exists for — while identically aged counts still rendered unlabeled as live whenever the refresh was merely pending. Age now decides live vs "Last known" and the failure branch keeps whatever the host last proved; "Worktree list unavailable" is reserved for a catalog that never loaded. Header stats summed every entry ever cached, so removing a desktop left its lifetime numbers in the total for the rest of the session. totalHomeStats now sums the hosts still paired, which also covers removal from the host screen. wireHostSubscriptions is the effect body moved verbatim out of useEffect; react-doctor's effect-needs-cleanup false-positives on `subscribe` inside one and the changed-code gate has no working suppression path (an inline directive reads as unused to the plugin-less scan). --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
5a2b329d8d |
chore(mobile): bump to 0.0.37 (versionCode 10) (#12365)
Completes the 0.0.37 release attempted on 2026-08-03 (run 30791649691 failed on the version assertion). Ships the post-0.0.36 transport fixes: relay session recovery when the LAN endpoint is unreachable (#12344, #11368, #11465, #11690) and honest worktree-catalog failure states (#12235) — the released-app defect class verified live tonight. Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
e8d3043107 |
fix(mobile): stop unrenewed-grace rotation churn and gate cadence gaps (#12426)
- skip proactive rotation when the resume confirmation reports renewed=false (a re-resume provably returns the same unchanged deadline; rotating churned one session replacement per clamp floor, ~60/hour, until a fresh credential) - armCredentialReprobe under a held gate mints the tick's pass token so the effective reprobe cadence stays 60s..15min instead of doubling to ~30min - registerFailure honors scheduleRetry=false in gate branches: no reprobe timer is armed while backgrounded/stopped; foreground resume re-arms - extract RelayRetryDelays and supervisor test fakes into their own modules (max-lines) Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
e59a319ffe |
fix(sidebar): keep each project's entry-point workspace visible under "Hide sleeping" (#12257)
"Hide sleeping" swept each project's main workspace out of the sidebar as soon as it had no live PTY, browser tab or agent — even with "Hide default branch" off. For a project whose only row is that workspace (a folder workspace, a fresh clone, a detached-HEAD main), the entire project vanished with no in-place way back. Adds a shared `isSleepingSweepExemptWorkspace` predicate keyed on `isMainWorktree` rather than the branch name, so folder workspaces (no branch), detached-HEAD mains, and SSH rows whose head/branch are blanked while a provider is disconnected all stay put. Wired into `computeVisibleWorktreeIds` (sidebar, Cmd+1-9, workspace board), the jump palette's duplicate inline pass, and mobile's `filterWorktrees`. Ships default-on with an escape hatch: a persisted `alwaysShowDefaultBranchWorkspace` setting surfaced as "Except default branch" under "Hide sleeping". Explicit "Hide default branch" still wins, since it filters before the sleeping sweep. Mobile reads the setting but never writes it back, so a desktop opt-out can't be clobbered by a filter tap before the ui.get roundtrip lands. Combines the two PRs open against #8873. #8966's exempt set is a strict subset of this one, so its production diff was subsumed rather than ported; its jump-palette render harness and e2e spec were carried over, and are the only such coverage here. Fixes #8873 Closes #8966 Co-authored-by: Rod Boev <rod.boev@gmail.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
5bd2f59d29 |
fix(runtime): open files from sibling workspaces (#11369)
* feat(runtime): match files to workspace owners * fix(runtime): resolve terminal paths through sibling workspaces * fix(editor): route restored sibling workspace files * fix remote sibling file ownership routing * fix(editor): migrate restored sibling file owners * fix(editor): revalidate restored owner activation * docs(review): record PR 11369 correction evidence * fix(editor): reject collision before activation prep * docs(review): record PR 11369 final correction * fix(editor): retain projected reconciliation narrowing * chore(review): keep verification artifacts out of PR * fix(editor): harden restored owner migration * fix(runtime): resolve workspace root terminal paths --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
9ba293cb74 |
fix(mobile): keep relay runtime recovery alive without direct connectivity (#12374)
* fix(mobile): keep relay runtime recovery alive without direct connectivity A phone paired over the relay whose direct LAN endpoint is unreachable (e.g. a Tailscale IP with Tailscale off) could lose the runtime channel permanently: the reconnect controller's recovery gates parked with no timer and no logs, the supervisor snapshotted relay credentials once at start (dying silently if the read failed and dialing stale tokens after rotation), and the only path that cleared a rejected-credential gate required a working direct connection. Field symptom: home card shows "Connected - Orca Relay" (or "Can't connect - check Tailscale") while the host page sits at zero worktrees forever. - gates (fresh-credential, external-signal) now arm a slow 60s reprobe instead of parking; each gated attempt re-reads the durable credential bundle and adopts it when its version is fresher than the rejected one - supervisor start no longer dies for the process lifetime when the initial Keychain read fails or the bundle is expired - every recovery decision now reaches logcat and the in-app connection log ([relay] lines); previously the whole relay dial path was silent - direct-return probing extracted to mobile-direct-return-probe.ts, credential selection to mobile-relay-credential-selection.ts Regression suite mirrors the field failure (rejected outer credential, unreadable bundle at start, expired bundle, E2EE rejection without a UI nudge) plus real-rpc-client failover integration tests; the four deterministic scenarios fail on the previous code. * fix(mobile): adopt durable relay credentials by outcome, not version Adversarial review caught two blockers in the version-comparison rule: renewals extend expiresAt without bumping current.version, and a re-pair restarts the version counter — both left the durable bundle unadopted and reproduced the original outage. Selection now adopts the disk bundle exactly when it yields a dialable (unexpired, non-rejected) credential while memory does not, which also keeps revoked versions unresurrectable. Also from review: the gate reprobe cadence now escalates 60s -> 15min ceiling with 0.75-1.25x jitter (no fleet phase-alignment, no permanent one-minute beacon); clearing a gate drops its timer, pending tick, and cadence so an orphaned reprobe cannot swallow the next fast backoff; the reprobe tick token is only minted while its gate still holds; and a merely missing/expired bundle uses a plain cooldown instead of the fresh-credential gate so it cannot force rotations on direct reconnects. New regression tests (all red on the previous code): renewal without a version bump, re-pair with a restarted counter, orphaned-timer backoff swallowing, escalating gated cadence, and background/foreground recovery after an E2EE rejection. * fix(mobile): reset gated relay cadence on app resume Review round 2: an escalated fresh-credential gate kept its cadence across background/foreground, so reopening the app could wait out a 15-minute tick (measured 11.25min to first attempt after a 2h background) — indistinguishable from the outage itself. A resume now resets the streak even when it cannot lift the credential gate, and a successful direct connection does the same in resetForDirectConnection. Also: the streak now advances once per fired tick instead of once per armed-delay computation (three arms per cycle escalated 60s -> ceiling in ~7 minutes instead of the documented eight steps); delay computation is a pure read. * fix(mobile): rotate relay sessions on resume expiry, not attach deadline Live phone verification of the failover fix exposed a second defect the old latch had been masking: the relay-hello's leaseExpiresAt is the cell's attach-reservation deadline (now + 10s for resumes, credential-store.ts:213 server-side), but the supervisor scheduled proactive rotation from it with a 30s margin clamped to 1s — so every relay runtime session force-replaced itself ~1s after connecting (measured every ~2.5s on device, 253 dials per 5 simulated minutes in the red test). Any RPC slower than the cycle could never complete, which is the "Worktree list unavailable" symptom. The session now captures resumeExpiresAt from the hello (updated by the resume confirmation) and rotation keys off it. Test fakes previously used a 120s lease, which is why no suite ever reproduced the loop; they now mirror the production 10s attach deadline, and a churn regression holds one session across 5 minutes with direct unreachable. * fix(mobile): clamp lease rotation delay on both ends Adversarial review of the resume-expiry rotation fix caught an int32 setTimeout overflow: production resumeTtlMs is 30 days, and 30d - 30s = 2,591,970,000ms exceeds INT32_MAX, so Node (and vitest's fake timers) clamp the timer to 1ms — 3001 relay dials and credential writes in 3 simulated seconds, ~2500x worse than the churn being fixed. The delay is now clamped to [60s, 6h]: the ceiling makes overflow unreachable regardless of server TTL (a harmless re-resume every 6h on long sessions), and the floor bounds any bad deadline to one forced rotation per minute instead of a sub-second loop — which also disarms the Math.max(1000, ...) landmine for return-unchanged-grace resumes whose stored expiry can be arbitrarily near. Also from review: getLeaseExpiresAt is renamed getAttachDeadlineAt (it had zero production callers left; the plausible name is how the churn bug happened), the expired-vs-missing bundle cases now log distinct strings, and both test fakes use production constants (10s attach deadline, 30-day resume TTL) — fictional fake values hid all three defects in this subsystem. The four forced-rotation lease tests are retimed to the 60s floor with direct pinned unreachable so return probes cannot race their windows. * style(mobile): merge duplicate imports in relay failover test * style(mobile): use T[] array syntax in credential selection --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |