mirror of
https://github.com/stablyai/orca.git
synced 2026-09-22 16:02:32 +00:00
99d19d4635c40e23b2ef546136b22fa3efeb7367
1529
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
99d19d4635 | feat(vm): add provisioned root recipe contract (#14352) | ||
|
|
953cfab635 |
fix(terminal): make the bold font weight its own setting (#14368)
* fix(terminal): make the bold font weight its own setting
Deriving bold as max(700, regular + 200) silently destroyed bold. A family
exposes only a few real faces: the monospace the default chain resolves to on
macOS has exactly two, splitting at 600. Measured by rasterizing each weight to
a canvas — 100-500 are byte-identical (ink 3023) and 600-900 are byte-identical
(ink 3855), at every weight the same advance. So any base weight at or above 600
put both values in the same face and bold stopped existing, on 4 of the 9
positions the slider offers, with no error and nothing the user could do.
Arithmetic cannot fix it — on a two-face family there is no heavier face to
escape to. So bold is now user-owned: a new terminalFontWeightBold setting with
its own control, defaulting to 700. The default pair (500/700) straddles the
boundary, so existing profiles render exactly as before; a collision is now a
choice the user can see and undo.
The old test asserted 800 -> {800, 900} as 'keeps bold heavier', which is where
this hid: numerically heavier, identically rendered.
* fix(terminal): surface bold face collisions accurately
|
||
|
|
11cd2b4310 |
revert(ssh): back out #13326 and #13928 — reconnect loses every tab (#14361)
* Revert "fix(daemon): stop killing live coding agents when the daemon can't report its sessions (#13928)" This reverts commit |
||
|
|
281cc77e79 |
fix(ai-vault): make the merged scan stamp independent of leg order (#14270)
* fix(ai-vault): make the merged scan stamp independent of leg order
The all-host merge picked its stamp with a strict `stampMs > latestMs` and
echoed the winning leg's verbatim string. Two legs reporting the same instant
in different legal ISO shapes ("...:05Z" vs "...:05.000Z") therefore resolved
by position in the results array, i.e. by host-enumeration order (local, then
SSH, then runtime). The prior lexicographic max was order-independent, so this
was a regression with no test covering it.
A merge has no single scan instant, so its stamp is derived data rather than
any one leg's string: return the canonical ISO form of the newest accepted
instant. That is order-independent and format-independent, and drops a
variable instead of adding a tie-break branch.
Also share one request resolver between main and the renderer so the
renderer's merged-scope predicate is equivalent to main's routing by
construction, rather than by a comment that overclaimed it.
* docs(ai-vault): scope the merged-predicate comment to the desktop IPC path
The replacement comment still asserted the result is always several hosts'
legs. The paired web transport drops executionHostScope and serves one host,
so 'all' there is a single scan. State that the predicate is deliberately
over-inclusive and why erring the other way would be unsafe.
* test(ai-vault): pin the merged-stamp Date range boundary
new Date(ms).toISOString() throws RangeError outside +/-8.64e15. That is
unreachable only because Date.parse applies TimeClip, so the NaN guard alone
constrains the argument. Nothing pinned that. Dropping the guard now fails
these two cases with the RangeError they exist to prevent.
* refactor(ai-vault): route session-title scope through the shared resolver
The last character-for-character copy of the request-scope default. Leaving
it would make the shared resolver the single source of truth for two of three
sites, which is the drift this change exists to remove. No behavior change.
|
||
|
|
2f0c33757d |
fix(worker-start): match Codex effort ceilings (#14281)
Honor the advertised reasoning-effort ceilings for Codex models, preserve conservative unknown-model handling, and localize the new ultra effort label. |
||
|
|
0824ed39ea |
fix(terminal): clear the SGR pen on hidden-output restore and abandon (#14241)
* fix(terminal): clear the SGR pen on hidden-output restore and abandon The hidden-delivery gate drops renderer-bound PTY bytes while a pane has no visible view. The renderer's xterm is a separate emulator from the daemon model, so when the dropped span contains the sequence closing an attribute run (e.g. the ESC[22m ending a bold run) the renderer's pen stays latched while the daemon model stays correct. Neither recovery path cleared it: - buildMainModelSnapshotReplayWrites reset the pen on the two alt-screen branches but not on the normal-buffer branch, and replayed scrollbackAnsi ahead of the reset it did emit, so replayed content inherited the stale pen. - abandonHiddenOutputRestoreAndDrainPendingForeground declares the dropped bytes unrecoverable (it writes a user-visible warning) and then drained the queued foreground chunks straight into xterm under that same unknown pen. Add RESET_GRAPHIC_RENDITION and emit it ahead of replayed content in every branch, and on both abandon exits. The existing profiles all clear DEC mode bits and none touched SGR. * fix(terminal): also restore charset designation after a dropped-byte gap A gap can strand more than the pen: a dropped `ESC(B` leaves line-drawing selected and ordinary text renders as box characters. Route both recovery paths through one RESET_AFTER_BYTE_GAP profile covering SGR + charset. Deliberately not a soft reset (DECSTR): xterm's DECSTR wipes kitty flags and stacks (terminal-kitty-keyboard-mode-tracker applySoftReset), which would silence Option chords for a live agent that negotiates them only at startup. Reset what a gap strands and no running TUI re-asserts on its own; leave the rest to its next repaint. * fix(terminal): close the emulator state gap where the drop is announced The restore-needed marker is the single point where "renderer-bound bytes were dropped" is known. The handler already resets the transport's cross-chunk parser state there for exactly this reason — a partial escape spanning the gap would corrupt the next chunk. The emulator carries state across chunks in the same way, so reset it in the same place. That makes restore, abandon and overflow all start from a known pen by construction, instead of each recovery path having to remember. * fix(terminal): fully ground byte-gap recovery state * fix(terminal): reset state when remote restore re-arms * fix(terminal): keep the gap reset on the warning abandon path The reset had been folded into an else of the unavailable-warning branch, so the primary abandon path relied on the marker's earlier reset still standing. It does not always: this function captures a replayingSnapshot, so it can run after a partially-applied replay has already moved the pen, and the warning itself is plain text carrying no SGR. Restore the unconditional write, guarded only against the remote re-arm which writes its own. * fix(terminal): scope the byte-gap reset to the pen and skip it under flood Two regression risks in the widened recovery reset, both removed: - The profile had grown to cancel partial escapes, close OSC 8 and re-designate all four ISO 2022 registers. Each changes what a live TUI sees on a path that runs in production, and none has a reported symptom behind it — a legitimately line-drawing TUI that does not re-designate after recovery would render box characters as ASCII. Scope back to SGR, which is what the field reports show. - The marker-time reset ran before the flood-backpressure guard, so a flood wrote one reset per marker in exactly the case that guard exists to damp. Move it after; the flood path repaints through buildMainModelSnapshotReplayWrites, which grounds the pen itself, so no coverage is lost. Coverage verified non-vacuous: blanking RESET_AFTER_BYTE_GAP fails 5 tests across all four paths (replay branches, marker, abandon-with-warning, remote re-arm). |
||
|
|
b908b55f6d |
feat(worktrees): add per-source visibility controls (#14189)
* feat(worktrees): add per-source visibility controls * fix(worktrees): explain unsupported visibility hosts * fix(worktrees): keep add location form inline * fix(worktrees): align source visibility across runtimes * test(worktrees): cover Windows drive-relative roots |
||
|
|
4882eeb8ac |
rm git shim: neutralize stale wrappers without a host gate (#14255)
* Revert "fix terminal attribution shim removal edge cases (#14187)"
This reverts
|
||
|
|
4c5f818187 |
refactor(skills): remove the unreachable Skills page and the file count it rendered (#14259)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
3984023375 |
feat(workspaces): make workspace board shortcut a toggle (#14240)
* feat(workspaces): make workspace board shortcut a toggle The workspace.openBoard command only opened the board; pressing the bound shortcut again was a no-op, so closing required Escape, the toolbar button, or collapsing the sidebar. Bind the shortcut bridge event to the existing toggleWorkspaceBoard so one shortcut both opens and closes. Rename the bridge event to TOGGLE_WORKSPACE_BOARD_EVENT and retitle the command "Toggle Workspace Board". The action id stays workspace.openBoard to preserve users' stored keybinding overrides. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(keybindings): assert new toggle/open/close search keywords Cover the search-keyword additions from the toggle rename, per CodeRabbit review on #14240. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * style(keybindings): wrap workspace board search keywords for oxfmt --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> |
||
|
|
3ab8b6a117 |
fix(ssh): stop SSH reconnect from multiplying terminals and resuming agents twice (STA-3077) (#13326)
* fix(ssh): stop reconnect from grafting panes and stacking remote leases Reconnecting an SSH-backed workspace added terminal panes the user never opened, and the remote host accumulated shells nobody was using — one report went from 2 to 19 to 20 relay PTYs across three reconnects (STA-3077). Two root causes, both in the store. Reattach could create UI. `persistPtyBinding` has four creating branches — mint a tab, mint a root leaf, split the root and graft a leaf, mint a layout. All four are load-bearing for `pty:spawn`, which can beat the renderer's debounced layout writer, but none of them is appropriate on reattach, where the pane either already exists or is gone for good. Add `mayCreate`, defaulting true so the spawn path is untouched; every creating branch already sets `terminalMembershipChanged`, so refusing is a check rather than a new code path. Lease identity had no pane key. `upsertSshRemotePtyLease` matched on `(targetId, ptyId)` alone, so a pane that re-leased under a new relay id left its predecessor live with nothing to retire it, and the next reattach fanned out over both. One pane now keeps at most one live lease. Superseded leases are marked `expired` rather than terminated: losing a lease is not proof the shell died, so the remote process is deliberately left running. Tests assert observable behavior rather than mechanism, so they stay valid under any implementation that fixes this. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the terminal session behavior contract Properties stated as observable behavior rather than mechanism, so an oracle written against them survives a change of implementation. Records the weaker, correct form of the timer rule — a timer may never be the sole cause of a destructive action — because recovery budgets and scratch-file age gates are correct code that an absolute ban would condemn. Also notes which mechanisms are deliberately not required, so each has to earn its place rather than arrive with an architecture. Co-authored-by: Orca <help@stably.ai> * fix(ssh): heal duplicate pane leases that predate pane-keyed supersession Pane-keyed supersession stops new duplicates, but it does nothing for installs that already carry the ones STA-3077 accumulated — the report behind this reached 20 live leases across a handful of panes, and every reconnect fanned out over all of them. Retire the stale duplicates once per reattach pass, keeping the newest lease for each pane under a total order so two hosts resolve a tie the same way. As with supersession, retired leases are marked `expired` rather than terminated: their remote shells are deliberately left running, because a lease we chose not to revive is not evidence the shell died. The relay-session store stubs gain the new method. Note the gap this leaves open: those shells keep running and are no longer reachable from the app, so the "accumulates unused shells" half of the report needs a visible recovery surface rather than a silent kill. Co-authored-by: Orca <help@stably.ai> * fix(terminal): stop respawning a shell that is still running A pane that failed to reattach spawned a fresh shell. Because the restored session id came along, the replacement resumed the same agent session, and two processes appended to one transcript — reported repeatedly, up to five concurrent resumes of a single session. Two defects fed it. The relay reported a source that merely needed re-establishing as `SSH_SESSION_EXPIRED`. The shell was still running; only its output source was gone. Give that outcome its own error so it stops reading as "the session no longer exists". The reattach failure handler then treated every error as proof of death. It checked for expiry and, in the else branch, took the identical action — so the check bought nothing and a transport fault, a timed-out call, or a wedged relay all respawned. Respawn now requires proof: an explicit host expiry or a not-found PTY. Anything else, including an error we have never seen before, is unresolved, leaves the shell running, and keeps the binding for a later reattach. Two existing tests asserted the old behavior. One threw a bare error as scaffolding to reach the spawn-adoption door; it now throws proof, which is what it meant. The other pinned the expiry mapping itself, and now asserts the outcome fails closed *without* being reported as expiry. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record what makes a retention bound safe Shortening a grace period is the wrong lever. Measuring process time and gating reclamation on an independent observation are what make one safe, and they are what deployed systems actually do. Also records that lifecycle belongs in the attach reply rather than a delivered event — that is what removes the need for a durable per-consumer cursor to guarantee an exit is never lost. Co-authored-by: Orca <help@stably.ai> * test(terminal): assert the empty-failure case without an empty Error A thrown empty value exercises the same property — a failure carrying no usable message is not proof the session is gone — and does not trip the empty-error-message lint. Co-authored-by: Orca <help@stably.ai> * fix(ssh): let the durable pane binding outrank recency when retiring leases Choosing the newest lease for a pane is wrong whenever a newer lease exists that no pane is bound to: it retires the lease the pane is actually attached to, detaching a live terminal instead of healing it. Two changes. Arbitration now prefers the lease matching the pane's durable binding, across both the SSH-target and local partitions, falling back to recency only when no binding names either candidate. And supersession at upsert time now defers rather than expiring a bound predecessor. When a lease arrives for a pane that is still bound to a different PTY, the binding has not caught up yet, so both stay live and reattach arbitrates once the binding is available. Co-authored-by: Orca <help@stably.ai> * fix(ssh): roll back a lease retirement whose durable write fails `flush()` logs and swallows write errors, so a failed write left these leases retired in memory while disk still called them attached — and the pane bindings scrubbed alongside them stayed scrubbed. Use `flushOrThrow` and restore both the lease states and the affected session partitions when it throws, reporting nothing retired. Co-authored-by: Orca <help@stably.ai> * test(ssh): prove pane and remote PTY cardinality across reconnects Counts the shells the relay actually hosts, on the container, rather than inferring them from app state — that is the census the report was based on. Asserts the PIDs are unchanged, not merely the count, so a kill-and-respawn cannot pass. Every pane streams before the transport is severed: an idle pane sends no recovery checkpoint, so only a live source comes back needing re-establishment, which is the outcome that used to read as expiry. Co-authored-by: Orca <help@stably.ai> * fix(ssh): actually pass mayCreate:false from the reattach binding write The `mayCreate` guard was correct and had no production caller, so the reattach path still went through the creating branches and grafted panes back. `restoreReattachedPtyRuntime` is that call site — RC3 in the original diagnosis — and it now refuses to create. Binding moves ahead of runtime registration, because registering first would surface a pane the user never opened before the refusal landed. A refusal leaves the remote shell running and reattachable; a *thrown* write stays unknown and still registers, so a failed disk write cannot detach a live pane. Adds an oracle over the call site itself. The store-level tests all passed while the fix was inert, because they called the store directly — only pinning the wiring catches that. Co-authored-by: Orca <help@stably.ai> * fix(terminal): apply the respawn-requires-proof rule to both reattach paths connectPanePty has two near-verbatim reattach blocks — one keyed on the deferred SSH session, one on the restored session — and only the second was fixed. The first still checked for expiry and then respawned unconditionally anyway, so a transport fault there resumed the same agent session a second time. Also keep the wire token out of the pane. The main-process bridge only special-cases expiry, so a source-restore failure crossed IPC as raw `SSH_SOURCE_RESTORE_REQUIRED: <id>` text and surfaced to the user. It correctly does not respawn; it just should not read like that. Co-authored-by: Orca <help@stably.ai> * test(ssh): state plainly that the reconnect spec is a forward guard It was run against an unfixed tree and passed, so it does not prove the STA-3077 fixes and should not be read as if it does. A clean severed transport does not reproduce the field conditions — accumulated duplicate leases, or a source returning needing re-establishment. It keeps its place as a forward guard: it counts the shells the relay actually hosts and pins their PIDs, so a later change that grafts a pane or respawns a shell fails here. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record that a guard must be pinned at its call site A refusal that exists and is never passed is indistinguishable from no refusal, and store-level tests cannot tell the difference — they call the store directly. Learned from `mayCreate`, which was correct and had no production caller for several commits. Co-authored-by: Orca <help@stably.ai> * fix(ssh): park one PTY's exhausted delivery recovery instead of dropping the channel A per-PTY recovery budget running out disposed the whole relay channel, so one PTY that could not re-prove its delivery aborted every in-flight filesystem and git request on that host and stalled every sibling pane. A retry count is not proof of anything, and it certainly is not proof about the other sessions sharing the channel. Exhaustion now parks that PTY's delivery. The remote shell keeps running, its lease stands, and the next relay open reattaches it with a fresh delivery generation — the parked state is cleared on teardown and the generation changes on reconnect, so a reconnect recovers it. The consecutive-attempt ceiling goes away entirely; the per-generation one is what bounds the retry cost, and the second ceiling only existed to reach the channel drop sooner. Tradeoff worth stating: the failing pane used to self-heal within seconds because the forced reconnect wiped all rejection state, and it now stays frozen until the next relay open. That is a worse outcome for that one pane and a much better one for every other session on the host, and reconnecting is user-reachable. Co-authored-by: Orca <help@stably.ai> * fix(pty): let liveness say unknown instead of forcing it to say dead `IPtyProvider.hasPty` returned a boolean, so a provider whose inventory was empty for reasons that have nothing to do with the session — socket down, cache never hydrated, provider generation just constructed — had no way to say so and answered "absent". Its own siblings already knew better: `probePtyLiveness` and the runtime's `PtyController.hasPty` were both already `boolean | null`, with consumers branching on null correctly. The lie was injected at exactly one interface. Now three-valued, and each provider answers unknown where it cannot prove absence: the daemon adapter off-socket, the SSH provider before a completed listing, the router when any adapter cannot answer, and the degraded provider rather than fabricating a verdict. `terminal_gone` requires unanimous proven absence. Also fixes a real cold-start bug this surfaced: `pty:hasPty` never awaited the daemon-swap startup promise, though the sibling `probePtyLiveness` bridge already did, so before the swap the local provider answered an authoritative false for every daemon-owned id. Net +27 production lines. The plan behind this predicted -92 on the strength of deleting the renderer's dead-session reconcile path; that code is live (`pty-connection.ts` imports it), so nothing was deleted. Expressing a third value where there were two costs lines, and a deletion that is not real is not worth manufacturing. Co-authored-by: Orca <help@stably.ai> * docs(terminal): track the terminal-session correctness handoff package The package was untracked under a gitignored `docs/**`, with the un-ignore rules living only in an uncommitted .gitignore edit — a single `git clean -xdf` would have destroyed the authoritative plan. The 814-path construction snapshot is now pushed as `nwparker/react185-authority-snapshot` too; it had no remote ref. Co-authored-by: Orca <help@stably.ai> * test(ssh): make the reconnect settle window actually wait The settle poll reused a matcher the assertion 15 lines above had already satisfied, and Playwright's poll engine probes immediately and returns as soon as the matcher passes — so it observed the same state twice and elapsed 0ms. A shell grafted a second or two after reattach reported ready slipped through into the next cycle. Reviewer was right on #13111. Test-only; no production change. Co-authored-by: Orca <help@stably.ai> * test(ssh): census both durable session partitions on reconnect Adds a second reconnect scenario and a helper that reads pane records from the local partition as well as the ssh host partition. That split matters: the reattach binding call passes no hostId, so a grafted pane lands in the LOCAL partition and an oracle reading only the host partition passes whether or not the guard is present. Both tests remain forward guards. The second one was reported as discriminating and did not reproduce: with `mayCreate: false` removed from the call site and the app rebuilt, both still passed. Its induction races `pty:kill` against a severed transport, so when the kill lands the lease is cleaned up and there is nothing left to graft. The handoff README is corrected to say so rather than claim a journey. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the user decision relaxing G6 G6 becomes minimise-and-justify rather than strictly net-negative. The deletion budget the plan assumed does not exist: an entrypoint-rooted import graph found 51 of 53 candidate files reachable and instantiated on live paths, leaving 263 deletable LOC against roughly +1,021 to offset. Correctness may still not be traded for line count. Co-authored-by: Orca <help@stably.ai> * test(terminal): add discriminating oracles for restart, daemon, skew and namespaces Six parallel streams, each required to fail with its guard removed rather than merely pass. Local restart proves the OS process itself survives, by reading `ps -o lstart=` for the shell's own pid. That matters: with the quit path made destructive, the tab, leaf and pty ids all came back byte-identical while the shell underneath was a new process — every existing restart spec would have stayed green. Two separate guards were removed to redden it, and the second reddens only the stale-operation case. Daemon restart discriminates by reverting three-valued `hasPty`; version skew now covers publication semantics and confirms the new `SSH_SOURCE_RESTORE_REQUIRED` token mutates nothing on an old client; two-host isolation censuses both containers. Deletes `src/relay/pty-source-replay-index.ts` — 201 production lines with no importer outside its own test, verified against an entrypoint-rooted import graph rather than a name grep. Five namespace tests are skipped, not passing: they reproduce a defect still live on main where folder-workspace ids compare equal with the instance suffix stripped. PR #12474 fixes it; they are its oracle. Co-authored-by: Orca <help@stably.ai> * test(ssh): induce the reattach graft deterministically instead of racing a kill The previous induction closed a pane while the transport was severed and relied on `pty:kill` FAILING so the lease outlived the pane record. It does not fail: with the provider already torn down, `pty:kill` takes its tombstone branch and marks the lease terminated, and `reattachKnownPtys` filters terminated leases out of the fan-out — so the reconnect never visited the PTY the test was about. It passed on both trees. Seed the precondition instead. Spawn a real remote PTY on a leaf that never becomes a pane, then roll the host partition back to its pre-spawn snapshot, leaving a live lease and a live remote shell that no durable pane owns. No failure races a success. Adds a vacuity guard that is independent of the tree under test: the lease's own `lastAttachedAt` must advance, proving the fan-out actually visited this lease before the pane census is trusted. Verified on this machine under an isolated TMPDIR, since the e2e harness keys its seeded-repo pointer on a machine-global tmpdir path: guard present passes, guard removed fails with the phantom leaf grafted into the local partition, guard restored passes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): propose one authoritative binding identity Every defect this program has touched is the same defect: identity compared with the wrong key, or not compared at all. Lease keyed without the pane, reattach using a creating write, folder-workspace ids compared with the instance suffix stripped, local mutating IPC carrying only an id, a live shell classified as expired, liveness unable to say unknown. Proposal: one branded binding type built from fields that already exist and are already persisted, constructible only from an authoritative source, carried by mutating operations, compared by one shared function. Makes a wrong-key comparison a type error rather than the next incident. Under adversarial review, including against the open issue corpus. Not accepted. Co-authored-by: Orca <help@stably.ai> * fix(pty): refuse mutating operations aimed at a superseded PTY `pty:write`, `pty:writeAccepted` and `pty:resize` accepted any id. The renderer queues input, so a keystroke buffered before a reattach landed on whatever PTY had since taken the pane — and a resize reshaped the successor's shell. Main already tracks `ptyPaneKey` and `paneKeyPtyId` in lock-step, so their disagreement is proof the caller's id was superseded. No wire change, no renderer change, nothing added to the input payload. An id with no recorded pane stays permitted: unowned and orphaned PTYs are unknown, not stale, and unknown never authorizes refusing an explicit operation. That is also what keeps orphan cleanup working — those ids have no pane by construction. The tests pin the CALL SITES, not the predicate. A capability that exists and is never called is indistinguishable from no capability, which is exactly how `mayCreate` sat inert here for several commits with every test green. Co-authored-by: Orca <help@stably.ai> * fix(pty): fence signals at a superseded PTY, and pin why kill is exempt A signal means "interrupt my pane", so delivering one to a PTY the pane has already replaced is a misdirected interrupt. Fence it with the same lock-step proof used for write and resize. `pty:kill` stays deliberately unfenced and a test now pins that: a superseded PTY is orphaned, and reclaiming it is exactly what the orphan-cleanup callers ask for. Refusing there would break the operation that reclaims leaked shells — the opposite of the intent. The fence sits at the IPC boundary, above `tryGetProviderForPty`, so it covers local, daemon and SSH rather than the local path alone. Co-authored-by: Orca <help@stably.ai> * test(terminal): poll the pane binding read so a slower host cannot flake it `readPaneBinding` took a single unpolled read of a DOM dataset attribute immediately after a renderer reload, while its sibling helper polls the same data for 15s. On a native Linux host both tests failed every run with 'No bound terminal pane is mounted' while the app was demonstrably healthy — the screenshot showed the terminal restored with a live prompt and the boot PID echoed. The assertion is unchanged; it is only awaited. Nothing is weakened. Found by running this spec on native Linux rather than assuming macOS behaviour generalises. Co-authored-by: Orca <help@stably.ai> * test(terminal): make the restart identity spec run on Windows too Both probes were POSIX-only and unconditional: `echo ...=\$\$` for the shell's own pid, and `ps -o lstart=` for its start time. Running the spec on a real Windows host proved it dies before reaching either guard, so Journey 1's Windows half was unprovable rather than merely unproven. PowerShell exposes the same two facts as `$PID` and `Get-Process` StartTime. The start time still matters on both platforms for the same reason: a PID alone cannot separate a survivor from a reused number. Still green on macOS. The Windows path is written from the host probe and has not itself been executed end to end — that is the next thing to run there, not a claim being made here. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the fence's real gap and what peer designs taught Marks the client-constructed binding proposal as rejected with the three false claims that sank it, and records what shipped instead. States the shipped fence's actual limitation rather than leaving it implied: it compares a binding, not an incarnation, so a respawn under a reused ptyId passes. The obvious remedy is wrong here — the agent-create id is deterministic by design so a replayed create stays idempotent, and randomising it would trade this narrow gap for a duplicate-spawn bug. Also records the ranked lessons from four comparable agent IDEs, chiefly that a typed end-reason at end time is what stops a user quit from looking like a resume candidate. Co-authored-by: Orca <help@stably.ai> * docs(terminal): promote Journey 1 to proven on all three platforms The oracle now runs natively on macOS, Linux and Windows, and its discrimination was watched on each: a mutation reddens it, a restore greens it. On Linux and Windows both mutations were run, and the second reddens only the stale-operation test — so the journey's two clauses are proved independently rather than jointly. Windows is the new evidence. The PowerShell branches added blind at ebffb85a848 executed correctly on their first run: `$PID` expanded to real integers, which also proves the pane shell there is PowerShell-family rather than Git Bash, and `Get-Process StartTime` returned kernel start times 5.4s apart — so a recycled pid could not have passed as a survivor. First journey promoted in this program. The other twelve are unchanged, and the residual limit on "every stale exact operation" is recorded rather than glossed. Co-authored-by: Orca <help@stably.ai> * test(terminal): add discriminating oracles for the daemon, skew and multi-host journeys Daemon: replaces a spec that modelled only a client restart and never crossed the daemon boundary, whose successor generation owned nothing so "the live successor is neither killed nor replaced" was vacuous. The PTY leader is now a real login shell reporting `$$` back through the production write path, resolved to a kernel start time. Two mutations each redden exactly one of the three clauses, on macOS and Linux: reverting three-valued `hasPty` reddens only the unknown-not-dead clause; widening the sole-provider fallback reddens only the stale generation clause. Skew: reverting the restore-required publication to expiry reddens 4 of 5 new tests while the legacy control stays green — the regression this branch fixed is now caught if reintroduced. Multi-host: restoring `mux.dispose('connection_lost')` reddens sibling isolation on one host. It does NOT redden across hosts, and that is recorded rather than glossed: a mux belongs to one relay session per target, so its dispose cannot cross a host boundary. Journey 4's cross-host clause rests on isolation-by-construction, not on a mutation. No production code changes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record journey evidence that falls short of promotion Four journeys now have discriminating oracles but none meets its full stated scope, and each shortfall is named rather than rounded up. Journey 2 is one WSL run from promotion. Journey 12's tests are in-process, so they do not close the live-skew gap the original ledger named. Journey 4's cross-host clause cannot be proven by mutation at all — a mux is per target, so its dispose cannot cross hosts, and the cross-host test stayed green under the mutation that reddens siblings. Journey 13 measured one dimension of ten, on lifted predicates rather than through real IPC. Co-authored-by: Orca <help@stably.ai> * docs(terminal): promote Journey 2 to proven on macOS, Linux and physical WSL The oracle runs on every environment the journey names, and is clause-selective on all three: reverting three-valued `hasPty` reddens only the unknown-not-dead clause, and widening the sole-provider fallback reddens only the stale-generation clause. Selectivity in WSL was established rather than assumed. The spec runs serially, so a red first test reports the others as "did not run" — they were re-run alone under the same mutation and stayed green. Also records that an Orca WSL-mode terminal now starts on that host at all, which it could not before: the distro had no provisioned default Unix user, so every interactive launch blocked on first-run setup. One diagnosis from the WSL run is corrected here rather than repeated: the unrelated `local-pty-shell-ready` failure was attributed to bash 5.3.9, but macOS runs the same bash version and passes 67/67. The trigger is environmental to that distro, and the underlying defect is that the spec pins an absolute count of OSC markers it does not own. Co-authored-by: Orca <help@stably.ai> * docs(terminal): correct the WSL provider-suite diagnosis The WSL run blamed bash 5.3.9 for the unrelated `local-pty-shell-ready` failure. macOS runs the same bash version and passes 67/67, so the version is not the cause — the trigger is environmental to that distro, and the underlying defect is that the spec asserts an absolute count of OSC markers it does not own. Co-authored-by: Orca <help@stably.ai> * test(runtime): unskip the workspace-namespace oracles now their fix has merged These five reproduced a defect that was live on main: folder-workspace ids were compared with the instance suffix stripped, so two workspaces sharing a directory read as the same namespace. They were committed skipped, pointing at the PR that fixes it. That PR is merged, and they pass. Verified they still bite: restoring the suffix-stripping comparison reddens exactly these five and leaves the other four green. An oracle written before its fix, held skipped, and confirmed against the fix after the merge — rather than deleted and rewritten from the answer. Co-authored-by: Orca <help@stably.ai> * test(ssh): add MaxSessions, lazy-discovery and paired-skew oracles Three journeys attempted; none promoted, and the reasons are recorded in the ledger rather than rounded up. MaxSessions=1 against real OpenSSH, with the cap read back from `sshd -T` rather than assumed, and remote pids read on the container two independent ways that must agree, each carrying its kernel start time. Two disjoint mutations discriminate — one reddens only the reconnect clause, the other only the two restart clauses. But the disconnect clause is a forward guard: four separate guard removals left it green, so nothing shipped is load-bearing for it. Lazy discovery samples sshd's own accept log and live session census across a 22s window with the in-use host as a positive control. No mutation reddens its third clause alone — the real cross-host lease scoping is load-bearing, but removing it breaks the sibling host during setup, so the failure carries no clause information. The paired-runtime skew spec pairs two real processes at different versions and refuses to run rather than degrade into a same-version pairing that would look green and prove nothing. No production code changes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record why the duplicate-resume fix was not built I recommended adding a typed end-reason so a user quit stops looking like a resume candidate, then went to implement it and stopped. `SleepingAgentSessionRecord` already carries three fields that each exist to stop something resuming that should not have — `origin`, `restoreOnTabOpenOnly`, and `automaticResumeBlockedBy` — each traceable to its own incident, consulted at 22 non-test sites. A fourth predicate, however well typed, is the fifth containment cycle. The designs without this bug do not have a better flag; they resume only on an explicit action, into a new terminal id, and make two agents in one terminal unrepresentable in the schema. The first of those is a product decision about whether automatic resume stays a feature, so it is the user's call rather than mine. Co-authored-by: Orca <help@stably.ai> * docs(terminal): reconcile G6 with the recorded decision and assess its clauses G6's body still demanded strictly-negative production LOC after the user relaxed it to minimise-and-justify, so the gate had two conflicting pass conditions and no single truth value. Its body now points at that decision. Assessed the remaining clauses against the branch rather than assuming. Two fail structurally: more than one identity comparison and mutation admission path still exist, and `terminal-input-quarantine.ts` is still reachable from two production files. Records why the quarantine is not subsumed by the superseded-PTY fence, which I had assumed and checked. The fence refuses writes aimed at a stale ptyId; the quarantine guards the user's next keystrokes landing on the successor under its current, correct id — a case the fence never sees. Removing it needs the recovery path to surface a different shell as unresolved, not a deletion. Co-authored-by: Orca <help@stably.ai> * docs(terminal): the input quarantine is load-bearing, not superseded G6 lists "no superseded quarantine remains reachable" and this module was assumed to be one. Disabling its single call site reproduces the hazard it exists for — `cho hi; rm -rf x` reaching the shell — so deleting it without a replacement re-opens command execution. The replacement was costed by building it rather than estimated: +26 production LOC to thread the incarnation, ~+33 complete, and the cross-remount state it needs outlives the destroyed pane so it becomes a module about the size of the one deleted. Floor is roughly +140 to delete 88, and it would add a second identity comparison to a gate already failing for having more than one. The decisive part is that the route is not uniformly available: remote runtime results carry no incarnation, old hosts cannot be made to publish one, and mixed versions are the normal state. A paired client reads unknown, which this program's own rule says is not proof — so either every remote reattach surfaces unresolved, or a fallback is needed and the only correct fallback is this module. Whether to amend the clause or accept something weaker on remote hosts is a user decision, so the clause verdict is left as failing rather than quietly reclassified. Co-authored-by: Orca <help@stably.ai> * refactor(runtime): collapse duplicate identity comparisons G6 requires one identity comparison; five implementations existed across two concepts. Worktree-namespace identity had two: `runtimeWorktreeIdsEqual` and `runtimeWorktreeIdentityKey` independently re-derived repoId plus normalized path. Equality now derives from the key, so the comparison and the sleep / mutation-queue keying cannot drift into two different rules — which is exactly how the suffix-stripping bug reached production once. Pane identity had three byte-identical leaf-UUID comparisons, in orchestration `db.ts`, `lifecycle-reconciliation.ts`, and `orchestration-legacy-process-identity.ts`. One copy moved to `stable-pane-id.ts`, which already owns `PaneKey`, `parsePaneKey` and `makePaneKey` and which all three already imported. No new module, no branded type, no parallel comparison. Net -14 production lines. The namespace oracle still bites: restoring the filesystem parser inside the identity key reddens exactly its five cases. The raw counts are not the actionable set, and the classification is worth recording: of 409 non-test `worktreeId` comparisons, 71 are typeof guards and 81 are sentinel tag checks. Most of the remainder are renderer predicates over store rows where both operands are the same main-minted id, so normalizing there would widen equality rather than correct it. Co-authored-by: Orca <help@stably.ai> * refactor(terminal): finish a half-done fixture move and audit the rest `xterm-bypass-event-fixture.ts` and `__fixtures__/xterm-bypass-event.ts` were byte-identical apart from an import path. The `__fixtures__` copy had zero importers and the live copy compiled as production — someone started the move and left both. Dead copy deleted, live one moved, its three test importers updated. Audited the wider G6 clause by importer rather than filename: 32 test-only files, roughly 3,300 LOC, currently compile as production; 4 of the 36 candidates have real production importers and are correctly placed. The list is recorded in the goalposts. Those 32 are almost all older than this program and outside the terminal surface, so sweeping them belongs in its own change rather than inside a terminal PR. The clause stays failing, with the remaining files named. Co-authored-by: Orca <help@stably.ai> * docs(terminal): the fixture clause already holds where it matters Checked what the build emits rather than reasoning from file paths. None of the 32 test-only fixtures appears in `out/` — Rollup drops them because no production entrypoint reaches them. On "compiles into the shipped product", this clause holds today. On the other reading it cannot be closed by moving files at all: both production tsconfigs use bare `include` globs with no `exclude`, so a `__tests__/` directory matches exactly like any other path, as does every `*.test.ts` in the repo. Relocating 32 fixtures would remove nothing from typecheck scope. A sweep was started and stopped once this was verified, rather than landing 32 moves across areas this program does not own for no gain. If the intent is that typecheck scope should exclude test code, that is a repo-wide tsconfig change with a different owner. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add plain-language design and test overviews Two reviewable documents with diagrams, written so someone with no prior context can follow what breaks, why, and what changed. The design overview explains the five things stacked behind one terminal rectangle, the 2 -> 19 -> 20 report, the three root causes, and the rule underneath all of them: unknown is not dead. The test overview explains why a green test proves nothing on its own, the four-step mutation proof we adopted, and — the part worth reviewing hardest — an honest account of what could not be proven and why, including the properties that are true by construction and therefore have no guard to remove. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add a self-contained visual report of the design and its evidence Pre-renders every diagram to inline SVG in both themes so the report opens offline and stays sharp when zoomed. States the gate/journey score and the retractions alongside the fixes, so the unproven half is as visible as the proven half. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the finalized two-plane architecture decision Adopts the data-plane proposal and adds the control-plane track it does not cover: re-key ownership by pane, split orphan inventory out, then delete the compensating code. Records that the host-authority alternative was refuted and that the shipped keystroke fence is inert on the reattach path. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add the design brief the review counsel works from Separates verified code facts from unverified leads so reviewers attack the design rather than a reconstruction of it, and records which simpler alternatives were already refuted and why. Co-authored-by: Orca <help@stably.ai> * docs(terminal): report the design counsel's outcome and the live respawn bug it found Three review rounds across two models replaced the two-record split with one leaf-keyed record, deleted attach-time pane identity, and made orphans a connect-time projection. Records that a shipped gesture still turns a healthy remote shell into a duplicate agent resume, and that the renderer classifier in that chain treats an error-message shape as proof of death. Co-authored-by: Orca <help@stably.ai> * docs(terminal): correct the report — the respawn proof gate guards a minority path A final review traced every auto-respawn route. The primary one converts the reattach failure into a boolean before any classifier sees it, so the shipped proof gate never runs there. Records that two of the six shipped changes are narrower than claimed, and why their tests could not have caught it. Co-authored-by: Orca <help@stably.ai> * docs(terminal): explain the landed design on its own terms One leaf-keyed ownership record, orphans computed at connect, and replacement shells only on positive proof — with the shipping order and the one product trade the design asks the owner to accept. Co-authored-by: Orca <help@stably.ai> * docs(terminal): rewrite the design explainer in plain English The first version assumed the reader knew the codebase. Reframed around two bugs, two fixes and one decision, with the jargon replaced by pane / program / note / helper and a five-word glossary for what could not be avoided. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop reading an identity mismatch as a dead shell The relay reports a pane-identity mismatch by saying the pty was not found, but it found it — comparing identity is how it noticed. Publishing that as expiry made the renderer clear the binding and cold-restore with agent resume, so a live shell gained a second agent on one transcript. Reachable today by detaching a pane into a new tab, which changes the tab the relay froze at spawn. Mismatch now carries its own token and the classifier refuses it as proof. Genuine absence still expires, so a shell that really went away is not stranded. The three failure tokens move to src/shared: main published them and the renderer decided respawn on them, from two copies that had drifted apart. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop sending pane identity on reattach The relay froze pane identity at spawn, so moving a pane to another tab made it refuse a live shell — and refuse by saying 'not found'. The comparison is presence-guarded, so not sending the fields disarms it on every relay version including ones already installed on hosts: no wire change, no redeploy. Nothing is lost. It existed to catch a relay restart recycling pty-N for a new shell, and in exactly that case pane and tab both still match, so it accepted the wrong shell anyway. The incarnation the attach returns is what distinguishes those, and it already crosses the wire. Removes the whole client-side apparatus: the expected-identity type, its per-lease derivation, its map, and the parameter threaded through four layers. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add tracked goalposts for the new design Each goalpost is a behaviour with an oracle and the mutation that must redden it, so 'proven' cannot be claimed from a green test. Records the anti-inert rule as a first-class goalpost, since three guards in this program passed their tests while sitting off the route production takes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record that the recovery grant is dead code, deleting a design step The lease stores a relay-native pty id and the caller passes the app form, with a raw equality comparison between them, so the 30s grant cannot fire for a real SSH pane. The death rule that existed to referee it is deleted rather than built, and the dead path itself becomes a removal. Co-authored-by: Orca <help@stably.ai> * docs(terminal): keep the full design detail in the repo It only existed in an ephemeral job directory, so the plain-English explainer had no durable source for its specifics — record shape, death rule, reattach algorithm, migration order and the 25 oracles. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add a resume prompt for a clean session Points at the goalposts as the contract, names the three goalposts whose oracles are already written and red, and carries the process rules that were learned the expensive way — prove guards reachable, verify mutations land, commit per step, and never let a subagent write production files in a shared worktree. Co-authored-by: Orca <help@stably.ai> * test(ssh): add the failing oracles for goalposts S3, S4 and S5 Intentionally RED: 14 clauses that fail against current behaviour and go green under the changes named in new-design-goalposts.md. The branch is held unmerged, so red here means unimplemented, not broken. Each was verified to fail for the right reason and to flip green under the identified fix, which was then reverted. Each pins the producer as well as the consumer, so no clause can pass vacuously if its route is ever severed — the failure mode that let three earlier guards ship inert. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop fabricating an exit when a reattach fails A failed attach never proves the shell exited. The relay answers not-found for a pane-identity mismatch and for any id it merely cannot hand back, so treating it as death sent the pane a synthetic `pty:exit { code: -1 }`, cleared provider state, deleted ownership and expired the lease — four claims about a process we know nothing about, on a shell that is usually still running. Collapse every failure into the non-destructive branch that already existed a few lines above (`restoreRequired = 'reattachAttemptsExhausted'` + wakeRecovery). A branch collapse, not a new mechanism: goalpost S3. Two tests pinned the deleted premise and are INVERTED rather than patched, so the new intent stays covered: - ssh-relay-orphan-abandon-paths: "retires the lease without a kill when the relay proves the PTY is gone" -> "leaves the shell running when the relay only reports the PTY as not found". Its comment claimed attach verifies liveness before answering not-found; it does not. - ssh-relay-session: "invalidates and broadcasts remote PTYs that cannot reattach" -> "leaves an unreattachable remote PTY alone while its sibling reattaches". Also repairs two clauses left red by |
||
|
|
ede69ffc7f |
perf(skills): bound and share skill discovery scans (#14204)
Skill discovery re-walked every skill root on every window focus, pane mount, and connected client. The root set was already bounded; what was not bounded was how often and how redundantly it was walked. - Focus called refresh(true), bypassing every cache down to a disk walk. - The process that owns the disk had no cache and no in-flight dedup. - Panes with different cwds each re-walked the same 12 home roots. - Fan-out inside a scan was unbounded, and every package was walked twice (once to find SKILL.md, once to count its files, node_modules included). Adds one coalescing primitive — in-flight dedup plus a short TTL behind a bounded LRU — used for per-target dedup below both the IPC and RPC entry points, per-root sharing on the native path, and whole-result reuse on the WSL path. A scan may publish only while it still owns its pending slot, so a scan begun before an invalidation can never re-cache a pre-mutation result. Bounds per-skill fan-out to the existing candidate concurrency limit, and bounds the package file walk by depth with a node_modules prune. Focus now reads through a 15s freshness window; explicit signals (install completed, Settings Refresh, native-chat Retry, terminal exit) set a new optional `refresh` wire field that bypasses every cache, including on remote runtimes. Measured on a 32-concurrent-scan burst across 8 workspaces: 134,880 -> 2,956 filesystem calls and 1080ms -> 45ms, same 31 skills returned. |
||
|
|
0f51d0b3bb |
Fix native chat image marker position handling (#14162)
* fix(native-chat): handle image markers in any position * fix(native-chat): preserve image caption whitespace * test(native-chat): cover marker boundary spacing * fix(mobile): normalize image echo reconciliation * fix(mobile): use idiomatic tail access * refactor(native-chat): share image echo matching * perf(native-chat): avoid unchanged block copies |
||
|
|
0ed6db77cf |
fix(mobile): open agent-cited external chat files (#14166)
* fix(mobile): open agent-cited external chat files * fix(mobile): keep cited external files read-only * refactor(mobile): derive cited-file mode from provenance * fix(mobile): accept sentence-final cited paths * fix(mobile): preserve cited SSH grant scope * refactor(file-links): share location suffix parsing |
||
|
|
af7dcdc196 |
feat(dashboard): identify SSH and remote hosts (#14177)
* feat(dashboard): identify SSH and remote hosts * fix(dashboard): resolve host labels consistently * fix(dashboard): reuse host server icon * test(dashboard): guard host label refresh cost * test(dashboard): satisfy runtime environment shape |
||
|
|
7c93aed6dc |
Fix fsync of read-only files on POSIX (#14235)
* fix(files): fsync read-only files on POSIX * test(e2e): add golden E2E tests for POSIX profile index fsync Validates that profile index files are properly persisted on POSIX systems, including with restrictive umask settings. These are release-blocking golden tests for Linux and macOS. * test(terminal): wait for fish child ownership before stdin write Fish 4.8 withdraws DECSET 2031 before spawning the child, so the shell-contracts harness could send hello into an intermediate prompt and hang waiting for CHILD-READ. Wait for the child's CHILD-READY marker and answer split DA1/CPR/OSC queries across chunk boundaries. * test(e2e): verify profile index persists to disk with restrictive umask Strengthen the POSIX fsync test to verify the rebuilt index is actually written to disk and has correct permissions under a restrictive umask, not just cached in memory. |
||
|
|
a73f122c61 |
fix(store): keep project catalog identity when repo.addedAt is 0
* fix(store): keep project catalog identity when repo.addedAt is 0 projectHostSetupProjectionFromRepos used `repo.addedAt || now`, so a missing or zero addedAt stamped Date.now() into createdAt/updatedAt on every refresh. #13803's reconcile then treated the project as changed and never reused it. Use a finite check with a stable 0 fallback, and treat 0 as unknown in merge so a persisted createdAt is not wiped. Refresh-identity tests use production-shaped nested fixtures plus structuredClone and go red if the fallback is reverted. Co-authored-by: Orca <help@stably.ai> * type(shared): return readonly setups from getProjectHostSetupsForProject The helper already takes a readonly catalog and returns a filter subset. Mark the return readonly so callers cannot mutate a live setups array. Co-authored-by: Orca <help@stably.ai> * style: oxfmt projection files and drop stale addedAt comment oxfmt --check failed on the addedAt identity commit. Also stop claiming the projection still restamps Date.now() when addedAt is 0. Co-authored-by: Orca <help@stably.ai> * fix(store): treat createdAt 0 as unknown when merging sibling repos Repo order decided a project's createdAt: a zero-addedAt repo seeded the accumulator with 0, and the merge only treated the *incoming* addedAt as unknown, so min(0, 100) kept 0 when the unknown sibling came first. Share unknown-aware mergeCatalogCreatedAt/mergeCatalogUpdatedAt helpers and apply them on both sides of the projection merge and of the renderer's cross-host mergeProjectCompatibilityProject, which had the same 0-poisoning via Math.min(base.createdAt, overlay.createdAt). Co-authored-by: Orca <help@stably.ai> * test(store): pass a valid updateProject payload in createdAt merge case updateProject only accepts localWindowsRuntimePreference. The new unknown-vs-known createdAt test used displayName and failed typecheck. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
d41cb21e94 |
Sta 4062 folder note rollback (#14232)
* fix(persistence): keep folder-workspace notes across a build rollback normalizeFolderWorkspaces rebuilds each FolderWorkspace field-by-field, so the inline diffComments field #14112 added is dropped by any build that predates it — and the next full-state write makes the loss durable with no user edit. Move the on-disk home to an optional top-level PersistedState.folderWorkspaceDiffComments, which older builds round-trip untouched through their {...defaults, ...parsed} load spread and omit-style getDurableState(). load() hydrates it onto the records and deletes it from Store state; buildStateToSave() is the only producer. The in-memory FolderWorkspace shape, and therefore every IPC/RPC/renderer/mobile path, is unchanged. Co-authored-by: Orca <help@stably.ai> * fix(persistence): prefer inline folder notes over a stale map entry Hydrate preferred a non-empty folderWorkspaceDiffComments entry over non-empty inline notes. A rollback to a notes-capable #14112 build writes notes inline and leaves the older map untouched, so re-upgrading deleted everything authored while rolled back. Inline now wins when present; the map only fills a stripped record. Co-authored-by: Orca <help@stably.ai> * Extract folder workspace diff comments to dedicated module Moves normalizeFolderWorkspaceDiffComments and collectFolderWorkspaceDiffComments from persistence.ts to a new folder-workspace-diff-comments.ts module for better code organization. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
c61f29639c |
fix(workspaces): recognize pasted Linear issue URLs (#14190)
* fix(workspaces): recognize pasted Linear issue URLs * fix(workspaces): resolve legacy Linear URLs and keep typed-name fallback Saved API-key Linear workspaces often omit organizationUrlKey, so pasted issue URLs never called fetchLinearIssue. Probe those unknown-org workspaces and accept only a matching issue URL. Keep "Use … as workspace name" visible in Smart Entry when a Linear URL owns the results. sourceIntent still focuses the issue row. Co-authored-by: Orca <help@stably.ai> * fix(workspaces): jump to existing worktrees from pasted task URLs Cmd+J now treats GitHub, GitLab, Jira, and Linear issue/PR URLs as decisive search, so pasting one lists already-linked worktrees first and keeps a create preview underneath. GitHub/GitLab/Jira still hand the raw URL to the composer for cross-project detection. Co-authored-by: Orca <help@stably.ai> * fix(workspaces): resolve GitHub issue titles in Cmd+J URL paste Pasted GitHub issue/PR URLs now fetch the title for the create preview, same as Linear. Existing linked worktrees stay selected first so Enter jumps; create still hands the raw URL to the composer. Co-authored-by: Orca <help@stably.ai> * fix(workspaces): attach resolved GitHub items from Cmd+J create Pasting a GitHub issue/PR URL into Cmd+J already previewed the title, but Enter still opened the composer with the raw URL. Await the in-flight lookup and hand the linked work item through, matching Linear and Task page create. Also fix CI type, lint, and focus-routing source checks. Co-authored-by: Orca <help@stably.ai> * fix(workspaces): add Cmd+J task-URL locale keys Unblock PR CI localization catalog checks and make the Linear lookup-miss e2e wait out the resolving state before advancing to the agent field. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
585dd6d3a9 |
fix terminal attribution shim removal edge cases (#14187)
* fix(terminal): fully retire attribution shim * fix(terminal): harden shim tombstone path lookup |
||
|
|
58a926170c |
perf(renderer): index the repo catalog compat merge and keep catalog identity (#13803)
* perf(renderer): index project host ownership on repo catalog refresh mergeFetchedProjectCompatibilityForHost runs synchronously inside a zustand set() on every repo catalog refresh and was O(projects x (setups + repos)): getProjectHostIds / getExplicitProjectHostIds rescanned all setups and all repos once per project, and mergePreviousProjectMetadata rebuilt a full catalog repo key-set per project. Precompute the indexes once and hand each helper only its own project's slice. No ownership logic is reimplemented — the same resolvers run, just on pre-sliced inputs, memoized per project object. 200 repos 1.1 ms -> 0.2 ms (5.2x) 600 repos 11.2 ms -> 0.7 ms (15.3x) 1200 repos 43.0 ms -> 1.6 ms (26.6x) Verified against 4,000 randomized differential cases (duplicate repo ids, dangling sourceRepoIds, empty-repoId and orphan setups, dropped setups, local/SSH/runtime host mixes) comparing output element provenance and order against the previous implementation, not just deep structure. Each of the three index helpers was negative-controlled: breaking any one of them turns the differential red. Co-authored-by: Orca <help@stably.ai> * perf(renderer): keep the repo filter array identity across no-op refetches Three fetch sites unconditionally reallocated filterRepoIds (`s.filterRepoIds.filter(...)`) on every repo catalog refresh, even when nothing was pruned. Six identity-sensitive subscribers select this array, including App.tsx at the root, so every refresh woke the whole tree. retainValidFilterRepoIds returns the input when every id is still valid; `.every` short-circuits so the common case allocates nothing at all. It lives in its own module rather than ui.ts: repos.ts does not import ui.ts and this codebase has a documented circular-slice-import hazard. The store type widens to `readonly string[]`, which takes three more declaration edits (visible-worktrees, add-repo-skip-finalization, the setter). PersistedUI.filterRepoIds deliberately stays `string[]`: main owns that array, and widening it would fail the ui-state schema-parity assertion unless the zod schema gained `.readonly()`, which Object.freeze()s every parsed `ui.set` payload arriving from a paired client. Copy at the App.tsx IPC boundary instead — one allocation per debounced 150ms persist, not per render. Co-authored-by: Orca <help@stably.ai> * perf(renderer): reconcile projects and host setups across no-op refetches mergeFetchedProjectCompatibilityForHost always allocates — sourceRepoIds is rebuilt per project and fetched setups arrive freshly cloned over IPC — so `projects` and `projectHostSetups` lost both array and element identity on every catalog refresh. #13770's identity work covered only projectGroups and folderWorkspaces. Reconcile both at the merge's single return site (five callers) against `previous`, keyed on what each merge already dedups by: project.id, and getProjectHostSetupOwnerKey for setups. setup.id is deliberately not the setup key — the repo-derived fallback sets id = repo.id, so one repo on two hosts yields two setups sharing an id, and keying on it would splice the wrong host's routing metadata into a row. filterSetupsForPrunedRepoRows had to stop returning `[...setups]` unconditionally: three callers feed its result in as `previous`, so the throwaway copy defeated the reconcile even though element identity survived. Rather than add a third parallel deep-equality helper, generalize reconcileFetchedRepos (#13744) into reconcileCatalogRows and delete repos.ts's isPlainCatalogObject + areCatalogEntriesEqual (#13770). The two were the same algorithm modulo the null-prototype branch; the survivor takes the union of both prototype guards, which is strictly more permissive and so can only reconcile more, never mask a change. repos-catalog-merge-identity.test.ts and repo-identity-reconcile.test.ts stay green unmodified as the proof. Element-level reuse means a consumer mutating a Project or ProjectHostSetup in place would corrupt the previous render's object, so RepoSlice['projects'], ['projectHostSetups'] and ProjectHostSetupProjection widen to readonly. That is the safety mechanism, not polish — it caught two real in-place accumulators (ai-vault-scope-paths and the runtime-host purge in worktrees.ts), which now annotate mutable locals per the repo's readonly precedent. No casts added. Co-authored-by: Orca <help@stably.ai> * test(renderer): cover repo catalog refresh identity A dedicated file — repos.test.ts is already at the max-lines ceiling and these are a distinct concern. The fixture is production-shaped on purpose. A scalar-only repo reconciles even when the structural compare is broken, which is how an earlier version of this work shipped green while being completely inert; addedAt is pinned non-zero because project-host-setup-projection falls back to `repo.addedAt || now`, so a zero timestamp stamps Date.now() into every projection and nothing ever reconciles. Every mocked fetch returns a structuredClone so identity can never match by accident. Every identity assertion was verified red against reverted source: - drop reconcileCatalogRows from the merge return -> no-op-refetch identity and per-element reuse fail - filterSetupsForPrunedRepoRows back to `[...setups]` -> no-op-refetch identity fails - retainValidFilterRepoIds back to `.filter(...)` -> both filter-identity cases fail - areValuesEqual reduced to `a === b` -> all three reconcile cases fail - areValuesEqual forced to `true` -> all six change-propagation cases fail - setup key switched to setup.id -> the two-hosts-one-repo-id case fails Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
63a0462b1b |
Stop deriving interactive prompt for OMP ask tool
Why: OMP ask maps to blocked sidebar state but should not trigger native prompt rendering. Only Pi's ask_user_question needs the interactivePrompt payload for live card display. |
||
|
|
49606af8f8 |
Update src/shared/agent-hook-listener.test.ts
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> |
||
|
|
f04895f8b4 |
fix(agent-status): map omp ask tool to blocked sidebar state
The omp agent-hook path could never reach the blocked state: the ask-question gate in normalizePiCompatibleEvent required agentType === 'pi', isAskUserQuestionTool only matched ask_user_question/request_user_input (omp's tool is named 'ask'), and extractPiToolFields only derived interactivePrompt for Pi. An omp agent parked on its ask tool therefore showed 'working' in the sidebar with no question card. Widen all three gates to include omp so tool_call / tool_execution_start for 'ask' maps to blocked and carries the pending-question envelope for the live card. |
||
|
|
e8044b1b30 |
fix(windows): restore fresh-profile startup after durable fsync (#14173)
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> Co-authored-by: DHTheOne <238933622+DHTheOne@users.noreply.github.com> Co-authored-by: Anton Tupitsyn <70199858+PLUTONYY@users.noreply.github.com> Co-authored-by: 7loop <48346764+7loop@users.noreply.github.com> |
||
|
|
c86418eaad |
rm git shim (#14141)
* rm git shim Drops the terminal git/gh PATH wrapper and its settings toggle. Renames the no-marker shell-ready launch config after what it does. Co-authored-by: Orca <help@stably.ai> * rm git shim: clear stale state from older installs Deletes the orphaned wrapper dir and scrubs inherited env/PATH, so a daemon that outlives the upgrade cannot keep seeding it. Drops a now-unread spawn option. Co-authored-by: Orca <help@stably.ai> * rm git shim: cover the daemon and headless paths Scrub after the PATH prepends (they re-read process.env on the sparse daemon env) and run the cleanup above the serve branch so remote hosts get it too. Retry a locked removal; match PATH case-insensitively. Co-authored-by: Orca <help@stably.ai> * rm git shim: keep the scrub final Refuse to re-prepend a legacy entry during agent-teams PATH promotion, which runs after the scrub. Cover the removal guard. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
6ac39b7331 |
fix(worktrees): make hidden agent worktrees recoverable from the visibility dialog (#13652)
* fix(worktrees): make hidden agent worktrees recoverable from the visibility dialog A discovered agent scratch worktree (.claude/worktrees, .gsd-workspaces) was a one-way door: the non-Orca visibility toggle never reveals scratch by design (#9388), the inbox never announces it, and the dialog listed nothing — so once hidden it was unreachable from the UI while sitting on disk. The per-path import exception has outranked the hidden rule all along; no surface offered it. The dialog now refetches an authoritative list on open (a stale snapshot must not read as 'nothing hidden'), lists hidden importable worktrees, and offers a per-row Show wired to the existing inbox import action, which already merges the import + baseline and rolls back on a failed refresh. When the list cannot be read the dialog says so and offers a retry instead of claiming the repo has nothing. No new settings, schema, or persistence: recovery rides entirely on importedExternalWorktreePaths, which every host already stores and validates. The repo-wide toggle is untouched and still never reveals scratch. Fix #10324 Co-authored-by: Gldywn <14254051+Gldywn@users.noreply.github.com> * fix(worktrees): honest scan states and race-safe row actions in the visibility dialog - row Show stays disabled until the open-time authoritative scan settles; a click mid-scan could join the pre-write refetch and read success off a list computed before the import landed, a silent no-op on slow hosts - checking/failed indicators follow the scan state alone, so a warm older snapshot cannot present stale rows as current with no failure indication - ownership-neutral section copy: non-scratch rows are listed too when the repo-wide switch is off * fix(worktrees): close the retry race window and clear stale failure state - Try again is locked while a row import is in flight; a retry scan started before the import's write lands can absorb the import's own refetch and report success off a pre-import list - a successful row import (which requires a successful authoritative refetch) clears an earlier failed open-time scan instead of leaving a contradictory alert over the refreshed list - zh: 智能体 for agent (代理 reads as network proxy); polite live region on the checking hint * fix(worktrees): serialize visibility dialog actions * fix(worktrees): clarify persistent visibility policy * fix(worktrees): clarify hidden worktree list * Explain hidden worktree defaults * Show agent worktrees with Always show * fix(worktrees): bound visibility dialog state and rendering * fix(worktrees): preserve visibility mutation fences across dismissal * fix(worktrees): scope visibility mutations by host --------- Co-authored-by: Gldywn <14254051+Gldywn@users.noreply.github.com> |
||
|
|
9772da844a |
perf(renderer): publish agent status bursts once (#13974)
Agent-status events fan out multiplicatively: every event pays an O(worktrees x tabs) pane-routing scan and its own zustand publication, and WorktreeList's unconditional sortEpoch subscription re-renders the whole sidebar root on each one. A 256-pane reconnect replay meant 256 full sidebar re-renders. Fold a burst into one status publication (plus one generated-title and one tab-title publication - three total, not 2N) via transactAgentStatuses, and build the pane-routing index once per batch. Coalescing is only safe for level-triggered consumers. Two edge-triggered ones needed work: - useAutomationDispatchEvents diffed entry.state, so a swallowed intermediate `done` lost the run's completion output. It now walks the newly appended stateHistory rows and reads the completed turn's output from the entry-level lastCompletedAssistantMessage slot (one message per pane - putting it on every history row would retain 20 transcripts per live status and reprise the renderer OOM in #9872). - pty-connection's native-Windows ConPTY reset compared `state` against a closure-local previous value, so a coalesced done -> [working, done] never re-fired RESET_TERMINAL_CURSOR_STYLE / RESET_KITTY_KEYBOARD_PROTOCOL and left the pane with kitty keyboard protocol armed, corrupting plain input. It now tracks the newest completed turn's start across stateHistory, which also keeps same-turn `done` repaints from re-resetting. The transaction commits as a MERGE patch of the keys the fold changed, not a REPLACE of its snapshot: a batched action reaching another slice through get() writes straight to the real store, and a REPLACE would silently revert it. The shadowed `set` is typed without zustand's `replace` parameter so no call site can reintroduce semantics the merge cannot express. |
||
|
|
cb41f3c5f4 |
fix(terminal): require agent identity for guarded sends (STA-4028) (#14092)
* fix(terminal): require agent identity for guarded sends (STA-4028) Quarter-circle spinner glyphs (U+25D0-U+25D3) started classifying a title as "working" in #13925, and that status alone authorized guarded agent sends — which auto-submit with Enter — into any pane whose TUI animates those generic progress frames. Keep the glyphs as an activity signal, but stop treating a title whose only agent evidence is a quarter circle as proof an agent owns the pane: send authorization now falls through to recognized-agent identity in the title or a recognized foreground process. Braille-spinner behavior is unchanged. * fix(terminal): preserve verified managed busy identity * fix(terminal): bind busy identity to process incarnation * chore(test): avoid duplicate terminal gate suite |
||
|
|
798b9b3d6b |
fix: persist review notes for folder workspaces (#14112)
* fix: persist review notes for folder workspaces * test: satisfy duplicate import lint |
||
|
|
90b8554fc9 | fix(terminal): recover readiness after startup exec (#14027) | ||
|
|
ebe5125476 | Gate paired web creation actions by provider (#13909) | ||
|
|
09ec516ae5 | fix(editor): index WSL watcher aliases per batch (#14015) | ||
|
|
2249330acf |
fix(browser): bulk clear cookies during native import (#13966)
* fix(browser): bulk clear cookies during native import * fix(browser): separate Google cookie import warning * fix(browser): report encrypted Google cookie exclusions * fix(browser): skip key lookup for excluded cookies * test(browser): cover excluded-cookie key bypass * perf(browser): keep plaintext key scan cheap |
||
|
|
a81224614a |
fix(agents): detect a live OpenCode pane from its native OC | session title (#13957)
* fix(agents): detect a live OpenCode pane from its native OC | session title OpenCode publishes `OC | <session>` as its OSC title, which carries no agent-name token. detectAgentStatusFromTitle gates status on a whole-token name match, so it returned null and every status consumer read a live OpenCode pane as a plain shell: no "Send notes to" entry, no title-derived sidebar row, and a title that the runtime's agent-presence check scored as neutral. Identity already resolved (getAgentLabel returns OpenCode); only activity was missing. Treat the native marker as a live idle agent, placed after the spinner and glyph checks so the decorated frames pinned by #8940 keep their status, and accept it in the send-readiness gate the way Claude's U+2733 prefix is accepted -- only a running OpenCode TUI ever publishes it. * fix(agents): require spaced `OC | ` marker for native OpenCode detection Unspaced pipes like `OC|Build` match other tools and would mistakenly route non-OpenCode panes as send targets. Enforce literal ` | ` as OpenCode emits it. Also clarify that only spinner decorations carry working status, not keywords in the session summary. * fix(agents): require OpenCode foreground process to validate native mark OpenCode's native `OC | ` title marker now requires an active OpenCode process to authorize agent sends, preventing false detection when the marker is left on shell prompts. Extends wrapper prefix matching (ssh, tmux, etc.) and spinner glyph support. Adds permission-prompt blocking signals for guarded writes. |
||
|
|
7319d59a10 |
Make worker completion and cleanup authoritative (#13927)
* fix: require authoritative worker completion verdicts * Harden federated settlement replay * fix(orchestration): reconcile dead retained workers * docs: record SSH worker release coverage * test(e2e): exercise worker settlement and release CLI * docs: register combined orchestration CLI oracle * test(orchestration): pin pre-ack attachment state |
||
|
|
971d9548b5 |
docs(agent-status): make design references self-contained (#13902)
docs/design/agent-status-over-ssh.md was cited from ~10 source files but does not exist in the repo. Replace each pointer with the invariant the code actually relies on so the knowledge survives without the doc. Renderer-side citations (useIpcEvents.ts, agent-status-types.ts) are left for the concurrent batching change that owns those files. Co-authored-by: Orca <help@stably.ai> |
||
|
|
5bee7b5ce9 |
P1 STA 3887 design Preview Kitty IME (#13940)
* fix(terminal): carry kitty flags through Preview snapshots and pair rele Preview was omitting the live kitty mirror from the IME bridge and dropping kitty flags from snapshots, so every commit was evaluated at flags 0. A TUI that negotiated bit-3 (report_all_keys_as_escape_codes) would receive the legacy raw text it declined. Now the snapshot carries proven kitty flags beside their sequence boundary, the forwarder reads flags once per commit, and bit-1 (report_event_types) commits are paired with exactly one release regardless of keyup/insertText ordering. Snapshot authorities expose only the active screen's proven flags, so an old host's absent field stays unknown rather than downgraded to a manufactured zero. * fix(terminal): sync kitty flags and IME releases across snapshots * trim wordinesss * fix(terminal): settle owed IME release before fresh same-key press When a keyup is lost and the same key is pressed again, settle the stale record's owed release instead of discarding it — this maintains correct IME state during recovery. Also refine Kitty flag propagation to only carry proven baselines across snapshots, and tighten related comments. * fix(terminal): gate kitty flags on sequence boundaries - Remote snapshots only include flags when seq is present - Daemon uses parsed flags value when defined - Ensures correct flag ordering in snapshot replay |
||
|
|
5ea7df1a5b |
fix(terminal): make DECSET 2031 subscriptions silent (#13904)
fish arms `CSI ?2031h` before painting each prompt and withdraws it when it hands the tty to a child — a ~1ms window. Orca answered that subscribe with `CSI ?997;Nn` across a 1-3ms renderer hop, so the reply landed after the withdrawal and was read as stdin by the next child, corrupting `brew`/`npx` `[y/N]` prompts. The reply is not stale by Orca's own view when written (measured staleReplies: 0), so no suppress-the-stale-reply scheme can close this — the information needed to suppress does not exist yet. Nothing asked for the reply either. The Contour spec says a terminal "should only send out the DSR when the palette has been updated"; Ghostty (Termio.zig:729 — force=true reachable only from the ?996n DSR), iTerm2 (VT100Terminal.m:995 — flag only) and xterm.js (InputHandler.ts:2035 — flag only) all emit nothing on the DECSET. So stop entering the race: record the subscription, answer nothing. Of 17 real programs measured under a pty, only fish, tmux, claude and opencode subscribe; none block on a reply, and answering produces one redundant palette re-query and zero rendering difference. tmux is the only one that sends `?996n`, which Orca still answers. - Subscribes are record-only at all four emitters (live scan, hidden-gate fact, parked byte watcher, parked responder — the last is deleted, it only replied). - `?996n` answers, the subscription registry, and the theme-flip push are unchanged. `paneLastThemeMode` is still seeded at subscribe so the next appearance re-apply is not read as a flip. - Replay grammar carries `?2031l` alongside `?2031h`, so a late-attaching remote client no longer registers a subscription the TUI already retired. Also closes fish-integration gaps found alongside: `unset` (which fish lacks) becomes `set -e` on paths parsed by the client's login shell, `config.fish` is parsed for agent-home detection, and bracketed-paste startup delivery is made consistent across local/daemon/relay. Regression test drives real fish 4.7.1 under node-pty and asserts on what the child process reads; it fails against pre-fix code with the exact payload from the issue. CI installs fish 4 and fails loudly rather than skipping. Closes #9993 Co-authored-by: Orca <help@stably.ai> |
||
|
|
137e724119 |
fix(agent-title): treat Claude Code quarter-circle spinners as working (#13889) (#13925)
* fix(agent-title): treat Claude Code quarter-circle spinners as working
Claude Code 2.1.228 swapped its busy OSC title spinner from braille
(U+2800-U+28FF) to quarter circles (U+25D0/U+25D1). Orca recognized a busy
Claude title only by braille codepoints, so the new frames matched nothing.
The summary-bearing busy frame ("<glyph> Say hi in one word") carries no
"claude" name token, so it resolved to no-status. The tracker's "idle or
permission followed by no-status means the agent exited" rule then fired
mid-turn, confirmPtyAgentExit confirmed it, and the chat surface routed
exitChat -- kicking the tab to the terminal view on every message.
Widen the accepted glyph set via a shared containsAgentSpinnerGlyph helper.
Agent-specific braille frame shapes (Grok, Pi, synthetic Cursor) stay pinned
to their own glyph set.
Fixes #13889
* fix(agent-title): satisfy static analysis and trim scope
|
||
|
|
991a3fe963 |
chore(lint): update oxlint to 1.77 and enable no-op cleanup rules (#13901)
Enable eleven oxlint rules that simplify code without changing behavior, and fix
every existing violation. Each candidate was gated on measured cost rather than
assumption, so rules that regressed runtime performance or type checking were
dropped instead of suppressed.
typescript/no-redundant-type-constituents is the largest addition: 113 sites, no
autofix. Dead constituents are deleted. Where the redundant literal existed to
document intent (`string | 'all'`), it is preserved as `(string & {})`, which
keeps the autocomplete hint the original code was reaching for instead of
flattening it away. The rule also caught a broken import —
remote-shared-control-retirement-probe.ts pulled RuntimeStatus from
src/shared/types, which does not export it, so the type silently degraded to
`any`; no tsconfig covers that file, so tsc never saw it.
oxlint stays at 1.77.0 rather than 1.78.0 because .npmrc sets
minimum-release-age=4320 and 1.78.0 is younger than that window.
Rules evaluated and rejected, with what disqualified each:
- prefer-string-raw: String.raw is a runtime call, not a literal (184x slower)
- prefer-string-replace-all: 26% slower
- text-encoding-identifier-case: ~5% slower, reproducible
- prefer-spread: [...str] is 110% slower than split('') and differs on surrogates
- no-implicit-coercion: `!!x` narrows types and `Boolean(x)` does not (22 tsc errors)
- prefer-arrow-callback: arrows are not constructible, breaking `new` on mocks
- object-shorthand: rewrites source text asserted by a tracked reliability gate
- switch-case-braces: pushes ten files past max-lines, which cannot be suppressed
- no-useless-switch-case: drops `case undefined:` that switch-exhaustiveness-check needs
- arrow-body-style: 115 violations have no fix, and it breaks max-lines
- newline-after-import: false-positives on the leading-semicolon ASI idiom
electron-vite-output-contract asserted on the literal
Object.prototype.hasOwnProperty.call text; retarget it to Object.hasOwn, which
rejects inherited keys identically.
|
||
|
|
4c2a10d157 |
Fix smart sort ranking of done agents by completion time (#13899)
* Fix smart sort ranking of done agents by completion time Completed entries stayed in the Done sort class indefinitely when same-state writes refreshed updatedAt without moving stateStartedAt. Introduce agentEntryCompletionAt() to use actual completion time for both age display and sort eligibility, ensuring consistent aging regardless of hook updates. * Fix smart sort ranking of done agents by completion time Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
55beeebd25 |
Allow directly search in google (#13863)
* Allow direct search with configured search engines Users can now search directly from the tab creation menu using their configured search engine. Forced search mode (`?` prefix) skips file and tab matching for guaranteed search. Refactored tab-entry operations into focused modules for clarity: forced-search parsing, network-safe selection, keyboard focus, copy strings, empty options, and props types. Search routes through the same workspace browser tab opening mechanism used for URLs, with safe title and query presentation that doesn't retain sensitive details. * Allow direct search with configured search engines - Permit search and URL navigation while file index loads; require explicit selection only when needed, not automatic opening - Block malformed IPv6 addresses in bracket notation to prevent misclassification - Support Kagi private-session links via searchUrlOptions - Improve error handling with accessibility: show error messages in status region, disable input during submission, display loading state - Fall back to local browser tab creation when remote creation fails instead of throwing; avoids remote availability blocking local search/navigation - Add error translations for all supported locales (es, ja, ko, zh) * Allow direct search with configured search engines - Extract tab create entry lifecycle to key-driven component remounting, replacing conditional state reset with useEffect cleanup - Consolidate network tab entry classification and request building into reusable helpers, eliminating duplicate logic - Simplify owner resolution by inlining logic directly into openWorkspaceBrowserTab - Replace custom surrogate-pair handling with native String.toWellFormed() for search queries - Disable explicit URL classification to prioritize search-engine queries over raw URLs * Allow direct search from quick-open tab bar entry - Single-token queries keep file matches ranked above search (quick-open intent) - Multi-word phrases promote search to top, since they cannot be file paths - Arm network actions once file index fails or text is unambiguous search - Cache prepared file index to avoid re-processing per keystroke - Generate specific tab titles (e.g. 'example.com/docs') instead of generic 'Open URL' - Surface opening workspace when launching browser tabs remotely * Use readOnly instead of disabled for pending search input Maintain keyboard focus during submission so arrow/Escape navigation continues to work. Use aria-busy to indicate loading state accessibly. Also fixes button hover styling when disabled and cleans up error message handling in the classifier. * Treat bare searches as prompts; refine path and IP classification - Bare search queries (e.g., "?") no longer display as error rows - Path prefixes with existing matches are no longer blocked mid-keystroke - Private IPv4 addresses now use http, public addresses use https - Ambiguous inputs with non-numeric ports fall through to search instead of blocking - Improve diagnostics by logging failure reasons in openFailure |
||
|
|
09c8597fb7 | perf(orchestration): route federated reads over shared control (#13814) | ||
|
|
e77e1fe850 |
fix(claude): guard cold-restore resume selectors (#13868)
* fix(claude): guard cold-restore resume selectors Persisted Claude default args or a custom command can carry their own --resume/-r/--continue/-c selectors (a bare picker default or a stale id). Cold restore appended the authoritative --resume <id> after them, typing a command with competing selectors into the restored pane (#12982). buildAgentResumeStartupPlan now routes Claude through a selector guard that tokenizes the base with the existing startup tokenizer, strips selectors in option position only (value-taking options keep dash-leading values), and appends exactly one authoritative selector, inserting before Claude's own -- terminator when present. Splicing is span-based so untouched bytes stay verbatim, wrapper commands are left alone, and any tokenization failure falls back to the previous append-only behavior. Launch paths, other agents, persistence, and the wire are unchanged. * fix(claude): harden resume selector guard against false matches Round-1 review findings: locate the claude executable by command position (index 0, after a wrapper --, or behind NAME=value assignments) so an argument merely ending in /claude can never be mistaken for it; stop matching the joined -r<id> form, which was ambiguous with dash-leading option values and forced an unmaintainable arity table (now deleted). Ambiguous shapes degrade to the pre-guard append-only behavior. * fix(claude): fail resume guard open on chained shell syntax Round-2 review findings: an unquoted operator or newline after the claude token means the base chains other commands, and splicing across that boundary handed the selector to the wrong command — detect it and fall back to plain appending. Also recognize claude behind PowerShell's & call operator, decouple the test oracle from the implementation's selector predicate, add Windows tokenizer span tests, and rename the module after its public API. * fix(claude): flag bare shell operators inside the tokenizers Round-3 review findings: the guard's operator scan compared raw source to token value, so one quote or escape anywhere in a token hid a shell-active operator outside the quotes and the splice crossed a live command boundary, losing the resume entirely. Both tokenizers now flag tokens carrying an unquoted, unescaped operator byte (or a word-leading # comment on posix/powershell) on their spans, where quote state actually lives, and the guard fails open on that flag. Also strengthens the redirect fail-open test to carry a stale selector, re-tokenizes each raw span in the shell span tests, and documents agent-resume-argv-drop as codex-only. * fix(claude): flag expansions and clamp separator backoff Round-4 review findings: unquoted multi-token expansions (backtick, $(, ${) split across whitespace, so removing only the recognized selector token left a broken construct tail — both tokenizers now raise the span flag (renamed bareShellSyntax) for those openers, on cmd also for operators between single quotes, which cmd does not treat as quoting. The separator backoff is clamped to the previous token's span end so a token ending in an escaped space can no longer donate its escape to the appended selector. * fix(claude): treat cmd single-quoted regions as unmodelable Round-5 review finding: cmd.exe has no single-quote syntax, so the Windows tokenizer's grouping of a single-quoted region diverges from what cmd parses — literal argv like 'claude ...--resume... old' was being read as a real selector and stripped, and a literal '--' as claude's terminator. Flag any cmd single-quoted token as bareShellSyntax so the guard fails open. * fix(claude): flag quoted expansions and scope assignment prefixes Round-6 review findings: the span flag was only evaluated in the unquoted branch, so an expansion opener inside double quotes went unflagged — and inside $(…)/backticks a nested quote re-opens a context this tokenizer does not model, so the splice could cut mid-construct (syntax error, or a silently mutated substitution body). Both tokenizers now flag those, and the flag is renamed divergesFromShell to say what it means. Restrict the NAME=value command-position prefix to posix, where that syntax exists. Drops two branches proven dead. * fix(claude): model shell-literal escapes and scan the whole base Round-7 review findings: (1) the divergence scan started after the claude token, so an expansion opened in a prefix — $(x; npx -- claude --resume s) — had its closer spliced away, producing a base bash cannot parse; it now covers every token including the executable, exempting only PowerShell's leading call operator. (2) posix drops a double-quoted backslash the shell keeps literal, and the Windows escape branch ran inside quoted regions where cmd/PowerShell keep the escape byte literal — both now flagged, so a literal can never be misread as a selector. (3) an unquoted line continuation hid a selector inside a token and skipped the newline gap check. Also removes a third provably dead branch and collapses the cut floor into the cut itself. * fix(claude): flag escapes the tokenizer models but the shell removes Round-8 review findings, all one family — escapes whose token value hides a selector the shell would see: a double-quoted line continuation (bash deletes both bytes), posix $'…'/$"…" quoting, a windows escaped newline, and a trailing unpaired escape. The last one was previously written off as pre-fix-identical, but once stripping happens the dangling escape swallows the separator and no exact --resume reaches claude at all — strictly worse than appending, so it must fail open. Also folds the three gap predicates into one scan. * fix(claude): stop over-flagging a literal dollar sign Round-9 review findings from both lanes: inside double quotes only $( and ${ open an expansion — $' and $" are literal there — and a trailing $ was flagged unconditionally because JS ''.includes('') is true. Both made the guard fail open on modelable bases, leaving the stale selector to compete, so #12982 went unfixed for them. Separately, cmd strips ^ before the child re-splits on the bare whitespace, so an escaped separator hides two real arguments and must fail open rather than drop one. * fix(claude): fail open on cmd caret-quotes and bare PowerShell syntax Round-10 review findings, both Windows-only (a bash oracle cannot reach them): cmd strips a caret before a quote and the child's parser then reads a bare quote delimiter, so the tokenizer's word boundaries stop matching argv — one case turned a working resume into no resume at all, another let a stale selector survive the splice. And bare (…)/{…} are live PowerShell syntax in argument position, so splicing through them emitted unbalanced output that PowerShell cannot parse. * fix(claude): fail open on the PowerShell stop-parsing token Round-11 review finding: after a bare --%, PowerShell passes the rest of the line to the child literally, so the guard stripped a real selector and then appended quoting that arrives as literal bytes — claude ends up with no exact --resume at all, worse than leaving the stale one. Quoted "--%" and cmd, where the token is ordinary, still splice. * fix(claude): model cmd backslash-escaped quotes Round-11 review finding: an odd run of backslashes before a quote makes it a literal byte to the child's CommandLineToArgvW parser, not a delimiter, so the tokenizer's word boundaries stopped matching argv. Orca manufactures that pattern itself — quoteStartupArg wraps every token in quotes without escaping a trailing backslash — so a pasted Windows path was enough to move the selector into a desynced region and leave claude with no resume flag. Also replaces a caret test case that was byte-identical before and after its own fix, and merges two stacked comment blocks. * fix(claude): fail open on PowerShell double-quoted escape sequences Round-12 finding: PowerShell expands backtick escapes only inside double quotes, so a sequence there produces a token value argv never sees — the guard could strip "-`r" plus the argument after it. Also narrows the stop-parsing comment: a quoted --% can engage stop-parsing before a parameter token, where the base is already mangled either way. * fix(claude): flag PowerShell escape sequences in bare arguments too Round-13 finding: the previous commit gated on quote === '"', but PowerShell's tokenizer calls Backtick() from ScanGenericToken, so it expands these sequences in unquoted arguments as well — bare -`r really is a control character, not -r. The guard read it as a selector and dropped it plus the argument after it. Widening to all PowerShell contexts measures 0 under-flag and 0 over-flag across the full printable matrix; the backtick-escaped-space idiom still splices. Also swaps a test case that was byte-identical with and without its own fix. * fix(claude): drop a token-leading PowerShell backtick before whitespace Round-14 observations, all pre-existing and measured: PowerShell drops a token-leading backtick together with the whitespace after it, emitting no token, so the tokenizer's extra token shifted the locator; and a backtick before a bare CR is a line continuation too. Flagging both takes the lane's 329k-base sweep from 87 bad to 0 with no new failures and the must-splice list byte-unchanged. Also corrects a comment that no longer listed every PowerShell divergence. * docs(claude): correct the bare-CR rationale in the tokenizer comment Round-15 verified against a real PowerShell 7.6.4 engine: a backtick before a bare CR is not a line continuation there — pwsh keeps the CR in the token. The flag stays because 5.1 is unverified and failing open costs nothing, but the comment now says that rather than claiming continuation. |
||
|
|
077f5a11cd |
feat(github): create stacked pull requests (#13750)
Adds GitHub stacked pull request creation: a contextual "Stack this PR above #N" option that appears only when the selected base branch has an open PR, plus the main-process stack preflight and registration. Also reworks the create-review composer for cohesion: shadcn Checkbox and Label primitives, base label above a full-width searchable combobox with attached results, keyboard navigation, and a unified field skin, spacing and typography scale. Verified end to end against real GitHub: extending an existing stack and creating a new one. |
||
|
|
63271a5933 |
feat(bitbucket): connect Bitbucket from Settings and create pull requests (#5832)
* feat(bitbucket): connect Bitbucket from Settings with encrypted credential storage Bitbucket Cloud was the only review provider with no in-app auth: GitHub and GitLab delegate to the gh/glab CLIs, but Bitbucket has no comparable first-party CLI, so the only option was ORCA_BITBUCKET_* env vars plus a restart (discussion #5364). Adds a Connect/Edit/Disconnect flow on the Bitbucket integration card, modeled on Linear and Jira: - Credentials are verified against /user before they are persisted, so a dead token is rejected inline instead of silently stored. - The secret is encrypted with safeStorage (0600 plaintext fallback when no OS keyring); non-secret metadata lives in a separate plaintext file so status reads render the connected account without decrypting. Opening Settings therefore never triggers a keychain prompt. - Env vars keep precedence over stored credentials, so existing headless and SSH setups are unaffected. Env-managed connections hide Disconnect. - connect/disconnect reset the preflight cache, so no relaunch is needed. The Bitbucket card moves to its own file to stay under the tsx max-lines cap. * feat(bitbucket): support creating pull requests from Orca Bitbucket was the only configured provider whose Create button reported "This repository provider does not support creating a pull request from Orca" — supportsReviewCreation was false and the forge provider had no createReview, so even a correctly authenticated setup was blocked. Adds createBitbucketPullRequest against POST /repositories/{ws}/{repo}/ pullrequests, using the same env-first / stored-credential resolution as PR lookups (extracted into resolve-auth.ts so both share one path). Bitbucket Cloud has no draft pull requests, so a draft request is rejected with a clear message rather than silently publishing a live PR. * fix(bitbucket): hide the draft toggle where drafts do not exist, plus review fixes Bitbucket Cloud has no draft pull requests, so the composer no longer offers the toggle for it and forces the flag off at submit — better than failing after the user has filled the form in. Review fixes: - writeFileSync's `mode` only applies when it creates the file, so rewriting a credential kept whatever permissions it already had. chmod after every write, for the secret and the metadata. - An explicit ORCA_BITBUCKET_API_BASE_URL now wins over a stored base URL. Env precedence is per-setting, not all-or-nothing. - Enter in the credentials dialog only submits from a text field, so it no longer hijacks Cancel and the docs link. - Replace the chmod-based delete-failure test with a mocked unlinkSync: file modes are not portable to Windows and elevated runners unlink anyway. * fix(bitbucket): stop a merged pull request from blocking the branch's next one Reported on #5832: with a merged PR on a branch, Create reported "Pull request already exists" and offered no way forward. The branch lookup queries every PR state and returns the most recently updated one, so a merged PR came back as the branch's current review and eligibility blocked on it. Bitbucket only discarded such a match on the repo default branch (#9171), while GitHub already drops any merged PR it matched by branch alone — "a merged PR without an explicit link is just a historical branch match, not implicit review context". Applies that rule to Bitbucket. An explicitly linked review still resolves through the linked-number fallback, so merging a PR Orca knows about keeps showing it. * fix(bitbucket): add bitbucket to the shared review-creation provider list Reported on #5832: on a Bitbucket repo with no existing PR, Create still said "This repository provider does not support creating a pull request from Orca", even after the forge provider gained createReview. There are two capability lists. Enabling supportsReviewCreation on the forge provider was necessary but not sufficient — the blocker and the whole renderer read the separate shared list, which never included bitbucket. Adds it, gives Bitbucket its own provider name so review copy stops saying "GitHub", and asserts the two lists agree so they cannot drift apart again. * fix(bitbucket): persist pull request links after creation * fix(bitbucket): fetch linked pull requests by number first * fix(i18n): use generated Bitbucket integration keys * test(bitbucket): cover forge creation delegation * fix(bitbucket): fall back when linked pull request is stale * docs(bitbucket): explain notFoundIsNull and fix a garbled permissions comment notFoundIsNull arrived without the rationale its sibling flag carries, and reads as a bare `true` at the only call site that opts in. * fix(bitbucket): address review findings before merge Two of these made the feature unusable in real setups: - Create PR checked GitHub authentication for Bitbucket. isProviderAuthenticated fell through to isGitHubAuthenticated, which was unreachable while Bitbucket could not create reviews at all. Anyone with Bitbucket connected but no `gh auth login` got auth_required with no way forward. - The draft flag was only gated in ChecksPanel, not the two SourceControl call sites. With "create as draft" saved as a default, the composer hides the toggle for Bitbucket, so the flag could not be cleared and creation failed every time. Bitbucket now ignores draft instead of rejecting it. Also: - Blocked-create copy said "GitHub is not authenticated. Run gh auth login" on Bitbucket repos, in both the main-process and renderer paths. - A decryption failure resolved to an anonymous config and queried anyway; a private repo answers 404, which reads as "no pull request" and offers Create for a branch that already has one. Requests now fail closed. - Hiding non-open implicit branch matches was too broad: a declined PR became permanently invisible off the default branch. Scoped to merged, restoring the default-branch rule (#9171) for the rest. - A failed disconnect rejected unhandled and the card silently re-rendered as connected; a partial delete left the secret live in memory for the session. - The credentials dialog refused to open on a remote runtime, so a local repo could never store a credential. Now only the storage note changes, matching the Jira dialog. --------- Co-authored-by: devatnull <59279509+devatnull@users.noreply.github.com> |
||
|
|
e6eec11b9f | fix(wsl): preserve single-letter POSIX terminal cwd (#13859) | ||
|
|
1f8fd5c44e | fix(editor): guard local WSL path aliases | ||
|
|
d35fcc9e1c |
test(setup): cover Windows forward-slash sequencing (#13557)
Co-authored-by: OrcaWin <alpha-eng@stably.ai> |