mirror of
https://github.com/stablyai/orca.git
synced 2026-10-01 00:02:10 +00:00
e06cdbb4ee0442599eaa7ecef759b8a600bb2cf5
562
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e06cdbb4ee | fix(orchestration): wait for Claude composer render (#14342) | ||
|
|
11cd2b4310 |
revert(ssh): back out #13326 and #13928 — reconnect loses every tab (#14361)
* Revert "fix(daemon): stop killing live coding agents when the daemon can't report its sessions (#13928)" This reverts commit |
||
|
|
8460a63c61 |
fix(orchestration): retry silent mail pointers (#14332)
* fix(orchestration): retry silent mail pointers * fix(orchestration): bound mail pointer repair |
||
|
|
e525f3fe15 | fix(mobile): stop republishing stale launch agent identity (#14244) | ||
|
|
b908b55f6d |
feat(worktrees): add per-source visibility controls (#14189)
* feat(worktrees): add per-source visibility controls * fix(worktrees): explain unsupported visibility hosts * fix(worktrees): keep add location form inline * fix(worktrees): align source visibility across runtimes * test(worktrees): cover Windows drive-relative roots |
||
|
|
9cfa00d665 | Fix federation terminal settlement retries and legacy admission (#14105) | ||
|
|
3ab8b6a117 |
fix(ssh): stop SSH reconnect from multiplying terminals and resuming agents twice (STA-3077) (#13326)
* fix(ssh): stop reconnect from grafting panes and stacking remote leases Reconnecting an SSH-backed workspace added terminal panes the user never opened, and the remote host accumulated shells nobody was using — one report went from 2 to 19 to 20 relay PTYs across three reconnects (STA-3077). Two root causes, both in the store. Reattach could create UI. `persistPtyBinding` has four creating branches — mint a tab, mint a root leaf, split the root and graft a leaf, mint a layout. All four are load-bearing for `pty:spawn`, which can beat the renderer's debounced layout writer, but none of them is appropriate on reattach, where the pane either already exists or is gone for good. Add `mayCreate`, defaulting true so the spawn path is untouched; every creating branch already sets `terminalMembershipChanged`, so refusing is a check rather than a new code path. Lease identity had no pane key. `upsertSshRemotePtyLease` matched on `(targetId, ptyId)` alone, so a pane that re-leased under a new relay id left its predecessor live with nothing to retire it, and the next reattach fanned out over both. One pane now keeps at most one live lease. Superseded leases are marked `expired` rather than terminated: losing a lease is not proof the shell died, so the remote process is deliberately left running. Tests assert observable behavior rather than mechanism, so they stay valid under any implementation that fixes this. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the terminal session behavior contract Properties stated as observable behavior rather than mechanism, so an oracle written against them survives a change of implementation. Records the weaker, correct form of the timer rule — a timer may never be the sole cause of a destructive action — because recovery budgets and scratch-file age gates are correct code that an absolute ban would condemn. Also notes which mechanisms are deliberately not required, so each has to earn its place rather than arrive with an architecture. Co-authored-by: Orca <help@stably.ai> * fix(ssh): heal duplicate pane leases that predate pane-keyed supersession Pane-keyed supersession stops new duplicates, but it does nothing for installs that already carry the ones STA-3077 accumulated — the report behind this reached 20 live leases across a handful of panes, and every reconnect fanned out over all of them. Retire the stale duplicates once per reattach pass, keeping the newest lease for each pane under a total order so two hosts resolve a tie the same way. As with supersession, retired leases are marked `expired` rather than terminated: their remote shells are deliberately left running, because a lease we chose not to revive is not evidence the shell died. The relay-session store stubs gain the new method. Note the gap this leaves open: those shells keep running and are no longer reachable from the app, so the "accumulates unused shells" half of the report needs a visible recovery surface rather than a silent kill. Co-authored-by: Orca <help@stably.ai> * fix(terminal): stop respawning a shell that is still running A pane that failed to reattach spawned a fresh shell. Because the restored session id came along, the replacement resumed the same agent session, and two processes appended to one transcript — reported repeatedly, up to five concurrent resumes of a single session. Two defects fed it. The relay reported a source that merely needed re-establishing as `SSH_SESSION_EXPIRED`. The shell was still running; only its output source was gone. Give that outcome its own error so it stops reading as "the session no longer exists". The reattach failure handler then treated every error as proof of death. It checked for expiry and, in the else branch, took the identical action — so the check bought nothing and a transport fault, a timed-out call, or a wedged relay all respawned. Respawn now requires proof: an explicit host expiry or a not-found PTY. Anything else, including an error we have never seen before, is unresolved, leaves the shell running, and keeps the binding for a later reattach. Two existing tests asserted the old behavior. One threw a bare error as scaffolding to reach the spawn-adoption door; it now throws proof, which is what it meant. The other pinned the expiry mapping itself, and now asserts the outcome fails closed *without* being reported as expiry. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record what makes a retention bound safe Shortening a grace period is the wrong lever. Measuring process time and gating reclamation on an independent observation are what make one safe, and they are what deployed systems actually do. Also records that lifecycle belongs in the attach reply rather than a delivered event — that is what removes the need for a durable per-consumer cursor to guarantee an exit is never lost. Co-authored-by: Orca <help@stably.ai> * test(terminal): assert the empty-failure case without an empty Error A thrown empty value exercises the same property — a failure carrying no usable message is not proof the session is gone — and does not trip the empty-error-message lint. Co-authored-by: Orca <help@stably.ai> * fix(ssh): let the durable pane binding outrank recency when retiring leases Choosing the newest lease for a pane is wrong whenever a newer lease exists that no pane is bound to: it retires the lease the pane is actually attached to, detaching a live terminal instead of healing it. Two changes. Arbitration now prefers the lease matching the pane's durable binding, across both the SSH-target and local partitions, falling back to recency only when no binding names either candidate. And supersession at upsert time now defers rather than expiring a bound predecessor. When a lease arrives for a pane that is still bound to a different PTY, the binding has not caught up yet, so both stay live and reattach arbitrates once the binding is available. Co-authored-by: Orca <help@stably.ai> * fix(ssh): roll back a lease retirement whose durable write fails `flush()` logs and swallows write errors, so a failed write left these leases retired in memory while disk still called them attached — and the pane bindings scrubbed alongside them stayed scrubbed. Use `flushOrThrow` and restore both the lease states and the affected session partitions when it throws, reporting nothing retired. Co-authored-by: Orca <help@stably.ai> * test(ssh): prove pane and remote PTY cardinality across reconnects Counts the shells the relay actually hosts, on the container, rather than inferring them from app state — that is the census the report was based on. Asserts the PIDs are unchanged, not merely the count, so a kill-and-respawn cannot pass. Every pane streams before the transport is severed: an idle pane sends no recovery checkpoint, so only a live source comes back needing re-establishment, which is the outcome that used to read as expiry. Co-authored-by: Orca <help@stably.ai> * fix(ssh): actually pass mayCreate:false from the reattach binding write The `mayCreate` guard was correct and had no production caller, so the reattach path still went through the creating branches and grafted panes back. `restoreReattachedPtyRuntime` is that call site — RC3 in the original diagnosis — and it now refuses to create. Binding moves ahead of runtime registration, because registering first would surface a pane the user never opened before the refusal landed. A refusal leaves the remote shell running and reattachable; a *thrown* write stays unknown and still registers, so a failed disk write cannot detach a live pane. Adds an oracle over the call site itself. The store-level tests all passed while the fix was inert, because they called the store directly — only pinning the wiring catches that. Co-authored-by: Orca <help@stably.ai> * fix(terminal): apply the respawn-requires-proof rule to both reattach paths connectPanePty has two near-verbatim reattach blocks — one keyed on the deferred SSH session, one on the restored session — and only the second was fixed. The first still checked for expiry and then respawned unconditionally anyway, so a transport fault there resumed the same agent session a second time. Also keep the wire token out of the pane. The main-process bridge only special-cases expiry, so a source-restore failure crossed IPC as raw `SSH_SOURCE_RESTORE_REQUIRED: <id>` text and surfaced to the user. It correctly does not respawn; it just should not read like that. Co-authored-by: Orca <help@stably.ai> * test(ssh): state plainly that the reconnect spec is a forward guard It was run against an unfixed tree and passed, so it does not prove the STA-3077 fixes and should not be read as if it does. A clean severed transport does not reproduce the field conditions — accumulated duplicate leases, or a source returning needing re-establishment. It keeps its place as a forward guard: it counts the shells the relay actually hosts and pins their PIDs, so a later change that grafts a pane or respawns a shell fails here. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record that a guard must be pinned at its call site A refusal that exists and is never passed is indistinguishable from no refusal, and store-level tests cannot tell the difference — they call the store directly. Learned from `mayCreate`, which was correct and had no production caller for several commits. Co-authored-by: Orca <help@stably.ai> * fix(ssh): park one PTY's exhausted delivery recovery instead of dropping the channel A per-PTY recovery budget running out disposed the whole relay channel, so one PTY that could not re-prove its delivery aborted every in-flight filesystem and git request on that host and stalled every sibling pane. A retry count is not proof of anything, and it certainly is not proof about the other sessions sharing the channel. Exhaustion now parks that PTY's delivery. The remote shell keeps running, its lease stands, and the next relay open reattaches it with a fresh delivery generation — the parked state is cleared on teardown and the generation changes on reconnect, so a reconnect recovers it. The consecutive-attempt ceiling goes away entirely; the per-generation one is what bounds the retry cost, and the second ceiling only existed to reach the channel drop sooner. Tradeoff worth stating: the failing pane used to self-heal within seconds because the forced reconnect wiped all rejection state, and it now stays frozen until the next relay open. That is a worse outcome for that one pane and a much better one for every other session on the host, and reconnecting is user-reachable. Co-authored-by: Orca <help@stably.ai> * fix(pty): let liveness say unknown instead of forcing it to say dead `IPtyProvider.hasPty` returned a boolean, so a provider whose inventory was empty for reasons that have nothing to do with the session — socket down, cache never hydrated, provider generation just constructed — had no way to say so and answered "absent". Its own siblings already knew better: `probePtyLiveness` and the runtime's `PtyController.hasPty` were both already `boolean | null`, with consumers branching on null correctly. The lie was injected at exactly one interface. Now three-valued, and each provider answers unknown where it cannot prove absence: the daemon adapter off-socket, the SSH provider before a completed listing, the router when any adapter cannot answer, and the degraded provider rather than fabricating a verdict. `terminal_gone` requires unanimous proven absence. Also fixes a real cold-start bug this surfaced: `pty:hasPty` never awaited the daemon-swap startup promise, though the sibling `probePtyLiveness` bridge already did, so before the swap the local provider answered an authoritative false for every daemon-owned id. Net +27 production lines. The plan behind this predicted -92 on the strength of deleting the renderer's dead-session reconcile path; that code is live (`pty-connection.ts` imports it), so nothing was deleted. Expressing a third value where there were two costs lines, and a deletion that is not real is not worth manufacturing. Co-authored-by: Orca <help@stably.ai> * docs(terminal): track the terminal-session correctness handoff package The package was untracked under a gitignored `docs/**`, with the un-ignore rules living only in an uncommitted .gitignore edit — a single `git clean -xdf` would have destroyed the authoritative plan. The 814-path construction snapshot is now pushed as `nwparker/react185-authority-snapshot` too; it had no remote ref. Co-authored-by: Orca <help@stably.ai> * test(ssh): make the reconnect settle window actually wait The settle poll reused a matcher the assertion 15 lines above had already satisfied, and Playwright's poll engine probes immediately and returns as soon as the matcher passes — so it observed the same state twice and elapsed 0ms. A shell grafted a second or two after reattach reported ready slipped through into the next cycle. Reviewer was right on #13111. Test-only; no production change. Co-authored-by: Orca <help@stably.ai> * test(ssh): census both durable session partitions on reconnect Adds a second reconnect scenario and a helper that reads pane records from the local partition as well as the ssh host partition. That split matters: the reattach binding call passes no hostId, so a grafted pane lands in the LOCAL partition and an oracle reading only the host partition passes whether or not the guard is present. Both tests remain forward guards. The second one was reported as discriminating and did not reproduce: with `mayCreate: false` removed from the call site and the app rebuilt, both still passed. Its induction races `pty:kill` against a severed transport, so when the kill lands the lease is cleaned up and there is nothing left to graft. The handoff README is corrected to say so rather than claim a journey. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the user decision relaxing G6 G6 becomes minimise-and-justify rather than strictly net-negative. The deletion budget the plan assumed does not exist: an entrypoint-rooted import graph found 51 of 53 candidate files reachable and instantiated on live paths, leaving 263 deletable LOC against roughly +1,021 to offset. Correctness may still not be traded for line count. Co-authored-by: Orca <help@stably.ai> * test(terminal): add discriminating oracles for restart, daemon, skew and namespaces Six parallel streams, each required to fail with its guard removed rather than merely pass. Local restart proves the OS process itself survives, by reading `ps -o lstart=` for the shell's own pid. That matters: with the quit path made destructive, the tab, leaf and pty ids all came back byte-identical while the shell underneath was a new process — every existing restart spec would have stayed green. Two separate guards were removed to redden it, and the second reddens only the stale-operation case. Daemon restart discriminates by reverting three-valued `hasPty`; version skew now covers publication semantics and confirms the new `SSH_SOURCE_RESTORE_REQUIRED` token mutates nothing on an old client; two-host isolation censuses both containers. Deletes `src/relay/pty-source-replay-index.ts` — 201 production lines with no importer outside its own test, verified against an entrypoint-rooted import graph rather than a name grep. Five namespace tests are skipped, not passing: they reproduce a defect still live on main where folder-workspace ids compare equal with the instance suffix stripped. PR #12474 fixes it; they are its oracle. Co-authored-by: Orca <help@stably.ai> * test(ssh): induce the reattach graft deterministically instead of racing a kill The previous induction closed a pane while the transport was severed and relied on `pty:kill` FAILING so the lease outlived the pane record. It does not fail: with the provider already torn down, `pty:kill` takes its tombstone branch and marks the lease terminated, and `reattachKnownPtys` filters terminated leases out of the fan-out — so the reconnect never visited the PTY the test was about. It passed on both trees. Seed the precondition instead. Spawn a real remote PTY on a leaf that never becomes a pane, then roll the host partition back to its pre-spawn snapshot, leaving a live lease and a live remote shell that no durable pane owns. No failure races a success. Adds a vacuity guard that is independent of the tree under test: the lease's own `lastAttachedAt` must advance, proving the fan-out actually visited this lease before the pane census is trusted. Verified on this machine under an isolated TMPDIR, since the e2e harness keys its seeded-repo pointer on a machine-global tmpdir path: guard present passes, guard removed fails with the phantom leaf grafted into the local partition, guard restored passes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): propose one authoritative binding identity Every defect this program has touched is the same defect: identity compared with the wrong key, or not compared at all. Lease keyed without the pane, reattach using a creating write, folder-workspace ids compared with the instance suffix stripped, local mutating IPC carrying only an id, a live shell classified as expired, liveness unable to say unknown. Proposal: one branded binding type built from fields that already exist and are already persisted, constructible only from an authoritative source, carried by mutating operations, compared by one shared function. Makes a wrong-key comparison a type error rather than the next incident. Under adversarial review, including against the open issue corpus. Not accepted. Co-authored-by: Orca <help@stably.ai> * fix(pty): refuse mutating operations aimed at a superseded PTY `pty:write`, `pty:writeAccepted` and `pty:resize` accepted any id. The renderer queues input, so a keystroke buffered before a reattach landed on whatever PTY had since taken the pane — and a resize reshaped the successor's shell. Main already tracks `ptyPaneKey` and `paneKeyPtyId` in lock-step, so their disagreement is proof the caller's id was superseded. No wire change, no renderer change, nothing added to the input payload. An id with no recorded pane stays permitted: unowned and orphaned PTYs are unknown, not stale, and unknown never authorizes refusing an explicit operation. That is also what keeps orphan cleanup working — those ids have no pane by construction. The tests pin the CALL SITES, not the predicate. A capability that exists and is never called is indistinguishable from no capability, which is exactly how `mayCreate` sat inert here for several commits with every test green. Co-authored-by: Orca <help@stably.ai> * fix(pty): fence signals at a superseded PTY, and pin why kill is exempt A signal means "interrupt my pane", so delivering one to a PTY the pane has already replaced is a misdirected interrupt. Fence it with the same lock-step proof used for write and resize. `pty:kill` stays deliberately unfenced and a test now pins that: a superseded PTY is orphaned, and reclaiming it is exactly what the orphan-cleanup callers ask for. Refusing there would break the operation that reclaims leaked shells — the opposite of the intent. The fence sits at the IPC boundary, above `tryGetProviderForPty`, so it covers local, daemon and SSH rather than the local path alone. Co-authored-by: Orca <help@stably.ai> * test(terminal): poll the pane binding read so a slower host cannot flake it `readPaneBinding` took a single unpolled read of a DOM dataset attribute immediately after a renderer reload, while its sibling helper polls the same data for 15s. On a native Linux host both tests failed every run with 'No bound terminal pane is mounted' while the app was demonstrably healthy — the screenshot showed the terminal restored with a live prompt and the boot PID echoed. The assertion is unchanged; it is only awaited. Nothing is weakened. Found by running this spec on native Linux rather than assuming macOS behaviour generalises. Co-authored-by: Orca <help@stably.ai> * test(terminal): make the restart identity spec run on Windows too Both probes were POSIX-only and unconditional: `echo ...=\$\$` for the shell's own pid, and `ps -o lstart=` for its start time. Running the spec on a real Windows host proved it dies before reaching either guard, so Journey 1's Windows half was unprovable rather than merely unproven. PowerShell exposes the same two facts as `$PID` and `Get-Process` StartTime. The start time still matters on both platforms for the same reason: a PID alone cannot separate a survivor from a reused number. Still green on macOS. The Windows path is written from the host probe and has not itself been executed end to end — that is the next thing to run there, not a claim being made here. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the fence's real gap and what peer designs taught Marks the client-constructed binding proposal as rejected with the three false claims that sank it, and records what shipped instead. States the shipped fence's actual limitation rather than leaving it implied: it compares a binding, not an incarnation, so a respawn under a reused ptyId passes. The obvious remedy is wrong here — the agent-create id is deterministic by design so a replayed create stays idempotent, and randomising it would trade this narrow gap for a duplicate-spawn bug. Also records the ranked lessons from four comparable agent IDEs, chiefly that a typed end-reason at end time is what stops a user quit from looking like a resume candidate. Co-authored-by: Orca <help@stably.ai> * docs(terminal): promote Journey 1 to proven on all three platforms The oracle now runs natively on macOS, Linux and Windows, and its discrimination was watched on each: a mutation reddens it, a restore greens it. On Linux and Windows both mutations were run, and the second reddens only the stale-operation test — so the journey's two clauses are proved independently rather than jointly. Windows is the new evidence. The PowerShell branches added blind at ebffb85a848 executed correctly on their first run: `$PID` expanded to real integers, which also proves the pane shell there is PowerShell-family rather than Git Bash, and `Get-Process StartTime` returned kernel start times 5.4s apart — so a recycled pid could not have passed as a survivor. First journey promoted in this program. The other twelve are unchanged, and the residual limit on "every stale exact operation" is recorded rather than glossed. Co-authored-by: Orca <help@stably.ai> * test(terminal): add discriminating oracles for the daemon, skew and multi-host journeys Daemon: replaces a spec that modelled only a client restart and never crossed the daemon boundary, whose successor generation owned nothing so "the live successor is neither killed nor replaced" was vacuous. The PTY leader is now a real login shell reporting `$$` back through the production write path, resolved to a kernel start time. Two mutations each redden exactly one of the three clauses, on macOS and Linux: reverting three-valued `hasPty` reddens only the unknown-not-dead clause; widening the sole-provider fallback reddens only the stale generation clause. Skew: reverting the restore-required publication to expiry reddens 4 of 5 new tests while the legacy control stays green — the regression this branch fixed is now caught if reintroduced. Multi-host: restoring `mux.dispose('connection_lost')` reddens sibling isolation on one host. It does NOT redden across hosts, and that is recorded rather than glossed: a mux belongs to one relay session per target, so its dispose cannot cross a host boundary. Journey 4's cross-host clause rests on isolation-by-construction, not on a mutation. No production code changes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record journey evidence that falls short of promotion Four journeys now have discriminating oracles but none meets its full stated scope, and each shortfall is named rather than rounded up. Journey 2 is one WSL run from promotion. Journey 12's tests are in-process, so they do not close the live-skew gap the original ledger named. Journey 4's cross-host clause cannot be proven by mutation at all — a mux is per target, so its dispose cannot cross hosts, and the cross-host test stayed green under the mutation that reddens siblings. Journey 13 measured one dimension of ten, on lifted predicates rather than through real IPC. Co-authored-by: Orca <help@stably.ai> * docs(terminal): promote Journey 2 to proven on macOS, Linux and physical WSL The oracle runs on every environment the journey names, and is clause-selective on all three: reverting three-valued `hasPty` reddens only the unknown-not-dead clause, and widening the sole-provider fallback reddens only the stale-generation clause. Selectivity in WSL was established rather than assumed. The spec runs serially, so a red first test reports the others as "did not run" — they were re-run alone under the same mutation and stayed green. Also records that an Orca WSL-mode terminal now starts on that host at all, which it could not before: the distro had no provisioned default Unix user, so every interactive launch blocked on first-run setup. One diagnosis from the WSL run is corrected here rather than repeated: the unrelated `local-pty-shell-ready` failure was attributed to bash 5.3.9, but macOS runs the same bash version and passes 67/67. The trigger is environmental to that distro, and the underlying defect is that the spec pins an absolute count of OSC markers it does not own. Co-authored-by: Orca <help@stably.ai> * docs(terminal): correct the WSL provider-suite diagnosis The WSL run blamed bash 5.3.9 for the unrelated `local-pty-shell-ready` failure. macOS runs the same bash version and passes 67/67, so the version is not the cause — the trigger is environmental to that distro, and the underlying defect is that the spec asserts an absolute count of OSC markers it does not own. Co-authored-by: Orca <help@stably.ai> * test(runtime): unskip the workspace-namespace oracles now their fix has merged These five reproduced a defect that was live on main: folder-workspace ids were compared with the instance suffix stripped, so two workspaces sharing a directory read as the same namespace. They were committed skipped, pointing at the PR that fixes it. That PR is merged, and they pass. Verified they still bite: restoring the suffix-stripping comparison reddens exactly these five and leaves the other four green. An oracle written before its fix, held skipped, and confirmed against the fix after the merge — rather than deleted and rewritten from the answer. Co-authored-by: Orca <help@stably.ai> * test(ssh): add MaxSessions, lazy-discovery and paired-skew oracles Three journeys attempted; none promoted, and the reasons are recorded in the ledger rather than rounded up. MaxSessions=1 against real OpenSSH, with the cap read back from `sshd -T` rather than assumed, and remote pids read on the container two independent ways that must agree, each carrying its kernel start time. Two disjoint mutations discriminate — one reddens only the reconnect clause, the other only the two restart clauses. But the disconnect clause is a forward guard: four separate guard removals left it green, so nothing shipped is load-bearing for it. Lazy discovery samples sshd's own accept log and live session census across a 22s window with the in-use host as a positive control. No mutation reddens its third clause alone — the real cross-host lease scoping is load-bearing, but removing it breaks the sibling host during setup, so the failure carries no clause information. The paired-runtime skew spec pairs two real processes at different versions and refuses to run rather than degrade into a same-version pairing that would look green and prove nothing. No production code changes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record why the duplicate-resume fix was not built I recommended adding a typed end-reason so a user quit stops looking like a resume candidate, then went to implement it and stopped. `SleepingAgentSessionRecord` already carries three fields that each exist to stop something resuming that should not have — `origin`, `restoreOnTabOpenOnly`, and `automaticResumeBlockedBy` — each traceable to its own incident, consulted at 22 non-test sites. A fourth predicate, however well typed, is the fifth containment cycle. The designs without this bug do not have a better flag; they resume only on an explicit action, into a new terminal id, and make two agents in one terminal unrepresentable in the schema. The first of those is a product decision about whether automatic resume stays a feature, so it is the user's call rather than mine. Co-authored-by: Orca <help@stably.ai> * docs(terminal): reconcile G6 with the recorded decision and assess its clauses G6's body still demanded strictly-negative production LOC after the user relaxed it to minimise-and-justify, so the gate had two conflicting pass conditions and no single truth value. Its body now points at that decision. Assessed the remaining clauses against the branch rather than assuming. Two fail structurally: more than one identity comparison and mutation admission path still exist, and `terminal-input-quarantine.ts` is still reachable from two production files. Records why the quarantine is not subsumed by the superseded-PTY fence, which I had assumed and checked. The fence refuses writes aimed at a stale ptyId; the quarantine guards the user's next keystrokes landing on the successor under its current, correct id — a case the fence never sees. Removing it needs the recovery path to surface a different shell as unresolved, not a deletion. Co-authored-by: Orca <help@stably.ai> * docs(terminal): the input quarantine is load-bearing, not superseded G6 lists "no superseded quarantine remains reachable" and this module was assumed to be one. Disabling its single call site reproduces the hazard it exists for — `cho hi; rm -rf x` reaching the shell — so deleting it without a replacement re-opens command execution. The replacement was costed by building it rather than estimated: +26 production LOC to thread the incarnation, ~+33 complete, and the cross-remount state it needs outlives the destroyed pane so it becomes a module about the size of the one deleted. Floor is roughly +140 to delete 88, and it would add a second identity comparison to a gate already failing for having more than one. The decisive part is that the route is not uniformly available: remote runtime results carry no incarnation, old hosts cannot be made to publish one, and mixed versions are the normal state. A paired client reads unknown, which this program's own rule says is not proof — so either every remote reattach surfaces unresolved, or a fallback is needed and the only correct fallback is this module. Whether to amend the clause or accept something weaker on remote hosts is a user decision, so the clause verdict is left as failing rather than quietly reclassified. Co-authored-by: Orca <help@stably.ai> * refactor(runtime): collapse duplicate identity comparisons G6 requires one identity comparison; five implementations existed across two concepts. Worktree-namespace identity had two: `runtimeWorktreeIdsEqual` and `runtimeWorktreeIdentityKey` independently re-derived repoId plus normalized path. Equality now derives from the key, so the comparison and the sleep / mutation-queue keying cannot drift into two different rules — which is exactly how the suffix-stripping bug reached production once. Pane identity had three byte-identical leaf-UUID comparisons, in orchestration `db.ts`, `lifecycle-reconciliation.ts`, and `orchestration-legacy-process-identity.ts`. One copy moved to `stable-pane-id.ts`, which already owns `PaneKey`, `parsePaneKey` and `makePaneKey` and which all three already imported. No new module, no branded type, no parallel comparison. Net -14 production lines. The namespace oracle still bites: restoring the filesystem parser inside the identity key reddens exactly its five cases. The raw counts are not the actionable set, and the classification is worth recording: of 409 non-test `worktreeId` comparisons, 71 are typeof guards and 81 are sentinel tag checks. Most of the remainder are renderer predicates over store rows where both operands are the same main-minted id, so normalizing there would widen equality rather than correct it. Co-authored-by: Orca <help@stably.ai> * refactor(terminal): finish a half-done fixture move and audit the rest `xterm-bypass-event-fixture.ts` and `__fixtures__/xterm-bypass-event.ts` were byte-identical apart from an import path. The `__fixtures__` copy had zero importers and the live copy compiled as production — someone started the move and left both. Dead copy deleted, live one moved, its three test importers updated. Audited the wider G6 clause by importer rather than filename: 32 test-only files, roughly 3,300 LOC, currently compile as production; 4 of the 36 candidates have real production importers and are correctly placed. The list is recorded in the goalposts. Those 32 are almost all older than this program and outside the terminal surface, so sweeping them belongs in its own change rather than inside a terminal PR. The clause stays failing, with the remaining files named. Co-authored-by: Orca <help@stably.ai> * docs(terminal): the fixture clause already holds where it matters Checked what the build emits rather than reasoning from file paths. None of the 32 test-only fixtures appears in `out/` — Rollup drops them because no production entrypoint reaches them. On "compiles into the shipped product", this clause holds today. On the other reading it cannot be closed by moving files at all: both production tsconfigs use bare `include` globs with no `exclude`, so a `__tests__/` directory matches exactly like any other path, as does every `*.test.ts` in the repo. Relocating 32 fixtures would remove nothing from typecheck scope. A sweep was started and stopped once this was verified, rather than landing 32 moves across areas this program does not own for no gain. If the intent is that typecheck scope should exclude test code, that is a repo-wide tsconfig change with a different owner. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add plain-language design and test overviews Two reviewable documents with diagrams, written so someone with no prior context can follow what breaks, why, and what changed. The design overview explains the five things stacked behind one terminal rectangle, the 2 -> 19 -> 20 report, the three root causes, and the rule underneath all of them: unknown is not dead. The test overview explains why a green test proves nothing on its own, the four-step mutation proof we adopted, and — the part worth reviewing hardest — an honest account of what could not be proven and why, including the properties that are true by construction and therefore have no guard to remove. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add a self-contained visual report of the design and its evidence Pre-renders every diagram to inline SVG in both themes so the report opens offline and stays sharp when zoomed. States the gate/journey score and the retractions alongside the fixes, so the unproven half is as visible as the proven half. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the finalized two-plane architecture decision Adopts the data-plane proposal and adds the control-plane track it does not cover: re-key ownership by pane, split orphan inventory out, then delete the compensating code. Records that the host-authority alternative was refuted and that the shipped keystroke fence is inert on the reattach path. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add the design brief the review counsel works from Separates verified code facts from unverified leads so reviewers attack the design rather than a reconstruction of it, and records which simpler alternatives were already refuted and why. Co-authored-by: Orca <help@stably.ai> * docs(terminal): report the design counsel's outcome and the live respawn bug it found Three review rounds across two models replaced the two-record split with one leaf-keyed record, deleted attach-time pane identity, and made orphans a connect-time projection. Records that a shipped gesture still turns a healthy remote shell into a duplicate agent resume, and that the renderer classifier in that chain treats an error-message shape as proof of death. Co-authored-by: Orca <help@stably.ai> * docs(terminal): correct the report — the respawn proof gate guards a minority path A final review traced every auto-respawn route. The primary one converts the reattach failure into a boolean before any classifier sees it, so the shipped proof gate never runs there. Records that two of the six shipped changes are narrower than claimed, and why their tests could not have caught it. Co-authored-by: Orca <help@stably.ai> * docs(terminal): explain the landed design on its own terms One leaf-keyed ownership record, orphans computed at connect, and replacement shells only on positive proof — with the shipping order and the one product trade the design asks the owner to accept. Co-authored-by: Orca <help@stably.ai> * docs(terminal): rewrite the design explainer in plain English The first version assumed the reader knew the codebase. Reframed around two bugs, two fixes and one decision, with the jargon replaced by pane / program / note / helper and a five-word glossary for what could not be avoided. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop reading an identity mismatch as a dead shell The relay reports a pane-identity mismatch by saying the pty was not found, but it found it — comparing identity is how it noticed. Publishing that as expiry made the renderer clear the binding and cold-restore with agent resume, so a live shell gained a second agent on one transcript. Reachable today by detaching a pane into a new tab, which changes the tab the relay froze at spawn. Mismatch now carries its own token and the classifier refuses it as proof. Genuine absence still expires, so a shell that really went away is not stranded. The three failure tokens move to src/shared: main published them and the renderer decided respawn on them, from two copies that had drifted apart. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop sending pane identity on reattach The relay froze pane identity at spawn, so moving a pane to another tab made it refuse a live shell — and refuse by saying 'not found'. The comparison is presence-guarded, so not sending the fields disarms it on every relay version including ones already installed on hosts: no wire change, no redeploy. Nothing is lost. It existed to catch a relay restart recycling pty-N for a new shell, and in exactly that case pane and tab both still match, so it accepted the wrong shell anyway. The incarnation the attach returns is what distinguishes those, and it already crosses the wire. Removes the whole client-side apparatus: the expected-identity type, its per-lease derivation, its map, and the parameter threaded through four layers. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add tracked goalposts for the new design Each goalpost is a behaviour with an oracle and the mutation that must redden it, so 'proven' cannot be claimed from a green test. Records the anti-inert rule as a first-class goalpost, since three guards in this program passed their tests while sitting off the route production takes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record that the recovery grant is dead code, deleting a design step The lease stores a relay-native pty id and the caller passes the app form, with a raw equality comparison between them, so the 30s grant cannot fire for a real SSH pane. The death rule that existed to referee it is deleted rather than built, and the dead path itself becomes a removal. Co-authored-by: Orca <help@stably.ai> * docs(terminal): keep the full design detail in the repo It only existed in an ephemeral job directory, so the plain-English explainer had no durable source for its specifics — record shape, death rule, reattach algorithm, migration order and the 25 oracles. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add a resume prompt for a clean session Points at the goalposts as the contract, names the three goalposts whose oracles are already written and red, and carries the process rules that were learned the expensive way — prove guards reachable, verify mutations land, commit per step, and never let a subagent write production files in a shared worktree. Co-authored-by: Orca <help@stably.ai> * test(ssh): add the failing oracles for goalposts S3, S4 and S5 Intentionally RED: 14 clauses that fail against current behaviour and go green under the changes named in new-design-goalposts.md. The branch is held unmerged, so red here means unimplemented, not broken. Each was verified to fail for the right reason and to flip green under the identified fix, which was then reverted. Each pins the producer as well as the consumer, so no clause can pass vacuously if its route is ever severed — the failure mode that let three earlier guards ship inert. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop fabricating an exit when a reattach fails A failed attach never proves the shell exited. The relay answers not-found for a pane-identity mismatch and for any id it merely cannot hand back, so treating it as death sent the pane a synthetic `pty:exit { code: -1 }`, cleared provider state, deleted ownership and expired the lease — four claims about a process we know nothing about, on a shell that is usually still running. Collapse every failure into the non-destructive branch that already existed a few lines above (`restoreRequired = 'reattachAttemptsExhausted'` + wakeRecovery). A branch collapse, not a new mechanism: goalpost S3. Two tests pinned the deleted premise and are INVERTED rather than patched, so the new intent stays covered: - ssh-relay-orphan-abandon-paths: "retires the lease without a kill when the relay proves the PTY is gone" -> "leaves the shell running when the relay only reports the PTY as not found". Its comment claimed attach verifies liveness before answering not-found; it does not. - ssh-relay-session: "invalidates and broadcasts remote PTYs that cannot reattach" -> "leaves an unreattachable remote PTY alone while its sibling reattaches". Also repairs two clauses left red by |
||
|
|
0ed6db77cf |
fix(mobile): open agent-cited external chat files (#14166)
* fix(mobile): open agent-cited external chat files * fix(mobile): keep cited external files read-only * refactor(mobile): derive cited-file mode from provenance * fix(mobile): accept sentence-final cited paths * fix(mobile): preserve cited SSH grant scope * refactor(file-links): share location suffix parsing |
||
|
|
00cab82fc0 |
Fix terminal split source incarnation and rejection cleanup (#14238)
* fix(terminal): fence split source incarnation * fix(terminal): retire rejected split safely - Track retired rejected PTYs to prevent synthetic exits from landing after split rejection completes. - Validate splits using persisted incarnation IDs only, allowing restored sessions without incarnation maps to work correctly. - Reduce stop timeout from 10s to 2s to avoid stalling on unreachable hosts. |
||
|
|
501337454c |
perf(runtime): stop rescanning every repo every 30s with a Git-admin fingerprint (#14207)
The main-process worktree resolution cache expired on wall-clock time: the whole-fleet snapshot has a 1s TTL, so any poller faster than 1Hz recomputed it, and every 30s the per-repo scan cache expired and shelled out `git worktree list` for every registered repo. A production trace recorded 4,272 `git worktree` invocations over 3h27m across 10 repos. In-Orca mutations are already event-driven, so the 30s TTL existed only to discover changes made outside Orca. Before re-running an expired scan for a local, non-WSL repo, read a cheap subprocess-free Git-admin fingerprint (admin dir entries, per-checkout HEAD and its ref tip, gitdir/locked/existence per entry, packed-refs and reftable stamps). If it matches the fingerprint captured at the cached scan's start, extend the cache without spawning Git. A real scan still runs every 5 minutes so anything the probe cannot see still reconciles. Measured: 600 -> 60 `git worktree list` spawns on the reported workload (10 idle repos, 1Hz polling, 30 simulated minutes), and main-thread event-loop stall of 2.69ms -> 0.01ms per refresh. External worktree add/remove/move/lock/checkout/commit discovery stays bounded at 30s. SSH repos, WSL-routed repos, folder workspaces, agent-scratch repos, and any repo whose layout the probe cannot read keep today's behaviour exactly. Design: docs/reference/worktree-scan-fingerprint.md |
||
|
|
6ac39b7331 |
fix(worktrees): make hidden agent worktrees recoverable from the visibility dialog (#13652)
* fix(worktrees): make hidden agent worktrees recoverable from the visibility dialog A discovered agent scratch worktree (.claude/worktrees, .gsd-workspaces) was a one-way door: the non-Orca visibility toggle never reveals scratch by design (#9388), the inbox never announces it, and the dialog listed nothing — so once hidden it was unreachable from the UI while sitting on disk. The per-path import exception has outranked the hidden rule all along; no surface offered it. The dialog now refetches an authoritative list on open (a stale snapshot must not read as 'nothing hidden'), lists hidden importable worktrees, and offers a per-row Show wired to the existing inbox import action, which already merges the import + baseline and rolls back on a failed refresh. When the list cannot be read the dialog says so and offers a retry instead of claiming the repo has nothing. No new settings, schema, or persistence: recovery rides entirely on importedExternalWorktreePaths, which every host already stores and validates. The repo-wide toggle is untouched and still never reveals scratch. Fix #10324 Co-authored-by: Gldywn <14254051+Gldywn@users.noreply.github.com> * fix(worktrees): honest scan states and race-safe row actions in the visibility dialog - row Show stays disabled until the open-time authoritative scan settles; a click mid-scan could join the pre-write refetch and read success off a list computed before the import landed, a silent no-op on slow hosts - checking/failed indicators follow the scan state alone, so a warm older snapshot cannot present stale rows as current with no failure indication - ownership-neutral section copy: non-scratch rows are listed too when the repo-wide switch is off * fix(worktrees): close the retry race window and clear stale failure state - Try again is locked while a row import is in flight; a retry scan started before the import's write lands can absorb the import's own refetch and report success off a pre-import list - a successful row import (which requires a successful authoritative refetch) clears an earlier failed open-time scan instead of leaving a contradictory alert over the refreshed list - zh: 智能体 for agent (代理 reads as network proxy); polite live region on the checking hint * fix(worktrees): serialize visibility dialog actions * fix(worktrees): clarify persistent visibility policy * fix(worktrees): clarify hidden worktree list * Explain hidden worktree defaults * Show agent worktrees with Always show * fix(worktrees): bound visibility dialog state and rendering * fix(worktrees): preserve visibility mutation fences across dismissal * fix(worktrees): scope visibility mutations by host --------- Co-authored-by: Gldywn <14254051+Gldywn@users.noreply.github.com> |
||
|
|
cb41f3c5f4 |
fix(terminal): require agent identity for guarded sends (STA-4028) (#14092)
* fix(terminal): require agent identity for guarded sends (STA-4028) Quarter-circle spinner glyphs (U+25D0-U+25D3) started classifying a title as "working" in #13925, and that status alone authorized guarded agent sends — which auto-submit with Enter — into any pane whose TUI animates those generic progress frames. Keep the glyphs as an activity signal, but stop treating a title whose only agent evidence is a quarter circle as proof an agent owns the pane: send authorization now falls through to recognized-agent identity in the title or a recognized foreground process. Braille-spinner behavior is unchanged. * fix(terminal): preserve verified managed busy identity * fix(terminal): bind busy identity to process incarnation * chore(test): avoid duplicate terminal gate suite |
||
|
|
798b9b3d6b |
fix: persist review notes for folder workspaces (#14112)
* fix: persist review notes for folder workspaces * test: satisfy duplicate import lint |
||
|
|
a868b090e0 | fix: connect mobile emulator in folder workspaces (#14009) | ||
|
|
a81224614a |
fix(agents): detect a live OpenCode pane from its native OC | session title (#13957)
* fix(agents): detect a live OpenCode pane from its native OC | session title OpenCode publishes `OC | <session>` as its OSC title, which carries no agent-name token. detectAgentStatusFromTitle gates status on a whole-token name match, so it returned null and every status consumer read a live OpenCode pane as a plain shell: no "Send notes to" entry, no title-derived sidebar row, and a title that the runtime's agent-presence check scored as neutral. Identity already resolved (getAgentLabel returns OpenCode); only activity was missing. Treat the native marker as a live idle agent, placed after the spinner and glyph checks so the decorated frames pinned by #8940 keep their status, and accept it in the send-readiness gate the way Claude's U+2733 prefix is accepted -- only a running OpenCode TUI ever publishes it. * fix(agents): require spaced `OC | ` marker for native OpenCode detection Unspaced pipes like `OC|Build` match other tools and would mistakenly route non-OpenCode panes as send targets. Enforce literal ` | ` as OpenCode emits it. Also clarify that only spinner decorations carry working status, not keywords in the session summary. * fix(agents): require OpenCode foreground process to validate native mark OpenCode's native `OC | ` title marker now requires an active OpenCode process to authorize agent sends, preventing false detection when the marker is left on shell prompts. Extends wrapper prefix matching (ssh, tmux, etc.) and spinner glyph support. Adds permission-prompt blocking signals for guarded writes. |
||
|
|
7319d59a10 |
Make worker completion and cleanup authoritative (#13927)
* fix: require authoritative worker completion verdicts * Harden federated settlement replay * fix(orchestration): reconcile dead retained workers * docs: record SSH worker release coverage * test(e2e): exercise worker settlement and release CLI * docs: register combined orchestration CLI oracle * test(orchestration): pin pre-ack attachment state |
||
|
|
5bee7b5ce9 |
P1 STA 3887 design Preview Kitty IME (#13940)
* fix(terminal): carry kitty flags through Preview snapshots and pair rele Preview was omitting the live kitty mirror from the IME bridge and dropping kitty flags from snapshots, so every commit was evaluated at flags 0. A TUI that negotiated bit-3 (report_all_keys_as_escape_codes) would receive the legacy raw text it declined. Now the snapshot carries proven kitty flags beside their sequence boundary, the forwarder reads flags once per commit, and bit-1 (report_event_types) commits are paired with exactly one release regardless of keyup/insertText ordering. Snapshot authorities expose only the active screen's proven flags, so an old host's absent field stays unknown rather than downgraded to a manufactured zero. * fix(terminal): sync kitty flags and IME releases across snapshots * trim wordinesss * fix(terminal): settle owed IME release before fresh same-key press When a keyup is lost and the same key is pressed again, settle the stale record's owed release instead of discarding it — this maintains correct IME state during recovery. Also refine Kitty flag propagation to only carry proven baselines across snapshots, and tighten related comments. * fix(terminal): gate kitty flags on sequence boundaries - Remote snapshots only include flags when seq is present - Daemon uses parsed flags value when defined - Ensures correct flag ordering in snapshot replay |
||
|
|
7a58f72814 |
Fix remote-runtime terminal pane split authority (#13867)
* fix(runtime): preserve remote terminal pane splits * test(runtime): tighten split authority guards * fix(runtime): revalidate transient split sources --------- Co-authored-by: E2E Test <e2e@test.local> |
||
|
|
991a3fe963 |
chore(lint): update oxlint to 1.77 and enable no-op cleanup rules (#13901)
Enable eleven oxlint rules that simplify code without changing behavior, and fix
every existing violation. Each candidate was gated on measured cost rather than
assumption, so rules that regressed runtime performance or type checking were
dropped instead of suppressed.
typescript/no-redundant-type-constituents is the largest addition: 113 sites, no
autofix. Dead constituents are deleted. Where the redundant literal existed to
document intent (`string | 'all'`), it is preserved as `(string & {})`, which
keeps the autocomplete hint the original code was reaching for instead of
flattening it away. The rule also caught a broken import —
remote-shared-control-retirement-probe.ts pulled RuntimeStatus from
src/shared/types, which does not export it, so the type silently degraded to
`any`; no tsconfig covers that file, so tsc never saw it.
oxlint stays at 1.77.0 rather than 1.78.0 because .npmrc sets
minimum-release-age=4320 and 1.78.0 is younger than that window.
Rules evaluated and rejected, with what disqualified each:
- prefer-string-raw: String.raw is a runtime call, not a literal (184x slower)
- prefer-string-replace-all: 26% slower
- text-encoding-identifier-case: ~5% slower, reproducible
- prefer-spread: [...str] is 110% slower than split('') and differs on surrogates
- no-implicit-coercion: `!!x` narrows types and `Boolean(x)` does not (22 tsc errors)
- prefer-arrow-callback: arrows are not constructible, breaking `new` on mocks
- object-shorthand: rewrites source text asserted by a tracked reliability gate
- switch-case-braces: pushes ten files past max-lines, which cannot be suppressed
- no-useless-switch-case: drops `case undefined:` that switch-exhaustiveness-check needs
- arrow-body-style: 115 violations have no fix, and it breaks max-lines
- newline-after-import: false-positives on the leading-semicolon ASI idiom
electron-vite-output-contract asserted on the literal
Object.prototype.hasOwnProperty.call text; retarget it to Object.hasOwn, which
rejects inherited keys identically.
|
||
|
|
077f5a11cd |
feat(github): create stacked pull requests (#13750)
Adds GitHub stacked pull request creation: a contextual "Stack this PR above #N" option that appears only when the selected base branch has an open PR, plus the main-process stack preflight and registration. Also reworks the create-review composer for cohesion: shadcn Checkbox and Label primitives, base label above a full-width searchable combobox with attached results, keyboard navigation, and a unified field skin, spacing and typography scale. Verified end to end against real GitHub: extending an existing stack and creating a new one. |
||
|
|
e790266546 | fix(windows): show first window before shell PATH hydration (#13799) | ||
|
|
5e3a2d25f7 |
Filter remote workspaces by creating device (#13718)
* Filter remote workspaces by creating device * Fix workspace origin filter reconnect behavior * Shorten workspace origin filter label * Refine remote workspace filter UX --------- Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com> |
||
|
|
c1e75477f3 |
Fix static analysis page stuck in loading state (#13674)
* Fix static analysis page stuck in loading state - Bound check-details requests with 30s timeout, matching remote RPC budget - Track request IDs to discard stale responses when context changes - Propagate githubRepository through store and components for proper routing - Add retry button for failed check-details loads - Improve accessibility with ARIA labels for loading and error states * Fix static analysis page stuck in loading state When an open check-details tab's repository is removed, the loading state would continue indefinitely because the fetch was still being triggered. Prevent the fetch call in this scenario to unblock the UI. Also migrates translation keys to obfuscated identifiers. * Fix static analysis page stuck in loading state Add deadline-based timeouts and request ID tracking to prevent stale responses from freezing the checks panel. Include abort signal propagation throughout the request chain and provide retry UI for failed check details loads. * fix(checks): prevent loading state from getting stuck on retry - Consolidate mount checks into a helper function - Details now clear when a new request begins - Add i18n strings for retry status |
||
|
|
ee7e743813 | perf(orchestration): skip redundant federation acknowledgments (#13671) | ||
|
|
550eabe921 | feat(workspaces): review preserved branches after bulk delete (#13693) | ||
|
|
fc7387b28b | perf(runtime): index PTY worktree inventory lookups (#13690) | ||
|
|
f2984e2230 |
fix(terminal): skip a too-wide alt frame on snapshot replay (#13014)
* fix(terminal): skip a too-wide alt frame on snapshot replay Reopening a parked worktree could paint a stale full-width TUI frame through a narrower viewport, leaving clipped gutter fragments and mid-word omissions until the live application repainted. Replay pins xterm to the snapshot grid so soft-wrapped normal-buffer history stays exact (#7279), then the post-replay fit returns the pane to its container grid. Alternate buffers have no scrollback and do not reflow; an absolutely positioned frame remains at its capture layout, so narrowing exposes only clipped portions of those fixed-grid rows. Skip only the visual frame when its capture is wider than the grid the fit will land on. The alt buffer is still entered and cleared, so the resize signal lands on a clean screen that the live application can repaint. Equal-width and wider restores retain the frame, while normal history always replays at its capture grid before fitting. The target width comes from proposeDimensions, not terminal.cols: an unfitted pane can still read xterm's default grid even when its actual container matches the capture. * fix(terminal): preserve offline SSH prepaint frame * fix(terminal): drop a too-wide daemon alt frame on reattach The renderer-only gate did not run on the user-visible remount path. Instrumentation showed that daemon-connectResult-snapshot won before the model-snapshot branches, so the composed snapshot painted in full and the later narrower fit exposed its stale fixed-grid frame clipped at the new viewport. The daemon branch could not omit only the visual frame while main sent one merged string. Publish the normal-buffer/mode prefix and visual alt frame as additive optional metadata while retaining the merged snapshot for mixed-version fallback. New renderers can keep history and restore state without painting a frame captured for a wider grid. Replay ordering remains capture-grid, write, then fit so normal-buffer soft wrapping stays exact. The application owns the foreign-width alt frame and repaints it after the resize signal rather than Orca trying to transform an absolutely positioned screen. * fix(terminal): preserve split daemon snapshot payload * fix(terminal): preserve live state when dropping alt frame * fix(terminal): restore DECOM cursor state exactly * fix(terminal): preserve ordinary snapshot bytes * test(terminal): cover fixed-grid alt replay resize * fix(terminal): repaint after dropping mismatched frames A hidden snapshot can omit an alternate-screen frame before the pane has a measurable target grid. If reveal later lands on the capture grid, a same-size PTY resize emits no SIGWINCH, so pulse the local PTY size whenever that frame was skipped.\n\nKeep performSafeFit's measurable-pane contract intact, publish daemon snapshot prefix and frame as explicit optional strings, and fall back to the merged payload when either field is absent. Cold owner-gone restores now omit a mismatched frame while retaining history and fresh-shell reset treatment; offline SSH preconnect remains unchanged.\n\nPin the vendored SerializeAddon out-of-range-row behavior used to capture live SGR state. |
||
|
|
8da362919e |
perf(runtime): batch legacy worker recovery persistence (#13649)
* perf(runtime): batch legacy recovery persistence * fix(runtime): preserve concurrent recovery state * fix(runtime): require durable recovery retry |
||
|
|
84bd306949 |
perf: Stop unchanged worktree refresh churn (#13662)
* fix: stop unchanged worktree refresh churn * fix: preserve smart sort telemetry recomputations * fix: preserve duplicate worktree host identities * perf: skip reconciled catalog traversal * test: strengthen worktree refresh regressions |
||
|
|
ec7e3ea477 |
fix(terminal): prevent paired activity renderer starvation (#13508)
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> |
||
|
|
ce3ec4d5ce |
fix(repo-icon): keep a renamed fork's own owner avatar (#12271)
* fix(repo-icon): keep a renamed fork's own owner avatar Fork repos always took the upstream owner's avatar, so a renamed fork showed its parent project's logo. Same-name forks (personal copies) still prefer the upstream owner; renamed forks now keep their origin owner across auto-detect, the startup backfill, and the settings avatar refresh. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(repo-icon): re-read repo state before backfill avatar write The startup backfill computed icon updates from a pre-loop snapshot, so an icon chosen in settings while the upstream/origin probes were pending could be clobbered. Re-read the repo after the probes and only migrate an icon that is still the auto-detected GitHub avatar. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(repo-icon): own the fork avatar rule in one shared selector The renamed-fork rule was written out twice — once in the main-process auto-detect and once in the renderer refresh — so the two copies could drift. Move it next to `githubAvatarIcon` as `githubAvatarSlug`, which collapses the renderer resolver to a single unbranched path. Also stop swallowing a rejected origin probe: it cannot tell a renamed fork from a same-name one, so degrading to the upstream owner would flip a renamed fork's stored avatar back to the parent's. Letting it propagate keeps the stored icon, matching how the non-fork path already behaved. Adds coverage for the startup backfill, the third decision point the fix claims, which had none. * test(repo-icon): cover pending backfill icon change --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
f5d934a3f5 |
3c5f0060 (#13496)
* refactor(runtime): extract pure path-candidate, review-branch, and folder-workspace helpers from orca-runtime.ts Mechanical move of three closed, pure module-scope clusters out of orca-runtime.ts (37,608 -> 37,207 lines) into domain-named siblings: - terminal-output-path-candidates.ts: PTY output path harvesting and the recent-candidate history bound (3 entry points + 15 private callees). - selected-review-branch.ts: forge-agnostic selected-review predicates and lookup hints (GitHub/GitLab/Bitbucket/Azure DevOps/Gitea). - runtime-folder-workspace.ts: folder-workspace id math and the repo+meta -> Worktree projection. Bodies are token-identical to their previous form; the only production changes are the moves, the new import statements, and `export` keywords. The no-control-regex suppression travels with the path-candidate scanning that needs it. No max-lines suppression was added and the ratchet is unchanged. Adds characterization tests for the two clusters that had no direct coverage; the path-candidate cluster keeps its existing tests, repointed at the new module. * fix flaky timer on CI |
||
|
|
73efab98f7 | perf(orchestration): keep drift Git off main thread (#13440) | ||
|
|
158212b8b3 |
feat(github): add PR comment reactions (#13470)
* feat(github): add PR comment reactions * fix(github): harden comment reaction updates --------- Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com> |
||
|
|
75650936f4 | fix(source-control): hydrate Windows shell profile PATH (#13418) | ||
|
|
b075a95b06 |
Strip liveness gate from AI Vault session delete (#13279)
* Strip liveness gate from AI Vault session delete Delete now requires only path validation + user confirmation — no process roster, no liveness check, no quiescence, no ownership ledger. Co-authored-by: Orca <help@stably.ai> * Remove obsolete AI Vault liveness delete reliability gate Session delete no longer checks process liveness, so drop the manifest entry that still referenced the deleted test files. * minor fix --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
5df2ddbc9c |
perf(ai-vault): isolate tab title resolution (#13377)
* perf(ai-vault): isolate tab title resolution * fix(ai-vault): preserve background scan caches * fix(ai-vault): resolve nested worker from chunks |
||
|
|
6397668271 | Add manual artifact sharing from HTML and Markdown views (#13369) | ||
|
|
3ec48a74d5 |
Gate artifact publishing behind off-by-default capability (#13368)
* fix(artifacts): gate agent artifact publishing behind an off-by-default capability Public artifact sharing was reachable by any agent through `orca artifacts share`: the Artifacts settings toggle only controlled sidebar visibility, and nothing in the main process checked a capability before minting a public URL. Add `artifactSharingEnabled` (default off) and enforce it in ArtifactCloudService.share/update — before auth, network, or the share-record write — so the CLI, relay-forwarded remote CLI, and IPC paths are all denied. The denial carries a stable `artifact_sharing_disabled` code plus next steps through the RPC error allowlist, so the CLI prints actionable guidance. list, unshare, and delete stay ungated: turning publishing off must not strand already-published links. The capability is absent from the `settings.update` RPC schema, so an agent cannot grant it to itself — only the desktop UI can. Co-authored-by: Orca <help@stably.ai> * fix(artifacts): gate agent artifact publishing behind an off-by-default Publishing is blocked until enabled in Settings → Artifacts. CLI preflights the capability before reading files to avoid unnecessary uploads. RPC surface rejects capability grants so callers cannot self-grant. UI shows opt-in workflow and recovery path when publishing is off. Web clients mirror the host's setting read-only. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
970696a008 |
fix(sidebar): show Cursor rows and stop a stray "claude" title hijacking OpenCode (#12466)
* fix(sidebar): show Cursor rows and stop a stray "claude" title hijacking OpenCode Two defects in the same title-resolution path. **#10258** — Cursor's only native OSC title is the literal `cursor agent`, which both title trackers dropped unconditionally. A hookless Cursor pane therefore had neither a status entry nor any title carrying Cursor identity, so the worktree card showed nothing at all. **#8940** — two owner-blind paths let an incidental `claude` token anywhere in an OpenCode session or task title outrank the pane's known owner, so the tab icon and sidebar row flipped to Claude Code. #10258: let the literal through exactly once as identity, so a restored or mobile tab keeps its Cursor row instead of vanishing. #8940: require an *identity frame* — after stripping status decoration the title must PRESENT Claude, not merely mention it — before a Claude title may reclaim a pane from its prior identity, and make the sidebar row builder owner-aware. > These two are in one PR because they share the `ownerAgentType` plumbing through `buildTitleDerivedAgentRow` — split apart, neither half compiles on its own. Fixes #10258 Fixes #8940 Co-authored-by: Orca <help@stably.ai> * test(e2e): add recordable proof for sidebar-agent-row-identity Fails on origin/main, passes on this branch. Test: sidebar keeps a Cursor pane visible and an OpenCode pane out of Claude Code hands Co-authored-by: Orca <help@stably.ai> * fix(terminal): preserve restored Cursor identity * test(terminal): cover restored Cursor redraw suppression * refactor(terminal): tighten Cursor identity handling and Claude frame matching Review follow-ups on the title-resolution path: - pty-transport dropped a native Cursor literal that main emits whenever a non-Cursor title preceded it, re-introducing the #10258 blank row in the renderer path. The pre-filter now projects the predecessor the drain will actually see, and defers to the drain gate while facts are still queued. - applyTrackedPtyTitle threaded the cursor flag through 12 sites, including ptyRecordChanged bookkeeping the sole caller ignores. Force the status null once, and the activity-gated effects fall out unchanged. - isClaudeIdentityFrameTitle missed a multiplexer-wrapped Claude title ("zsh | Claude Code"), costing a genuine Claude pane its identity. Reuse the ' | ' segment split that agent-title-owner already had inline. - Keep title normalization on launchAgent: it only rewrites within an identity group (OMP wraps Pi), so a split does not make it wrong, and the hook-row path normalizes the same way. - Drop the tab.ptyId tracker fallback, which read a pty that the pane identity check had just rejected. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
06144ca26c |
fix(worktrees): guard lineage pruning from failed scans (#13248)
* fix(worktrees): guard lineage pruning from failed scans * fix(worktrees): back off failed local resolution scans |
||
|
|
c991bb27d3 | Add account-backed artifact sharing (#13012) | ||
|
|
6ea99f6607 |
fix(runtime): isolate same-path folder workspace PTY identity (#12474)
Folder-project workspace ids (`repoId::/path::workspace:<uuid>`) were compared via the suffix-stripping `splitWorktreeIdForFilesystem`, so every workspace sharing one directory compared equal at ~35 runtime call sites — PTYs leaked between siblings and paired/mobile clients hung on "Loading terminal". Both identity helpers now use the suffix-preserving `splitWorktreeId`, `stopTerminalsForWorktree` routes through the shared helper, and `findResolvedWorktreeIdForPath` gains a `targetWorktreeId` tie-break. Co-authored-by: dgk-dev <dgk-dev@users.noreply.github.com> |
||
|
|
2f30eb9af5 |
fix(ai-vault): block deletion of live sessions (#13108)
* fix(ai-vault): block deletion of live sessions * fix(ai-vault): retain external session authority --------- Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com> |
||
|
|
96a6c93757 |
fix(orchestration): allow fresh agents to take over legacy runs (#12896)
* fix(orchestration): allow fresh agents to take over legacy runs * test(orchestration): cover forged takeover evidence * fix(orchestration): retain renderer launch authority --------- Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com> |
||
|
|
ddf58d6d6a |
fix(terminal): restore preserved remote PTYs after host relaunch (#12990)
* fix(terminal): foreground preserved daemon PTYs * fix(terminal): keep snapshot sequence domains distinct * test(terminal): use active reconnect control * test(terminal): await reconnect control activation * test(terminal): validate reconnect with fresh control * test(terminal): tighten host restart evidence * fix(terminal): retry preserved PTY attach after inventory * fix(terminal): retry attach after overlapping inventory --------- Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com> |
||
|
|
2b42de1f52 |
fix(orchestration): wake coordinators with mail pointers (#12988)
Wake idle Run coordinators with durable orchestration mail pointers while keeping message payloads in the store until check consumes them. Preserve waiter, Cursor, restart, real Codex title, and PTY replacement behavior.\n\nPart of #12953. |
||
|
|
a025a71447 |
fix(orchestration): deliver pending mail to already-idle agents (#12584)
Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com> |
||
|
|
8ddf575fe6 |
Revert "Remove source control group order preference (#12785)" (#12955)
This reverts commit
|
||
|
|
74e1678d4b |
test(runtime): pin the fail-closed contract for bare Cursor titles (#12816)
Reverts the classification change and keeps behavior at base. cursor-agent's native OSC title is the bare literal "Cursor Agent" and never carries a status word, so it names the agent without proving one is present. The title tracker drops it live, so main records it only when the stale-working timer strips the spinner off the synthesized "⠋ Cursor Agent" — and that fires both when Cursor parks idle and when cursor-agent exited and the shell reclaimed the pane. The two states are observationally identical: same title, same null foreground read. Classifying it as an agent therefore removes a refusal rather than adding evidence. Guarded sends auto-submit Enter, so the false positive types into the user's shell. A null foreground is also not "unreadable" on the default local provider, which returns null when the pty is gone. hasPty, probePtyLiveness, hasChildProcesses and inspectProcess were each checked as corroborating signals; none separates alive-with-agent from alive-with-shell when the foreground read is unavailable. Tests pin every no-evidence branch fail-closed and document the mechanism, so both attempted fixes fail loudly if reintroduced. Real gap tracked in #12946. |