mirror of
https://github.com/stablyai/orca.git
synced 2026-09-23 08:02:31 +00:00
c3b8c145e2e060da170a300151ebd1160c045243
53
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
11cd2b4310 |
revert(ssh): back out #13326 and #13928 — reconnect loses every tab (#14361)
* Revert "fix(daemon): stop killing live coding agents when the daemon can't report its sessions (#13928)" This reverts commit |
||
|
|
3ab8b6a117 |
fix(ssh): stop SSH reconnect from multiplying terminals and resuming agents twice (STA-3077) (#13326)
* fix(ssh): stop reconnect from grafting panes and stacking remote leases Reconnecting an SSH-backed workspace added terminal panes the user never opened, and the remote host accumulated shells nobody was using — one report went from 2 to 19 to 20 relay PTYs across three reconnects (STA-3077). Two root causes, both in the store. Reattach could create UI. `persistPtyBinding` has four creating branches — mint a tab, mint a root leaf, split the root and graft a leaf, mint a layout. All four are load-bearing for `pty:spawn`, which can beat the renderer's debounced layout writer, but none of them is appropriate on reattach, where the pane either already exists or is gone for good. Add `mayCreate`, defaulting true so the spawn path is untouched; every creating branch already sets `terminalMembershipChanged`, so refusing is a check rather than a new code path. Lease identity had no pane key. `upsertSshRemotePtyLease` matched on `(targetId, ptyId)` alone, so a pane that re-leased under a new relay id left its predecessor live with nothing to retire it, and the next reattach fanned out over both. One pane now keeps at most one live lease. Superseded leases are marked `expired` rather than terminated: losing a lease is not proof the shell died, so the remote process is deliberately left running. Tests assert observable behavior rather than mechanism, so they stay valid under any implementation that fixes this. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the terminal session behavior contract Properties stated as observable behavior rather than mechanism, so an oracle written against them survives a change of implementation. Records the weaker, correct form of the timer rule — a timer may never be the sole cause of a destructive action — because recovery budgets and scratch-file age gates are correct code that an absolute ban would condemn. Also notes which mechanisms are deliberately not required, so each has to earn its place rather than arrive with an architecture. Co-authored-by: Orca <help@stably.ai> * fix(ssh): heal duplicate pane leases that predate pane-keyed supersession Pane-keyed supersession stops new duplicates, but it does nothing for installs that already carry the ones STA-3077 accumulated — the report behind this reached 20 live leases across a handful of panes, and every reconnect fanned out over all of them. Retire the stale duplicates once per reattach pass, keeping the newest lease for each pane under a total order so two hosts resolve a tie the same way. As with supersession, retired leases are marked `expired` rather than terminated: their remote shells are deliberately left running, because a lease we chose not to revive is not evidence the shell died. The relay-session store stubs gain the new method. Note the gap this leaves open: those shells keep running and are no longer reachable from the app, so the "accumulates unused shells" half of the report needs a visible recovery surface rather than a silent kill. Co-authored-by: Orca <help@stably.ai> * fix(terminal): stop respawning a shell that is still running A pane that failed to reattach spawned a fresh shell. Because the restored session id came along, the replacement resumed the same agent session, and two processes appended to one transcript — reported repeatedly, up to five concurrent resumes of a single session. Two defects fed it. The relay reported a source that merely needed re-establishing as `SSH_SESSION_EXPIRED`. The shell was still running; only its output source was gone. Give that outcome its own error so it stops reading as "the session no longer exists". The reattach failure handler then treated every error as proof of death. It checked for expiry and, in the else branch, took the identical action — so the check bought nothing and a transport fault, a timed-out call, or a wedged relay all respawned. Respawn now requires proof: an explicit host expiry or a not-found PTY. Anything else, including an error we have never seen before, is unresolved, leaves the shell running, and keeps the binding for a later reattach. Two existing tests asserted the old behavior. One threw a bare error as scaffolding to reach the spawn-adoption door; it now throws proof, which is what it meant. The other pinned the expiry mapping itself, and now asserts the outcome fails closed *without* being reported as expiry. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record what makes a retention bound safe Shortening a grace period is the wrong lever. Measuring process time and gating reclamation on an independent observation are what make one safe, and they are what deployed systems actually do. Also records that lifecycle belongs in the attach reply rather than a delivered event — that is what removes the need for a durable per-consumer cursor to guarantee an exit is never lost. Co-authored-by: Orca <help@stably.ai> * test(terminal): assert the empty-failure case without an empty Error A thrown empty value exercises the same property — a failure carrying no usable message is not proof the session is gone — and does not trip the empty-error-message lint. Co-authored-by: Orca <help@stably.ai> * fix(ssh): let the durable pane binding outrank recency when retiring leases Choosing the newest lease for a pane is wrong whenever a newer lease exists that no pane is bound to: it retires the lease the pane is actually attached to, detaching a live terminal instead of healing it. Two changes. Arbitration now prefers the lease matching the pane's durable binding, across both the SSH-target and local partitions, falling back to recency only when no binding names either candidate. And supersession at upsert time now defers rather than expiring a bound predecessor. When a lease arrives for a pane that is still bound to a different PTY, the binding has not caught up yet, so both stay live and reattach arbitrates once the binding is available. Co-authored-by: Orca <help@stably.ai> * fix(ssh): roll back a lease retirement whose durable write fails `flush()` logs and swallows write errors, so a failed write left these leases retired in memory while disk still called them attached — and the pane bindings scrubbed alongside them stayed scrubbed. Use `flushOrThrow` and restore both the lease states and the affected session partitions when it throws, reporting nothing retired. Co-authored-by: Orca <help@stably.ai> * test(ssh): prove pane and remote PTY cardinality across reconnects Counts the shells the relay actually hosts, on the container, rather than inferring them from app state — that is the census the report was based on. Asserts the PIDs are unchanged, not merely the count, so a kill-and-respawn cannot pass. Every pane streams before the transport is severed: an idle pane sends no recovery checkpoint, so only a live source comes back needing re-establishment, which is the outcome that used to read as expiry. Co-authored-by: Orca <help@stably.ai> * fix(ssh): actually pass mayCreate:false from the reattach binding write The `mayCreate` guard was correct and had no production caller, so the reattach path still went through the creating branches and grafted panes back. `restoreReattachedPtyRuntime` is that call site — RC3 in the original diagnosis — and it now refuses to create. Binding moves ahead of runtime registration, because registering first would surface a pane the user never opened before the refusal landed. A refusal leaves the remote shell running and reattachable; a *thrown* write stays unknown and still registers, so a failed disk write cannot detach a live pane. Adds an oracle over the call site itself. The store-level tests all passed while the fix was inert, because they called the store directly — only pinning the wiring catches that. Co-authored-by: Orca <help@stably.ai> * fix(terminal): apply the respawn-requires-proof rule to both reattach paths connectPanePty has two near-verbatim reattach blocks — one keyed on the deferred SSH session, one on the restored session — and only the second was fixed. The first still checked for expiry and then respawned unconditionally anyway, so a transport fault there resumed the same agent session a second time. Also keep the wire token out of the pane. The main-process bridge only special-cases expiry, so a source-restore failure crossed IPC as raw `SSH_SOURCE_RESTORE_REQUIRED: <id>` text and surfaced to the user. It correctly does not respawn; it just should not read like that. Co-authored-by: Orca <help@stably.ai> * test(ssh): state plainly that the reconnect spec is a forward guard It was run against an unfixed tree and passed, so it does not prove the STA-3077 fixes and should not be read as if it does. A clean severed transport does not reproduce the field conditions — accumulated duplicate leases, or a source returning needing re-establishment. It keeps its place as a forward guard: it counts the shells the relay actually hosts and pins their PIDs, so a later change that grafts a pane or respawns a shell fails here. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record that a guard must be pinned at its call site A refusal that exists and is never passed is indistinguishable from no refusal, and store-level tests cannot tell the difference — they call the store directly. Learned from `mayCreate`, which was correct and had no production caller for several commits. Co-authored-by: Orca <help@stably.ai> * fix(ssh): park one PTY's exhausted delivery recovery instead of dropping the channel A per-PTY recovery budget running out disposed the whole relay channel, so one PTY that could not re-prove its delivery aborted every in-flight filesystem and git request on that host and stalled every sibling pane. A retry count is not proof of anything, and it certainly is not proof about the other sessions sharing the channel. Exhaustion now parks that PTY's delivery. The remote shell keeps running, its lease stands, and the next relay open reattaches it with a fresh delivery generation — the parked state is cleared on teardown and the generation changes on reconnect, so a reconnect recovers it. The consecutive-attempt ceiling goes away entirely; the per-generation one is what bounds the retry cost, and the second ceiling only existed to reach the channel drop sooner. Tradeoff worth stating: the failing pane used to self-heal within seconds because the forced reconnect wiped all rejection state, and it now stays frozen until the next relay open. That is a worse outcome for that one pane and a much better one for every other session on the host, and reconnecting is user-reachable. Co-authored-by: Orca <help@stably.ai> * fix(pty): let liveness say unknown instead of forcing it to say dead `IPtyProvider.hasPty` returned a boolean, so a provider whose inventory was empty for reasons that have nothing to do with the session — socket down, cache never hydrated, provider generation just constructed — had no way to say so and answered "absent". Its own siblings already knew better: `probePtyLiveness` and the runtime's `PtyController.hasPty` were both already `boolean | null`, with consumers branching on null correctly. The lie was injected at exactly one interface. Now three-valued, and each provider answers unknown where it cannot prove absence: the daemon adapter off-socket, the SSH provider before a completed listing, the router when any adapter cannot answer, and the degraded provider rather than fabricating a verdict. `terminal_gone` requires unanimous proven absence. Also fixes a real cold-start bug this surfaced: `pty:hasPty` never awaited the daemon-swap startup promise, though the sibling `probePtyLiveness` bridge already did, so before the swap the local provider answered an authoritative false for every daemon-owned id. Net +27 production lines. The plan behind this predicted -92 on the strength of deleting the renderer's dead-session reconcile path; that code is live (`pty-connection.ts` imports it), so nothing was deleted. Expressing a third value where there were two costs lines, and a deletion that is not real is not worth manufacturing. Co-authored-by: Orca <help@stably.ai> * docs(terminal): track the terminal-session correctness handoff package The package was untracked under a gitignored `docs/**`, with the un-ignore rules living only in an uncommitted .gitignore edit — a single `git clean -xdf` would have destroyed the authoritative plan. The 814-path construction snapshot is now pushed as `nwparker/react185-authority-snapshot` too; it had no remote ref. Co-authored-by: Orca <help@stably.ai> * test(ssh): make the reconnect settle window actually wait The settle poll reused a matcher the assertion 15 lines above had already satisfied, and Playwright's poll engine probes immediately and returns as soon as the matcher passes — so it observed the same state twice and elapsed 0ms. A shell grafted a second or two after reattach reported ready slipped through into the next cycle. Reviewer was right on #13111. Test-only; no production change. Co-authored-by: Orca <help@stably.ai> * test(ssh): census both durable session partitions on reconnect Adds a second reconnect scenario and a helper that reads pane records from the local partition as well as the ssh host partition. That split matters: the reattach binding call passes no hostId, so a grafted pane lands in the LOCAL partition and an oracle reading only the host partition passes whether or not the guard is present. Both tests remain forward guards. The second one was reported as discriminating and did not reproduce: with `mayCreate: false` removed from the call site and the app rebuilt, both still passed. Its induction races `pty:kill` against a severed transport, so when the kill lands the lease is cleaned up and there is nothing left to graft. The handoff README is corrected to say so rather than claim a journey. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the user decision relaxing G6 G6 becomes minimise-and-justify rather than strictly net-negative. The deletion budget the plan assumed does not exist: an entrypoint-rooted import graph found 51 of 53 candidate files reachable and instantiated on live paths, leaving 263 deletable LOC against roughly +1,021 to offset. Correctness may still not be traded for line count. Co-authored-by: Orca <help@stably.ai> * test(terminal): add discriminating oracles for restart, daemon, skew and namespaces Six parallel streams, each required to fail with its guard removed rather than merely pass. Local restart proves the OS process itself survives, by reading `ps -o lstart=` for the shell's own pid. That matters: with the quit path made destructive, the tab, leaf and pty ids all came back byte-identical while the shell underneath was a new process — every existing restart spec would have stayed green. Two separate guards were removed to redden it, and the second reddens only the stale-operation case. Daemon restart discriminates by reverting three-valued `hasPty`; version skew now covers publication semantics and confirms the new `SSH_SOURCE_RESTORE_REQUIRED` token mutates nothing on an old client; two-host isolation censuses both containers. Deletes `src/relay/pty-source-replay-index.ts` — 201 production lines with no importer outside its own test, verified against an entrypoint-rooted import graph rather than a name grep. Five namespace tests are skipped, not passing: they reproduce a defect still live on main where folder-workspace ids compare equal with the instance suffix stripped. PR #12474 fixes it; they are its oracle. Co-authored-by: Orca <help@stably.ai> * test(ssh): induce the reattach graft deterministically instead of racing a kill The previous induction closed a pane while the transport was severed and relied on `pty:kill` FAILING so the lease outlived the pane record. It does not fail: with the provider already torn down, `pty:kill` takes its tombstone branch and marks the lease terminated, and `reattachKnownPtys` filters terminated leases out of the fan-out — so the reconnect never visited the PTY the test was about. It passed on both trees. Seed the precondition instead. Spawn a real remote PTY on a leaf that never becomes a pane, then roll the host partition back to its pre-spawn snapshot, leaving a live lease and a live remote shell that no durable pane owns. No failure races a success. Adds a vacuity guard that is independent of the tree under test: the lease's own `lastAttachedAt` must advance, proving the fan-out actually visited this lease before the pane census is trusted. Verified on this machine under an isolated TMPDIR, since the e2e harness keys its seeded-repo pointer on a machine-global tmpdir path: guard present passes, guard removed fails with the phantom leaf grafted into the local partition, guard restored passes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): propose one authoritative binding identity Every defect this program has touched is the same defect: identity compared with the wrong key, or not compared at all. Lease keyed without the pane, reattach using a creating write, folder-workspace ids compared with the instance suffix stripped, local mutating IPC carrying only an id, a live shell classified as expired, liveness unable to say unknown. Proposal: one branded binding type built from fields that already exist and are already persisted, constructible only from an authoritative source, carried by mutating operations, compared by one shared function. Makes a wrong-key comparison a type error rather than the next incident. Under adversarial review, including against the open issue corpus. Not accepted. Co-authored-by: Orca <help@stably.ai> * fix(pty): refuse mutating operations aimed at a superseded PTY `pty:write`, `pty:writeAccepted` and `pty:resize` accepted any id. The renderer queues input, so a keystroke buffered before a reattach landed on whatever PTY had since taken the pane — and a resize reshaped the successor's shell. Main already tracks `ptyPaneKey` and `paneKeyPtyId` in lock-step, so their disagreement is proof the caller's id was superseded. No wire change, no renderer change, nothing added to the input payload. An id with no recorded pane stays permitted: unowned and orphaned PTYs are unknown, not stale, and unknown never authorizes refusing an explicit operation. That is also what keeps orphan cleanup working — those ids have no pane by construction. The tests pin the CALL SITES, not the predicate. A capability that exists and is never called is indistinguishable from no capability, which is exactly how `mayCreate` sat inert here for several commits with every test green. Co-authored-by: Orca <help@stably.ai> * fix(pty): fence signals at a superseded PTY, and pin why kill is exempt A signal means "interrupt my pane", so delivering one to a PTY the pane has already replaced is a misdirected interrupt. Fence it with the same lock-step proof used for write and resize. `pty:kill` stays deliberately unfenced and a test now pins that: a superseded PTY is orphaned, and reclaiming it is exactly what the orphan-cleanup callers ask for. Refusing there would break the operation that reclaims leaked shells — the opposite of the intent. The fence sits at the IPC boundary, above `tryGetProviderForPty`, so it covers local, daemon and SSH rather than the local path alone. Co-authored-by: Orca <help@stably.ai> * test(terminal): poll the pane binding read so a slower host cannot flake it `readPaneBinding` took a single unpolled read of a DOM dataset attribute immediately after a renderer reload, while its sibling helper polls the same data for 15s. On a native Linux host both tests failed every run with 'No bound terminal pane is mounted' while the app was demonstrably healthy — the screenshot showed the terminal restored with a live prompt and the boot PID echoed. The assertion is unchanged; it is only awaited. Nothing is weakened. Found by running this spec on native Linux rather than assuming macOS behaviour generalises. Co-authored-by: Orca <help@stably.ai> * test(terminal): make the restart identity spec run on Windows too Both probes were POSIX-only and unconditional: `echo ...=\$\$` for the shell's own pid, and `ps -o lstart=` for its start time. Running the spec on a real Windows host proved it dies before reaching either guard, so Journey 1's Windows half was unprovable rather than merely unproven. PowerShell exposes the same two facts as `$PID` and `Get-Process` StartTime. The start time still matters on both platforms for the same reason: a PID alone cannot separate a survivor from a reused number. Still green on macOS. The Windows path is written from the host probe and has not itself been executed end to end — that is the next thing to run there, not a claim being made here. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the fence's real gap and what peer designs taught Marks the client-constructed binding proposal as rejected with the three false claims that sank it, and records what shipped instead. States the shipped fence's actual limitation rather than leaving it implied: it compares a binding, not an incarnation, so a respawn under a reused ptyId passes. The obvious remedy is wrong here — the agent-create id is deterministic by design so a replayed create stays idempotent, and randomising it would trade this narrow gap for a duplicate-spawn bug. Also records the ranked lessons from four comparable agent IDEs, chiefly that a typed end-reason at end time is what stops a user quit from looking like a resume candidate. Co-authored-by: Orca <help@stably.ai> * docs(terminal): promote Journey 1 to proven on all three platforms The oracle now runs natively on macOS, Linux and Windows, and its discrimination was watched on each: a mutation reddens it, a restore greens it. On Linux and Windows both mutations were run, and the second reddens only the stale-operation test — so the journey's two clauses are proved independently rather than jointly. Windows is the new evidence. The PowerShell branches added blind at ebffb85a848 executed correctly on their first run: `$PID` expanded to real integers, which also proves the pane shell there is PowerShell-family rather than Git Bash, and `Get-Process StartTime` returned kernel start times 5.4s apart — so a recycled pid could not have passed as a survivor. First journey promoted in this program. The other twelve are unchanged, and the residual limit on "every stale exact operation" is recorded rather than glossed. Co-authored-by: Orca <help@stably.ai> * test(terminal): add discriminating oracles for the daemon, skew and multi-host journeys Daemon: replaces a spec that modelled only a client restart and never crossed the daemon boundary, whose successor generation owned nothing so "the live successor is neither killed nor replaced" was vacuous. The PTY leader is now a real login shell reporting `$$` back through the production write path, resolved to a kernel start time. Two mutations each redden exactly one of the three clauses, on macOS and Linux: reverting three-valued `hasPty` reddens only the unknown-not-dead clause; widening the sole-provider fallback reddens only the stale generation clause. Skew: reverting the restore-required publication to expiry reddens 4 of 5 new tests while the legacy control stays green — the regression this branch fixed is now caught if reintroduced. Multi-host: restoring `mux.dispose('connection_lost')` reddens sibling isolation on one host. It does NOT redden across hosts, and that is recorded rather than glossed: a mux belongs to one relay session per target, so its dispose cannot cross a host boundary. Journey 4's cross-host clause rests on isolation-by-construction, not on a mutation. No production code changes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record journey evidence that falls short of promotion Four journeys now have discriminating oracles but none meets its full stated scope, and each shortfall is named rather than rounded up. Journey 2 is one WSL run from promotion. Journey 12's tests are in-process, so they do not close the live-skew gap the original ledger named. Journey 4's cross-host clause cannot be proven by mutation at all — a mux is per target, so its dispose cannot cross hosts, and the cross-host test stayed green under the mutation that reddens siblings. Journey 13 measured one dimension of ten, on lifted predicates rather than through real IPC. Co-authored-by: Orca <help@stably.ai> * docs(terminal): promote Journey 2 to proven on macOS, Linux and physical WSL The oracle runs on every environment the journey names, and is clause-selective on all three: reverting three-valued `hasPty` reddens only the unknown-not-dead clause, and widening the sole-provider fallback reddens only the stale-generation clause. Selectivity in WSL was established rather than assumed. The spec runs serially, so a red first test reports the others as "did not run" — they were re-run alone under the same mutation and stayed green. Also records that an Orca WSL-mode terminal now starts on that host at all, which it could not before: the distro had no provisioned default Unix user, so every interactive launch blocked on first-run setup. One diagnosis from the WSL run is corrected here rather than repeated: the unrelated `local-pty-shell-ready` failure was attributed to bash 5.3.9, but macOS runs the same bash version and passes 67/67. The trigger is environmental to that distro, and the underlying defect is that the spec pins an absolute count of OSC markers it does not own. Co-authored-by: Orca <help@stably.ai> * docs(terminal): correct the WSL provider-suite diagnosis The WSL run blamed bash 5.3.9 for the unrelated `local-pty-shell-ready` failure. macOS runs the same bash version and passes 67/67, so the version is not the cause — the trigger is environmental to that distro, and the underlying defect is that the spec asserts an absolute count of OSC markers it does not own. Co-authored-by: Orca <help@stably.ai> * test(runtime): unskip the workspace-namespace oracles now their fix has merged These five reproduced a defect that was live on main: folder-workspace ids were compared with the instance suffix stripped, so two workspaces sharing a directory read as the same namespace. They were committed skipped, pointing at the PR that fixes it. That PR is merged, and they pass. Verified they still bite: restoring the suffix-stripping comparison reddens exactly these five and leaves the other four green. An oracle written before its fix, held skipped, and confirmed against the fix after the merge — rather than deleted and rewritten from the answer. Co-authored-by: Orca <help@stably.ai> * test(ssh): add MaxSessions, lazy-discovery and paired-skew oracles Three journeys attempted; none promoted, and the reasons are recorded in the ledger rather than rounded up. MaxSessions=1 against real OpenSSH, with the cap read back from `sshd -T` rather than assumed, and remote pids read on the container two independent ways that must agree, each carrying its kernel start time. Two disjoint mutations discriminate — one reddens only the reconnect clause, the other only the two restart clauses. But the disconnect clause is a forward guard: four separate guard removals left it green, so nothing shipped is load-bearing for it. Lazy discovery samples sshd's own accept log and live session census across a 22s window with the in-use host as a positive control. No mutation reddens its third clause alone — the real cross-host lease scoping is load-bearing, but removing it breaks the sibling host during setup, so the failure carries no clause information. The paired-runtime skew spec pairs two real processes at different versions and refuses to run rather than degrade into a same-version pairing that would look green and prove nothing. No production code changes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record why the duplicate-resume fix was not built I recommended adding a typed end-reason so a user quit stops looking like a resume candidate, then went to implement it and stopped. `SleepingAgentSessionRecord` already carries three fields that each exist to stop something resuming that should not have — `origin`, `restoreOnTabOpenOnly`, and `automaticResumeBlockedBy` — each traceable to its own incident, consulted at 22 non-test sites. A fourth predicate, however well typed, is the fifth containment cycle. The designs without this bug do not have a better flag; they resume only on an explicit action, into a new terminal id, and make two agents in one terminal unrepresentable in the schema. The first of those is a product decision about whether automatic resume stays a feature, so it is the user's call rather than mine. Co-authored-by: Orca <help@stably.ai> * docs(terminal): reconcile G6 with the recorded decision and assess its clauses G6's body still demanded strictly-negative production LOC after the user relaxed it to minimise-and-justify, so the gate had two conflicting pass conditions and no single truth value. Its body now points at that decision. Assessed the remaining clauses against the branch rather than assuming. Two fail structurally: more than one identity comparison and mutation admission path still exist, and `terminal-input-quarantine.ts` is still reachable from two production files. Records why the quarantine is not subsumed by the superseded-PTY fence, which I had assumed and checked. The fence refuses writes aimed at a stale ptyId; the quarantine guards the user's next keystrokes landing on the successor under its current, correct id — a case the fence never sees. Removing it needs the recovery path to surface a different shell as unresolved, not a deletion. Co-authored-by: Orca <help@stably.ai> * docs(terminal): the input quarantine is load-bearing, not superseded G6 lists "no superseded quarantine remains reachable" and this module was assumed to be one. Disabling its single call site reproduces the hazard it exists for — `cho hi; rm -rf x` reaching the shell — so deleting it without a replacement re-opens command execution. The replacement was costed by building it rather than estimated: +26 production LOC to thread the incarnation, ~+33 complete, and the cross-remount state it needs outlives the destroyed pane so it becomes a module about the size of the one deleted. Floor is roughly +140 to delete 88, and it would add a second identity comparison to a gate already failing for having more than one. The decisive part is that the route is not uniformly available: remote runtime results carry no incarnation, old hosts cannot be made to publish one, and mixed versions are the normal state. A paired client reads unknown, which this program's own rule says is not proof — so either every remote reattach surfaces unresolved, or a fallback is needed and the only correct fallback is this module. Whether to amend the clause or accept something weaker on remote hosts is a user decision, so the clause verdict is left as failing rather than quietly reclassified. Co-authored-by: Orca <help@stably.ai> * refactor(runtime): collapse duplicate identity comparisons G6 requires one identity comparison; five implementations existed across two concepts. Worktree-namespace identity had two: `runtimeWorktreeIdsEqual` and `runtimeWorktreeIdentityKey` independently re-derived repoId plus normalized path. Equality now derives from the key, so the comparison and the sleep / mutation-queue keying cannot drift into two different rules — which is exactly how the suffix-stripping bug reached production once. Pane identity had three byte-identical leaf-UUID comparisons, in orchestration `db.ts`, `lifecycle-reconciliation.ts`, and `orchestration-legacy-process-identity.ts`. One copy moved to `stable-pane-id.ts`, which already owns `PaneKey`, `parsePaneKey` and `makePaneKey` and which all three already imported. No new module, no branded type, no parallel comparison. Net -14 production lines. The namespace oracle still bites: restoring the filesystem parser inside the identity key reddens exactly its five cases. The raw counts are not the actionable set, and the classification is worth recording: of 409 non-test `worktreeId` comparisons, 71 are typeof guards and 81 are sentinel tag checks. Most of the remainder are renderer predicates over store rows where both operands are the same main-minted id, so normalizing there would widen equality rather than correct it. Co-authored-by: Orca <help@stably.ai> * refactor(terminal): finish a half-done fixture move and audit the rest `xterm-bypass-event-fixture.ts` and `__fixtures__/xterm-bypass-event.ts` were byte-identical apart from an import path. The `__fixtures__` copy had zero importers and the live copy compiled as production — someone started the move and left both. Dead copy deleted, live one moved, its three test importers updated. Audited the wider G6 clause by importer rather than filename: 32 test-only files, roughly 3,300 LOC, currently compile as production; 4 of the 36 candidates have real production importers and are correctly placed. The list is recorded in the goalposts. Those 32 are almost all older than this program and outside the terminal surface, so sweeping them belongs in its own change rather than inside a terminal PR. The clause stays failing, with the remaining files named. Co-authored-by: Orca <help@stably.ai> * docs(terminal): the fixture clause already holds where it matters Checked what the build emits rather than reasoning from file paths. None of the 32 test-only fixtures appears in `out/` — Rollup drops them because no production entrypoint reaches them. On "compiles into the shipped product", this clause holds today. On the other reading it cannot be closed by moving files at all: both production tsconfigs use bare `include` globs with no `exclude`, so a `__tests__/` directory matches exactly like any other path, as does every `*.test.ts` in the repo. Relocating 32 fixtures would remove nothing from typecheck scope. A sweep was started and stopped once this was verified, rather than landing 32 moves across areas this program does not own for no gain. If the intent is that typecheck scope should exclude test code, that is a repo-wide tsconfig change with a different owner. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add plain-language design and test overviews Two reviewable documents with diagrams, written so someone with no prior context can follow what breaks, why, and what changed. The design overview explains the five things stacked behind one terminal rectangle, the 2 -> 19 -> 20 report, the three root causes, and the rule underneath all of them: unknown is not dead. The test overview explains why a green test proves nothing on its own, the four-step mutation proof we adopted, and — the part worth reviewing hardest — an honest account of what could not be proven and why, including the properties that are true by construction and therefore have no guard to remove. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add a self-contained visual report of the design and its evidence Pre-renders every diagram to inline SVG in both themes so the report opens offline and stays sharp when zoomed. States the gate/journey score and the retractions alongside the fixes, so the unproven half is as visible as the proven half. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the finalized two-plane architecture decision Adopts the data-plane proposal and adds the control-plane track it does not cover: re-key ownership by pane, split orphan inventory out, then delete the compensating code. Records that the host-authority alternative was refuted and that the shipped keystroke fence is inert on the reattach path. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add the design brief the review counsel works from Separates verified code facts from unverified leads so reviewers attack the design rather than a reconstruction of it, and records which simpler alternatives were already refuted and why. Co-authored-by: Orca <help@stably.ai> * docs(terminal): report the design counsel's outcome and the live respawn bug it found Three review rounds across two models replaced the two-record split with one leaf-keyed record, deleted attach-time pane identity, and made orphans a connect-time projection. Records that a shipped gesture still turns a healthy remote shell into a duplicate agent resume, and that the renderer classifier in that chain treats an error-message shape as proof of death. Co-authored-by: Orca <help@stably.ai> * docs(terminal): correct the report — the respawn proof gate guards a minority path A final review traced every auto-respawn route. The primary one converts the reattach failure into a boolean before any classifier sees it, so the shipped proof gate never runs there. Records that two of the six shipped changes are narrower than claimed, and why their tests could not have caught it. Co-authored-by: Orca <help@stably.ai> * docs(terminal): explain the landed design on its own terms One leaf-keyed ownership record, orphans computed at connect, and replacement shells only on positive proof — with the shipping order and the one product trade the design asks the owner to accept. Co-authored-by: Orca <help@stably.ai> * docs(terminal): rewrite the design explainer in plain English The first version assumed the reader knew the codebase. Reframed around two bugs, two fixes and one decision, with the jargon replaced by pane / program / note / helper and a five-word glossary for what could not be avoided. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop reading an identity mismatch as a dead shell The relay reports a pane-identity mismatch by saying the pty was not found, but it found it — comparing identity is how it noticed. Publishing that as expiry made the renderer clear the binding and cold-restore with agent resume, so a live shell gained a second agent on one transcript. Reachable today by detaching a pane into a new tab, which changes the tab the relay froze at spawn. Mismatch now carries its own token and the classifier refuses it as proof. Genuine absence still expires, so a shell that really went away is not stranded. The three failure tokens move to src/shared: main published them and the renderer decided respawn on them, from two copies that had drifted apart. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop sending pane identity on reattach The relay froze pane identity at spawn, so moving a pane to another tab made it refuse a live shell — and refuse by saying 'not found'. The comparison is presence-guarded, so not sending the fields disarms it on every relay version including ones already installed on hosts: no wire change, no redeploy. Nothing is lost. It existed to catch a relay restart recycling pty-N for a new shell, and in exactly that case pane and tab both still match, so it accepted the wrong shell anyway. The incarnation the attach returns is what distinguishes those, and it already crosses the wire. Removes the whole client-side apparatus: the expected-identity type, its per-lease derivation, its map, and the parameter threaded through four layers. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add tracked goalposts for the new design Each goalpost is a behaviour with an oracle and the mutation that must redden it, so 'proven' cannot be claimed from a green test. Records the anti-inert rule as a first-class goalpost, since three guards in this program passed their tests while sitting off the route production takes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record that the recovery grant is dead code, deleting a design step The lease stores a relay-native pty id and the caller passes the app form, with a raw equality comparison between them, so the 30s grant cannot fire for a real SSH pane. The death rule that existed to referee it is deleted rather than built, and the dead path itself becomes a removal. Co-authored-by: Orca <help@stably.ai> * docs(terminal): keep the full design detail in the repo It only existed in an ephemeral job directory, so the plain-English explainer had no durable source for its specifics — record shape, death rule, reattach algorithm, migration order and the 25 oracles. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add a resume prompt for a clean session Points at the goalposts as the contract, names the three goalposts whose oracles are already written and red, and carries the process rules that were learned the expensive way — prove guards reachable, verify mutations land, commit per step, and never let a subagent write production files in a shared worktree. Co-authored-by: Orca <help@stably.ai> * test(ssh): add the failing oracles for goalposts S3, S4 and S5 Intentionally RED: 14 clauses that fail against current behaviour and go green under the changes named in new-design-goalposts.md. The branch is held unmerged, so red here means unimplemented, not broken. Each was verified to fail for the right reason and to flip green under the identified fix, which was then reverted. Each pins the producer as well as the consumer, so no clause can pass vacuously if its route is ever severed — the failure mode that let three earlier guards ship inert. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop fabricating an exit when a reattach fails A failed attach never proves the shell exited. The relay answers not-found for a pane-identity mismatch and for any id it merely cannot hand back, so treating it as death sent the pane a synthetic `pty:exit { code: -1 }`, cleared provider state, deleted ownership and expired the lease — four claims about a process we know nothing about, on a shell that is usually still running. Collapse every failure into the non-destructive branch that already existed a few lines above (`restoreRequired = 'reattachAttemptsExhausted'` + wakeRecovery). A branch collapse, not a new mechanism: goalpost S3. Two tests pinned the deleted premise and are INVERTED rather than patched, so the new intent stays covered: - ssh-relay-orphan-abandon-paths: "retires the lease without a kill when the relay proves the PTY is gone" -> "leaves the shell running when the relay only reports the PTY as not found". Its comment claimed attach verifies liveness before answering not-found; it does not. - ssh-relay-session: "invalidates and broadcasts remote PTYs that cannot reattach" -> "leaves an unreattachable remote PTY alone while its sibling reattaches". Also repairs two clauses left red by |
||
|
|
501337454c |
perf(runtime): stop rescanning every repo every 30s with a Git-admin fingerprint (#14207)
The main-process worktree resolution cache expired on wall-clock time: the whole-fleet snapshot has a 1s TTL, so any poller faster than 1Hz recomputed it, and every 30s the per-repo scan cache expired and shelled out `git worktree list` for every registered repo. A production trace recorded 4,272 `git worktree` invocations over 3h27m across 10 repos. In-Orca mutations are already event-driven, so the 30s TTL existed only to discover changes made outside Orca. Before re-running an expired scan for a local, non-WSL repo, read a cheap subprocess-free Git-admin fingerprint (admin dir entries, per-checkout HEAD and its ref tip, gitdir/locked/existence per entry, packed-refs and reftable stamps). If it matches the fingerprint captured at the cached scan's start, extend the cache without spawning Git. A real scan still runs every 5 minutes so anything the probe cannot see still reconciles. Measured: 600 -> 60 `git worktree list` spawns on the reported workload (10 idle repos, 1Hz polling, 30 simulated minutes), and main-thread event-loop stall of 2.69ms -> 0.01ms per refresh. External worktree add/remove/move/lock/checkout/commit discovery stays bounded at 30s. SSH repos, WSL-routed repos, folder workspaces, agent-scratch repos, and any repo whose layout the probe cannot read keep today's behaviour exactly. Design: docs/reference/worktree-scan-fingerprint.md |
||
|
|
95ff19f35f | chore(repo): ignore generated Clawpatch state (#14144) | ||
|
|
9c091cf77e |
chore(perf): add renderer agent-status benchmark harness (#13905)
Splits run-idle-cpu-benchmark.mjs into a scale fixture, an in-page timing probe, and process sampling, and records a measured origin/main baseline so the agent-status batching slice has an auditable before. The agent-status write workload is not included: it needs setAgentStatuses, so it lands with the store slice. |
||
|
|
d50adec2d2 |
feat(ai-vault): isolate scanning from terminal workloads (#13411)
* feat(ai-vault): isolate scanning in service processes * fix(ai-vault): retire idle service processes * fix(ai-vault): discard unverified cache processes * fix(ai-vault): clear relay sidecar cancel watchdog on acknowledgement A cancelled relay call is settled before its 2s cancel watchdog is armed, so the acknowledgement path bailed out of settle() before clearing the timer. The watchdog then faulted a healthy sidecar two seconds after every aborted scan, killing whatever request had since become active. * fix(ai-vault): clear the pending restart before scheduling another recordFault overwrote this.timer, stranding a restart that dispose() could no longer cancel. * refactor(ai-vault): drop the orphaned first-prompt IPC wrapper session-first-user-prompt-handler.ts now owns this entry point and routes through the service; the copy left in the read module had no callers. * fix(ai-vault): retry a faulted cold start before surfacing it A slow first start surfaced a raw 'did not become ready' error to the caller even though the supervisor was already respawning. Requeue an unsent call once onto the scheduled respawn instead. Also stop arming the cancellation watchdog for a call the child never received: no acknowledgement is coming, so it killed a healthy service and stalled the lane. Invalidation bookkeeping and ready-waiter construction move to the state module to stay under the max-lines cap. * fix(ai-vault): give relay title reads their own lane Before this branch the relay read title files directly, concurrently with scans. Routing both through one sidecar lane put title resolution behind a list scan that may run up to 130s, so SSH tab titles could lag minutes behind. Split cache and interactive lanes in both the relay client and the sidecar entry, mirroring the desktop service. Also: clear the ready deadline on fault, so a sidecar that dies before ready cannot fault its healthy replacement five seconds later; retry an unsent call once across a respawn; and skip the cancellation watchdog for a call the sidecar never received. Restart/circuit bookkeeping moves to its own module, mirroring the desktop policy, to stay under the max-lines cap. * fix(ai-vault): degrade relay title resolution on sidecar failure listSessions already returns a host issue when the sidecar is unavailable; titles propagated the raw RPC error instead. Return no titles so callers fall back to preview text, and keep cancellation propagating. * fix(ai-vault): scrub the service child environment The children are forked with a 384 MiB heap cap and no loader, but both spawn sites handed them the full parent environment, so an exported NODE_OPTIONS silently raised the cap or --require'd code into them. Allowlist both, following the plugin worker. The desktop child keeps the eleven agent-root overrides it resolves its own roots from; the relay sidecar takes remoteHome and hostPlatform from its init message and so needs none of them. Both children share one priority module while they share this one. * fix(ai-vault): soft-disable relay vault when the service is missing A missing service threw out of the constructor, so a Vault wiring bug would abort relay startup and take every PTY on the host with it. The unsupported-platform branch three lines above already treats a Vault failure as a soft disable; do the same here. Threading the service through the two handlers instead of a field also retires the definite-assignment assertion the throw was propping up. * fix(ai-vault): drain consumed cache invalidations invalidatedPaths was re-applied in every request's finally and never drained, so once N paths had been invalidated every later request paid N evictions for the life of the process; the 4096 cap only bounded how bad that got. The re-apply exists to cover a read that overlapped the invalidation, so drain once nothing is executing. Clearing unconditionally would drop the re-apply for a request still running on the other lane. * fix(ai-vault): keep a busy child through slow invalidation acks invalidate() reused the 5s ready budget as its acknowledgement deadline and killed the child on expiry, so a delete issued during a large scan could kill a healthy process mid-scan and burn a slot toward the restart circuit. Fault only when nothing is executing. Fork IPC ordering already puts the invalidation ahead of any later request, so a busy child owes no ack here, and the 130s/15s request deadlines still catch a wedged one. The start-retry predicate moves to the state module to stay under the line cap, matching the shape the relay client already uses. * fix(ai-vault): report a failed local scan as a host issue A local-scope scan let its error escape to the renderer, which paints it over the session list. Service supervision now produces those errors, so "AI Vault service restart circuit is open." replaced the list. Route local scope through the degradation the all-hosts leg and every SSH leg already use, so it lands as a retryable host issue row instead. Same result shape either way, so no IPC or wire contract changes. * test(ai-vault): cover the relay restart circuit transitions The relay policy shipped without tests. Pin both circuit edges, the aging-out case, the forced-refresh reopen the relay has and the desktop does not, and the backoff schedule. * fix(ai-vault): keep the OpenCode roots in the service child env The scrubbed allowlist dropped XDG_DATA_HOME and OPENCODE_DB, which the child reads to locate the OpenCode store and database. The pre-PR worker thread inherited them, so a user who sets either lost every OpenCode session. * test(ai-vault): anchor the service spawn env assertion |
||
|
|
17cfc968cf |
Revert the terminal IME composition-ownership change (#13282)
* Revert "test(ime): restore coverage the composition-ownership change removed (#13168)" This reverts commit |
||
|
|
17b3dff3c4 |
refactor(terminal): return IME composition ownership to xterm (#13128)
* fix(terminal): return IME composition ownership to xterm * fix(mobile): derive terminal input from native replacement ranges * test(mobile): record iOS Japanese IME traces * fix(mobile): preserve native IME replacement ranges * fix(xterm): flush queued application input after IME commit * test(terminal): pin Korean intermediate commit * test: pin Windows IME shortcut ownership * test: replay IBus number candidate commit * fix: preserve native macOS input-method punctuation * refactor(terminal): remove stale mac focus override * fix(mobile): preserve soft keyboard deletion ranges * fix: keep IME-owned palette chords in renderer * fix: stop carried IME shortcuts at renderer owner * fix: preserve carried IME shortcut dispatch * fix: narrow main-owned shortcut actions * test(mobile): pin Japanese IME replacement traces * test(terminal): retain paired native IME trace * fix(chat): preserve browser IME composition ownership * fix(chat): retain macOS IME confirm gesture * fix(chat): expire unmatched IME confirm carry * fix(chat): isolate IME confirmation expiry * fix(chat): retain active IME confirmation * refactor(terminal): remove dead composition handler * feat(ime): add shared Enter-ownership seams for CJK composition The confirming Enter of a CJK composition arrives as two keydowns and the orderings differ by platform: Windows/Linux redispatch the unmarked Enter/13 before keyup, macOS delivers keyup first. A guard reading only isComposing or keyCode 229 misses the redispatch, so surfaces submitted on a confirm. Adds useImeEnterGestureOwnership (carry token, next-frame expiry), a shared ImeEnterGuardedForm for native implicit submission, and the cmdk seam covering 18 CommandInput surfaces at one site. A chorded Enter arms the carry but is never swallowed — the reverse would eat a user's deliberate Cmd/Ctrl+Enter. Both failure modes are pinned by ime-enter-gesture-ownership-contract.test.ts. Co-authored-by: Orca <help@stably.ai> * refactor(terminal): consolidate native input listeners and parked-screen owner Extracts the shared native-input listener installer and renames the parked-screen detector for what it actually does, replacing per-call-site duplication. The listener installer keeps a forgetOptionKeyLocationOnBlur flag so per-window semantics are preserved rather than flattened. Net deletion; no behaviour change intended. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin recorded IME shapes as regression tests Nine regression tests built from hashed affected-platform captures, each with a paired ordinary negative and a discriminating mutation verified to take the file from all-passing to exactly one failure. Covers the Windows MS-Korean Shift family (#12179, #11878, #12151, #11946, #12152) and the Korean TUI line-break rows (STA-3237, STA-3222, STA-3129). STA-3237 pins the empirical 3-Shift / 2-active-composition / 2-newline ratio the device run established — the third Shift produces nothing because Space has already committed. That ratio is not derivable from a static capture. Co-authored-by: Orca <help@stably.ai> * fix(ime): guard Enter-commit surfaces against CJK confirm Applies the Enter-ownership guards across the surfaces whose Enter commits something: publishes, clones, pairs, installs, posts, or persists. Tiered deliberately rather than uniformly. Irreversible and remote-effect sites take the carry token, which also blocks the unmarked redispatch. Locally reversible sites take the oracle check with a one-line comment naming the residual, because a spurious commit there costs one undo. Three numeric fields are left unguarded with the reason in-code: Chromium blanks number inputs at compositionstart, so a confirm-Enter only ever reaches an empty-draft reset. Measured with a CDP probe rather than assumed — a guard that cannot fire is noise. Co-authored-by: Orca <help@stably.ai> * test(ime): teeth-check the Enter guards on every guarded surface One suite per guarded surface, each verified by deleting the guard and confirming the test fails. A green guard test without that check is unverified, not verified. Two shapes pass vacuously in happy-dom and are avoided here: native implicit form submission never fires, and blur() is inert on an unfocused element. Both made "the commit did not happen" assertions pass with the guard removed, so the suites assert the guard's contract directly instead. Co-authored-by: Orca <help@stably.ai> * fix(mobile): keep iOS Korean commits whole through the live-input path iOS Korean reports isComposing: false on every event, so it bypasses the composition guard entirely. The strict owner rejected UIKit's transformed post-change field and sent only the leading jamo — the reported symptom. Prefers the authoritative same-event field text over the predicted text when the supplied operation cannot produce it. Generic: no Korean special-case, no locale classifier, no normalization. Adds the RN-target-keyed submit carry alongside it. Co-authored-by: Orca <help@stably.ai> * test(e2e): make IME capture harnesses fail loudly instead of silently Four instruments recorded silence as success, so a void run scored as a clean one: - readTerminalImeBoundaryTrace returned an empty trace when the probe never installed, making every "nothing leaked" negative pass vacuously - summarizeLatencies([]) returned a perfect zero distribution that passed all three latency thresholds - the macOS Vietnamese spec pinned an input-source ID that does not exist, and failed as though the operator had chosen the wrong source - the expectedLineCount=1 prefix property was undocumented and one edit from silently downgrading a PTY assertion Input sources now resolve by enumeration and name the near-matches on failure. Co-authored-by: Orca <help@stably.ai> * test(terminal): cover Cangjie cancellation and fix a cross-namespace assertion Adds #11951's recorded Cangjie cancel shape to the existing cancellation suite, which covered Pinyin and Sogou but not Cangjie. One keystroke then Backspace arriving as deleteContentBackward with data: null, so the stale preedit is the only thing a fallback could replay. Verified against the historical pre-6cd944c62b3 bundle: the positive fails with ['尸'] where [] is expected, while the ordinary negative stays green. Also fixes the Vietnamese spec, which asserted a TIS-space input-source ID against getKeyboardInputSourceId(). Those two Orca APIs report the same source in different namespaces — TIS nests it under VietnameseIM, the app API does not. The resolver stays as an installation precondition; the assertion matches the leaf. Co-authored-by: Orca <help@stably.ai> * test(e2e): add a real-IME macOS arm for the Korean chord commit The existing korean-ime-terminal-shift-enter-commit spec synthesizes composition over CDP: Input.imeSetComposition sets the preedit directly and Input.insertText performs the commit. Asserting the IME produced events you injected yourself is circular, so that spec cannot certify real-IME behaviour. This arm selects 2-Set Korean via TIS, reads it back live, and injects through System Events key codes, so the OS owns the preedit, the commit instant, and isComposing. PTY byte expectations are preserved verbatim. Covers 2 of the original 4 cases by design. The other two are the Windows/Linux redispatch-before-keyup ordering, which macOS cannot produce and which cannot be selected -- the OS decides it. Reintroducing synthesis to "restore coverage" would reintroduce the circularity. Co-authored-by: Orca <help@stably.ai> * test(e2e): assert the macOS chord arm at the PTY boundary, not the renderer The byte expectations were transcribed from korean-ime-terminal-shift-enter-commit :364/:383, which assert against onData -- a renderer boundary where the terminator is CR. This spec reads the PTY child, where the tty has already converted CR to LF. Names both forms per row rather than swapping the constant, so the conversion reads as evidence that the capture reached past the renderer, as #11936 and #11951 record. Ctrl+Enter's CSI-u sequence is unaffected and is identical at both boundaries. Co-authored-by: Orca <help@stably.ai> * test(e2e): measure composer-to-onData latency and stop dropping IME keystrokes Two defects in the echo latency probe. It hooked onWriteParsed and onRender but never onData, so it measured key->parse->render echo rather than the composer-vs-onData delta the latency rows need. Adds a third hook feeding its own sample set. And `event.key.length !== 1` silently dropped IME keystrokes: Pinyin and Cangjie keydowns arrive as key:'Process' (length 7). Replayed over the captured corpus, the old filter accepted 580 of 4137 Chinese IME keydowns -- it was discarding 80% of them. The new filter matches the shape the owner itself branches on. Attribution charges each onData to the latest keydown rather than a FIFO head, because composing jamo emit no onData at all and a queue would credit a whole composition to its first keystroke. The consumer now asserts sample count before any percentile, so a zero-sample run cannot render as a flawless distribution. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin the WSL shifted-jamo newline shape for #11919 In Korean 2-set, Shift types ordinary letters -- the double consonants and the compound vowels. Each such keystroke reaches Chromium as key='Process', keyCode=229, shiftKey=true. The v1.4.163 classifier matched exactly that pattern with no code guard, so it called those keystrokes Enter, rewrote them to a synthetic Shift+Enter, and injected a newline into the middle of the word -- with no Enter key pressed. That is why the reporters said "no modifier key pressed": they had not chorded Shift+Enter, but they had pressed Shift, to type the double consonant. Asserts the row's own recorded capture: 40 immediate keydowns, exactly 3 of them Shift-carrying inside a single syllable, and an onData stream with one newline per Enter press and none mid-word. Two ordinary negatives keep it from being a blanket mute -- the same session's non-IME keydowns still reach shortcut policy, and an ordinary Shift+Enter still resolves through the real policy. Co-authored-by: Orca <help@stably.ai> * test(terminal): pin the composition commit lag that made Korean type one behind macOS Korean 2-Set commits syllable N only when the first jamo of N+1 arrives, so compositionend and compositionstart land in the same task. A composition-start handler cancelled the pending finalizer that was the only path to triggerDataEvent and ended the session without emitting bytes, so every committed syllable reached onData exactly one syllable late and the backlog cleared only at a Space or Enter. Types continuously with no Enter and no Space -- either would flush the backlog and hide it -- and samples onData at every syllable boundary. Paired with a length-matched ASCII arm that stays green throughout, so the positive is a fact about composition rather than about timing in general. Bisected to a single call site across five builds: pristine, 1.4.155 and 1.4.162 pass, 1.4.163 fails, removing the one call repairs it, restoring it fails identically. That window is exactly the reporter's "started immediately after updating". Co-authored-by: Orca <help@stably.ai> * test(mobile): cover the send-queue abort that silently drops queued keystrokes One failed send in use-terminal-live-input-commit aborts every keystroke queued behind it, with the error swallowed by .catch(() => false). The existing test resolves(true) on every send, so the failure branch was uncovered. Four arms: the abort itself, an ordinary negative on the healthy path, a throwing sender, and a liveness control proving the queue recovers once the chain settles. Deleting the abort takes 4 passed to 3 failed, with the ordinary negative correctly surviving. Scope is stated in the docblock: this is a transport send-queue abort, reachable only via a real disconnect or RPC error. REQUEST_TIMEOUT_MS is 30s, so latency alone cannot reach the branch — consistent with #7094's symptom class, not proven to be its cause. * test(terminal): pin that daemon snapshot/restore cannot disturb a composition Two independent reporters attributed broken Korean composition to the always-on PTY daemon repainting terminal state over the preedit. The attribution is wrong on ancestry — the daemon shipped three months before the version both call good — but the boundary was never actually tested. Runs the real applyMainBufferSnapshot choreography against a live composition, including the full 2J/3J/H wipe plus the resize and alt-screen branches. textarea.value, selectionStart/End, compositionView.textContent and .active all survive byte-identical, and interleaving a restore between every jamo of 문제 still commits 문제 at onData. Also pins that the uncommitted preedit is absent from the captured snapshot: it lives in the textarea, never the buffer, so a restore has nothing stale to echo back. Injecting one textarea.value = '' into the restore fails exactly the three restore-boundary tests. * test(terminal): pin that Cmd tears down a composition where Ctrl and Shift do not xterm's composition keydown exempts only keyCode 16/17/18 (Shift/Ctrl/Alt) plus 20/229. macOS Meta — 91/93/224 — is absent, so a Cmd press mid-composition takes _finalizeComposition(false): the overlay goes dark and never recovers, because compositionstart is not re-fired. The user composes the rest of the word blind. Linux and Windows users press Ctrl and are exempt. xterm already has a Meta-aware modifier predicate in wasModifierKeyOnlyEvent, so this is an internal inconsistency rather than a deliberate choice. Owns no reported row and is version-neutral: 5/5 on both 1.4.162 and 1.4.163. The branch is unexercised in all 328 recorded traces, so this is a hazard pin, not a regression guard. Only the teardown is asserted; the likely duplicated commit needs a compositionend the IME kept alive across the Cmd, which no capture contains. Deleting the exemption fails exactly the three paired negatives; adding Meta to it fails exactly the two Cmd arms. * test(native-chat): characterize preedit loss when a question card replaces the composer An AskUserQuestion card fully replaces the composer by design, but the in-flight composition goes with it: the composer unmounts before compositionend reaches it, so the preedit is never committed to the draft. The committed text survives only because the draft is cached and restored via defaultValue. Node identity changes, value 'abc' is preserved, the 가 is gone. Drives the real NativeChatView -> SessionGate -> InteractiveCard -> questionActive swap -> Composer -> ComposerField, flipped by writing the same store field an AskUserQuestion hook event writes. Flipping questionActive to false fails exactly this test and nothing else across 639 native-chat tests, so the path was entirely unguarded. CHARACTERIZATION TEST: it asserts the loss. Fixing the defect — committing the preedit before the swap, or keeping the composer mounted — will make this file fail. Update the expectations to the new contract rather than working around them. Owns no reported row. #12118/STA-3219 flicker is keyed to token counters, which provably do not remount, and a question card arrives once per question. * test(terminal): pin the duplicated commit when Meta interrupts a composition _finalizeComposition(false) sends textarea.value.substring(start, end) but cannot clear the IME-owned textarea, so a later compositionend re-sends the same range. Meta reaches that path because CompositionHelper exempts only Shift/Ctrl/Alt; xterm's own wasModifierKeyOnlyEvent covers Meta four ways, so the omission is an internal inconsistency rather than a choice. Companion to the modifier-exemption guard, which deliberately pins only the overlay teardown. This pins the data consequence. HAZARD PIN: owns no reported row. The trigger is unverified on hardware — no capture in the corpus contains a Meta-during-composition gesture, and whether macOS keeps the composition alive across it is unmeasured. The duplication follows from the code given that sequence; whether users reach the sequence is the open half. An earlier premise that Space (keyCode 32) reaches this path was refuted by a corpus scan: 0 of 731 evidence files carry a keyCode-32 Space while composing, against 171 at 229, and 229 returns early. * test(terminal): characterize the syllable lost when the textarea blurs mid-composition CoreBrowserTerminal._handleTextAreaBlur clears the helper textarea unconditionally — "Text can safely be removed on blur" — while CompositionHelper._finalizeComposition reads the committed text back out of that same value from a deferred timeout. By the time it runs the value is empty, the substring is '', and triggerDataEvent never sees the syllable. xterm checks composition state in _syncTextArea and omits the same check here. Six cases. Blurring mid-composition loses the syllable in every ordering, including compositionend-before-blur, which is Chromium's real order — so it is not an ordering artifact. A bare textarea.blur() with no Orca code loses it too, which places the owner upstream: Orca's unguarded release on outside pointerdown is one trigger, not the cause. Committing 한 then blurring mid-가 yields ['한'] where ['한','가'] is correct: one syllable gone, surrounding text intact. Teeth checked by inverting — adding an Orca-side composition guard flips exactly the three cases that route through the release path and leaves the bare-blur and no-blur cases green, which is the scope split: a fix in regular-terminal-focus-ownership alone would not close this. HAZARD PIN, but unlike the others this one has a real production injector — clicking outside the terminal mid-composition. Owns no reported row. The shape matches #9738's report; the injector does not, and a shape match with a mismatched injector is not an owner. * test(terminal): say which arm the STA-3237 fixture came from The recorded keydowns are wave 4's A-shift-unmarked-only — the arm that emits no PTY bytes. Nothing in the file said so, so two readers concluded the row's events fail the owner's predicate and that STA-3237 and STA-3222 were different defects. They share an owner; the arm that fires is Process/229+Shift, absent from this bubble-phase trace because the owner claims it in the capture phase. Also corrects "code-blind": the v1.4.163 policy emits \x1b\r only for a shift-only key:'Enter', and a jamo keydown reaches that branch solely via the isTerminalImeProcessEnter rewrite. The mock is deliberately wider so the ownership guard stays under test if that rewrite moves. Comments only — no assertion, fixture value, or mock behaviour changed. * test(e2e): track the input-source selector the macOS specs shell out to Five tracked macOS IME specs ran `swift .tmp/select-input-source.swift`, a file that is gitignored and existed only on one machine. Anyone else checking out the repo — or the same machine after .tmp is cleaned — could not run them, and they are the capture drivers for the macOS rows that are blocked waiting for exactly those runs. Moves it to tests/e2e/ beside its callers. The chord spec now resolves it from __dirname rather than reaching two levels up into .tmp. * test(terminal): pin the CJK repaint decision against the reporter's own output #12164 comment 1 and #5921 report agent output with double-width glyphs rendering duplicated character-by-character while ASCII in the same line stays clean. No IME, no composition, no keystroke — the user never types the CJK. Segmenting all three verbatim samples into maximal same-risk-class runs gives 33 runs and zero violations of "this run is corrupted iff the production detector flags it": 17 wide runs all corrupted, 16 narrow runs all byte-identical. The paired negative is co-located in the same line rather than in a separate run — the reporter supplied it without knowing. Doubling is asserted as present, not uniform: 자바스크립트 and 시스템 each leave a jamo undoubled, which is a repaint-region boundary artifact rather than a per-character transform. The discriminating arm is in the test rather than a source mutation: |
||
|
|
06780260c0 |
test(remote-runtime): run an old client and an old server against current code (#12682)
Mixed versions are the normal state of the remote-server feature: users update clients and servers independently. Until now nothing tested that. Every cross-version claim was made by code reading plus unit tests with hand-written old/new shapes — enough to catch design problems, not enough to catch a real skew regression. This runs the REAL protocol implementations from two builds against each other in one process: the actual host methods and RPC dispatcher on one side, the actual renderer multiplexer on the other, with a transport that reproduces the production asymmetry — each side decodes with its OWN codec and drops frames whose opcode it does not know. A frame survives only if the RECEIVING build understands it, which is what makes this level sufficient without launching two apps. The old side is a genuine checkout extracted from the release tag; the extracted client was confirmed to lack a symbol that exists only on main. Journey: subscribe, first snapshot, input reaching the process, live output, hide/reveal snapshot, transport drop, resubscribe, input landing again — across old->new, new->old, and a current/current control. Every step ends on an observed-state barrier; no sleeps. The oracle asserts the recorded step list, the exact 16-frame named sequence, negotiated capabilities, the exact input the host wrote to the PTY, rendered content, and zero decoder-rejected frames. A host method the stub lacks is recorded by name and asserted empty, so a harness gap cannot masquerade as a wire break. Detection is proven per violation shape, and it attributes each to the correct side: an unnegotiated opcode goes red only where a decoder would reject it, a removed published field goes red only where an old client consumes it, and a legal additive field stays green in all three pairings so the harness will not cry wolf on safe changes. It also documents the three compatibility rules in docs/reference/remote-wire-compatibility.md, linked from AGENTS.md, since they previously existed only as folklore — notably that "decoders reject unknown opcodes" is true for the desktop decoder but NOT for mobile, which silently drops them. Deliberately scoped: terminal stream only. The session-tab sync channel is not covered, nor agent-session publications, file/Git RPCs, mobile E2EE framing, or the relay transport. Two version points, so a regression introduced and reverted between them is invisible. CI selection was verified rather than assumed — `vitest list` confirms 0 matches under the shard's exclude and 4 under the dedicated job — because a lane silently running zero tests is precisely how a host-side defect escaped CI earlier in this series. Closes STA-3469. |
||
|
|
ed7849eb7b |
fix(worktrees): stop silently switching existing Windows setup scripts to Git Bash (#12406)
* fix(worktrees): stop silently switching existing Windows setup scripts to Git Bash #6967 derived the Windows setup-runner shell from `terminalWindowsShell`. On upgrade, any Windows user whose terminal preference resolved to Git Bash had their existing `orca.yaml` setup script (and issue command) handed to bash instead of cmd.exe. Scripts authored against the cmd runner — `copy`, `xcopy`, `set VAR=value`, `if errorlevel 1`, `%VAR%`, backslash paths — broke with no migration and no warning, and the failure looked like Orca broke the project. The conflation is also wrong in the steady state: a terminal preference is per-user, so two people on the same repo got different interpreters for the same orca.yaml and no project could write a setup script that worked for all of its Windows contributors. The interpreter is now a property of the script, declared the standard way: a leading `#!` line. Native Windows keeps the historical `.cmd` runner unless the script declares a POSIX shell, so no existing script changes behavior. `resolveSetupRunnerShell` keeps its role as the feasibility gate — a bash runner still requires the terminal to resolve to Git Bash, since the launch command is typed into that shell and uses MSYS `/c/...` paths. `buildWindowsRunnerScript` now drops a leading `#!` line rather than `call`ing it, so a declared-bash script that falls back to cmd (Git Bash missing) fails on a real setup line instead of aborting on errorlevel at line one. WSL worktrees, POSIX platforms, and SSH hosts are untouched. * fix(worktrees): keep the cmd setup runner launchable from a Git Bash pane Adversarial review of this PR found that pinning the runner format per script reopened issue #6896 one layer down. - `WorktreeSetupLaunch.shell` had been redefined to mean "the format the runner file was written in". `resolveSetupRunnerCommand` consumes it as "the shell that types the launch command", so a Git Bash terminal with a batch setup script produced `cmd.exe /c "C:\...\setup-runner.cmd"` typed into a bash pane, where MSYS rewrites the `/c` switch into a drive path: cmd opens interactively and setup never runs. `shell` is the terminal's family again; the runner file's .cmd/.sh extension carries the format, and a batch runner launched from a POSIX pane reuses the existing PowerShell ProcessStartInfo launcher. - The cmd runner dropped a leading `#!` line and ran the rest as batch, so a bash script reaching cmd (PowerShell/cmd terminal, or any SSH-to-Windows host) got its interpreter-agnostic prefix executed before failing mid-way. It now prints why and exits 1 without running anything. - A `#!` line's option flags were discarded: `#!/usr/bin/env -S bash -euo pipefail` lost pipefail because the runner is launched as `bash <path>`. The generated posix runner now replays declared flags via `set` and drops the duplicate interpreter line. - Docs cover the per-user setup command in repository hook settings, which goes through the same `#!` rule, and describe what the `#!` line does and does not select. Tests: composed launch command for a POSIX pane + cmd runner (hooks, shared runner command, setup sequencing gate, observed-setup signal), the cmd runner's shebang refusal, and shebang flag replay. Each fails with the source reverted. * fix(worktrees): replay only real `set` flags and keep the gate in the pane's shell Two round-2 review findings: - `#!/bin/bash -l` replayed `set -l`, which exits 2 and aborted the runner under its own `set -e` before a single setup line ran (all platforms). Only the flags `set` documents are replayed now; a bare `-o` with no option name is dropped instead of dumping the shell-option table. - The wait-for-setup gate picked its language from the runner file, so a batch runner launched from a Git Bash pane got the PowerShell gate while the agent startup command was already POSIX-quoted — `Invoke-Expression` cannot parse `'\''`. The gate now follows the pane; the runner still launches through the ProcessStartInfo launcher, never through bash. --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
5c0195af64 |
Bound remote watcher fan-out and defer File Explorer refreshes (#11908)
* batch remote watcher events and defer File Explorer refreshes Remote filesystem watcher events now batch with the shared 150ms trailing and 500ms max-wait window, coalescing per-path like local events. File Explorer tree and directory refreshes are scheduled with debounce and transport-aware concurrency caps (16 local, 8 runtime, 4 SSH). Stale directory cache tracking prevents trusting collapsed listings skipped by full refresh; they are re-read on re-expansion. Relay implements a 15-minute idle-only grace cap for zero-PTY relays via PTY pool lifecycle tracking, independent of explicitly configured grace time. * fix(watch/relay): bound remote watcher fan-out and read the live relay grace Three P1 fixes from the SSH/remote freeze audit: - Remote watchers now debounce on the same 150/500 window as local ones (finding D), and every teardown path drops the trailing flush timer instead of letting it fire into a dead watch. The deferred send is wrapped so a frame disposed mid-window can't escape as a fatal main-process exception. - File Explorer refreshes are scheduled and concurrency-capped rather than fanned out unbounded over expanded dirs (finding C). Local transports use a zero window, since main already coalesced the burst. - relay.startGrace reads ptyHandler.configuredGraceTimeMs instead of the launch-time argv closure, so a grace raised after launch is honored. The branch selection moves to relay-grace-branch.ts because relay.ts has no exports and calls main() at import, making it untestable. Consequence: a host-sleep relay holding zero PTYs now exits after the idle cap. Pinned by test and documented in docs/reference/relay-grace-time-reconfiguration.md. Also drops the duplicated 150/500/5000 constants in the runtime-RPC batcher in favor of the shared window module. * docs(relay): correct grace-reconfiguration line numbers after the relay.ts edit Co-authored-by: Orca <help@stably.ai> * refactor(file-explorer): use useMemo for paths; remove relay reference Replace manual ref-based caching with proper React hooks for content-stable path memoization. Remove outdated relay grace-time reference documentation from code review cycle. * rm design doc * fix(remote-watcher): prevent stranded timer after close An in-flight provider receive can land after the batch is torn down. Without a guard, pushing events to a closed batch would re-arm a timer that would never be cleared, stranding the task indefinitely. Track the closed state and skip pushes after close(). Relay.ts comment clarifies why pool watches remain registered during grace-period shutdown deferral — the socket server stays listening so a reconnecting client can cancel the grace and resume. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
ad1e58d966 |
chore: declutter top-level repo layout (#11890)
Remove one-off incident docs and committed test-results noise, move dev/repro/bench tools under tests/tools, and relocate i18next config into config/ so the GitHub root scrolls to the description faster. |
||
|
|
49cfbf014c |
fix(skills): stop OS sidecars marking an untouched skill as modified (#11471)
* fix(skills): stop OS sidecars marking an untouched skill as modified Package identity compared a live user directory against a tree read from a clean checkout, so anything the OS deposited counted as drift. One Finder visit writes .DS_Store, which sorts before SKILL.md and misaligns the index-aligned snapshot comparison — the copy became 'unrecognized', was reported as "may be modified... Remove it", and left out of the update. Running the update could not clear it either: the updater compares its lock to the source and never reads disk, so it correctly reports "up to date" and writes nothing. Ignore OS-authored names on both sides of the comparison. The generator half is not hypothetical: a stray sidecar in a working tree made the committed artifacts read as stale, failing lint for that developer. Scoped to OS-authored names only. Tolerating unexpected files in general would let an injected payload ride along beside a clean SKILL.md; these are safe because an official SKILL.md never references them, so no agent can be routed into one. Mode bits are deliberately untouched — that would weaken identity for real scripts. * fix(skills): keep guarding a directory or link wearing an OS metadata name The name-only skip dropped any entry matching an OS metadata name, so a directory named .DS_Store or ._scripts took its whole subtree out of identity and a symlink wearing one stopped tripping the link guard — a skill hiding either read as pristine. The OS writes these as plain files only, so the entry type decides, still ahead of the case-fold map. Also compares both walkers over the same fixture: an asymmetric skip is worse than none, since one side would bake in content the other can never observe. * chore: ignore the OS metadata names skill identity already skips Both skill-identity walkers ignore these names, but .gitignore covered only .DS_Store and Thumbs.db — so a stray ._SKILL.md showed as untracked and `git add -A` could commit it. That is the one way the two walkers can disagree: the disk walker skips such a file while the git-tree producer (collectGitPackageFiles, used by the unreferenced --rebuild-from-tags path) does not, so a committed sidecar would make released history and observation describe different content. Ignoring them keeps that asymmetry unreachable rather than adding a second skip to the released-history path, which is load-bearing and provably never sees one today: no committed sidecar exists on any ref. Nothing tracked matches the new patterns. * chore: correct the skill-identity ignore comment The previous wording claimed these names cannot be committed, which overstates what .gitignore provides: `git add -f` and `git apply --index` both bypass it, so a cherry-pick, rebase or fork branch already carrying a sidecar is unaffected. That clause was load-bearing — it was the stated reason for leaving the released-history producer unhardened — so it should not read as a structural guarantee. Also fixes the producer count (three, not two: two disk walkers plus the git-tree producer, which does not skip) and says plain file, since the skip is isFile()-gated so a directory or link wearing the name is still walked. |
||
|
|
a0944cc129 |
fix(linux): restore Ubuntu 20.04 launch — pin node-pty glibc symbols + add glibc/libstdc++ packaging gate (#9902) (#10019)
* fix(linux): restore Ubuntu 20.04 launch by pinning node-pty glibc symbols (#9902) The bundled node-pty pty.node is compiled from source in release CI on ubuntu-latest (glibc 2.39). glibc's 2.32-2.34 libpthread/libutil merge relocated openpty/forkpty (GLIBC_2.34) and pthread_sigmask (GLIBC_2.32) into libc under new symbol versions, so the from-source build bound to versions absent on Ubuntu 20.04 (glibc 2.31). The main process imports node-pty at startup, so the app crashed on launch. pty.node is the sole blocker (Electron needs GLIBC_2.25; other native modules <= 2.17). - Patch node-pty: a .symver shim pins the 3 symbols to their pre-merge version (GLIBC_2.2.5 x64 / GLIBC_2.17 arm64), and Linux-only ldflags force libutil.so.1/libpthread.so.0 back into DT_NEEDED. Guarded to Linux; macOS/Windows untouched. - Add a packaging gate (verify-linux-glibc-floor.cjs, afterPack): reads each bundled native binary's objdump -p version needs and fails the Linux build if any strong GLIBC_/GLIBCXX_/CXXABI_ node exceeds stock Ubuntu 20.04 (glibc 2.31 / GLIBCXX_3.4.28 / CXXABI_1.3.12). Catches GLIBC_ABI_DT_RELR, rejects GLIBC_PRIVATE, skips weak needs, fail-closed. - Docs + tests; the lazy sherpa-onnx speech prebuilt (GLIBCXX_3.4.29, never loaded at launch) is a documented libstdc++-floor exemption. * fix(linux): assert DT_NEEDED provider deps in the glibc-floor gate Harden the packaging gate (flagged in adversarial re-eval): the version-floor check alone can false-pass if the patch's forced `-l:libutil.so.1` ever silently drops — the pinned openpty@GLIBC_2.2.5 still resolves from libc's compat alias at build time, but fails to load on Ubuntu 20.04 where openpty/forkpty live only in libutil. The gate now also asserts that any binary importing openpty/forkpty keeps libutil.so.1 in DT_NEEDED. Validated on a real symver-pinned .so with libutil dropped (now fails) vs. present (passes). Documents the recommended real-host smoke-test follow-up. |
||
|
|
1def694e80 |
Add native chat skill picker with host-aware discovery (#9480)
* Add native chat skill and command picker with host-aware discovery Adds a unified, keyboard-first skill and command picker to native chat that: - Uses agent-native invocation syntax (slash for Claude/OpenClaude/Grok, dollar for Codex) - Discovers skills only on the pane's execution host (local, WSL, SSH-unavailable, or runtime) - Groups or separates commands and skills per agent configuration - Deduplicates by canonical path but preserves visibility through all contributing roots - Handles IME composition, loading states, and errors without claiming PTY-level control - Records picker telemetry (open, item accepted, send classification, discovery outcomes) - Extends shared agent profiles to define per-agent skill grammars and source ownership * Remove obsolete reference and design documentation Clean up stale design specs, implementation plans, and investigation notes from docs/reference/. These documents predate the current implementation and are no longer actively maintained or referenced by the codebase. * Extract shared skill discovery utilities and add skill invocation envelo - Move skill comparison and source classification to shared module for native/WSL reuse - Extract display text sanitization to prevent control/zero-width character spoofing - Add native-chat command envelope parser and surfacer for skill invocations - Extend discovery timeout backstop to account for WSL metadata read sequence * Localize skill picker UI for Spanish, Japanese, Korean, Chinese Translate skill picker UI strings including commands, skills, loading states, error messages, and scope labels for the new skill picker feature across four language locales. * Fix skill picker bugs and improve code robustness - Fix i18n plural handling: rename `count` to `sourceCount` to prevent unintended plural-key resolution in localized strings - Fix skill discovery array mutations: copy `root.providers` to prevent bugs during dedup merge - Fix image attachments being silently dropped when message text starts with /skill or agent prefix - Extract `quoteBashString` utility for WSL command code reuse across builders - Add line-separator safety characters (0x2028/0x2029) to skill display filter - Remove stale doc reference links and clarify inline comments * Add reference docs for git compatibility and headless Linux server setup Track previously untracked operational guides in `docs/reference/` that explain Git binary compatibility requirements across host types and how to run `orca serve` on headless Linux. Update AGENTS.md and README.md to link to these references. |
||
|
|
8ac6d79abf | chore: drop local notes/ and pr-evidence/ from the tree (#8677) | ||
|
|
e84a8ddec9 |
Terminal performance initiative: pipeline fixes + term-speed-2 revival + PTY flow control (integration branch) (#7214)
* Skip legacy hidden skip grammar assertions
* Fix hidden TUI snapshot test setup
* Fix sleep wake history test contract
* Fix hidden delivery startup gate helper
* Fix hidden Latin skip branch predicate
* Fix hidden synchronized split-boundary replay
* Stabilize remote runtime mixed subscription test
* Keep hidden startup query parser active during window
* Stabilize raw emoji golden restore width
* Stabilize raw emoji golden fixture completion
* Keep terminals responsive under agent output load
* Add frozen-terminal repro harness and silent-drop regression tests
Investigation harness for the frozen-terminal reports (Discord
#performance, issue #2836): pane shows content, shell alive, daemon
output.log flat while typing.
- e2e: renderer crash -> auto-reload recovery and three restart/restore
shapes (live daemon, SIGSTOP-wedged daemon, daemon killed between
launches), each probing input at both drop layers. Post-crash phases
drive the renderer from the main process because a crashed target
severs Playwright's CDP session even though the app recovers.
- e2e helpers: layer-discriminating probes (direct pty.write vs
transport input, plus pty:listSessions ownership-rebuild revival).
- unit repro: vendored xterm 6.1.0-beta.287 WriteBuffer permanently
wedges when a sync throw escapes a write-completion callback or a
custom parser handler (xterm-write-buffer-stall.repro.test.ts).
- unit repros for both silent input-drop layers: main drops writes for
a live PTY once ptyOwnership loses the id (revived by listSessions),
and the renderer transport stays unbound after a failed connect.
- pty.test.ts: unregister every leaked SSH provider id in afterEach so
module-level provider state cannot leak across tests.
Co-authored-by: Orca <help@stably.ai>
* Harden xterm write pipeline against sync-throw wedge that freezes panes
A synchronous exception escaping xterm's WriteBuffer loop permanently
wedges that terminal: _innerWrite has no try/catch around the parse
action or the write-completion callback, the tail re-schedule never
runs, and write() only re-arms on an empty buffer. The pane stops
rendering and, if a replay was in flight, the replay guard latches and
pty-connection's onData silently eats every keystroke — matching the
field reports (Discord #performance, issue #2836: content visible,
shell alive, daemon output.log flat). Both vectors verified against
vendored xterm 6.1.0-beta.287 in xterm-write-buffer-stall.repro.test.ts.
Three layers of defense:
- Guard every write-completion callback Orca hands xterm at the two
choke points (writeForegroundTerminalChunk, writeBackgroundTerminalChunk),
with settle and onParsed guarded separately so a WebGL/renderer
failure during viewport settle cannot starve the replay-guard release.
- Guard all throwing-capable custom parser handlers (DA1, OSC 10/11,
CSI ?h/?l mode reports, OSC 52 clipboard, OSC 7 cwd), degrading a
throw to "not handled" — same escape class as
terminal-link-provider-guard.ts.
- Replay-guard watchdog: each engagement releases exactly once, from
xterm's completion or a 10s watchdog, so a lost completion (wedged
pipeline, disposed-terminal race) cannot latch the guard on a live
pane; replayIntoTerminalAsync resolves on either path so restore
chains cannot hang. Force-releases record a crash breadcrumb.
All guard trips record rate-capped crash breadcrumbs, so the next field
occurrence names the throwing stack instead of failing silently.
Co-authored-by: Orca <help@stably.ai>
* Cap unbounded terminal output buffers in main and the foreground queue
Field evidence (Discord #performance / #2836): renderer memory climbs to
~1.5 GB and terminals freeze; a force reload does not help until memory
recovers. Two unbounded buffers matched that shape:
- Main-process pendingData grew by string concatenation without bound
while the renderer could not receive (frozen, starved, mid-reload) —
main-heap bloat a renderer reload cannot clear. Now capped at 2 MB per
PTY: past the cap the buffered bytes are dropped and the entry stays
O(1) until the renderer ACKs again, then a droppedOutput sentinel is
delivered and the pane repaints from the authoritative main-owned
buffer snapshot (existing hidden-output restore path) instead of
continuing a stream with a silent gap.
- The renderer output scheduler capped only hidden-pane backlogs; the
foreground path could queue a visible pane's flood without bound when
the drain could not keep up. The 2 MB cap now applies to every
foreground enqueue branch too, with a foreground-specific skip notice.
Verified: new main-side cap test (starve → flood → sentinel → normal
flow resumes), renderer sentinel-to-snapshot-restore test, two
foreground scheduler cap tests; full pty/terminal-pane/pane-manager
suites (1981 tests) and typecheck pass.
Co-authored-by: Orca <help@stably.ai>
* Make replay-guard stall release probe-certified instead of time-based
The previous stall watchdog blindly released the input guard after 10s.
If a replay were genuinely still parsing on a starved machine, that
early release could leak xterm's auto-replies into the shell — and into
agent TUIs, where a leaked ESC reads as the user pressing Escape.
Replace the blind release with a probe: when a completion looks
overdue, enqueue an empty write behind the replay. xterm parses writes
in order, so every outcome is provably safe:
- probe parses after the replay completion ran: normal release already
happened; probe is a no-op.
- probe parses but the replay completion never ran: all replay bytes
have parsed, no further auto-replies can exist — the completion was
genuinely lost. Release + breadcrumb.
- probe never parses (bounded wait): the pipeline is wedged, and a dead
parser can never emit auto-replies, so releasing cannot leak input.
Release + breadcrumb naming the pane as needing recovery.
While the probe is pending — a slow-but-alive replay — the guard now
HOLDS instead of releasing early; that case is pinned by a regression
test.
Co-authored-by: Orca <help@stably.ai>
* Scale output backlog caps with the scrollback setting and breadcrumb drops
The 2 MB pending-output caps were flat, which risked dropping lines a
50k-row scrollback user would have retained. Both caps (main pendingData
and the renderer output queue) now derive from one shared policy:
max(2 MB, scrollbackRows x 120 chars) — 2 MB at the 5k default, 6 MB at
the 50k max. The main side reads the setting live via getSettings; the
renderer scheduler is configured where the terminal lifecycle already
reads the scrollback setting.
Every drop now records a rate-limited crash breadcrumb with dropped and
cap sizes (terminal_output_backlog_dropped in the renderer,
terminal_pending_output_dropped in main — no pty ids, session ids can
embed workspace paths). Field drop frequency and size decide whether the
cap constants need raising, replacing theory with data (#2836, #7017).
Backlog skip notices are now cap-agnostic since the limit varies.
Co-authored-by: Orca <help@stably.ai>
* Extract breadcrumb recording into a collection-safe leaf module
Playwright loads spec imports at collection time, and e2e specs import
terminal-module constants (e.g. terminal-attention.spec.ts pulls
POST_REPLAY_MODE_RESET from layout-serialization, whose chain reaches
replay-guard). The breadcrumb import added to the terminal modules made
that chain reach crash-diagnostics.ts, whose top-level import.meta.hot
and webview-registry import crash Playwright's transform
("ReferenceError: exports is not defined in ES module scope") — every
e2e shard failed at collection before running a single test.
Move recordRendererCrashBreadcrumb into crash-breadcrumb-recorder.ts
(type-only imports, no import.meta) and point the terminal modules and
their test mocks at it; crash-diagnostics re-exports for existing
callers. Full e2e suite collects again (262 tests / 94 files); unit
suites, typecheck, lint green. No runtime behavior change.
Co-authored-by: Orca <help@stably.ai>
* Add cross-terminal pipeline benchmark (DSR-fenced throughput + latency probe)
Run inside any terminal (Orca pane, iTerm2, Ghostty, Terminal.app, VS Code)
to measure its full byte path. DSR round-trip latency at idle and under a
paced agent-TUI load, plus fenced throughput over four deterministic
fixtures. The DSR fence forces 'all bytes parsed' before the clock stops so
xterm.js-class ingest queues can't flatter the result.
First piece of the terminal performance initiative's measurement rig.
Co-authored-by: Orca <help@stably.ai>
* Add terminal performance initiative plan
Working plan for the orca-performance branch: verified architecture
findings, workstreams (baselines, #7153 validation, term-speed-2 revival
with merge-scout numbers, stall fixes, flow control, rig extensions,
utilityProcess router, telemetry), benchmark protocol, sequencing, and
baseline-relative success criteria.
Co-authored-by: Orca <help@stably.ai>
* Add cross-terminal baseline results (Orca 1.4.91 prod vs Terminal.app vs Ghostty)
Headline: Orca DSR latency under 1MB/s agent-TUI load is p50 134ms / p99 292ms
vs 0.45ms (Terminal.app) and 0.21ms (Ghostty). Idle latency is fine (0.69ms
p50) — the problem is queueing under load, not the pipeline hop. agent-tui
fenced throughput: Orca 2.0 MB/s vs Terminal.app 37 MB/s, Ghostty 78 MB/s.
Co-authored-by: Orca <help@stably.ai>
* Add pipeline-loss decomposition benches (headless xterm + daemon ingest)
Both isolate layers of the 51x agent-tui gap found in baseline-jul02:
bare @xterm/headless parses agent-tui at 103 MB/s and daemon Session
ingest (emulator + pending-output recording + fanout) at 103 MB/s —
on the byte stream the full Orca pipeline delivers at 2.0 MB/s.
Parser and daemon are exonerated; the loss is in main per-chunk
processing, delivery/ACK pacing, or renderer layers above xterm.
Co-authored-by: Orca <help@stably.ai>
* Record baseline + decomposition findings in initiative plan
Co-authored-by: Orca <help@stably.ai>
* Add dev-build orca-performance bench result (confounded: dev mode, 282-col window, 3MB fixtures)
DSR under load p50 161ms — the #7139/#7150 branch does not move the
under-load latency class. Expected in hindsight: DSR replies are ordered
within the output stream, so the metric measures output-queue depth;
cooperative drain paces input responsiveness but cannot reorder the queue.
Shrinking the queue itself (producer flow control, task 6) and raising
agent-tui throughput (task 9) are the levers for this number.
Co-authored-by: Orca <help@stably.ai>
* Record dev-build #7153 check in findings log
Co-authored-by: Orca <help@stably.ai>
* Parse-clock high-priority terminal drains instead of fixed-nap dripping
Attribution (task #9): the drain loop wrote at most 2x16KB then slept
4/16ms regardless of parse speed — an isolation bench (new
pane-terminal-output-scheduler-throughput.bench.test.ts) measures that
drip at 1.9 MB/s background / 27 MB/s foreground against xterm's
~103 MB/s parse rate, matching the baseline-jul02 end-to-end numbers
(agent-tui 2.0 MB/s in prod 1.4.91).
Fix: high-priority (visible-pane) drains now re-arm on xterm's
parse-completion callback and carry 8 writes per tick; the isolation
ceiling rises 27 -> 117.6 MB/s (parse-limited). Background cadence is
deliberately unchanged (2 MB/s drip protects the focused pane; hidden
delivery is term-speed-2's job). DRAIN_TIME_BUDGET_MS still bounds
per-tick work, preserving #7139's cooperative-drain intent.
Validation: 621 scheduler/guard/pty tests green, typecheck clean.
Co-authored-by: Orca <help@stably.ai>
* Record task #9 attribution + parse-clock fix in findings log
Co-authored-by: Orca <help@stably.ai>
* Findings: 51x loss attributed to O(tail) retained-tail redraw path in main onPtyData
Co-authored-by: Orca <help@stably.ai>
* Window the retained-tail redraw path to the cursor's reach
Attribution (findings log 2026-07-03): main's onPtyData consumed ~93% of
the event loop under an agent-TUI flood, and the dominant term was
appendNormalizedToMultilineTailBuffer + finalizeRetainedTerminalRows
materializing ~2x tail-length row objects plus a per-row trailing-space
regex on every chunk — 0.888ms/chunk at the 2,000-line cap, on every
Claude-Code-shaped frame (cursor-up + erase-below).
The multiline algorithm now runs on a suffix window sized by the chunk's
maximum upward cursor excursion (plus the inherited redraw cursor and a
safety margin); the untouched prefix is shared by reference with a cheap
last-char trailing-space check to match the reference trim. Pathological
full-height cursor-ups fall back to the unwindowed implementation, which
is kept verbatim and exported as the reference for the 500-case
differential fuzz (retained-tail-redraw-window.equivalence.test.ts).
Micro-bench at a full 2,000-line tail: 0.888 -> 0.073 ms/chunk (12x).
1,415 runtime tests green, typecheck clean.
Co-authored-by: Orca <help@stably.ai>
* Add dev bench results: parse-clock and windowed-tail fixes
Co-authored-by: Orca <help@stably.ai>
* Record windowed-tail partial win + next-cycle recipe in findings log
Co-authored-by: Orca <help@stably.ai>
* Findings: remaining whale is the per-chunk blocked-reason check (~85% of onPtyData post-fix)
Co-authored-by: Orca <help@stably.ai>
* Throttle the terminal wait-blocked check off the PTY hot path
Post-windowed-tail attribution (findings log 2026-07-03): the blocked-
reason complex — two full-tail buildTerminalWaitText builds plus
toLowerCase and multi-pattern scans per chunk, existing only to stamp
waitBlockedAt — consumed ~85% of onPtyData's remaining cost (~700-790ms/s
under an agent-TUI flood).
The check now runs at a 50ms cadence over coalesced chunks (PTY chunk
boundaries are arbitrary, so coalescing preserves semantics), with a
trailing-edge timer so burst-final state is always evaluated, and an
immediate bypass when the incoming chunk (plus a 31-char split carry)
contains a prompt keyword — so actionable-prompt stamping stays
per-chunk-immediate while keyword-free flood frames skip the complex
entirely. Previous wait text is cached per pty instead of rebuilt, and
state is cleared at both pty teardown sites.
1,415 runtime tests green (including the cross-chunk prompt test, which
exercises the keyword bypass), typecheck and lint clean.
Co-authored-by: Orca <help@stably.ai>
* Findings + results: three stacked fixes unlock the pipeline (agent-tui 16x, DSR-load p50 161->18.8ms in dev)
Co-authored-by: Orca <help@stably.ai>
* Add producer flow-control design to initiative plan
Co-authored-by: Orca <help@stably.ai>
* Findings: revival branch green but perf-gated — daemon Session ingest regressed 103->40-48 MB/s (chain emulator restructure); merge blocked until blockedfix parity
Co-authored-by: Orca <help@stably.ai>
* Pre-filter daemon OSC/mouse scanners for introducer-free chunks
Skips the scan-tail copy and full-chunk walks when a chunk cannot contain
an OSC or private-mode sequence (single native includes() checks), with
split-sequence correctness preserved via explicit tail retention. Strictly
positive micro-optimization on the daemon per-chunk path; 641 daemon tests
green (1 pre-existing WSL failure unrelated).
Co-authored-by: Orca <help@stably.ai>
* Retract confounded daemon conviction; mandate load-controlled A/B protocol for the revival merge gate
Co-authored-by: Orca <help@stably.ai>
* Record A/B gate pass in findings log; add A/B result JSONs
Co-authored-by: Orca <help@stably.ai>
* Add producer-side PTY flow control (watermarks + protocol v19)
Main now pauses the actual PTY when a pane's renderer-pending backlog
crosses the 256KB high watermark and resumes once it drains below the
32KB low watermark (wide hysteresis band so a draining queue cannot flap
pause/resume per flush slice). node-pty pause() stops the pty fd read, so
the kernel/ConPTY buffer fills and a flooding shell blocks on write —
flood-induced buffered lag becomes shell blocking instead of unbounded
main-process buffering (terminal-performance-initiative §5).
Transport: new fire-and-forget pausePty/resumePty daemon notifications
(protocol v19; 18 added to PREVIOUS_DAEMON_PROTOCOL_VERSIONS), routed
DaemonServer -> TerminalHost -> Session -> subprocess pause()/resume().
LocalPtyProvider pauses node-pty directly. Router/degraded providers
forward; IPtyProvider gains optional pauseProducer/resumeProducer.
Safety invariants:
- Lost-resume failsafe: daemon Session auto-resumes 5s after a pause with
no matching resume; main re-asserts the pause at most once per 5s while
still above the high watermark, so a lost resume can never wedge a shell
and a sustained flood stays throttled.
- Resume on every teardown path: Session kill/exit/dispose/detach; main
releases on pty exit and on window-destroyed bookkeeping wipes; the
adapter owes paused sessions a resumePty on the next connect after a
socket drop.
- Providers without support (SSH relay, legacy protocol <= v18) no-op
silently, and the scrollback-scaled pending-output cap still bounds
main memory when pause is unavailable.
- Kill switch: PRODUCER_FLOW_CONTROL_ENABLED in ipc/pty.ts flips the
whole mechanism off in one line.
daemon-errors.ts is split out of types.ts to stay under the max-lines cap.
Tests: watermark transitions/hysteresis/re-assert (controller unit),
lost-resume failsafe + resume-on-kill/exit/dispose/detach (session),
notification routing + v18 gating + reconnect owed-resume (adapter),
direct pause/resume (local provider), and a flood test asserting pause
fires once, pending stays bounded at HIGH + one chunk, and resume fires
once after drain (ipc/pty).
Co-authored-by: Orca <help@stably.ai>
* Findings: flow control merged; definition-of-done accounting; prod verification re-scoped to packaged RC
Co-authored-by: Orca <help@stably.ai>
* Fix stray brace from revival merge in long-table-scroll-restore e2e spec (broke e2e transform in CI)
Co-authored-by: Orca <help@stably.ai>
* Prod verdict: v1.4.121-rc.0 bench — DSR-load p50 134->18.6ms (7.2x), agent-tui 2.0->11.2 MB/s, idle at Terminal.app parity; pipeline now cadence-bound
Co-authored-by: Orca <help@stably.ai>
* Recover terminal output delivery after system sleep
Root cause: main gates every pty:data send on a global + per-PTY
in-flight counter that only renderer ACKs decrement. If ACKs are lost
across a system suspend, the counters pin at the cap and every PTY —
old and newly created — is silently gated forever while output piles up
in pendingData. A focus-preserving display wake also fires no renderer
focus/visibilitychange events, so terminal wake recovery (and the WebGL
context-loss latch clear) never runs. Only a renderer reload recovered.
Three fixes:
- ACK-stall watchdog (src/main/ipc/pty.ts): if sends stay gate-blocked
for 10s with zero ACK progress while the renderer webContents is
alive, warn once, reset the in-flight delivery counters, and flush
held pendingData. Armed lazily on the first gate-blocked send and
disarmed by every ACK, so it can never fire under healthy heavy load.
- Renderer lifecycle reset now also zeroes the in-flight counters — a
reload/navigation destroys the renderer dispatcher, so outstanding
ACKs can never arrive and stale counters would gate the new renderer.
- System-resume wake IPC: main relays powerMonitor 'resume' as
system:resumed to live windows (plus forceRepaint); preload exposes
ui.onSystemResumed; the terminal wake-recovery hook runs the same
recovery path as window focus/visibilitychange.
Co-authored-by: Orca <help@stably.ai>
* VS Code head-to-head: Orca beats/ties 5 of 6 metrics (16x idle, 5x styles-stress, better p99); load p50 gap attributed to ACK window + timer-clamped drain cadence
Co-authored-by: Orca <help@stably.ai>
* Schedule zero-delay terminal drains via MessageChannel
Chromium clamps nested setTimeout(0) to ~4ms, stacking dead gaps onto
every parse-clocked drain tick; the explicit 4ms high-priority re-arm
interval added more. A posted message is still a macrotask — input and
paint are serviced between posts — so cooperative yielding survives
without the clamp. Generation-tokened cancellation; vitest keeps the
timer path (fake timers can't advance channel posts) plus a real-timer
smoke test for the channel path. Standing-queue target: VS Code's ~7ms
class (measured us 18.6ms, them 7.18ms, same rig).
Co-authored-by: Orca <help@stably.ai>
* Cut daemon and main PTY batch windows 8ms -> 2ms
At 9% pipeline utilization the DSR-under-load latency is fixed batching
windows, not queue depth (proved by the MessageChannel drain lever
moving nothing). Both hops charged an expected half-window per chunk;
2ms keeps burst coalescing at negligible IPC overhead (~500 msgs/s
worst case vs MB/s payloads).
Co-authored-by: Orca <help@stably.ai>
* Findings + tests: batch windows were the DSR-load gap (19->8.0ms dev); timing tests updated to 2ms windows
Co-authored-by: Orca <help@stably.ai>
* Fix PR CI and guard resume relay during shutdown
Co-authored-by: Orca <help@stably.ai>
* Chain e2e specs 6/6 green — gate x drain validation debt paid
Co-authored-by: Orca <help@stably.ai>
* Replace ack-stall watchdog with cumulative ACKs + solicited delivery resync
Design review: the 10s blind-reset watchdog decided correctness from a
wall-clock threshold. Rework piece 1 into a deterministic two-part design
(pieces 2 and 3 — lifecycle-reset counter zeroing and powerMonitor wake
IPC — are unchanged):
- Cumulative ACKs (TCP-style): the renderer dispatcher now tracks a
monotonic per-pty total of processed chars (terminal-pty-ack-gate) and
sends it on every ACK alongside the legacy per-chunk delta. Main keeps
per-pty sentChars/ackedChars and max-merges received totals — idempotent
and reorder-tolerant, so a lost ACK self-heals when any later ACK
arrives instead of becoming permanent in-flight debt. Provider
(SSH/daemon) backpressure is credited only the derived delta, clamped,
never negative. Main tolerates both payload shapes keyed by field
presence (dev hot-reload can mix renderer/main versions); totals reset
on pty exit and renderer lifecycle reset on both sides.
- Solicited resync (replaces the blind reset): when new pty data arrives
while that pty's delivery is fully gated and no probe is outstanding,
main sends pty:requestDeliveryResync; the renderer replies with its
cumulative totals and main reconciles via max-merge, then flushes held
pendingData. Event-triggered, verified-state recovery — no wall-clock
threshold decides correctness. The only timer is a 5s request/response
hygiene timeout that clears the outstanding flag and logs one
diagnostic warn per silent streak; it never mutates counters (a
renderer that cannot answer has dead IPC — reload is the only cure).
The 10s corrective watchdog is deleted.
Co-authored-by: Orca <help@stably.ai>
* Starting point: prior agent's garble differential fuzz harness
Three files recovered (were untracked) from a prior agent killed by API
outages, plus a trivial curly-brace lint fix in the op dispatcher so the
pre-commit hook passes:
- src/shared/agent-tui-ansi-fuzz-stream.ts (seeded agent-TUI byte-stream gen)
- src/shared/terminal-restore-parity-fixture.ts (renderer-parity fixture)
- src/main/daemon/headless-emulator-fidelity.fuzz.test.ts (suite 1: differential
HeadlessEmulator vs @xterm/headless reference on identical bytes)
Co-authored-by: Orca <help@stably.ai>
* Suite 1 findings: two new serialize round-trip bugs (B bold-loss, C cursor)
Scanned seeds 1..2000. Beyond the pre-documented serialize wrap-null-cell bug
(A, 27 seeds, tolerated), the fuzz surfaced two NEW real @xterm/addon-serialize
0.15.0-beta.287 round-trip defects, both of which garble a revealed hidden pane:
- Bug B (seeds 435, 770, 1321): serializing a dim cell followed by a bold-only
cell emits \x1b[1;22m; SGR 22 clears bold too, so restored bold is lost.
Minimal repro: '\x1b[2mA\x1b[22m\x1b[1mB' -> restored 'B' loses bold.
- Bug C (seeds 454, 1696): a final content row filled to the right margin leaves
xterm wrap-pending; the serializer's relative cursor restore lands one column
short. Minimal repro: '0123456789\x1b[3;5H' at cols=10 -> cursor x=3 not x=4.
Both isolated to pure serializer replay (no Orca preamble), confirming upstream.
Parity fixture verified faithful to the renderer pane's buffer options. Each is
pinned as a standalone it.skip repro; full evidence + classification in
notes/garble-fuzz-divergences.md. Seed 113 (handoff's DECSC/DECRC case) does not
diverge on the current harness. No production code changed.
Co-authored-by: Orca <help@stably.ai>
* Add perf prerelease update check modifier
Co-authored-by: Orca <help@stably.ai>
* Suite 2: hidden-reveal seq-reconciliation fuzz + two new snapshot bugs (D, E)
Property-tests the reveal seq-reconciliation byte-stitch (getChunkDataAfterSnapshot
/ reconcileChunkAgainstRestoredSnapshot in pty-connection.ts), mirrored exactly:
N=200 seeded hide/reveal scenarios with a rich agent-TUI hidden prefix snapshot
and an append-only racing tail, chunked with seq/rawLength meta, seq-domain
restarts, unmetered chunks and droppedOutput markers. Asserts snapshot-at-S +
reconciled tail == snapshot-of-everything (seq-neutral) and == always-visible
(end-to-end). Runtime ~5s at 200; FUZZ_ITERATIONS override documented.
Two NEW real snapshot-limitation garbles found while building it, both distinct
from suite 1's serialize bugs and pinned as standalone it.skip repros:
- Bug D: the DECSC saved-cursor register is not serialized. A hidden TUI that
saves the cursor (ESC 7 / CSI s) and restores it on reveal (ESC 8 / CSI u)
lands the restore at home. Repro: 'AB\x1b7\x1b[4;10HCD' + '\x1b8X' -> 'XB' vs 'ABX'.
- Bug E: a snapshot taken mid-escape-sequence (a PTY read split an escape) drops
the partial sequence (it's parser state, not buffer), so the tail's
continuation renders literal. Repro: 'AB\x1b[3' + 'mCD' -> 'ABmCD' vs 'ABCD'.
Fired on ~24% of the corpus (tolerated + counted via prefixEndsMidSequence).
The append-only-tail design isolates seq reconciliation from these and the Bug C
cursor cascade. Full evidence + fix directions in notes. No production changes.
Co-authored-by: Orca <help@stably.ai>
* Suite 3: 25-cycle park/reveal drift e2e test
Extends terminal-hidden-view-parking.spec.ts with a deterministic 25-cycle
park->reveal test on a static rich alt-screen TUI frame (box drawing, SGR
colors, wide CJK/emoji). Baselines against the frame after the first snapshot
restore (so both sides pass through identical machinery — the alt-screen restore
correctly drops normal-buffer scrollback, which is contract not garble), then
asserts every subsequent reveal reproduces it byte-for-byte with no accumulated
drift and no hidden-skip banner. Exercises the real renderer teardown +
HeadlessEmulator snapshot restore + PTY reattach path the fuzz suites model in
isolation. Passes in ~29s (electron-headless, workers=1).
Co-authored-by: Orca <help@stably.ai>
* Fix two serialize round-trip bugs garbling hidden-terminal snapshot restore
BUG B (addon patch): @xterm/addon-serialize's SGR diff emitted bold/dim set
params before the shared intensity reset 22, so "1;22" wiped a freshly set
bold and a bare "22" dropped a still-set bold/dim. Patched via pnpm
patchedDependencies (config/patches) to diff bold+dim as one intensity
group with the clearing 22 emitted first. Other flag pairs (4/24, 3/23,
7/27, ...) have dedicated resets and were verified unaffected.
BUG C (Orca-side hardening): the addon restores the cursor with relative
moves computed from where it assumes replay leaves the cursor; a final row
filled exactly to the right margin leaves replay wrap-pending and the
restore lands one column short. New shared
serializeWithAbsoluteCursor appends an absolute CUP from the source
terminal's authoritative cursor at every restore/replay serialize site
(daemon/runtime HeadlessEmulator.getSnapshot, renderer mobile snapshot
serializer, shutdown layout capture). It skips empty snapshots and
wrap-pending sources so it never changes already-correct behavior.
Round-trip repros + non-regression coverage in
src/main/daemon/terminal-snapshot-serialize-roundtrip.test.ts (verified
failing with the fixes stashed). buildRehydrateSequences extracted to its
own module to keep headless-emulator.ts under the max-lines budget.
Co-authored-by: Orca <help@stably.ai>
* Gates: tolerate+count Bugs B/C in deep mode; drop inverse from reconciliation tail
- Fidelity suite: add snapshotHasSelfCancellingBoldReset (Bug B) and
isMarginWrapPendingCursorOffByOne (Bug C) predicates so the corpus tolerates +
counts them like Bug A. FUZZ_ITERATIONS=2000 is now green (~113s) and fails
only on genuinely new divergences; each tolerance keeps its <50% degeneracy
guard. Default 300 unchanged (~17s).
- Reconciliation suite: drop SGR 7 (inverse) from the append-only tail. Inverse
marks trailing blanks with an inverse-fg the serializer round-trips slightly
differently by capture depth — a Bug-B-class serialize nuance, not seq
reconciliation. FUZZ_ITERATIONS=1000 is now green; default 200 unchanged.
- Notes updated: every bug class is both pinned (skipped repro) and tolerated in
its corpus; combined default runtime ~19s.
Regex uses String.fromCharCode(27) to stay oxlint no-control-regex clean.
Co-authored-by: Orca <help@stably.ai>
* Keep RC update checks off perf prereleases
Co-authored-by: Orca <help@stably.ai>
* Fix snapshot DECSC register loss (Bug D) and mid-escape boundary drop (Bug E)
Bug D: the serialized screen cannot carry the VT100 DECSC saved-cursor
register, so a hidden ESC 7 followed by a post-reveal ESC 8 restored to
home and clobbered live cells. The snapshot epilogue now re-saves at the
source's saved position before the final absolute CUP
(readSavedCursorRegister + serializeWithAbsoluteCursor; the active
buffer's own register, so alt screens carry theirs). Position-only by
design; never-saved terminals are left untouched.
Bug E: a PTY read ending mid-escape leaves the sequence in the emulator's
parser, so serialize dropped it and the racing tail's continuation bytes
rendered literally after reveal (~24% of the fuzz corpus). The emulator
now tracks the unparsed trailing partial at ingest
(terminal-partial-escape-tail.ts, committed post-parse like the mouse
mirror) and ships it as TerminalSnapshot.pendingEscapeTailAnsi.
applyMainBufferSnapshot writes it LAST, after POST_REPLAY reset — any
later ESC would abort the dangling sequence. Seq accounting is unchanged:
the tail is a suffix of bytes the snapshot seq already counts, so
reconcile slicing needs no adjustment.
Fuzz suites: unskip the Bug B/C repros (fixed on this branch) and the new
D/E repros; remove the B/C/E tolerance predicates so regressions fail
loudly. Only Bug A (upstream wrap null-cell) stays tolerated + counted.
Green at FUZZ_ITERATIONS=2000 (fidelity) and 1000 (reconciliation).
Co-authored-by: Orca <help@stably.ai>
* Count suffixed RC tags (rc.N.perf) in the shared rc counter — second suffixed cut collided with the first
Co-authored-by: Orca <help@stably.ai>
* Classify suffixed rc tags (rc.N.perf) as rc telemetry identity in release builds
The build-identity guard only knew vX.Y.Z and vX.Y.Z-rc.N, so suffixed
perf RCs cut fine but every platform build refused the tag and the
releases published empty.
Co-authored-by: Orca <help@stably.ai>
* Cut the hidden-restore flood feedback loop (A) + query carve-out on drops (B)
(A) Under a foreground flood, the hidden-output-restore loop re-fetched
snapshots endlessly: each synchronous applyMainBufferSnapshot starved ACK
processing, main pinned at the in-flight cap, dropped at the pending cap,
and every droppedOutput/modelRestoreNeeded marker re-armed another
restore until the flood ended (rc.7.perf DSR timeouts).
- Restore loop: a foreground live-chunk queue overflow now abandons the
restore immediately (the stream is outrunning snapshot fetch+replay),
with a 3-iteration hard cap + lifecycle warn as backstop.
- Re-arm gate: drop markers/sentinels and reconcile seq-gaps on a visible
pane during its own in-flight/just-abandoned restore no longer re-arm;
live bytes write through and ONE deferred repaint (2s after the last
backpressure signal) heals the gap. Hidden-pane gate semantics are
unchanged.
- Query salvage: discarding queued restore bytes (overflow/refetch) now
extracts DSR/CPR/DA/OSC-color queries and replays them to xterm so
replies still flow.
(B) Main-side: dropOversizedPendingPtyData carves reply-eliciting query
sequences out of the dropped buffer (and out of post-drop latched data,
bounded) and ships them on the droppedOutput sentinel, so DSR probes
survive bulk drops. Query scanning moved to
src/shared/terminal-reply-query-extraction.ts, shared verbatim with the
renderer's hidden-startup query extraction.
Co-authored-by: Orca <help@stably.ai>
* ACK terminal output at parse-drain, not dispatcher enqueue (C)
The renderer credited main's per-PTY in-flight window the moment a
pty:data chunk entered the dispatcher, so the 512KB window meant "bytes
received", never "bytes parsed". Under flood the renderer write queue
grew unbounded behind instant ACKs; main saw no backpressure, crossed
the pending cap, and bulk-dropped output (rc.7.perf DSR timeouts).
Crediting is now parse-deferred: each delivery carries a fire-once
credit (deliverPtyDataWithDeferredAck); the pane's first scheduler write
claims it (writeTerminalOutput.ackCredit) and the output scheduler fires
it when the bytes are consumed — after terminal.write in the
parse-clocked drain, or on ANY discard path (backlog cap replacement,
discardTerminalOutput, disposed-terminal drops, flush recovery).
Deliveries that never reach the scheduler (reconcile drops, restore
queueing, pre-mount eager buffer) settle at handler return, so the
invariant holds: every delivered chunk credits exactly once, parsed or
discarded. E2E ack-gate hold/release and delivery-resync semantics are
unchanged (all crediting still routes through ackPtyData).
Main-side equilibrium: with ACKs at parse cadence, in-flight becomes
true backpressure — pendingData stays near the 256KB producer-pause
watermark, far under the >=2MB drop cap, so bulk floods block the shell
(node-pty pause) instead of dropping.
Co-authored-by: Orca <help@stably.ai>
* Synthesize salvaged query replies directly instead of replaying into xterm
The 10MB dev bench proved the write-back salvage insufficient: a
pending-cap drop always triggers a snapshot restore, whose replay guard
swallows xterm auto-replies and whose discardTerminalOutput races away
still-queued query writes — the salvaged DSR died both ways and the
fence still timed out.
Salvage now answers directly on the input path (immune to both): CPR
(CSI 6n) from the live buffer via transport.sendInput, DA1 with the
renderer's canned response, OSC color probes via the existing direct
responder. Rare queries (DECRQM, DA2) keep the best-effort xterm
replay.
Co-authored-by: Orca <help@stably.ai>
* Untrack branch-added bench result JSONs (20 files); keep numbers in the findings log
Files stay on disk; main's 7 pre-existing results are untouched.
Co-authored-by: Orca <help@stably.ai>
* Branch guide: document merge-not-rebase sync strategy and conflict pattern
Co-authored-by: Orca <help@stably.ai>
* Merge origin/main (#7316 tab-strip click-vs-drag fix); adapt #7290 recovery-reload tests to this branch's dual did-finish-load listeners
The three tests grabbed the FIRST did-finish-load listener; on this branch
the renderer delivery-gate reset registers before the orphan sweep, so the
sweep tests exercised the wrong handler (one failing, two vacuously green).
They now fire all listeners like a real reload.
Co-authored-by: Orca <help@stably.ai>
* Fix branch CI lint: split pane-interaction functions out of artificial-opencode-terminal-load.spec (815>800 lines), modernize perf-html-report script
No max-lines disable per repo rules; extracted to
artificial-opencode-pane-interactions.ts. toReversed() and
import.meta.filename replace reverse()/fileURLToPath.
Co-authored-by: Orca <help@stably.ai>
* Fix Windows update-relaunch killing the live terminal daemon
On a Windows update relaunch the daemon can be wedged past every RPC
budget (final checkpoint flush + installer/AV disk pressure), so the 3s
health check AND the 5s session-list hello both time out while sessions
are still alive - and the launcher failed closed, killing the daemon and
every terminal session it owned.
- Adopt an unresponsive daemon whose pipe still accepts a raw
connection; a new rejected health state keeps replacing daemons that
answered and refused the handshake (never adoptable).
- Give Windows pid files a real startedAtMs (daemon self-reports it in
the ready IPC message) and verify it via CIM CreationDate piggybacked
on the existing command-line query, so the pid-recycling guard is no
longer inert on win32.
- Only delete legacy daemon pid/token files when the pid-file process is
provably dead; deleting a live daemon''s token made its sessions
permanently unadoptable after a protocol bump.
- Capture agent resume records every 60s in the renderer (skipping
unchanged records) so hard kills still leave a fresh resume record.
* Heal blank terminals when main→renderer push delivery dies (renderer-pull delivery watchdog)
Field evidence (v1.4.121-rc.0 debug snapshot, 2026-07-06): a wedged window
held 530,115 un-ACKed in-flight chars — one PTY pinned at the 512KiB per-PTY
high water plus a fresh terminal's 245-char prompt that was sent and never
consumed — while the user ran the snapshot over invoke from that same window.
Main→renderer push delivery (pty:data and every sibling channel) was dead;
renderer→main→renderer invoke was alive. Upstream precedent for
one-directional IPC death: electron#37067 (suspected Mojo pipe disconnect,
stalled as need-info). Every terminal goes blank, new terminals are born
blank, and only a renderer reload recovered.
The existing recovery layers cover the OTHER variants of this bug family and
structurally cannot reach this one:
- The xterm write-pipeline sync-throw guards, output-buffer caps, and
probe-certified replay-guard release (#7150 family) run only after bytes
arrive in the renderer — here they never do. (The pending cap did work as
designed in the field: ~2.1MB pendingDroppedChars, bounded main heap.)
- Cumulative ACKs self-heal lost ACK messages and the solicited delivery
resync reconciles verified totals (4647df86a; #7260 on main) — but the
resync probe, the powerMonitor wake relay, and the droppedOutput restore
markers all ride main→renderer push, the direction that is dead. The
probe's unanswered path deliberately only logs.
This adds the missing lane, renderer-initiated and ridden entirely over
invoke — the direction the field snapshot proved alive:
- terminal-delivery-watchdog.ts: 15s heartbeat, free while output flows.
Hot-path cost is one Map upsert per received chunk; a tick does no IPC
unless the terminal plane was silent for the whole interval and a PTY
still expects delivery. Two consecutive silent ticks with main reporting
ACK-starved in-flight confirm the wedge; heals are one-shot per 60s
cooldown so a persisting wedge cannot repaint-storm.
- pty:reportRendererDeliveryState (invoke): always max-merges the renderer's
cumulative processed totals (a free extra repair lane for the lost-ACK
variant); with heal:true — and only after main has itself seen ≥10s of ACK
silence — writes off bytes the renderer provably never received
(received ≤ acked < sent; a received-but-unparsed backpressure window is
never written off), drops that PTY's pendingData (snapshot covers
everything ≤ markerSeq, hidden-drop parity), credits provider flow
control, and returns restore markers in the reply.
- The renderer re-attaches all push listeners (cures a detached-listener
variant outright; a safe no-op against a dead channel) and routes the
pulled markers through the existing pty:modelRestoreNeeded machinery —
panes repaint from the main-owned buffer snapshot with zero push
delivery involved.
- Field discrimination built in: the heal warn logs
ipcRenderer.listenerCount('pty:data') (listener detached vs channel dead)
with the full delivery snapshot, so the next occurrence names the root
cause without asking the user to run anything in a console.
Repro harness: the exposeStore-gated __terminalDeliveryWatchdog hook
blackholes pty:data ahead of the dispatcher — the field failure in
miniature (no receive count, no ACK credit, no dispatch).
terminal-push-delivery-loss-recovery.spec.ts proves the wedged output
repaints while the blackhole is still engaged and live flow resumes after
release, with no reload. Unit suites pin the watchdog state machine
(zero IPC under flow, two-tick confirm, cooldown), the dispatcher reattach
seam, and the main-side write-off semantics.
Perf: nothing added to main's send/flush path; the renderer data path gains
one integer/Map update per chunk; idle cost is one ~100-byte invoke per 15s
only during total terminal silence. Terminal perf e2e suite (typing
latency, redraw freeze, output scheduler, hidden TUI restore, artificial
opencode load) passes on this change; no watchdog activity occurs under
ack-gate pressure scenarios because receive-progress gates the heartbeat.
Co-authored-by: Orca <help@stably.ai>
* Expose the hidden-yet-visible delivery-gate contradiction in the debug snapshot
The v1.4.124-rc.2.perf blank-terminal field snapshot showed a different
state than the v1.4.121 transport wedge: no delivery gating at all
(ackGatedFlushSkipCount 0, in-flight 38KB, far under every cap) but TWO
ptys hidden-delivery-gated with 78MB dropped as hidden. The aggregate
counters cannot say whether the pane the user was staring at was one of
the gated ones — the one number that separates "normal background
dropping" from "main is starving a visible pane because the reveal
unmark never fired".
Add hiddenDeliveryGatedVisiblePtyCount / hiddenDeliveryGatedActivePtyCount
(overlap of the gate's hidden set with the renderer's visible/active
reports — a contradiction that must be zero) to the delivery debug
snapshot, and a once-per-minute warn when hidden-gated bytes are dropped
for a pty the renderer reports visible or active, with the full snapshot
attached. Zero cost outside the debug read and the already-dropping path.
Co-authored-by: Orca <help@stably.ai>
* Unlatch the hidden-delivery gate when user input disproves a stuck document.visibilityState
macOS occlusion tracking can wedge document.visibilityState at 'hidden'
after display sleep and never fire another visibilitychange. The hidden-
delivery gate then keeps dropping renderer-bound bytes for panes the user
is looking at (field snapshot 2026-07-06, v1.4.124-rc.2.perf: 78MB dropped
across 2 pane-level-visible ptys with a fully healthy transport), and every
recovery path (window focus, system-resume relay, backlog recovery) re-ran
syncHiddenRendererPtyDelivery only to recompute the same stale predicate —
nothing could ever clear the gate. The user sees a frozen terminal; typing
echo is dropped in main; only a reload recovers.
Real user input while the document claims hidden is a physical
contradiction: keystrokes and clicks only reach a focused, on-screen
window. stale-document-visibility.ts latches that proof, runs each pane's
existing visibilitychange resync (gate unhide + hidden-output snapshot
restore), and hands authority back to the occlusion tracker on the next
genuine visibilitychange. No timers; the failure bias is safe — a wrong
latch can only restore pre-gate delivery cost, never drop bytes. Hot path
unchanged: the foreground predicate still returns on the same single
comparison while the document is visible.
tests/e2e/terminal-stuck-occlusion-recovery.spec.ts pins the wedge
(visibilityState pinned hidden -> output dropped, not painted; the
hiddenDeliveryGatedVisiblePtyCount field discriminator reads >0) and the
recovery (one Shift keypress repaints the missed output from the main-owned
snapshot, no reload, while visibilityState still reads hidden). Negative
control verified: the spec fails without this fix. Typing-latency perf
gate passes; terminal-pane unit suites 366/366.
Co-authored-by: Orca <help@stably.ai>
* Add a one-paste terminal freeze report: __orcaTerminalFreezeReport()
Every field report of the frozen-terminal family so far has needed
follow-up asks (console output, main logs, second snapshots) because each
capture showed one process's counters at one instant. This makes a single
DevTools command sufficient: `await window.__orcaTerminalFreezeReport()`
returns renderer state (document.visibilityState + the stale-visibility
override, pty:data listener count, delivery-watchdog totals), main's debug
snapshot extended with a per-pty delivery table (sent/acked/pending, hidden
vs visible-set membership, last send/ACK ages, window focus flags, power
suspend/resume ages, app version), and bounded breadcrumb rings from BOTH
processes recording the transitions that matter: gate marks/unmarks,
visibilitychange and stale-visibility latches, watchdog stalls and heals,
restore markers, heal write-offs, and renderer lifecycle resets (so "user
already reloaded" is visible in the history).
Costs stay off the data path: breadcrumbs record only rare transitions into
a 100-entry ring with same-kind coalescing (a flood costs one slot per
second); the per-pty table is built only when the snapshot is read; the
per-send bookkeeping adds one Date.now() to existing accounting writes. Pty
ids are redacted to their `@@` suffix because daemon session ids embed
worktree paths. The report assembles over invoke IPC — the direction proven
alive in every observed wedge — and a failing invoke is captured as data
instead of sinking the report.
The stuck-occlusion e2e now also pins the report end-to-end: after the
wedge + keystroke recovery, the report must carry the stale-visibility
latch and gate transitions in the renderer ring, gate-mark/unmark in main's
ring, and a populated per-pty table. Suites: pty.test.ts 258, terminal-pane
1705, shared ring 5; typing-latency perf gate passes.
Co-authored-by: Orca <help@stably.ai>
* perf(daemon): keep-tail thin hidden panes' stream so agent floods never bury typing (STA multi-workspace lag)
Hidden panes are exempt from pendingData flow control (main gate-drops
their bytes after ingestion), so N background agents ran unbounded ahead
on the one shared daemon->main stream socket (measured 192MB user-space
backlog) and visible-pane echo waited FIFO behind it — typing appeared
seconds late whenever several agents burst on a loaded machine
(8x512KB/s + 12 CPU spinners: p50 293ms fix-off; 12x1MB/s: 6.1s).
Mechanism (replaces producer pacing — no reveal catch-up, ever):
- Shallow socket write gate (128KB) + per-session fairness bypass bounds
echo latency by construction; kernel-flush refill sentinel keeps held
bulk draining at full speed (drain-only refill capped at ~8MB/s).
- Backgrounded sessions' queued output is keep-tail dropped (newest
512KB kept, in-order dataGap replaces the middle); a ~2MB GLOBAL
budget shrinks per-session keep-tails (floor 64KB) so a worktree
switch never waits behind the aggregate. Reply-eliciting query bytes
(DSR/DA/OSC probes) are salvaged from dropped spans.
- Notifications are structurally lossless: the daemon runs the same
shared scanners main uses (bell/OSC 133/pr-link/2031) over every byte
BEFORE drop decisions and relays facts in byte order; ordered
background markers hand scan authority back and forth, seeded with the
emulator's partial escape tail so a sequence split across the handoff
neither phantom-fires nor goes missing. Titles/agent-status stay
main-side (kept-tail convergent).
- Main: background = hidden AND no remote view subscriber (a live
mobile/web view is never thinned); on dataGap main resets cross-chunk
parse carries, drops the headless mobile mirror, and reuses the
hidden-drop model-restore marker.
Wire: three new stream events, tolerated within protocol v19 (old mains
ignore unknown events; old daemons never see the trigger). Kill
switches: ORCA_DAEMON_BACKGROUND_STREAM_DROP=0,
ORCA_DAEMON_SHALLOW_SOCKET_GATE=0.
A/B (pnpm bench:multi-workspace-typing): 8x512KB/s + 12 CPU workers
p50 293ms/p90 647ms -> 15/21ms (= baseline); 12x1MB/s 6,146ms -> 20ms;
light loads unchanged; zero missing echoes. Latin hidden-restore e2e
green (probe-verified aggregate-drain root cause). New deterministic
repro harness: tests/e2e/terminal-multi-workspace-typing-latency.spec.ts
+ CPU pressure workers.
Co-authored-by: Orca <help@stably.ai>
* diag(terminal): breadcrumb WebGL context-loss/atlas + wake triggers into freeze report
Silent instrumentation (memory ring only, no new console lines) so the next
post-wake garble report attributes itself. Adds:
- shared/terminal-webgl-diagnostics.ts: lib-safe sink so pane-webgl-renderer
(lib) can record without importing the components-layer ring; wired to the
ring in terminal-freeze-breadcrumbs.
- webgl-context-loss crumb at onContextLoss, webgl-atlas-reset crumb at the
atlas registry reset — the pair that distinguishes 'atlas corrupted' from
'missed repaint'.
- wake-recovery:<source> crumb (focus/visibilitychange/system-resumed) with the
clearGlyphAtlases decision; source in the kind so distinct triggers don't
coalesce.
- per-pane WebGL state (getAllPaneRenderingDiagnostics) in the freeze report.
Gates: typecheck 0 errors; terminal suites 328 files pass; oxlint clean.
Co-authored-by: Orca <help@stably.ai>
* fix(lint): use Number.parseInt/parseFloat in terminal-view-attributes
oxlint unicorn(prefer-number-properties) flagged 24 global parseInt/parseFloat
calls in the terminal-view-attributes feature (
|
||
|
|
6c5582ced7 |
Remove accidental prod-release-scan output file and ignore future scans (#8037)
- Delete prod-release-scan-1.4.131-rc2-output.md, a scan report that was accidentally committed - Add prod-release-scan-*.md to .gitignore to prevent recurrence |
||
|
|
417723411e |
perf(source-control): stop gh rate-limit storms and idle git-status spawn churn (#7595)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
2f3c7e8660 | Fix Windows slow startup and multi-minute UI freezes (#7225) (#7266) | ||
|
+3 |
36277801e4 |
Make remote hosts first class: concurrent multi-host workbench (#5071)
* Restore the outlined server card for host headers Feedback: the bordered card with the server glyph made it clearer that a host section is a separate machine, not just another group. Bring that back while keeping the recent quieting: no status dot when healthy (marks only for connecting/blocked/error/disconnected), no 'This computer' detail on the local host, and collapse/menu/count behavior unchanged. Co-authored-by: Orca <help@stably.ai> * Anchor host badge to its label, indent rows under host cards Sidebar polish from review: - The count badge sat in dead space between the label and the hover-only chevron/menu; it now hugs the label like repo headers - Rows under a host card get a left inset so projects and workspaces visibly belong to the machine above them - A host whose only visible row is a collapsed repo group counted 0 while the group badge said 9; host counts now fall back to header counts for groups contributing no visible items Co-authored-by: Orca <help@stably.ai> * Two-tier sticky headers: pinned host card above pinned group header When scrolling inside a host section, the host card now stays pinned at the top (z-30) while project/status group headers hand off beneath it (z-20, offset by the pinned card height). The host is the outer hierarchy level, so it is the most persistent context — previously the first repo header replaced it, losing 'which machine am I on' exactly when it mattered. The pinned card keeps its collapse/menu/warning affordances. Handoff rules: the next host card pushes the previous one out at the viewport top; a group pins only once it reaches the slot beneath the host card, and a previous host's group can never pin under the next host. Without host sections the logic degrades to the original single-tier behavior. Co-authored-by: Orca <help@stably.ai> * Revert host-section row indent The two-tier sticky host card now provides continuous 'inside this machine' context at any scroll depth, making the static indent redundant — and it cost 12px of sidebar width on every row while making multi-host layouts misalign with single-host ones. Host cards bracketing their sections plus the pinned header carry the ownership signal on their own. Co-authored-by: Orca <help@stably.ai> * Checkpoint multi-host sidebar and project-first notes Co-authored-by: Orca <help@stably.ai> * Add project-first compatibility persistence Co-authored-by: Orca <help@stably.ai> * Expose project host setup APIs Co-authored-by: Orca <help@stably.ai> * Group sidebar rows by project setup Co-authored-by: Orca <help@stably.ai> * Document project-first host model discussion Co-authored-by: Orca <help@stably.ai> * Resolve workspace creation through project host setups Co-authored-by: Orca <help@stably.ai> * Stamp workspace ownership with project host setup Co-authored-by: Orca <help@stably.ai> * Add project host setup existing folder API Co-authored-by: Orca <help@stably.ai> * Summarize project-first host model discussion Co-authored-by: Orca <help@stably.ai> * Add project host setup CLI commands Co-authored-by: Orca <help@stably.ai> * Allow CLI worktree creation by project host setup Co-authored-by: Orca <help@stably.ai> * Add workspace host setup picker Co-authored-by: Orca <help@stably.ai> * Add project host setup settings summary Co-authored-by: Orca <help@stably.ai> * Make project host setup settings navigable Co-authored-by: Orca <help@stably.ai> * Stabilize project host setup settings selector Co-authored-by: Orca <help@stably.ai> * Add project host existing-folder setup form Co-authored-by: Orca <help@stably.ai> * Update project host model implementation status Co-authored-by: Orca <help@stably.ai> * Keep projects outermost in default sidebar view Co-authored-by: Orca <help@stably.ai> * Update project-first sidebar status Co-authored-by: Orca <help@stably.ai> * Show host context in project sidebar groups Co-authored-by: Orca <help@stably.ai> * Show unavailable hosts in workspace run target Co-authored-by: Orca <help@stably.ai> * Import missing project host from composer Co-authored-by: Orca <help@stably.ai> * Clone project host setup from composer Co-authored-by: Orca <help@stably.ai> * Persist project host setup method Co-authored-by: Orca <help@stably.ai> * Clone project hosts over SSH Co-authored-by: Orca <help@stably.ai> * Improve SSH clone cancellation cleanup Co-authored-by: Orca <help@stably.ai> * Backfill workspace project host ownership Co-authored-by: Orca <help@stably.ai> * Gate project host setup runtime capability Co-authored-by: Orca <help@stably.ai> * Preserve independent project host setups Co-authored-by: Orca <help@stably.ai> * Add project host setup update API Co-authored-by: Orca <help@stably.ai> * Add project host setup delete API Co-authored-by: Orca <help@stably.ai> * Add project host setup create API Co-authored-by: Orca <help@stably.ai> * Expose project host setup lifecycle in renderer store Co-authored-by: Orca <help@stably.ai> * Handle independent project host setups in settings Co-authored-by: Orca <help@stably.ai> * Add pending host setup action in project settings Co-authored-by: Orca <help@stably.ai> * Show pending project host setup status in composer Co-authored-by: Orca <help@stably.ai> * Report pending setup state in workspace target resolution Co-authored-by: Orca <help@stably.ai> * Use shared host registry for project setup choices Co-authored-by: Orca <help@stably.ai> * Add settings clone flow for project host setups Co-authored-by: Orca <help@stably.ai> * Gate unavailable project host setup options Co-authored-by: Orca <help@stably.ai> * Gate unavailable project setup hosts in settings Co-authored-by: Orca <help@stably.ai> * Stream SSH clone progress to renderer Co-authored-by: Orca <help@stably.ai> * Update project host model status notes Co-authored-by: Orca <help@stably.ai> * Add CLI project host setup clone command Co-authored-by: Orca <help@stably.ai> * Make add project host aware Co-authored-by: Orca <help@stably.ai> * Complete project host setup validation Co-authored-by: Orca <help@stably.ai> * Recover floating workspace terminal WebGL atlas on reopen (#5069) Co-authored-by: Orca <help@stably.ai> * Fix stale terminal daemon spawn health (#5064) Co-authored-by: Orca <help@stably.ai> * Suspend floating workspace terminal WebGL while the panel is closed (#5073) Co-authored-by: Orca <help@stably.ai> * Fix source control branch compare base (#5074) Co-authored-by: Orca <help@stably.ai> * Fix workspace-creation tour panel clipped by the Create Worktree dialog (#5078) * Fix workspace-creation tour panel clipped by the composer dialog The tour panel portals into dialog/sheet content that clips overflow, but its position was clamped against the window viewport. With the Project field spanning nearly the dialog's full width, the panel landed past the dialog's right edge and overflow-hidden cut it down to a sliver. Clamp hosted panels within the host's bounds instead, so the panel flips below the target and stays fully visible. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Add JSDoc docstrings to satisfy CodeRabbit docstring coverage check Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * Test hosted contextual tour overlay positioning --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> * release: v1.4.56 * Handle buffer overflows gracefully and truncate diffs fairly (#5083) - Gracefully fall back to file-name summaries when staged diffs exceed node/ssh execution maxBuffer limits, preventing generation failures. - Split oversized diffs by file and allocate budget via water-filling, ensuring single huge files do not starve smaller human changes. - Clip truncated diff sections on line boundaries to avoid half-lines. * Wrap AI generation controls with tooltips and clean i18n dependencies (#5087) - Wrap the AI generation button in a tooltip so users can see the disabled reason or the action description on hover. - Add unit tests verifying tooltip triggers and aria-label safety. - Simplify memo dependencies in settings metadata and worktree palette by using 'useTranslation()' to handle language-change rerenders directly without needing 'i18n.language'. * fix: address review findings (#5088) * Fix localization in repository hooks and base ref suggestion toast (#5089) * Fix localization in base ref toast and custom hook description - Localize the "commit"/"commits" plural nouns in the base ref toast. - Translate missing suggestion toast strings for JA, KO, and ZH locales. - Pass `{{artifact_url}}` as a literal template variable to translate calls to prevent i18next from treating it as a dynamic placeholder. * Fix localization reactivity in RepositoryHooksSection Move static variables containing translation calls into helper functions and subscribe to translation updates using useTranslation. This ensures that localized options, descriptions, and error messages refresh dynamically when the user changes the UI language. * Fix task page labels after language changes (#5086) Co-authored-by: Orca <help@stably.ai> * release: v1.4.57 * Fix automation tabs showing a shell instead of the live agent (#5099) * Fix automation tabs showing a shell instead of the live agent Opening a background automation's terminal tab showed a bare shell while the agent (Claude) kept running headless — the sidebar updated but the pane was attached to the wrong PTY. On first mount the restored ptyId equals the tab ptyId, and isSessionOwnedByWorktree() returns true for it, so connectPanePty routed the still-live eagerly-spawned PTY into the daemon-reattach branch (transport.connect({ sessionId })), which spawns a fresh shell and orphans the live agent PTY instead of adopting it via attach()+replay. Part A: gate the deferred reattach on the absence of a live eager buffer. A live eager buffer means the PTY is a still-running local session to adopt (attach + replay), not a daemon session to re-connect. Daemon reattach and remote PTYs are unaffected (gated on the eager buffer). Part B: publish never-mounted background automation tabs into the runtime graph (gated on a live eager buffer) so the live agent PTY binds to its real tab instead of surfacing as an orphan `pty:<id>` terminal — fixing `orca terminal list`, the CLI, and automation session-reuse. Adds a characterization test (fails on the old code, passes now) and a runtime-graph publish test. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Harden eager PTY tab adoption Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: Orca <help@stably.ai> * Fix i18n label spacing in menus and settings (#5108) * fix i18n label spacing * Fix localized account runtime labels Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: Orca <help@stably.ai> * Improve localization catalog sync workflow (#5110) Co-authored-by: Orca <help@stably.ai> * Add Warp terminal theme import (#4714) Co-authored-by: Orca <help@stably.ai> * release: v1.4.58 * Tidy README badge layout * Handle integration credential decrypt failures (#4683) Co-authored-by: Orca <help@stably.ai> * Fix git repo telemetry for repo adds (#5121) Co-authored-by: Orca <help@stably.ai> * Add feature interaction usage bucket telemetry (#5119) Co-authored-by: Orca <help@stably.ai> * Reset WebGL glyph atlases globally to stop cross-terminal glyph corruption (#5122) Co-authored-by: Orca <help@stably.ai> * perf(windows): fix 60s startup ACL walk and OpenCode streaming freeze, with benchmark harnesses (#5124) * release: v1.4.59-rc.0 * Fix packaged shell PATH order (#5125) Co-authored-by: Orca <help@stably.ai> * Add Floating Workspace contextual tour (#5062) * Add floating workspace contextual tour Co-authored-by: Orca <help@stably.ai> * Clarify floating workspace tour intro copy Co-authored-by: Orca <help@stably.ai> * Differentiate floating workspace tour steps instead of repeating examples Co-authored-by: Orca <help@stably.ai> * Lead floating workspace tour with the user benefit Co-authored-by: Orca <help@stably.ai> * Pitch floating workspace tour around cross-repo agents Co-authored-by: Orca <help@stably.ai> * Refine floating workspace tour step 1 copy Co-authored-by: Orca <help@stably.ai> * Anchor floating workspace tour step 2 on the minimize control Co-authored-by: Orca <help@stably.ai> * Restore floating workspace tour step 2 Co-authored-by: Orca <help@stably.ai> * Anchor floating workspace tour steps on New Terminal and New Markdown Note Co-authored-by: Orca <help@stably.ai> * Retitle floating workspace tour step 2 as scratchpad Co-authored-by: Orca <help@stably.ai> * Add why-comments for tour selector fallback and placement flipping Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> * Fix source control compare base ambiguity (#5127) Co-authored-by: Orca <help@stably.ai> * release: v1.4.59-rc.1 [rc-slot:2026-06-10-15] * release: v1.4.59 * Default-driven create-project flow: name-first form with sensible defaults (#5115) Co-authored-by: Orca <help@stably.ai> * Redesign Connect integrations (#4531) Co-authored-by: Orca <help@stably.ai> * Expose E2E store via build mode * File search match counts (#5085) * Add matchCount to SearchFileResult for accurate per-file hit counts Co-authored-by: Orca <help@stably.ai> * Add file search match count design * rm design doc --------- Co-authored-by: Orca <help@stably.ai> * fix: address review findings (#5139) * perf(windows): avoid blocking daemon pid checks (#5137) * release: v1.4.60-rc.0 * release: v1.4.60 * Preserve core workflow terms in English and apply CJK spacing (#5141) * Preserve core workflow and product terms in English across locales Update translation policy to prevent localization of key terms such as "Agent", "Commit", "Markdown", and "Terminal". This ensures consistent jargon and product branding. Introduce CJK-Latin term spacing to keep these Latin terms legible when combined with CJK text, while adjusting Korean particle spacing. Also add overrides to prevent network proxy settings from being mistranslated as "Agent". * Preserve repo terminology in English and localize source control labels Treat "repo" and "repos" (and their capitalized forms) as brand terms that should remain in English/Latin across CJK and Spanish locales. Update translation files and policies to replace translated words like "repositorio" or "リポジトリ" with "repo"/"repos", and fix an issue where latin brand terms could be incorrectly matched as substrings in larger words during cleanup. Additionally, externalize and localize the "Staged Changes", "Changes", and "Untracked Files" section labels in the source control sidebar. * UX (#5143) * UX/copy tweaks (#5142) * UX/copy tweaks * UX/copy tweaks * Fix missed star UI translations (#5148) * fix: make windows ssh relay deploy survive session teardown (#5136) * Add option to remove child projects when deleting repo groups (#4702) Co-authored-by: Orca <help@stably.ai> * fix: remove checks panel response badge (#5147) * Add read-only `orca linear` CLI with trusted launch-prompt pointer (V1) (#5126) Co-authored-by: Orca <help@stably.ai> * Add AI Vault session history ## Summary - add AI Vault session scanning and resume command construction - add the Agents sidebar panel with filtering, grouping, copy/open actions, and local resume launch - support dragging saved sessions onto terminal split panes ## Validation - pnpm run lint - pnpm run typecheck - pnpm exec vitest run --config config/vitest.config.ts src/main/ipc/register-core-handlers.test.ts src/main/ai-vault/session-scanner.test.ts src/renderer/src/components/right-sidebar/ai-vault-session-filters.test.ts src/renderer/src/lib/ai-vault-session-drag.test.ts src/renderer/src/lib/launch-ai-vault-session.test.ts * Default agent launches to yolo permissions mode (#5145) * Default agent launches to yolo mode * test: update launch default validations * Fix Claude usage refresh error copy (#5155) Co-authored-by: Orca <help@stably.ai> * Move workspace board to sidebar bottom toolbar (#5146) Co-authored-by: Orca <help@stably.ai> * Rebuild contextual tour positioning on floating-ui; fix hosted dialog placement and arrow seam (#5154) Co-authored-by: Orca <help@stably.ai> * Fix missing spaces in cross-repo switch dialog (#5158) * Fix Ctrl+Tab switcher selection on release (#5116) * Fix additional i18n spacing regressions from #4995 (#5159) * Refine add project selection styling (#5160) Co-authored-by: Orca <help@stably.ai> * improve chinese localization (#5162) * Fix floating workspace needing two clicks after app switch (macOS) (#5128) * Autofocus feedback textarea when Send Feedback dialog opens (#5164) * fix: address pr-bug-scan validated finding from #4683 (#5151) Isolated CredentialDecryptionError per-item in Linear getClients (client.ts:518) and Jira getClients (client.ts:373) on the 'all' selection so one bad credential no longer collapses healthy workspaces Co-authored-by: orca-bug-scan-bot <orca-bug-scan-bot@stably.ai> * fix: enable claude agent teams by default (#5168) * Refresh Jira and Linear status after credential errors (#5169) * fix: address pr-bug-scan validated finding from #4683 Isolated CredentialDecryptionError per-item in Linear getClients (client.ts:518) and Jira getClients (client.ts:373) on the 'all' selection so one bad credential no longer collapses healthy workspaces * Refresh Jira and Linear status to clear stale credential errors Ensure stale credential decryption errors are cleared from the store status once a successful API read completes. By updating the check in shouldRefreshStatusAfterRead to trigger when a credentialError is currently set, successful issue or list fetches will trigger a status check and remove stale error flags. --------- Co-authored-by: orca-bug-scan-bot <orca-bug-scan-bot@stably.ai> * Hide internal context from AI Vault titles (#5175) * Fix detached HEAD publish actions (#5173) * Keep freshly split terminal pane mounted if newborn PTY exits early (#5171) Prevent a newly split pane from collapsing immediately if its PTY exits during initial setup before any output is received or input is sent. This ensures a failed startup session remains visible to the user. * Route task PR queries by upstream source (#5176) * Route task PR queries by upstream source Implements the routing described in docs/tasks-pr-upstream-source.md so task PR and issue queries stay scoped to the selected source. * rm design doc * Prevent stale PR refreshes from restoring unlinked review state (#5180) - Pass `worktreeId` to `fetchPRForBranch` to track active worktree context - Ignore inflight or queued PR fetches if the worktree has been unlinked - Include linked PR/MR metadata in the checks panel snapshot key to trigger updates immediately on link/unlink events * Fix Claude agents management status detection (#5179) Co-authored-by: Orca <help@stably.ai> * fix: address review findings (#5177) * Allow resolving selected review comments with AI (#5184) * Allow resolving selected PR/MR review comments with AI Users can now select specific unresolved review comments or threads in the Checks panel sidebar, queue them, and trigger an AI agent to address them, marking resolved threads on the host upon agent launch. - Adds checkboxes and action/send buttons to select and queue comments. - Builds a structured, robust prompt with sanitized comment metadata. - Optimistically marks threads resolved on launch with rollback on error. - Supports both GitHub PRs and GitLab MRs. * Consolidate PR comment selection state and eliminate effects Combine independent selection states and context-tracking into a single state object. Derive active selection data and prune ineligible comments during render using useMemo instead of relying on asynchronous useEffect synchronization hooks. * Improve source control action dialog layout and recipe saving UX (#5153) * Improve source control agent action dialog layout and recipe UX - Constrain dialog and scroll area heights to prevent viewport overflow. - Add variable chips to easily insert the base prompt with tooltip previews. - Keep the recipe save controls visible when a recipe is already saved, showing informational status text instead of hiding them. - Update localized copy across multiple languages and reduce textarea rows. - Add unit tests for the variable chip preview and save target visibility. * Fix recipe-saved check in source control action dialog * Evaluate only the selected save target instead of checking all available targets, as the action only writes to the selected target. * Update daemon PTY adapter test fake PID to prevent collision with real host OS processes during runtime directory lookups. * fix: remove unsupported agent launch defaults (#5185) * Update Chinese and Japanese translations for worktrees and fixes (#5187) - Correct awkward Chinese translation of "fix" ("使固定") to "修复" and "基本的" to "主工作树" (main worktree). - Improve Japanese translation of "fix" from physical repair ("修理") to software correction ("修正"). * Embed hosted review creation composer directly in Checks panel (#5140) * Embed hosted review creation composer directly in the Checks panel - Replaces the modal pull request/merge request creation dialog with an inline composer embedded in the empty state of the Checks sidebar. - Extracts and moves pull request generation state to a dedicated store slice so AI-generated details are persisted across sidebar unmounts. * Fix hosted review composer feedback * Combine file search and file explorer right sidebar tabs (#5182) Unifies file discovery and tree navigation under a single Explorer domain, simplifying the right sidebar activity bar and reducing tab clutter. * Replaces the standalone 'search' activity bar tab with a nested 'search' subview inside the File Explorer tab * Introduces 'rightSidebarExplorerView' ('files' | 'search') state to manage the active subview inside the Explorer * Adds a search button to the File Explorer toolbar and a back button to the search subview for seamless transition * Exposes 'showRightSidebarFiles' and 'showRightSidebarSearch' store actions to route and seed search queries/include patterns * Adapts file explorer keybindings, git status polling, and external workspace watchers to respect the active subview * Maps legacy persisted search tab state to the new explorer search view for backward compatibility * release: v1.4.61-rc.1 * Add multi-repo folder workspaces (v1) (#5172) Co-authored-by: Orca <help@stably.ai> * release: v1.4.61-rc.2 * Hide unavailable project hosts in worktree composer Co-authored-by: Orca <help@stably.ai> * Remove inline project host setup from composer Co-authored-by: Orca <help@stably.ai> * Mark imported project host setup methods Co-authored-by: Orca <help@stably.ai> * Fix rebase merge fallout Co-authored-by: Orca <help@stably.ai> * Disable unavailable Add Project hosts Co-authored-by: Orca <help@stably.ai> * Compact Add Project host selector Co-authored-by: Orca <help@stably.ai> * Hide redundant SSH target chooser Co-authored-by: Orca <help@stably.ai> * Browse SSH clone destinations Co-authored-by: Orca <help@stably.ai> * Avoid local clone defaults for SSH hosts Co-authored-by: Orca <help@stably.ai> * Polish host-aware Add Project flows Co-authored-by: Orca <help@stably.ai> * Polish remote host add project flows Co-authored-by: Orca <help@stably.ai> * Remove redundant host kind chips Co-authored-by: Orca <help@stably.ai> * Fix remote project setup UX gaps Co-authored-by: Orca <help@stably.ai> * Fix multihost workspace composer project identity Co-authored-by: Orca <help@stably.ai> * Finish host context merge repair Co-authored-by: Orca <help@stably.ai> * Continue host context checklist implementation Co-authored-by: Orca <help@stably.ai> * Route Linear and Jira tasks by source context Co-authored-by: Orca <help@stably.ai> * Preserve Linear task source context in history Co-authored-by: Orca <help@stably.ai> * Scope task retry state by source context Co-authored-by: Orca <help@stably.ai> * Route GitHub drawer reads by source context Co-authored-by: Orca <help@stably.ai> * Guard GitLab selectors with repo context Co-authored-by: Orca <help@stably.ai> * Guard GitHub metadata selectors Co-authored-by: Orca <help@stably.ai> * Route GitHub task row actions by source context Co-authored-by: Orca <help@stably.ai> * Update GitHub source-context checklist status Co-authored-by: Orca <help@stably.ai> * Show host ownership for CLI provider accounts Co-authored-by: Orca <help@stably.ai> * Persist GitLab task detail source context Co-authored-by: Orca <help@stably.ai> * Show host scope for provider API budgets Co-authored-by: Orca <help@stably.ai> * Preserve Jira task source context Co-authored-by: Orca <help@stably.ai> * Scope Jira optimistic task patches Co-authored-by: Orca <help@stably.ai> * Resolve task PR bases on run host Co-authored-by: Orca <help@stably.ai> * Record Jira task workspace usage Co-authored-by: Orca <help@stably.ai> * Scope Linear optimistic task patches Co-authored-by: Orca <help@stably.ai> * Scope GitHub optimistic task patches Co-authored-by: Orca <help@stably.ai> * Clean host copy in onboarding flows Co-authored-by: Orca <help@stably.ai> * Preserve automation CLI run context Co-authored-by: Orca <help@stably.ai> * Add automation CLI source context selector Co-authored-by: Orca <help@stably.ai> * Clarify unavailable task source hosts Co-authored-by: Orca <help@stably.ai> * Surface host model runtime capability skew Co-authored-by: Orca <help@stably.ai> * Use SSH host copy in reconnect dialog Co-authored-by: Orca <help@stably.ai> * Show host context in task source picker Co-authored-by: Orca <help@stably.ai> * Mark task source display complete Co-authored-by: Orca <help@stably.ai> * Clarify provider account host selection Co-authored-by: Orca <help@stably.ai> * Guard task source switching boundary Co-authored-by: Orca <help@stably.ai> * Mark task source diagnostics persisted Co-authored-by: Orca <help@stably.ai> * Mark base resolution host boundary Co-authored-by: Orca <help@stably.ai> * Clarify external automation source states Co-authored-by: Orca <help@stably.ai> * Harden project host compatibility projection Co-authored-by: Orca <help@stably.ai> * Finish host copy audit Co-authored-by: Orca <help@stably.ai> * Add provider host scope controls Co-authored-by: Orca <help@stably.ai> * Show task source account labels Co-authored-by: Orca <help@stably.ai> * Show automation run context in CLI Co-authored-by: Orca <help@stably.ai> * Scope Jira task cache lookups by source Co-authored-by: Orca <help@stably.ai> * Seed workspace creation from task source context Co-authored-by: Orca <help@stably.ai> * Explain disabled external automation actions Co-authored-by: Orca <help@stably.ai> * Surface task source runtime capability gaps Co-authored-by: Orca <help@stably.ai> * Persist automation run context from UI saves Co-authored-by: Orca <help@stably.ai> * Require workspace run capability for setup hosts Co-authored-by: Orca <help@stably.ai> * Disable automation runs for stale host setup Co-authored-by: Orca <help@stably.ai> * Route GitHub drawer metadata by source host Co-authored-by: Orca <help@stably.ai> * Guard runtime project setup mutations by host model Co-authored-by: Orca <help@stably.ai> * Route PR page metadata by repo host Co-authored-by: Orca <help@stably.ai> * Route PR mention metadata by repo host Co-authored-by: Orca <help@stably.ai> * Route GitHub Project edits by view source Co-authored-by: Orca <help@stably.ai> * Clarify runtime automation disabled states Co-authored-by: Orca <help@stably.ai> * Guard runtime automation backend dispatch Co-authored-by: Orca <help@stably.ai> * Preserve GitLab task source identity Co-authored-by: Orca <help@stably.ai> * Remove redundant SSH target row in add project Co-authored-by: Orca <help@stably.ai> * Add task source provider availability reasons Co-authored-by: Orca <help@stably.ai> * Surface task provider preflight availability Co-authored-by: Orca <help@stably.ai> * Record local GitHub task source verification Co-authored-by: Orca <help@stably.ai> * Record Linear task source verification Co-authored-by: Orca <help@stably.ai> * Show automation source context in details Co-authored-by: Orca <help@stably.ai> * Record remote capability negotiation coverage Co-authored-by: Orca <help@stably.ai> * Record local add project create verification Co-authored-by: Orca <help@stably.ai> * Scope Linear cached task reads by source Co-authored-by: Orca <help@stably.ai> * Preserve PR generation host ownership Co-authored-by: Orca <help@stably.ai> * Route git operations by owner host Co-authored-by: Orca <help@stably.ai> * Route delete warnings by worktree owner Co-authored-by: Orca <help@stably.ai> * Route editor drops by worktree owner Co-authored-by: Orca <help@stably.ai> * Route agent draft paste by tab owner Co-authored-by: Orca <help@stably.ai> * Route file explorer requests by worktree owner Co-authored-by: Orca <help@stably.ai> * Document remaining host context gaps Co-authored-by: Orca <help@stably.ai> * Check runtime task source provider auth Co-authored-by: Orca <help@stably.ai> * Validate automation source availability Co-authored-by: Orca <help@stably.ai> * Route remaining UI requests by owner host Co-authored-by: Orca <help@stably.ai> * Route quick open file listing by worktree owner Co-authored-by: Orca <help@stably.ai> * Route typed GitHub lookups by source host Co-authored-by: Orca <help@stably.ai> * Centralize automation run identity fallback Co-authored-by: Orca <help@stably.ai> * Surface unsupported task source providers Co-authored-by: Orca <help@stably.ai> * Document automation legacy repo compatibility Co-authored-by: Orca <help@stably.ai> * Record live host model verification Co-authored-by: Orca <help@stably.ai> * Quiet disconnected SSH polling Co-authored-by: Orca <help@stably.ai> * Verify task drawer source boundaries Co-authored-by: Orca <help@stably.ai> * Verify GitLab repo source selectors Co-authored-by: Orca <help@stably.ai> * Route automations through owning host Co-authored-by: Orca <help@stably.ai> * Update host context verification checklist Co-authored-by: Orca <help@stably.ai> * Run remote automations headlessly in serve mode Co-authored-by: Orca <help@stably.ai> * Keep setup guide entry stable during refresh Co-authored-by: Orca <help@stably.ai> * Keep setup script prompt stable during host switches Co-authored-by: Orca <help@stably.ai> * Deduplicate Tasks project picker sources Co-authored-by: Orca <help@stably.ai> * Use project identity for Tasks picker dedupe Co-authored-by: Orca <help@stably.ai> * Add Tasks source host switcher Co-authored-by: Orca <help@stably.ai> * Refine Tasks source picker disclosure Co-authored-by: Orca <help@stably.ai> * Polish Tasks source picker hover Co-authored-by: Orca <help@stably.ai> * Open Tasks source menu on hover Co-authored-by: Orca <help@stably.ai> * Match Tasks source submenu hover behavior Co-authored-by: Orca <help@stably.ai> * Open Tasks source submenu from project row hover Co-authored-by: Orca <help@stably.ai> * Group automation project hosts Co-authored-by: Orca <help@stably.ai> * Tighten automation project picker density Co-authored-by: Orca <help@stably.ai> * Show selected host in Tasks project picker Co-authored-by: Orca <help@stably.ai> * Hide host labels for single-host project pickers Co-authored-by: Orca <help@stably.ai> * Use saved remote server names in host pickers Co-authored-by: Orca <help@stably.ai> * Use standard add project start for remote servers Co-authored-by: Orca <help@stably.ai> * Use saved host labels in workspace surfaces Co-authored-by: Orca <help@stably.ai> * Route remote browser tabs through runtime hosts Co-authored-by: Orca <help@stably.ai> * Keep sidebar project-first across grouping modes Co-authored-by: Orca <help@stably.ai> * Polish multi-host remote runtime UX Co-authored-by: Orca <help@stably.ai> * Fix CI lint and remove design notes Co-authored-by: Orca <help@stably.ai> * Fix CI test failures Co-authored-by: Orca <help@stably.ai> * Fix Windows CLI path expectation Co-authored-by: Orca <help@stably.ai> * Fix CI renderer test expectations Co-authored-by: Orca <help@stably.ai> * Fix remaining verify test failures Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> Co-authored-by: Bryant Ung <bryant.ung@outlook.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> Co-authored-by: Borja <3930245+BorjaLL@users.noreply.github.com> Co-authored-by: Parker Rex <me@parkerrex.com> Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> Co-authored-by: Trevin Chow <trevin@trevinchow.com> Co-authored-by: buf0-bot[bot] <252831055+buf0-bot[bot]@users.noreply.github.com> Co-authored-by: orca-bug-scan-bot <orca-bug-scan-bot@stably.ai> |
||
|
|
8bc15e1439 | fix: address review findings (#5088) | ||
|
|
d09a7fa28b | Add Simplified Chinese, Korean, and Japanese UI localization (#5046) | ||
|
|
dd302d3e91 |
Remove H1/H2/H3 badges from markdown TOC (#4832)
* Remove H1/H2/H3 badges from markdown TOC Replace heading level badges with pure indentation for visual hierarchy. - Remove 'H1', 'H2', 'H3' badge spans from MarkdownTocRow - Increase indentation per depth level for clearer visual distinction - Remove unused .markdown-toc-level CSS class Design doc: docs/remove-h1-h2-prefix-from-toc.md Fixes #4821 Co-authored-by: Orca <help@stably.ai> * Include h1 headings in markdown table of contents Previously h1 headings were excluded, treating them as a document-level root. Now h1 entries appear as TOC items alongside h2/h3, giving users a more complete and consistent navigation outline. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
56f761d93d | Ignore local React perf audit ledger (#3417) | ||
|
|
30a09f3bd9 |
Add mobile terminal shortcut bar customization (#3012)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
41a2d7ba90 | Target source-control actions to publish branches (#2612) | ||
|
|
d88b62f115 |
Gate task providers by availability (#2189)
* Gate task providers by availability Implement provider availability gating documented in docs/task-provider-availability.md. * Remove task provider availability design doc * Restore task source after provider availability checks - Keep saved GitLab/Linear defaults from being lost while provider checks hydrate - Ignore stale Linear status responses after connect or workspace changes * fix: address review findings |
||
|
|
e957b86197 |
Add configurable Open In applications (#2110)
* feat: add configurable Open In menu Implements configurable Open In applications in worktree context menus with persisted settings, preload/API wiring, and renderer controls/tests.\n\nDesign doc: docs/configurable-open-in-menu.md * fix: address review findings |
||
|
|
df44d8aff1 | chore: ignore ephemeral docs (#2004) | ||
|
|
ac818caae5 | Reduce GitHub GraphQL rate-limit spend (#1974) | ||
|
|
a22717bb35 |
Refactor runtime app architecture (#1878)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
cb58b109e0 | feat: add workspace space analyzer (#1877) | ||
|
|
0273e38240 | Fix start-from-PR display refresh (#1751) | ||
|
|
0f54103dda |
Add native computer-use automation (#1683)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
cbf99a1c98 |
feat(sidebar): allow manual drag-and-drop reordering of repos (#1686)
* feat(sidebar): allow manual drag-and-drop reordering of repos Users can now drag repo headers in the sidebar to reorder them. The custom order is persisted to disk and survives restarts. Includes design doc at docs/manual-repo-reorder.md. Co-authored-by: Orca <help@stably.ai> * fix: scope post-drag click swallow to dragged repo header Avoid silently eating unrelated clicks if one races between pointerup and the failsafe teardown. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
fd86e1869a |
feat(telemetry): PR 2 — transport (client, validator, burst cap, IPC, build gate) (#1374)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
8d1a6cbbb2 |
chore: remove tracked skills-lock.json (#1255)
Follow-up to #1254. The lockfile tracks skill source/hash metadata which now lives alongside the skill sources in the internal delivery repo; keeping a parallel copy here in the public tree is misleading (entries can drift) without being useful to anyone. - Remove skills-lock.json from the tree. - Gitignore it so local tooling can still write one without making git status dirty. Co-authored-by: Orca <help@stably.ai> |
||
|
|
c85f487ebf |
chore: move agent skills to internal delivery (#1254)
Private skill files are now stored and versioned in a separate internal repo; a setup hook populates .claude/skills and .agents/skills with machine-local symlinks instead. This keeps contributor-facing patterns (agent skills in the tree) while allowing some skills to be developed privately. - Remove tracked skill files (62 files across 7 skill dirs). - Gitignore the now machine-local skill directories. - Add an env-gated hook to orca.yaml scripts.setup. Runs a setup script whose path lives in $ORCA_INTERNAL_DEV_SETUP when present; silently no-ops for public contributors. No new contributor-facing requirements: the hook is optional, the env var is only set by internal tooling, and the setup itself runs in the worktree that Orca just created. Co-authored-by: Orca <help@stably.ai> |
||
|
|
cfc444242c |
fix(win32): resolve EPERM on userData writes, batch-file spawn failures, and native dep rebuild (#1152)
* chore: update .gitignore to include stackdump and .serena, enhance pre-commit script * fix(win32): resolve EPERM on userData writes and batch-file spawn failures Three Windows-specific issues prevented Orca from running correctly on machines where Chromium resets the userData DACL during startup: 1. **EPERM on userData writes** — Chromium's BrowserWindow constructor calls SetNamedSecurityInfo on the userData folder with a Protected DACL. When propagated to child directories the ACEs carry the Inherit-Only flag, meaning they apply to children-of-children but NOT to the directories themselves. Any file write inside codex-runtime-home, agent-hooks, or similar subdirectories fails with EPERM. Fix: grant an explicit Full Control ACE (OI)(CI)(F) on userData and all existing children before BrowserWindow is created (icacls /T /C). Explicit ACEs survive future DACL propagation from the parent. Per-write EPERM retries in fs-utils and installer-utils serve as the backstop for directories created after startup. 2. **Batch-file spawn failures** — resolveCodexCommand() can return a .cmd or .bat path (e.g. codex.cmd installed via npm). Node's spawn() cannot execute batch scripts directly without shell:true, but shell:true with an args array triggers DEP0190 because args are concatenated rather than escaped. Both service.ts and codex-fetcher.ts were affected. Fix: detect .cmd/.bat paths and route through cmd.exe /c explicitly, which is equivalent to what shell:true does internally but avoids the deprecation warning and arg-escaping hazard. 3. **Native dep rebuild failure** — electron-builder install-app-deps does not expose the ignoreModules option. On Windows dev machines without the full VC++ / Python toolchain, cpu-features (an optional dep of ssh2) fails to build with node-gyp, aborting the entire postinstall step. Fix: replace electron-builder install-app-deps with a thin wrapper script (scripts/rebuild-native-deps.mjs) that calls @electron/rebuild's JS API directly with ignoreModules: ['cpu-features'] on Windows. ssh2 detects the missing native module and falls back to pure-JS automatically. Refactoring: extract shared win32-utils.ts with getIcaclsExePath(), getCmdExePath(), isWindowsBatchScript(), isPermissionError(), grantDirAcl(), and getSpawnArgsForWindows() to eliminate five instances of duplicated SystemRoot path construction and two near-identical EPERM retry blocks. Reduce startup icacls calls from three sequential blocking /T invocations to one, removing up to 20 s of potential startup delay. * fix(win32): address review feedback on ACL and spawn helpers - Fall back to SID via `whoami /user` when `USERNAME` is unset so `grantDirAcl` works under services, CI, and hardened envs instead of silently no-op'ing. - Use a 60s timeout for recursive `icacls /T` walks; the 10s cap could starve on large userData trees and silently fail the startup grant. - Pass `windowsHide: true` to `icacls` and the cmd.exe-routed Codex spawns so no console window flashes in the packaged GUI app. - Add `/d` to `cmd.exe /c` invocations to disable AutoRun registry commands — safer default for background spawns. - Drop unused `createRequire`/`require` from rebuild-native-deps.mjs. - Add `@electron/rebuild` as an explicit devDependency; relying on the electron-builder transitive was brittle under pnpm. - Fix two misleading "Re-enable inheritance" comments that describe behavior opposite to what the code actually does (explicit ACL grant). - Add unit tests for `isWindowsBatchScript`, `getSpawnArgsForWindows`, and `isPermissionError` to lock in Windows batch detection + cmd.exe routing. Co-authored-by: Orca <help@stably.ai> * fix(win32): unify PTY spawn through /d and document cmd.exe safety - fetchViaPty now uses getCmdExePath() and /d /c, matching the rest of the codebase instead of hand-rolling 'cmd.exe' + ['/c', ...]. - getSpawnArgsForWindows gains a SAFETY note: when the .cmd/.bat branch is taken, cmd.exe re-parses the combined command line, so callers must only pass trusted/literal args. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
f5ad7aa249 |
feat(settings): import Ghostty config with preview, color overrides, opacity and blur (#1001)
* feat(settings): add Ghostty config import Add safe one-shot Ghostty import with preview and success summary, and probe documented Ghostty config paths before applying changes. Refs #958 * chore(git): ignore atl artifacts * feat(settings): map font-weight, cursor-blink and focus-follows-mouse from Ghostty Adds three safe direct-mapping Ghostty keys that have clear equivalents in GlobalSettings: font-weight, cursor-style-blink, and focus-follows-mouse. Refs #958 * feat(settings): expand Ghostty import to support colors, opacity and option-as-alt - Add TerminalColorOverrides type grouping 21 optional xterm ITheme fields - Add terminalBackgroundOpacity, terminalPanePaddingColor, terminalPaddingBalance to GlobalSettings - Extend parser to collect repeated keys as string[] (needed for palette lines) - Map background-opacity, background, foreground, cursor-color, selection-background/foreground, palette (0-15), window-padding-color, window-padding-balance, macos-option-as-alt in mapper - Merge terminalColorOverrides into xterm ITheme at theme resolution; apply opacity as rgba() with allowTransparency enabled - Accept hex colors with or without leading # (Ghostty omits it) - Fix preview diff to use deep equality for object values so already-applied color overrides no longer reappear on next import * feat(settings): support background-blur-radius and window-padding-color extend in Ghostty import - Map background-blur-radius > 0 to windowBackgroundBlur: true; apply vibrancy on macOS and backgroundMaterial acrylic on Windows at window creation (blur requires restart — no hot-reload IPC exists) - Accept window-padding-color = extend/background as valid Ghostty values; both map to default Orca padding behavior (undefined field) instead of landing in unsupportedKeys - Split mapper.test.ts into domain-scoped describes to stay under 300-line limit * feat(settings): add Window section to Terminal settings panel Expose terminalBackgroundOpacity, windowBackgroundBlur, terminalPaddingBalance, terminalPanePaddingColor, and terminalColorOverrides in the settings UI so imported Ghostty values can be viewed and changed manually. - New TerminalWindowSection component (extracted from TerminalPane to stay under the 400-line limit) - Collapsible color overrides sub-section with ColorField for all 21 xterm ITheme fields grouped as base, ANSI normal, and ANSI bright - Reset button clears all color overrides at once - Window blur toggle shows restart-required note (blur applies at window creation, no hot-reload IPC exists) - Search entries added for all new controls * feat(settings): expand Ghostty import with scrollback, padding, divider, cursor and word-chars keys Map 9 additional Ghostty keys to Orca settings: - split-divider-color → terminalDividerColorDark + terminalDividerColorLight (single value applies to both; Ghostty has no dark/light distinction) - unfocused-split-opacity → terminalInactivePaneOpacity (direct float 0-1) - scrollback-limit → terminalScrollbackLimit; applied to xterm scrollback option - window-padding-x / window-padding-y → terminalPaddingX/Y; applied as CSS vars --pane-padding-x / --pane-padding-y in terminal.css - cursor-text → terminalColorOverrides.cursorAccent (xterm ITheme field) - bold-color → terminalColorOverrides.bold (persisted; xterm ITheme has no bold field yet — stored for future xterm upgrade) - cursor-opacity → terminalCursorOpacity; blended into cursor rgba at theme resolution time - selection-word-chars → terminalWordSeparator; applied to xterm wordSeparator - mouse-hide-while-typing → terminalMouseHideWhileTyping field added; renderer application deferred (needs per-pane disposable + global mousemove listener) * feat(settings): expose new Ghostty-imported settings in Terminal Settings UI - Window section: scrollback limit, horizontal/vertical padding, hide mouse while typing toggle, cursor text and bold color in Color Overrides - Cursor section: cursor opacity NumberField - Advanced section: word separators text input - Search entries added for all new controls - terminalDividerColorDark/Light and terminalInactivePaneOpacity skipped — already present in Theme and Pane Styling sections respectively * feat(terminal): implement mouse-hide-while-typing per pane Register terminal.onData → cursor:none and mousemove → restore, scoped to the pane container element. Uses the existing IDisposable per-pane pattern (same as selectionDisposablesRef). Cleans up on pane close and effect teardown. * refactor(settings): address ghostty import code review findings - Centralize GhosttyImportPreview type in shared/types (remove duplicate from mapper) - Fix parser to strip inline comments without breaking hex color values (#1a1a1a) - Extract HEX_COLOR_RE to shared/color-validation to avoid duplication - Remove redundant Number.isNaN checks after Number.isFinite (4 sites) - Replace unsafe catch-all assignment with explicit font-family branch - Migrate 280-line if-chain in mapGhosttyToOrca to FIELD_PARSERS registry - Add human-readable setting labels in GhosttyImportModal via setting-labels map - Add clarifying comment in index.ts re JSON.stringify undefined behavior * fix(settings): harden ghostty import from judgment-day review - Surface readFile errors in GhosttyImportPreview.error instead of showing misleading 'No config found' on permission denied - Guard handleApply against double-apply when already applied - Strip surrounding quotes from parsed config values (font-family) - Return null from palette handler when all entries fail validation - Inform user when background-blur-radius radius is not preserved - Add valuesEqual key-order stability via stableStringify - Normalize hex colors to #-prefixed format across all color mappers - Reject blank values before numeric parsing (Number('') === 0 trap) - Remove selection-word-chars mapping (inverted xterm semantics) - Guard window-padding-x/y against negative integers - Reactive mouse-hide-while-typing on existing panes when setting toggles - Merge terminalColorOverrides on import instead of replacing * fix(ghostty): drop broken imports, tighten parsing, prompt restart for blur Review found three high-impact issues in the Ghostty import: scrollback-limit semantics are inverted/rescaled (Ghostty is bytes with 0=unlimited, xterm is rows with 0=disabled), and window-padding-color + window-padding-balance set CSS custom properties (--pane-padding-color, --pane-padding-balance) that have no consuming rule anywhere in the tree — so users confirming "changes" to those keys would see nothing happen. Because none of the three keys have a safe mapping today, drop them from the import and remove the dead UI controls + GlobalSettings fields + CSS var plumbing. The mapper now lists them as unsupportedKeys alongside the existing window-decoration / keybind / custom-shader entries. Other fixes in the same review: - allowTransparency now clears when background-opacity returns to 1 (prior code only ever set it to true, leaving a stale flag with measurable render cost). - background-blur-radius = 0 no longer emits a misleading "radius value not preserved" note (0 cleanly maps to blur=false with no radius to lose). - Add a 1 MB size cap on the config read so a pathological or symlinked file cannot OOM the main process. - Make handleApply async and surface IPC errors inline in the modal instead of flipping straight to "Import complete" on failure. - Type settings.previewGhosttyImport as Promise<GhosttyImportPreview> in preload so shape drift is caught at compile time. - Make stableStringify recursive so future nested settings round-trip cleanly through valuesEqual. - Restrict window-padding-x/y and background-blur-radius to decimal ints; prior code accepted exponent notation (1e10 sails through Number.isInteger) and would have landed absurd values in the store. - Window blur now shows a "Restart required" banner with a Restart now button when the setting differs from the mount-time snapshot, mirroring the ExperimentalPane daemon pattern. Blur only applies at BrowserWindow creation on macOS/Windows. Co-authored-by: Orca <help@stably.ai> * feat(settings): move Ghostty import trigger to Terminal section header Per review feedback: the "Import from Ghostty" row was taking its own slot in the Terminal settings list alongside real configuration sections. Move the trigger into the Terminal section's header (upper-right corner) as a headerAction, next to the Terminal heading — it's a one-shot action, not a setting. - SettingsSection gains an optional `headerAction` slot rendered to the right of the section title/description. - The useGhosttyImport hook is lifted from TerminalPane into Settings.tsx so the section header button (owned by Settings.tsx) and the modal (still rendered inside TerminalPane) share one state instance. - TerminalPane drops its own "Import" section + the TERMINAL_GHOSTTY_IMPORT search entry group is no longer referenced there. - Button carries the official Ghostty mark as a 16x16 icon so it reads clearly as a cross-app import even before users parse the label. useGhosttyImport now accepts `GlobalSettings | null` so the parent can call it above the pre-load spinner guard without violating hook ordering; the apply path no-ops until settings arrive. Related test file updated to pass the new `ghostty` prop and to assert the trigger is *not* rendered inside TerminalPane anymore. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
933d59f165 | diff-comments: copy-only flow in Source Control (#884) | ||
|
|
222d70e063 | feat: add idempotent E2E test suite with headless Electron support (#671) | ||
|
|
2688c0fee3 |
feat(editor): preserve Cmd+B for bold in markdown editor (#802)
* feat(editor): preserve Cmd+B for bold in markdown editor Carve out bare Cmd/Ctrl+B from the main-process before-input-event interceptor when the TipTap markdown editor is focused, so its bold keymap can run instead of toggling the left sidebar. Focus state is mirrored from renderer to main via a one-way IPC send, with default-deny resets on crash/navigate/destroy and sender validation so only the main window's webContents can mutate the flag. * fix: add oxlint max-lines disable to createMainWindow.ts |
||
|
|
bced9b5308 |
refactor: reorganize project structure for build assets and scripts (#419)
Move build resources to resources/build/, icon source files to resources/icon-source/, scripts to config/scripts/, and patches to config/patches/ for a cleaner top-level directory layout. |
||
|
|
5d6fb3923f |
feat: add file duplicate to file explorer context menu (#351)
* feat: add file duplicate option to file explorer context menu * fix: address review findings |
||
|
|
902e2271b9 |
fix: handle orphaned worktree deletion with disk cleanup (#109)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
cd14506142 |
chore: add CLAUDE.md and gitignore package-lock.json (#63)
* feat: use Enter to submit and Shift+Enter for line break in comment editor Closes #48 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * chore: add CLAUDE.md and gitignore package-lock.json Enforce pnpm usage for AI agents and prevent stray npm lock files. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
8a231c46f6 | gitignore | ||
|
|
8f5f07b221 | Shadcn |