mirror of
https://github.com/stablyai/orca.git
synced 2026-09-29 16:02:50 +00:00
caa3a729885fab3d43debd025dff2281eb03e8d7
1127
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
caa3a72988 | Merge fix-send-federation-transcript into integrate-fixes | ||
|
|
58adae375e |
fix(transcript): wire the attested WSL distro into exact worker session selection
A WSL pane's PTY is local (connectionId null), so every WSL hook status was filtered out and worker-read always fell back to screen scraping. Pass the same distro expression the headless terminal state already uses. Also deletes the dead v1 transcript_pin archive path (nothing has written version 1, and its reader read a possibly-remote path from the local filesystem) along with the endOffset thread it was the only caller of, and adds Archived:/Liveness:/Worker: lines so a released archive read no longer prints identically to a live one. |
||
|
|
58efd45c33 |
fix(federation): negotiate structured read by method, not advertisement
- close an exited remote worker's terminal before labelling the release closed_exited_terminal; a host-certified exit keeps the verdict when the kill stops nothing - replace the structured-read and fleet-snapshot capability probes with the optimistic call plus method_not_found, so hosts that serve federationReadOutput without advertising it stop downgrading to a scrape - drop forceProbe so an unchanged peer at an unchanged epoch probes once - distinct fleet reasons for home-side budget exhaustion and peer_changed; keep a host-supplied unverifiable reason and lastObservedAt for exited - delete three compile-time-true self-capability checks, the advertised but never-read federation-release capability, and decode the pull page with zod |
||
|
|
92e1124c78 | Merge branch 'fix-send-federation-transcript' into integrate-fixes | ||
|
|
994f01c20a | Merge branch 'fix-ergonomics' (early part) into integrate-fixes | ||
|
|
80024119a9 | Merge fix-worker into integrate-fixes | ||
|
|
ca53fff0b7 | test(orchestration): pin the settled arm of the legacy recovery plan | ||
|
|
6f20b25663 |
refactor(orchestration): name the worker-release harness as test support
The harness is test-only code sitting among the rpc/methods production modules; *.test-support.ts is the repo's existing marker for that. The five folds the review asked for are not possible: every merge target is already at 292-300 effective lines against the 300 cap. |
||
|
|
6d91a1c438 |
fix(orchestration): fence automatic resume for a settled, unreleased worker pane
Between worker_done and release the pane still holds a resumable provider session, but listLegacyWorkerTerminalRecoveryRows selected live worker states only, so nothing fenced it: reopening the workspace after a restart re-ran `codex resume <session>` on a finished worker. Settled workers whose terminal resource is still owned and neither released nor retained now join the recovery rows, the plan marks those panes settled so they never compete for adoption, and any fence the plan no longer claims is lifted — on release, retain, user takeover and dispatch prune. An unreadable plan fails closed and lifts nothing. Ported from PR #17651, adapted to this branch's worker_terminal_resources ownership model. |
||
|
|
b1c56e2cbe |
fix(orchestration): stop worker receipts from contradicting themselves
worker-show spread the raw worker row beside its parsed copies, so a reader got residual_resources (a JSON string) next to residualResources (an array), plus host_scope as JSON-inside-JSON and two authority hashes with no consumer. Parse once, emit camelCase once, and withhold the hashes. worker-show also published only PTY liveness, so an agent that died at a trust prompt read live there while worker-list called it unverifiable -- and worker-list's nextAction pointed back at worker-show. Both now publish the same fleet projection. worker-list's projection.resource restated fields the row already carried, and the unconfirmed-stop sentence doubled a terminator on an already-punctuated reason. |
||
|
|
3053efe9bd |
test(orchestration): pin the watermark, restart-scan, lifecycle and downgrade fixes
Adds regression coverage that fails on
|
||
|
|
f9d95574e3 |
fix(orchestration): project fleet liveness once, on the evidence clock
Fleet liveness measured staleness against status.receivedAt, the delivery clock a relay reconnect restamps, so an hour-stale agent read live after every reconnect; it now uses evidenceObservedAt when the host supplies it. worker-attention-context kept a second, divergent copy of that projection which ignored host scope entirely and could not say exited. It calls the shared projectLiveness now, with hostScope carried on WorkerAttentionFacts. refreshOrchestrationFleetLivenessAttention re-derived categories from liveness alone and dropped the unverifiable an unproven outcome contributed; it re-runs the one attention projection instead, using the outcome now exposed on the worker. The byte-identical active-sibling predicate is one shared fragment. |
||
|
|
4da7866b41 |
fix(runtime): re-check the prompt binding on every replay and warn on a swallowed Enter
- compare the recorded terminal binding on every prompt replay, not only the --wait-submit ones, so a replayed receipt cannot claim observation:supported for an incarnation that is gone - normalise `terminal` to a pane key in replayStableCallerParams so a re-minted handle replays instead of failing request_mismatch - warn (exit 0) when a supported send never reached turn_started, naming --retry-request <id> --wait-submit <seconds> - collapse the write-only stages to input_accepted | turn_started - move the PTY-keyed correlation state into AgentPromptRequestCorrelation (pty-keyed arrays, not NUL-joined string keys); the abandoned scan in lifecycle claim allocation now continues instead of returning |
||
|
|
5974c500e0 |
fix(orchestration): release the mailbox watermark with the DB reservation
The watermark was set before stageMailboxPointerEnter, and neither the staging-failure nor the DB-error exit cleared it, so a mailbox whose claim was stolen parked every later delivery forever. Also restores the restart scan's pointer-pending union and dispatch: handles, aligns the pointer-enter index predicate with its query, adds messages.pointer_enter_pending to the v33 skew columns, makes v34's mailbox_handle downgrade-safe, and completes the lifecycle graph. |
||
|
|
0719a459f9 |
fix(orchestration): reject worker-start --terminal on the coordinator's own pane
A coordinator adopted as its own worker answers its own dispatch preamble; the --terminal target is now rejected by handle and by resolved pane key with a typed terminal_is_coordinator error naming the next action. |
||
|
|
dcc540102d |
fix(orchestration): give a dispatched worker a concrete follow-up read cadence
The dispatch address is durable but never interrupts a worker, so a preamble that only listed `check --terminal` produced workers that never read one coordinator follow-up. Name the checkpoints and the pre-worker_done read. |
||
|
|
233fb0e92e |
fix(orchestration): derive worker terminal state once and page by rowid
The SQL CASE in worker-terminal-inventory-counts duplicated deriveWorkerTerminalListState and diverged on an owned, unreleased resource with no worker_dispatches row: SQL returned NULL where TS returns 'active', so worker-list --terminal-state active dropped a row it labelled active. The SQL copy is gone; filtering and counting now read the one TS projection. worker-list also ordered by COALESCE(w.created_at, d.created_at) while the fence pinned d.rowid, so a worker registering between pages re-emitted its row. Order and fence now share d.rowid, and a pinned filtered cursor counts over its own row extent instead of a live scan. |
||
|
|
08bea361f9 |
fix(orchestration): never settle a worker terminal the dispatch no longer owns
worker-release's dead-process shortcut called settleDeadWorkerTerminalRelease without requireArchive, and the guard only excluded ownership_state='released', so a user takeover, an external terminal, or a transferred resource was marked released and its output archive was lost. Both release guards now read one (ownership_state, release_state) -> action table, and settlement always requires the durable archive. |
||
|
|
6f2dfa7e3e |
Rehome orchestration-v3 runtime and rpc changes into main's split modules
Main split OrcaRuntimeService, rpc/methods/orchestration.ts, and the runtime test harness while this branch was open. Move the branch's prompt-request correlation, PTY liveness history, fleet snapshot, pointer submit target, federation pairing-revision, and mutation replay-nudge changes into the split modules; split cli/handlers/terminal.ts under the max-lines bar. |
||
|
|
07fa28a3ae |
Merge origin/main into orchestration-v3
Runtime and rpc hunks from the pre-split monolith still need rehoming into main's split modules; terminal.ts max-lines follow-up pending. |
||
|
|
6c4797ca9f |
perf(runtime): stop the expired-SSH-lease sweep from rescanning every tab layout (#18409)
* perf(runtime): stop the expired-SSH-lease sweep from rescanning every tab layout The `runtime:syncWindowGraph` IPC handler is the most expensive thing the main process does: measured on a real session it costs 20.7 ms per call at 0.71 calls/sec, which is 1.47% of wall and ~17% of all main-thread JS. 76% of that sits in one subtree: `getHydrationTargets` -> `hasRuntimeOwnedPtyCandidate` -> `getRecentExpiredSshLease` -> `findTerminalTabIdForLeaf`. Three pieces of pure waste, none of which change an answer: 1. `getRecentExpiredSshLease` evaluated its cheapest and most selective filter LAST. `SSH_PANE_RECOVERY_GRACE_MS` is 30 s, so nearly every stored expired lease fails it — but only after the predicate had already resolved the lease's leaf to its current tab, which is the expensive part. The freshness and reattach-eligibility gates now run first; the predicate is otherwise identical and side-effect free, so the selected lease is unchanged. 2. The sweep ran once per tab. `workspaceSessionWorktreeHasRuntimeOwnedPtyCandidate` asked "does a recent expired lease name THIS tab" for every tab in a worktree, and each ask re-read and re-filtered the whole lease list. It now resolves the worktree's recoverable tab ids once, lazily, so a worktree whose first tab already owns a serve/SSH pty still never sweeps. 3. `findTerminalTabIdForLeaf` allocated a `Set` and walked a whole pane tree per tab to answer one leaf lookup. It now reads a leafId -> tabId index built once per layouts record and reused until a layout object is replaced, which keeps first-tab-wins ordering identical. Measured by replaying a real 414-worktree / 801-tab / 137-lease session: 2.51 ms -> 0.27 ms per publish for this subtree, a 9.3x cut. No user-facing trade-off: same leases selected, same tabs reported recoverable, same SSH pane recovery affordance. * fix(runtime): revalidate the leaf membership index on root identity persistPtyBinding grafts a leaf by assigning `layout.root` on the SAME layout object inside the SAME layouts record, so the layout-identity revalidation kept serving an index blind to the grafted leaf and findTerminalTabIdForLeaf answered `undefined` where the pre-index linear scan answered the tab. That fed the SSH reattach fence (restoreReattachedPtyRuntime) and the expired-lease pane recovery resolver, both of which then fall back to the frozen lease tabId. Membership is a pure function of the root tree and no writer mutates a node in place, so root identity is the exact revalidation key — same O(tabs) pointer compare, no new cap, cadence or staleness window. * perf(runtime): resolve a leaf's tab by scan instead of a cached membership index Fix #3 of this PR cached a leafId -> tabId map per layouts record and revalidated it by comparing every root reference on every read. It was the only mutable cross-call state in the change, the only piece carrying a staleness invariant, and it had already needed one follow-up fix (1c23c544) after a layout-identity key turned out to be blind to `persistPtyBinding`'s in-place `layout.root` graft. The index was never what produced the measured win. After fix #1 moves the freshness gate first, the reporter's replay never calls `findTerminalTabIdForLeaf` at all — every stored expired lease is older than the 30 s recovery grace, so the entire 24.9 ms -> 0.9 ms comes from fixes #1 and #2, both of which are unchanged. `findTerminalTabIdForLeaf` is now an allocation-free scan over the existing `layoutContainsLeafId`, which short-circuits on the first matching leaf instead of materialising a Set per tab. Same answers, same first-tab-in-record-order semantics, no revalidation key, nothing for a writer to invalidate. Re-measured on the same 414-worktree / 801-tab / 137-lease replay (process.cpuUsage deltas, median of 3; wall clock is useless on this box): scenario main index scan all leases stale (replay) 24.86 0.87 0.88 ms/publish one lease inside the grace 24.31 1.08 1.04 ms/publish all 137 inside the grace 18.04 3.71 4.31 ms/publish The measured win is unchanged. Only the synthetic worst case — every one of 137 leases expiring inside the same 30 s window — pays for the cache's absence, and even there the two ranges overlap because the index's own revalidation is O(tabs) per lookup. Removes 208 net lines. `terminal-leaf-tab-resolution.test.ts` keeps the parity cases and adds the guard the cache needed: a subtree replaced in place after an earlier read must be visible to the next one. That test fails against the index. * docs(runtime): say why the leaf scan keeps Object.keys 'Allocation-free' overstated it — Object.keys does allocate one key array. A guarded for...in trades that for a hasOwn call per tab and measures slower, so record the reason the next reader does not re-litigate it. |
||
|
|
ef9e9f3fd9 |
perf(main): take the idle ownership poll off the main thread and batch pending marker probes (#18425)
* perf(main): take the idle ownership poll off the main thread and batch pending marker probes The runtime-metadata ownership watch ran existsSync + readFileSync + JSON.parse on the main thread every 10s for the life of the process. Move it to fs/promises with an ENOENT catch (dropping the existsSync pre-check, a TOCTOU race anyway) and guard overlapping ticks. The base-directory poller's pending `.git` marker probes ran serially, costing D x latency per tick for up to 300 ticks. Route them through the same forEachWithConcurrency bound the full scan already uses. * test(runtime): pin that a shutdown-straddling ownership read cannot republish CodeRabbit flagged the async read resuming after stop(). The cleared activeTransports guard already neutralizes it; this test pins that guard rather than the interval teardown. |
||
|
|
07e50e9513 |
perf(terminal): scan only new tail lines for the wait-blocked sentinel (#18437)
* perf(terminal): scan only new tail lines for the wait-blocked sentinel The wait-blocked scan must prove a signal is ABSENT, so it could not early-exit and re-tested all 2000 retained lines with a 13-alternative regex on every scan (20/s per streaming PTY) even though only ~20 lines were new. Index the matching line indices per tail-array identity and carry them across appends, testing only the lines each append produced. Also carries the retained character total and the redraw prefix's right-trimmed state across appends, so a saturated tail is no longer re-summed and re-scanned per chunk. * perf(terminal): build the carried tail window and its match index from one constructor |
||
|
|
34222e0137 |
perf(orchestration): project explicit columns so the graph publish stops recompiling SQL (#18420)
* perf(orchestration): cache the prepared statements the graph publish recompiles SyncDatabase refuses to cache any `SELECT *` — node:sqlite can build the first row after a schema change from stale column names — so every wildcard read in the orchestration DB recompiles its SQL on each call. The graph publish runs that fan-out once per pane, ~0.7 times a second, forever. Add a per-connection prepared-statement cache scoped to the orchestration DB, whose schema is frozen in the constructor (createTables/migrate/trigger) and whose resets are DELETE-only, and route the buildByPaneKey -> getForHandle -> getRecent path through it. 5 publishes over 2 panes: 30 compilations -> 2. * perf(orchestration): project explicit columns so the existing cache covers the hot path Replaces the branch's second statement cache. The six graph-publish reads were uncacheable only because they were spelled `SELECT *` / `SELECT t.*`, which SyncDatabase refuses to cache (node:sqlite can build the first row after a schema change from stale column names). Spelling the projection out from type-checked column tuples makes them cacheable by the SyncDatabase LRU that is already merged, already bounded, and already clears on DDL — so the WeakMap and its documented cross-connection ALTER hazard both go away. Drift is caught at build time: `satisfies readonly (keyof Row)[]` plus an `Exclude<keyof Row, Cols[number]> extends never` assertion pins list vs type at tsc, and a PRAGMA table_info test against a freshly migrated OrchestrationDb pins list vs schema. Same win, verified: 6 compilations per publish -> 2 total then 0, identical to the WeakMap branch; 92/96/91 us CPU per 2-pane publish before, 11-12 us after on both. |
||
|
|
5d8532f6d3 |
fix(worktrees): resolve the execution host at both worktree-create entry points (#18545)
Two entry points create the same workspace and disagreed about how to read its host. `orca-runtime-create-managed-worktree.ts:63` resolved through `getRepoSshConnectionId` and then normalized the row; the `worktrees:create` IPC handler branched on raw `repo.connectionId` (`register-worktree-create-handlers.ts:66-69`). So a repo naming its owner only as `executionHostId: 'ssh:<target>'` created remotely through the runtime and ran `git worktree add` on the client against a remote path through IPC (#11163). Same repo, two entry points, different answers. Both now take one route, resolved through the existing layer (`getRepoExecutionHostId` -> #18296's `resolveGitRouteForHost`). No new resolver. The row normalization on the `ssh` variant is kept, and it is a **workaround, not the pattern**. `createRemoteWorktree` and its callees re-read `repo.connectionId!` at five depths in `ipc/worktree-remote.ts` (1627, 1847, 1848, 1865, 2029), so the resolved connection has to reach them through the field they already read. It travels only as far as that object does — anything downstream that re-reads the row from the store still sees the unnormalized one, and it cannot express the `runtime:` refusal on its own. Proper fix, deliberately not done here: give that pipeline an explicit connection parameter and delete `repo.connectionId!` from it so every reader becomes a compile error, the technique #18307/#18325 used. That is a change inside a 2800-line module plus its callers, and it wants its own PR. Three answers that used to collapse into one, now distinct at both entry points: - `executionHostId: 'ssh:*'` with no `connectionId` -> that SSH host (IPC used to create locally); - `executionHostId: 'local'` with a surviving `connectionId` -> local, since a local row cannot nest an SSH namespace. This is what `getRepoSshConnectionId` and therefore the runtime sibling already answered; IPC used to go remote; - `runtime:<env>` -> refused. Its worktree is created by that environment's own server and the SSH target on its repo row is that server's nested one, addressable only as (environmentId, targetId). The renderer already routes runtime-environment creates over `worktree.create` RPC rather than this IPC channel, so reaching either entry point with one is a routing mistake. Matches `workspace-cleanup-git-route` and `runtime-git-command-target`. Folder-workspace creation is untouched on both sides: it is a registration, not a filesystem create, so the route is resolved after that branch on the IPC side, and on the runtime side only the agent trust write consumes it — where a `runtime:` host now yields `null` instead of the nested target, so the write stops going to a same-named target in this client's table. No wire or persistence change: the normalized row is a local value passed to the create pipeline, never stored, and `CreateWorktreeResult` is untouched. |
||
|
|
3c91404820 |
fix(worktrees): route managed worktree removal by resolved execution host (#18529)
`removeManagedWorktree` resolved its host once — for the metadata prune (`cleanupHostId ?? getRepoExecutionHostId(repo)`) — and then read raw `repo.connectionId` for every step that touches the filesystem: the `git worktree list` deciding whether the path is registered, the provider handed to the unregistered-removal branch, the registered-remote-vs-local fork, and the PTY/history teardown. One function, two spellings. For a row naming its owner only as `executionHostId: 'ssh:<target>'` — the exact class #18296 names — the list ran on the client against a remote path, `removeRuntimeUnregisteredWorktree` was entered with `provider: null`, and the metadata was pruned under `ssh:<target>` while a same-named *local* directory was the one considered for deletion (#11163). #18358 made this reachable: it migrated the cleanup scan, so `executionHostId`-only rows now surface as removable candidates, but removal did not move with it. Routing is now one answer for the whole removal, taken from the host the prune already used, through #18296's host-keyed dispatch. The ambiguous `provider: SshGitProvider | null` carrier is deleted from the callees rather than supplemented, so every remaining reader is a compile error in the typed modules that do the destructive work (`runtime-unregistered-worktree-removal`, `runtime-registered-remote-worktree-removal`, `runtime-worktree-filesystem`). The orchestrator itself carries `@ts-nocheck` from its mechanical split, so that guarantee does not reach it — tests cover it instead. Every change is in the refusing direction; nothing became more aggressive: - an `ssh:` host with no registered provider throws instead of deleting a client-side path (`requireSshGitProvider` already threw for rows that spelled the same host as `connectionId`); - `runtime:<env>` throws rather than dialling a same-named target in this client's namespace, matching `workspace-cleanup-git-route` and `runtime-git-command-target`; - the folder-workspace teardown resolves its connection instead of reading the raw field, so a `runtime:` row stops dialling the wrong namespace. No wire or persistence change: `removeWorktreeMetadataAndHistory` already took the resolved host, and the removal RPC result shape is untouched. |
||
|
|
98e77ef1a7 |
feat(mobile): structured native Codex chat (#18074)
* feat(mobile): finalize structured native Codex chat * fix(mobile): close structured chat lifecycle gaps * wip(mobile): fence stale structured inventory and bound operation-id retention Fence local structured-session inventory and subscription responses with a sync generation so a toggle-off clear, reconnect restore, or retry cannot apply a mirror from a superseded instance. Bound mobile ambiguous operation-ID retention at 128 with unmount cleanup. Staged on the reconcile branch only: the sync module is now 312 lines and needs a real split before this can reach the PR head. * fix(ci): split the structured session-tabs sync and give static analysis mobile types The local structured session-tabs sync module outgrew the 300-line cap once it took on generation fencing, so split it along its real seams instead of raising the cap: the generation/cursor fence, snapshot projection, snapshot apply, inventory refresh, and the subscription loop. The original path stays as a barrel so no importer moves. Repoint the host-session-mirror settle census at the apply module, which owns two receipts now — the snapshot it mirrors in, and the toggle-off teardown that retracts what it published. The teardown receipt is named rather than anonymous so the pin says which direction it settles. The changed-code quality gate lints mobile files and resolves their types from mobile/node_modules, but mobile is a separate pnpm project that the root install never populates, so every mobile type degraded to an `error` type and the gate reported phantom findings. Install mobile dependencies in static analysis when the diff touches mobile, gated on a new classifier output. * fix(mobile): let a slow capability handshake still reach connected The mobile capability update is an advisory whose result is discarded, yet an unanswered one was fatal while an explicit rejection was tolerated. A 5s timeout on the direct client force-closed the socket, and on the relay path it failed `confirmResume` before `connected` was ever published, so a consistently slow link redialled forever. Both paths now share one helper that settles every ambiguous outcome (timeout, mid-flight drop) like a rejection and rejects only when the frame never reached the wire — the one case nothing else recovers from, since the socket's own desync force-close is gated on already being connected. The generation guard still keeps a replaced session from connecting. Retained structured-session operation ids were capped at 128 with oldest-first eviction, but every retained id belongs to a send whose outcome is unknown, so eviction turned a user's retry into a second message on the host. Bound the map by expiry against the id's own embedded timestamp instead, mirroring the host's operation ledger, so no id is released while the host would still honour it. Also give the mobile CI install the root install's lockfile drift guard (mobile's lockfile carries patchedDependencies a silent rewrite would drop), gate mobile_dependencies on should_run, and key the pnpm store cache on both lockfiles. * refactor(mobile): extract the relay pending-request registry The merge composed two independently-sized changes — this branch's capability handshake settle and main's dial-stage tracking — pushing the relay session file to 304 lines against a 300 cap. Neither side broke it alone. Move the in-flight request registry (id generation, tracking, settlement, and reject-all with its delivery-ambiguity marking) into RelayPendingRequests, matching the existing collaborator pattern alongside RelayDialStageTracker and RpcSessionLivenessWatchdog. No behavior change. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
95eed52801 |
fix(cli): report which hosts a worktree listing covered, and stop the cap starving remote ones (#18417)
`orca worktree list` returned zero of 24 SSH worktrees at the default limit (#18104). Rows are resolved repo by repo, so every SSH repo's rows land contiguously at the end of the fleet order — the 24 remote rows sat at indices 496-520 of 521 and a plain `slice(0, 200)` never reached them. The omission was not fully silent: text output printed `truncated: showing 200 of 521` and JSON carried `totalCount` / `truncated`. What was missing is that the omission was *categorically every remote host* — no host column, no `hostScope`, nothing to distinguish "200 of 521" from "one host is entirely absent". Per docs/reference/ssh-execution-boundary.md, a listing that does not name its scope reads as absolute. Adopt the mechanism `terminal list` already has rather than inventing a second one: - `RuntimeTerminalListHostScope` becomes an alias of a shared `RuntimeListingHostScope`, now also carried (optional, so old hosts are unaffected) on `worktree.list` and `worktree.ps` results. - `src/shared/host-balanced-listing-page.ts` round-robins the row cap across hosts and returns the survivors in the caller's original relative order, so the page stays a subsequence of the unbounded listing and nothing downstream re-sorts. An uncapped listing is returned unchanged. - `worktree list` / `worktree ps` text output gains a `host=` column and the same trailing `scope:` line `terminal list` prints. Third defect, same mechanism: `hostScope.omittedHostIds` is built from the runtime's own bookkeeping, so it names `runtime:` ids for servers that are no longer paired — 6 of 9 in the recorded QA run hard-error when queried. Since `hostScope` is *the* documented way to complete a partial listing, that makes the mechanism unreliable for its intended use. Annotate rather than filter. Dropping an id would shrink what the listing admits it did not cover, and the boundary doc requires a listing to name its gaps — the gap is real whether or not this machine can name the host that owns it. `src/cli/omitted-host-scope-selectors.ts` resolves each omitted id against this machine's pairing store and the runtime's SSH-target registry and attaches the exact flag that reaches it, or `null` marked "not selectable from this machine". This is a client-side annotation: nothing new goes over the wire, it answers "can I select it" and never "is it up", and the SSH round trip is only paid when an `ssh:` host was actually omitted. No `--host` filter was added; the host column plus scope line covers the reported need without a new selector axis. |
||
|
|
a35451f5b9 |
fix(relay): stop self-closing the control socket on unknown messages (#18400)
* fix(relay): stop self-closing the control socket on unknown messages
The desktop control client tore its own relay control WebSocket down with
code 4401 "unknown control message" for any well-formed control frame it
did not recognize. handleMessage() funneled everything that was not
ping / conn-open / drain / a tracked request reply into
failProtocol('unknown control message'), which closes the socket and
orphans the origin.
Three real frames hit that branch:
- A relay reply that arrives after the desktop's 10s request deadline
already deleted the pending entry. Relay control operations run DB
transactions that can exceed 10s under load, so resolveMessage() finds
no waiter and returns false.
- A control-error carrying no reqId (or an unknown one), including the
relay's own 'unknown_control_message' reply to a host command it could
not route.
- A newer relay's opcode that this build predates.
Fleet telemetry shows ~15 of these closes per day across app versions
1.4.175..1.4.197, so it is version-agnostic. The self-close was also far
more costly than the message that caused it: the relay session dropped to
'orphaned' and answered the phone with HOST_OFFLINE (4404) for the orphan
grace window, then the desktop had to re-register through the director's
503 reconnect throttle, stretching a single stray frame into minutes of
mobile downtime.
Per docs/reference/remote-wire-compatibility.md Rule 2, an unknown but
well-formed control frame must be dropped, not treated as fatal. Log and
ignore it; malformed JSON, binary frames, and messages before activation
still close as protocol violations.
Adds unit tests for the unknown-opcode drop, the timed-out-reply drop,
and the preserved malformed-frame teardown.
* docs(relay): correct the ignore rationale, drop the Rule 2 misattribution
Rule 2 of remote-wire-compatibility governs the SENDER of a new terminal-
stream opcode and treats the receiver's silent drop as a hazard, not a
mandate. Reframe the comment around the actual justification: the decoder
convention of dropping unknown frames, the control channel's lack of an
opcode negotiation step, and the incident cost asymmetry.
|
||
|
|
4cc0b8de61 |
perf(hot-paths): delete allocation-only work in sort, explorer, monaco, rpc, snapshots (#18372)
* perf(hot-paths): delete allocation-only work in sort, explorer, monaco, rpc, snapshots * fix(perf): revert snapshot revision fast-path — same revision can carry a new session * perf(hot-paths): drop the unproven rpc buffer rewrite, dedupe the equality helpers - Revert the unix-socket chunk-carry change. Its comment claimed it avoided O(n^2) rescans, but chunks is reset to [remainder] every data event, so the join plus the tail byteLength is two passes where the old code did one; benchmarks showed no win. It also moved consumed-frame bookkeeping out of the closure, so a synchronous throw from the handler would re-dispatch frames. - project-host-compatibility: fold the two byte-identical array comparators into one generic arraysEqualByJson. - smart-attention: drop the leftover byTab.size === 0 branch that returned the same value as the line after it. |
||
|
|
d05dd8ef50 |
fix(source-control): route hosted reviews by resolved execution host (#18382)
`ForgeProvider.createReview(repoPath, input, connectionId, options)` and the `connectionId` on `ForgeProviderRepositoryContext` carried the same collapse the five prior migrations closed: `string | null` spells "genuinely local", "runtime host" and "could not resolve" with one value. Because it was decided two layers up -- `repo.connectionId ?? null` at the `hostedReview:*` IPC handlers and in `RuntimeHostedReviewCommands` -- a row naming its owner only as `executionHostId: ssh:<target>` ran the whole review path against this machine's copy of a remote path (#11163): `git rev-parse`, `git status`, the base-on-remote ref probe, the upstream divergence read, and `gh`/`glab` with no host flags. Replace it with a required `ExecutionHostId` threaded from the decision point through the contract, routed by #18296's `resolveGitRouteForHost`. The parameter is removed rather than added beside, so all five implementations -- GitLab, GitHub, Bitbucket, Azure DevOps, Gitea -- and every caller became a compile error. None of these families carries `@ts-nocheck`, so unlike #18325 that guarantee is real here; `orca-runtime-file-commands.ts` does, but it only constructs `RuntimeHostedReviewCommands` with unchanged deps. Also fixed at the sites: - The branch cache scoped entries on `connectionId ?? ''`, so two rows at one path on different hosts shared one cached review, one backoff deadline and one invalidation. Keyed on the resolved host now, as #18377 did for its probe key. - `hostedReview:create` resolved shared symlink paths and normalized worktree paths off the raw field, so an `executionHostId`-only SSH row read `orca.yaml` and `resolve()`d a remote POSIX path on the client. Those ask the file-holder question -- `getRepoSshConnectionId` -- not the dialable one. - An SSH host with no provider now refuses inside the git-state layer instead of reaching the local branch, keeping "remote and unreachable" distinct from "local" (docs/reference/ssh-execution-boundary.md). `runtime:` is a routing mistake inside `hostedReviewSshConnectionId` -- that environment's server runs its own git, and the SSH target on its repo row is nested in that server's namespace, so dialing it here reaches a same-named box of ours. But store-backed callers ask `getRepoHostedReviewExecutionHostId` first, which is "what may this client dial" and answers `local` for a `runtime:` row. That is deliberate and matches #18377: the runtime registration controller only adopts a `runtime:` stamp onto a row with no `connectionId` (`runtimeRepoMatchesExecutionHost` refuses to match an SSH row), so the checkout really is in this process and refusing would regress a runtime server creating reviews for its own rows. No wire change. `connectionId` on `CreateHostedReviewArgs`, `CreateStackedHostedReviewArgs` and `HostedReviewCreationEligibilityArgs` in src/shared/hosted-review.ts is untouched -- every host already ignores it in favor of the repo row, and removing it from the request types would only churn the schema older clients still populate. The main-side eligibility input `Omit`s it so nothing on this side can read the ambiguous field again. |
||
|
|
573537ecd4 |
feat(cli): make terminal close the canonical workspace teardown (#18073)
* fix(runtime): recover stale session owners and await retirement * fix(runtime): preserve session hydration and smoke compatibility * test(runtime): cover empty and unindexed session owners * feat(cli): make terminal close the canonical workspace teardown * fix(preload): align ssh termination result type * test(runtime): assert folder hydration owner * fix(runtime): fence legacy terminal stop by worktree host * fix(preload): reconcile ssh result import with main * fix(runtime): keep same-id sibling hosts out of workspace close The stale-owner fallback in the session controller re-routed any worktree whose catalog partition had no tabs to whichever other partition held tabs. Only `runtime:` environment ids rotate across relay restarts; `repoId::path` legitimately repeats across hosts, so an SSH workspace close could retire the local copy's tabs and resume records, or flip owners mid-close and strand the SSH PTY. Restrict the fallback to runtime hosts, and pin the session partition once per workspace close so record clearing targets the partition that owned the tabs. * test(runtime): give the cross-host close fixture a real resume record * fix(preload): take main's ssh-bridge import order so the merge stays duplicate-free |
||
|
|
316ec38f67 |
fix(repos): route icon and remote-identity probes on a resolved execution host (#18377)
`detectRepoIcon`, `detectRepoIconAndUpstream`, `detectGitHubAvatarIcon`, `detectRepoFileIcon` and `probeGitRemoteIdentity` took a `connectionId`-shaped parameter threaded down from their callers. That shape spells "runtime host", "unresolved" and "genuinely local" all as one falsy value, and because it is a *parameter* each caller decided independently what to pass — a wrong answer was invisible at the boundary. Replace it with a required `ExecutionHostId` and route through #18296's `resolveGitRouteForHost` / `resolveFilesystemRouteForHost`. The parameter is removed rather than added beside, so every caller became a compile error. No new resolver, no wire change: nothing these modules return carries a host id. Fixed at the call sites: - `repo-git-remote-identity-enrichment` read `repo.connectionId` raw, so a row minted with only `executionHostId: ssh:<t>` ran `git remote -v` against this machine's copy of the path (#11163), and a `runtime:` row handed its *nested* SSH target to this client's dispatch table — a same-named box of ours. - Its location key had the same collapse, so two rows at one path on different hosts shared a probe, an abort controller and a backoff deadline. - `runtime-repository-fork-backfill` guarded on `repo.connectionId`, so an `executionHostId`-only SSH row had its upstream read off the client. `runtime:` is refused inside the modules (this process does not execute another environment's git or filesystem), but store-backed callers ask `getSshTargetIdForExecutionHost` — "what may this client dial" — so a `runtime:` row keeps the probe this process has always run for it. Registering and cloning stay `local` on purpose: those controllers do the filesystem work here, whatever host id is stamped on the row (see `assertCloneHostIsSupported`). |
||
|
|
968dbd905f |
perf(renderer): take the English catalog and the xterm WebGL addon off the boot graph (#18326)
* perf(renderer): take the English catalog, xterm WebGL addon and emoji data off the boot graph The renderer's boot graph — the entry chunk plus its 331 modulepreload links, all fetched and evaluated before first paint — carried three payloads nothing needs at that moment. `en.json` (644 KB) was an eager i18next resource, but every renderer string goes through `translate(key, fallback)` and `en` resolves that inline default, so most of the catalog was dead weight. The renderer now bundles a generated `en-runtime-required.json` holding only the 2,583 of 13,828 entries a default cannot reproduce: plural-suffixed keys, keys whose catalog value differs from a call site's default, and keys no call site references with a literal default. `en.json` stays the translator source and the input to the four lazy catalogs. `@xterm/addon-webgl` (243.6 KB) and `emojibase-data` (170 KB) are now primed right after the React root renders instead of statically imported. The load stays eager and `attachWebgl` stays synchronous — it reads the resolved constructor — so no terminal ever falls back to the DOM renderer for a frame. `isPluginPanelTabKey`/`isQualifiedPluginKey` move to schema-free sibling modules, re-exported from `plugin-manifest.ts`. This evicts the plugin manifest schema graph from the boot chunk but measures ~0 KB, because six other shared modules still put zod on the boot path. Boot graph: 332 chunks / 5107.2 KB -> 336 chunks / 4161.5 KB (-945.7 KB, -18.5%). A new ratchet parses the built index.html and fails if `en.json`, `@xterm/addon-webgl` or `emojibase-data` is preloaded again; it runs at the end of every `build:electron-vite`. * chore(i18n): pin the generated English subset to LF and mark it generated * fix(i18n): make the runtime-catalog gate merge-robust and prime emoji data in tests CI builds the merge of a PR with main, so a byte-for-byte comparison against a committed generated file fails the moment any unrelated PR adds a translate() call — which is what happened here. The check now asserts the property that actually matters instead of byte equality: every runtime-required entry is shipped, and nothing shipped disagrees with en.json. Entries that stopped being required are dead weight, never a wrong string, so they are reported and tolerated. Failures now name the offending keys rather than saying "stale". The generator itself was already deterministic (plain code-unit sort, no locale collation, order-independent set construction); a test now pins that a reversed call-site walk produces byte-identical output. Test fixes for the catalog prune and the deferred emoji load: - browser-search / NativeChatSupportedAgents asserted key presence on the renderer's runtime resource. The durable contract is en.json — the renderer deliberately no longer bundles entries a call site default reproduces — so they assert against the translator catalog. - Four emoji tests typed a shortcode in the same tick as mount, before the catalog the hook primes on mount resolves. Not reachable by a human; the tests now await the prime. * revert(renderer): keep the emoji shortcode catalog statically imported Deferring emojibase-data introduced a window that did not exist before: until the dynamic import settled, getPrimedEmojiShortcodeEntries returned [], so exactShortcodeIndex built an empty map and replaceCompletedWorkspaceEmojiShortcode returned null — leaving a typed `:wink:` in the field literally, and persisting it as the workspace display name. Pre-change the shared catalog was statically imported, so the first call at any tick returned full data. The window is reachable by anything that dispatches input in the same task as the field's mount effect — Playwright/CDP in the e2e suite and agent automation both do, and the WorktreeMetaDialog test failure was exactly that, producing 'Feature 😉' instead of 'Feature 😉'. Nothing that resolves a shortcode can be async without that race, and a wrong persisted name is not an acceptable trade for 166.7 KB, so the deferral is reverted rather than papered over in the tests. The boot-graph ratchet drops its emojibase-data probe and records why. Boot graph: 5108.9 KB -> 4329.9 KB (-779.0 KB, -15.2%), down from -945.7 KB. * fix(terminal): make the deferred WebGL addon load recoverable and refit on late attach Two defects the deferral introduced, neither possible with a static import. A failed load latched the DOM renderer for the whole session. `.then(onOk, onError)` settles fulfilled, so the memoized promise was cached forever with a null constructor: attachWebgl's re-prime got the cached promise back, and resetTerminalWebglSuggestion — the documented "GPU setting changed, retry" path — could not clear it either. The rejection path now clears the memo, latches the queued panes the way a failed construction does so they retry at a recovery boundary rather than every frame, and caps attempts so a genuinely missing chunk is not re-fetched forever. The recovery boundary re-arms it. The queued-attach drain skipped the refit. Every other late-attach path pairs attach with a refit because the grid was measured under DOM cell metrics and WebGL floors the device cell width. Post-deferral, openTerminal's attachWebgl queued and returned, the initial fit rAF then measured DOM metrics and sized the PTY from them, and the addon attached with no refit — a persistently narrow PTY and an unpainted right gutter, not a one-frame flicker. Both paths now go through one attachWebglAndRefit pairing so they cannot diverge again. Regression tests cover both, and each was verified to fail without its fix. The addon-load state machine moves to terminal-webgl-addon-loader.ts and the viewport presentation helpers to pane-viewport-present.ts, keeping pane-webgl-renderer.ts under the 300-line budget without a suppression. |
||
|
|
f2ddf7779f |
fix(ssh): pick the eligible expired lease, not the first one matching a pane (#18366)
`getRecentExpiredSshLease` selected the first `expired` lease matching the pane coordinates and left eligibility to the caller. Only `recoverTerminalPane` asked, and it asks id-qualified, where lease identity `(targetId, ptyId)` already makes the match unique -- so that check could never fire on a lease a different one shadowed. The two unqualified callers never asked at all: `workspaceSessionWorktreeHasRuntimeOwnedPtyCandidate` and `hasRecentExpiredSshLeasePane` both take a bare `!== null`. `(worktreeId, tabId, leafId)` is not unique. `supersedeSiblingLeasesForPane` exists because a pane accumulates leases as it re-leases under new relay ids, and it stamps `supersededBy` on an already-expired predecessor precisely so the predecessor stops counting. Inside the 30s SSH_PANE_RECOVERY_GRACE_MS window those two readers still counted it: a pane whose only recent lease is a superseded or relay-id-recycled corpse was reported as runtime-owned and preserved for recovery, and `recoverTerminalPane` then refuses it. Where an eligible successor also exists, the predecessor is stored first and shadowed it. Apply the existing `sshRemotePtyLeaseAllowsReattach` inside the selection, so the reader answers with the first ELIGIBLE orphan or nothing, and all three callers agree on what `expired` authorizes. `recoverTerminalPane`'s own check becomes unreachable and is folded into the comment on the branch that now covers it. Scope: an over-report in headless/mobile reconciliation, not a wrong-route readoption -- the id-qualified recovery path already refused these leases. No wire change and no host-semantics change: `expired` still says only that the client lost its route, and nothing here asserts a remote shell died. Coverage lives in a `*.test.ts`: config/vitest.config.ts, the config CI runs, includes only `*.test.ts`, so the orca-runtime-tests/*.spec.ts neighbours would never execute. |
||
|
|
d5803bdbc4 |
feat(ssh): host-stamped remote foreground identity (#18078)
* docs: add SSH agent identity implementation plan * feat(ssh): host-stamped remote foreground identity * fix(runtime): preserve unfenced inspect call shape * perf(ssh): traverse foreground descendants linearly * fix(ssh): bound retired PTY evidence records * test(ssh): cover retired incarnation retention * fix(ssh): make remote process inspection total * Split SSH identity build hot spots * Fix process table snapshot module split * test(ssh): update process inspection expectations * docs: drop the SSH identity plan from the PR The design doc does not belong in the product repo; it stays out of the shipped tree while the implementation carries its own comments. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
7ea213cf8c |
perf(main): remove four per-chunk/per-waiter hot-path costs in PTY and terminal-wait (#18315)
Four independent wastes on the main process, none of which changes behavior: - One shared 2s sweep replaces one setInterval per terminal-wait waiter. 20 waiters allocated 20 handles and 10 main wakeups/s independent of output; now 1 handle and 0.5 wakeups/s. Same cadence, same per-waiter checks in the same order, same resolve semantics; the foregroundPollInFlight latch moved into the waiter's poll entry unchanged and each entry still interleaves its own foreground read, so one slow ps cannot delay another waiter. - SIGWINCH's `ps` for Orca's own row is memoized. It reads this process's controlling tty, which is invariant for the process lifetime, and feeds exactly one guard. Exec count per 4-pane tab switch drops 16 -> 8. The call stays synchronous: making it async would reorder SIGWINCH against subsequent writes. - The wait-blocked carry retains chunks with a running char count instead of concatenating and re-slicing a 256KB window on every chunk, and joins once at scan time. runWaitBlockedCheck receives a byte-identical `appended`. - maxUpwardCursorReach no longer compiles a RegExp per redraw chunk, and containsTerminalVerticalLineControl walks with charCodeAt instead of minting a one-char string per position. |
||
|
|
37694d9896 |
fix(memory): close two per-id map reaper gaps and ratchet the pty-exit reaper (#18320)
`onPtyExit` deletes ~25 per-PTY maps but never `ptyLifecycleGenerationById`, so every PTY that ever ran left one entry behind for the life of the main process. Safe to delete because `getPtyLifecycleGeneration` lazily mints from the monotonic `nextPtyLifecycleGeneration` — a re-read after the delete returns a strictly newer number, never a reused one, so no stale frame can be accepted. `warnedLostHandlerPtyIds` outlived the buffered data it describes when the LRU cap evicted that data, and because the warn is once-per-id it also suppressed a legitimate re-warn on a fresh accumulation for that same id. `ambiguousOwnerWarnedWorktreeIds` was a module-scope Set with no delete anywhere, while both worktree teardown paths prune ~20 sibling collections. Not pruning also suppressed a legitimate re-warn for a recreated worktree id. Adds a ratchet that reads every per-PTY-keyed collection off a real runtime instance and requires each to be deleted by the reaper, cleaned by a helper the reaper calls (verified against that helper's source), self-clearing per in-flight operation, or explicitly justified as retained. |
||
|
|
720c3299ba |
fix(ssh): require a host death certificate before recreating a pane, and unstick expired leases (#18013)
* fix(ssh): match an expired lease on where its leaf lives now, not its frozen tab A lease freezes tabId at write time, but detachTerminalPaneToTab moves a live pane, so the stored tab is the one the pane LEFT. getRecentExpiredSshLease required lease.tabId === tabId, which is wrong in both directions: a viewer on a stale mirror matched under the abandoned coordinates (and resolvePersistedStable PaneOwner then reads an empty layout for that tab, so adoptStablePane is skipped entirely and a fresh shell is spawned over a possibly-live one, binding the same leaf in two tabs), while a viewer using the pane's real coordinates matched nothing and got terminal_not_recoverable. Resolve the leaf's current tab the way restoreReattachedPtyRuntime already does and compare against that, falling back to the frozen tabId only when nothing can say where the leaf lives. Both workspace partitions are read because SSH spawns bind into ssh:<target> while reattach binds into local. * fix(ssh): let a proven reattach take an expired lease back to attached #17965 authorized reattach from `expired` but the state machine refused the transition back, so a lease that reattached and proved itself alive stayed `expired` forever. That silently exempted a demonstrably running remote shell from `ssh:reset` (skips `expired`), from the SSH_TERMINATE_RECONNECT_REQUIRED ownership fence in `ssh:terminateSessions` (marks it not-owned), and from the quit-time `detached` sweep, and made it permanently ineligible to win supersession so its own successors never retired their predecessors. Only the id-qualified caller carries per-pty proof: markSshRemotePtyLeases AttachedAsync is fed the relay's `attachedLeaseIds`, so an unqualified bulk mark over a whole target still cannot revive `expired`. `terminated` stays absorbing. Re-entering `attached` drops supersededBy/relayIdRecycled, since route retirement belongs to the shell that lost the pane and this one just proved it is not that shell — the same invariant upsertSshRemotePtyLease enforces. * fix(ssh): make the pane-recovery liveness gate refuse without positive evidence of life The gate refused only `live` and `unverifiable` and passed on `null` — but the register is an in-memory Map, so `null` is equally what a fresh app start, a never-asked host and a certified death look like. Absence of evidence was reading as authorization to spawn a shell over a possibly-live remote process: `!pty.connected` is cleared for every PTY a dropped relay owned, and `expired` only ever says the CLIENT lost its route. - `exited` is now RETAINED rather than deleted, so the register is three-valued in the map as well as in the type. Its one writer is a host-delivered exit frame — an exit with a real code, or an explicit `hostExitConfirmed` — which records the certificate instead of merely dropping the doubt. - `recoverTerminalPane` refuses on `live` and `unverifiable`, and deliberately does NOT demand a positive `exited`. The only answer that ever reaches this gate is a reachable relay reporting no such id, and that is a union: pty.attach throws not-found for an unknown id with no liveness check, and a relay restart makes every previously minted id unknown (ids carry a per-start `ptyIdMintEpoch`). No writer of `exited` co-occurs with a reattachable `expired` lease either — a host-delivered exit frame tombstones the lease `terminated` — so requiring one would close the gate permanently. - `handlePtyReattachFailure`'s not-found branch publishes `code: -1` to the renderer and does not call `runtime.onPtyExit`. The relay's not-found answer is not a death certificate, and #17963's ratchet on the same branch pins that. - The inventory's `observed === false` hunk keeps dropping doubt rather than asserting a death: `pty.listProcesses` returns the relay's CURRENT session map, so a restarted relay omits every previously minted id whether or not those shells died — the same union, one hop away. A live or unprovable pane refuses; a disowned one still recovers. No wire change. The gate's ratchets live in terminal-pane-recovery-liveness-gate.test.ts: config/vitest.config.ts — the config CI runs — matches only `*.test.ts`, so cases placed under orca-runtime-tests/*.spec.ts would never execute. * fix(ssh): gate paired-viewer pane recovery on the narrowed session-gone predicate isSshSessionGoneError landed on the IPC transport, which never calls terminal.recoverPane. The one caller that does — recoverExpiredHostPane in the paired-viewer transport — still triggered on a bare SSH_SESSION_EXPIRED substring, so the identity-mismatch reply (the relay found a LIVE PTY under that id owned by another pane, which is evidence of presence) still asked the HUB to replace the pane, putting a second agent on one transcript. Main already refuses the respawn on that same reply; this makes the two agree. A pane whose shell genuinely died is unaffected: plain SSH_SESSION_EXPIRED still matches. The mismatch reply now surfaces as an error instead of a respawn. * test(persistence): update the reattach ratchet for expired-lease reclaim markSshRemotePtyLeasesAttachedAsync is id-qualified, so a named pty that proved itself alive now returns to attached instead of staying expired. |
||
|
|
08c7152ab6 |
fix(ssh): compare a lease's relay pty id against the pane's app id (#17969)
`getRecentExpiredSshLease` compared the stored lease ptyId (relay form, written through `toStoredPtyId` -> `toRelaySshPtyId`) raw against the runtime's app-form `pty.ptyId`, so `'pty-3' === 'ssh:target@@pty-3'` never held and `recoverTerminalPane` refused every real SSH pane. Normalize with the same tolerant helper the binding reader already uses, now shared as `toComparableRelaySshPtyId`. Switching the path on is only safe on top of #17957 (respawn gated on the runtime liveness verdict), #17965 (`expired` no longer withdraws bindings) and #17966 (supersession and id recycling carry their own marks). `recoverTerminalPane` additionally refuses a lease those marks disqualify, so it acts only on an `expired` lease that means "reattach gave up". The path's outcome is a reattach, not a respawn: `createTerminal` calls `adoptStablePane` first, which attaches attach-only to the retained binding and only falls through to a fresh shell once the host itself answers that the PTY is absent. |
||
|
|
57681ecd09 |
fix(remote): resolve the spawn cwd, the node manager dir, the vault host and the scrollback seed (#17952)
* fix(remote): resolve workspace cwd, mise Node, host scope, and TUI scrollback honestly #15296 relay: a folder workspace id (`folder:<uuid>`) carries no path, so the worktree-id split yielded nothing and $HOME silently won. Resolve the spawn cwd through worktreeId -> ORCA_WORKSPACE_ROOT -> host default, and refuse an agent spawn outright when a folder workspace names a root this host cannot resolve. #11733 ssh: generalize the NVM dotfile scrape into `orca_dotfile_dirs` and drive mise off `MISE_DATA_DIR` / `XDG_DATA_HOME` instead of a hardcoded `$HOME/.local/share/mise`. #13713 ai-vault: an unresolvable workspace host is `unverifiable`, not local. Widen the default scope to every host rather than scanning the client's own history and reporting "No agent sessions found". #6106 terminal: hydration asked the renderer for `scrollback: 0` while an alt-screen TUI was up, which drops the normal buffer's shell history rather than the TUI bytes. Drop the flag; readers already split the two buffers apart. * fix(remote): stop the relay answering host questions for a guest execution host Three findings from review of the spawn-cwd resolver, all the same shape: a path question answered against the wrong host, or with the wrong key. - resolveRelaySpawnCwd refused an agent launch whenever a folder workspace named a root that did not stat on the relay. But relayHostDirectoryExists stats the relay's *own* filesystem, and the relay supports WSL shells, so a folder workspace on a Windows relay launching into WSL now threw where it previously spawned -- contradicting the function's own doc comment, which says an absent path for that exact host pair is a miss, not a refusal. Thread the shell's execution host in and demote the refusal to a miss when the spawn does not run on the relay's filesystem. - requireRelaySpawnCwd's doc claims both call sites route through one resolver so the fence can never be keyed on a directory the spawn won't use, but the fence key was still computed with the non-stripping splitWorktreeId while the cwd used splitWorktreeIdForFilesystem. For a `::workspace:<uuid>` id those disagree by construction, in adjacent lines: the removal fence guarded a path no spawn ever enters. Same defect in shutdownForWorktreePath and the revive path; all three now use the filesystem split. - The remote Node probe expanded `$HOME` and `~/` prefixes out of a dotfile assignment but not `$XDG_DATA_HOME`, so `MISE_DATA_DIR=$XDG_DATA_HOME/...` was used as a literal directory name. Add the case arm, defaulting to the POSIX `$HOME/.local/share` the seed value already uses -- sshd's exec channel usually has no XDG_DATA_HOME at all. |
||
|
|
64dac75d9b |
fix(ssh): stop respawning panes on client-side-only absence evidence (#17957)
* fix(ssh): stop respawning panes on client-side-only absence evidence Three respawn gates acted on evidence weaker than host-attested exit. Per docs/reference/ssh-execution-boundary.md, loss of contact, a failed reattach, an identity mismatch and absence from a client map are all `unverifiable`, never `exited`. Gate 1 (ipc-pty-connect.ts): "belongs to SSH connection" is minted by the id router from a pure client-side string compare, before any relay is asked, and still returned `sessionExpired: true` -> fresh PTY + agent resume. After an SSH target re-adoption the "other" connection is the same machine, so that puts a second `claude --resume` on the transcript the surviving PTY still owns. Now returns undefined with no error, which routes the pane to recoverUnverifiableDirectSshReattach (remount + reattach, no shell restart) and keeps #7661's no-red-toast outcome. Gate 3 (ssh-reconnect-pane-retry.ts): `!tabPtyId` read `tab.ptyId`, which is only the single-pane fallback for legacy attach. It diverges from the real records deterministically: workspace-terminal-reconnect fills ptyIdsByTabId from the leaf map but writes tab.ptyId only when a tab-level id survives, and clearTransientTerminalState nulls tab.ptyId on every hydrated row. Both leave live leaf PTYs with a null fallback field, arming a generation bump onto the fresh-spawn path. Now consults ptyIdsByTabId and the layout leaf map too; a tab with no PTY in any record still retries. Gate 2 (recoverTerminalPane): an `expired` lease plus `!pty.connected` authorized createTerminal. Every writer of `expired` records that the CLIENT lost its route, not that the shell died. Now also requires the runtime's own liveness verdict to be neither `live` nor `unverifiable`, and ssh-relay-session records markPtyLivenessLive at the persistPtyBinding refusal, which is reached only after pty.attach succeeded. See the report for why this branch is currently unreachable for SSH panes. * fix(ssh): let the respawn gate see the relay's own absence answer Gate 3 refused to respawn a pane whose records still named a PTY, which is right for a transport drop and wrong for a killed relay: after the relay is SIGKILLed and comes back, the leaf map still holds `pty2:<dead-epoch>:1` while the new relay answers that it has no such id. #18017's "replaces the pane only when the host proves the session is gone" regressed on exactly that. The gap was not the predicate, it was its inputs. `handlePtyReattachFailure` already distinguishes the three reattach outcomes and only its not-found branch publishes anything — a lost link and an identity mismatch send nothing. But it published `pty:exit { code: -1 }`, and `-1` is the sentinel every reader resolves to `stop_unverified`, so the one branch holding positive host evidence of absence arrived looking exactly like loss of contact. The renderer had no host answer at all, which the gate's own comment conceded. The exit now carries `livenessVerdict: 'exited'` beside the unchanged `-1`, so the code keeps meaning "no provable status" for every existing reader while the verdict rides its own field. A store bridge records those ids in `hostAttestedAbsentPtyIds` regardless of whether a pane is mounted to hear it — during reconnect none is — and the gate stops counting a recorded id the host has disowned. Settled when a PTY answers to that id again, because a redeployed relay renumbers from pty-1. This narrows #17963, which pinned the same exit as unverified on the grounds that not-found cannot separate "verified the pid is dead" from "my session map never had this id". Everything #17963 protects is untouched: `-1` still fails isProvenProcessExit, so the tab is not closed, the pane's leaf binding is not dropped on exit, and markUnverifiedPtyLoss still fires. Only the reconnect respawn gate reads the new field, and only for an id whose sole channel — the relay that answered — has disowned it, which no client can reach again under any verdict. That is the reading ssh-pty-relay-absence-verdict.test.ts already pins for the spawn path; the reconnect path now agrees with it. Rejected: parsing the relay's mint epoch out of `pty2:<epoch>:<n>`. It needs the current epoch on the wire (a capability-negotiated relay change), it has no answer for legacy `pty-N` ids, and a relay that comes back with zero PTYs gives the client no epoch to compare against. Rejected: clearing the leaf record outright, because the remote workspace snapshot re-hydrates those ids after the clear and the gate would refuse again. * refactor(ssh): name the relay-disowned signal for disownership, not exit |
||
|
|
946627f2ce |
fix(runtime): route runtime filesystem commands by resolved execution host (#18325)
`ResolvedRuntimeFileTarget` carried `connectionId?: string` and no host id, so `undefined` spelled three different answers at once — "runtime: host", "unresolved" and "genuinely local". Its sole resolver read `store.getRepo(worktree.repoId)?.connectionId` and never looked at `worktree.hostId`, which outranks every repo row, so one arbitrarily chosen row decided the execution host for ~30 filesystem dispatches. This is #18307's defect in the same file family; it was deliberately left out of that PR rather than doubling an already-36-site diff. The target now carries `executionHostId: ExecutionHostId` (never null, never optional), resolved through `resolveWorktreeHostRouting` — the same adapter #18307 added — and dispatched through #18296's `resolveFilesystemRouteForHost`. Dispatch sites call `requireRuntimeFileProvider`, where `null` means exactly one thing: the host is `local` and the read happens here. Four answers that used to collapse into one: - `ssh:x` with a rival row on `ssh:y` — routes to x. Previously the first row won. - `local` with a surviving `connectionId` — a row contradicting itself; no SSH connection is handed out. - `runtime:<env>` — throws `ExecutionHostNotDispatchableError`. Its repo row's connection names a target in the *server's* namespace; reading it here reaches a same-named target on this client. - rival rows disagreeing with no worktree host — `worktree_execution_host_unresolved`, matching the launch and Git paths rather than guessing a row. Two further reads stop degrading. `assertRuntimeFileMutationExpectation` recomputed the host from `connectionId`, so a client's host expectation could pass against a host the workspace never named; it now compares the resolved host. And the cross-workspace terminal tap coalesced `knownWorkspaceTarget?.connectionId ?? connectionId`, so a sibling workspace resolved as `local` inherited the origin worktree's SSH target and statted a local path on the remote box; a non-optional host id replaces rather than coalesces. An unreachable SSH host still throws `SSH_FILESYSTEM_PROVIDER_UNAVAILABLE_MESSAGE`; loss of contact is never evidence of locality (docs/reference/ssh-execution-boundary.md). Quick-open listing and path search keep degrading to empty for an unreachable host — that is a false negative, not a local answer — and now do so only for a host that really is remote. The whole `runtime-file-commands-*` family carries `@ts-nocheck` from a mechanical class split, so removing the field could not raise the compile errors that made #18307 safe. `runtime-file-command-target.ts` is deliberately checked, and a ratchet test stands in for the errors the family cannot produce. No wire change: `ResolvedRuntimeFileTarget` is main-process internal, and the SSH watcher-release and grant keys are byte-identical to before. |
||
|
|
21210aad34 |
fix(native-chat): make structured Codex launches race-resistant (#18251)
* fix(native-chat): cancel close-racing structured launches * fix(native-chat): make structured launches observable and recoverable * fix(native-chat): reconcile merged session tab publications * refactor(native-chat): unify host snapshot versioning * refactor(native-chat): complete launches from host snapshots * fix(native-chat): replay unknown launches by intent * fix(native-chat): guard duplicate launches and bound sync recovery * test(native-chat): type owner fixture * test(native-chat): type owner fixture * fix(native-chat): back off structured session resubscribe * fix(native-chat): fence delayed local session snapshots * fix(native-chat): retry initial session sync safely * fix(native-chat): refresh before sync retry * test(native-chat): cover folder sync cursor cleanup * fix(native-chat): retry failed structured session subscriptions --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
d5750648c2 |
fix(runtime): route runtime Git by resolved execution host, not repo connectionId (#18307)
`RuntimeGitTarget` carried `connectionId?: string` and no host id, so `undefined` spelled three different answers at once — "runtime: host", "unresolved", and "genuinely local". Its sole resolver read `store.getRepo(worktree.repoId)?.connectionId` and never looked at `worktree.hostId`, which outranks every repo row, so one arbitrarily chosen row decided the execution host for 36 downstream dispatches. The target now carries `executionHostId: ExecutionHostId` (never null, never optional), resolved through the shared rule that landed with #17909/#17919 and dispatched through the host-keyed routes from #18296. Dispatch sites call `requireRuntimeGitProvider`, where `null` means exactly one thing: the host is `local` and the command runs here as free functions. Four answers that used to collapse into one: - `ssh:x` with a rival row on `ssh:y` — routes to x. Previously the first row won, which is the reproduced cross-host leak. - `local` with a surviving `connectionId` — a row contradicting itself; no SSH connection is handed out. - `runtime:<env>` — throws `ExecutionHostNotDispatchableError`. Its repo row's connection names a target in the *server's* namespace; dialling it here reaches a same-named target on this client. - rival rows disagreeing with no worktree host — `worktree_execution_host_unresolved`, matching the launch path rather than guessing a row. An unreachable SSH host still throws `SSH_GIT_PROVIDER_UNAVAILABLE_MESSAGE`; loss of contact is never evidence of locality (docs/reference/ssh-execution-boundary.md). `resolveWorktreeLaunchHost` keeps its exact signature and now delegates to `resolveWorktreeHostRouting`, the same resolution answering "which host is this on" rather than "what may this client dial" — the git target needs the first question because `local` and `runtime:` are two different non-SSH answers. No wire change: `RuntimeGitTarget` is main-process internal, and the SSH and local model-discovery host keys are byte-identical to before. `RuntimeFileTarget` has the same defect in ~30 filesystem dispatches and is deliberately left for a follow-up. |
||
|
|
9cda5a9dc0 |
fix(worktrees): stop a resolved-worktree snapshot answering for repos it never saw (#18295)
* fix(worktrees): stop a resolved-worktree snapshot answering for repos it never saw `listResolvedWorktrees` caches one fleet-wide snapshot for RESOLVED_WORKTREE_CACHE_TTL_MS (1s) and reuses it on time alone. Nothing invalidates it when a repo is registered, so for up to a second after a repo row lands, every caller reads a snapshot computed before that repo existed -- and reads the gap as a verdict. The visible failure is the SSH skill install. `resolveSkillSshTarget` resolves a workspace-scope destination through that snapshot, so installing into a worktree on a host connected moments earlier threw `skill-install-workspace-not-found`: the client asserting a remote workspace is absent on the strength of client-side bookkeeping that had never looked at the host. That is the shape `docs/reference/ssh-execution-boundary.md` rules out -- absence from a client-side set is not evidence about the execution host. It made `tests/e2e/ssh-skill-installation.spec.ts:108` fail 3 runs in 4 locally and deterministically in the Docker SSH lane, where connect-then-install lands inside the one-second window every time. The snapshot now carries the repo-registration revision it was computed under and is only reused while that revision still holds. The counter is the one `bumpLocalWorktreeScanGeneration` already advances on every repo add, removal and update, so the check is O(1) and cannot drift from the mutation sites. * fix(worktrees): key the snapshot on repo mutations only, not on generation reads Two things the headless-reattach lane surfaced. The revision I keyed the snapshot on was `generationSequence`, which `getLocalWorktreeScanGeneration` also advances when it mints a key for a repo id nothing has scanned yet. That is a read, not a mutation, so a read path could discard a snapshot that was still perfectly valid -- the mirror image of the staleness this fixes, and a way to make a lookup fail that would otherwise have succeeded. The counter now advances only where the scan generation is actually bumped: repo add, removal, update, and scan-cache invalidation. Separately, `pty-restore-record-seeding.test.ts` primed the cache by writing its private `resolved` field with a literal spelling out `worktrees`, `platformByRepoId` and `expiresAt`. That literal is a second copy of the cache's freshness contract, so adding a field to the real entry left the fake one failing the check: the primed snapshot was rejected, resolution fell through to a real scan, and the headless fixture -- which has no git -- got `selector_not_found`. It now primes through `getSnapshot` so the cache stamps its own entry and the two cannot drift again. The revision never moved during that test (0 before and after), so nothing was being invalidated; the fake entry simply never satisfied the contract. |
||
|
|
d084a2a36a |
fix(ssh): decide remote-vs-local from the resolved execution host, not a raw field (#18294)
`repoIsRemote` read `repo.connectionId` directly. That is one of four spellings of host ownership, so the predicate was wrong in both directions: a row carrying only `executionHostId: 'ssh:<target>'` read as local and got the Linux-only `orca-ide` rename it cannot resolve through the relay shim, while a row that declares itself `local` with a stale `connectionId` read as remote and lost the rename it needs on a Linux desktop. The predicate now resolves the host first and asks "does an SSH target hold this row's files" via `getRepoSshConnectionId`. That keeps a `runtime:` host's nested SSH target remote (that machine reaches the files through its own relay shim) while a runtime with no nested target - a full Orca install - stays local, as do WSL and local. Its call sites did not all want that question: - The four launch-scope sites in main already hold the resolved PTY route on `TerminalWorkspaceLaunchScope.connectionId`. `scope.repo` is documented display metadata and can be a row from a different host than the worktree names, so they now read the route they will actually spawn on. A launch shape that disagrees with its own route is the bug, not a second predicate. - `launchAgentInNewTab` picked its repo row with a host-blind `store.repos.find`, so a worktree that names its own host could be shaped by another host's row. It now resolves through `getConnectionIdFromState`, the same rule the file already used for transcript readability. - `resolveAgentBackgroundLaunchHost` derived the route, the trust write and the launch shape from three reads of the raw field; one resolution now feeds all three. Also converts the raw `repo.connectionId` agent-detection probe eight lines above `buildWorktreeStartupForDraft`'s launch shape, which #17919 deferred precisely because converting it alone would have left that file internally inconsistent. Tests cover two distinct SSH hosts (a single-host fixture passes even when the answer comes off the wrong row, which is how the `ssh:m4air` -> openclaw leak survived review) and a `runtime:` host carrying a nested SSH target. |
||
|
|
7c94d12190 |
fix(ssh): route four host-blind seams through the resolved execution host (#17919)
* fix(host-routing): resolve the execution host before reading a connection Three issues in one defect class: a resolver reads one spelling of one arbitrarily chosen row instead of resolving the worktree's execution host, so something local answers a question about a remote. returned that row's connectionId. With duplicate repo rows for one repo id it could pair a runtime owner with a client-owned SSH connection. It now resolves through the same ambiguity-aware index getRuntimeEnvironmentIdForWorktree uses, prefers the repo row for the host the worktree names, and derives the connection from the resolved host. Conflicting rows return `undefined` (this module's documented "cannot determine the host"), never `null`. `store.getRepo(worktree.repoId)?.connectionId ?? null`. `getRepo` is host-blind and the same repo id can exist on local, SSH and runtime hosts, so a remote worktree could spawn its PTY on the client with the remote cwd. resolveWorktreeLaunchHost picks the row for the worktree's host and reads the connection off that host; conflicting rows are unresolved, not local. session-partition owner maps that contradict each other. Both now compute through one shared function whose argument records the divergence. No behaviour change on either side: converging needs a read-both migration, since both partitions hold real data written by shipping builds. * fix(host-routing): keep nested SSH connections resolvable under a runtime host getRepoSshConnectionId read only the resolved execution host, so a repo row owned by a runtime that reaches a nested SSH target (connectionId: ssh-*, executionHostId: runtime:*) resolved to no connection — answering 'local' for a remote worktree, the same defect #17909 fixed in the other direction. * fix(host-routing): resolve both sides of the execution host through one rule The renderer resolver leaked between two different SSH hosts: a worktree on `ssh:m4air` whose only indexed repo row belonged to `openclaw` answered 'openclaw', because the host-scoped lookup missing fell through to an id-only one. Main's resolver, in the same change, answered 'm4air' — two resolvers, one right and one wrong, on identical input. Both sides now adapt one shared rule (`worktree-execution-host-resolution.ts`): the worktree's own host outranks every repo row, and a row on a different host is never evidence about this one. The renderer's WeakMap index becomes the memoizing adapter it always was; `resolveWorktreeLaunchHost` becomes main's mapping of unresolved onto its throw. Settles the rule the change previously answered two ways. `getRepoSshConnectionId` and `getSshTargetIdForExecutionHost` disagreed for a runtime host carrying a nested `connectionId`; they now compose, so the execution host is the single authority. On a `runtime:*` row that field is a paired HUB's private SSH target, spread through by `repoWithFetchedOwner` and unaddressable from this client — the project-first successor of the row nulls it for exactly that reason. That also fixes the `kind !== 'ssh'` fallback, which fired for `local`: a row declaring itself local handed out an SSH connection. * fix(ssh): resolve the execution host in the worktree scan and managed create The worktree scan and createManagedWorktree both picked remote-vs-local from repo.connectionId, so a row stamped only executionHostId: 'ssh:*' was scanned and created on the client against a remote path. The folder branch returns before the check, so its agent-trust write landed locally too. Refs #11163 * fix(ssh): stop over-rejecting and refusing SSH hosts the process owns runtimeRepoMatchesExecutionHost rejected an unstamped SSH repo against its own ssh:<connectionId>, so repo-add/clone dedupe could register a second row for a path the host already owns. assertHostIsSupported made the CLI/runtime RPC refuse --host ssh:* while the same process's IPC handler routed it correctly; setupExistingFolder now shares that registration. Clone still refuses, because nothing in this process clones onto an SSH host. Refs #11163 * test(ssh): retarget the SSH host-setup guard spec at the substitution it prevents setupProjectExistingFolder now registers the remote path through the same addRemoteRepoFromPath the desktop IPC uses, so it fails on the host's terms (connection not registered) rather than a categorical refusal. The local clone/probe side effects it exists to catch are still asserted absent. Refs #11163 * fix(cli): require an absolute path when setting a project up on an SSH host Routing --host ssh:* to the remote registration made relative paths newly reachable there, and they were resolved against the client cwd — registering a path that names the wrong machine. Refs #11163 * fix(repos): read the SSH registry directly so the runtime stays Node-bootable Routing runtime project setup through addRemoteRepoFromPath dragged ipc/ssh -- and its 25-module electron graph -- into the runtime bundle. ssh-target-registry already exists for exactly this; ipc/ssh only re-exports it. * fix(ssh): close the agent-launch and session-export host-blind twins Three sites left on the legacy spelling, all the same shape as the ones this branch already fixed: - `launchAgentTerminal` did `getRepo(worktree.repoId)` then wrote agent trust with that row's `connectionId`. Host-blind, so a repo id carried by two SSH hosts wrote a remote path into the *client's* Codex/Cursor/Copilot config and the agent on the host never saw the trust. Every sibling call site already passes the resolved `workspace.connectionId`; this was the last that did not. - `targetForWorktree` (workspace-session export) fell back to the same host-blind read, so a session could be published to a machine that never owned the worktree. Unresolvable ownership now exports to nobody. - `addRemoteRepoFromPath` minted `connectionId`-only rows while being the routing path this branch adds, so it kept creating rows in exactly the spelling the branch works around. It now stamps `toSshExecutionHostId(connectionId)` at creation; `reassignSshTargetId` already migrates both spellings, so target rename stays correct. Tests cover two *different* SSH hosts throughout — the case none of the earlier duplicate-row tests had, all of which were local-vs-ssh or runtime-vs-ssh. |
||
|
|
c61ca56a9b |
fix(ssh): resolve the worktree's execution host instead of guessing from one repo row (#17909)
* fix(host-routing): resolve the execution host before reading a connection Three issues in one defect class: a resolver reads one spelling of one arbitrarily chosen row instead of resolving the worktree's execution host, so something local answers a question about a remote. returned that row's connectionId. With duplicate repo rows for one repo id it could pair a runtime owner with a client-owned SSH connection. It now resolves through the same ambiguity-aware index getRuntimeEnvironmentIdForWorktree uses, prefers the repo row for the host the worktree names, and derives the connection from the resolved host. Conflicting rows return `undefined` (this module's documented "cannot determine the host"), never `null`. `store.getRepo(worktree.repoId)?.connectionId ?? null`. `getRepo` is host-blind and the same repo id can exist on local, SSH and runtime hosts, so a remote worktree could spawn its PTY on the client with the remote cwd. resolveWorktreeLaunchHost picks the row for the worktree's host and reads the connection off that host; conflicting rows are unresolved, not local. session-partition owner maps that contradict each other. Both now compute through one shared function whose argument records the divergence. No behaviour change on either side: converging needs a read-both migration, since both partitions hold real data written by shipping builds. * fix(host-routing): keep nested SSH connections resolvable under a runtime host getRepoSshConnectionId read only the resolved execution host, so a repo row owned by a runtime that reaches a nested SSH target (connectionId: ssh-*, executionHostId: runtime:*) resolved to no connection — answering 'local' for a remote worktree, the same defect #17909 fixed in the other direction. * fix(host-routing): resolve both sides of the execution host through one rule The renderer resolver leaked between two different SSH hosts: a worktree on `ssh:m4air` whose only indexed repo row belonged to `openclaw` answered 'openclaw', because the host-scoped lookup missing fell through to an id-only one. Main's resolver, in the same change, answered 'm4air' — two resolvers, one right and one wrong, on identical input. Both sides now adapt one shared rule (`worktree-execution-host-resolution.ts`): the worktree's own host outranks every repo row, and a row on a different host is never evidence about this one. The renderer's WeakMap index becomes the memoizing adapter it always was; `resolveWorktreeLaunchHost` becomes main's mapping of unresolved onto its throw. Settles the rule the change previously answered two ways. `getRepoSshConnectionId` and `getSshTargetIdForExecutionHost` disagreed for a runtime host carrying a nested `connectionId`; they now compose, so the execution host is the single authority. On a `runtime:*` row that field is a paired HUB's private SSH target, spread through by `repoWithFetchedOwner` and unaddressable from this client — the project-first successor of the row nulls it for exactly that reason. That also fixes the `kind !== 'ssh'` fallback, which fired for `local`: a row declaring itself local handed out an SSH connection. |
||
|
|
b8b7a6be9d |
fix(activity): persist the agents unread filter and grouping (#18255)
* fix(activity): persist the agents unread filter and grouping The Agents view's "Show unread threads only" toggle and Group-by select were plain component state in the sidebar and the Activity page, so both reset on every mount — including app restart — while their neighbours in the same toolbar (compact rows, show child agents) survived via the persisted UI store. Promote both to `agentsReadFilter` / `agentsGroupBy` persisted UI preferences, wired through the same seams as `agentsCompactMode`: shared type, default, strict client RPC schema, pairing-local field census, web read pin, store contract/actions, and hydration normalizers that reject unknown values. Both consumers now read the store, so the sidebar and the Activity page share one filter the way they already share compact mode. * refactor: centralize thread filter value domains Establish filter and groupby value domains as the single source of truth, with types derived from them to prevent drift between valid values and their normalizers. Extract common validation logic into a shared isMember helper to keep the two normalization functions in sync. * refactor: centralize thread filter value domains Consolidate filter value definitions in agents-view-thread-filters and use them in Zod schema validation to ensure consistent, persistent serialization of filter state. |