Commit Graph
910 Commits
Author SHA1 Message Date
Brennan Benson 8e13485c9b fix(stats): count agent sessions from hook transitions, not OSC titles (STA-2445) (#14657)
* test(stats): dual-record the OSC-title detector against canonical hook transitions

Adds an AgentSessionTransitionRecorder that derives agent-session start/stop
boundaries from agent-hook status transitions, and a side-by-side comparison
that feeds both pipelines into a real StatsCollector.

Nothing is rewired yet — this commit only measures the delta:

                                 title detector   canonical
  hook-only agent                       0             1
  braille-spinner non-agent TUI         1             0
  one agent, one reconnect              2             1
  totals                                3             2

Refs #10201, STA-2445.

* fix(stats): count agent sessions from hook transitions and delete the title detector

Switches StatsCollector off AgentDetector and onto the canonical agent-hook
status stream, then removes the detector and its raw-PTY invocation.

- main/index.ts subscribes the recorder to subscribeEnrichedStatus and
  subscribePaneStatusClear, next to where StatsCollector is constructed.
- orca-runtime.ts no longer feeds raw PTY bytes to a stats detector.
- StatsCollector keys sessions on a stable pane key, not a per-spawn ptyId.

Fixes #10201, refs STA-2445.
2026-08-18 00:35:50 -07:00
Brennan Benson eb0ec39242 fix(runtime): stop one unreachable relay from freezing every workspace as active (STA-517) (#14649)
* fix(runtime): stop one unreachable relay from freezing every workspace as active

The worktree.ps liveness refresh is the only thing that retires an exited PTY,
and mobile renders "active" straight off the summary it produces. Its aggregate
inventory ran every provider through Promise.all with no per-provider deadline,
so a single SSH relay that rejected — or simply did not answer inside the 3s
budget, since a relay list runs to the mux's own 30s default — cost the runtime
the whole inventory. Nothing was ever proven dead, so every retained pane kept
reporting hasHostSidebarActivity/liveTerminalCount, and the SSH workspaces stayed
"active" on mobile for as long as the connection stayed unreachable.

Settle each SSH provider independently and forward the caller's deadline, so
local and healthy relays are still reconciled. A provider that does not answer is
unknown, not empty: the runtime's existing hasPty rescue keeps its panes. A local
failure still fails the aggregate, matching pty:listSessions.

The restored-orchestration-authority sweep now runs after that rescue, so a pane
the controller still vouches for keeps its handle instead of losing it to a
listing that merely omitted it.

STA-517

* test(runtime): assert the provider scope, not the exact arity, of inventory calls

These assertions exist to prove which provider scope the inventory asked for.
Forwarding the caller's deadline added a second argument, which broke them on
arity alone. Match the scope argument and require a numeric deadline beside it,
so the intent is preserved and the budget is covered too.

* test(runtime): type the inventory mock's scope parameter

A bare `async () =>` mock types mock.calls as an empty tuple, so reading the
scope argument off it fails typecheck. Declare the parameter the runtime
actually passes.
2026-08-18 00:35:16 -07:00
Neil e39def3825 fix(repo-identity): bound git remote-identity probes and retire them with their repo (#15196)
* fix(repo-identity): bound git remote-identity probes and retire them with their repo

The local `git remote -v` probe ran with no timeout and no signal, and the
runner only arms its kill timer when a timeout is passed, so a hung NFS/SMB
cwd or a wedged `wsl.exe -d <distro>` left the promise unsettled and the child
alive. Because the sweep is sequential and dedupes per location, that one
wedged location stalled enrichment for every other repo.

- probeGitRemoteIdentity/detectGitRemoteIdentity take `{ signal, timeoutMs }`;
  local reads get the 5s background local-git-read budget, SSH gets a budget
  under the relay's 30s request timeout. Timeouts/aborts still map to
  `unavailable`, never `no-remote`, so they cannot clear a resolved identity.
- In-flight probes are tracked with an AbortController and retired (aborted +
  dropped) when their location is no longer backed by a repo, so a re-added
  repo is not poisoned by the dead entry and a retired probe cannot re-seed a
  retry deadline.
- Added the missing sweep-level guard so repos:list / projects:list /
  projectHostSetups:list coalesce into one pass instead of stacking one
  sequential sweep per list IPC.

STA-4452

* refactor(repo-identity): bound the enrichment listener set to stable caller references

Every call site allocated a fresh onChanged closure, so the Set that notifies
coalesced sweeps deduped nothing: during a chain that never quiesces it grew one
entry per list IPC and multiplied the repos:changed broadcast. Hoist the closures
to stable references in ipc/repos.ts and OrcaRuntimeService, and state the
contract on the set.

Also: guard the synchronous retirement call so the fire-and-forget entry point
keeps its no-throw contract, drop the placeholder promise in favour of building
the in-flight entry in one shot, and make the coalescing test use a shared list
reference plus a distinct runtime reference so it detects both stacked passes and
a dropped caller.
2026-08-17 20:43:06 -07:00
Neil 0e96b82e44 fix(mobile): keep phone tab selection across host snapshots
* fix(mobile): keep phone tab selection across host snapshots

Preserve device-owned tab focus across ordinary host republications while explicit follow navigation remains authoritative. Retire closed selections across clients so stale snapshots cannot resurrect tabs.

* fix(mobile): acknowledge session tab closes

* fix(mobile): avoid tombstones for uncommitted closes

* fix(web): implement session close IPC stubs

* refactor: simplify mobile tab close flow

* fix: bound session tab close confirmation
2026-08-17 18:26:52 -07:00
Brennan Benson 8cd338357e fix(runtime): tear down terminal subscriptions only through their owning registration (#14992)
* fix(runtime): tear down terminal subscriptions only through their owning registration

A terminal.subscribe teardown was keyed on `${terminal}:${clientId}`, which is
stable across reconnects. cleanupSubscription invokes whichever cleanup currently
owns that key, so after a mobile reconnect rebound the id, a late teardown from the
dead connection killed the replacement stream and the terminal froze (STA-4510).

Add registerOwnedSubscriptionCleanup, returning a registration handle whose
releaseIfCurrent no-ops once the id has been rebound, and route all 12 teardown
call sites in the three terminal.subscribe branches through it. Register-time
eviction now also targets the owner it captured rather than re-resolving the key.

terminal.unsubscribe gains the same ownership rule via connectionId, matching the
runtime.clientEvents.unsubscribe precedent: make-before-break migration sends the
unsubscribe over the old session after the new one has already rebound the id.

The teardown-by-key pattern predates the bug; #7490 made it reachable by binding
the exit-waiter to the per-socket abort signal, so every socket close now runs it.

Tests use a faithful subscription-registry double; the ad-hoc Map stubs they
replace never evicted the prior generation, which is why no test caught this.

* test(runtime): migrate the remaining subscription stubs to the faithful registry

streaming.test.ts and terminal-provider-snapshot-sequence.test.ts still stubbed
registerSubscriptionCleanup only, so terminal.subscribe's owned registration was
undefined and teardown never fired. Assert the registration is retired rather than
spying on cleanupSubscription, which the owned path now calls internally.

* fix(runtime): guard the lease-only presence release and use the shared registry double

Review findings on this PR:

- The lease-only branch's compensating handleMobileUnsubscribe ran unconditionally
  after a rebind. Presence is keyed (ptyId, clientId) with no refcount and `closed`
  is exactly the post-rebind state, so a superseded handler deleted the replacement
  subscriber's presence. Gate it on the registration still being current. This also
  gives SubscriptionRegistration.isCurrent its production caller.

- The STA-4510 regression test hand-rolled a second registry copy that diverged from
  the shared double (no try/catch around cleanup, no in-flight join). Import the
  shared double instead, so the test proving the bug uses the same fidelity as the
  rest of the suite.

- cleanupSubscriptionIfOwnedByConnection treats an absent connectionId as authority.
  That is the connection-less local unix-socket path, not an oversight; say so.

* fix(runtime): drop the lease-only presence guard; it disabled a real compensation

Review pass 2 showed the guard added in the previous commit was wrong.

registration.isCurrent() is always false whenever `closed` is true: either our own
cleanup ran, in which case cleanupSubscriptionAndWait already deleted the map entry,
or a rebind replaced it. So the guard did not distinguish the two cases — it made the
compensating handleMobileUnsubscribe unreachable, which is precisely the
resurrect-after-cleanup case that line exists to handle.

The scenario the guard was meant to fix is also not reachable: the lease-only call
passes no viewport, and both !viewport paths in handleMobileSubscribeInternal return
with no await, so the subscribe resolves in microtasks and a socket close cannot win it.

Revert the guard, and drop SubscriptionRegistration.isCurrent with it — it had no
remaining production caller and shipping unused runtime API invites exactly this.

Also from review:
- cleanupSubscriptionIfOwnedByConnection reported false for an id with no registration,
  conflating 'refused, another connection owns it' with 'already gone'. Report true.
- Note on the registry double that it mirrors production and can drift; the runtime
  tests pin real behavior, the doubles only pin routing.

* fix(runtime): make the unsubscribe refusal observable and restore test-double parity

Review pass 3:

- The registry test double omitted production's 'unregistered id is already gone'
  early-out, so it returned refused where production returns gone. The legacy
  terminal.unsubscribe path reaches that branch with a never-registered bare id, so a
  routing test would have locked in the inverse of production.
- terminal.unsubscribe ORed the bare-id and composite results. Registrations always use
  the composite, so the bare id reported 'already gone' and masked a real ownership
  refusal: a stale connection was correctly refused but told unsubscribed: true. The
  composite answer is authoritative when we try it.
- Pin why the lease-only compensating handleMobileUnsubscribe is deliberately unguarded:
  it is safe only because that call passes no viewport and returns with no await.
- Cover the no-registration branch, which had no test.

* fix(runtime): report an unsubscribe refusal without masking a real teardown

Review pass 4 disproved the previous commit's rationale. A clientless legacy-JSON
stream registers under the bare terminal id, not the composite, so either call in
terminal.unsubscribe can be the real teardown. Overwriting reported false after a
destructive success; the earlier OR reported true after a refusal. AND over the calls
that actually ran is the honest aggregate: false needs a genuine refusal.

Tests:
- cover the bare-id path, which the ownership tests never exercised
- pin the microtask invariant the unguarded lease-only compensation depends on. The
  first version of that test was vacuous: it raced two setTimeout(0) timers, and the
  earlier-registered one always won, so it passed with an await injected. It now drains
  microtasks only and goes red under that mutation.
2026-08-17 16:13:45 -07:00
Brennan BensonandQA 64de8dd637 fix(workspaces): delete on the confirmed host, and make both hosts' rows selectable (STA-4343) (#15013)
* fix(workspaces): host-qualified workspace deletion (STA-4343, STA-4448)

Squashed integration of PR #14606 + the codex review-loop output, replayed
onto current main. Granular history preserved on brennanb2025/sta-4343-review-full.

Fixes the regression from #13413: a workspace id is repoId::path with no host
component, so the same repo at the same path on two hosts published one id for
two workspaces, and deletion routed by that id landed on whichever host routing
preferred - usually the ACTIVE one, not the row the user confirmed.

- removeWorktree takes a REQUIRED host-qualified WorktreeRemovalTarget; omitting
  the host is a type error. All destructive callers migrated.
- Projections dedup on (host, id), so two hosts render as two selectable rows
  while the createWorktree/fetchWorktrees race duplicate still collapses.
- Ephemeral VM cleanup is host-scoped. It matched on bare workspaceId, so the
  host-scoped delete path destroyed the SURVIVING host's VM and its unpushed
  filesystem - a leak fix that had become data destruction.
- Selection, keyboard routing, lineage grouping and Space rows carry host
  identity end to end; fixing the executor dedupe alone would have turned
  one-row intent into deleting both hosts.

Files split to stay under max-lines rather than raising any cap.

* refactor: split files that crossed max-lines

The review-loop commits used --no-verify, so the pre-commit hook never
enforced the caps. Extracted cohesive units rather than raising any limit:
renderer teardown, delete-with-toast, pinned-group rows, host-scope helpers,
workspace-kind predicates, filter actions, kanban drag selection, the
renderer removal result type, and the native-chat persistence tests.

* refactor(workspaces): extract cleanup deletion-phase selector

Clears the last max-lines violation and the import-type side effect the
changed-code gate flagged.

* refactor(sidebar): track the delete-dialog extraction modules

* fix(workspaces): preserve host identity across remaining surfaces

* fix(sidebar): re-carry host through the rewritten palette result model

#15170 replaced PaletteSearchResult while this PR was open. Re-applied the
host qualification on top of the new model instead of taking either side:
results carry worktreeHostId again, and the board filter keys its matched
set on host identity rather than the bare id.

Known gap, documented in the board test rather than deleted: searchWorktrees
resolves evidence through a `documents` map keyed by BARE worktree id, so two
same-id host rows collapse before this code sees them. Closing that belongs
with the palette work.

* test(cmd-j): pin the palette collision gap instead of asserting the old model

The palette collision test asserted two host-qualified rows, which #15170's
rewrite made unreachable: item ids are bare again and worktreeMap is id-keyed.

Rewritten to assert what holds — activation always names a host — and to pin
the defect it exposes: two same-id rows render on ONE command value, so React
sees duplicate keys and a click on the first row activates the second row's
host. That reproduces on main, so it is pre-existing, not from this PR. Pinned
rather than deleted so fixing it must update this test.

---------

Co-authored-by: QA <qa@local>
2026-08-17 15:57:07 -07:00
Jinwoo Hong 0bedeea642 fix(orchestration): expose unsupervised dispatch lanes (#15105) 2026-08-17 13:53:26 -07:00
Brennan Benson 32ee3b0536 reland(browser): route every cookie-import write through CDP identities, and never clear what it will not write back (#15030)
* reland(browser): restore CDP-identity cookie-import writes (#14729)

Reverts the revert 3c8410d927 (#14942) to restore the reviewed bf6dc6fcba
tree. The defect that caused that revert is NOT fixed by this commit — it is
fixed in the commits that follow, so the delta a reviewer must scrutinise
stays small instead of hiding inside a 2000-line re-add.

Conflict resolutions, all keep-both:
- browser-cookie-import.ts: two hunks around the zero-import early return.
  #14683's undecryptable warning and decryptedCookies.length === 0 condition
  are kept alongside the reland's old-client gate and partitionSkippedCookies.
  The old-client gate stays above the early return, and so above any mutation.
- en.json: both string sets (#14683's undecryptable copy and partitionSkipped).
- crashpad-capture.test.ts: took HEAD wholesale; unrelated to this ticket.

getStoragePath, and its positive-write assertion moves from cookies.set to the
CDP identity store — adapted, not weakened, and now strictly stronger because
it also asserts cookies.set carries no imported user data. Every decryption
assertion is untouched.

* fix(browser): add registrableFamily, one definition of a cookie family

STA-4300, change 1 of the reland fixes. No behaviour change yet — this only
adds the helper the later commits derive every skip scope from.

Deriving "family" inline in several places is what let the removal scope and
the write set disagree in STA-4090 and STA-4170, so there is exactly one.

The IP check runs on normalizeCookieDomain's OUTPUT, never the raw string.
psl treats an IPv4 literal as a dotted DNS name — psl.parse('127.0.0.1').domain
is '0.1' — and Chromium accepts many spellings of one address. new URL() inside
normalizeCookieDomain canonicalises 127.1 / 2130706433 / 0x7f.1 / 010.0.0.1 and
a trailing dot to a dotted quad first, so isIP() then catches all of them.
Mutation-proved: moving the check ahead of normalisation reddens the five
alternate-spelling cases and nothing else.

Bracketed IPv6 needs its own branch because isIP('[::1]') is 0; without it the
value falls through to psl, which throws, which returns the host — the right
answer for the wrong reason.

Returns null for a bare public suffix: naming `com` as a family would preserve
an entire TLD from removal and silently turn an import into a no-op.

* fix(browser): plan every cookie's fate before any jar mutation

STA-4300, change 2. Adds planImportWrites — pure, no I/O — so the write set and
every removal scope can derive from one value instead of being computed twice.
Not yet wired into the import paths; that is the next commits.

TWO passes, and the second is not optional. Family-atomic skip is a property of
the whole input: with a readable mixed.example row BEFORE an unreadable
sub.mixed.example row, a per-row guard emits the readable one before anything
knows the family will be skipped. Pass 1 classifies and collects the skipped
families; pass 2 re-filters the provisional writes.

Mutation-proved with a faithful one-pass guard (consult skippedFamilies as you
go): the readable-before-unreadable case goes red and the unreadable-before-
readable case still passes. That asymmetry is why both orders are tested and
why only the first is the named detector.

Family closure is at the registrable boundary because that is the boundary the
removal scope actually expands to: importedDomainScopes turns an imported
mixed.example into descendant roots, so a readable apex cookie drags a skipped
subdomain's live session into the removal scope (STA-4300 §2b).

hasUnrepresentableSkip surfaces a skip whose family cannot be named. A family we
cannot name is one we cannot exclude from the removal plan, and clearing a
family we cannot protect is the P0 this ticket exists to stop — so the caller
refuses before mutating rather than proceeding and hoping.

* fix(browser): derive path A's removal scope from the write plan (STA-4300 §2b)

Change 3. importValidatedCookies now plans before it opens the jar, and the
replacement scope and the write set are the SAME array.

bf6dc6fcba filtered replacementDomains per exact cookie
(partition.status !== 'unreadable'). That is not sufficient:
replaceCookiesForImportedDomains expands each imported domain into its
descendant roots, so a readable cookie on mixed.example pulls sub.mixed.example
into the removal scope — and a skipped cookie living there has its session
removed with nothing written back. Same erasure as the P0, different path.

plan.writes is family-closed, and because path A passes the very same array to
replaceCookiesForImportedDomains and to writeImportedCookies, the two sets
cannot drift apart. That is stronger than keeping them in sync: there is only
one set.

Also here, both before any mutation:
- an unrepresentable skip (registrableFamily → null) refuses the import, because
  a family we cannot name is one we cannot exclude from the removal scope;
- the old-client gate now keys off plan.skips rather than re-deriving the
  condition.

Counters: partitionSkippedCookies is a BREAKDOWN of skippedCookies — the
unreadable rows plus their family-suppressed siblings — added in exactly once,
so totalCookies === importedCookies + skippedCookies keeps holding. Counting it
separately is how a summary silently stops adding up.

Full src/main/browser suite: 69 files, 729 tests, green.

* fix(browser): split path B into scan and emit, and never stage a skip (STA-4300)

Change 4 — this is the commit that fixes the shipped P0.

bf6dc6fcba did:
    decryptedCookies.push({ ..., partition })   // write set built HERE
    if (partition.status === 'unreadable') { ... continue }   // guard AFTER

so an unreadable row discovered late could not retract a sibling already
emitted, while removeTransplantableCookies cleared the ENTIRE jar. The mutation
set was the whole jar and the write set a strict subset of it, which is how a
mixed readable/unreadable source emptied a populated jar and repopulated only
part of it.

Now the row loop SCANS only. Nothing is emitted inside it — not decryptedCookies,
not domainSet, not a staging row, not the imported count. Between scan and emit,
planImportWrites closes the skip over whole registrable families, and emit walks
the plan. Each scanned candidate carries its raw source row because
buildChromiumCookieInsertParams needs it; a record holding only the derived
fields compiles and then silently cannot stage.

Also between scan and emit, both before any mutation:
- an unrepresentable skip refuses the import;
- any skipped family calls disableStaging. A staged image is a whole-database
  replacement on the next start, so it cannot express "preserve this family";
  main's existing memoryFailed arm then reports restart-fallback-unavailable.

Test contract change, deliberately stronger: "never stages a partitioned row
whose ancestor bit is unreadable" asserted that an image WAS registered and
merely omitted the row. That image would still have erased the preserved family
on replay. It now asserts no image is registered at all — plus a new paired case
proving a no-skip import still stages, so "disable on skip" cannot be
implemented as "disable always" unnoticed.

Honest note on detection: reverting the emit filter currently reddens only the
staging assertion, because every end-to-end fixture in this module still starts
with an EMPTY jar — where clear-then-write-all and clear-then-write-some are
indistinguishable. That is the blind spot that let this ship. The populated-jar
suite in the next commit is what actually detects the erasure.

Full src/main/browser: 69 files, 730 tests, green.

* fix(browser): never remove a family the import declined to write (STA-4300 I2)

Change 5, and the one that makes preservation observable.

removeTransplantableCookies takes preserveFamilies. removableCookieEntries
filters those families out, which keeps them out of the removal plan AND — since
the CDP snapshot is taken from that same list — out of the restore set, so a
preserved coordinate is never submitted to any mutation at all.

When anything is preserved, the bulk clearData path is not used. clearData
removes everything outside excludeOrigins, and its own comment concedes a
rejection may already have emptied part of the jar; a preserved family is
deliberately absent from the snapshot, so a partial delete followed by a
rejection would destroy it with no identity able to restore it. Handing a
dynamically derived preserve list to that primitive cannot be made safe, so a
skip-bearing clear runs the frozen per-coordinate plan — which already excludes
those families — as its primary path instead. Nothing is preserved on an
ordinary import, so that path keeps today's single clearData call unchanged.

New populated-jar suite. Every end-to-end fixture in this module starts with
cookies.get returning [], and against an empty jar "clear then write all" and
"clear then write some" are indistinguishable — which is why four independent
gates missed the P0. These fixtures start populated.

Mutation-proved, and this is the first detector in the change set that catches
the actual erasure rather than a proxy:
- drop the preserved-family filter          → 6 of 8 red
- use bulk clearData while preserving       → 6 of 8 red, including
  "leaves a preserved family untouched", which is the erasure itself

IPv4-literal and single-label host families are covered explicitly: psl reads
127.0.0.1 as the dotted DNS name '0.1', so a wrong family here would fail to
match the preserve set and erase the live loopback session.

Full src/main/browser: 70 files, 738 tests, green.

* test(browser): pin STA-4300 end to end in real Electron, both write paths

Ports the pre-implementation reproduction into the PR. Verified RED against
main's product code and GREEN with the fix, on both paths — I re-ran it both
ways rather than assuming.

Two assertions are INVERTED rather than deleted, and both are now stronger. The
repro proved the defect by showing cookies.set was called WITHOUT a partitionKey;
after the reland the import write path does not touch cookies.set for user data
at all, so seeing the imported cookie there is itself the regression. The
fixture's internal guard is inverted the same way.

Non-vacuity, asserted in-run rather than in a report: ok:true with a nonzero
importedCookies, the importer reaching its terminal step, the source really
carrying its partition fields, and a CONTROL cookie written via CDP *with* a
partitionKey and read back with it — so "partitionKey absent" cannot be confused
with "the oracle cannot see partitions". electronVersion proves Electron ran.

Fixture correction worth recording, because the original reproduction was
invalid on this path and I did not catch it on first read: readJsonCookiePartition
reads a NESTED partitionKey object, but the fixture wrote topLevelSite and
hasCrossSiteAncestor at the entry's top level, where the importer never looks.
The file/paste case therefore failed on main for a fixture-shape reason rather
than for the defect — and its own non-vacuity check validated what the fixture
wrote instead of what the parser reads, which is exactly how a check that looks
rigorous proves nothing. Only the native half was ever a real reproduction.

Partition keys use a schemeful SITE, not an origin: Chromium canonicalises
https://app.example.com to https://example.com, measured via CDP.

* fix(browser): fail closed on lossy cookie identity recovery

* fix(browser): preserve unreadable families before decrypt

* fix(browser): reject malformed cookie partition sites

* fix(browser): validate recovered cookie partition sites

* fix(browser): reject empty JSON partition identities

* fix(browser): fail closed and serialize cookie imports

* fix(browser): reject malformed CDP opaque flags

* fix(browser): roll back outcome-unknown cookie writes

* revert(browser): move import serialization out to STA-4601

Reverts only the concurrency-serialization half of f7c27b71ab so #15030 stays
scoped to the STA-4300 reland. That commit mixed two concerns; this removes one
of them and keeps the other.

REMOVED (moves to STA-4601, a separate pre-existing P1):
- acquireCookieMutationLock / withCookieMutationLock and the mutationLocks map;
  clear.ts goes back to withCookieClearLock
- the import-wide lock in path A (mutationOwner on the target, acquire/release
  around replace + writes + rollback)
- the widened lock around path B's clear + memory writes
- browser-cookie-import-concurrency.test.ts

KEPT (genuinely STA-4300 scope, not concurrency):
- partitionKeyOpaque handling in readJsonCookiePartition and the CDP snapshot.
  An opaque partition key cannot be represented faithfully, so it reads as
  unreadable and its family is preserved — exactly the rule this ticket adds.
  194e57a628's refinement of that shape stays too.

The concurrent-import interleaving is real and reachable (nothing serialises
imports per partition, and the clear lock is released before the memory writes),
but it predates this change and is not caused or worsened by it. It gets its own
PR and its own review rather than riding a P0 reland whose history is that every
additional fix introduced a new defect.

src/main/browser: 71 files, 772 tests, green. Electron oracle still passes on
both write paths and still goes red against main's product code.
2026-08-17 12:44:39 -07:00
Brennan Benson 60805f5c45 fix(agent-status): preserve restored child provenance (#15082)
* fix(agent-status): preserve restored child provenance

* fix(agent-status): preserve restored completion context

* fix(agent-status): retain child boundary across OSC
2026-08-17 11:10:38 -07:00
Jinwoo Hong 646e9b692b fix(mobile): render pairing QR at scanner-safe scale (#15058) 2026-08-17 10:29:18 -07:00
Jinwoo Hong be07b43a2b fix(orchestration): enforce honest recipient routing (#14964) 2026-08-17 01:50:17 -07:00
Jinwoo Hong 7c798907c5 fix(skills): harden cross-host bundle installs (#15000) 2026-08-17 00:55:04 -07:00
Jinwoo Hong 30b76e6edb fix(orchestration): enforce task dispatch state invariant (#14961) 2026-08-17 00:52:10 -07:00
Brennan Benson 7afce2ea41 fix(ssh): stop reporting a confirmed kill when the SSH provider is gone (#14977)
* fix(ssh): stop reporting a confirmed kill when the SSH provider is gone

A detached relay PTY is designed to outlive the provider that addressed it
(it ignores SIGHUP and ships with an unlimited grace), so "the SSH provider
is no longer registered" is lost contact, never evidence the remote process
stopped. Both stop primitives in the PTY controller returned `true` from
that branch, and every caller downstream reported the fabricated success:
the CLI printed "PTY killed.", worker-stop settled the dispatch as stopped,
and — because the stop "succeeded" — the unstopped-PTY gate never ran, so
worktree removal walked straight past a live remote agent.

`kill`/`stopAndWait` now still tombstone the local lease but report an
unconfirmed stop and record why, using the three-verdict vocabulary the
worktree teardown gate already spoke (`live` / `unverifiable` / `exited`),
promoted out of that module into `src/shared/pty-liveness-verdict.ts`.
The close receipt, the CLI wording, worker-stop and the removal gate all
read that verdict instead of inferring an exit from silence.

The same rule fixes the mirror-image defect: the aggregate inventory only
enumerates registered providers, so a dropped relay clears `connected` for
every remote PTY at once. The sweep now separates the provider answering
"absent" (an exit) from no provider being able to answer (lost contact), so
worker-stop stops claiming `exited` from a disconnect.

The `connected` wire field is unchanged in meaning and shape.

* fix(orchestration): apply the same honesty to the federation stop path

The federation host runs its own copy of the worker observation and stop
logic, with the same two defects: `inspectRemoteAttachment` read a dropped
relay's `connected: false` as `exited`, and `federationStop` settled the
dispatch as stopped from a close it never confirmed — relaying a fabricated
success all the way home to the coordinator.

Two guards also had to move so the honest verdict does not become a new
refusal. `federationRead` gated on `status !== 'running'`, which would have
rejected a connected terminal the moment a stop lost contact with it; it now
gates on `status === 'exited'`, which is equivalent for every pre-existing
status given the two guards beside it. Local `workerStop` likewise still
attempts the close when the verdict is `unverifiable` — losing contact is a
reason to report the outcome honestly, never a reason to stop trying.

The show observations now carry the reason alongside the status, so a bare
`unverifiable` is actionable. Both are new optional fields.

* fix(ssh): preserve unconfirmed stop verdicts across consumers

* fix(ssh): use canonical live verdict wording

* fix(ssh): refuse wrong-host teardown verification

* test(orchestration): confirm worker release teardown

* fix(orchestration): negotiate honest worker stop receipts

* fix(agent-teams): fence uncertain teammate respawns

* fix(ssh): avoid duplicate missing-provider teardown

* fix(orchestration): preserve archives across release retries

* fix(ssh): preserve verdicts across synthetic kill exits

* fix(ssh): preserve liveness evidence across teardown

* fix(agent-teams): replace panes only after confirmed stop

* fix(ssh): distinguish host exits from relay loss

* fix(ssh): narrow concurrent inventory verdicts

* fix(orchestration): serve archives after uncertain release

* fix(orchestration): expose unverifiable read liveness

* test(ssh): align liveness assertions with verdicts

* fix(ssh): preserve host scope across inventory failures
2026-08-17 00:11:19 -07:00
Neil 3bb87ff93b reland(shell): one portable Unix startup dialect, with both revert causes fixed (#15018)
* reland: portable startup-shell dialect, with the two revert causes fixed

Relands #14863 (reverted by #14975) with fixes for both regressions the
revert cited.

1. History GC deleted folder-workspace shell history. The live set was built
   from `getAllWorktreeMeta()` alone, but a folder workspace's PTY carries
   `folder:<id>` as its worktree id, so every live folder workspace looked
   orphaned. `getKnownWorktreeIdsForHistoryGc` now unions in
   `getFolderWorkspaces()`. Both consumers — the history-directory prune and
   the fish-history sweep — read that one set, so the fix covers bash, zsh and
   fish history alike. The directory prune had this gap since #1524; #14863
   only widened its blast radius to fish files.

2. A copied Codex resume command aborted under `set -u`. Its leading clear
   statement has to test `$fish_pid`, and that unbound expansion takes the
   whole line — including the agent launch — down with it. Copied text runs in
   a shell Orca never spawned, so nothing can seed that variable first. The
   removal now rides on the agent itself as `env -u`, which needs no shell
   syntax and no expansion. Verified byte-identical under `set -u` in sh,
   bash, zsh, dash, ksh and fish.

   `env` cannot run the `cd` builtin, and a child `cd` would not move the
   agent, so the prefix is placed on the agent rather than on the whole
   `cd … && agent` chain. cmd and PowerShell have no nounset hazard and keep
   their clear ahead of the `cd`, which preserves `cd … && agent` — a failed
   `cd` still cannot launch the agent in the wrong directory.

* fix(history-gc): stop three more paths from deleting live shell history

Found by adversarial review of the reland. All three are the same class as
the bug that caused the revert: a live set that is missing a category of
real workspace, so the GC reads it as orphaned.

1. Profiles. The history root is `userData/terminal-history`, which has no
   profile segment, but the Store the GC consults is per-profile. So after a
   profile switch the live set condemned every other profile's history — and
   fish history, which lands in the user's own fish data dir, is shared by
   every profile on the machine. The live set now unions in the inactive
   profiles' worktrees and folder workspaces, read from their data files. A
   profile whose ids cannot be read reports the empty set rather than one
   that condemns real history.

2. No empty-set guard on the tree scan. `sweepOrphanedFishHistoryFiles`
   refuses an empty live set because it cannot be told apart from a store
   that failed to hydrate; the directory scan, which deletes more, had no
   such guard. A store that fell back to default state would have taken
   every worktree's bash and zsh history with it, across all roots including
   WSL. Four existing tests passed `new Set()` and relied on "empty means
   everything is orphaned" — exactly the behavior being removed — so they
   now pass a real live set.

3. Relay fish history. The relay isolates its history tree under its own
   root but wrote fish history into the shared fish data dir under the
   desktop naming, keyed by the CLIENT's worktree ids. On a machine running
   both Orca and a relay host, the desktop sweep deleted remote sessions'
   history once it went stale. Relay files are now `orca_relay_<hash>`,
   which the sweep's pattern deliberately does not match; the relay still
   deletes them by exact name when the worktree goes away.

* fix(resume): enforce the env-removal invariants instead of documenting them

Both found by adversarial review; both were unreachable from today's callers
and silent if reached, which is exactly how they would survive to a caller
that does reach them.

- A pinned CODEX_HOME and the removal named the same variable, and `env -u`
  strips what the assignment just set — so the agent would have resumed
  against the real home and not found the session. The removal list now
  excludes any name the prefix pins, keeping the assignment authoritative as
  the old `clear…; CODEX_HOME=x agent` ordering did. Same fix in the git-bash
  twin. The PowerShell branch already clears before it assigns, so it was
  never affected.

- Placement was keyed on the platform while the grammar it selects is keyed
  on the shell, so `platform: 'linux'` with `shell: 'powershell'` emitted
  POSIX `env -u` into a PowerShell line. PowerShell now routes to the
  PowerShell builder whatever the host, and the POSIX/cmd split below asks
  the shell rather than the platform.
2026-08-16 23:51:53 -07:00
Jinwoo Hong 77ef6bb9ee fix(terminal): verify agent prompt submission (#14962) 2026-08-16 23:34:18 -07:00
Jinjing 81d7f9b24e refactor: split db.ts under 400 lines (#14979)
* refactor: split db.ts under 400 lines

* rm plan

* fix(orchestration-db): add safety guards to database operations

Add status guards to UPDATE statements to prevent late operations from
overwriting changes made by concurrent requests. Validate mutation results
to surface silent no-ops. Extract circuit-break threshold, add transaction
wrapping, and sanitize untrusted input. Bump schema version to v28.

* Add transactional safety to dispatch and message operations

Wrap dispatch failures and batched message updates with SAVEPOINTs to
ensure atomicity and idempotency:
- Dispatch failures now check status guards and roll back if the
  related task update fails, preventing partial state corruption
- Message batches (across multiple 500-id chunks) roll back entirely
  if any batch fails, avoiding partial mutations
- Add tests verifying idempotency and atomicity under failure conditions

* Handle concurrent writes and improve transactional safety

- Remote question answering: add classification check before and after UPDATE to safely detect concurrent modifications. Prevents false success when the UPDATE loses a race.
- Transaction rollback: wrap in try-catch to prevent errors from masking the original failure.
- Question thread reset: use status update instead of deletion to preserve message references.

* test: add answer replay and race condition edge case coverage

Add test cases for answer replay idempotency, conflict detection, and a race condition between concurrent answer updates in federation relay. Also verify local question state transitions during orchestration reset.
2026-08-16 22:44:43 -07:00
Brennan Benson 8ca4ed945e feat(terminal): report execution host and listing scope in terminal list (#14973)
* feat(terminal): report execution host and listing scope in terminal list

`orca terminal list` returned rows with no host identity and no statement
of what the listing covered, so a scoped listing that saw nothing read as
"nothing exists anywhere" — an agent reported a live remote worker dead.

Each row now carries an optional `executionHostId` derived from the PTY id
(SSH and paired-runtime ids embed their owner), and the result carries an
optional `hostScope` naming the hosts covered and the known hosts skipped.
Both are surfaced in `--json` and in the human-readable CLI output, where
an absent field renders as `unknown` rather than `local`.

Both row builders route through one resolver, so the rule lives in one place.

* fix(terminal): preserve unverifiable host scope

* fix(terminal): fail closed on unverifiable hosts

* test(terminal): name unverifiable scope explicitly

* perf(terminal): keep graph hydration host scans narrow

* fix(terminal): reject blank foreign host owners

* fix(terminal): validate inferred inventory hosts

* fix(terminal): preserve paired folder host scope

* fix(terminal): keep inventory host inference typed

* fix(terminal): disclose paired folder hosts
2026-08-16 22:13:03 -07:00
Brennan Benson 5e9e38fa75 fix(agent-status): announce Claude turn complete while background work runs (#14580)
* fix(agent-status): announce Claude turn complete while background work runs

Lead Stop/StopFailure already ends the turn, but resolveClaudePaneState keeps
the pane working for subagents, background shells, and session crons. That
erases the working→done edge that mints the completion banner: subagent turns
notify late with a stale body, and shells/crons never notify at all.

Stamp turnCompletedAt on the gated lead Stop, announce immediately from that
row, and pair the later all-clear done to the same end time so it cannot
double-fire or collapse consecutive turns onto the pinned stateStartedAt.

Fixes #13245

* fix(agent-status): suppress stamped turn replays

* fix(agent-status): notify paired clients at turn end

* Fix late-paired completion notification arming

* test(agent-status): make the notification-id test fail on the pre-fix ordering

The stored row inherited the helper's default codex agentType while the event
named claude, so agentSnapshotMatchesExplicitTitle dropped it and
freshStoredAgentStatus was undefined — the assertion held under either side of
the `??`. Name the stored row's agent so the pinned working row survives and
the snapshot-first precedence is what the test actually pins.

* Suppress stamped completion tail replays

* Prevent cross-coordinator title replays

* Preserve stamped tails across fallback signals

* Keep remount replay state while sibling lives

* Scope OSC turn stamp preservation

* Forward paired host completion stamps

* Deduplicate paired completion tails

* Bind paired completion tails to their turn

* Preserve paired completion tail ownership

* Keep paired tail replay state across remounts

* Seed paired recovery without replaying completions

* Seed startup replay and release stale fallback dedupe

* fix(notifications): preserve stamped OSC repaints

* fix(notifications): retain paired client turn boundary

* fix notifications module import safety

* chore: keep main integration focused

* fix: preserve completion re-enable boundary
2026-08-16 20:47:02 -07:00
Jinjing f070033156 Revert "refactor(shell): one portable Unix startup dialect instead of shell d…" (#14975)
This reverts commit b6ea3f17a9.
2026-08-16 16:48:55 -07:00
Neil b6ea3f17a9 refactor(shell): one portable Unix startup dialect instead of shell detection (#14863)
Orca had to guess which shell would parse a queued command line, then emit
syntax for it. Guessing is unreliable for a remote or WSL host, and every
dialect-dependent function is a place to get it wrong.

Replace the guess. Everything emitted for a Unix shell is now built to be
correct in sh, bash, zsh, dash, ksh and fish alike, so no detection is needed:

- quoteStartupArg emits backslashes as "\\" and apostrophes as "'" between
  single-quoted runs. Both families read that identically, unlike the sh '\''
  idiom, which fish silently halves and which makes a trailing backslash a
  hard syntax error.
- clearEnvCommand emits a self-contained fish/sh branch. It deliberately does
  NOT call a helper defined by Orca's shell wrappers: Orca wraps only zsh, bash
  and fish, so an `sh`/`dash`/`ksh` login shell launches unwrapped — and the
  same text is copied to the clipboard and pasted into shells Orca never
  spawned. In both, a helper would be `command not found`, which is the exact
  failure this exists to avoid. Two guarded statements rather than `A && B ||
  C`, because fish's `set -e` returns non-zero for an already-unset variable
  and would fall through to the sh branch; a trailing `true` pins the status,
  since this is the last statement of a launch line and the prompt renders it.
- One tokenizer for Unix. The input is a settings string the shell never
  parses, so parsing it per-shell only made the same setting mean different
  things in different workspaces.

AgentStartupShell loses its 'fish' and 'unix' members, and the three
login-shell resolvers, the fish tokenizer and the agentEnv.SHELL probe go with
them.

Per-worktree shell history now actually works:

- zsh on macOS was a no-op. /etc/zshrc assigns HISTFILE unconditionally before
  any wrapper Orca controls, so the injected value was already gone — and with
  ZDOTDIR still pointing at Orca's wrapper dir, history landed inside it. The
  intended path rides ORCA_HISTFILE and is restored after user config.
  Fixes #11044.
- fish keeps history in its own data dir keyed by session name, since it
  ignores HISTFILE and has no custom-directory knob. Files are deleted rather
  than truncated, a symlinked ~/.local/share no longer disables cleanup, and a
  GC sweep reclaims orphans whose meta.json is gone. The sweep refuses an empty
  live-worktree set (indistinguishable from a store that failed to hydrate) and
  skips files younger than GC_MIN_AGE_MS, mirroring the tree GC's guard against
  the live-set snapshot race.

Verified against real shells rather than asserted as strings:
startup-shell-portability.live-shell.test.ts runs 194 assertions across
sh/bash/zsh/dash/ksh/fish, and zsh-scoped-histfile.live-shell.test.ts drives a
real login zsh through /etc/zshrc. Both are vacuity-checked. The same quoting
corpus was replayed byte-exact on Linux, where /bin/sh is dash.
2026-08-16 15:28:50 -07:00
Jinwoo Hong 9e3e583a83 feat(relay): prefer the closest available region (#14366) 2026-08-16 13:51:04 -07:00
Jinwoo Hong fa9b20cb41 feat(skills): reland private bundle sharing safely (#14934) 2026-08-16 13:45:54 -07:00
Jinjing 3c8410d927 Revert "fix(browser): route every cookie-import write through CDP identities …" (#14942)
This reverts commit bf6dc6fcba.
2026-08-16 12:33:45 -07:00
Jinjing ef1224c4f7 Revert "Preserve OpenCode session across command completion, control SessionS…" (#14943)
This reverts commit 1da1bdc01c.
2026-08-16 12:33:28 -07:00
Brennan Benson bf6dc6fcba fix(browser): route every cookie-import write through CDP identities (STA-4300) (#14729)
* test(browser): repro STA-4300 CHIPS partition downgrade on native cookie import success path

* fix(browser): route every cookie-import write through CDP identities (STA-4300)

* test(browser): cover partition fidelity on both import write paths (STA-4300)

* fix(browser): never stage unreadable cookie partitions (STA-4300)

* fix(browser): gate partition skips on client support (STA-4300)

* test(browser): anchor the client partition-skip capability assertion (STA-4300)

* fix(browser): skip Firefox CHIPS cookies whose Chromium identity cannot be rebuilt (STA-4300)

* fix(browser): tolerate a Firefox schema without originAttributes (STA-4300)

* docs(browser): correct a comment that still named the removed cookies.set write

* fix(browser): block lossy Firefox cookie import on old clients

* fix(browser): read Firefox CHIPS from schema flag
2026-08-16 11:46:45 -07:00
Jinjing 763b1febeb Revert "feat(skills): add private bundle sharing (#14401)" (#14913)
This reverts commit 757fae28d7.
2026-08-16 10:39:57 -07:00
Jinjing 1da1bdc01c Preserve OpenCode session across command completion, control SessionStart emission (#14866)
* Preserve OpenCode session across command completion

- Add session start events and launch token tracking to establish session boundaries
- Defer retiring launch authority until OpenCode process actually exits, not just when a command finishes
- Fence previous tokens after restarts to prevent status updates from stale sessions
- Maps SessionStart as a session boundary for proper turn/state management

* Emit SessionStart only from OpenCode, not mimo-code

Restrict SessionStart lifecycle events to OpenCode exclusively. Mimo-code no longer emits SessionStart, as it should rely on OpenCode for session boundary signals. This prevents duplicate lifecycle events that could interfere with pane authority tracking and session state management. Also tighten foreground process result validation to reject stale results after title observation changes, fixing a race where a delayed foreground read from a previous cycle would incorrectly retire authority.
2026-08-16 10:01:27 -07:00
Jinwoo HongandE2E Test 757fae28d7 feat(skills): add private bundle sharing (#14401)
Co-authored-by: E2E Test <e2e@test.local>
2026-08-16 02:36:18 -07:00
Jinjing 8b04e060fa refactor(persistence): extract modules to half persistence.ts (#14252)
* refactor(persistence): extract modules to half persistence.ts

* refactor(persistence): tighten the extracted operations seam

Review follow-ups on the module extraction, all behavior-neutral.

The extracted operations read and mutate the Store's state object in place, but
every seam typed it as a bare PersistedState, so nothing at the boundary said a
caller must pass the live reference — a future caller handing over a clone would
have its writes silently dropped. Name that contract: StoreOwnedPersistedState
carries it to every operations interface and every mutating free function.
normalizePersistedPaneIdentityState and backfillFolderScopeConnectionIds stay on
PersistedState; they build a fresh state rather than mutating the Store's.

The six *PersistenceOperations wrappers were constructed per delegate call. They
are stateless today, so this was inert, but any future instance state would be
lost between calls. Memoize them, and mark state and gitUsernameCache readonly
so the compiler enforces the single-assignment invariant memoizing them relies
on.

Also: restore flushSshPtyConsumerRecovery, whose inlining left its rationale
duplicated at both call sites; document that migrateWorktreeIdentity's boolean
gates the caller's save, since the extracted function kept no docs of its own;
and merge a duplicate shared/types import that was failing lint under
--deny-warnings.

* delete plan doc

* refactor(persistence): add error recovery and improve field cleanup

- Rollback failed migrations to prevent corrupted state that blocks retry
- Gracefully skip malformed entries in normalization instead of aborting
- Strip retired fields to prevent orphaned state and sync issues

* refactor(persistence): drop the redundant persistence- filename prefix

The extracted modules already live in src/main/persistence/, so name
them after the domain they own. Point leftover shared/types imports
at the real type modules while touching those files.

* refactor(persistence): optimize lookups and fix unsanitized updates

- Use Maps instead of repeated array searches for O(1) lookups
- Apply sanitized updates instead of raw input in ui-state-update
- Compare fields directly rather than JSON strings to avoid false dirty states from persisted key ordering differences

* refactor(persistence): group modules into lifecycle folders

Move the 42 flat persistence modules into six folders named for what the
module does, and lift the Store class out of the barrel so persistence.ts
becomes an 8-line public surface.

Bodies are unchanged: every moved file diffs clean against HEAD once import
blocks are excluded. Only import specifiers were rewritten, by resolving each
one to an absolute path and mapping it through the move map.

Store keeps its existing max-lines suppression; its baseline entry is repathed
rather than re-added. Its 119-method public API sets a ~525-line floor, so it
cannot meet the 400-line cap without breaking the API for 153 importers.

* Sanitize worktree visibility sources and preferences on hydration

Ensure invalid or corrupted data from disk (untracked whitespace,
relative paths, bogus preference values) is cleaned during load
rather than corrupting the in-memory store.
2026-08-15 21:41:01 -07:00
Neil 73aa5d0ca7 refactor(daemon,runtime): split daemon, pty and rpc modules under the max-lines budget (#14834)
Splits the nine oversized modules in the daemon/provider/runtime domain into
focused per-concern files and drops their max-lines baseline entries.

- daemon: `Session` decomposes into an output plane (emulator, pending-output
  buffer, client fan-out), a producer-pause controller, a shell-ready barrier and
  a termination controller; `DaemonClient` into socket connect, hello handshake,
  ndjson readers, pending-request settlement, listener registry and notify
  settlement; `daemon-health` into pid-file parsing, process identity,
  stale-kill, TCC attribution and bundle staleness; `shell-ready` into the marker
  constant and the bash/zsh rcfile generators.
- providers: local-pty shell-ready wrapper generation, wrapper root, startup
  command and bash rcfile split out of local-pty-shell-ready.
- runtime: `Coordinator` sheds DAG convergence, decision gates, escalation
  triage, the runtime contract, the stale-base flag and task dispatch; the
  files/git/github rpc modules split into per-domain method groups.

Behavior-preserving: the extracted units keep their original construction order,
guards and timer lifetimes, and every RPC method name is still registered.
Test `vi.mock` surfaces were re-partitioned to follow the moved symbols.
2026-08-15 19:36:03 -07:00
Neil bc28107864 refactor(hooks,relay): split agent hook services and relay under the max-lines budget (#14725)
The four agent hook services, the main hooks module, and the two relay modules
each carried a file-level `eslint-disable max-lines` and ran 365-628 counted
lines against a 300-line budget. AGENTS.md calls for splitting rather than
suppressing, and config/max-lines-baseline.txt is a shrink-only ratchet, so this
removes all seven suppressions and prunes their entries (341 -> 334).

Pure move, no behavior change. Each hook service splits into its managed script
source, its config/bundle serialization, and its remote-install path, keeping the
per-agent integrations independent: copilot, amp, antigravity and hermes each
retain their own getManagedScript rather than sharing one, because each emits a
different script body for a different agent. Merging them by name would have
been a behavior change, not a refactor.

For antigravity the suppression's stated rationale -- that local install, Windows
wrapper generation, status cleanup, and SSH remote install must share one event
list and managed-command matcher so stale-hook cleanup cannot drift by platform
-- is now enforced structurally instead: both install paths call
buildInstalledConfig + createAntigravityManagedCommandMatcher over the single
ANTIGRAVITY_EVENTS catalog, with the graph a strict DAG.

Also registers the six new antigravity/ and copilot/ modules in
config/tsconfig.cli.json. That project uses a curated `include` list rather than
a glob, so an unlisted module fails `tsc -p config/tsconfig.tc.cli.json` with
TS6307 even though the entire unit suite passes.

Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(remaining failures are pre-existing load flakes in untouched files, green when
re-run serially), no new runtime import cycles, and no lint suppression added.
2026-08-15 18:25:37 -07:00
Neil 83117f2860 refactor(integrations): split issue-tracker clients under the max-lines budget (#14704)
The GitLab, GitHub, Jira and Linear integration modules, their two IPC
registrars, and the shared GitHub project types each carried a file-level
`eslint-disable max-lines` and ran 351-614 counted lines against a 300-line
budget. AGENTS.md calls for splitting rather than suppressing, and
config/max-lines-baseline.txt is a shrink-only ratchet, so this removes all
eight suppressions and prunes their entries (341 -> 333).

Pure move, no behavior change. Each client is cut along the seam it already
had: per-operation modules for the issue APIs (create / update / comment /
field options), and for Jira the request queue, site credential store,
authenticated request, and site identity. The two IPC registrars keep their own
handlers and delegate the rest to per-domain sub-registrars, so they remain
real entry points rather than re-export shims.

The IPC surface is proved intact rather than assumed: comparing (method,
channel) multisets between HEAD and the split gives 52 registrations across 52
distinct channels on both sides.

Provider-neutrality is preserved -- GitLab and GitHub keep separate, parallel
module layouts rather than being merged behind a shared abstraction.

Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(the one remaining failure is a pre-existing load flake in an untouched file,
green when re-run serially), no new runtime import cycles among 744 modules,
and no lint suppression added anywhere.
2026-08-15 18:17:20 -07:00
Neil 15e1ba3f84 refactor(ipc): split main-process IPC modules under the max-lines budget (#14703)
The six oversized src/main/ipc modules each carried a file-level
`eslint-disable max-lines` and ran 427-671 counted lines against a 300-line
budget. AGENTS.md calls for splitting rather than suppressing, and
config/max-lines-baseline.txt is a shrink-only ratchet, so this removes all six
suppressions and prunes their entries (341 -> 335).

Pure move, no behavior change. Each file is cut along the seams it already had:
pet splits into format allowlist / storage paths / symlink-safe copy / bundle
manifest + import; filesystem-auth into path-containment primitives, the
config-derived allow-list, and the git-registered root cache; notifications into
sound selection, native lifecycle, permission probe, and burst cooldown;
crash-reporting into renderer error reports, breadcrumbs, and sender.

The IPC surface is proved intact rather than assumed: comparing (method,
channel) multisets between HEAD and the split gives 49 registrations across 49
distinct channels on both sides. filesystem-auth's security boundary keeps its
acyclic layering -- containment primitives, then allow-list, then root cache,
then path-resolution orchestration -- with no layer gaining a back-edge.

Also keeps clipboard-ipc-handlers.test.ts under the 800-line test budget. The
split had briefly added a redundant vi.mock for isENOENT (byte-identical to the
real implementation) that pushed it to 801; the mock is dropped in favor of the
real function, with realpath added to the existing node:fs/promises mock.

Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(the one remaining failure is a pre-existing load flake in an untouched file,
green when re-run serially), no new runtime import cycles among 617 modules,
and no lint suppression added anywhere.
2026-08-15 18:08:54 -07:00
Neil c8fe5fc8c1 refactor(browser): split browser and browser-IPC modules under the max-lines budget (#14697)
The five oversized src/main/browser modules and src/main/ipc/browser.ts each
carried a file-level `eslint-disable max-lines` and ran 377-654 counted lines
against a 300-line budget. AGENTS.md calls for splitting rather than
suppressing, and config/max-lines-baseline.txt is a shrink-only ratchet, so
this removes all six suppressions and prunes their entries (341 -> 335).

Pure move, no behavior change. cdp-ws-proxy is decomposed into collaborating
objects rather than free functions because its state is genuinely
per-connection: every collaborator is a private readonly instance field built
in the constructor with live closures over `this`, so per-connection state
stays per-connection. Likewise the screencast pacer's isClosed/isStopping and
snapshot capture's getSeq are live thunks, not values captured at wiring time,
so guards inside already-armed timers still observe a later stop().

browser-guest-ui.ts is renamed to browser-guest-shortcut-forwarding.ts: after
the split it exports exactly one function, setupGuestShortcutForwarding, so the
old name no longer described its contents.

Also restores a single `webContents.debugger` read in the screencast path. The
extraction had left three reads where the original had one; the accessor is
stable today, so this is not a behavior fix but it removes a latent divergence.

Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(remaining failures are pre-existing load flakes in untouched files, each green
when re-run serially), no new runtime import cycles, and the IPC channel set
diffed identical before/after with all 23 handlers still trust-gated.
2026-08-15 17:59:29 -07:00
Jinwoo Hong d2ffe1f362 fix(terminal): settle CLI prompts for Claude and Codex (#14608) 2026-08-15 15:45:17 -07:00
Jinwoo Hong 2eb3e11327 fix(terminal): make close and handles incarnation-stable (STA-4327) (#14590) 2026-08-15 01:25:09 -07:00
Neil 9367169888 refactor(tests): split every oversized test file off the max-lines suppression list (#14728)
* refactor(tests): split oversized test files off the max-lines suppression list

Every `*.test.ts`/`*.spec.ts` that carried an `eslint/oxlint-disable max-lines`
directive is now split into focused, behavior-scoped suites that fit the 800-line
test budget, with shared setup extracted into co-located `*-test-harness.ts` /
`*-test-fixtures.ts` modules (300-line budget). 83 files became ~930; the largest
output is 797 effective lines. `orca-runtime.test.ts` is intentionally untouched.

Test bodies were moved by scripted line-range slicing rather than retyped, so
assertions are byte-identical. The only permitted body edits were mechanical
rebinding where a shared value moved into a harness (e.g. `tmpHome` ->
`homes.tmpHome`).

Registries that enumerate test files were updated in lockstep:
- config/max-lines-baseline.txt: pruned 341 -> 258 entries (all 83 removed).
- config/reliability-gates.jsonc: 33 gates repointed at the split files, with
  assertionRefs split per file where a gate's coverage now spans several.
- .github/workflows/pr.yml: the real-zsh lane now lists the 4 split files that
  actually exercise zsh, so they keep running in the dedicated shell lane.

Also renamed agent-hooks `server-test-fixtures.ts` to `server.test-fixtures.ts`
so the global-fetch call-site audit keeps skipping it, and added `.js` extensions
to the CLI suites' dynamic harness imports (node16 resolution) to unbreak
`build:cli`.

Verification: full suite 52,449 passing vs 52,448 at baseline with zero
assertions lost; `pnpm lint`, `pnpm typecheck`, and `pnpm build:cli` all exit 0;
the terminal-pane e2e spec runs 31/31 headless.

* refactor(tests): split hook-idle arbitration suite that oxfmt pushed over budget

The pre-commit oxfmt pass reflowed pty-connection-hook-idle-arbitration.test.ts
to 811 effective lines, 11 over the test budget. Split the hook-completion side
effect and replacement-agent veto cases into their own suite; both files now sit
well under the cap and the 15 tests are unchanged.

* test: port upstream test changes into the split files after rebase

Rebasing onto main surfaced 27 tests that main had added to files this branch
deleted, plus edits to tests that had already moved. Taking the deletion side of
those modify/delete conflicts would have dropped that coverage silently, so each
upstream change is ported into the split file that now owns the behavior — for
example main's six orchestration mailbox tests land across orchestration-runs,
-send, and -check.

Also repoints `orchestration.notification-mailbox-consistency`, a gate main added
after this branch's gate remap, at those same three split files, and re-prunes
the max-lines baseline against main's (257 entries).

Verified: all 27 upstream test titles present; full suite 52,761 passing with the
only diff vs baseline being 12 tests main itself removed and 3 that moved from
skipped to passing; lint and typecheck exit 0.

* fix(test): flush pending continuations before tearing down terminal test globals

CI shard 5/16 failed on both Node 24 and 26 with `ReferenceError: window is not
defined` from pty-connection.ts, surfacing through
pty-connection-daemon-snapshot-replay.test.ts.

The reattach/settle chains `await` a real promise and then touch `window.api`.
Under fake timers those continuations cannot run, so they only become schedulable
once restoreTerminalTestGlobals() switches back to real timers — which previously
happened immediately before `delete globalThis.window`, so a late continuation
threw and failed the whole file. Flush async ticks in that window instead.

This is latent in the source rather than new: the pre-split 25k-line file kept
running other tests after these, which gave the chains time to settle before
teardown. Splitting the file moved teardown directly behind them.

* fix(test): keep an inert window after terminal test teardown instead of deleting it

The async-tick flush was not enough: the reattach/settle chain can resolve after
teardown regardless of how long we drain, so CI shard 5/16 still failed with
`ReferenceError: window is not defined` from pty-connection.ts.

A real renderer never loses `window`, so deleting it was the artificial part.
Swap in an inert proxy whose properties resolve to callables and whose calls
resolve to undefined, making a late `window.api.pty.*` call a harmless no-op.
The next test replaces it wholesale via installTerminalTestGlobals(), and no test
asserts that `window` is absent.
2026-08-15 00:54:20 -07:00
Brennan Benson 375b735e9c fix(agent-launch): preserve cold Codex startup drafts (#14688)
* fix(agent-launch): preserve cold Codex startup drafts

* fix(agent-launch): honor startup draft readiness budgets
2026-08-14 23:16:29 -07:00
Brennan Benson ab9d1a29a9 fix(worktree): never reissue a generated workspace name (#14350)
* fix(worktree): never reissue a generated workspace name

Generated workspace names were deduped only against currently-live
worktrees, so deleting a workspace returned its name to the pool. A later
workspace could draw the same name, land on the same directory path, and
inherit the previous occupant's agent conversation history — coding-agent
CLIs key their prompt history and transcripts by cwd.

Names are now retired permanently per repo. The registry is written in
main with the name Git actually used (the create loop can advance past a
requested name on collision), and seeded once per run from workspace
directories and surviving agent transcript buckets so already-spent names
are excluded from the start. Suggestions degrade to -2, -3 variants
instead of recycling, and those variants retire too.

User-typed names are untouched: retirement filters suggestions only.

* fix(mobile): honor retired workspace names, on one shared implementation

Mobile hand-duplicated the desktop name-suggestion algorithm and deduped
only against live workspaces, so a phone could still be offered a name
whose deleted workspace left agent conversation state behind at that path.

Both platforms now call one shared selector in src/shared, so the two can
no longer drift. The host publishes retired names as an optional field on
the existing worktree.list response, and mobile fetches them per selected
repo while the create sheet is open — mirroring the desktop hook.

Mobile never calls worktree.list for its catalog (it uses worktree.ps,
which carries rows only), so this is a targeted request rather than a
change to the catalog or its cache. Hosts predating the field omit it and
mobile falls back to live-only dedupe, which is the pre-change behavior.

* fix(worktree): close retirement consistency gaps

* test(worktree): cover retirement runtime contracts

* fix(worktree): retire generated collision names

* fix(worktree): enforce retired names at creation

* refactor(ai-vault): extract the Claude project-dir encoder

The bucket-name encoder and its scope-boundary check were private to the
session scanner, so a second consumer had to reimplement them — and got the
per-character encoding wrong. Move both to a shared module with direct tests.

* fix(worktree): make the retirement seed scan actually match buckets

The bucket encoder collapsed runs of non-alphanumerics while the real one
emits a dash per character, so every dot-path bucket missed and the Windows
default workspace root (C:\...) matched nothing at all. Reuse the shared
encoder and its boundary check, which also stops a repo absorbing a sibling
whose path merely shares its prefix.

Also:
- Derive the workspace leaf by stripping the known encoded parent instead of
  guessing from trailing dash segments, which retired the parent directory's
  name whenever a workspace was named numerically.
- Reuse isAutoGeneratedCreatureBranchName so the -10 and -100 tiers retire.
- Drop the .codex/sessions root: Codex keeps the cwd inside the transcript
  rather than in a directory name, so the scan could only ever see a year
  folder. Reading transcript contents is not a trade this feature justifies,
  so the gap is documented instead.
- Honor CLAUDE_CONFIG_DIR, which relocates the bucket root.
- Delete the unused retirableLeafName export.

Tests write buckets with the real per-character encoding against a fake home,
covering POSIX, dot-directory, Windows drive and WSL UNC roots; all three
platform cases fail against the previous encoder.

* fix(worktree): retire only generated names, keyed by cwd namespace

Two problems in the host-side registry.

Retirement fired for every create, including names the user typed. The
creature pool contains ordinary words — orca, runner, sole, molly, oscar — so
typing a retired 'nautilus' silently produced directory and branch
'nautilus-2' and burned the name for good. Creates now carry an explicit
nameWasGenerated flag; both the skip and the retire are gated on it, and it
defaults to false so CLI and automation callers are unaffected.

The registry was keyed by repo id, but both readers already discarded the id
and unioned by the cwd collision key, because the collision this prevents is
on the path. Keying by that namespace directly fixes several things at once:
entries no longer orphan when a repo is removed, remove/re-add no longer loses
every retirement for an unchanged path, the missing removeProject prune is
moot, and the backfill promise no longer merges into only the first repo id it
saw. The feature is unreleased, so no migration is needed.

Also:
- Memoize the collision key. It runs computeWorktreePath, which for a WSL repo
  is a blocking execFileSync('wsl.exe') whose failure path is uncached, and
  the previous code recomputed it once per repo on every create and every
  listRetiredNames call.
- Drop retiredNamesByRepo from the worktree list result. It had no readers and
  leaked onto 'orca worktree list --json', and its awaited backfill sat on CLI
  selector resolution. The dedicated listRetiredNames RPC keeps its consumers.
- Make the three RuntimeStore methods required. RuntimeStore is file-private
  with two constructors, so the 'older embedders' the optionality protected do
  not exist, and the optional chain silently returned no retirements.
- Revert the unrelated forceDeleteBranch rewrite, and make room under the
  file's line budget by extracting the create-args mapping instead.

* fix(worktree): send name provenance and stop gating Create on the fetch

Desktop and mobile now mark a create as generated-name only when the user
typed nothing and the composer fell back to the suggestion, so the host knows
which names it may retire.

Remove the retired-names loading gate from every create path. The host already
skips retired candidates before doing any git work, so the client gate bought
nothing while it could disable Create for the length of a full mobile
reconnect ladder (the wait had no timeout) and blank the desktop button
between queued creates. The suggestion still waits; the button never does.

Also make the web client call worktree.listRetiredNames instead of hardcoding
an empty list — the method is registered and mobile-allowlisted, so the
comment claiming no wire call existed was wrong — and filter the mobile
response to strings so a malformed row cannot throw during normalization.

* fix(worktree): key retirement by repo id and prune it with the repo

Reverts the collision-key storage key. It was a function of workspaceDir,
nestWorkspaces, worktreeBasePath and repo.path, so toggling any one of those
orphaned every retirement for every affected repo at once — trading a rare
churn (remove/re-add) for a common one. The read path already unions by cwd
namespace at query time, so cross-repo sharing never depended on the storage
key.

Instead, address the growth and orphaning directly:
- Drop the registry in removeProject, and in removeProjectForHost once the last
  host's copy of the repo id is gone, alongside the sparse-preset deletes that
  already follow this convention.
- Bound each repo's registry. The cap sits far above the 552-name pool because
  evicting inside it would reissue a name whose agent state is still on disk;
  only -2/-3 tier accumulation can ever reach it.
- Carry retirements through profile transfer, re-keyed to the destination repo
  id and dropped from the source, mirroring sparsePresetsByRepo.

Separately, fix the backfill merge: the scan promise is cached per cwd
namespace, but it closed over the first repo id that triggered it, so a second
repo in the same namespace received nothing. The scan stays shared; the merge
moves out of the cached promise and runs for whichever repo asked.

Local repos re-seed on re-add through that backfill. SSH repos do not — the
scan cannot see the execution host — which is now stated in the module.

* docs(worktree): spell out why the retirement bound sits above the pool

Names the trap directly: the neighbouring 50/200 bounds cap histories, so
lowering this one to match them would silently start reissuing names whose
agent state is still on disk. Also states that oldest-first eviction is a
deliberate least-bad choice rather than a neutral one.

* fix(worktree): send name provenance from the web runtime client

This client hand-enumerates worktree.create params, so the new optional field
was silently dropped and typecheck could not see it. On web and paired-desktop
the host therefore never received it: generated names were never retired, and
the host-side skip that backstops a stale suggestion was disabled too. The same
client does fetch retired names for suggestions, so it was filtering against a
registry nothing ever wrote to.

The test asserts both directions, and fails without the fix.

* fix(worktree): retire names that took more than one collision suffix

isAutoGeneratedCreatureBranchName strips exactly one trailing -N, which is
right for auto-rename eligibility but wrong here. Once the pool is spent the
suggester emits nautilus-2, and a collision on that yields nautilus-2-3 —
which a single strip leaves as nautilus-2, not a pool name, so retirement
no-opped at exactly the tier where every base name is already gone. Strip
repeated suffixes locally rather than moving the auto-rename predicate.

* perf(worktree): keep the retirement backfill off the blocking WSL probe

The backfill runs on composer repo-select, not just at create time, and it
derived the probe path synchronously — which for a WSL repo with a mirrored
workspace dir reaches getWslHome and its blocking execFileSync('wsl.exe').
A stopped distro froze the main process for up to 5s on composer open.

Adds an async twin of computeWorktreePath and uses it for the probe. Resolving
the home there also warms the shared cache, so later sync callers are free.

Also stops memoizing the collision key when the WSL home is still unresolved:
only the success path is cached upstream, so caching the fallback namespace
would strand the repo there for the rest of the session.

* fix(worktree): hold retired names across a refresh instead of blanking

refreshKey changes on every workspace-list mutation, so create-multiple
refetches after each create and the hook returned an empty list until the
refetch landed — precisely the window in which resetForNextCreate clears the
name field and a fresh suggestion is drawn. Keep the previous answer while
revalidating and reset only when the repo changes; a failed refresh keeps what
was already loaded rather than un-retiring everything.

Also makes the returned array referentially stable, so the suggestion memo
downstream stops rerunning on every refetch.

* refactor(worktree): put the retired-name cache rules on one implementation

The desktop and mobile hooks that fetch retired names had already drifted
four ways. The transports genuinely differ (IPC vs RPC), but the caching
rules must not, and mobile's copy reset to [] on any error -- which
un-retires every name for the rest of the sheet session, the one outcome
retirement exists to prevent.

Moves the rules into src/shared/worktree/retired-name-cache: response
normalization, the never-leak-across-repos rule, and the hold-previous-on-
failure rule. Pure, no React, because src/shared is on the main process's
import graph. Each platform keeps its own transport and effect.

Mobile moves up to desktop's behavior: it now holds the previous answer
through a failed refresh, and refetches when the workspace list changes
instead of never refetching after mount.

Also drops the unused `loading` return. Neither platform consumed it; its
only consumer was the Create-button gate reviewed out earlier, and removing
it makes that regression unexpressible.

* fix(worktree): import shared types from their real modules

Main dropped the src/shared/types barrel, so the retirement module's import
resolved locally but not against the PR's merge base.

* refactor(worktree): bound the retirement registry by tier compaction, not eviction

Retirement is a correctness guarantee — a spent name's directory may still hold
agent conversation state keyed by that cwd — so the 2000-entry cap was the wrong
shape: reaching it handed a name back. At the owner's measured rate (~6.6 pool
names retired per day in one repo) the cap was ~9 months out.

Names come from a fixed 552-entry pool and the suggester only reaches tier N+1
once every tier-N name is taken, so a completed tier is exactly a set that no
longer needs listing. A row is now a watermark plus the names above it: reads
answer at-or-below the watermark with no lookup, and compaction drops the 552
entries the watermark now covers. Bounded at one pool per repo forever, with no
eviction and nothing un-retired.

Tiers can complete out of order (a create-time collision can spend `nautilus-2`
while tier 1 is open), so compaction loops and higher-tier names simply wait.

The RPC result carries the watermark beside the names as a new field; a client
predating it reads the names only and under-retires the compacted tiers, which
degrades to the pre-retirement behavior rather than breaking.

* fix(worktree): preserve generated name retirement across failures
2026-08-14 22:18:36 -07:00
fsdwenandBrennan Benson 190d153194 fix: detect external git init on folder projects and upgrade to git repo (#11480)
* Remove scheduled triggers from E2E and README badge workflows

* fix: detect external git init on folder projects and upgrade to git repo

Closes #11477

Three root causes fixed:
1. buildWorktreeBaseDirectoryWatchTargets continued for folder repos
   - now register parent dir as base watch target so poller sees .git creation
2. No path re-evaluated kind after registration
   - add tryUpgradeFolderRepo, checks .git on structural change, calls
   store.updateRepo(repoId, { kind: 'git' })
3. No IPC signal after store update
   - emit repos:changed so frontend git polling re-evaluates

* revert: drop the base-watch-target approach to folder-project git detection

Registering dirname(repo.path) as a base watch target makes the existing poller readdir
the parent directory and stat every sibling, so a project under the home directory scans
the whole home directory on every poll. Replaced by a per-repo .git poll in the following
commits.

* fix(folder-projects): upgrade to a git repo when an external git init lands

Folder projects were registered once as kind: folder and never re-evaluated, so
running `git init` outside Orca left them without any git affordances until a
restart (#11477).

Poll `<repo>/.git` for each local folder project on the base-watcher cadence and
flip kind to git when the marker appears, matching a freshly added git project
(explicit externalWorktreeVisibility, prepared worktree root) before notifying the
renderer and resyncing the base watchers.

One stat per folder project per tick, parked while the window is hidden, backed off
to 30s while no folder project exists.

* fix(folder-projects): reuse the shared repo-change notifier and stop reading the store at attach

Perf audit follow-ups: reading getRepos() synchronously in attachMainWindowServices
broke every test in that file and put O(repos) hydration on the startup path, and a
bare repos:changed send skipped the paired-client broadcast (#11994). Also invalidate
the authorized-roots cache the way the runtime's own folder->git path does.

* fix(folder-projects): keep the project's workspace visible when git's root differs from the stored path

Electron QA found the golden path breaking for a folder project whose path traverses
a symlink: Add Project stores git roots as rev-parse reports them, folder projects
keep the raw path, so after the upgrade the root checkout reads as an *external*
worktree and externalWorktreeVisibility: 'hide' hid the project's only workspace.
Only set 'hide' when git's toplevel matches the stored path.

Also gate the upgrade on isGitRepo so a stray .git file cannot flip a project, and
switch the tests to real git init so both guards are exercised against real git.

* fix(folder-projects): refuse non-root folders and stop re-probing git for a rejected marker

Review round found three real defects:
- A folder project inside another repo's work tree upgraded with repo.path pointing at a
  non-root subdirectory, because git accepts any path inside a work tree. Refuse unless
  git's toplevel resolves to the project directory itself.
- A .git git keeps rejecting re-ran two synchronous git spawns every 2s forever. Cache the
  verdict against the marker's stat signature and re-probe only when the marker changes.
- The poll kept probing after the window was destroyed (macOS keeps the app alive with no
  window), so idle out there instead.

Tests: build the symlink explicitly instead of relying on macOS TMPDIR being one, so the
spelling-mismatch case runs on Linux and Windows CI too; count real git probes; assert
per-project stat counts instead of a modulus; make the idle-backoff test observe the
interval it names.

* fix(folder-projects): refuse the upgrade when it would destroy the project's workspaces

Reproduced in the app: a folder project with extra workspaces went from three sidebar
rows to one within ~2s of an external git init, and their lineage was pruned.

A folder project's extra workspaces are worktreeMeta rows keyed
repoId::path::workspace:<uuid>, and only the folder branch of the worktree listing knows
those keys. Flipping kind moves the repo onto the git branch, which lists git worktree
list (one path) and prunes every lineage id under the repo that is not in it.

Migrating that meta belongs to the listing code that owns both shapes, not to this watch,
so refuse the upgrade for those projects. They keep working exactly as they do today.

* test(folder-projects): wait for the stat count instead of a fixed number of ticks

A tick that spawns git can outrun a fixed wall-clock wait on a loaded machine, so the
rejected-marker test failed roughly one run in six. Poll for the stat count with a
deadline; the load-bearing assertion (git probed exactly once) is unchanged.

* fix(folder-projects): wake git upgrade checks on catalog changes

* docs(folder-projects): align upgrade polling rationale

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-14 18:47:14 -07:00
Jinwoo HongandJinwoo-H 1d8eaa81c7 fix(orchestration): preserve mailbox delivery identity (#13717)
* test(orchestration): reproduce mailbox pointer mismatch

* fix(orchestration): align pointers with actionable mailboxes

* fix(orchestration): harden pointer reservation lifecycle

* fix(orchestration): bound mailbox reconciliation

* fix(orchestration): make pointer staging restart-safe

* test(orchestration): expect scoped dispatch index

* fix(orchestration): guard skewed inbox indexes

* refactor(orchestration): extract mailbox notification lifecycle

* fix(orchestration): settle mailbox pointer writes

* fix(orchestration): bound mailbox recovery work

* fix(orchestration): fence inactive mailbox snapshots

* test(orchestration): cover mailbox notification boundary

* test(orchestration): pin STA-4325 delivery identity

* fix(orchestration): preserve mailbox delivery identity

* fix(orchestration): harden mailbox settlement

* test(orchestration): make mailbox gates self-contained

* fix(orchestration): preserve paged mailbox ownership

---------

Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
2026-08-14 18:36:51 -07:00
Brennan Benson 78d5920446 fix(orchestration-cli): point dropped mutations at --retry-request (#14586)
* fix(orchestration-cli): guide dropped mutations to idempotent retry

* test(orchestration-cli): preserve read-only drop message

* fix(orchestration): harden mutation replay identity

* fix(orchestration): preserve replay across remints

* fix(orchestration): defer local mutation identity
2026-08-14 18:11:12 -07:00
9cb04b5a35 fix(orchestration): stop fencing fresh-run callers and accept revoked coordinator after takeover (#11582)
* fix(orchestration): stop fencing fresh-run callers and accept revoked coordinator after takeover

LegacyCoordinatorAuthority.resolve was forcing the adopted legacy run for
every orchestration preflight and throwing legacy_read_only at any caller
that could not prove legacy-coordinator identity — including fresh-run
coordinators with no connection to the legacy system. Now only callers
previously known to the legacy run are fenced; unknown fresh-run callers
fall through to the normal current-run handler.

isLegacyCoordinatorHandle returned only the committed principal's handle,
so after a takeover that revoked the principal, worker_done/escalation to
the new coordinator was rejected with "not a retained coordinator". Now
both the retained legacy handle and the current run binding's handle are
accepted, so workers can deliver lifecycle mail to either coordinator.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(orchestration): deliver legacy lifecycle mail to the replacement coordinator

After `run-use --takeover-legacy` revokes the old coordinator principal, a
retained legacy worker addressing worker_done/escalation/ask at the new
coordinator was rejected with "not a retained coordinator", so the dispatch
stayed open forever.

Add a recipient-side permit, isLegacyCoordinatorDeliveryTarget, that also
accepts the current Run binding's coordinator handle. Both handles already
route to run:<id> via resolveLegacyWorkerCoordinatorDelivery.

isLegacyCoordinatorHandle stays narrow: #11745 reused it as the caller-side
fence jurisdiction, where widening it fences MORE callers and replaces an
actionable run_required with dead-end legacy_read_only guidance.

Co-Authored-By: Leonardo <leonardo.marciano@toolzz.me>

* fix(orchestration): keep the delivery permit in step with the takeover router

Round 1 review fixes on top of the recipient-side permit.

isLegacyCoordinatorDeliveryTarget accepted any handle bound as the Run's
coordinator, but resolveLegacyWorkerCoordinatorDelivery only promotes to
run:<id> once the legacy principal is no longer committed. bindRun leaves a
committed principal alone when it rebinds without a takeover over live legacy
work, so a coordinator restarting inside the legacy pane produced a permitted
send that routed legacy_direct to a handle no reader can see: current-contract
inboxes require current_delivery, and legacy mail requires a principal on that
handle. Gate the binding branch on the same takeover test the router uses, so
that send goes back to request_mismatch instead of vanishing.

Tests: cover the ask call site (it had none — reverting question.ts alone
failed nothing), assert the takeover bind landed inside the helper rather than
two assertions downstream, and replace the not-legacy_read_only assertion on
the fence guard with the concrete outcome it means to protect.

Drop the fresh-run coordinator tests: they guard #11745/#11802, already on
main and already covered by orchestration-legacy-fence-jurisdiction and
orchestration-11745-regression-verification, and they pass with this fix
reverted. Rename the file to what it now contains.

* refactor(orchestration): drop the unreachable pane-key clause from the delivery permit

The permit's takeover branch must mirror resolveLegacyWorkerCoordinatorDelivery,
which tests only the principal status. Every runs-table write sets
coordinator_handle and coordinator_pane_key together, so the extra pane-key
term never fires — and if it ever did it would deny mail the router would have
promoted to the readable run mailbox.

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-14 18:03:07 -07:00
Jinwoo Hong ff9bc0f079 fix(orchestration): wait for Codex composer render (#14575) 2026-08-14 12:56:35 -07:00
Brennan Benson 83e2123582 Add global worktree visibility source defaults (#14276)
* Add global external worktree visibility defaults

* Expand global worktree visibility source defaults

* Fix host-scoped visibility settings races

* Fix global worktree visibility integration

* Enable source visibility defaults on mobile

* Polish external worktree settings navigation

* Clarify inherited worktree visibility settings

* feat(sidebar): replace the inherited-visibility switch with a Show/Hide picker

Each source row now shows a two-segment Show / Hide control preselected to the
global setting, and explains itself only where the project actually disagrees:
an "Overriding global setting: <value>" card names the value being ignored.
Picking the segment global already holds drops the override instead of pinning
a duplicate, so the same control both overrides and reverts, retiring the
separate "Use global" link. The dialog footer now lists every inheritable
source with its global value.

* fix(sidebar): preserve reset for matching visibility overrides
2026-08-14 12:15:58 -07:00
Langning ZhangandJinwoo-H 92b6ffd17d Terminate renderer graph reload generations and contain disposed-frame notifications (#14070)
* fix(runtime): terminate renderer graph reload generations

* fix(runtime): harden renderer reload teardown

* fix(runtime): fence renderer graph publication ownership

* test(runtime): register renderer graph reload gate

* test(runtime): record live reload validation

* fix(runtime): ignore cancelled renderer navigations

* chore: preserve main formatting during branch sync

* chore: satisfy changed-code quality gate

* fix(runtime): restore cancelled renderer reloads

* fix(runtime): preserve committed reload fencing

* test(runtime): prove cancelled reload timeout

* docs(reliability): record reload cancellation oracle

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
2026-08-14 01:21:00 -07:00
Neil b82a8791f5 perf(runtime): resolve an explicit worktree id without scanning every repo (#14399)
`resolveWorktreeSelector` resolved every selector kind from the whole-fleet snapshot, so a targeted `id:<repoId>::<path>` lookup fanned `git worktree list` across every registered repo to answer a question about one of them. With a cold scan cache -- app startup, or the first lookup after a mutation clears the snapshot -- that is one subprocess per repo, ~17ms each, to find a worktree whose owning repo the id already names. Measured on a ten-repo fleet: one `id:` lookup scans 10 repos before and 1 after.

Scope only `id:`. Every other selector kind is matched across the fleet and its `selector_ambiguous` contract is defined over all repos, so scoping `branch:`, `name:`, `issue:`, or a bare selector would silently pick a winner where they correctly refuse today. A test pins that: `branch:main` across ten repos still throws `selector_ambiguous` and still scans all ten.

Lineage stays correct because edges are intra-repo by construction. The scoped path returns null and falls back whenever that does not hold: a repo id registered on several execution hosts, an unknown repo id, or a worktree the scoped scan does not contain. A warm fleet snapshot always wins.

Row resolution moves out of orca-runtime.ts into repo-worktree-row-resolution.ts, which owns no state -- the cache-aware scan and folder-workspace stamping are injected. orca-runtime.ts ends up 65 lines shorter than before despite the added feature.
2026-08-14 01:15:30 -07:00
Jinjing a6a64439a0 fix(terminal): keep split error when rejected cleanup throws (#14463)
Wrap kill and retireRejectedPty so a cleanup failure cannot replace
the original split-authority error or skip the remaining teardown.
2026-08-14 00:47:45 -07:00
NeilandOrca 2100fb2553 fix(runtime): cap remote git.diff and file previews at the transport budget (#14160)
* fix(runtime): cap remote git.diff and file previews at the transport budget

A remote or mobile user who opens the diff of a large image loses their whole
WebSocket, not just that request: the E2EE channel closes with 1013 when a reply
exceeds the 4 MiB outbound envelope. Two producers can exceed it unaided.

git.diff/branchDiff/commitDiff cap text with MAX_RENDERED_DIFF_COMBINED_CHARACTERS
(6M chars) -- a *renderer* budget that sits above the transport limit -- and return
base64 for previewable binaries bounded only by MAX_GIT_SHOW_BYTES, so a 10 MiB PNG
changed in place is ~26.7 MiB in one envelope. files.readPreview inlines base64 up
to 10 MiB, and mobile calls it for every image tab.

Both now measure against a budget derived from the outbound limit. The check sits in
orca-runtime-git.ts, downstream of the dedupe and of both the SSH-provider and local
branches, so a payload forwarded verbatim by an old relay is covered by the same code
and src/relay needs no change. Local and in-process callers pass no budget and keep
full fidelity.

Measuring raw bytes would not work, which is the whole reason this needs a module.
JSON escaping turns one control byte into six (\u00XX), and binary-buffer.ts sniffs
only for NUL in the first 8 KiB -- so a NUL-free file of 0x01-0x1f bytes is classified
as *text*, would pass a raw-byte cap, and would then blow the envelope. The budget is
escape-aware, with a three-branch fast path that keeps normal diffs at two native
byteLength calls and scans only the ambiguous band.

The SSH branch of readFileExplorerPreview had the same raw-vs-escaped gap: its stat
gate sizes base64 binaries, but text crossed unbounded. It now honours the same
decoded-text limit the local branch already enforced.

No wire change: GitDiffResult is untouched -- no third kind, no new field. Old clients
see an error for one request instead of a dropped connection. diff_too_large joins the
structured passthrough codes and lands on an existing error arm in both mobile
consumers and the desktop remote path; file_too_large was already handled on both.

Instruments the 1013 close, which nothing measured before, so the incidence this cap
is meant to drive to zero is finally observable. `emitter` separates a producer size
bug from a wedged link.

Known regression: remote image previews between ~3.096 and ~3.146 MB now return
file_too_large. They only intermittently worked before -- above ~3.0 MB they killed
the socket -- so this trades intermittent connection loss for a consistent error.

Test: 10281 passed in src/main/runtime + src/shared + src/main/git; mobile 3427
passed. Each of the six budget-enforcement sites is independently mutation-killed.
Escaping fixtures cover newline-dense, control-char, CJK, lone-surrogate and base64
content against native JSON.stringify. tsc clean for node, web and cli; oxlint clean.

Co-authored-by: Orca <help@stably.ai>

* fix(runtime): harden remote reply transport budgets

* test(runtime): cover desktop remote preview budgets

* test(runtime): close telemetry review gaps

* chore(shared): repoint budget imports after the shared/types barrel removal

Upstream #14447 dropped the shared/types barrel; GitDiffResult now lives in
git-diff-compare-types and GlobalSettings in global-settings-types.

Co-authored-by: Orca <help@stably.ai>

* fix(ssh): surface an over-cap preview read as file_too_large

The stream reader aborts an over-cap read with StreamProtocolError, whose numeric
code falls through mapRuntimeError to a generic runtime_error carrying the raw
"Reported totalSize N exceeds client cap M" string. Neither preview client
recognizes that: runtime-file-client.ts and mobile-file-preview-response.ts both
key on file_too_large. It also made the two file_too_large guards directly below
the read unreachable on the streaming path.

Gives the cap its own error type so the caller can translate it, keeping the
bandwidth saving the cap exists for. A genuine protocol fault still propagates
unmasked.

Found by the readiness review. Mutation-verified: removing the translation fails
exactly the new test.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-14 00:24:57 -07:00