Commit Graph
10363 Commits
Author SHA1 Message Date
Jinwoo-H eb76d52ff4 test(e2e): pin the hosted mobile WebView SSH spec's unrun status
The spec runs in no CI job. It cannot: it needs an iOS simulator and a Docker
daemon on one runner, and no GitHub runner has both — the whole
hosted-mobile-webview e2e family is manual-only, not just this spec. It already
carries the fallback the finding asks for: two test.skip lines naming both
missing prerequisites, so it never reports a green skip as coverage.

What was only a YAML comment is now a checked contract: the exclusion from the
changed-spec lane, the named manual entry points for the dev and packaged runs,
and the spec's own refusal to run without either prerequisite. A rename or a
dropped skip now reddens instead of quietly changing what CI covers.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-04 02:15:42 -04:00
Jinwoo-H e02097ad8e docs(browser): record why the navigation screencast event needs no gate
Traced the receiving side rather than assuming: handleStreamingResponse passes
every screencast result straight to the listener, use-mobile-browser-stream
casts the payload with no schema parse, and handleBrowserScreencastEvent is an
if/else-if chain with no else and no throw. A build predating the member drops
it silently, so Rule 3 is satisfied without a capability gate.

Pin that with a test, since the reasoning only holds while the handler stays
tolerant of an unrecognized type.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-04 02:12:45 -04:00
Jinwoo-H 446ebc354b perf(mobile-web): bound the gzip chunk cache
gzipChunks was keyed on the client-chosen `length` and cleared only when a new
package fingerprint verified — which on a shipped desktop is never, since the
bundle is fixed for the life of the install. A client walking every offset at
all eight permitted lengths caches each source byte 36 times, so request
parameters alone could retain ~100 MB in the main process with no cap and no
TTL. acquireRead bounds in-flight bytes, not retained ones.

Keep the key (dropping `length` would mean gzipping per request and losing the
cache) and bound the map instead: an LRU over total compressed bytes with a
16 MiB cap, reset alongside the existing fingerprint clear.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-04 02:10:23 -04:00
Jinwoo-H b5fe05188f fix(packaging): stop packing the mobile-web bundle into app.asar
`files` is exclusion-only, so electron-builder's default `**/*` packed both
out/mobile-web-rnw and out/mobile-web-rnw-export on top of the
Resources/mobile-web copy mobileWebExtraResource makes. Only that resource copy
is ever read: resolveMobileWebPackageRoot joins process.resourcesPath when
packaged and never consults the asar. At the 10 MiB build budget that is up to
~20 MiB of dead bytes per installer, and -export is the pre-hardening Metro
artifact that still contains the eval/new Function forms
disableRuntimeCodeGeneration strips.

The new assertion drives the real FileMatcher rather than pinning the pattern
strings, so it proves the trees are excluded, and re-asserts the extraResources
copy so the exclusion cannot orphan the feature.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-04 02:08:33 -04:00
Jinwoo-H bac4bea36e fix(mobile-rpc): allowlist files.unwatch and derive the required pairs
files.watch was mobile-allowlisted but files.unwatch was not, so the teardown
RpcClient sends on cancel was rejected as forbidden and the source-control
view leaked a @parcel/watcher handle per workspace visit. The client's cancel
swallows the rejection, so it was silent.

The hand-maintained MOBILE_STREAMING_CLEANUP_RPC_METHODS list existed for
exactly this class and still missed it. The allowlist test now derives the
subscribe -> unsubscribe pairs from buildServerSubscriptionUnsubscribe's own
map and requires the unsubscribe to be allowlisted and registered whenever the
subscribe is allowlisted; the hand list keeps only the cleanups that map does
not cover.

The leak was bounded: files.watch registers its cleanup with the connection id
and runtime-rpc-lifecycle calls cleanupSubscriptionsForConnection on socket
close, so watchers were released when the socket died, not before.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-04 02:07:03 -04:00
Jinwoo-H f1cfb2c6ee fix(terminal): guard binary query-reply frames like terminal.send
Opcode 18 set inputKind:'query-reply' on the client's word alone, and that
kind forces clientId undefined in sendTerminalStreamInput, so the frame
skipped beginMobileInputFloor entirely. A phone that is not the elected
reply authority could write arbitrary bytes into a desktop-driven pane
unfloored.

Lift the terminal.send guard into terminal-query-reply-guard.ts and apply it
at all three sites. A binary frame that fails is dropped, matching the JSON
path, which throws on shape and returns accepted:false on authority — neither
demotes the bytes to ordinary input. The page-side sender now runs the same
isTerminalQueryReply grammar check the native sender does, so the phone never
emits a non-grammar query reply on either opcode.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-04 02:04:45 -04:00
Jinwoo-H 8ec4d59331 Merge remote-tracking branch 'origin/main' into mobile-rearch
Resolves #16239's shared-client terminal identity against the hybrid split: the
hosted page has no native client, so identity readiness is a flag
(hostClientIdentityReady) rather than a non-null clientId, and the bridge
terminal operations keep their workspaceId/terminalId/clientId contract.
Adds getClientId to the disabled hosted client context and keeps the
mobile-web extra resource alongside main's new emoji shortcode dataset.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-04 01:32:48 -04:00
Neil 8854b5ded5 perf(ssh): coalesce concurrent git.listWorktrees reads (#18419)
* perf(ssh): coalesce concurrent pty.inspectProcess and git.listWorktrees reads

Both were the only reads in their provider class with no in-flight dedupe while
their siblings already had it. Route them through the existing
InFlightPromiseDedupe, keyed per (relay pty id, incarnation) and per repoPath,
scoped to the provider instance so two hosts never share an entry. The worktree
listing clears from invalidateGitReads(), and a signalled read keeps its own
request so one caller's abort cannot cancel its joiners' scan.

In-flight only, no TTL: the relay does answer inspectProcess from a 500ms
TTL-cached process table, but a client TTL would compound with it rather than
match it, so it is not a free win.

* perf(ssh): drop the inspectProcess half, ratchet per-read host observations

The `git.listWorktrees` dedupe ships unchanged. The `pty.inspectProcess` dedupe
is reverted: the host mints one `observationEpoch` per request and the pane
foreground reader commits it per read, so two overlapping probes sharing one
reply make the second read a stale replay and `admitRemoteForegroundEvidence`
rejects it -- a would-be `live` identity read becomes `unverifiable`. The pane
foreground tracker overlaps its own probes by design (cancel-and-reissue after
a 350 ms settle), so that path is reachable.

Adds a ratchet that fails when the dedupe returns, driving the real reader
through the real provider operations.

* fix(ssh): move the inspect ratchet to the provider it guards

The ratchet lived under src/renderer and imported
src/main/providers/ssh-pty-provider-rpc-operations, dragging the whole
main-process graph into config/tsconfig.tc.web.json (TS6307).

Split it: the request-counter ratchet moves next to the provider it
pins, and the renderer file keeps the why -- a shared host observation
degrades the second overlapping read to unverifiable -- against the real
reader with no cross-project import.

* docs(ssh): document the worktree-list coalescing contract
2026-09-03 22:03:05 -07:00
f36c03e84a fix(windows): make the install-dir ACL repair rescue the launch it runs in (#18361)
* fix(windows): repair the poisoned install-dir ACL before the window, not after

The install-dir LPAC ACL poison (electron/electron#51761) still costs every
affected machine at least one crash: the probe that detects it is
setImmediate-deferred and answers 0.9-3.0s in, while createMainWindow runs
synchronously in the same frame and its renderer dies at init 48-1373ms later.

- Persist the poison verdict the moment the probe reports it, and await the
  repair (bounded at 20s) before any window is created on a launch that already
  carries the marker.
- Do not engage the GPU safe-graphics fallback while the install-dir ACL verdict
  is poisoned or still outstanding. Safe graphics does not rescue a poisoned
  tree, and --in-process-gpu removes the GPU child, erasing the sibling-death
  evidence that identifies the shape (4 field reports landed in 'misc' this way).
- Clear the safe-graphics marker once the repair lands, so a repaired machine
  stops launching software-rendered for the rest of that build.
- Give the repair marker a bounded retry budget: it was written on failure and
  matched regardless of outcome, so one transient failure pinned a machine to
  'marker-hit' for the life of that version.

* test(windows): pin the install-dir ACL repair against the real icacls binary

* fix(windows): stop the install-DACL verdict from outliving the evidence

Adversarial review round 1. Five blocking findings, all addressed.

1. gpu-lifecycle guard had only a source grep (green with the polarity
   inverted). The stated justification -- that gpu-lifecycle's import graph
   cannot be driven in-process -- was wrong: mocking `electron` plus
   `@electron-toolkit/utils` imports it fine. Replaced with
   gpu-lifecycle-install-dir-acl-guard.test.ts, which drives the real
   handleGpuChildCrash against a stub tracker. All four cases go red when the
   guard is flipped to `if (!isInstallDirAclSuspect())`.

2. A clean probe verdict retired the on-disk marker but not the in-memory
   `poison` verdict, so a machine the probe just proved healthy kept
   suppressing the GPU safe-graphics fallback and kept the dialog accusing the
   install folder -- permanently, since a `status:'failed'` probe deliberately
   keeps the marker. A positive clean reading now latches `installDirReadClean`,
   drops the verdict, and outranks a repair result that lands after it (a
   'failed' from a repair with nothing left to fix must not re-accuse).
   'repaired' is kept: it is not a contradiction and it is what tells the user
   to reload.

3. `noteWindowsInstallDirAclProbePending()` ran on every `openMainWindow` while
   the probe is once-per-process, so every tray/second-instance reopen armed a
   15s window in which `recordGpuCrash` was never called at all -- on healthy
   machines. `probeWindowsInstallDirAcl` now reports whether THIS call
   dispatched, and only a dispatch arms the grace window.

4. The pre-window ordering guarantee was defeatable and untested.
   `focusExistingMainWindow` opens a window whenever there is none and the app
   is ready -- true for the whole 20s gate, which is exactly when a user
   double-clicks the shortcut again. Added a `canOpenWindow` seam (same
   'pending' semantics as the existing `!app.isReady()` case) wired to
   `isBlockingInstallDirAclRepairInFlight()`, plus
   windows-install-dir-acl-startup-wiring.test.ts pinning the await ahead of
   both window-creation paths and both new call sites.

5. windows-install-dir-acl-repair.win32.test.ts was absent from the pr.yml
   win32 allowlist, so it ran nowhere. Added.

Also from the non-blocking list:
- The repair no longer clears a `userConfirmed: true` safe-graphics marker;
  "keep safe graphics" is a user choice, not Orca's automatic latch.
- `repairWindowsInstallDirPackageAcl` now reports its dispatch too, so a second
  entry into the gate resolves immediately instead of eating the full 20s
  budget waiting on an `onDone` that is never coming.
- The gate is wrapped in try/catch/finally, matching the contract the probe
  documents as mandatory for anything upstream of window creation.

Rebutted, not applied:
- "Gate should be conditioned on app.isPackaged." A dev launch only carries the
  poison marker if a dev launch actually probed that tree and found the
  signature, in which case the dev renderer is dying the same way and the
  repair is exactly what is needed. The adjacent `isPackaged` check guards a
  packaged-only early-window optimisation, not a correctness boundary.
- "Fold the poison marker into the repair marker's `outcome`." They answer
  different questions with different lifetimes. The repair marker is a retry
  budget (`attempts >= 3` disables the repair for that version) and is never
  cleared; the poison marker is cleared by a successful repair and by a clean
  probe. A `'pending'` outcome written before the attempt would bump `attempts`,
  so three launches killed mid-repair would permanently disable a repair that
  never once ran icacls to completion.

* fix(windows): keep counting GPU crashes while the install-DACL verdict is pending

Adversarial review round 2. Both blocking findings addressed.

1. handleGpuChildCrash early-returned on isInstallDirAclSuspect() BEFORE
   recordGpuCrash, so the crash left no trace in the 30s rolling window. The
   suspect window is armed on every win32 non-serve launch, and the field
   bundles put it at 0.8-1.7s after main_window_created on hosts whose DACL is
   clean (matchesPoisonSignature=false) -- squarely inside the 2.1-6.2s
   bad-driver bursts this repo already pinned in
   gpu-crash-fallback-field-sessions.test.ts. A healthy machine with a failing
   driver could lose an entire coalesced burst and never engage safe graphics.

   The crash is now always recorded; only the engagement consults the verdict,
   and it waits for the verdict rather than acting on the suspicion
   (waitForInstallDirAclVerdict, resolved by the probe's onDone or by the
   existing 15s grace, whichever lands first).

   Deviation from the review's suggested shape, deliberately: awaiting the
   verdict before persisting anything reintroduces the exact race
   gpu-fallback-engagement.ts documents -- Chromium aborts the whole browser
   process on the 6th GPU crash, ~1.3s after the 3rd, which is less than the
   probe takes to answer. So the unconfirmed marker is written up front and
   withdrawn if the verdict comes back poisoned. A machine killed mid-wait
   still comes back software-rendered, and its marker is unconfirmed, which is
   the state the repair's own clear already retires.

   gpu-lifecycle-install-dir-acl-guard.test.ts now drives the real
   GpuCrashFallbackTracker and the real engagement path (the restart prompt
   firing is the signal) instead of a stub tracker, and covers the case the
   previous suite could not express: a burst that lands entirely inside the
   pending window still engages once the probe reports clean. Four reverts go
   red -- restoring the pre-record guard (2 tests), dropping the wait, dropping
   the post-wait re-check, and dropping the pre-wait marker write (2 tests).

2. The round-1 evidence block quoted commits, a test name and pass counts that
   no longer exist, and its real-icacls Windows run predated the commit that
   rewrote the gate. Re-run at this commit; counts and the live-Windows result
   are restated in the handoff rather than carried forward.

Also from the non-blocking list:
- 'marker-hit' conflated "already repaired" with "retry budget spent", because
  hasMarkerFor matches outcome === 'repaired' too. The result now carries
  alreadyRepaired, and the recovery maps that to stage 'repaired' -- so a launch
  killed between a successful repair and its marker clear no longer tells the
  user the folder needs an administrator, no longer latches
  isInstallDirAclSuspect() for the session, and does retire the poison marker.

Not applied, with reasoning:
- "clearGpuFallbackMarker narrowed to userConfirmed === false leaves the target
  population software-rendered after a repair." The summary was overstated and
  is corrected, but the narrowing stands: a userConfirmed marker now requires a
  clean DACL verdict, because the restart prompt that writes it is exactly what
  the gate above withholds while the install is a suspect. The population this
  family targets can no longer reach confirmMarker while poisoned.
- "writeInstallDirAclPoisonMarker re-stamps on a budget-exhausted machine
  forever." True, but on that machine the tree really is still poisoned and the
  gate resolves immediately ('skipped', no icacls spawn, no 20s wait), so the
  marker is telling the truth. Retiring it would be wrong; only a clean probe
  reading should.

* fix(windows): register the real-icacls spec and stop its teardown racing icacls

Two ratchets were red:
- windows-lane-tree-removal-boundary: the win32 spec's afterAll used raw
  rmSync on a tree two icacls.exe children had just rewritten DACLs on, which
  is the EPERM race removeTreeSync exists for.
- win32-test-lane-registration: the spec was in the pr.yml argv but not in
  WINDOWS_PACKAGE_TESTS, so a future diff touching only test files would not
  select package_windows and the spec would self-skip on ubuntu and report
  success.

* fix(windows): re-arm the GPU fallback latch when the install-DACL verdict withholds it

recordGpuCrash reports the threshold crossing exactly once and latches `engaged`.
handleGpuChildCrash consumes that report before consulting the DACL verdict, and
installDirAclClearsGpuFallback then discards it — so nothing could ever engage
safe graphics again in that process. A machine whose tree the repair fixes and
whose driver is genuinely broken stayed hardware-accelerated through an unbounded
crash loop, with no prompt and no marker.

disengage() releases only the one-shot latch; the crash window is untouched, so a
real driver burst is still never erased. Test is RED without the re-arm.

* fix(windows): keep the safe-graphics marker while an install-DACL repair is in flight

The gate dispatches a repair without arming the probe clock, so
waitForInstallDirAclVerdict() returns immediately and the withdrawal deleted the
marker inside Chromium's FATAL window (crash 6 lands ~1.3s after crash 3, well
inside the 20s gate). The process then died mid-repair, spent no attempt, and
relaunched hardware accelerated into the same gate — spawning the same GPU
children, FATALing again, forever.

Hold the marker while poison.stage is 'pending' so that launch comes back
software rendered and the next gate runs to completion. Still not engaged this
launch, so --in-process-gpu does not erase the sibling-death evidence. A
terminal verdict has no next step to rescue, so it still withdraws. Both new
tests are RED without the retention.

* fix(windows): stop a repaired marker outranking a fresh poison verdict

The probe reads the install DACL and finds it poisoned; `startRepair` dispatches;
`markerHitFor` sees a repair marker recording `outcome: 'repaired'` for the same
installDir+appVersion and reports `alreadyRepaired`, which the recovery module maps
to stage 'repaired'. So the launch that just proved the tree poisoned runs no icacls,
deletes the poison marker that arms the next launch's pre-window gate, clears the
suspect flag so `--in-process-gpu` can engage on a tree safe graphics cannot rescue,
and tells the user "Orca repaired the permissions."

Reachable whenever the tree is re-poisoned after one successful repair of the same
version, and whenever a repair reports success without clearing the tree — the silent
icacls no-op this module exists to document.

A DACL reading taken this launch now outranks the marker: `probeConfirmedPoisoned`
stops `outcome: 'repaired'` short-circuiting the repair. The attempt budget still
bounds it, so an unrepairable tree does not re-spawn icacls forever. The pre-window
gate does not set the flag — it acts on a marker from an earlier launch, not on
evidence of its own, so a recorded repair still outranks it there.

Also drives the GPU-fallback re-arm test through a repair that actually completes
'repaired', rather than a later clean probe, which is the route the review exercised.

* fix(windows): make the pre-window ACL gate act on the poison evidence it fired on

The gate fired on a poison marker — an earlier launch's DACL reading that nothing has
retired — but withheld `probeConfirmedPoisoned` from the repair, so a repair marker
recording an older success still short-circuited it. On the three-launch shape the gate
exists for (repair succeeds; tree is re-poisoned; the next launch's probe records the
poison but dies before writing its repair marker) the gate ran no icacls, deleted the
poison marker that arms every later gate, un-suspected the tree so --in-process-gpu could
engage, and told the user "Orca repaired the permissions." `applyInstallDirAclProbeVerdict`
then swallowed that launch's own reading behind `if (poison) return`.

Both callers of `startRepair` hold outstanding poison evidence, so the flag is now
unconditional (renamed `poisonEvidenceOutstanding`) and `marker-hit` means only that the
attempt budget is spent. The probe guard is narrowed to an in-flight gate repair: a reading
taken after the gate finished re-arms the poison marker and downgrades a claimed repair.

Also: withholding safe graphics now ends with the repair budget. A machine whose attempts
are spent while the signature persists was denied safe graphics on every launch for the
life of that appVersion — and had its marker deleted each time — including the healthy
installs the probe's flag-blind ACE match over-matches, where the driver really is broken.

Non-blocking, same lane: re-read `isQuitting` after the up-to-15s verdict wait, and skip
the recovered-launch prompt when the ACL gate retired the marker read before whenReady.

* fix(windows): stop a timed-out gate repair outranking a later poison reading

The gate's 20s budget expires while icacls runs on under its own 120s cap, so
the probe can read the tree poisoned while that repair is still in flight. Its
success claim then deleted the poison marker, un-suspected the tree and told the
user their permissions were fixed. The reading is now latched and outranks it.

* fix(windows): stop a gate repair claim pre-empting this launch's probe reading

Round-7 adversarial findings, both driven against the real modules:

- isInstallDirAclSuspect returned false the moment the pre-window gate set
  stage 'repaired', short-circuiting ahead of the probe-pending grace check.
  The GPU children die 48-1373ms after window creation while the probe
  answers 0.9-3.0s in, so an icacls that silently no-opped (exit 0, tree
  untouched) opened exactly that interval to --in-process-gpu on a
  still-poisoned tree - and a 'keep safe graphics' answer then pinned a
  userConfirmed marker no later repair may clear, with the poison marker
  already deleted so no later launch gates. The claim now stays provisional
  until this launch's probe corroborates it or the grace window lapses.

- A probe reading that disproves a 'repaired' claim re-armed the poison
  marker but never restored the unconfirmed safe-graphics marker the claim
  had cleared, so the next launch relaunched hardware-accelerated into the
  re-armed gate. The clear is now captured and handed back on disproof.

* test(windows): pin the nested and update-inherited grants against real icacls

The live spec asserted the grant landed on the root-level module file only.
It now also pins that the flagless /T pass reaches a nested file carrying
its own protected DACL (the shape app.asar.unpacked and node_modules have),
and that a file written after the repair inherits the (OI)(CI) root grant -
the stated reason that grant form exists.

* fix(windows): keep the recovered-launch prompt silent while the tree is the suspect

Round-8 fresh-eyes finding, driven against the real modules: the prompt
re-read the marker the pre-window gate may have retired, but never consulted
isInstallDirAclSuspect() - so after a FAILED gate (tree still a live suspect,
window blank behind the 10s reveal fallback, Keep as both defaultId and
cancelId) a 'keep it' answer pinned a userConfirmed marker no later repair
may clear, on the exact victim class the repair cannot help. The guard now
covers both gate outcomes; staying silent leaves the marker unconfirmed,
which a successful repair still retires.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
Co-authored-by: OrcaWin <alpha-eng@stably.ai>
2026-09-03 21:39:34 -07:00
Neil 36354f1742 perf(remote): read the repo catalog once per publish, not once per worktree (#18410)
* perf(remote): read the repo catalog once per publish, not once per worktree

`remoteWorkspace:setForConnectedTargets` costs 13 ms of main-thread time per
call at 0.48 calls/sec — 0.63% of wall on a real session, the second most
expensive IPC handler in the main process.

Almost all of it is one line. `exportRemoteWorkspaceSession` asks
`isTargetWorktree(worktreeId)` once per worktree in the session, and that
callback called `targetForWorktree(store, ...)`, which called
`store.getRepos()` — and `getRepos()` maps `hydrateRepo` over every repo row.
So publishing to one SSH target re-hydrated the whole repo catalog once per
worktree, then threw a fresh `createRepoRowExecutionHostLookup` (which itself
`filter`s the catalog per lookup) away each time.

The lookup is now built once per handler invocation and shared across targets:
the rows cannot change inside one synchronous projection, and they are the same
for every target.

On the session that surfaced this — 413 worktrees, 13 repos, 1 connected target
— that is 413 catalog hydrations (5369 `hydrateRepo` calls) per publish reduced
to 1 (13 calls). The repo normaliser reached through `hydrateRepo` was the #2
self-time function in a 30 s main-process CPU profile at 0.25%.

No user-facing trade-off: identical ownership resolution, identical exported
session, identical stale-revision handling.

* perf(remote): resolve each worktree's owning target once per publish

Follow-up on the same handler: hoisting `store.getRepos()` removed the repeat
hydration, but the ownership resolution itself was still repeated once per
connected target.

`targetForWorktree` computes a connection id from the repo catalog alone — only
the final `=== targetId` differs — so exporting to N targets ran the identical
resolution N times over every worktree key, and the projection asks the question
once per key of `tabsByWorktree`, `activeTabIdByWorktree`,
`lastVisitedAtByWorktreeId` and `defaultTerminalTabsAppliedByWorktreeId`.

Resolutions are now memoised for the life of one publish, keyed on
`(worktreeId, executionHostId)` because both participate in resolution.

Test asserts 6 worktree keys resolve 6 times across 2 targets instead of 12.

* perf(remote): skip the session and repo reads when no hydrated target is connected

Hoisting the catalog read made a zero-connected-target publish pay for a full
repo hydration it never did before. Return early instead.
2026-09-03 21:38:03 -07:00
OrcaWin f40e94d844 Revert "docs: document localization workflow" (#18571)
This reverts commit 912463c278.
2026-09-03 21:37:52 -07:00
OrcaWin 5a2bbef9d1 Revert "docs: add Ukrainian README translation" (#18570)
This reverts commit 963839aa4f.
2026-09-03 21:37:49 -07:00
Neil 0593f4e0ad perf(persistence): stop writing every worktree metadata row twice (#18451)
* perf(persistence): stop writing every worktree metadata row twice

`setWorktreeMetaForHost` assigns one object to both `worktreeMeta` and
`worktreeMetaByIdentity`, so the profile serialized every metadata row twice.
On a measured 3.64 MB install, 1,347 of 1,349 locator rows were byte-identical
to their identity twin.

The serializer now omits a `worktreeMeta` row the identity map can rebuild, and
the load path rebuilds it — reinstating the shared object reference `JSON.parse`
splits in two. A row is only omitted when exactly one alias claims the locator
and that alias names exactly one identity key, so the rebuild is a pure function
of the file with no winner selection to disagree about.

Omission rather than an in-value sentinel: a downgraded build reads a non-object
`worktreeMeta` value as corruption and deletes that locator's lineage companions
with it. An absent key is a shape every build already tolerates, and it falls
back to the untouched identity map.

* fix(persistence): keep the lineage maps when a profile file has no worktreeMeta key

The rebuild returned `parsed.worktreeMeta` untouched when it was not a plain
record, so an absent key became an explicit `worktreeMeta: undefined` that
outranked the defaults spread. `normalizeWorktreeLinkedItemMetadata` reads a
non-object `worktreeMeta` as corruption and wipes that file's
`worktreeLineageById` and `workspaceLineageByChildKey` with it, then marks the
state dirty so the wipe is persisted. Before this branch the spread supplied
`{}` and the lineage survived.

Also pins the two raw-file readers the projection made load-bearing: the
history GC recovering projected ids from the alias keys (a miss deletes shell
history a live workspace is using), and the profile-transfer read rebuilding
the omitted locator rows (a miss transfers workspaces with no metadata).

* perf(persistence): drop the identity twin, not the locator row

Reverses the projection direction: `worktreeMeta` stays complete on disk and
`worktreeMetaByIdentity[K]` is omitted instead, only when the locator row
regenerates K by construction (`wt2:<hostId>:<instanceId>`) and the two rows are
equal. That removes the format marker, the deliberate-absence list, both raw-file
reader patches and every downgrade hazard, because "alias present, identity row
absent, locator derives it" is a shape every shipped build already heals to
exactly this state.

Keeps 88% of the byte win (541 KB vs 613 KB) and the whole heap-sharing win.

* chore(persistence): keep the derivation predicate module-private

* test(persistence): pin that a file with no identity map never gains one

Answers the review ask that the projection's absent-key contract be asserted on the
bytes, not inferred: a profile whose file carries no `worktreeMetaByIdentity` must
still not have one after a load+flush.
2026-09-03 21:30:45 -07:00
Neil 6c4797ca9f perf(runtime): stop the expired-SSH-lease sweep from rescanning every tab layout (#18409)
* perf(runtime): stop the expired-SSH-lease sweep from rescanning every tab layout

The `runtime:syncWindowGraph` IPC handler is the most expensive thing the main
process does: measured on a real session it costs 20.7 ms per call at 0.71
calls/sec, which is 1.47% of wall and ~17% of all main-thread JS. 76% of that
sits in one subtree: `getHydrationTargets` -> `hasRuntimeOwnedPtyCandidate` ->
`getRecentExpiredSshLease` -> `findTerminalTabIdForLeaf`.

Three pieces of pure waste, none of which change an answer:

1. `getRecentExpiredSshLease` evaluated its cheapest and most selective filter
   LAST. `SSH_PANE_RECOVERY_GRACE_MS` is 30 s, so nearly every stored expired
   lease fails it — but only after the predicate had already resolved the
   lease's leaf to its current tab, which is the expensive part. The freshness
   and reattach-eligibility gates now run first; the predicate is otherwise
   identical and side-effect free, so the selected lease is unchanged.

2. The sweep ran once per tab. `workspaceSessionWorktreeHasRuntimeOwnedPtyCandidate`
   asked "does a recent expired lease name THIS tab" for every tab in a
   worktree, and each ask re-read and re-filtered the whole lease list. It now
   resolves the worktree's recoverable tab ids once, lazily, so a worktree whose
   first tab already owns a serve/SSH pty still never sweeps.

3. `findTerminalTabIdForLeaf` allocated a `Set` and walked a whole pane tree per
   tab to answer one leaf lookup. It now reads a leafId -> tabId index built
   once per layouts record and reused until a layout object is replaced, which
   keeps first-tab-wins ordering identical.

Measured by replaying a real 414-worktree / 801-tab / 137-lease session:
2.51 ms -> 0.27 ms per publish for this subtree, a 9.3x cut.

No user-facing trade-off: same leases selected, same tabs reported recoverable,
same SSH pane recovery affordance.

* fix(runtime): revalidate the leaf membership index on root identity

persistPtyBinding grafts a leaf by assigning `layout.root` on the SAME
layout object inside the SAME layouts record, so the layout-identity
revalidation kept serving an index blind to the grafted leaf and
findTerminalTabIdForLeaf answered `undefined` where the pre-index linear
scan answered the tab. That fed the SSH reattach fence
(restoreReattachedPtyRuntime) and the expired-lease pane recovery
resolver, both of which then fall back to the frozen lease tabId.

Membership is a pure function of the root tree and no writer mutates a
node in place, so root identity is the exact revalidation key — same
O(tabs) pointer compare, no new cap, cadence or staleness window.

* perf(runtime): resolve a leaf's tab by scan instead of a cached membership index

Fix #3 of this PR cached a leafId -> tabId map per layouts record and revalidated
it by comparing every root reference on every read. It was the only mutable
cross-call state in the change, the only piece carrying a staleness invariant,
and it had already needed one follow-up fix (1c23c544) after a layout-identity
key turned out to be blind to `persistPtyBinding`'s in-place `layout.root` graft.

The index was never what produced the measured win. After fix #1 moves the
freshness gate first, the reporter's replay never calls `findTerminalTabIdForLeaf`
at all — every stored expired lease is older than the 30 s recovery grace, so the
entire 24.9 ms -> 0.9 ms comes from fixes #1 and #2, both of which are unchanged.

`findTerminalTabIdForLeaf` is now an allocation-free scan over the existing
`layoutContainsLeafId`, which short-circuits on the first matching leaf instead of
materialising a Set per tab. Same answers, same first-tab-in-record-order
semantics, no revalidation key, nothing for a writer to invalidate.

Re-measured on the same 414-worktree / 801-tab / 137-lease replay
(process.cpuUsage deltas, median of 3; wall clock is useless on this box):

  scenario                     main     index     scan
  all leases stale (replay)   24.86     0.87     0.88 ms/publish
  one lease inside the grace  24.31     1.08     1.04 ms/publish
  all 137 inside the grace    18.04     3.71     4.31 ms/publish

The measured win is unchanged. Only the synthetic worst case — every one of 137
leases expiring inside the same 30 s window — pays for the cache's absence, and
even there the two ranges overlap because the index's own revalidation is O(tabs)
per lookup.

Removes 208 net lines. `terminal-leaf-tab-resolution.test.ts` keeps the parity
cases and adds the guard the cache needed: a subtree replaced in place after an
earlier read must be visible to the next one. That test fails against the index.

* docs(runtime): say why the leaf scan keeps Object.keys

'Allocation-free' overstated it — Object.keys does allocate one key array.
A guarded for...in trades that for a hasOwn call per tab and measures slower,
so record the reason the next reader does not re-litigate it.
2026-09-03 21:21:22 -07:00
Jinwoo Hong 4101505b6b fix(cloud): recalibrate the relay monitor's exhausted-retry freeze to a measured bar (#18569)
* fix(cloud): recalibrate the relay monitor's exhausted-retry freeze to a measured bar

The pre-drain dry-run froze at minute one on relayPostgresRetryExhausted: 0
in every run since #18521 reached the director, blocking the cell roll that
carries the same fix. #18521 made contended request-path waiters fail fast
(500 ms) instead of succeeding slowly, so exhaustion is now a steady
contention rate: 236/236 five-minute windows non-zero over 23 h; post-#18521
p50 42 / p90 147 / max 220 fleet-wide; the 2026-08-23 incident peaked at 467.
300 clears every measured healthy window and stays under the incident shape.
/v1/assign 503 share was unchanged by #18521 (13.9% vs 12.3%).

* test(cloud): pin the exhausted-retry freeze boundary at exactly 300

* docs(cloud): reword relay comments that still described the zero exhausted-retry bar
2026-09-04 00:20:28 -04:00
Neil 1c4c6b7fec perf(startup): stop queueing window creation behind the proxy apply and i18n (#18436)
* perf(startup): stop queueing window creation behind the proxy apply and i18n

Three independent, measured startup wins, all free:

1. Park the initial Chromium proxy apply on `mainProcessState` instead of
   awaiting it mid-`initializeReadyFoundation`. `setProxy` still starts at the
   identical moment; the default-session request guard (which holds, not
   cancels) is what actually fences fetchers on it, so only window creation
   stops waiting. Runtime launch still awaits it before the desktop relay and
   before every headless-serve fetcher.
2. Run `initializeMainProcessI18nAndMenu` concurrently with
   `initializeMainProcessRuntimeLaunch`. Nothing in window creation reads a
   translated string or the native menu.
3. Load `emojibase-data` in main through `createRequire` on first use instead
   of a static import, keeping 166 KB of JSON off `out/main/index.js` and its
   ~2 ms parse off every launch. The renderer keeps its eager copy unchanged.

out/main/index.js 7,210,071 -> 7,040,147 bytes. No renderer behaviour changes.

* fix(packaging): ship the emoji shortcode dataset main lazily requires

app.asar carries no node_modules, so main's bare requires resolve only out of
Resources/node_modules. emojibase-data is a devDependency and is not in the
packaged runtime allowlist, so the new createRequire in
deferred-emoji-shortcode-dataset.ts threw MODULE_NOT_FOUND in every packaged
build — breaking sanitizeWorktreeName, and with it workspace creation.

Copy the single 166 KB dataset (not the 49 MB package root) into
Resources/node_modules, and gate every createRequire'd bare specifier in
src/main against the packaged resource plan. verifyPackagedMainRuntimeDeps
cannot catch these: the bundler renames the require binding.

* test(proxy): fail CI when a main-process fetcher escapes the default-session guard

The hoist relies on installElectronProxyRequestGuard(session.defaultSession) holding every app-owned request until the persisted proxy lands. Nothing enforced that every fetcher actually lands on defaultSession. Two source-anchored rules do now: no net.fetch/net.request may name a session/partition, and every non-net .fetch( call site is counted against an allowlist.

* test(proxy): close the shorthand and chained-receiver holes in the fetch call-site audit

The audit caught `net.request({ session: x })` and `ident.fetch(`, but not the two
shapes a real regression is just as likely to take: the `{ url, session }` shorthand
that both `net.request` overloads accept, and a receiver with no bare identifier
(`session.fromPartition(...).fetch(`, `ctx.session.fetch(`). Rule 1 now also matches
the shorthand key; rule 2 scans every `.fetch(` and excludes only a literal
`net`/`globalThis`/`global` receiver. Audited counts are unchanged (2/2/1).

* fix(startup): scope the deferred emoji loader to the projects that own it

TS6307: the composite web project lists src/main/ipc/worktree-logic.ts, which
now imports the deferred dataset loader, and the shared lazy test reached into
src/main from a project that has no src/main files. Add the loader to
tsconfig.tc.web.json and move the cross-project case into a src/main test.

Also close the last two review gaps: gate the runtime-RPC startup failure
dialog (the only launch-phase translateMain reader) on a published i18n
barrier so a concurrent i18n phase cannot leave a non-English user with the
English fallback, and let the fetch call-site audit match `net.fetch (url)`.
2026-09-03 21:19:06 -07:00
Neil ef9e9f3fd9 perf(main): take the idle ownership poll off the main thread and batch pending marker probes (#18425)
* perf(main): take the idle ownership poll off the main thread and batch pending marker probes

The runtime-metadata ownership watch ran existsSync + readFileSync + JSON.parse on
the main thread every 10s for the life of the process. Move it to fs/promises with an
ENOENT catch (dropping the existsSync pre-check, a TOCTOU race anyway) and guard
overlapping ticks.

The base-directory poller's pending `.git` marker probes ran serially, costing
D x latency per tick for up to 300 ticks. Route them through the same
forEachWithConcurrency bound the full scan already uses.

* test(runtime): pin that a shutdown-straddling ownership read cannot republish

CodeRabbit flagged the async read resuming after stop(). The cleared
activeTransports guard already neutralizes it; this test pins that guard
rather than the interval teardown.
2026-09-03 21:16:04 -07:00
Neil 0d42e3fc99 perf(persistence): stop double-traversing the persisted session at load (#18458)
* perf(persistence): stop double-traversing the persisted session at load

normalizeLoadedProfileState is the largest measured startup cost that scales
with profile size, and almost all of it is zod-validating the 1.9 MB
workspaceSession blob.

Two redundant traversals removed, with no change to what is accepted:

- The salvage containers wrapped `z.record(z.string(), z.unknown())` /
  `z.array(z.unknown())` around a transform that re-validates every entry
  itself, so zod validated and copied each map and array before the real
  per-entry parse even started. The containers now apply the same guards zod
  applied (`isPlainObject` plus its enumerable-symbol-key rejection,
  `Array.isArray`) and walk the input once.
- The two recursive layout node schemas were plain unions, so every split node
  of every restored terminal and tab-group layout re-tried the leaf branch.
  They discriminate on `type`, which has the same accept and reject set.

Cold parse of a 413-worktree / 801-tab profile: 52.5 -> 45.0 ms CPU
(-14.3%, median of 25 interleaved processes). Steady state: 9.8 -> 7.1 ms.

* test(persistence): pin absence-stays-fatal for the salvaging containers

The comment on salvagingArray claimed a bare container in a z.object shape
would read a missing key as an absence unless wrapped. Not true on zod 4.5.4:
a bare transform sets neither optin nor optout, and handlePropertyResult only
swallows an absent key's issues when a field is both. Assert it instead of
documenting it, and correct the comment.

Also add the .js extension the node16 CLI project needs on the two dynamic
imports these tests added, which broke `pnpm tc`.
2026-09-03 21:15:28 -07:00
Neil 71721a6eef perf(renderer): narrow the App-root badge and terminal pty-set subscriptions (#18444)
The unread dock badge held the App root subscribed to `tabsByWorktree`, so every
agent title frame re-rendered the whole shell for an integer that had not moved.
The terminal snapshot-capability memo was keyed on the same raw maps plus
`terminalLayoutsByTabId`, so title frames and active-leaf moves rebuilt the whole
pty-id set — work its own value key then discarded.

Both now gate on the exact fields their consumer reads, compared in place.
2026-09-03 21:11:21 -07:00
Neil 6a5aa1904f perf(renderer): load the project-location and feedback dialogs on click (#18440)
* perf(renderer): load the project-location and feedback dialogs on click

Both are reachable only from an explicit click, but their chunks sat on the
renderer boot graph and were fetched and parsed on every launch. Route them
through the existing `lazy-with-retry` helper, keeping each trigger eager so the
click target still exists, and keep the mount sticky once opened so the dialog's
own close animation and repeat opens are unaffected.

Renderer boot graph 4,473,242 -> 4,424,142 bytes (-49,100 B / -47.9 KiB).

Trade-off: the first open per session now waits on a local chunk fetch —
measured at ~0.53 ms (project location) and ~0.26 ms (feedback) of read plus V8
parse/compile, warm page cache.

* test(renderer): flush the lazy set-location chunk in the ready-target test

Without the flush this case only passed because an earlier test in the file
had already resolved the shared lazy chunk; it fails under -t filtering.

* perf(renderer): warm the lazy dialog chunks on their precursor

Both deferred dialogs have a guaranteed, strictly-earlier precursor: the
composer only renders "Set location" for a needs-setup host that can take
one, and Send Feedback only exists inside an open help menu. Warm each
chunk there with a swallowed `import()` (the `preloadCommentMarkdown`
pattern) so the fetch/parse happens while the user is reading the picker
or the menu, not on the click.

Boot graph is unchanged in kind: `import()` never enters modulepreload,
so the win holds at -48,958 B (was -49,100 B before the warm; the 142 B
is the warm's own source on an already-preloaded chunk).

* test(renderer): make the composer warm guard's no-mount assertion real

The mock stubbed SetProjectLocationDialog as `() => null`, so the
"warming must not mount the dialog" assertion could never fail — the
testid it looked for was not rendered under any condition. Render a
marker unconditionally instead, matching the sidebar guard, so the
assertion actually pins the behaviour.

Verified non-vacuous: forcing the lazy element to mount eagerly now
fails with "expected <div /> to be null" rather than passing.

* fix(renderer): latch the lazy dialog mounts in state instead of during render

React Doctor's ref-mutated-during-render rule failed static analysis on both sticky-mount latches. Use the useState mount-flag idiom already in NewWorkspaceComposerModal (addProjectMounted), set from the open handler.
2026-09-03 21:11:02 -07:00
Neil 558f57de58 perf(source-control): sort branch entries before filtering, gate projections by view mode (#18426)
* perf(source-control): sort branch entries before filtering, gate projections by view mode

Two dead-work fixes in the Source Control file projection.

1. filterAndSortSourceControlPathEntries copied and re-sorted the uncapped
   branch entry list with Intl.Collator on every keystroke. Sort once on
   branchEntries, filter after: Array#filter preserves order and
   compareFileNames is a total order, so filter(sort(x)) === sort(filter(x)).

2. The tree projection was built in list mode and the list projection in tree
   mode, then discarded. Gate each memo on sourceControlViewMode and return a
   shared empty projection, matching the combined-diff file tree precedent.

* docs(source-control): drop the total-order premise from the projection sort argument

The sort-before-filter swap does not need compareFileNames to be a total
order. A stable Array#sort places each element by (comparator result,
original index) and Array#filter disturbs neither, so filter(sort(x)) ===
sort(filter(x)) for any self-consistent comparator -- which the previous
filter-then-sort already required. Restating that removes a shared-module
property (the code-unit tie-break in file-name-sort.ts) from this hook's
correctness argument instead of defending it.

Also record on the EMPTY_* singletons that the gates and both branching
consumers read one sourceControlViewMode prop in one synchronous render,
so the off-mode value cannot reach the screen, and warn against deriving
the mode from a separate store read.

New guard: matches filter-then-sort under a comparator that is not a
total order. It ties every path sharing a top-level directory over 300
entries and fails against a correct-but-unstable sort. With the duplicate
paths removed from ORDERING_FIXTURE the pre-existing equivalence test
passes under that same mutant, so this is the only test that pins
stability.

No behaviour change: counters over first render + 8 keystrokes at n=2000
are identical before and after (34685 compareFileNames calls, 0 tree
builds in list mode).

* refactor(source-control): freeze the empty branch-tree singleton

Object.freeze([]) matches the other three empty projection singletons and
the combined-diff-file-tree precedent; readonly types keep it honest.
2026-09-03 21:05:55 -07:00
Neil 6815fed6d6 perf(worktrees): converge the trash sweep instead of re-walking doomed trees (#18429)
* perf(worktrees): converge the trash sweep instead of re-walking doomed trees

`transientLockRemovalOptions()` only asked for `maxRetries` on Windows, and
`removeHostTree`'s retry ladder was gated on `process.platform === 'win32'`.
A concurrent writer is not Windows-specific: Spotlight/`mds`, a scanner, or a
live process writing under the tree surface the same EBUSY/ENOTEMPTY/EPERM on
macOS and Linux. So on POSIX the startup sweep got exactly one attempt per
entry, failed, and re-issued the same guaranteed-to-fail walk on every launch.

- Extend the retry policy to every platform. Windows keeps its error set,
  its message fallback, and its delays; the message fallback stays
  Windows-only because POSIX always sets a code.
- Persist a per-entry failure ledger in the trash root so a repeatedly
  failing entry is retried on a 15m/1h/6h ladder rather than on every launch.
  Nothing is abandoned: the ladder clamps, records are pruned when the entry
  goes, and a torn ledger fails open to a full sweep.
- Defer the sweep behind first paint, so its recursive readdir/rm no longer
  competes with window creation and worktree-catalog hydration.

* fix(worktrees): keep Node's per-level rm retries Windows-only

Node's rimraf hands every child back to the retrying entry point
(`_rmchildren` -> `rimraf`), so `maxRetries` is applied once per directory
level and compounds: a permanently-failing leaf at depth d costs roughly
`retryDelay * 36 * 9^(d-1)`. Measured on macOS against one `chflags uchg`
file at depth 2, `{recursive, force}` rejected in 1 ms while
`{maxRetries: 8, retryDelay: 150}` had not settled after 5 minutes.

Handing those options to POSIX removals turned every `removeHostTree` on a
worktree residue (`node_modules/.pnpm/...`, a dozen levels deep) into a
promise that never settles -- wedging the serialized trash-deletion queue,
hanging the sweep on its first failing entry so no backoff is ever recorded,
and leaving the unregistered-worktree removal IPC pending forever.

Keep the cross-platform retry where this PR put it -- the bounded outer
ladders that re-issue one whole `rm` against the same already-chosen path --
and restore `transientLockRemovalOptions()` to Windows-only `maxRetries`.

Also guard the deferred first-window task: off whenReady's promise chain a
synchronous throw is an uncaughtException, which the pipe-error guard
re-throws fatally.

* fix(worktrees): make host tree removal see through Electron's asar shim

The 267 stranded trash entries were not a concurrent-writer race. Electron
patches `fs` so a `*.asar` file reports `isDirectory() === true`, so Node's
recursive `rm` descends into the archive, `rmdir`s a real file, and fails the
parent with ENOTEMPTY — deterministically, on every attempt. Every worktree
that has run `pnpm install` carries a `default_app.asar`, which is why every
residue stopped at the same path.

Route `removeHostTree` through `original-fs` (Electron's unpatched `fs`, with a
`node:fs/promises` fallback outside Electron) instead of retrying a failure that
can never succeed. `removalPath`, `rmOptions` and the Windows retry ladder are
byte-identical to `main`.

Reverts the POSIX retry ladder, the `isTransientRemovalError` widening, the
sweep backoff ledger and the inverted `does not retry host removal failures
outside Windows` ratchet — none of them were fixing the actual failure.

* fix(worktrees): drop the stray orchestration test and bundle the asar guard like production

`orchestration-statement-compilation.test.ts` belongs to #18420 and was swept
into this branch by accident. It imports `./prepared-statement-cache`, which
does not exist here, so `tsc -p config/tsconfig.node.json` failed on this
branch. Removed; typecheck is clean again.

The Electron asar guard pre-externalized `original-fs` in its own Vite build,
which is not what the shipped bundle does. Mirror `isExternalMainModule` from
electron.vite.config.ts instead, so the guard also proves the production
bundler leaves `createRequire(__filename)('original-fs')` as a runtime require
— if that ever became a static import or got folded, production would silently
degrade to the shimmed `fs` while the old test kept passing.
2026-09-03 21:04:54 -07:00
Neil d247d6441b perf(startup): overlap the runtime capability refresh with the session-tabs inventory (#18460)
The startup structured-session restore chained `runtime:getStatus` before
`session.tabs.listAll`, but the capability value is discarded at that call site —
it only seeds the module cache later launch flows read, and the inventory fetch
never reads it. On a profile with 413 worktrees / 801 tabs that serial leg cost a
median 109 ms of the did-finish-load -> renderer-startup-hydration-done window.

Issue both calls concurrently. `Promise.all` still resolves only after both
settle, so the capability cache is populated no later than before.
2026-09-03 21:04:04 -07:00
Neil 949c9d3353 perf(worktrees): classify each worktree once, defer the SSH meta index, drop the conflict-path probe (#18433)
* perf(worktrees): classify each worktree once, defer the SSH meta index, unserialise conflict probes

Three redundancies on the worktree-catalog and git-status read paths:

- buildDetectedGitWorktrees ran mergeWorktree + toDetectedWorktree twice for
  every visible row. Discovery backfill returns the same meta object when it
  wrote nothing, and both builders are pure over it, so skip the second pass on
  identity.
- The SSH worktree-meta index parsed every worktree id on the host, then threw
  it away whenever the provider was connected. Build it lazily, memoised.
- Unmerged `u` records were resolved one fs.access at a time. Resolve the prefix
  the cap can reach with 8-way concurrency, keyed by record index so Git's
  output order and error precedence are unchanged.

* perf(git): read the porcelain worktree mode instead of probing conflicted paths

Every porcelain-v2 `u` record already carries `mW`, the working-tree mode Git
stat'ed for that row: `000000` means the conflicted path is absent. Reading it
replaces the per-conflict `fs.access`, so the bounded-concurrency resolver, its
`= 8` cap, and the order/error-precedence invariant are unnecessary rather than
cheaper. `access()` stays only as a fallback for a malformed `mW`, so
`parseUnmergedEntry` keeps its signature and neither status-read.ts nor the
relay loop changes.

Also corrects two fixtures that encoded `mW=100644` for a file that does not
exist, which real Git never emits.

* fix(test): import the conflict parser statically so the CJS cli project compiles
2026-09-03 21:02:11 -07:00
Neil c11c6878c1 perf(persistence): stop dead SSH leases pinning metadata, retire unreachable tombstones (#18430)
Two unbounded-growth fixes in the persisted profile, which is re-serialised in
full on every save.

`collectPersistedWorkspaceOwners` registered every SSH lease's worktreeId as a
live persisted owner with no state filter, so a route-retired lease — the
operator-close `terminated` tombstone, or an `expired` row already marked
`supersededBy`/`relayIdRecycled` — pinned its worktree's metadata row
permanently. The prune gate's own doc names that failure: "Rows pinned by a
persisted session are never removable, so the repetition cannot even make
progress." Reuses `sshRemotePtyLeaseAllowsReattach`, the predicate that already
decides which leases still name a route.

`sshRemotePtyLeases` had no pruning path at all: removal happens in three
explicit places, none age- or state-based, so `terminated` rows accumulated
forever (137 rows / 54 KB on the reported profile, ~38/day from one target).
Marking a lease `terminated` scrubs its pane bindings in the same write, so once
no persisted binding names the id the row routes nothing — reattach, pane
recovery, the orphan sweep, `ssh:reset` and `ssh:terminateSessions` all behave
identically on an absent row. Delete it then, gated on that reachability check
because a lease freezes its tabId and the tab-qualified scrub cannot reach a
pane that was detached into a new tab.

`expired` rows are deliberately untouched, superseded ones included:
`sweepOrphanedRelayPtys` reads those ids as its leave-alone list, so dropping one
would authorize stopping a remote shell that supersession left running on
purpose (docs/reference/ssh-execution-boundary.md).
2026-09-03 20:59:53 -07:00
Neil ab32c2c0c5 perf(startup): stop the persistence milestone from timing its own details closure (#18439)
`logPersistenceStartupMilestone` resolved the lazy `details` closure before
reading `performance.now()`, so the 1.6 MB `JSON.stringify` that
`persistence-load-done` uses to report `workspaceSessionBytes` was billed to the
milestone it measures. Snapshot `t` first.

Diagnostics output is unchanged; only the recorded timestamp moves.
2026-09-03 20:57:57 -07:00
Neil 6415b1dc22 perf(images): probe raster headers instead of decoding whole payloads, memoize repo icon validation (#18421)
* perf(images): measure raster headers from a probe and memoize repo icon validation

`getRepos()` re-sanitizes every repo on every call, and an uploaded/file repo
icon costs a full base64 decode of its data URI each time. Three fixes:

- `writeQuartet` destructured a mutable array, which sends V8 through the
  iterator protocol once per four input characters; index reads plus a length
  counter produce identical bytes.
- `decodeBase64Prefix` decoded the whole payload despite only the first bytes
  being needed. `exceedsRasterImagePreviewLimits` now probes 64 bytes and
  widens x16 until the header measures, and only re-runs the original
  full-payload decode when the verdict would suppress a preview.
- `sanitizeRepoIcon`'s src validation is memoized per icon source with a
  bounded FIFO map, reusing the `memoizeTitleClassification` idiom (now a
  shared `memoizeByStringKey`).

* perf(images): key icon-validation memo on the persisted icon object

Replaces the per-source 64-entry FIFO string-key memo with a WeakMap keyed on
the persisted repoIcon object that hydrateRepo already receives, storing
{src, source, supported} and re-checking both fields on a hit.

Retention becomes zero by construction (entries die with state.repos[i].repoIcon),
so there is no cap to evict live icons and no dead icon strings held after a repo
or icon is replaced. The identity re-check makes an in-place mutation unable to
serve a stale verdict. Drops bounded-string-key-memo.ts and reverts the collateral
terminal-title-classification-memo refactor.
2026-09-03 20:56:29 -07:00
Neil 97e5eb8886 perf(paths): guard the no-op regex passes on the path-comparison hot path (#18418)
* perf(paths): guard the no-op regex passes and hoist the loop-invariant root

`normalizeRuntimePathForComparison` ran two whole-string regex passes on every
call — `/\/+/g` and `/\/+$/` — that cannot change a path with no doubled slash
and no trailing slash, which is nearly every path. `parseWslUncPath` likewise
folded backslashes and ran an anchored UNC regex over every POSIX path.

Substring/char-code probes skip all of them, and a
`createRelativePathInsideRootResolver` factory (mirroring the existing
`createNormalizedPathInsideOrEqualMatcher`) folds a fan-out's root once instead
of once per candidate. Outputs are unchanged; a seeded 200k-path differential
fuzz against a pre-guard copy proves it.

* perf(paths): drop the root hoist, land the guards alone

The three in-module guards are the whole win: 5000-op batches, CPU time,
median of 9 --- normalize 541 -> 239 ns/op, relativePathInsideRoot
1778 -> 899, isPathInsideOrEqual 1076 -> 572, parseWslUncPath 57 -> 14.

The loop-invariant root hoist added 176 ns/op on top of that (899 -> 723)
at 7 hand-picked call sites, and cost a new exported factory whose
input contract is the opposite of the one next to it, plus a function
substitution in worktree/ownership.ts. Not worth 0.9 ms per storm.

Prod diff: 2 files. New ratchet pins the single-factory surface.

* docs(paths): point the fixture header at the real guards test
2026-09-03 20:54:59 -07:00
Neil 7ed86a98ae perf(ipc): index worktree owners instead of rescanning the repo list per lookup (#18416)
Two hot lookups rescanned a whole table once per repo.

`getLocalRepoForRegisteredWorktree` (59 IPC call sites, including Quick Open
keystrokes and every File Explorer expand) walked the entire worktree-meta table
once per repo. One pass now collects the owning repo ids, built lazily so a repo
whose own path matches still never touches the table.

`createRepoRowExecutionHostLookup` re-filtered the repo array on every `byId` /
`byHost` call. Rows are grouped into a Map once at construction, preserving
repo-list order so `rows[0]` still picks the same owner.
2026-09-03 20:41:17 -07:00
Neil 63aee7f1ee perf(terminals): spend one inspection start on a whole cadence round (#18438)
The inspection rate limiter counted panes when it should have counted host
observations. `MAX_INSPECTION_STARTS_PER_SECOND = 8` is global, and it was
spent one pane at a time, so N due panes meant an effective per-pane period of
max(tier, N/8 seconds) — ~37.5s at 300 panes for a pane the code polls at
750ms. Agent-completion latency degraded monotonically as panes were added.

Every local pane's inspection resolves out of the same TTL-and-in-flight-
deduped process-table capture, so a whole round of them is one host
observation. The queue now drains all shared-observation tasks as one round on
one start, launched in a single tick. Remote panes each cost their own
execution-host round trip and stay admitted one at a time. Both the budget and
the cadence tiers are numerically unchanged.

Disposed tasks are also compacted out in one pass instead of a splice per drop,
so the per-round predicate cost is linear rather than quadratic at pane scale.

No IPC, preload, wire, or main-process change: each pane keeps its existing
per-pane `pty:inspectProcess` invoke.
2026-09-03 20:37:44 -07:00
Neil 07e50e9513 perf(terminal): scan only new tail lines for the wait-blocked sentinel (#18437)
* perf(terminal): scan only new tail lines for the wait-blocked sentinel

The wait-blocked scan must prove a signal is ABSENT, so it could not early-exit and re-tested all 2000 retained lines with a 13-alternative regex on every scan (20/s per streaming PTY) even though only ~20 lines were new. Index the matching line indices per tail-array identity and carry them across appends, testing only the lines each append produced.

Also carries the retained character total and the redraw prefix's right-trimmed state across appends, so a saturated tail is no longer re-summed and re-scanned per chunk.

* perf(terminal): build the carried tail window and its match index from one constructor
2026-09-03 20:34:42 -07:00
Neil a711cb8b60 perf(renderer): gate the tab strip's worktree subscriptions and fix the orchestration batch's self-invalidating cache (#18428)
* perf(renderer): gate the tab strip's worktree subscriptions and stop the orchestration batch invalidating itself

Two store-subscription hot paths.

The tab strip subscribed to projects/repos/worktreesByRepo for the Windows shell
menu's local project runtime, which is never built unless that menu is on. On
macOS/Linux every worktree write therefore re-rendered and re-committed every
mounted tab strip. Gate the three on the condition that already gates their only
consumer.

The runtime-orchestration batch keyed its cache on agentStatusByPaneKey identity,
which `agentStatus:set` replaces by definition, so it missed 100% of the time on
the only event that calls it. Key on the paneKey -> worktreeId pairs the batch
actually reads instead, and hang the requested-id array off the existing
activeWorkspaces memo so the O(worktrees) prologue stops running per event.

* refactor(renderer): make the orchestration batch's cache key its build's only inputs

buildRuntimeBatch no longer receives agentStatusByPaneKey/retainedAgentsByPaneKey.
It takes a RuntimeBatchInputs record whose paneWorktreeIds projection is its whole
view of those maps, and that same record is the cache key, so the key cannot drift
from the read set. Adds a guard asserting one read per orchestrated pane per map.

* refactor(renderer): move the orchestration projection key onto the shared index

The batch builder and `worktree-agent-orchestration-index.ts` were near-duplicate
implementations of the same attribution walk, and both had the self-invalidating
`liveSource === agentStatusByPaneKey` gate. Fixing only the batch left the index —
which every mounted WorktreeCard hits on every `agentStatus:set` — still rebuilding
per publication.

Put `paneWorktreeIds` on the index instead and reduce the batch to a `.get`-compatible
view of it. That deletes the whole `requestedWorktreeIds` apparatus the batch fix needed
(the `worktreeIds` memo threading, the optional `selectDashboardOrchestration` param, the
`uniqueWorktreeIdsByInput` WeakMap and its no-mutation contract, `getRequestedTabMembership`),
leaves one builder guarded by the index's randomized oracle test, and extends the fix to
the sidebar.

The projection is memoised on the live/retained map identities so it is computed once per
publication rather than once per card, and a successful ordered compare adopts the new array
so the remaining cards compare by identity.
2026-09-03 20:25:22 -07:00
Neil 34222e0137 perf(orchestration): project explicit columns so the graph publish stops recompiling SQL (#18420)
* perf(orchestration): cache the prepared statements the graph publish recompiles

SyncDatabase refuses to cache any `SELECT *` — node:sqlite can build the first
row after a schema change from stale column names — so every wildcard read in
the orchestration DB recompiles its SQL on each call. The graph publish runs
that fan-out once per pane, ~0.7 times a second, forever.

Add a per-connection prepared-statement cache scoped to the orchestration DB,
whose schema is frozen in the constructor (createTables/migrate/trigger) and
whose resets are DELETE-only, and route the buildByPaneKey -> getForHandle ->
getRecent path through it. 5 publishes over 2 panes: 30 compilations -> 2.

* perf(orchestration): project explicit columns so the existing cache covers the hot path

Replaces the branch's second statement cache. The six graph-publish reads were
uncacheable only because they were spelled `SELECT *` / `SELECT t.*`, which
SyncDatabase refuses to cache (node:sqlite can build the first row after a schema
change from stale column names). Spelling the projection out from type-checked
column tuples makes them cacheable by the SyncDatabase LRU that is already merged,
already bounded, and already clears on DDL — so the WeakMap and its documented
cross-connection ALTER hazard both go away.

Drift is caught at build time: `satisfies readonly (keyof Row)[]` plus an
`Exclude<keyof Row, Cols[number]> extends never` assertion pins list vs type at tsc,
and a PRAGMA table_info test against a freshly migrated OrchestrationDb pins list
vs schema.

Same win, verified: 6 compilations per publish -> 2 total then 0, identical to the
WeakMap branch; 92/96/91 us CPU per 2-pane publish before, 11-12 us after on both.
2026-09-03 20:10:01 -07:00
Jinwoo-H 4d4e361f52 Revert "fix(mobile): surface a dropped hosted tab subscription instead of spinning forever"
This reverts commit ca7c7a281d. The root cause of the dropped subscription is
still unknown; surfacing the drop only trades one symptom for another.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-03 23:07:21 -04:00
Jinwoo-H ca7c7a281d fix(mobile): surface a dropped hosted tab subscription instead of spinning forever
A rejected session snapshot (invalid shape, oversize, or a failed post) cancelled
the shell-side subscription silently, the bridge client discarded the late error
because the subscribe request had already resolved, and the page kept "Loading
tabs" with no error and no retry. The shell now posts an error on the subscribe
request id, the client routes it to the subscription's onError (which already
falls back to polling), and the session screen shows Retry after two consecutive
failures. Reuses the existing response opcode, so mixed versions degrade to
today's behaviour.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-03 22:52:21 -04:00
Jinwoo-H 979cf30247 feat(mobile): opt-in inspectable Android release build for hosted WebView devtools
`-PorcaInspectableRelease=true` marks the release variant debuggable through an
Expo config plugin and flips a build config field the shell's inspection policy
reads. The OS debuggable flag stays a hard requirement, so a shipped production
APK can never be inspected regardless of the Gradle property. Default builds are
byte-for-byte unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-03 22:52:21 -04:00
Jinwoo-H 738132bf9a fix(mobile): address hosted markdown tabs by id so external files load
Tabs opened from outside the worktree carry an absolute path; the hosted
snapshot stripped it and the read payload was rejected before any request.
The shell now resolves the file from its own session.tabs.list by tab id,
carries isDirty through, clamps oversized reads to read-only instead of
failing, and the retry state names the error code.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-03 21:35:33 -04:00
Jinwoo-H 54de839211 fix(mobile): no empty repo placeholders while workspaces are still loading
Placeholder sections for repos with no rows were built before the paged
worktree list landed, so every repo flashed as a count-0 header. Gate them
on rows loaded, and let the hosted host state reuse the in-memory cache so
returning from a session keeps the last complete list.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-03 21:35:33 -04:00
Jinwoo-H 5bc7e95040 fix(mobile): keep the resume route across a hosted package swap
A desktop update swaps the hosted page under the user; the shell reset the
remembered route to the workspace list on every session id, so a session
only came back through cold resume after the list mounted and fetched.
Scope the memory to the host so init replays the session route directly.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-03 21:35:33 -04:00
Jinwoo-H fc1e78f862 fix(mobile): open network diagnostics through the shell from the hosted host list
The hosted page pushed /connection-log page-locally; the page has no such
route and cannot show the shell's transport log anyway. Hand the shell a
connectionLog native route instead, and fence the class with a reachability
test over the hosted module graph. Re-pair goes through repairPairing too.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-03 21:35:33 -04:00
Shahar MorandMerge Sim 7106101ed2 fix(mobile): restore terminal input when reopening worktrees (#16239)
* fix(mobile): restore terminal input when reopening worktrees

* test(mobile): update session parity facts

* refactor(mobile): split host client hooks

* chore: restore localization formatter scope

* fix(mobile): retain RpcClient type import

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 18:32:53 -07:00
Jinwoo Hong 11aace8dec fix(relay): reject malformed percent-escapes on upgrade instead of throwing (#18547)
decodeURIComponent on the /v1/connect/ and /v1/host/data/ path segments threw
URIError out of the http 'upgrade' listener, which is uncaught and kills the
relay process. Any client that sends GET /v1/connect/% could take down a cell
(and every connection on it) or a director instance. Pre-existing since the
splice landed (orca-cloud #20); not introduced by the import.

A malformed escape now takes the existing 4xx reject branch. The blackbox test
sends three malformed connect targets and one host-data target to the real
server and asserts no uncaughtException fires and a well-formed upgrade still
gets 101 afterwards; reverting either site fails it.
2026-09-03 21:31:53 -04:00
Neil 8c1a28d39c fix(i18n): repair French locale drift breaking static analysis (#18550)
The French UI locale landed with two catalog drifts that fail `static
analysis` on every PR in the repo:

- `fr.json` carried 15 keys absent from `en.json` (and from every other
  locale), so `verify:localization-catalog` rejected it. They are stale
  entries generated against an older `en.json` snapshot; none is
  referenced anywhere in the source.
- `settings.appearance.language.french` had no call site supplying a
  literal default, which promotes it to a boot-bundle-required entry
  that `en-runtime-required.json` does not ship, so
  `verify:localization-runtime-catalog` rejected it.

Registering the key in settings search alongside its siblings fixes the
runtime-catalog failure at its source and closes the real gap the drift
exposed: French was the only supported language not findable in settings
search.

`en-runtime-required.json` is deliberately untouched — the sync script
regenerates it wholesale and would drop 925 entries the check itself
documents as harmless.
2026-09-03 18:24:43 -07:00
Jinwoo Hong dd9eaa9585 fix(cloud): retry the committed-winner collision codes in relay schema startup (#18553)
* fix(cloud): retry the committed-winner collision codes in relay schema startup

`CREATE TABLE IF NOT EXISTS` only checks the name before the catalog inserts, so
the loser of a concurrent CREATE fails in one of two ways depending on timing:
on the catalog unique index (23505, which the startup retry already handled) or,
when the winner has committed by the time the loser reaches TypeCreate /
heap_create_with_catalog, on the name check those routines repeat (42710
duplicate type, 42P07 duplicate relation). The predicate treated the latter as
fatal, so a director could fail startup on a table it was about to find present.

This is what turned `postgres-schema-concurrency-postgres.test.ts` red on main
and on every relay PR (CI's shared runner loses the race more often than a dev
box): a throwaway diagnostic run in CI reported 42710 from TypeCreate and 42P07
from heap_create_with_catalog as the only rejection reasons.

Treat 42710/42P07 as retryable for `CREATE TABLE IF NOT EXISTS` and 42P07 for
`CREATE [UNIQUE] INDEX IF NOT EXISTS`; every other statement shape still fails
fast. The concurrency test now runs ten rounds and reports the loser's SQLSTATE
instead of a bare boolean.

* chore(cloud): allowlist the RFC 6455 example Sec-WebSocket-Key for upgrade tests

Cloud Verify's Secret scan runs gitleaks over --all refs, so the raw-socket
upgrade test on fix/relay-upgrade-malformed-uri (#18547) trips every cloud PR's
scan until its allowlist reaches main. Land the allowlist here first.
2026-09-03 21:20:27 -04:00
Brennan BensonandMerge Sim 90780acb85 refactor(agents): one pane-identity resolver behind six thin adapters (tranche 0) (#18243)
* feat(agents): pane-identity canonical adapter, comparison telemetry, inventory ratchet phase 1

* fix(agents): preserve canonical coverage provenance

* refactor(agents): unify pane identity adapters for tranche 0

* fix(agents): keep title resolver cache-free after rebase

* Fix ladder tranche zero review findings

* fix(agents): restore title classifier memoization

* fix(agents): fence unknown canonical evidence sources

* docs: drop the ladder plan and decision table from the PR

Design docs stay out of the shipped tree; the code carries its own comments
and the decision table lives in the test fixture.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 18:07:56 -07:00
Neil 8463dcb7b9 fix(terminal): make wrapped-line search rewind iterative and bound its scans (#18402)
Patches @xterm/addon-search so one very long un-newlined line no longer overflows the stack, freezes the renderer, or goes unsearched. Submitted upstream as xtermjs/xterm.js#6149 (issue #6148); drop the patch once a release ships it. See the PR for measurements and the differential fuzz.
2026-09-03 17:59:16 -07:00
Ihor 963839aa4f docs: add Ukrainian README translation
iho <4000375+iho@users.noreply.github.com>
2026-09-03 17:33:02 -07:00
foXaCe 49d6d35b16 feat(i18n): add French UI locale
foXaCe <290678+foXaCe@users.noreply.github.com>
2026-09-03 17:32:59 -07:00
hwantage 48cb575db5 feat(i18n): localize Orca Account settings and navigation to Korean
hwantage <82494320+hwantage@users.noreply.github.com>
2026-09-03 17:32:55 -07:00
Trevin Chow 912463c278 docs: document localization workflow
tmchow <517103+tmchow@users.noreply.github.com>
2026-09-03 17:32:51 -07:00