Commit Graph
9722 Commits
Author SHA1 Message Date
Merge Sim ecb97c8f5a docs(native-chat): clarify tested implementation head 2026-08-31 17:44:33 -07:00
Merge Sim 26f500c1ea docs(native-chat): pin final readiness head 2026-08-31 17:44:33 -07:00
Merge Sim c510811c6a docs(native-chat): pin readiness report to final head 2026-08-31 17:44:33 -07:00
Merge Sim e82b8cc64a docs(native-chat): record successful Electron proof 2026-08-31 17:44:33 -07:00
Merge Sim aa24d48f78 docs(native-chat): refresh exact-head readiness report 2026-08-31 17:44:33 -07:00
Merge Sim 1dc0a652b3 fix(native-chat): update structured chat capability copy 2026-08-31 17:44:33 -07:00
Merge Sim 4a9fca0ddd fix(native-chat): harden Claude structured launches 2026-08-31 17:44:33 -07:00
Merge Sim 68722cf395 fix(native-chat): prioritize structured routing before web tabs 2026-08-31 17:44:32 -07:00
Merge Sim 667184c951 test(native-chat): update Windows capability expectations 2026-08-31 17:44:32 -07:00
Merge Sim 8cf9bc92c0 fix(native-chat): recover unexpected Claude exits 2026-08-31 17:44:32 -07:00
Merge Sim 077fae0c8f refactor(native-chat): centralize structured session target resolution 2026-08-31 17:44:32 -07:00
Merge Sim 15d456a0e1 fix(native-chat): gate Windows capability on process identity 2026-08-31 17:44:32 -07:00
Merge Sim aabdaa55e7 fix(remote): route structured chat lifecycle to owning host 2026-08-31 17:44:32 -07:00
Merge Sim 4f8ba8ded4 fix(native-chat): route structured sessions by execution host 2026-08-31 17:44:31 -07:00
Merge Sim b7201a8051 fix(native-chat): harden WSL session routing 2026-08-31 17:44:31 -07:00
Merge Sim 584c8adc18 docs: refresh readiness evidence 2026-08-31 17:44:31 -07:00
Merge Sim 4f3aec0309 docs: normalize readiness report formatting 2026-08-31 17:44:31 -07:00
Merge Sim 6020f1ce44 docs: add structured chat readiness report 2026-08-31 17:44:31 -07:00
Merge Sim 8ac66d8714 feat(native-chat): enable Windows structured sessions 2026-08-31 17:44:31 -07:00
Merge Sim 9249df5a61 docs: record structured native-chat platform split 2026-08-31 17:44:31 -07:00
Merge Sim e95fb2cc86 feat(remote): route structured native chat to paired hosts 2026-08-31 17:44:31 -07:00
Merge Sim 42947128ec test(ssh): preserve structured session host verdict boundary 2026-08-31 17:44:30 -07:00
Merge Sim e4038f6d3d feat(native-chat): support structured Codex sessions in WSL 2026-08-31 17:44:30 -07:00
Merge Sim be6891a08e feat(native-chat): support structured Claude sessions 2026-08-31 17:44:30 -07:00
Merge Sim f289fee894 feat(native-chat): add Claude structured adapter 2026-08-31 17:44:30 -07:00
Jinwoo Hong 26031ca317 fix(browser): scroll oversized viewport presets (#17569)
* fix(browser): scroll oversized viewport presets

* fix(browser): preserve guest wheel scrolling at viewport edges

* fix(browser): keep viewport scroll state synchronized

* test: assert partial viewport wheel forwarding
2026-08-31 20:39:58 -04:00
Brennan BensonandMerge Sim 1a47b9ee85 fix(remote): distinguish SSH transport from runtime availability (#17710)
* fix(remote): distinguish transport from runtime availability

* fix(remote): preserve transport diagnostics for unavailable runtime

* fix(remote): propagate transport diagnostics to host setups

* fix(remote): keep unavailable runtimes out of ready setups

* fix(remote): preserve unavailable runtime state in settings

* fix(remote): preserve reconnecting runtime state

* fix(remote): guard stale settings connectivity

* fix(remote): preserve diagnostics after main merge

* fix(i18n): preserve translations during runtime status merge

* fix(remote): refresh settings row health from store

* fix(remote): refresh settings row health from store

* fix(remote): clear diagnostics generations in tests

* fix(settings): refresh runtime availability summary

* refactor(runtime): split status slice types

* refactor(runtime): reuse status app state type

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 17:37:53 -07:00
Jinjing 45c4823109 Format documentation with consistent line wrapping and table alignment (#17765)
Standardize MDX files across docs with:
- Remove trailing semicolons from import statements
- Wrap long lines and multi-line component props for readability
- Align Markdown table column separators
- Normalize text and JSX formatting for consistency
2026-08-31 17:24:35 -07:00
Neil fa0180dc61 perf(renderer): avoid combined-diff tree rebuilds during progressive loads (#17643)
* perf(renderer): avoid combined-diff tree rebuilds during progressive loads

* fix(renderer): preserve collapsed combined-diff tree boundaries

* perf(renderer): skip unfiltered combined-diff flatten when hiding viewed files

* fix(renderer): keep reordered viewed keys in the combined-diff delta

The incremental viewedSectionKeys delta walked indices issuing a delete
then an add, so a key added at index i and deleted as the previous key at
a later index was silently dropped. Fall back to a full recompute when any
index's key differs; the progressive-load fast path (stable keys, flipping
loading state) is unchanged.
2026-08-31 17:14:34 -07:00
Jinjing f546f53a4e docs: update Android APK link to 0.0.47 (#17764)
Update the README download links to the latest mobile Android release.
2026-08-31 17:03:39 -07:00
Neil c558d7e083 Activate terminal splits before inherited CWD resolution (#17601)
* perf(terminal): activate splits before cwd resolution

* test(terminal): prove split focus before cwd publish

* fix(terminal): release stale split cwd fence

* test(terminal): add visible split activation latency benchmark

* docs(reliability): clarify split benchmark provenance

* fix: preserve deferred split handoffs across remounts

* fix: fence late deferred split closes

* docs(reliability): record exact split benchmark runs

* test(reliability): fail benchmark on artifact write errors

* test(reliability): attribute split activation phases

* docs(reliability): record schema-v2 split benchmark

* refactor(terminal): collapse duplicated split-handoff and write-queue paths

- Drop the discardDeferredSplitPaneHandoff alias for its identical clear twin.
- Fold the deferred-cwd resolve/reject settle handlers into one applier.
- Extract settlePaneCwdDeferredSpawn for the repeated read-clear-write pattern.
- Share one head-index FIFO primitive between the ordinary and reply queues.

* fix(terminal): stop retaining a promise reaction per acknowledged write

Racing every accepted write against one queue-lifetime cancel promise kept a
reaction record alive until that promise settled: 200k acknowledged writes
retained 88.6MB, now 0.1MB. Give each in-flight write its own cancel, and
split the shared FIFO primitive into its own module.

Also sanitize the split-latency benchmark report at its single serialization
point so shared artifacts no longer carry the machine-local repo path or
unbounded cleanup error text.

* fix(terminal): settle deferred split input when the spawn is abandoned

An abandoned deferred spawn returns before transport.connect(), so nothing
drained the pre-connect buffer: sendInputAccepted's promise never settled and
a paste into that pane hung forever. Clear the buffer on the abandon fence.

Also re-derive the pre-connect retention cap from the clipboard-paste ceiling
rather than the 16MB single-write ceiling; it is held twice per pane across up
to 64 deferred splits, so 5.59M code units guarded the wrong thing.

* fix(terminal): release the deferred cwd fence on a rejected reattach

A daemon createOrAttach can turn an apparent fresh spawn into a reattach; when
that reattach is refused the spawn ends with deferredSplitSpawn/pendingCwd
still set, permanently arming the pre-bind detach refusal. The release no-ops
when a PTY did bind, so it only fires where the fence would otherwise leak.

The stale-generation return above is deliberately left alone: a newer connect
already owns the pane there, and the fence is not generation-scoped.
2026-08-31 16:45:36 -07:00
Neil 4ac8a8912c fix(startup): install the app environment with the userData decision (#17755)
src/main/index.ts decided where userData lives at module scope, then installed the
AppEnvironment port ~180 lines later inside the single-instance-lock block. Every
statement in that gap was a latent failure: a path resolve there either threw
'AppEnvironment not initialized' and killed the process, or — with the accessor
installed but the decision not yet run — would have memoized the pre-override
directory in getCanonicalUserDataPath() for the whole session.

The first outcome shipped. #16761/#16698/#17509 were one statement landing in that
gap and killing every macOS `orca serve` across 1.4.190-1.4.192; #16762 moved that
call but left the gap.

Install the port and capture the canonical path immediately after the two calls
that decide them, so the window is zero rather than small. Both are inert at this
point — ElectronAppEnvironment holds no state and calls `app` lazily per accessor,
and initDataPath only joins strings — so nothing that depended on the old position
moves with them. The secret store stays where its pre-ready Keychain note applies.

The throw is kept and still covers the case it should: resolving a path before the
decision has run.

Guarded by a source-level assertion that the decision, the install and the capture
stay adjacent.

Fixes #17750
2026-08-31 16:40:29 -07:00
Brennan BensonandMerge Sim 59facfb71e Show live tool progress in native chat (#17597)
* Show live tool progress in native chat

* fix(native-chat): scope live tool indicator to current turn

* fix(native-chat): settle orphaned live tool rows

* fix(native-chat): keep live tools running without lifecycle metadata

* fix(native-chat): keep working status stable during streaming

* fix(native-chat): anchor turn status below prompts

* fix(native-chat): preserve turn status and legacy tool activity

* fix(native-chat): limit turn status UI to structured Codex

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 16:31:30 -07:00
Neil 8ac1c6e2ac perf(git): bound ref and worktree scans (#17655)
* perf(git): bound ref and worktree scans

* fix(repo-search): clamp oversized ref limits

* fix(worktree): keep strict worktree listing unshared

The shared-scan re-export flipped every `listWorktreesStrict` caller from an
isolated subprocess to the coalesced scan. `git worktree prune` in the removal
recovery path does not bump the scan generation, so a post-prune verification
could join a pre-prune scan, see the stale row, and report a successful removal
as a stale registration. The same gap defeats the post-archive-hook rechecks
that exist to catch an external Git client locking the row.

Restore the unshared export and make coalescing opt-in via
`listWorktreesSharedStrict`, which existing callers already use deliberately.

* fix(git): separate a proven absent ref from a failed probe

`show-ref --verify --quiet` exits 1 for a missing ref, but so does `wsl.exe`
when its own launch fails, so reading any exit 1 as absence collapsed
`unverifiable` into `exited`. A genuine miss prints nothing while a wrapper
failure always explains itself, so require empty stderr alongside the exit
code; a runner that reports no stderr at all keeps its exit-code contract.

That same signal removes a spawn regression: `show-ref` is a direct-git read
under WSL, and the runner retried any numeric exit through the user's
interactive login shell. The replaced `for-each-ref` exited 0 on a miss, so
absence never retried; every absent probe now would. Treat a quiet exit 1 as
Git control flow and skip the fallback.

Also narrow the hosted-review suffix fallback: the replaced
`refs/remotes/*/<base>` could not cross a slash, but `show-ref -- <base>`
matches at any depth, so `origin/feature/main` answered a query for `main`
and submitted a review against a base the provider rejects.

Refresh the real-binary compatibility contract to the shipped excludes, and
assert exact probe concurrency rather than an upper bound so a regression to
serial probing fails.
2026-08-31 16:27:28 -07:00
Neil a5be10815c perf(terminal): drop the headless snapshot cache, keep the spawn lock (#17752)
Removes the snapshot memoization added in #17667 and everything that served
it: the epoch, markMutated/markWritten, the no-op resize gate, and the
HeadlessSnapshotCache module. The shared/exclusive spawn lock from the same PR
stays — it is the half that carries the measured win.

Why: a controlled A/B on merged main could not show the cache paying for its
memory. Same worktree, same sessions, reattach stable_adoption per session:

  cache + resize gate (as merged)   16 / 67 / 78 / 87 ms
  cache off, gate on                17 / 55 / 68 / 80 ms
  cache off, gate off               13 / 62 / 73 / 83 ms

Indistinguishable. Cold 4-tab activation was likewise unchanged (217-296ms
without the cache vs 226-309ms with), which is expected — a first attach is
always a miss.

The reason it under-delivers is a design fact the original PR missed: attach
does not request the full buffer. terminal-host-session-create.ts passes
resolveDaemonSessionScrollbackRows() — a deliberate 1000-row live window,
capped because unbounded retention once OOM-killed a host. Serializing 1000
rows is cheap, so there was little for a cache to save on that path.

Against that, the cache retained up to MAX_CACHED_SNAPSHOT_BYTES (4MB) per
entry across MAX_CACHED_SNAPSHOT_WINDOWS (2) entries per emulator, for the
session's lifetime, with no aggregate budget across sessions. It also carried
an invalidation contract that produced three separate over-invalidation bugs
during review (the parse fence, setCwd/setLastTitle, and the no-op resize).

The lock fix is unaffected and independently measured: the `options` phase,
which is pure queueing, went 0/125/212/291ms -> 1/1/1/1ms across a 4-tab
worktree activation and stays there.
2026-08-31 16:08:47 -07:00
Jinwoo Hong 8f15f217a2 Preserve user-set workspace names across branch changes (#17448)
* fix(worktrees): preserve user workspace names across branch changes

* test(worktrees): cover pinned rename metadata

* fix(workspaces): address display-name review edge cases

* fix(workspaces): keep automatic names fresh across refreshes

* fix(workspaces): preserve legacy CLI labels

* fix(workspaces): preserve display-name provenance across hosts

* fix(workspaces): honor legacy display-name provenance

* fix(workspaces): fence display-name refresh races

* fix(workspaces): accept peer renames from provenance-less hosts

The old-host preserve fence kept a pinned local label on every refresh,
which also suppressed a legitimate rename another client persisted
through the same host until app restart. Narrow it to labels the host
re-derived itself (branch short name, or path basename when detached);
any other changed label in a mode-less response is explicit meta a peer
wrote there. Stale prior-label responses stay covered by the downstream
staleness fence, in-flight writes by the pending fence.

* refactor(workspaces): unify display-name pin derivation

Three call sites (renderer optimistic update, local IPC updateMeta
handler, remote worktree.set handler) each restated the same formula;
a future edit to one would silently skew provenance between paths.
2026-08-31 19:08:05 -04:00
Jinwoo Hong b44ef1e59d fix(skills): narrow computer-use discovery boundary (#17736)
* fix(skills): narrow computer-use discovery boundary

* chore: remove merge-formatting noise

* fix(skills): name browser page automation surfaces
2026-08-31 18:57:52 -04:00
Neil 20a12a6a46 perf(codex): share one launch-prep hook install across a spawn burst (#17669)
* perf(codex): share one launch-prep hook install across a spawn burst

Codex launch prep runs a full managed-hook install on every local PTY
spawn, and both install lanes serialize globally per Codex home. Opening
a multi-pane worktree therefore paid N full installs back to back, and a
resumed Codex pane prepares twice. Concurrent spawns for the same runtime
home now share one run; the promise is dropped as soon as it settles, so
the next launch still re-reads hooks.json and the user's trust state.

Also split the `host_env` spawn-timing phase, which spanned the entire
Codex preamble and pinned that cost on the env builder that ran last.

* refactor(codex): unify the two hook-install single-flight lanes

Both the WSL and launch-prep lanes now share one generic in-flight helper
instead of duplicating the map bookkeeping. Also routes the WSL launch-prep
install through the serialized variant, which closes the same per-spawn
serialization gap on WSL that the native lane just got.

* refactor: extract the shared in-flight run dedupe

The codex hook service and the GitHub conflict-summary cache had grown
near-identical private copies of the same single-flight helper. Both now
use one module, which also keeps the hook service clear of the 300-line
budget. The shared copy keeps the identity check on clear so a late settle
cannot evict a newer entry for the same key.
2026-08-31 15:16:38 -07:00
Neil 4ca232c159 perf(terminal): cut cold worktree-switch attach latency (#17667)
* perf(terminal): let a worktree's terminal spawns run concurrently

The per-worktree terminal mutation guard was a FIFO mutex, so activating a
multi-tab worktree made each tab wait for every predecessor's whole spawn.
Measured with ORCA_PTY_SPAWN_TIMING=1 on a 4-tab worktree, the `options`
phase was a pure-queueing staircase: 0 / 125 / 212 / 291ms.

The invariant that guard protects is spawn-vs-sleep exclusion, never
spawn-vs-spawn. Replace it with a writer-preferring shared/exclusive lock:
spawns share, sleep still excludes, and a queued sleep blocks later spawns
so a stream of spawns cannot starve it into its 12s deadline.

Same worktree after: options = 2 / 3 / 12 / 34ms, and stable_adoption
flattens from 65/398/398/312ms to a steady ~253ms.

The existing folder-workspace control assertion asserted the FIFO behavior
this removes, so it is re-based on sleep-vs-spawn (which still queues) and
joined by a case asserting concurrent same-worktree spawns.

* perf(terminal): memoize the headless snapshot per mutation epoch

Attaching a viewer serializes the session's whole headless buffer
synchronously on the daemon event loop, so every reattach of a quiescent
session re-serialized identical bytes. Measured with ORCA_PTY_SPAWN_TIMING=1,
stable_adoption was 253-281ms per session on reattach.

Move snapshot assembly into HeadlessSnapshotCache and memoize its expensive
parts (the serialize, the OSC link walk, the frame-restore fields) on a
mutation epoch that every emulator state mutation bumps, so a cache hit is
byte-identical by construction rather than merely fresh-enough. The async
write path bumps on entry and again in the parse-completion callback, so a
snapshot taken mid-parse can never be retained.

Reattach of a quiescent session after: stable_adoption 12ms. Sessions with
output since their last snapshot re-serialize exactly as before.

Retention is capped: an entry is held for the session's lifetime once it goes
quiescent, and a renderer may request 50k scrollback rows, so oversized
payloads serve normally but are not retained. Cache hits clone nested values
so a caller mutating its snapshot cannot corrupt later ones.

* perf(terminal): keep the snapshot cache warm across zero-byte parse fences

Review follow-up. flushParsedWrites() is write(''), used purely as a parse
fence, and every getSettledSnapshot runs one — so the epoch bump on an empty
write evicted the attach entry on each checkpoint read, defeating the cache
for any session that gets checkpointed.

Zero bytes cannot mutate the buffer: the OSC and mouse-mode scans are no-ops
on '' and the partial-escape tail is idempotent, so skip the bump for empty
data. Real writes still bracket themselves, and any write a fence orders
behind has already bumped on its own completion.

Also invalidate on dispose, so a post-dispose read can never be served a
pre-dispose entry, and freeze the emulator's public method surface in a test:
the cache's correctness rests on every mutator calling markMutated(), which is
convention rather than a type, so a new method should be a deliberate decision
about invalidation instead of a silent stale-snapshot bug.

* refactor(terminal): apply elegance review to the attach-latency fixes

Reuse: waitForMutationGrant hand-rolled the deadline race that
settleBeforeDeadline (same directory, four existing callers) already owns.
Using it also picks up the timer.unref() the local copy lacked, which was
keeping the Node event loop alive for up to the 12s sleep deadline.

Simplify: drop the waiter `abandoned` flag. The timeout path sets it and
splices the waiter out in the same synchronous block, so no queued waiter can
ever be observed abandoned and both reads were unreachable. The splice is what
actually does the work; the comment now carries why that makes a
grant-after-timeout unrepresentable. drain() then collapses into its loop
condition and reuses markActive() instead of inlining it twice.

Extract markWritten() so the zero-byte parse-fence rationale lives at one
mutation gate instead of being restated at three call sites.

Match the daemon's byte-accounting convention: the retention cap is now
expressed in bytes with code-unit sizing, like MAX_COLD_RESTORE_CACHE_BYTES
next door, so the two retention budgets read in one unit. Same effective cap.

The public-surface guard test caught markWritten on the first run, which is
the behavior it was added for.

* perf(terminal): stop discarding the snapshot cache on non-memoized fields

Second elegance round, and it found the same class of bug as the parse-fence
one: cwd and lastTitle are read fresh on every build and were never memoized,
yet setCwd/setLastTitle bumped the epoch — discarding a whole serialize to
update a field the cache does not hold. OSC 7 cwd updates land on every `cd`,
so this was a live cost on exactly the busy sessions the cache targets. The
invariant is "every mutation of a memoized part", not "every state mutation";
both docblocks said the latter and are corrected.

Drop the dispose bump too. A post-dispose getSnapshot re-serializes the
disposed terminal to byte-identical content, so the bump bought nothing and
only reached into a disposed xterm — verified by probe, not assumed. Its test
asserted zero serializations after dispose, which no implementation could
violate; it passed with the bump deleted.

Also memoize rehydrateSequences (a string, so no clone needed) instead of
rebuilding it on every hit, derive the frameRestore type from
buildFrameRestoreSnapshotFields so a new field cannot flow through at runtime
while the type omits it, inline the single-caller resolve() into build(), and
move the write() entry bump after the sync early-return so the three bumps map
1:1 onto sync / async-pre-parse / async-post-parse.

The surface guard is sorted in source and renamed to say what it freezes: the
prototype, TS-private members included.

* fix(terminal): correct the fence justification and key the cache by window

The markWritten docblock claimed "zero bytes cannot mutate the buffer". That
is false, and I verified it: `_core.writeSync('')` drains xterm's pending queue
and applies it. The exemption is still correct, but for a different reason —
a fence cannot introduce an *unattributed* mutation, because any bytes it
drains belong to a queued async write whose own completion callback bumps
first. The two write regimes are exhaustive: with writeSync present nothing
can queue, without it every write is async and self-bumps. A comment asserting
a false invariant is worse than no comment, since the next change may rely on
it, so it now states the real one.

Key the cache by scrollback window instead of a single slot. Consumers ask for
different windows against the same emulator — attach passes the full window
while agent/text reads pass 0 — so one slot thrashed to a 0% hit rate whenever
they alternated, silently removing the benefit on runtime-side emulators. Two
entries cover every caller pair in the tree.

Drop the epoch counter: markMutated already nulls the retained entry and an
entry is only ever stored under the current epoch, so the comparison could
never fail. Invalidation is simply "clear the cache".

* refactor(terminal): name the lock sides shared/exclusive for a third caller

Rebasing onto main surfaced a semantic conflict the merge applied cleanly:
main added runWorktreeTerminalMutation (terminal orphan adoption, #17159) as a
third caller of the guard this PR changed. It needs the exclusive side —
adoption reconciles a worktree's terminal records, so it must not interleave
with a spawn registering a pty or with a sleep, which is exactly the semantics
it was written under when the guard was a plain mutex.

With three operations, naming the sides after two of them no longer fits, so
the kinds are now `shared` (spawn) and `exclusive` (sleep, adoption) — what
they do rather than who calls them.

* perf(terminal): do not invalidate the snapshot cache on a no-op resize

Re-measuring the final rebased build caught the cache barely working on the
path it exists for. Every attach re-asserts the pane's dimensions, and resize()
bumped unconditionally, so a reattach of a fully idle session missed its own
cached snapshot and re-serialized.

Measured on the same 4-tab worktree, reattach of sessions verified quiescent
(no buffer change over 4s), stable_adoption per session:

  before this commit:  98 / 382 / 395 / 418 ms
  after:               16 /  67 /  78 /  87 ms

A resize to the size already applied changes nothing the snapshot reads, so it
now returns early — which also stops it clearing restoredOscLinks, correct
since no rows shifted.
2026-08-31 15:16:24 -07:00
spellbook 18e8fe4770 fix(serve): install supervisor disconnect quit after app environment init (#16762)
Moves installServeSupervisorDisconnectQuit(isServeMode) out of module scope in
src/main/index.ts to just after setAppEnvironment() and initDataPath().

The call resolves the serve update handoff path through getCanonicalUserDataPath(),
which throws by design until the app environment accessor is installed. At module
scope that throw was unconditional on macOS whenever the CLI set
ORCA_SERVE_UPDATE_HANDOFF_PATH — which it does by default — so every `orca serve`
process died at startup before it could listen, and the supervising service manager
restarted it into the same crash. Reported in #16761, #16698 and #17509; shipped in
1.4.190 through 1.4.192.

Guards added so it cannot drift back: a source-level ordering assertion that also
pins the call synchronous and inside the single-instance block, and a runtime test
that keeps the real path resolver, since the existing suite mocks it and therefore
could never have caught this.

Fixes #16761
Fixes #16698
Fixes #17509
2026-08-31 14:58:10 -07:00
Jinwoo Hong 8bebcf5def fix(terminal): preserve large agent prompt pastes (#17718)
* fix(terminal): preserve large agent prompt pastes

* fix(terminal): guard oversized SSH PTY writes

* fix(terminal): avoid timer delay for generic sends

* fix(pty): propagate provider write refusals
2026-08-31 17:56:39 -04:00
Jinwoo Hong 02a7742406 fix(artifacts): raise desktop sharing limit to 5 MiB (#17708)
* fix(artifacts): raise desktop sharing limit to 5 MiB

* fix(artifacts): enforce recovery content limit

* fix(artifacts): bound recovery request envelopes

* fix(artifacts): clarify oversized request error
2026-08-31 16:51:11 -04:00
Neil aa658d28e3 test(updater): stop a slow module import from failing the next test (#17726)
* test(updater): stop a slow module import from failing the next test

`updater.ts` is 2.4k lines. Its first transform in a worker costs ~1.4s idle
but 45s+ when the machine is oversubscribed, which is past the 30s
`testTimeout`. Vitest cannot cancel the timed-out test body, so the abandoned
continuation went on to call `setupAutoUpdater` during the *next* test — with
the harness already reset — and failed it with:

    AssertionError: expected "vi.fn()" to be called 1 times, but got 2 times

That is the exact signature of the abandoned-instance timer flake fixed in
#17649/#17663, so a machine-load timeout reads as that regression returning and
sends the reader hunting in the wrong place.

Two changes, in `updater-test-module-loader.ts`:

- `loadUpdaterModule()` replaces every `await import('./updater')` in the suite.
  It records the test that asked for the module and throws if the import
  resolves after that test ended, stranding the continuation so the timeout
  stays the only reported failure. This removes the trap.
- `warmUpdaterModule()` imports the module once in `beforeAll`. The transform is
  cached across `vi.resetModules()` — only a file's first import pays it — so
  warming moves that one slow import onto the 60s `hookTimeout` and leaves every
  in-test import at re-evaluation cost (~25ms idle).

Measured on a 16-core mac, first vs later import in one file: 1439ms / 25ms
idle, 8339ms / 149ms under 40 CPU hogs, 45521ms / 15182ms under 400.

Under 400 hogs the suite went from 15 files and 22 tests failing (15 timeouts
plus 7 misleading assertion failures) to 23/23 files and 269/269 passing. Under
900 hogs it degrades into 14 plain `Hook timed out in 60000ms` failures and zero
assertion failures.

* fix: tighten the fence, surface its warning, stop patching timers on warm-up

Review findings on the loader:

Drop trackRealTimers() from warmUpdaterModule(). It was inert — updater.ts
arms no timers at module scope — and actively harmful for the 5 files that
build their own mocks and never call clearTrackedRealTimers(). Those files
previously had pristine timer globals; the warm-up installed a wrapper that
was never restored and whose armed-handle set grew unbounded.

Key the fence on TestRunner.getCurrentTest() instead of currentTestName.
Nothing ever clears currentTestName, so the fence only fired once the *next*
test had started; a continuation resolving during the timed-out test's own
teardown, or after the file's last test, was still handed the module. The
last-test case mattered: the harness afterAll has already cleared timer
tracking by then.

Emit the diagnostic through process.emitWarning. The throw lands on a promise
vitest already settled, so the message explaining why the continuation was
stranded was discarded and reached nobody — which was the entire payoff.

Widen the loader test's race margin 50ms -> 500ms. It gated on the
test-to-test transition completing in 50ms, so the regression test for a
contention bug could itself fail under contention.
2026-08-31 13:22:27 -07:00
Neil 08d95b9979 test(native-chat): remove initial-snapshot recovery race in watch-error test (#17722)
Root cause: the test cleared the injected tail-reader failure *before*
writing the recovered transcript line. The capped rotation retry loop is
still firing at that point, so a retry drain could succeed against the
still-empty file, consume the pending initial drain, and emit an empty
initial snapshot (`[], false, 0, undefined, undefined`). The later manual
watch callback then took the append path, and `u-recovered` never reached
onInitialSnapshot -- producing the CI failure
`expected [ false, +0, ...(5) ] to deeply equal ArrayContaining{...}`.

That empty-snapshot-then-append sequence is correct product behavior, so
this is a test bug: write the content first, then clear the failure, so no
drain can ever observe a readable-but-empty transcript. The assertion now
checks the exact recovered snapshot instead of a flattened
arrayContaining, so an empty recovery snapshot fails loudly.
2026-08-31 13:22:20 -07:00
Neil a381b47437 test(cursor): widen Windows hook spawn budget to fix ETIMEDOUT flake (#17721)
* test(cursor): widen Windows hook spawn budget to fix ETIMEDOUT flake

`package (windows)` failed once on an unrelated packaging PR with
`spawnSync cmd.exe ETIMEDOUT` at hook-service.test.ts:78. This is an
infrastructure-timing flake, not a logic race: the assertion is
`expect(result.error).toBeUndefined()` and `ETIMEDOUT` only means the
spawnSync `timeout` elapsed.

Cursor is the heaviest of the hook-service suites on Windows. Its managed
command is the PowerShell encoded launcher, so one hook run is
cmd.exe -> powershell.exe -> cursor-hook.cmd -> curl.exe: four process
creations, one of them a CLR start that installer-utils.ts itself
documents as ~300ms warm and "visibly slow". The sibling suites (codex,
grok, agent-hooks/installer-utils) spawn the .cmd directly and set no
per-spawn timeout at all, so 15s here was a one-off, not a convention.

Raise the per-spawn budget 15s -> 30s to match
WINDOWS_PROCESS_TEST_TIMEOUT_MS in src/shared/setup-agent-sequencing*.test.ts
and the 30-90s used by the real-subprocess tests in src/main/browser. 30s
is ~30x the warm cost of the chain, which leaves room for CPU contention
and Defender scanning of the freshly written .cmd on a packaging runner.

The default vitest testTimeout is also 30s, which would have become the
new binding constraint (the protocol case runs 16 chains back to back), so
give the four spawning cases 120s. That keeps ETIMEDOUT - which names the
stuck process - as the failure you see, instead of an opaque case timeout.

No product behavior changes and no end-to-end coverage of the Windows
launcher is removed.

* fix: halve the case timeout and correct the contention rationale

Review found the stated cause wrong. pr.yml runs "Test Windows-specific
boundaries" before "Build package inputs", so electron-builder is not
running. The real contender is that vitest invocation itself: ~25 files at
maxWorkers 4, including five real-Electron suites and two node-pty tests.

120s was over-provisioned. windows-hook-payload-delivery.test.ts drives the
identical PowerShell chain on the same job with a 60s case budget; 60s gives
the same property here (16 warm spawns plus one 30s outlier) and halves
time-to-signal on a genuinely stuck chain.

Also record that 30s deliberately exceeds the product's own
MANAGED_HOOK_TIMEOUT_SECONDS (10s) — this test gates launcher correctness,
not user latency, so the SLA is not the right bound. Left
windows-hook-payload-delivery.test.ts at 15s: its value is deliberate, set
to mirror Claude Code abandoning a hook at 10s.
2026-08-31 13:22:12 -07:00
Brennan BensonandMerge Sim 872bd51d47 fix(native-chat): reland large structured command results (#17720)
* fix(native-chat): preserve large structured command results (#17707)

* fix(native-chat): preserve large structured command results

* chore: place native chat validation artifacts under docs

* chore: drop stale root package config

* fix(native-chat): enforce rebuilt lifecycle append slots

---------

Co-authored-by: Merge Sim <sim@local>

* chore: omit native-chat reland planning docs

* fix(native-chat): remove journal store import cycle

* fix(native-chat): keep journal factory acyclic

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 13:12:19 -07:00
Neil dc5db4b01c fix(lint): merge duplicate process-table-snapshot imports (#17724)
main is red on `static analysis`: oxlint's code-quality pass runs with
--deny-warnings, and agent-foreground-process-batch.test.ts imports
'../../shared/process-table-snapshot' twice (lines 5 and 13), tripping
"Modules should not be imported multiple times in the same file".

Introduced by #17525. It blocks every open PR, none of which can go green
until this lands.
2026-08-31 13:04:40 -07:00
Brennan BensonandMerge Sim 9477b5fcbb feat(ssh): batch process evidence in PTY inventory (#17525)
* feat(ssh): batch process evidence in PTY inventory

* fix(ssh): accept Linux kernel process rows and make no-evidence polling push-driven

* fix(ssh): preserve process evidence polling semantics

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 12:44:41 -07:00
Brennan Benson 894ed75abb Revert "fix(native-chat): preserve large structured command results (#17707)" (#17719)
This reverts commit 5fe37729ea.
2026-08-31 12:34:49 -07:00
Brennan BensonandMerge Sim 5fe37729ea fix(native-chat): preserve large structured command results (#17707)
* fix(native-chat): preserve large structured command results

* chore: place native chat validation artifacts under docs

* chore: drop stale root package config

* fix(native-chat): enforce rebuilt lifecycle append slots

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 12:34:03 -07:00