Commit Graph
10202 Commits
Author SHA1 Message Date
Neil f03dc12b0d feat(worktree): prepare agent shells in retained composer drafts 2026-09-05 04:45:29 -07:00
Neil b268adef18 feat(runtime): prepare owner-fenced deferred terminals 2026-09-05 04:01:17 -07:00
Neil f0110601bd feat(daemon): negotiate deferred startup preparation and release 2026-09-05 03:49:32 -07:00
Neil f75781902d feat(daemon): hold composer startup until an owner-fenced release 2026-09-05 03:31:37 -07:00
Neil eae9f17236 refactor(pty): separate startup identity from eager shell argv 2026-09-05 03:12:06 -07:00
Neil b7375e090b docs(worktree): clarify successful preparation refill ownership 2026-09-05 03:02:52 -07:00
Neil 267bd8f094 fix(worktree): replenish active standby after successful creation 2026-09-05 02:56:12 -07:00
Neil 50071fc362 perf(worktree): retain an idle preparation for the active composer target 2026-09-05 02:42:50 -07:00
Neil 27bbd61b96 feat(worktree): expose checkout-only standby preparation 2026-09-05 02:30:33 -07:00
Neil d3ebcd4535 fix(worktree): preserve retained creation progress on submit 2026-09-05 02:15:08 -07:00
Neil 504f2a5d0a fix(worktree): preserve background startup across runtime requests 2026-09-05 01:55:32 -07:00
Neil b58849c71a fix(worktree): await queued Codex trust before agent startup 2026-09-05 01:47:50 -07:00
Neil da23e30e90 perf(worktree): prepare agent checkouts before composer submission 2026-09-05 01:41:49 -07:00
Neil 1f7392d2d7 perf(worktree): retain remote composer checkouts before terminal activation 2026-09-05 01:29:10 -07:00
Neil 14b360733b fix(worktree): preserve selection during background runtime startup 2026-09-05 01:28:40 -07:00
Neil 602bef4b6f perf(worktree): keep successful preparation off stale cleanup waits 2026-09-05 01:17:14 -07:00
Neil 766e30ca46 perf(worktree): activate completed composer workspaces without async gaps 2026-09-05 01:09:27 -07:00
Neil 3ff0d0e4c3 feat: retain background composer workspaces for blank terminal creation 2026-09-05 00:55:27 -07:00
Neil 89cf4433ac perf: preserve user Git checkout worker settings 2026-09-05 00:53:41 -07:00
Neil 045eb6c5b4 refactor: separate retained workspace creation from composer activation 2026-09-05 00:26:37 -07:00
Neil 6f0e63a6de test: combine main with worktree creation PRs 18793 and 18794
# Conflicts:
#	src/shared/child-process/process-tree-termination.test.ts
2026-09-04 23:50:48 -07:00
Neil 36a826ff48 fix(ssh): compile node-pty from the host's own Node headers instead of nodejs.org (STA-6674) (#18774)
* fix(ssh): compile node-pty from the host's own Node headers instead of nodejs.org

STA-6674: a Linux SSH host that cannot reach nodejs.org never came up. node-pty
ships no Linux prebuild, so npm hands it to node-gyp, and node-gyp's default is
to download node-v<ver>-headers.tar.gz before configuring. The host refused
that connection (ECONNREFUSED) and the relay deploy failed inside npm install,
which the UI showed only as "Disconnected".

Every official Node build and every version manager that unpacks one already
has those exact headers at <prefix>/include/node. Export node-gyp's nodedir to
that prefix, on every command that can compile node-pty (npm install, npm
rebuild, the cloexec patch's rebuild), when the shipped node_version.h matches
the running Node. Both npm_config_nodedir (node-gyp 10, Node 20) and
npm_package_config_node_gyp_nodedir (node-gyp >= 11.4) are set so every Node
the relay runs on reads it. A version mismatch leaves it unset, which is the
existing behaviour.

When a host is both header-less and offline, name that in the deploy error
instead of forty lines of gyp http output, with the two remedies.

Reproduced and verified with a Docker sshd whose nodejs.org resolves to
127.0.0.1, on node:24.12.0 (the user's version), node:20 and node:26:
ssh-relay-offline-node-headers.docker.test.ts.

* fix(ssh): fail loudly when node-gyp ignores the exported Node headers dir

The headers export relies on npm forwarding npm_config_nodedir /
npm_package_config_node_gyp_nodedir into lifecycle scripts. If a future npm
drops that, node-gyp would silently fall back to downloading, and an offline
host would fail with the same "install an official Node" diagnosis -- wrong,
since the host did ship headers.

The prefix now echoes ORCA-NODE-HEADERS:<dir|none> into the command's output
before the compile, and the download-failure diagnosis reads it back: an
exported dir plus a download attempt is reported as an Orca defect naming
the dir, not as a host problem. Nothing else changes when it works.

* fix(ssh): address review on the relay node-headers export

- Unset any inherited npm_config_nodedir / npm_package_config_node_gyp_nodedir
  before the conditional export, so a stale header dir from the remote profile
  cannot bypass the version check and build a wrong-ABI binding (CodeRabbit).
- Require `gyp ERR! configure error` and a real network errno in the
  headers-download matcher; node-gyp's fetch client logs retried attempts it
  recovers from, and a FetchError can be a non-2xx mirror answer (pullfrog).
- Say "no local headers matching its own version", since the probe also
  rejects a version mismatch, not only absent headers (CodeRabbit).
- Log the same diagnosis from the non-fatal `npm rebuild` fallback (CodeRabbit).
- Docker test waits for the SSH banner on the mapped port before connecting
  instead of trusting `docker run -d` (CodeRabbit).

* fix(ssh): read the node-headers marker from the host output, not the quoted command

execCommand rejects with `Command "<command>" failed (exit N): <output>`, and
<command> quotes the whole prefix, marker echo included. The first-match
parser hit that copy and returned `${ORCA_NODE_HEADERS_DIR:-none}"; ...` as a
"dir", so every real no-headers failure was misreported as an Orca defect
(measured by an independent Docker exercise of 609685e). Strip the exec-
failure head before scanning; keep first-match so gyp output cannot spoof it.

The unit fixture hid this by rejecting with `Command "npm install" failed`,
a string production never builds. It now rejects from the command the mock
actually received, and the Docker test gains a no-headers failure case on the
same offline fixture that asserts the host-remedy message.

Also unset NPM_CONFIG_NODEDIR (npm accepts either case), and narrow the claim:
a ~/.npmrc nodedir= is not overridable from the env (measured: empty env
override is ignored on npm 10 and 11), so it stays the operator's setting.
Copy the node binary via fs in the unit test so a failed copy fails the test.

* docs(ssh): state the header-mismatch refusal as a conservative default, not an observed crash

* docs(ssh): note why the exec-failure head regex may match lazily
2026-09-04 23:47:32 -07:00
Neil 58553bfe1c fix(recovery): fail a renderer recovery reload that never loads, instead of leaving a dead window (#18466) 2026-09-04 23:45:42 -07:00
Brennan BensonandMerge Sim 172aa1ac35 feat(native-chat): render agent file edits as inline diff cards (#18765)
* feat(native-chat): render agent file edits as inline diff cards

An agent's file edit rendered as a flat list of every removed line followed
by every added line, with no interleaving, no file header, and no line
numbers. A Codex edit on the transcript lane rendered no diff at all: the
patch arrives wrapped in the source string of its `exec` tool, which matched
none of the shapes the old parser looked for.

Adds one diff model shared by every edit shape the supported agents produce:

- `native-chat-edit-lcs` interleaves a snippet pair, falling back to a linear
  prefix/suffix diff above the quadratic guard.
- `native-chat-unified-patch` keeps the `@@` ranges as per-row line numbers
  instead of parsing them into display text and discarding them.
- `native-chat-begin-patch` recovers the `*** Begin Patch` envelope from the
  JavaScript string literal Codex sends it in, so that lane renders a diff.
- `native-chat-edit-normalize` folds all of it into one model, including the
  two Codex shapes that do not look like diffs: add and delete arrive as raw
  file content, and a rename is appended to the body as prose.

Claude reports an edit as a snippet pair, which cannot locate the change in
the file, so its result's resolved hunks are now carried on the tool-result
block and preferred when present. The field is optional, so an older client
reading a newer journal simply drops it. Where no resolved ranges exist the
gutter stays blank rather than showing a snippet-relative number, which would
read as a file position.

The card renders the verb from the observed change kind rather than the tool
name, pairs an edit's call and result into a single row, and takes its row and
gutter grounds from new tokens derived from the git status palette, replacing
the hardcoded Tailwind tints the old view used.

Desktop only; mobile chat keeps its existing renderer and parser untouched.

* fix(native-chat): stop the diff card from asserting an edit it cannot prove

Every defect here shares one failure mode: the card stated something the
input did not support, and stated it confidently.

Parsing:

- A hunk no longer ends on `--- `, `+++ ` or `\ No newline`. The first two
  are what a removed `-- comment` (SQL/Lua/Haskell) looks like once the
  marker is prepended, so they truncated the whole diff; the no-newline
  marker is emitted mid-hunk, between the removed old last line and the
  added new one. Real headers are recognised through `isFileHeaderPair`,
  lifted out of `native-chat-diff` so the rule has one home.
- A `*** Begin Patch` envelope with no `*** End Patch` is declined. With no
  closing marker `indexOf` returned -1 and the slice swallowed the rest of
  the command line, so `… +y" && echo ok` rendered as file content the
  agent never wrote.
- One splitter serves every shape, so a CRLF patch no longer keeps a `\r`
  on each row, in the phantom-row guard, or in the clipboard. It also
  tests for the trailing newline on the clipped body: on the un-clipped
  string that test deleted a real line whenever the slice fired.
- Truncation is carried from each slice site to the card, so content past
  the character cap can no longer render as a complete unchanged file with
  no "Diff truncated" footer.

Attribution:

- A failed or still-running edit renders no card. It kept the generic tool
  view, whose result block carries the provider's own error — the card had
  been drawing "Edited file +1 −1" from the input while hiding the red
  error body, which is worse than what preceded this feature.
- The result-as-patch fallback is scoped to `Diff`, the one tool whose call
  carries only a path. Any command tool's output could previously be read
  as a patch, so `git diff` through `exec` was reclassified as an edit of a
  file named "file" and its command line disappeared with the result.
- A whole-content write claims a creation only on evidence — the editor
  tool's own `create` command, or the provider reporting one. Overwriting a
  large existing file had always read as "Added file".
- `MultiEdit` reads its `edits[]`, and `NotebookEdit` leaves the set: it
  carries only the new cell source. Both previously fell through to the old
  renderer, so one turn could show two diff presentations at once.
- Snippet-relative numbers are dropped at the model layer rather than
  hidden by a zero-width gutter, which the flex min-width floor re-exposed
  on top of the marker and the first characters of the row.

The run memoizes its edit model, so a collapsed group no longer re-diffs on
every streaming token, and the card's copy button says what it copies.

* fix(native-chat): keep every edited file, and mark where the diff breaks

A run of hunks was concatenated into one flat row list, so the gutter jumped
from one region of the file to a distant one with nothing between them and
the reader saw two unrelated spans as one continuous block. Rows now carry
an explicit break: it holds no text and no position, counts toward neither
side of the change, is trimmed from the end where it would mark nothing, and
is left out of the copied text.

The patch envelope lost files, and lost them silently:

- An update chunk may carry no hunk header at all. The parser required one,
  returned nothing, and the caller dropped that file from a multi-file
  envelope with nothing to say it had gone. A header-less body now opens as
  a hunk of unknown position, and whether the rows are locatable is read off
  the rows themselves rather than off the header.
- The envelope's own control lines rendered as content rows in the card.
- A delete names its file and carries no body, which rendered as a card with
  an empty expandable row list. The header states the change and offers no
  disclosure behind it.
- The header patterns are anchored and `.` excludes a carriage return, so a
  CRLF envelope matched no header at all and produced no card whatsoever.
  The envelope is split on both newline forms once, up front, rather than
  each pattern having to tolerate the extra character.

A tool call's argument payload arrives as a string holding JSON. It was
passed along undecoded, which is the only reason this code carried a
hand-rolled string-literal unescaper. It is decoded once at the transcript
decoder now — defensively, since the transcript is untrusted, so anything
that is not a JSON object is left exactly as it arrived — and the unescaper
is gone. Recovering the envelope no longer guesses at argument names either:
it looks at the values, including the words of an argument vector, which is
where the envelope actually sits once the payload is decoded.

* fix(native-chat): only read a patch where a patch was actually run

Recovering the patch envelope from any value of a tool's payload meant a
write's own content was searched for one. A file documenting the patch
format rendered a card for the file its example names, while the file
actually written never appeared at all — the call and its result were
consumed by that card, so nothing was left to correct it. Two changes: the
envelope is recovered only for the tools that run one, never for a file
edit whose payload is content; and only patch- or command-bearing arguments
are searched, still including the words of an argument vector, which is
where the envelope sits when a command tool applies it.

The call payload is decoded back where it is needed rather than at the
transcript decoder. Decoding it there changed the shape every reader of a
tool's input sees, including the surface that recognises a question payload
from any tool by shape alone: a tool whose arguments happened to carry that
shape raised a question card pinned over the composer. That decode now
happens inside the envelope recovery, the one consumer that needs the
structure.

A card also states an edit as made, so it now takes evidence that it landed
— the provider reporting the call complete, or a result that is not an
error. A turn that stopped before its call was answered reported an edit
that may never have applied. This replaces the working-turn heuristic in the
view, so the rule lives in one place.

Two files still went missing. A multi-file patch has no per-file split, so
it rendered as one card under the first file's name, with the later files'
rows and their gutter numbers beneath it — a card asserting a false file
position. Patch text is now split on its file boundaries, one card per file,
each named by its own header, with a rename and a `/dev/null` side read from
the same headers. And an envelope section that names a file but carries no
body was dropped rather than reported, which is the same silent loss the
delete case was fixed for.

* fix(native-chat): type the patch-section scan and its test helper call

The section under construction was only ever assigned inside the helper that
opens one, which control-flow analysis does not see, so the variable stayed
narrowed to its initial null and reading a field off it did not compile. The
helper now only builds and records a section; the loop owns the assignment,
which also fixes a real leak in the fall-through row: it opened a section it
never made current, so the next row opened another one.

The multi-file case also passed a possibly-undefined slice to a helper that
takes an array or null.

* fix(native-chat): stop the patch lane naming files it cannot name

Splitting a patch into its files only ever looked for a boundary outside a
hunk, and nothing reopened that state once the first hunk began, so every
file after the first was swallowed as the first one's body. A `--- `/`+++ `
pair inside a hunk is now a boundary too, but only when a hunk header
follows it immediately: a removed `-- x` over an added `++ y` is never
followed by a column-0 header, which is what keeps the guard against reading
content as structure intact.

One producer cannot be recovered by any parser: it joins several files'
patches and keeps a count where the path goes, so nothing in what reaches
here names a file. That shape is refused rather than rendered under a name
no file has. Recovering the per-file paths belongs to the producer and is
filed separately.

A clipped body carries its own marker in its text, and the bound that clips
it is six times smaller than this module's, so it fires first. Read as
content, the marker became a numbered line of the file and the rows before
it were reported complete. It is recognised at the end of the text, removed,
and reported as the truncation it is — the footer says so and the copied
text no longer carries it.

Also: a move appended to the body as prose is now read as a rename on every
lane that carries the body as text, not just the one that also carries the
destination as a field, where it had been rendering as a numbered line of
the file it moved. The call's own path no longer wins over a rename's
destination, which is only ever in the header, and only sections that name a
file count toward deciding whether the call names the one file at hand. A
command that merely quotes an envelope — writing documentation about the
format — must now also invoke the tool that applies one. And two compared
directories are no longer called a rename: only a header that states both
sides as such is evidence of a move.

* fix(native-chat): anchor the move marker to its own line

The marker a producer appends to say where a file moved was matched anywhere
on the body's last line, so a row whose own content mentions a move was cut
in half at that point and the file it named claimed as the destination of a
rename that never happened. It is now anchored to the start of the final
line, on both lanes that carry the body as text.

The command that applies a patch envelope has a second spelling the runner
accepts and runs; requiring the first one refused a patch that really landed.
Both are accepted, still matched against whole argument words rather than the
payload at large.

A clipped diff also said so only under its own rows, where a collapsed card —
or one clipped down to no rows at all — showed nothing. It sits beside the
change counts now, which are visible either way.

* refactor(native-chat): tidy what the diff-card work left behind

The copy text is joined from every row of the diff, which a collapsed card
renders none of, and it was rebuilt on every render to seed a prop. It is
memoized on the rows, matching how the run memoizes its edit model.

The two scanners that read patch text kept the same file-section alternation
verbatim, so they could drift apart while both looking correct; there is one
definition now, beside the header-pair rule that already lives there.

Also: the row that marks a break between regions is built in one place, so it
is no longer exported; the move destination in the envelope reader was a
function-wide binding written and read within one iteration, which read as if
a move carried between sections; and a test comment named the wrong mechanism
for keeping a card collapsed.

Adds the missing pin on what the copy affordance actually copies.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-04 23:42:43 -07:00
Neil 149732df6d perf(persistence): stop the session write re-scanning and rebuilding unchanged state (#18739) 2026-09-04 23:39:02 -07:00
Neil 79cc5168af perf(worktree): narrow speculation to preparation overlap and startup 2026-09-04 23:37:42 -07:00
Brennan BensonandMerge Sim b0c67eaf88 feat(mobile): port the restructured native-chat turn status and live tool progress (#18761)
* feat(mobile): port the restructured native-chat turn status and live tool progress

Mobile chat had a single static "Agent is working" row and no live tool
activity, while the desktop restructure (#17597, #18705) replaced that with a
per-turn status row and a running-tool label. This brings mobile to parity and
puts the derivation in one place instead of two.

Shared (new, pure, RN-safe — desktop uses them as i18n fallbacks, mobile
directly, matching the native-chat-empty-state pattern):
- `native-chat-turn-status.ts`: duration formatting, label selection, the
  turn-timing state machine, and the active/settled split.
- `native-chat-tool-activity.ts`: command-tool classification, the running-tool
  label descriptor, and running-call selection.

Desktop now consumes both; `NativeChatWorkingStatus`, `NativeChatToolRun` and
`use-native-chat-turn-status` keep their existing behavior and strings.

Mobile gains the "Thinking" / "Working for 12s" / "Worked for 3m 4s" row with a
caret that discloses the turn's tool activity, the pulsing "Running npm test"
row with terminal-vs-wrench glyphs, and desktop's rule that a completed turn's
tool run hides behind the turn caret. The bridge lane is untouched and keeps its
three-dot indicator. Headings, quotes, code, lists and table cells are now
selectable.

Files at their max-lines cap were split rather than bumped: the tool-run subtree,
the prompt card, the session-lane wiring, and the turn-disclosure state each move
to their own module.

* perf(mobile): stop the turn-status rows from re-rendering the whole transcript

A streaming turn re-renders the chat list many times a second. The disclosure
wiring handed every row a fresh status object and a fresh toggle closure on each
of those renders, so `MobileNativeChatMessage`'s memo never held and every
visible row re-rendered per tick — including settled turns that had not changed.

Memoize the status selection on the timing map, and keep one stable toggle
handler per turn (pruned when a turn leaves the transcript) attached only to the
settled rows that can actually disclose anything. Now only the live turn's row
changes identity while the agent works.

* fix(mobile): keep the turn clock running when the optimistic echo is replaced

An accepted send renders as `pending-N` until the transcript echo lands under
its real message id. That flips the active turn key mid-turn, and the timing
reducer treated the new key as a new turn — so a turn that had reached
"Working for 8s" visibly restarted at "Working for 0s".

The reducer now carries the start over when the previous key names a turn that
has since left the transcript, which is exactly the echo-replacement case. A
genuinely new turn (the previous key still in the transcript) and a turn that had
already settled both keep their own clock; both are pinned by tests. Desktop does
not pass the new key and is unaffected.

* fix(mobile): keep the Tools toggle working on settled turns

Hiding a settled turn's tool run behind the turn caret (desktop parity) also
made the composer's global Tools control a no-op on every completed turn: the
run it wanted to expand was not rendered at all. Let that toggle override the
hiding, so it still reveals every run at once the way it did before.

* fix(mobile): re-key the turn timing instead of only carrying its start

The previous fix carried the start forward only while the turn was still
working. When the transcript echo landed after the turn had already settled,
the new key inherited nothing, the settled timing was pruned with the old key,
and the turn's "Worked for N" row disappeared entirely.

Move the timing onto the new key instead, which covers both orderings: an
in-flight turn keeps counting from its original start (and later settles against
it), and an already-settled turn keeps its duration. Both orderings are pinned.

* test(mobile): pin the structured turn-status wiring at the view level

Emulator QA could not reach the structured lane (mobile's Create Tab -> Codex
falls back to a terminal tab when agentSession.createSupport says unsupported),
so the view's own lane wiring had no coverage — the one seam between the shared
turn-timing reducer and the rendered rows.

Assert what the view hands each row: the live user turn gets a status object and
the three-dot indicator is gone on the structured lane; the bridge lane keeps the
indicator and gets no status; a finished turn settles to a numeric duration with
a toggle; and an assistant row never carries a status row of its own.

* fix(mobile): isolate structured chat turn state

* fix(mobile): let the capability RPC actually store what a phone advertises

`runtime.clientCapabilities.update` records the advertised set by assigning
`authenticatedSocket.clientCapabilities`, but the socket handed to the dispatcher
defined that property with a getter only. In strict mode the assignment throws
`TypeError: Cannot set property clientCapabilities ... which has only a getter`,
so the RPC answered `runtime_error` and the set was never stored.

The consequence is not subtle: `supportsStructuredAgentSessions` requires the
capability, so `projectSessionTabAgentStatus` removed every `agent-session` tab
from a phone that had advertised it correctly. A paired phone saw ZERO tabs on a
worktree whose only tab was a structured Codex chat — structured native chat was
unreachable on mobile over this transport, not just missing its new turn UI.

Give the socket a setter that writes through to the channel, which already owns
the set for the connection's lifetime, so later requests on the same socket see
it. Found while trying to capture emulator screenshots of the turn-status port:
two full QA runs reported the new UI "missing" because the phone could only ever
get a bridge/PTY tab.

* fix(mobile): carry the turn key instead of caching a handler in a ref

Builds on the scope-isolation fix: that kept (and extended) a ref that is
written during render — once to memoize a per-turn handler, once to prune dead
turns, once to reset on a scope change. React Doctor's "Ref mutated during
render" is what CI's `check:react-doctor:changed` was failing on (x2), and on
mobile it is a real hazard rather than a style note: react-freeze discards
renders, and a discarded render would leave the cache mutated.

Pass the settled turn's key down the row instead and let it call one stable
handler with it. That preserves both properties the cache was bought for — per
scope isolation, and identity stability so a streaming transcript does not
defeat the row's memo — with no ref writes and no pruning to get wrong. The
scope-keyed expanded set and the 128-turn cap are untouched; their tests move to
the new contract and one now pins handler identity across a re-render.

Note for future changes here: `check:code-quality:changed` does NOT cover this.
CI additionally runs the standalone react-doctor CLI, which has rules the oxlint
plugin config does not enable.

* fix: ship native chat status translations

* test(native-chat): pin the shared copy against the English catalog

The shared constants are desktop's i18n fallback and mobile's actually-rendered
string. If one changes without the other, desktop keeps rendering en.json while
mobile renders the constant — and nothing fails, because a fallback is only used
when the key is missing. That silent divergence is the exact thing the shared
module exists to prevent, and it is now reachable precisely because these strings
are runtime-required rather than statically extracted.

Assert every key in both shared copy objects matches en.json byte for byte, plus
the interpolation placeholders the catalog interpolates on.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-04 23:31:39 -07:00
Neil 3f84c358f0 fix: recover from an alternate shell install, and close the relay duplicate-echo gap (#18796)
* fix: recover from an alternate shell install, and close the relay echo gap

#18768: a startup profile that `exec`s a second install of the same shell keeps
the pid but loses the wrapper's ready marker, and the recovery probe rejected the
replacement's different canonical path -- costing plain Codex the full 15s
barrier. The probe now also accepts an install that the pane's own PATH resolves,
so a binary merely named bash/zsh outside it stays rejected.

#18767: the SSH relay left plain Codex on early startup delivery, displaying the
launch twice under a slow profile. The shell is the host's to know, so the relay
now folds it into the same rule the daemon uses, and the SSH background client
waits for the marker on any Codex launch. Bracketed paste is now gated on an
observed marker rather than on the intent to wait, so a fallback release on a
host shell that never publishes one submits raw.

The non-daemon local provider needs no change: it hands Codex to the wrapper's
own prompt hook and never writes it into the PTY.

* review: correct an overclaiming comment, announce a silent skip, drop a shim

Readiness review of #18796. The delivery comment claimed waiting "costs nothing",
which is only true on a host that arms the marker -- fish, sh, Windows and hosts
predating #18767 release on the fallback instead. Say so.

The alternate-install recovery tests skip on usrmerge hosts, where /usr/bin/bash
resolves back to /bin/bash; announce that rather than reading as coverage that
does not exist.

Import the line-editor predicate from shared directly instead of through a
re-export left on daemon/shell-ready.
2026-09-04 23:27:30 -07:00
Neil 17ec0d4f94 fix(worktree): retain index warming ownership until group exit 2026-09-04 23:14:45 -07:00
Neil b33d1972bc docs(relay): correct why the ConPTY teardown asset diverges from the desktop patch (#18636)
The divergence pinned by #18601 is real and worth keeping, but its stated reason
was wrong. It claimed the desktop patch carries the early conin placement "and
therefore the +2 File / +1 Process regression, measured against its exact
installed tree" -- i.e. that the shipped desktop app leaks because of its own
leak fix.

It does not, for any terminal a user opens. node-pty defaults `_useConptyDll` to
false. Every desktop site that opens a pane sets it true (`local-pty-utils.ts`
twice, `native-pty-spawn.ts`), as does the `windows-conpty-warmup.ts` warm-up, so
they take the `else` branch, where upstream already destroys the input socket. The
relay passes no such option (`src/relay/pty-handler.ts`) and takes the
`!useConptyDll` branch -- the one both this asset and the desktop patch edit.

The desktop is not entirely off that branch, though: the hidden rate-limit probes
in `src/main/rate-limits/claude-pty.ts` and `codex-pty-rate-limit-probe.ts` omit
the option, recur, and tear down through `kill()`, so the hunk is live there --
just never for a visible pane. Whether the early placement costs the same +2 File
/ +1 Process across a probe's lifecycle is unmeasured; the numbers in this comment
were taken on relay-style spawn/kill cycles, and the comment now says so. What is
settled is the replaced claim: not every Windows user, and not every terminal.

Those two probes were missed three enumerations running because they use
`await import('node-pty')`, which no static-import grep finds. The comment now
tells the next reader to grep for `node-pty` instead.

The measurement that produced the wrong claim was taken by a standalone harness
that passed no `useConptyDll` and so defaulted into the branch it was not trying
to measure -- the same standalone-is-not-the-real-host trap #18601's own body
warns about, one level down.

Also refreshes the self-exit paragraph, which #18635 made stale. That leak is now
fixed for the desktop, and the note records why the fix cannot reach a Windows
relay. The fix is mostly native (`src/win/conpty.cc`) and this asset only rewrites
`lib/*.js`, and all three delivery paths stop short of Windows: pnpm patches do
not cross the SSH boundary; `MATRIX_SLOTS` in `build-orcad-prebuilds.mjs` has no
win32 entry; and the one relay asset that does patch native source and rebuild on
the host (`node-pty-1.1.0-master-cloexec-patch.cjs`) returns
`skipped:unsupported-platform` for anything but linux/darwin. #18635's flat self-exit relay numbers were measured
against a locally rebuilt binary, so they describe the relay code path on a
patched tree, not the tree a relay host installs -- the note says so explicitly
rather than leaving the next reader to conflate them.

Assertion and hashes unchanged: the relay must still release conin after the
console-list fork, and a patch sync must still not copy the early placement onto
the relay's branch, where it does cost +2 File and +1 Process per terminal. Only
the justification changes, plus the test name, which said "like the desktop
patch" where it meant "unlike the desktop patch placement".
2026-09-04 22:40:38 -07:00
Neil 8c0e7eccd0 perf(worktree): overlap preparation and prestart blank terminals 2026-09-04 22:35:42 -07:00
Neil f8ad8fb35f test(worktree): cover optimized creation call signatures
Preserve explicit branch adoption, WSL callback routing and sparse cleanup expectations.
2026-09-04 22:35:42 -07:00
Neil 41b520259e ci(package): retry apt fetches and docker builds behind the Ubuntu mirror (#18797)
The package job builds three Docker images whose apt-get update/install
hit archive.ubuntu.com with no retry, timeout, or mirror fallback. When
the mirror is mid-sync every build dies in one of three ways:

- per-package fetch stalls (~64 s each, `Ign:` lines) until the runner's
  10-minute docker build timeout fires:
  https://github.com/stablyai/orca/actions/runs/33935104546/job/101221447425
- `apt-get update` exit 100 with `Hash Sum mismatch` on
  noble-updates/restricted/Packages.gz:
  https://github.com/stablyai/orca/actions/runs/33935104546/job/101226099525
- `apt-get update` exit 100 with `File has unexpected size ... Mirror
  sync in progress?`:
  https://github.com/stablyai/orca/actions/runs/33935244026/job/101231083497

Each Dockerfile now retries `apt-get update` up to five times with
Acquire::Retries and a 30 s HTTP timeout, clearing /var/lib/apt/lists
between attempts so a half-synced index is never reused, and passes the
same acquire options to `apt-get install`. Each runner script retries the
whole `docker build` once when the first attempt fails or times out.
2026-09-04 22:24:18 -07:00
Jinwoo Hong ba4bbacd6b fix(relay-ops): align the cloud-data freshness bar with Cloud Monitoring publish lag (#18798) 2026-09-05 01:18:21 -04:00
Neil f2d3d0f620 perf(worktree): remove redundant creation and terminal startup work 2026-09-04 22:00:56 -07:00
Jinwoo Hong 436ef827dd fix(browser): present Electron's own user agent so Cloudflare Turnstile clears (#18749)
Orca rewrote every browser session's UA to look like plain Chrome by stripping
the Electron and app tokens. That rewrite is what Cloudflare rejects: a Chrome
UA that ships no client hints reads as a spoof and Turnstile returns 600010,
while the same binary on the same IP clears every challenge with its stock UA.
PR #885 added the rewrite to fix 600010 and was treating a symptom it created;
issue #11518 later found the same rewrite is what broke Google sign-in.

- Keep the stock Electron UA on every partition. The webRequest handler now only
  owns the host-scoped Google auth Firefox switch, which stays unchanged.
- Delete the anti-detection script. Measured on Electron 43: plugins are already
  a real PluginArray, window.chrome exists, and navigator.webdriver is false even
  with the debugger attached, so three of its four premises were wrong, and the
  overrides it installed (instance-level webdriver, non-native Permissions.query,
  stubbed chrome.csi/loadTimes) are themselves published bot signatures.
- Stop attaching a CDP debugger to every browsing guest. Only the auth-UA detach
  listener remains, because a detach clears Chromium's standing UA override.
- Stop sending Runtime.enable into cross-origin iframes when the agent bridge
  auto-attaches. The challenge widget is one, nothing reads iframe Runtime
  events, and the Runtime domain's serialization side effect is the documented
  Cloudflare CDP tell.
- Add a real-Electron test proving the wire identity: stock UA to ordinary
  hosts, Firefox with no client hints to accounts.google.com.

Verified in the dev build: dash.cloudflare.com/login no longer shows
"There was a problem with verification" and scrapingcourse.com's managed
challenge clears, both failing deterministically before.

Fixes #13822
2026-09-04 23:47:15 -04:00
Jinwoo Hong 040c3e5b32 fix(browser): match loading surfaces to the Orca theme (#18738)
* fix(browser): theme unavailable guest surfaces without recoloring pages

* test(browser): keep generated loading evidence out of the PR diff

* test(browser): freeze recovery clock during artificial attach gate

* test(browser): await painted content after network recovery
2026-09-04 23:31:25 -04:00
Jinwoo Hong 974acc901c fix(relay-ops): retry freshness-only preflight failures on the first same-cap wave too (#18778) 2026-09-04 23:13:55 -04:00
Brennan BensonandMerge Sim cb7f7dd11a fix(native-chat): tell old mobile builds why a structured chat is missing (#18756)
* fix(native-chat): tell old mobile builds why a structured chat is missing

A structured native chat started on desktop was simply absent on a paired phone
running any shipped App Store build. The host strips every `agent-session` tab
from a client that does not advertise `agent-session.structured.v1`, and no
released mobile build advertises it — so the chat had no representation at all
and no way to explain itself.

Keep the row and retitle it instead of deleting it. The shipped client does not
filter unknown tab types and renders whatever title the host sends, so an old
build now shows the chat's slot with a title naming the fix. Nothing is removed,
so the tab order, groups and layout it belonged to are left intact.

The prompt is keyed on the capability for that specific agent, not on the
combined policy boolean: a capable phone whose desktop simply has the experiment
off would otherwise be told to take an update that cannot help it. Claude rows
are prompted too — mobile cannot render them yet and a later build can, so the
message is true for that client as well.

Restore is no longer gated on the caller's capability. It stayed gated on the
host setting, which is what decides whether there is anything to reach at all,
but gating on capability left an old client with nothing to project after a
desktop restart: neither the chat nor the prompt.

Tab titles are capped at 128px on one line in every shipped build, so the string
is sized for ~15 characters rather than a sentence.

Prompted rows are visible rows, so the host now permits all five session-tab
mutations on them, close included. That is intended: a mobile close runs the same
teardown as the desktop's own Close button.

* fix(native-chat): keep fallback tabs safe and truthful

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-04 20:08:39 -07:00
Brennan BensonandMerge Sim 5a1acfec17 fix(native-chat): make document paths and links clickable in chat (#18712)
* fix(chat): make assistant file paths clickable

* fix(chat): tighten native file link handling

* fix(chat): link prose-joined relative paths separately

* fix(chat): preserve complete Unicode file links

* fix(chat): preserve links before sentence punctuation

* test(chat): align structured session link props

* fix(native-chat): harden generated file links

* fix(native-chat): reject reference-number false positives

* test(native-chat): align structured session parity

* perf(native-chat): bound file-link detection

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-04 20:03:51 -07:00
Jinwoo Hong e2b70a5eba fix(relay-ops): retry transient admin-endpoint failures in same-cap verify and rehome jobs (#18769) 2026-09-04 22:04:21 -04:00
Neil df7b1028dd fix: use fallback shell readiness and assert fish startup timing (#18755)
* fix: classify effective startup shell and pin fish delivery timing

* test: clarify fallback fixture and remove unused readiness wrapper
2026-09-04 18:48:17 -07:00
Jinwoo Hong 30d7542bc5 fix(terminal): stop a hidden pane's unmeasured 80x24 from overwriting a live PTY's size on reattach (#18706)
* fix(terminal): stop a hidden pane's unmeasured 80x24 from overwriting a live PTY's size on reattach

A pane that mounts while display:none (app relaunch or update with the
floating terminal panel closed, a non-active floating tab, any
background tab) cannot fit its container, so it reattaches with xterm's
default 80x24. Main wrote those placeholder dims into `ptySizes`
unconditionally, both before and after the attach. A daemon attach never
resizes the live session, so the real PTY stayed at its wide grid while
main's hidden headless model was created (or reflowed after renderer
hydration) at 80 columns. Every byte the agent emitted while hidden was
parsed 80 wide; reveal restored that image into the pane: rows clamped
at column 80 with CHA fill, the status bar interleaved into response
text, and for alt-screen TUIs the whole screen stuck in an 80x24 corner
until a real resize forced a repaint. Scrollback damage was permanent.

Fix, main-side only:
- Pre-attach: seed `ptySizes` only for a genuinely fresh session id, or a
  measured request with nothing cached. A hidden reattach writes nothing.
- Commit: on reattach, record the provider's proven grid
  (`attachedGrid`, set only by the local provider whose attach really
  resizes), then the reply's `snapshotCols/Rows` (the daemon emulator's
  grid), then the size main already held, and only then the request.
- Reflow an already-created model to that grid after the seed block, so
  bytes that arrived before the reply no longer leave an 80x24 model.

Both the ipc and runtime spawn paths take the same authority module.
Renderer and wire formats are unchanged; `PtySpawnResult` is main-internal.

Reproduced deterministically: close the floating panel with Claude Code
streaming at 211x57, kill only the Electron main process so the daemon
survives, relaunch. Main's cache read 80x24 against an applied 211x57
and the reveal snapshot was 80 columns wide; replaying the recorded
bytes through an 80-column emulator reproduced the field screenshot.
Relaunch with the panel open, and a fresh spawn, keep the wide grid.

* fix(terminal): commit the adopted-claim reattach grid and reject non-integer provider grids

Review follow-ups. The runtime spawn path's adopted-claim branch returned before the size
commit, so an adoption attaching to a live session kept whatever the caller requested; it now
commits and reflows like every other reattach. The grid validator requires integers so a
malformed provider grid falls through to the cached size instead of reaching xterm.

* fix(terminal): reflow main's headless model onto the committed grid for every spawn, not only reattaches

A hidden attach whose daemon restarted comes back as a fresh session, and the
pre-attach seed is now withheld for unmeasured attaches, so a live byte that
created the model at 80x24 before the reply would have kept it there forever.

* refactor(terminal): let the provider's reattach flag pick the adopted-claim grid source

* fix(terminal): derive the adopted-claim reattach flag once for the size commit and the reservation

The SSH relay's adopted reply carries no isReattach, so the size commit
would have taken the request while the reservation was told it was an
attach. Normalize once so both agree.
2026-09-04 21:24:54 -04:00
github-actions[bot] 38bde20121 Update README downloads badge 2026-09-05 00:57:58 +00:00
Jinwoo Hong 74ad08ec66 fix(relay-ops): accept monitor evidence from an ancestor commit with identical monitor code (#18754) 2026-09-04 20:54:40 -04:00
Jinwoo Hong 0f5f5e6979 fix(relay-ops): retry a failed MIG inventory read once before calling a cell's power state unknown (#18740) 2026-09-04 20:54:33 -04:00
Brennan BensonandMerge Sim 746a6b4870 fix(orchestration): fence the dispatch CLI preamble so it stops rendering as headings (#18718)
* fix(orchestration): fence dispatch CLI preamble

* fix(orchestration): keep optional preamble sections out of Markdown headings

The sub-dispatch and base-drift sections end with a bare rule directly under a
paragraph, which Markdown parses as a setext H2, so the Chat UI rendered the
section's last sentence as a heading. The unfenced sub-dispatch commands also
lost their angle-bracket placeholders to the raw-HTML pass. Fence those
commands like the main CLI block and put a blank line before each closing rule.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-04 17:49:03 -07:00
Brennan BensonandMerge Sim 8096cb2803 fix(native-chat): render Claude structured chat through the same UI as Codex (#18743)
* fix(native-chat): render Claude structured chat through the same UI as Codex

Structured native chat is one shared, agent-agnostic component tree, but three
Codex-hardcoded terms on that path made a Claude session render differently.

- `showTurnStatus` was `agent === 'codex'`, which gated the whole structured
  presentation: the live tool-progress row, the completion check and activity
  grouping (`structuredActivityUi`), and the Thinking / Working for N /
  Worked for N status rows. Claude fell back to the legacy compact chrome.
- `runtimeContext` was likewise Codex-only, so a Claude transcript rendered
  images as filename chips instead of previews.
- The composer's structured slash menu always served the Codex catalog,
  ignoring `agent`. That disagreed with the dispatcher, which does branch per
  agent: Codex-only tokens offered to Claude missed the command guard and were
  sent to the model as literal prompt text, where Codex shows an error.

The first two gates landed Codex-first (#17597, #18266) before the Claude
structured lane existed; they were rollout scoping, not capability limits.
Neither `useNativeChatTurnStatus` nor `useNativeChatImageRuntimeContext` has
any agent-specific logic. The menu now reads `structuredSlashCommands(agent)`,
the function the dispatcher already used, so both read one list.

No new rendering logic: two gates removed and one existing shared function
reused. There are now zero `agent === 'codex' | 'claude'` branches in any
native-chat component.

* docs(native-chat): update structured turn status contract

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-04 17:40:20 -07:00
Brennan BensonandMerge Sim 86cd327749 Answer the structured-session support probe without installing the host (#18695)
* Answer the structured-session support probe without installing the host

`getStructuredAgentSessionCreateSupport` called `ensureStructuredAgentSessionHost()`
before answering, so a read-only "can you create a Codex session here?" question
performed the create route's lifecycle work: the first install opens the durable
agent-session record store, attaches the PTY write-gate record lookup and starts
the orphan-child reaper.

Ask the pure predicate instead. `supportsCreate` on the installed host resolves to
`adapterSupportsCreate`, which for the Codex adapter is exactly
`agent === 'codex' && supportsCodexStructuredLocation(location)` — no adapter
instance is needed to answer it.

Nothing is lost: the create/attach route still installs via `ensureStructuredHostInstalled`,
and startup restoration still installs and reconciles when a store is already persisted.

* test(runtime): cover structured support probe parity

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-04 17:11:41 -07:00
Jinjing 2e80972450 chore: update in-app Android APK link (#18745) 2026-09-04 17:03:01 -07:00