Commit Graph
11970 Commits
Author SHA1 Message Date
Alex-wangyang 47b8408e28 fix(zcode): wait for composer before first worker dispatch (#23374)
* fix: wait for ZCode composer before first worker dispatch

* fix: preserve readiness across renderer terminal adoption
2026-09-27 22:02:19 -07:00
Jinwoo Hong d33541c335 test(mobile): repin the RPC recording corpus to main after #22951 (#23535)
#22951 re-recorded the corpus with baseline set to its own branch commit,
which the squash left unreachable from main, so the RPC recording pin job
fails. Repin baseline to main's tip and re-record every golden; only the
baseline header moves.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 01:00:23 -04:00
Jinwoo Hong c8f9a65ff8 test(e2e): keep the Kitty-arming app alive in the Option-composed spec (#23495)
The host grounds Kitty flags a finished command left armed, so a bare
printf arm flipped back to 0 and raced the flags poll. Arm with a live
cat foreground and pass per-test flags into setup instead of re-arming.
2026-09-28 01:00:04 -04:00
Brennan Benson 934a2d44a0 fix(codex): a native chat's thread opens on the model the chat chose (#23532)
* fix(codex): a native chat's thread opens on the model the chat chose

* fix(codex): a resumed thread keeps its own saved model, provider and effort
2026-09-27 21:54:37 -07:00
Brennan Benson 78771646af fix(mobile): show one review sheet at a time so the review screen never freezes (#22951)
* fix(mobile): close the review sheet before opening Send Notes

On iPhone, Review Actions > Send Unsent Notes and Review Complete > Send Notes
opened the Send Notes sheet while their own sheet was still on screen. iOS
cannot present a second native sheet until the first has unmounted, so Send
Notes never appeared and every later tap on the review screen was swallowed
until the app restarted. Send Unsent Notes now uses the action sheet's
closeBeforePress, and the Review Complete drawer opens Send Notes from its
onAfterClose, the same sequencing the action sheet already uses.

* test(mobile): drive Send Unsent Notes through the real action sheet

The overflow test only checked the closeBeforePress flag, and nothing tested that the
action sheet actually defers such an action until it has closed. Press the real row and
assert Send Notes opens only from the sheet's after-close callback.

* fix(mobile): ignore a close request on a drawer that is already hiding

Android Back during a drawer's close animation restarted the hide, which
cancelled it, so the drawer never unmounted and its invisible Modal kept
swallowing every tap. Send Notes, which now opens after the review sheet
closes, never appeared either.

* refactor(mobile): give the review screen one sheet state so sheets cannot stack

The review screen kept five independent sheet flags that any caller could
set at any time. iOS cannot present a sheet while another is still on
screen (even mid-close), so any overlap froze every tap. Two openers could
still produce one: Review Complete appearing after the mark-reviewed save
landed on a sheet opened meanwhile, and the Send Notes list reopening a
sheet the user had already dismissed.

The five flags become one reducer that mounts at most one sheet. Switching
sheets closes the current one and shows the next only after its drawer
reports it has finished closing; Review Complete waits for the user's
sheet instead of closing it; a late send list only fills a Send Notes that
is still shown or queued. The two per-call-site sequencers this branch
added are removed in favour of it.

* test(mobile): repin the RPC goldens for the review sheet state

The review-actions recording adapter now drives the screen's sheet reducer
instead of the removed Send Notes setter, which moves adapterSha256 on the
14 goldens recorded through it. baseline is repinned to the refactor commit
so the recorder's product fence matches; no recorded body changed (only the
baseline and adapterSha256 header fields move).

* fix(mobile): settle a review sheet that closed before it was ever shown

A drawer mounts only once a commit shows it, so a sheet closed or displaced in
the same batch it opened in (e.g. Review Complete landing in the same frame as
a tap that opens another sheet) never sends onAfterClose. The sheet state then
waited on it forever and every later sheet on the screen stayed queued.

Track which sheet's drawer a commit actually showed and settle a closing sheet
that never reached the screen instead of waiting for a close it cannot send.

* fix(mobile): keep a queued Send Notes when a late Review Complete lands

A second Mark Reviewed save resolving while Review Complete was closing
toward Send Notes reopened Review Complete and dropped the Send Notes the
user had just asked for. A background opener now yields to any queued
user sheet, including behind a closing sheet of its own kind.

* refactor(mobile): present review sheets through one keyed drawer

iOS cannot present a native Modal while another is still presented, even
during its close animation. The review screen now renders all five sheets
through one KeyedBottomDrawer that alone decides what is presented: a
request for a different sheet hides the current one, and the next is
mounted only after its hide finished and a commit without any Modal has
landed. A request replaced before it was shown is never mounted.

The screen's sheet state now records only what the user asked for
(`requested`, plus a background Review Complete in `deferred`), so the
presented-sheet bookkeeping, the per-drawer close callbacks and the
screen-side mounted-sheet inference are gone.

BottomDrawer becomes a constant-key adapter over the same drawer, so the
app has one mount/close lifecycle. A hide that finished just before a
reopen is now ignored instead of latching, which used to swallow the next
close and leave an invisible Modal eating taps.

* test(mobile): repin the RPC goldens for the keyed review drawer

The review action adapter now drives the requested-sheet state, which
moves that family's adapterSha256 on its 14 goldens; the repin to the
keyed-drawer commit moves `baseline` on all 787. No golden body moved.
2026-09-27 21:38:29 -07:00
Jinwoo Hong 52dc32f9ea fix(mobile): the page pushes its terminal frame into the document from RN layout (#23079)
* fix(mobile): keep one mounted terminal frame under every session branch

The frame was keyed so react-native-web would attach its onLayout: a View
that gains onLayout after mount is never observed, and unkeyed the frame
reused the loading branch's View. One contentFrame View now wraps every
branch and carries the handler from the first mount, so no key is
needed. terminalFrame stays on the terminal branch, so nothing else is
clipped.

The parity pin moves: the key string leaves, one View and one style
reference arrive.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): push the page terminal's box into its document from RN layout

The page's document sized itself through a ResizeObserver on its host,
which duplicated RN layout and needed three rules in fit-scale: skip a
0-wide box, skip the last fitted box, and forget that box on any fit
request. The host View's onLayout now pushes into the mount, which is
the page's counterpart of the WebView's window resize.

react-native-web still lays a display:none screen out as 0x0 and its
return as the old box, so the mount treats neither as a change and keeps
answering the last real box while hidden, as a WebView keeps its size.
A fit asked for while hidden therefore lands at once, and fittedBox and
all three rules go. The zero-width wait in applyFitScale stays for a
host that mounts under a hidden screen before its first layout.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold a page terminal's fit while its host is hidden

A fit asked for under a covering screen read the last box, but where
xterm cannot measure cells on a display:none host (its DOM measure, with
no OffscreenCanvas) the retry loop ran out and committed scale 1, and
the show that followed was not a change, so 1 stuck.

The page's rect now says when the host is hidden, and the document holds
any fit asked for then as fitPending instead of committing. The mount
reports the same box coming back as 'shown', distinct from 'resized',
and the document runs a held fit on it and otherwise does nothing, so
pan and zoom still survive a plain hide and show. Native's window
resize reports 'resized' and is never hidden.

The mount also seeds its box from the host, so a first layout that
lands before the mount is not later mistaken for a resize. The hidden
text-scale case now asserts the resized column count, which stale
cells miss.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read the page terminal's grid once xterm has sized its cell

The init render check read the grid the moment .xterm-screen existed,
but xterm sizes its one-cell helper textarea only on a cursor move or
resize, after the replay drains. Under full-suite load the read won
that race and measured a 0-wide cell. In the failing runs the fit had
committed at 390/560 on the cell-width gate, so the page was right and
the read was early. It now waits for a sized cell: 8/8 under four-way
parallel load, where it was 4/8.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin that a held page terminal fit runs once

A held fit that lands on show must be spent: a second hide and show
re-running it would reset the pan and zoom the user set in between. The
case now hides and shows again and expects no new transform, which a
commit that stops clearing fitPending fails.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold the refit a narrow hidden text-scale change owes

A text-scale change whose new cells leave fewer than MIN_FIT_COLS in
the box skips the grid resize and returned before any fit. Base cleared
fittedBox there so the next box refit; with that gone, a hidden host
shown at the same box kept the old scale under the larger font. The
branch now asks for the fit while the host is hidden, which holds it
until show. A shown host, and every native one, returns as before.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): fit a text-scale change too large to resize the grid

A visible viewport too narrow for the new cells (280 px at 200%, 18
columns) skipped the resize and returned without a fit, so the larger
text overflowed the fit made for the smaller cells. The branch now fits
unconditionally; applyFitScale already holds the fit while hidden.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 00:06:35 -04:00
Brennan Benson 2119730ec0 fix(terminal-wait): unattended launches report Claude's trust dialog instead of timing out or typing into it (#22927)
* fix(terminal-wait): recognise Claude's workspace trust dialog as a blocking prompt

Claude's first-launch "trust this folder?" dialog parks the cursor above its
options, and the host's line tail drops the lines below it, which are the only
ones the trust matcher knew ("trust this folder", "Enter to confirm"). An
unattended Claude launch into a fresh folder therefore waited out its whole
budget and reported a timeout instead of the blocking prompt. The dialog's
opening question ("... one you trust?") survives in the tail, so it is now
recognised, pinned by a captured transcript of the real dialog.

* fix(terminal-wait): read the rendered screen for blocked prompts on the tui-idle poll

Claude's workspace-trust dialog parks the cursor on its highlighted option with a
cursor-up, and the host line tail deletes every retained row below the cursor, so
"Yes, I trust this folder" and "Enter to confirm" never reach the blocked-prompt
detector. In a live launch the dialog also arrives in 1024-byte reads, and the tail's
plain path blanks each line that ends in a carriage return before the newline, so
the opening question does not survive either. An unattended launch waited out its
whole budget and reported a timeout.

The runtime already feeds every PTY chunk into its own headless emulator. The tui-idle
poll now also runs the existing blocked-prompt rules over that emulator's visible
screen (no provider or host round trip), after the tail checks and before the
quiet-foreground idle fallback. The "one you trust" phrase is dropped: it only matched
when the whole dialog arrived in one chunk, wrapped away on narrow panes, and could
match prose.

Replays three live Claude 2.1.280 captures through the runtime: the dialog in one
chunk and in 1024-byte reads, a 60-column pane, and the dialog answered with "Yes",
which must report ready rather than blocked.

* fix(terminal-wait): skip the rendered-screen blocked check while the agent reports working

The screen check runs on every tui-idle poll, and the automation observer holds a tui-idle
wait open for a whole agent turn. A working Claude whose screen showed dialog wording (a diff
of the detector, say) was reported blocked where main kept waiting. The dialogs only the
screen reveals are start-up ones painted before any title, so a working title now vetoes it.

* fix(terminal-wait): settle weak tui-idle evidence only after a clean screen read

A shell auto-title (oh-my-zsh's `claude`, fish's `claude <cwd>`) names Claude before
its workspace trust dialog paints, and the line tail loses that dialog. The wait took
the bare name as rest, settled ready, and the launch typed its brief into the dialog.

Every tui-idle settle site now asks one evaluator for a verdict: blocked, strong ready,
working, weak ready or pending. The pre- and post-registration checks and the
title-change resolvers settle only blocked or strong ready; weak ready is left to the
poll, which settles it only once the rendered screen shows no blocker. A name-only
Claude title is held to the same quiet window as Codex and Devin, since Claude
announces rest with its own explicit title.

* test(serialize): record the answered Claude trust capture's known serializer divergences

The captured answered-dialog transcript added by this PR is replayed by the
serialize round-trip suite and diverges at 13 checkpoints. It diverges
identically on origin/main and on the pre-#22586 addon build (13 both-fail,
0 regressions): the live SGR pen leaks into the alt buffer, and an alt buffer
first entered after a shrink keeps hidden scrollback. Pin the count like the
other known captures and correct the comment, which called these upstream.
2026-09-27 20:59:02 -07:00
Jinwoo Hong cdf9f6e938 fix(mobile): native desktop-mode keyboard lift stays on screen; metrics carry the row pitch (#23070)
* fix(mobile): the session hears a row-pitch change in the keyboard metrics

Carried from 2eab7633013 and 62113b7ac30 with the `rowPitch` metric they need (from
f7d78405c7d, without its page-only lift). The document reports the drawn row pitch, cell height
times fit scale, and reports again when a fit commits a new scale. The metrics handler compared
four fields and dropped an update that moved only the pitch, so desktop display mode kept the
phone's 15 px pitch; it now compares every field.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): native desktop display mode lifts only what the keyboard hides of the grid

Carried from 097f5d8d33f, native rule only. Desktop display mode draws 40 rows about 270 dp
tall, and a lift by the covered strip (312 on Pixel_API_37) moved the whole grid under the
header. The lift is now also capped by the hidden strip of the drawn grid. A sweep shows it
equals base's arithmetic whenever the drawn grid fills the frame, which phone mode always does.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the desktop-mode anchor and the fit commit's re-emit by behaviour

A caret mid-grid in desktop display mode lifts only enough to clear its row, which the anchor
term decides; without it the lift was the whole hidden strip and no test noticed. The fit
commit's re-emit is now asserted by committing a fit and reading the pitch it reports, rather
than by the order of two lines in the source.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): give the fit test's viewport the whole rect type

The tests typecheck ratchet refused the double: the seam's rect carries `left` and `top` too.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): report keyboard-avoidance pitch after a pinch release

A pinch moves the drawn row pitch through userScale. A release onto the same
text-size preset skipped the refit, so no metrics were emitted and a write made
mid-gesture left the lift reading the transient pitch. applyTextScale now says
whether it scheduled the refit, and the release emits when it did not.

The text-scale frame also emitted a transient pitch after its resize, before
the fit it schedules commits and emits the final one. Drop that emit so a
preset change reports once.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): report text-scale pitch only through the fit commit

The pinch release emitted metrics itself whenever applyTextScale said no
refit was scheduled. A release onto a preset a settings change had just
applied saw an unchanged font size and emitted the pre-refit pitch, one
frame before the refit reported the real one.

A same-size applyTextScale now schedules the fit too, so every text-scale
change reports through commitFitScale. A pending refit supersedes that fit
by token. applyTextScale returns nothing and the release emits nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-27 23:55:21 -04:00
Brennan BensonandClaude bfe476f922 fix(native-chat): a message is accepted, then delivered (#22821)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 20:38:00 -07:00
Jinwoo Hong 992bad5375 refactor(agent-resume): give each resume-note reader its own named check (#23497)
* fix(terminal): stop a resume note from blanking a live remote agent pane

A host-mirrored remote terminal pane holding a client resume note with
origin 'live' and state 'done' (the idle anchor kept after every finished
turn) loaded blank and looped attach/disconnect several times a second.

- Mirrored web-terminal tabs ignore client resume notes at load; the host
  answers liveness at connect, so a note must not divert the attach.
- The empty-reattach retire rule never fires for remote PTYs: a remote
  disconnect only closes this viewer's stream, so the retry lands on the
  same live PTY and loops.
- The handler gets its own sleep-evidence check with the pre-#16308
  meaning, so the wake sweep's widening no longer leaks into it.

* test(agent-resume): pin how a finished turn's idle anchor reaches parking and the hibernation wake

#16308 let a live done note count as passive for the wake sweep. The park
exemption and the suppressed-exit wake shared that check without being
reviewed for it: a finished turn's idle anchor now lets its hidden tab park,
and any suppressed exit over it arms an in-place resume on reveal. Pin both
before the check is split so the refactor cannot change them silently.

* refactor(agent-resume): give each resume-note reader its own named check

isPassiveCompletedHibernationEvidence answered two different questions after
#16308 widened it for the wake sweep. Replace it with checks named for what
each reader asks, same answers as before:

- isFinishedTurnOwingNoResume: the wake sweep, background wake, preserved-pane
  ownership and the park exemption (a finished turn owes no resume).
- noteArmsHibernatedPaneWake: the suppressed-exit wake arm and the mobile wake
  latch, which must agree with each other.

The reattach handler keeps isHibernationDoneRecord from #23491.

* refactor(agent-resume): name the activation check for what it decides

isFinishedTurnOwingNoResume also returned true for worktree-sleep done
notes, which a pane mount does resume from (#9648). Its callers all ask
whether workspace activation leaves the note alone.
2026-09-27 23:06:36 -04:00
Jinwoo Hong a8797138ef fix(terminal): stop a resume note from blanking a live remote agent pane (#23491)
A host-mirrored remote terminal pane holding a client resume note with
origin 'live' and state 'done' (the idle anchor kept after every finished
turn) loaded blank and looped attach/disconnect several times a second.

- Mirrored web-terminal tabs ignore client resume notes at load; the host
  answers liveness at connect, so a note must not divert the attach.
- The empty-reattach retire rule never fires for remote PTYs: a remote
  disconnect only closes this viewer's stream, so the retry lands on the
  same live PTY and loops.
- The handler gets its own sleep-evidence check with the pre-#16308
  meaning, so the wake sweep's widening no longer leaks into it.
2026-09-27 22:52:06 -04:00
Brennan Benson 0b2dd0a99b fix(runtime): let the phone end a terminal stream by the request that opened it (#23006)
* fix(runtime): register phone terminal subscriptions when the request arrives

* fix(runtime): don't let an already-aborted terminal subscribe take the stream slot

* fix(runtime): let the phone end a terminal stream by the request that opened it

* fix(mobile): describe the relay sibling check accurately on this base
2026-09-27 18:51:14 -07:00
Brennan Benson bc7a17fef2 fix(runtime): register phone tab-list streams when the request arrives (#23045)
* fix(runtime): register phone tab-list streams when the request arrives

* fix(runtime): keep a worktree-wide tab unsubscribe off later subscribes

Tab-list subscribes now register when they arrive, so an older phone's
worktree-wide session.tabs.unsubscribe (no request id) could end a same-worktree
subscribe that arrived while the unsubscribe was still resolving the worktree.
Capture the registration version when the unsubscribe is dispatched, as
terminal.unsubscribe already does, and sweep only streams registered by then.

* test(runtime): cover a request abort during tab-list stream setup
2026-09-27 18:27:18 -07:00
Brennan Benson bf9194126e fix(mobile): release streams whose ready arrives after a replayed cancel (#22945)
* fix(mobile): release streams whose ready arrives after a replayed cancel

* test(mobile): cover a replayed browser stream replaced before its ready
2026-09-27 18:21:30 -07:00
Jinjing 5a8237a7a5 Browser tab placement review (#23462)
* Implement source-following placement for browser tabs

- Add afterTabId and executionHostId parameters to track and anchor new tabs after their source
- Resolve source browser pages to unified tab wrappers for background link opens
- Implement anchor-based insertion that respects pinned tab boundaries
- Stage source-adjacent rows before host RPC for paired browser creation
- Ensure duplicated tabs and background links place directly after their source
- Add comprehensive test coverage for placement scenarios across single/split groups

* Preserve tab strip scroll anchor when tabs are added

Keep the viewed tab stable on screen when other tabs are inserted around it.
Records the active tab's on-screen position before insertions and restores it
after, so the user's focused tab doesn't jump unexpectedly. Matches the behavior
of VS Code and Chrome.

* Refine browser tab placement: fix host handling and sort-order gaps

- Avoid reordering unified tabs when the computed order matches the current order
- Do not substitute execution host for browser tab wrappers; let createUnifiedTab apply its active-workspace fallback instead
- Always apply sort values after tab insertion to prevent sortOrder gaps from preview replacement
- Fix import path for paired browser tab creation and clarify registration comments

* refactor(tests): use paired-browser-tab-creator registry for browser tab

- Replace direct web-runtime-session mock with paired-browser-tab-creator pattern
- Add test-rig file naming convention to localization audit skip list
2026-09-27 18:08:09 -07:00
Neil dd90b3183f fix: identify partial-clone repositories from the correct remote (#23504) 2026-09-27 18:05:13 -07:00
Brennan Benson 6c56c0a3dc fix(browser): a failed SSH route keeps its card while the host redials (#23465)
* fix(browser): a failed SSH route keeps its card while the host redials

A browser route that already failed swapped its "SSH connection unavailable"
card for "Connecting" on every dial of its host, and re-ran prepare once per
dial cycle, so the card flickered and its buttons detached mid-click. Only a
route that is still preparing now waits on a dialing host; a failed route
keeps its card until the host actually connects, which re-derives it.

Retry and Try anyway land on preparing together with the new attempt, so a
press while the host dials waits for the connect instead of starting a
prepare the effect immediately cancels.

The escape-hatch e2e now makes the host truly unreachable before the
disconnect; it passed before only because the card stayed latched over a
host the terminal had already reconnected.

* fix(browser): an unrouted SSH route waits for a dialing host too

Only a failed or ready route is exempt from the host wait; a route that just
became routed (or still shows another target's page) has no answer for this
target, so it must not start a prepare that its own preparing write cancels.

The escape-hatch spec now reads the settled failure from the renderer store:
main's ssh:getState drops the entry on disconnect and on a failed connect, so
its status is null there and never matches the failure pattern.
2026-09-27 17:57:21 -07:00
Brennan Benson d6336be8db ci(e2e): run the SSH browser route e2e when its source changes (#23498)
A change to the SSH workspace browser route, its gate card, or the host-connection
phase it waits on selected no e2e specs, so the Docker SSH browser spec never ran
on the PR that changed it.
2026-09-27 17:51:16 -07:00
Brennan Benson 7a3735fcd2 fix(native-chat): let a multi-question ask record list its questions (#23451)
* fix(native-chat): let a multi-question ask record list its questions

A cancelled or pending structured ask with several questions showed only
"Asked: 2 questions", and its questions appeared nowhere in the transcript.
The row's subject now carries the question texts instead of a count, and
the row unfolds them as a list below its toggle unless the record already
lists them with their answers.

* fix(native-chat): keep a question list's open state off its first question

A pending Codex group is keyed by its first question, which becomes its own row once answered. Opening the list left that row expanded, and folding it could drop the toggle under the pointer. The list and a lone question now remember their open state separately.
2026-09-27 17:21:47 -07:00
Brennan Benson 89cf55dfc8 fix(agent-launch): report a launch that failed before spawning as failed, with its cause (#22913)
* fix(agent-launch): settle a launch that failed before spawning as failed, with its cause

A launch into an existing workspace whose terminal create threw before the
spawn request left this process (agent disabled, no launch command, runtime
unavailable) created nothing, yet agent.launchReplay recorded and answered it
as agent_session_operation_unknown. createTerminal now reports when it hands
the spawn to the pty controller; a failure before that point settles the
ledger row as failed and returns the original error. After the request leaves,
the outcome stays unknown: an SSH or daemon spawn whose reply was lost may
still have started.

* test(agent-launch): expect the spawn-dispatch hook on the launch's terminal create

* test(agent-launch): drive the pre-spawn failure with a missing launch command

A disabled agent is moving to a check made before either launch route runs,
so the tests use a failure that stays inside the terminal build.

* fix(agent-launch): keep the not-started verdict on the launch, not the shared error

A failed pane spawn rejects the same error object into the spawner (after its
request left) and into a concurrent create waiting on that pane (before its
own). Marking the error object globally let the waiting launch's verdict clear
the spawner's, recording a launch that may have started an agent as failed.
The launch now owns its dispatch tracker and carries the decision on its
execution error instead of re-deriving it from the error.

Also names the test's launch parameter type for the anti-slop audit.

* refactor(agent-launch): move launch failure classification into its own module

Rebasing onto the caller-selection change took agent-launch.ts past its line
limit; the failure-code helpers are a self-contained concern.
2026-09-27 16:37:04 -07:00
Brennan Benson da57f47353 fix(native-chat): say a refused chat write in plain words, and keep it with the write it describes (#22999)
* fix(native-chat): say a refused chat write in plain words, and keep it with the write it describes

A send's refusal now lives on the queued message and goes when that message is sent again or delivered, so a resend that succeeds after an agent restart no longer leaves a red line under the composer. Stop, answers, option and goal changes report a refusal once as a toast; a conversation command answers inline. One shared table turns every refusal code into copy with a next step, on desktop and mobile, instead of showing the host's diagnostic text.

* test(native-chat): pin the refused-then-restarted send in the order the live app sees it

* fix(native-chat): tell the phone to send a refused message again, not to press Retry

The phone puts a refused message back in the composer and has no Retry
control, so "Retry to send it again." named something that is not there. A
phone send now reads "Your message was not sent. Send it again."

Also types the rejection-cause test's empty submissions without a cast.

* fix(native-chat): keep why a message failed as a fact, and say only what is true

A queued message saved the words of its failure to local storage, so the copy
lived in users' data. It now saves the fact: the refusal code (with the host's
words only for a failed restart, which the host writes for people), the
provider's rejection reason, or that the host could not be reached. The Retry
row chooses the words when it shows the message. A saved failure this build
cannot read is dropped, and the row says only that the message was not sent.

The words are true for every host path behind each code:
- A refusal whose code does not say why (a cleared conversation, a pending
  question, a provider's own rejection) says only what did not happen, with no
  next step that could repeat the refusal.
- An unsettled owner no longer claims the agent was restarting; it says Orca
  could not confirm which agent process owns the chat.
- A code from a newer host says only what did not happen.

Desktop translates each sentence whole, with the shared English as the
fallback the phone shows as is, so the two never say it differently. The phone
no longer shows transport or host text when a request fails without a
refusal.

* fix(native-chat): never show a host refusal message, including a failed restart's

Every refusal code has at least one host path that writes its message for a
log or carries a marker. A replayed refusal, for one, says "Operation <id> was
already refused: <code>." because the ledger stores no message. A failed
restart's message also embeds the resume's own refusal text, which can be the
ledger's. So no code's message is safe to show, and the failed-restart
exception goes. The Retry row now says "The agent couldn't restart. Your
message was not sent." with no next step, because some restarts need a new
chat. The cause is still in the chat's own status row. The saved failure keeps
only the code.

The test lists one such emitter per code, so a code that later becomes safe to
show has to be argued against that list.

* fix(native-chat): offer only a next step that works, on the phone too

A message the agent turned away on the phone said "Retry to send it again", but the phone has no Retry control; it now says "Send it again." like every other refused phone message, built from the same sentences as the other notices.

The capacity refusal said Orca was handling too many requests "for this chat" and offered a retry. The limit is counted across every chat over a day, so an immediate retry would likely be refused again. It now says so, with no next step.

* chore: restore pnpm-lock.yaml

A local install rewrote it and it was committed by mistake.

* fix(native-chat): say only what is true for every host path behind a refusal

No refusal notice says how to try again any more: the control that sent
the write already does that, and a sentence per code was false for some
of the host situations the code covers. The phone keeps "Send it again."
where a resend under the same operation id can go through, and an older
host still gets "Update Orca, then try again."

A cause is named only for codes whose every emitter means it. The owner
codes and identity_required now say only what did not happen, the
outcome-unknown sentence no longer claims the write itself is in doubt
(a send behind an unconfirmed rewind is refused outright), and a request
that failed without a refusal is recorded as `failed` rather than
`unreachable`, since most such failures are host or compatibility errors.
The census test now pins the cause allowlist and the absence of retry
sentences, and the unused sentences leave every catalog.

* fix(native-chat): don't say a chat write failed when its request may have run

A desktop Stop, answer, option, goal or conversation command whose request threw
said the write did not happen, for every error. A timeout or a lost connection can
come after the host ran the write, so a timed-out /compact said "The command didn't
run." while it ran. Only an RPC error the host answers before running the method
proves that; any other failure now says Orca couldn't confirm what happened.

The phone already drew this line; both clients now read it from one shared rule.

* fix(native-chat): say a refused background-task stop is about the task, not the agent

A background-task stop goes through agentSession.cancel, so a refusal of it
said "The agent wasn't stopped." although the agent was never asked to stop.
The write kind now reads the cancel's scope and names the task or tasks.

* fix(native-chat): say a Stop refused over a moved-on question did not stop the agent

A Stop pressed while a question or approval is pending names that prompt, so the host can refuse it
as already resolved or stale. The notice then spoke only about the question and never said the agent
kept running.

* test(native-chat): check every refusal notice cell against one spec

Replaces the per-rule loops with one table of every failure (each refusal
code, a failed or unconfirmed request, and a code from a newer host) by
every write kind, and five rules each cell must keep: a certain refusal
says that write did not happen, once; a write that may have run never
says it did not; only "Update Orca" and the phone's resend where it can
work say to try again; a cause is named only for allowlisted codes; no
cell is empty or shows the host's text.

* fix(native-chat): the launch path drops a message's old failure when it sends it again

The first message of a new chat is sent by the launch path, which staged it
without removing the failure saved by an earlier attempt. A message that failed
through the outbox and was then refused again through the launch path showed the
earlier reason in its Retry row. Both paths now stage through one shared step
that removes it.
2026-09-27 16:34:42 -07:00
NeilandClaude b61797dce3 fix(agent-trust): bound the remaining Codex trust writes a launch waits on (#23380)
* fix(agent-trust): bound the remaining Codex trust writes a launch waits on

#23148 capped the trust write on the agentTrust:markTrusted IPC handler and the
worktree-remote startup path, but three call sites still awaited
markCodexProjectTrusted with no bound: the main-process worktree startup used by
orchestration worker-start, Codex quick-launch preparation, and Codex session
resume preparation. All three share the per-config.toml lane with hook installs
and app-server trust grants, which has no cap on queue depth, and the SSH writer
adds a resolveHome round trip plus unbounded SFTP over a possibly half-open
link. A wedged lane left the user clicking "start agent" with nothing happening
and no error.

Each site now routes through the existing awaitAgentTrustWriteWithinDeadline
helper with unchanged semantics for its two current callers. Abandoning at the
deadline never cancels the write, writes nothing else, and never substitutes a
local fallback, so the workspace stays untrusted and Codex raises its own trust
prompt. The await still sits ahead of the runtime-home resolution and hook
repair that the PTY spawn waits on, so the ordering #23148 established is kept.

Reliability only; there is no speed gain. The win is that a launch cannot hang.

Verified with 2,755 agent-trust/Codex/startup/runtime tests, the 53-test
reliability-gate command, the gate manifest check, typecheck:node, changed-code
quality, oxlint and oxfmt. Red-green: with the three sites reverted to a bare
await the new bounding tests hang to the 30s vitest timeout (3 fail/14 pass);
restored, 17 pass.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(reliability-gates): stop the trust-preflight gate claiming a bounded launch

The gate's performance budget said a stuck predecessor "degrades to the
agent's own trust prompt instead of an indefinite caller wait." That is false
at two of this branch's three new sites. In startup/codex-launch-preparation.ts
the next awaited call after the abandoned write is
codexHookService.prepareRuntimeHomeForLaunch; in
startup/codex-session-resume-launch.ts it is installForLaunchPrep or
refreshRuntimeUserHooksForLaunchPrep. Both reach
runExclusivelyForRuntimeAndSystemTrustConfig, which takes the same
runExclusivelyForCodexTrustConfig queue - FIFO per config.toml, no depth cap
and no timeout - on the runtime home and on ~/.codex/config.toml that
markCodexProjectTrusted itself takes. A wedged lane still stalls those two
launches one step later, with the abandoned write holding its queue slot ahead
of the hook step. Only markLocalWorktreeTrusted has nothing after its Codex
branch and so is bounded end to end.

The budget, invariant, oracle and coverage notes now say what is true: all five
markCodexProjectTrusted call sites are bounded, so no trust write hangs a
caller indefinitely, but the bound is on the write and not on the launch. A new
knownGaps entry records the residual stall and that the new launch-prep tests
mock codexHookService wholesale, so nothing in these suites can catch it. The
already-complete-write case stays labelled a timer control, not a red.

Wording only; no production code, test or count changed. The cited command
still reports 6 files / 53 tests, and the gate manifest check passes for 138
gates.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(reliability-gates): name the first unbounded re-entry, not the hook-service call

The trust-preflight gate cited codexHookService.prepareRuntimeHomeForLaunch as
the immediate next step after the abandoned write in
startup/codex-launch-preparation.ts. Tracing the module, the first awaited step
is ensureRealHomeHooksIfSelected. It is conditional: only when the target is not
WSL and runtimeHome.isHostSystemDefaultRealHomeSelected(launchEnv) is true does
it call ensureRealHomeCodexHookState, which chains behind that module's own
serial ensure promise and then acquires runExclusivelyForCodexTrustConfig
directly on ~/.codex/config.toml - the same lane markCodexProjectTrusted's inner
acquire takes. When the real home is not selected that call returns without
awaiting the lane, runtimeHome.prepareForCodexLaunchAsync is synchronous on the
non-WSL path, and only then is prepareRuntimeHomeForLaunch the first re-entry.
So the gate could miss a launch stall one step earlier than the one it named.

startup/codex-session-resume-launch.ts had the same omission: its hook branch
takes ensureRealHomeCodexHookState when the resume home is ~/.codex, and only
otherwise installForLaunchPrep or refreshRuntimeUserHooksForLaunchPrep. The
coverage notes also understated the mocking - codex-launch-trust-write-deadline
mocks ensureRealHomeCodexHookState as well as codexHookService and defaults the
real-home selection to false, so neither lane is driven by any suite here.

The performance budget, coverage notes and the residual-stall knownGap now name
the first re-entry on each path and say when it applies, rather than only the
later hook-service call.

Wording only; no production code, test or count changed. The cited command
still reports 6 files / 53 tests, and the gate manifest check passes for 138
gates.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(reliability-gates): scope the trust-preflight invariant to the local writes

The invariant claimed every Codex trust write a launch waits on is bounded or
abandoned at a deadline. The entry's own knownGaps contradicted that: the remote
write reached through markRemoteWorktreeTrusted in
runtime/runtime-worktree-agent-startup.ts is awaited with no timer, and it is a
Codex write, not an adjacent one - it reads TUI_AGENT_CONFIG[agent].preflightTrust,
which is 'codex' for the codex agent, and markRemoteAgentWorkspaceTrusted then
branches into markRemoteCodexProjectTrusted. Its one production caller,
markRemoteWorkspaceTrustedForAgent, is reached with a connectionId from seven
launch entry points; five await it directly before the createTerminal that spawns
the agent, and two gate the launch command they hand back for their caller to
spawn. It chains a session.resolveHome round trip plus SFTP realpath/read/mkdir/
write, none with a timeout, over a possibly half-open link - the same hang this
branch bounds locally.

The invariant now claims only the five local markCodexProjectTrusted sites and
names the remote exception inline instead of leaving it to knownGaps. Audited the
rest of the entry against the code in the same pass: performanceBudget said "no
trust write can hold its caller indefinitely" (now scoped to those five); the
"all five awaited Codex trust writes" tally is now "all five awaited local" and
points at the unbounded sixth; the remote gap no longer calls that path merely
"preset-agnostic and out of Codex scope"; oracle and coverageNotes now record
that nothing here drives markRemoteWorktreeTrusted, since the remote-preset
suite only checks what markRemoteAgentWorkspaceTrusted writes, never how long it
may take; and surfaces gains the remote writer that suite actually covers.

Re-verified the rest rather than assuming it: the launch-prep re-entry chain,
the FIFO no-cap no-timeout trust-config queue, the IPC handler capping both its
remote and local writes with no cross-branch fallback, and that
upsertProjectTrustLevel and upsertProjectTrustLevelInContent remain the only
producers of a project trust_level, so no further writer needs covering. The
already-complete-write case stays labelled a timer control, not a red.

Wording only; no production code, test or count changed. The cited command still
reports 6 files / 53 tests, and the gate manifest check passes for 140 gates.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 16:30:30 -07:00
Jinjing 29847641ab test(e2e): fix failures and improve stability (#23480)
- Narrow toolbar width and set window minimums for consistent testing
- Add node_modules symlink to fixture for ESM import resolution
- Exercise manual paging and fix button selector
- Preserve repo filters in reveal workflow
- Adjust timing strategy and increase test timeout
2026-09-27 16:17:06 -07:00
Neil fb51e6cb4c Speed up Markdown document discovery with bundled ripgrep (#23306) 2026-09-27 16:16:20 -07:00
Jinwoo Hong 433986fa3b fix(runtime): read Codex readiness from the live screen (#23475)
* fix(runtime): read Codex readiness from the live screen

Codex 0.157 runs in embedded mode when Orca passes `-c model_reasoning_effort=…`,
shows a startup warning, and repaints its header by cell diff
(`ESC[5;3Hdir ESC[5;7Hctory:`). The tui-idle body check read the line-folded
tail, which drops those cursor moves and reads `dirctory:`, so worker-start
timed out at agent_readiness while Codex sat idle at its prompt.

For Codex (or unknown) panes whose live emulator screen shows the Codex header,
the screen now decides readiness: `model:` and `directory:` present and neither
still `loading`. Otherwise today's text rules apply unchanged. All six tui-idle
satisfaction sites share one helper so they cannot disagree.

Fixes STA-8628 / #23241.

* refactor(runtime): read the Codex screen lazily and only for Codex panes

Pass the screen as a thunk so non-Codex panes never build the grid, hoist the
pane agent lookup, and share the unblocked-ready check with the Muse rule.

* fix(runtime): let the Codex screen only add readiness

A grid out of step with the PTY (size mismatch or a resize mid-paint) can garble
Codex's header. Keep every verdict the text rules give today and consult the
screen only when they say not ready, so no flow that settles today can stop.

* fix(runtime): read Codex readiness from the header box only

Chat below the header can mention "OpenAI Codex" or "model: loading", so the
screen rule now reads model/directory/loading only inside the header box.

The serialize round-trip sweep picks up every runtime fixture; record the
pre-existing header-border attribute divergence the new Codex 0.157 recordings
expose, which this change does not touch.
2026-09-27 18:48:23 -04:00
NeilandClaude 161bdf93c3 fix(bench): load benchmark modules under test through jiti (#23482)
`pnpm run bench:terminal-partial-escape-tail` and
`pnpm run bench:worktree-refresh-churn` both died at startup with
ERR_MODULE_NOT_FOUND. Each entrypoint static-imported a `src/` module with an
explicit `.ts` extension, but bare `node` type-stripping cannot resolve the
extensionless relative specifiers *inside* that module's graph
(`terminal-partial-escape-tail.ts` -> `./terminal-escape-introducer`,
`worktree-catalog-reconciliation.ts` -> `../../../../shared/structural-value-equality`).

Routes both through jiti, matching the four benchmarks that already load `src/`
TypeScript that way (`pty-source-ack-boundary`, `locale-collator-sort`,
`worktree-base-pending-marker`, `wsl-git-shell`). Adding the extension at each
import site was the alternative, but `config/tsconfig.node.json` does not set
`allowImportingTsExtensions`, so a `.ts` specifier in `src/shared` fails
typecheck with TS5097 -- and no file under `src/` uses that shape today.

Developer tooling only; no production code changed.

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 15:34:38 -07:00
Brennan Benson 993183afd7 fix(terminal): every explicit terminal close commits through one main transaction (#22929)
* fix(terminal): every explicit terminal close commits through one main transaction

A renderer save cannot shrink terminal membership once main owns a repo's
topology, so desktop tab and pane closes, CLI split-pane closes and mobile
split-pane closes only became durable when the killed process's exit retired
the surface. A close whose kill failed or threw, or whose exit was never
certified, came back after a reload.

Every close now reaches closeTerminalSurface: the renderer sends an explicit
intent for user and cleanup closes, the CLI and mobile split-pane closes commit
the pane after their stop, and the headless and relayed mobile closes reuse the
same commit. A failed flush keeps the in-memory removal and no longer cancels
the kill. Exit retirement is unchanged.

* fix(terminal): tell the desktop renderer to drop a split pane main closed

A CLI or mobile close of one pane in a split commits the pane in main, but the
desktop kept showing it until reload when no exit arrived to remove it. The
close now sends a leaf-addressed notice: a mounted pane closes by leaf id, and
a parked tab collapses its stored layout. Addressing by leaf makes the notice
and the renderer's exit handling no-ops after each other, which replaces the
numeric pane-id notice that could close the whole tab when the exit won.

* fix(terminal): a pane close never widens into a whole-tab close

A leaf-addressed close fell through to the whole-tab close whenever main's layout no longer
held that leaf as one of several. Main's exit handling retires an exited split pane from the
saved layout, so closing that pane afterwards (the exited-pane overlay's Close, or a CLI close
whose stop delivers the exit first) removed the whole tab, live sibling included, and the
next renderer save could not restore it. A pane close is now a no-op unless its leaf is in a
multi-pane layout.

Also updates two mobile split-close assertions to expect the leaf-addressed notice, and adds a
test that a relayed mobile close of a renderer-listed tab still reaches the renderer's pin guard.

* test(terminal): cover the PTY-handle branch of a CLI split-pane close

The existing CLI split test resolves its handle through the renderer graph, so the branch
that closes a runtime-owned pane by its PTY handle had no test failing without its commit.

* fix(terminal): a CLI pane close with an unconfirmed stop closes only that pane

`orca terminal close <handle>` on one pane of a split used to close the
whole tab, live sibling included, whenever that pane's stop could not be
confirmed (for example an unreachable SSH host). An unconfirmed stop is
unverifiable, not a reason to drop siblings: the close now commits only
that leaf, tells the renderer to drop that leaf, and leaves the owed kill
to the controller's existing SSH pending-kill path.

On a host where no renderer lists the tab, main now also removes the
closed pane from the paired-client snapshot (with its retirement proof),
since no exit may arrive to do it.

* refactor(terminal): one resolver decides whether a pane close becomes a tab close

Every explicit close now states its target as `{kind:'tab'}` or `{kind:'pane', leafId}`; no
optional leaf id silently means the whole tab. Main resolves a close it started in exactly one
place, reading the copy of the tab's panes its layout owner holds (the renderer-published layout
for tabs the desktop renderer lists, main's session layout otherwise). Only `last-pane` escalates,
through the existing tab path so the renderer's pin guard still runs; an unknown pane never widens.

- The CLI and phone paths drop their per-site sibling counts for the resolver.
- The notifier splits into a tab-only close and a leaf-addressed pane close.
- The headless tab closer takes a parent tab id, so a pane row cannot reach it.
- A phone close of one pane on a host with no desktop window now stops and closes only that pane.
- A phone close of one pane with no live process record closes that pane, not its tab.

* fix(cli): an unverifiable stop says the close happened

`orca terminal close` still exits 1 when the process stop cannot be verified, but its message now
says the terminal was closed and names the host's reason, instead of "close failed". It promises
that the kill retries on reconnect only when the SSH relay itself never answered the stop, the one
case a recorded kill order backs.

* fix(terminal): a phone pane close commits even when its kill fails

A paired client's close of one pane threw `terminal_close_failed` before committing anything when
the controller reported the kill failed, so the pane stayed. The kill is now best-effort, as it is
for a whole-tab close: the pane's removal always commits and the failure stays on the PTY's
liveness verdict.

* fix(terminal): a pane close widens only when a copy shows it is the last pane

The close resolver read an owner copy that records no panes as "the tab has
one pane", so a CLI close of one pane of a split, addressed while the
renderer listed the tab before publishing its panes, closed the whole tab.

Every copy now counts only if it records at least one pane, read in the
owner's order with the published rows as the last fallback, and a pane
close widens only when a copy lists that pane as the tab's only one. An
unsplit tab whose saved layout predates its pane still closes: its
published row names the pane.

* fix(cli): promise a kill retry only when the host recorded the kill

The close receipt inferred "the kill retries when the host reconnects" from
the stop reason's text, which a new transport message or a reworded error
would silently break.

An explicit close now records the replayable kill order when its stop goes
unconfirmed, before sending the follow-up kill (whose own failure is
recorded only once its RPC settles), and reports that on the receipt as an
optional `pendingKillRecorded`. The CLI promises the retry only from that
field, so an older host, which never sends it, gets no promise.

* test(pty): justify the controller cast the recorded-stop tests extend

* fix(terminal): parse the close target with typed narrowing

The low-evidence lint gate rejects Reflect.get and broad object parameters,
which failed static analysis. Narrow with 'in' checks instead and cover the
boundary parser's accept and reject cases.

* fix(terminal): a desktop tab close is not refused by a split that bound while it waited

The renderer has already removed and killed a tab it closes, so its close intent now skips the
owner fence phone and CLI closes use. Before, a split pane whose binding was admitted between the
close request and its durable write made main refuse the close, and the tab came back on the next
launch whenever its processes did not exit.

* chore(terminal): note that closedByLayoutOwner goes away once main owns the terminal layout

* test(terminal): reload the close-intent fixture through the SQLite profile store

Main now requires a SQLite profile-state authority for a writable Store, so the save-and-reload
close tests build and reopen their store through the shared SQLite test harness.
2026-09-27 15:29:22 -07:00
Brennan Benson f72079bd21 fix(mobile): send typed question answers as structured answers (#23458)
* fix(mobile): send typed question answers as structured answers

The phone packed a typed answer into agentSession.respondToQuestion's
`optionId`, a field capped at 1024 characters, so a long typed answer was
refused and never reached the agent.

The phone now builds per-question answers for single and grouped questions
and sends them as `answers` when the host advertises
agent-session.question-answers.v1, falling back to the packed `optionId`
for older hosts. An answer too long to pack is sent as `answers` while the
host's support is still unknown, and on a host known to predate the field
the phone asks the user to update the computer instead of sending an
answer the host must refuse. The host features the phone negotiates for
structured sessions (prompt cancel, question answers) travel as one
required object from the shared capability probe instead of one optional
boolean each.

* test(mobile): cover a long typed answer on a grouped question and the default question id

* fix(mobile): keep the chat controller's host support argument optional

* fix(mobile): refuse a grouped answer too long for an older host on the step that overflows

Packed grouped answers only grow, so an overflow on an earlier step could never reach an older host. Refusing it only at the final submit left the long text folded into the draft with no way to shorten it short of cancelling the question.

* fix(mobile): refuse a grouped step that leaves no room for the rest on an older host

* refactor(mobile): check the packed answer length once, at send
2026-09-27 15:17:38 -07:00
Jinwoo Hong 618a8b0758 fix(terminal): keep a dead app's input modes recoverable after a refuted proof (#23474)
A command's armed input modes were demoted at the first OSC 133;D whether or
not the foreground proof confirmed, and the alternate-screen trigger was spent
the same way. A full-screen agent's nested command shells leak their own D onto
the main PTY; the refuted proof for that stray D used up the only trigger, so
the app's real death later went unrecovered.

Every D now re-asks while a command's mode (or the alternate screen) is still
up; only the confirmed ground or the app's own disable clears it. Proofs that
can never succeed (wsl.exe, no Windows job reads) already refute immediately
without spawning anything. The accepted cost is one process read plus a bounded
output hold per D while a command-owned mode stays up and the proof keeps
refuting: a live agent leaking D, a stopped job, a nested subshell, or an rc
that execs another shell.
2026-09-27 18:17:04 -04:00
Jinwoo Hong 10807d8b45 test(e2e): stabilize terminal shortcut Kitty and Ctrl+C tests (#23472)
Shift+Enter: the host now grounds keyboard modes a finished command left
armed, so the test's one-shot `printf '\033[>1u'` was reset to 0 before the
poll. Keep the command alive via `read` while asserting, then let it pop
its own flags on Enter.

Ctrl+C: each test opens a fresh tab, and SIGINT during shell startup can
kill the shell and close the tab. Wait for a ready prompt before
interrupting. The existing post-interrupt assertions are unchanged.
2026-09-27 18:16:49 -04:00
Jinwoo Hong 0547b32224 test(usage): add long-context token fields to the Codex fixture (#23478)
#18759's test fixture predates #22360 adding required long-context token
fields to Codex session and location breakdowns, so main's typecheck fails.
2026-09-27 17:52:43 -04:00
leilei3167 838769d73e fix(linear): accept team keys that start with a digit (#23423)
Fixes #23422
2026-09-27 14:30:45 -07:00
Mr. ZengandNeil ef987e42d7 feat(analytics): persist local usage session identities (#18759)
Co-authored-by: Neil <neil@stably.ai>
2026-09-27 14:29:19 -07:00
Neil 32668c3417 perf: preserve matching Accounts settings sections during search (#23192)
* perf: preserve matching Accounts settings sections during search

* refactor(settings): route Accounts' section stack through SettingsSectionStack

Replaces the pane-local index-keyed wrapper with the shared SettingsSectionStack so the
search-stable identity rule lives in one place, and restates the read/subscription-count
oracles as node identity plus non-growing counts so a future freshness fix cannot read as
a regression. Applied on the branch head; the rebase onto main is still outstanding.
2026-09-27 13:51:56 -07:00
Jinjing eaa6fb71b0 Tab opening ranking (#23435)
* Add literal-basename file matching for tab opening

- Distinguish literal filename matches from fuzzy matches
- Prioritize search for short tokens (1-2 Latin characters)
- Sort literal matches naturally instead of by fuzzy score
- Reserve search option even when literal matches fill results

* Keep literal-basename file matches in tab entry options

Ensure literal-basename file matches appear in tab creation suggestions
as fallback options when exact-basename matches exist, improving file
discovery for users.
2026-09-27 13:50:12 -07:00
Neil a4b60c2ea3 test(windows): stop the NSIS capability probe failing on a slow PowerShell cold start (#23381) 2026-09-27 13:24:48 -07:00
Neil 3f1bf68c84 perf: preserve matching General settings sections during search (#23177) 2026-09-27 13:21:54 -07:00
Brennan Benson bdaf9a3ed3 fix(agent-launch): move only the requesting client's view to a launched tab (#22914)
* fix(agent-launch): move only the requesting client's view to a launched tab

A launch into an existing workspace from a paired client (a phone, or a
desktop client of a remote server) activated a new chat for every client and
the host, and recorded no selection for the caller's terminal. It now
publishes the chat without activating it for everyone and records the new tab
as the caller's own selection, the way that client's own tab tap does
(session.tabs.activate with caller navigation). In-process callers and
workspace-creating launches are unchanged. Selecting the tab is bookkeeping:
a failure there never fails the launch.

* fix(agent-launch): select a launched tab for its caller the way a create does

The launch recorded the caller's selection through the tab-tap path, which
refreshes the PTY inventory (twice for a chat) on the reply path and treats a
non-ready pane as a wake gesture that respawns its agent. Record it through
the same caller-navigation step session.tabs.createTerminal uses after its own
create: no host focus, no refresh, no materialize.

* test(agent-launch): prove a launched tab's selection spares subscribed clients and never respawns

* test(agent-launch): pin that the host's own desktop window still activates a launched chat

The desktop renderer reaches agent.launch as a runtime client with no paired device, so it must keep the host-wide activation rather than be treated as a paired caller.
2026-09-27 13:17:18 -07:00
Brennan Benson aeb96ab0e2 fix(native-chat): close only the reopened chat when main closes its tab (#22922)
After /clear the current chat keeps the window tab id derived from its first
session, so reopening that first session from history lands in a suffixed tab.
Main closed chat tabs in the window by re-deriving that id from the session,
which named the current chat instead of the reopened one: closing the reopened
tab (on the desktop or from a phone) also closed the current chat and stopped
its Claude.

Main now sends its own tab id, as it already does for every other tab kind, and
the window translates it through the host-to-window tab mapping it keeps for
mirrored chat tabs. The same translation lets a host-directed focus reach the
reopened tab. The close handoff marker is keyed on the id actually sent.
2026-09-27 13:10:59 -07:00
Brennan Benson 9795563695 fix(runtime): register phone terminal subscriptions when the request arrives (#22948)
* fix(runtime): register phone terminal subscriptions when the request arrives

* fix(runtime): don't let an already-aborted terminal subscribe take the stream slot
2026-09-27 12:35:52 -07:00
Brennan Benson 484598b18e fix(terminal): retry first-terminal seeding until a decision applies (#22919)
* fix(terminal): retry first-terminal seeding until a decision applies

The seeding effect marked a workspace as attempted before its async
activation check finished, so a StrictMode double run, an input change
mid-check, or a 'blocked' result left the workspace marked with no
terminal. Mark it only once an uncancelled, unblocked result is applied;
a rerun shares the gate's in-flight check.

* fix(terminal): read the closed-last-terminal row when the seed decision applies

The seeding callback runs after an async activation check, so a row captured
at render time can be stale by then. Read it from the store at decision time
and drop it from the effect dependencies, which no longer need to restart the
check when the row appears. Add tests for two separate empty checks and for a
last terminal closed while the check runs.
2026-09-27 11:32:41 -07:00
Brennan Benson d972bac38a fix(workspace): a restored workspace paints before its terminals reconnect (#22810)
* fix(workspace): a restored workspace paints before its terminals reconnect

After a restart the workspace content area stayed blank — no tab strip, no
pane — until the whole startup chain (SSH reconnect, PTY reconnect, legacy
worker recovery) had finished, even though the restored tab model had been in
the store for seconds. Only terminal panes need that chain, because a pane
binds a PTY on mount.

The workbench now mounts the active workspace from its hydrated tab model, and
holds its terminal tabs unadmitted (an empty admitted-tab restriction with no
deferral entry) until startup restoration has published PTY ownership; the
activation plan then replaces the hold. Chat, browser and editor panes mount
immediately.

Carried from #22293 (922c1c2626), renderer terminal files only.

* fix(workspace): track the startup terminal hold explicitly and keep its tabs watched

The startup hold was inferred from the restriction map's shape and released only
when a new hold was installed, so a workspace left for no workspace mid-startup
stayed mounted with no terminal panes and no background watchers. Track the held
workspace explicitly, release it on any pass where it is no longer held, apply it
after prune, and count held tabs as parked so their bells, titles, and
completions are still observed during the window.

Carried from #22293 (3fd421c41f), renderer terminal files only.

* test(workspace): cover the split surfaces' held-tab watcher wiring and fold the held branch

A held worktree now falls through to the existing else, which already resets the activation marker.

* fix(workspace): browser and editor panes wait for a connecting SSH host

A restored SSH workspace now mounts its panes before the host reconnects.
Derive one per-worktree host phase from the published SSH state, treating a
target startup restoration has not dialed yet as connecting. The browser route
holds its prepare until the host connects and re-derives a failed route on
connect or a new connection generation; the editor shows the connecting state
instead of a dropped connection and reloads once the host connects.

* fix(workspace): a held terminal tab shows that it is restoring

While startup restoration holds terminal panes unmounted, the visible tab's
slot was blank. Fill the slot of a visible held tab, including one created
during the hold, with a restoring state until its pane mounts.

* refactor(workspace): every pane reads one shared host-connection signal

The terminal host state now reads its SSH status from the same per-worktree
signal as the browser route and the editor, instead of resolving it on its own.
The signal carries the resolved status (an undialed target during startup
restoration reads as connecting), the owning environment, and a connection
epoch that changes on every reconnect, so no pane re-derives the no-entry case
or the reconnect generation. A nested target whose runtime cannot be seen is
reported as unverifiable rather than unavailable.

* fix(workspace): the terminal names the published SSH status from the shared signal

The shared host signal now carries the published status separately from its
phase. Only the phase reads a target startup restoration has not dialed yet
as connecting, so the terminal's reconnect overlay and error ownership keep
reporting what the host published, while the browser and editor still wait.
Tests pin that window for the terminal and pin that a host the client cannot
verify never holds the browser or editor.

* fix(startup): a degraded boot releases terminal startup restoration after reconnect

The degraded recovery path's successful reconnect set workspaceSessionReady
but never terminalStartupRestorationReady; only its failure branches did. Every
reader waiting on restoration then waited for the whole session: no fallback
terminal, no structured tab sync, the activation gate timing out, and an SSH
target never dialed at startup reading as connecting forever. Release the flag
once the reconnect finishes, unless a newer startup pass has taken over.
2026-09-27 11:31:11 -07:00
github-actions[bot] 6c75837750 Update README downloads badge 2026-09-27 18:30:40 +00:00
Brennan Benson b8ef5fcbc2 fix(native-chat): let the "Asked" row unfold a long question (#22941)
* fix(native-chat): let the "Asked" row unfold a long question

The question row clipped long questions to one line with no way to read
the rest. The row is now a toggle that wraps the full question in place.

* fix(native-chat): open a clipped question below its toggle, selectable

The first cut put the whole question inside the toggle button, where
Chromium will not select text: dragging or double-clicking in it selected
nothing and toggled the row instead. The open state also lived in the row,
so scrolling the row out of the windowed transcript, or answering the
question, folded it again.

- The full question now opens in a paragraph below the button, outside it,
  so it selects and copies like any other prose.
- The toggle is offered only when the question is actually clipped; a
  question that fits stays plain, selectable text as before.
- Open state is kept in the transcript's disclosure store, keyed by the
  message, so it survives windowing and the pending-to-answered remount.
- The chevron also shows on keyboard focus and on touch screens.
2026-09-27 11:18:13 -07:00
Jinjing 0a876b2f5a feat(activity): reveal threads in floating terminal workspace (#23239)
* feat(activity): reveal threads in floating terminal workspace

When selecting an activity thread from the floating terminal workspace, automatically open the floating panel if closed. Improves split pane focus handling to preserve active panes during thread reveals.

* fix(activity): enable floating terminal before revealing threads

Floating tabs persist even after disabling the feature, but the panel
ignores toggle events while disabled. Enable the setting first so that
subsequent toggle commands take effect.

* Use useLayoutEffect for floating terminal toggle binding

Enable-then-toggle callers dispatch events immediately, which passive useEffect rebind can miss. useLayoutEffect ensures the handler is bound synchronously before the next paint.
2026-09-27 10:48:48 -07:00
OrcaWinandm4air 27b823f934 ci: compile the E2E CLI once for all consumers (#23384)
* ci: share compiled CLI output across E2E consumers

* ci: preserve CLI setup and old-ref fallback for shared artifacts

* docs: record shared E2E CLI benchmark evidence

* docs: include final CLI reuse timing range

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 02:02:29 -07:00
Neil 069701446b fix(persistence): report a discarded cross-host repo update instead of failing silently (#23379)
`Store.updateRepo` takes an optional `hostId` that must be the row's own
`executionHostId` stamp. When it is anything else the lookup matches no row and
the method returns `null` — the same value a deleted row returns — so the write
is dropped with no error, no log, and no failing test. #22421 shipped exactly
that: identity enrichment addressed the probe host (`local`) instead of a
`runtime:` row's stamp, and every write was discarded for a full release.

Resolve the row through `findRepoRowForHostScopedWrite`, which logs once when the
id exists but only under other host stamps, naming the requested host and the
stored ones. The guard itself is unchanged: a matching host still writes, and a
client-local probe still cannot repair a peer's row.
2026-09-27 01:31:24 -07:00
OrcaWinandm4air c15f082031 ci: build independent Electron targets together for E2E (#23378)
* ci: reuse parallel Electron targets for E2E builds and guard cache action setup

* test: recognize the top-level cache repository preload

* docs: record E2E build timings and exact output parity

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 01:17:27 -07:00
Neil 17690e6b9a style: settle oxfmt 0.70 drift and stop formatting vendored licences (#23377)
The oxfmt 0.65 -> 0.70 bump landed without a repo-wide reformat, so 36 files
already in the tree no longer matched what the new version emits. Anyone running
`pnpm format` picked all of them up alongside their own change.

Also excludes `resources/licenses/**`: `oxfmt --write .` was rewriting the
vendored PCRE2 licence, turning its `*` redistribution bullets into `-`. Third
party licence text has to be reproduced verbatim, so formatting must not touch it.
2026-09-27 01:14:53 -07:00
93d8b1f042 fix(ssh): complete keyboard-interactive MFA prompt handling (#15588)
Honor SSH keyboard-interactive prompt echo and empty responses, reuse login
passwords without replaying rejected values, and stop cancelled or stale
credential requests from continuing authentication or restoring the cache.

Original implementation: Junho Kim (#8750).
Port and follow-up work: Allen (#15588).

Verified with 2,773 SSH/credential tests, full typecheck, changed-code quality,
a production Electron build, real-socket MFA fixtures, and rendered UI checks.

Fixes #8622

Co-authored-by: Junho Kim <arkimjh@illinois.edu>
Co-authored-by: microdaery <microdaery@gapp.nthu.edu.tw>
2026-09-27 01:12:29 -07:00