Commit Graph
12079 Commits
Author SHA1 Message Date
Neil ceffecf12a fix(setup): remove wait-for-setup helper text (#23799) 2026-09-28 22:49:03 -07:00
Neil f6324f242a chore(ci): stop auto-filing community PRs onto the project board (#23796)
Removes the Track Community PRs workflow. Community pull requests will no
longer be added to project stablyai/13 automatically.
2026-09-28 22:15:36 -07:00
Neil 2ea3fb1d46 perf(ci): take advisory unit-selection evidence off the gate (#23776)
selection_evidence is continue-on-error on both the job and its comparison step,
so it can never fail a PR -- it downloads the shard reports, compares selection
against the full results and uploads a review artifact. But a caller's
`needs: test` waits for every job in the called workflow, so living inside
unit-tests.yml it held verify for ~36s after the last shard finished.

It moves to its own reusable workflow called as a sibling, so it still runs on
every PR and still uploads its artifact, but verify no longer waits for it. It is
deliberately absent from verify's needs, and a contract test pins both that and
its advisory status so it cannot drift back onto the critical path.

Measured on a recent run: the shards finished, then selection_evidence ran 36s,
then verify 3s. Only the last of those gates anything.
2026-09-28 22:13:35 -07:00
Neil 7445feaca8 fix(terminal): only diagnose disk exhaustion from capacity errors
Merged after required checks passed.
2026-09-28 22:07:04 -07:00
Jinjing 18bd9cb1f6 test(status-bar): cover re-probing a roomier level after the tightest level's fit moves (#23793) 2026-09-28 21:47:23 -07:00
Jinwoo Hong c8af48d8a4 fix(runtime): settle a quiet Codex composer as ready on every version (#23765)
* fix(runtime): settle a quiet Codex composer as tui-idle on every version

Codex 0.158 dropped `model:`/`directory:` from its startup header, which both
Codex readiness rules require, so `worker-start --agent codex` timed out; an
idle Codex pane after a turn also had no readiness signal once the header left
the screen.

Generalize the Muse tier-1b lane into a quiet-ready-screen lane: a Codex (or
agent-unknown) pane whose live screen shows the empty composer placeholder,
no `to interrupt)` status row, no header `loading`, and no dialog wording in
its live window, settles once the stream has been quiet for the tui-idle
quiescence window. Additive only: the tier-1 rules and the Muse rule are
unchanged. Fixtures: codex 0.150.1-0.158.0 captures at 120x40, including
chunk-timed turns.

* refactor(runtime): anchor the Codex quiet lane to the empty composer line

Move the Codex screen rules into codex-terminal-readiness.ts and the quiet-screen
body beside isKnownReadyPromptBody. The composer rule now matches only the
`› Ask Codex to do anything` line and drops its dialog markers: every Codex
dialog replaces the composer, and an answer ending "Would you like to…?" above a
live composer must not hold the lane forever. The quiet lane checks quiescence
before reading the screen. Trim the redundant startup and untimed turn fixtures.

* fix(runtime): read Codex's busy row above the composer and scope the lane to codex panes

* fix(runtime): read only Codex's live status row above the composer
2026-09-29 00:39:01 -04:00
Brennan Benson c30c8f9d77 fix(native-chat): a message Claude folds into its running turn no longer splits the turn (#23621)
* fix(native-chat): treat a folded mid-turn Claude replay as a receipt, not a turn boundary

A message sent while a Claude turn runs is folded into the running turn by the
CLI and replayed mid-turn with the client uuid. The replay-driven opener treated
that replay as a new turn: the running turn was marked interrupted and the
'Worked for' bar split. The dispatch waiter now captures the open turn at write
time (sentDuringTurnId, volatile); a replay that adopts the client uuid while
that exact turn is still open settles delivery and opens no boundary. Plural
result user_message_uuids settle each waiter under its own uuid; every other
relation (miss, replaced turn, provider-resumed root, idle write, fresh replay
uuid) keeps the opener path. Replay/turn resolution split out of the dispatch
module to hold the line budget.

* fix(native-chat): state the fold receipt's uuid-adoption limit as unmeasured

The receipt admits only a replay that adopts the client uuid. The comment and
test named a fresh replay uuid as "the CLI starting the queued send's own
turn", but the measured miss case adopts the client uuid too, so adoption does
not tell a fold from a later turn. Say what the rule actually is: the measured
fold shape qualifies, and an unmeasured fresh-uuid replay keeps the opener path
it always had. No behavior change.

* fix(native-chat): decide Claude fold receipts from the provider's own request cycle

Measured (p3 captures, CLI 2.1.280): the CLI folds any send that arrives while
its request cycle runs — including one written before the first replay — and it
announces every new cycle (sequential turn, queued turn, background wake,
/compact) with a root system/init; each result names the sends its cycle ran in
user_message_uuids. So the fold decision now reads provider state: an adopted
replay while the open turn's cycle is still live (no root init since it opened)
is a delivery receipt. The send-time bookkeeping (sentDuringTurnId) is removed;
it missed the measured early-steer fold and guessed at what the provider states
outright. A lost result is covered by the init staleness mark, and CLIs below
the per-turn-init version floor (or with no reported version) keep today's
opener path via an explicit gate on the init frame's claude_code_version.

* fix(native-chat): read Claude fold membership from the cycle's own work

A finished background task revises its row before the wake cycle's init,
so the wake turn opened ahead of that init and was marked stale by it: a
send folded into the wake still split the bar and interrupted the wake
turn (captured p3-background-wake order). A send replayed after the
current cycle's first root work (send echo or model output) is the fold;
the cycle's first send is its opener. Init and every settle start a
cycle with no work.

* fix(native-chat): adopt the CLI version from any init frame, not only the startup proof

Live sessions prove the session from a SessionStart hook frame that arrives
BEFORE system/init, so startup facts read a frame with no claude_code_version
and the fold receipt's version gate never passed: the real app still split the
bar and marked the running turn interrupted while every fixture test — whose
harness proves startup with the init frame itself — stayed green. The version
is now adopted from whichever init frame carries it, at the same site that
already adopts the per-turn model report, and the startup proof never clobbers
a version a real init already reported. Pinned twice: a fixture run whose
startup proof is the SessionStart hook frame, and a real-CLI fold test (skipped
without a signed-in CLI) that fails on the unfixed branch at the interrupted
assertion and passes with the fix.

* test(native-chat): run the real-CLI fold test in the config dir it probed

The connection strips an inherited CLAUDE_CONFIG_DIR and inherited auth
vars, and with no launch env the pin compared against that same inherited
value and emitted nothing, so the fold test always ran against ~/.claude
whatever CLAUDE_CONFIG_DIR the availability probe checked. It also relied on
the user's own SessionStart hooks for the live proof order and on their
permission rules for the Bash steps; an isolated home hung at startup.
The test now launches with the probed home and env auth, and pins a
SessionStart hook and the sleep permission itself.

* fix(native-chat): drop the CLI-version gate and prove the fold against full adapter captures

The fold rule stands on frame-derived facts alone: an adopted client uuid
(the capability check, read off the replay itself) while the open turn's cycle
has done root work. The claude_code_version floor guarded only an unobserved
triple fault — an adopting CLI without per-turn init AND a lost result — and
its cost was a silent latch that already fired once live; all cliVersion
plumbing is removed. Mid-turn auto-compaction was measured (forced via
CLAUDE_CODE_AUTO_COMPACT_WINDOW): it emits status/compact_boundary and a
synthetic continuation but NO root init, so the init cycle reset stays and the
capture is pinned. The fake connection now defaults to the live startup shape
(SessionStart hook proof; one init when the first command starts a cycle;
capabilities on the initialize result), with init-at-startup an explicit
unmeasured opt-in. Full frame streams recorded through Orca's own adapter
against the real CLI are committed and replayed verbatim, asserting the
provider's membership fact (each result's user_message_uuids) rather than
design internals.

* test(native-chat): drop the removed CLI-version gate from fold test comments

Also reattaches the fake harness's initProof doc to initProof (it had
landed on contextUsage, replacing that field's own doc) and renames a
plural-result test whose title described a retired waiter it never
creates.

* test(native-chat): name the fold test's settings for their role

* chore: take main's lockfile back after the merge

* test(claude): route the real-CLI probe through the spawn chokepoint; give the slow-init test a live startup report

The shared real-CLI availability probe moved out of a test file, so the
child_process and CLI-runtime-pairing ratchets now scan it. It spawns through
runProcessSync with the CLI paired to its own node, as the structured launch does.

The provider-started test's CLI default now reaches startup the live way:
get_settings reports it, since system/init only arrives with the first command.
2026-09-28 21:22:47 -07:00
Brennan Benson 50a8ef18e4 fix(native-chat): keep a turn's bar on the prompt that opened it (#23573)
A message sent while a structured turn runs appears in the transcript at
once, so "the newest user message" is not the running turn's owner. The
live "Working for" bar moved to the mid-turn message and counted from the
earlier prompt's start, and a send queued behind the running turn counted
its wait twice: once in the previous turn and again from its own send.

Derive both from the host's turn records in one ordered pass:
- The running turn's bar belongs to the user message its lifecycle row
  names (resolved exactly as settled timing resolves it). A message sent
  mid-turn gets no bar until its own turn opens; a send folded into the
  running turn never gets one. Surfaces fall back to the latest user
  message only when the host names no opener.
- A turn counts from its send, but never before the previous turn in the
  journal ended (its recorded end, else its row's last host revision),
  capped at the turn's own start. The same origin feeds the live counter
  and the settled duration.

Desktop and mobile share the derivation; no wire, host, or storage change.
2026-09-28 21:22:22 -07:00
Brennan Benson c5fc0c6f26 fix(ci): keep a squash-merged RPC recording pin reachable through its pull request (#23720)
* fix(ci): keep a squash-merged RPC recording pin reachable through its pull request

Main's "RPC recording pin" check has been red since #22762: that branch pinned
the recording corpus to its own commit 03995ae, and the squash-merge left that
commit out of main's history. Every behaviour-change squash did the same, and
each needed a hand-made repin PR to clear it (#23565, #23535, #23046 and more).

The guard now accepts a pin that is either in this history or in the head of the
pull request whose squash wrote it into the manifest. It finds that pull request
from the `(#n)` subject of the commit that added the pin and fetches
`refs/pull/<n>/head`, which GitHub keeps after the branch is deleted. The
reproduce step uses the same lookup, so it can still check the pinned tree out.

* fix(ci): give the recording pin lookup room to walk a blobless clone

In CI's blobless clone, `git log -S` fetches the manifest's blobs one commit at a
time, a few seconds each. Under the 30 s process default the walk was killed after
a handful of manifest commits, which main's history already exceeds (up to 7
manifest commits between a pin landing and the next pin change), and the guard
then failed with an empty "Could not find the commit that pinned" error. The
lookup and the pull request fetch now carry explicit budgets and say when they
timed out.

The not-an-ancestor instruction now names the pull request whose head was
checked, or says the commit that pinned it names none.

Adds the two merge-preview shapes the guard runs on: a branch opened after a
squash resolves main's pin through the squash's pull request, and a branch whose
rebase dropped its own pinned commit fails on its pull request instead of on main.

* fix(mobile): tell a missing recording pin apart from product drift

After a squash the pinned commit can live only in its pull request's head, so a
clone that never fetched it makes `git diff --quiet <baseline>` exit 128. The
recorder reported that as "Product sources or lockfile differ from the pinned
main baseline", which sends the developer to repin a tree that may match. It now
prints git's error and the command that fetches the pin.

* fix(ci): ask GitHub which pull request holds a squash-dropped recording pin

The recording pin guard found the pull request that keeps a squash-dropped
pin by walking main's first-parent history for the commit that wrote the pin
into the manifest and reading "(#n)" off its subject. A merger who edits the
squash title loses the number, and the push to main turns red anyway. That
already happened on main: of the 22 squashes that left a pin outside main's
history, #21674's title had no "(#n)".

The guard now asks GitHub for the pull requests associated with the pinned
commit (GET /repos/{owner}/{repo}/commits/{sha}/pulls) and, for each in turn,
fetches refs/pull/<n>/head and accepts only when git proves the pin is an
ancestor of that head. GitHub only nominates candidates, so a wrong answer can
fail the guard but never pass it. The endpoint named the right pull request
for all 22 historical cases, #21674 included, and names none for commits a
force-push orphaned.

This removes the pickaxe walk, its 600 s budget and its lazy blob fetches in
a blobless clone, the first-parent subtlety, and the subject regex. A revert
that restores an older pull-request-only pin now resolves too, because the
lookup is by the pin itself rather than by the commit that last wrote it.

CI passes the job token to both guard steps and grants the job
pull-requests: read. Local runs work without a token on this public repo and
send GITHUB_TOKEN or GH_TOKEN when set. A failed lookup throws with the HTTP
status, and names the rate limit when an unauthenticated call is refused.
2026-09-28 21:17:07 -07:00
Brennan Benson e899809ff8 fix(mobile): unsubscribe session tabs by request on the direct connection (#22943)
* fix(mobile): unsubscribe session tabs by request on the direct connection

* fix(mobile): hold a direct session tabs unsubscribe until the first snapshot

The desktop registers a tab-list stream only as it emits the first snapshot. A direct
unsubscribe sent before that found nothing, and with per-request addressing no later
worktree-wide sweep collects the late stream, so it kept its desktop listener until the
socket closed. Hold the unsubscribe until the snapshot arrives, as the relay connection does.

* test(mobile): cover a held session tabs unsubscribe whose subscribe fails

* fix(mobile): keep the session tabs hold within the registry line budget after merging main

Move the pre-snapshot hold into the session tabs stream module, note that only older hosts need it, and give the unsubscribe test the real registration version now that a worktree-wide unsubscribe spares later streams.
2026-09-28 21:15:39 -07:00
Brennan Benson 2ae7c00e84 fix(opencode): keep OpenCode 2 panes Working across plugin reloads (#23700)
* fix(opencode): keep OpenCode 2 panes Working across plugin reloads

OpenCode 2 disposes and re-sets-up every plugin whenever its plugins dir
changes, while sessions keep running. The status plugin published a final
Idle on dispose, so a pane read Done mid-turn. Orca also rewrote the plugin
file on every PTY spawn, so opening any terminal triggered that reload.

Dispose now releases the factory's bookkeeping without publishing a
verdict; the next lifecycle event settles the pane, and Orca's ended-process
reconciliation still retires panes whose agent exited. The plugin file is
written only when its bytes differ.

* fix(opencode): skip rewriting an unchanged plugin in the SSH relay install too

The relay's canonical-config install still unlinked and rewrote the status plugin on every OpenCode launch over SSH, which restarts every plugin in a remote OpenCode 2 server. Share one install-currency check (lstat + the existing byte comparison) between the local and relay writers, and pin write-if-changed with mtime so the tests also fail on filesystems that reuse a freed inode.

* fix(opencode): keep the final Idle when OpenCode 1 tears its instance down

OpenCode 1 disposes a plugin only when it tears the instance down, and that
teardown cancels every running session, so the Idle published on dispose is
true there; the cancelled run's own idle may never reach the plugin. Only
OpenCode 2 disposes on a hot reload while turns keep running. The generated
module serves both hosts, so the OpenCode 2 setup() entry point now tells the
shared factory that sessions outlive disposal; the server() path keeps the
previous disposal behaviour, including the hand-off to a surviving factory.

* fix(opencode): compare a symlinked plugin by its target before rewriting

OpenCode 2 loads plugins through file-level symlinks and reads the revision
from the target's mtime, so a user whose Orca plugin file is a symlink (per-file
dotfile managers) failed the regular-file check and got a write through the
link, and a reload, on every spawn. The config-dir and relay installs now skip
the write when the resolved target already has Orca's bytes; when stale they
behave as before. Only the per-source overlay keeps the regular-file check,
since a link there mirrors a user entry. Installers also skip the write inside
a guarded block rather than returning early, so later install steps still run.

* test(opencode): skip the plugin symlink tests on Windows like their neighbours

Creating a file symlink on Windows needs Developer Mode or admin rights.

* test(opencode): stub fetch without a type assertion in the dispose host test
2026-09-28 21:09:13 -07:00
Neil 3976ad4c59 perf(test): remove obsolete structural snapshots (#23777) 2026-09-28 20:55:05 -07:00
Brennan Benson b776e9ac99 fix(browser): give a tab's identity one owner so viewport presets stop dropping client hints (#23718)
* fix(browser): give a tab's identity one owner so viewport presets stop dropping client hints

A desktop viewport preset installed a CDP user-agent override with no
userAgentMetadata. Chromium then drops navigator.userAgentData and every
sec-ch-ua header for that tab: a Chrome UA with no client hints. Identity was
decided separately by the session request hook, the Google sign-in switch and
the viewport code, and nothing decided per tab who it should claim to be.

resolveBrowserTabIdentity now derives it from the process identity mode,
whether the URL is a Google auth host, and whether a mobile preset is
requested. applyTabIdentity is the one writer: it keeps the WebContents UA on
the process or Firefox identity and clears the CDP override whenever that
layer already presents the identity. Viewport emulation only records the
requested preset; its metrics and touch steps log failures independently, so
a rejected step can no longer skip the identity restore.

* test(browser): read the presented identity instead of casting the guest stub

* test(browser): drop a comment that described desktop presets writing a UA

* fix(browser): keep same-document navigations and unapplied presets off the tab identity

A same-document navigation (pushState/replaceState) now never rewrites the
WebContents user agent. Chromium reloads a still-loading document when its
user agent changes, so an OAuth callback that strips its code with
replaceState after a redirect off Google sign-in was requested twice,
replaying the one-time code. Measured on Electron 43.7.5: the callback URL
hits the server twice with the write, once without.

The session request hook now derives the mobile identity from the CDP
override the tab actually holds instead of the requested preset. A preset
whose write never landed (debugger attach refused while DevTools is open, a
failed write, a detach) no longer puts the iPhone user agent and mobile
client hints on the wire while the document reports desktop.

* fix(browser): restore identity after a failed navigation without reloading the error page

did-fail-load fires while the failed URL's error page is still loading, and
WebContents.setUserAgent() at that moment makes Chromium reload it. After a
redirect onto or off the Google sign-in host (identity moved over CDP only),
the restore rewrote the WebContents UA there and replayed the failed request.
The restore now goes over CDP; the next navigation rewrites the WebContents UA.

* refactor(browser): let only a navigation start write the WebContents user agent

Two review rounds each found a caller that asked the identity writer to
rewrite the WebContents UA at a moment Chromium reloads or cancels the page
(a same-document navigation, a failed load). A boolean at every call site
left that decision to the callers. The writer now has two entry points:
presentTabIdentityAtNavigationStart, the only one that may write the
WebContents UA and only for a cross-document navigation, and
retargetTabIdentity, which goes over CDP only and serves redirects, failed
loads and preset changes. A table test pins the rule for every entry point.

* test(browser): reject touch emulation regardless of payload in the identity-restore test

The mock rejected only maxTouchPoints 0, so the test would stop exercising a
failed touch step once the touch payload is fixed.
2026-09-28 20:52:36 -07:00
Brennan Benson b4c19f12c4 fix(claude): run structured Claude under the POSIX provider supervisor (#23476)
* fix(codex): the provider supervisor outlives its provider group when stopped

A signalled supervisor forwards the signal to the provider group, escalates to
SIGKILL after the grace, and exits only once the group is gone, so recovery's
proof that the recorded pid is dead also proves the provider is. It refuses to
spawn when its parent is already not the owner named in its spec, and watches
that owner rather than whichever parent it first saw. The grace is a spec
field. Recovery's SIGTERM stage now outlasts the supervisor's own stop, since a
SIGKILL that lands first cannot be handled and leaves the group running.

* fix(codex): a closed owner pipe no longer ends the supervisor before its provider group

When Orca dies, the supervisor's stdout pipe has no reader. Provider output in the
window before the parent-death watch fired raised an unhandled EPIPE that exited the
supervisor with the provider group still running.

* fix(codex): bound the supervisor grace so recovery's SIGTERM stage always covers it

Recovery sized its SIGTERM stage from the default grace, so a launch with a
longer grace would be SIGKILLed mid-stop and orphan its group with no test
noticing. The spec now refuses any grace above one exported maximum, and
recovery derives its SIGTERM stage from that maximum.

* fix(codex): every supervisor stop asks the provider with SIGTERM first

Owner death, stdin end after the grace, and a signal to the supervisor now all
take one path: SIGTERM the provider group, SIGKILL it after the grace, and exit
only once it is gone. The signal handlers are registered before the provider
is spawned, so a stop that lands in the spawn window still reaps it. The
longest stop grows to two graces plus the reap wait, and both recovery's
SIGTERM stage and the connection's graceful close now wait that long before
forcing, since forcing the supervisor sooner can orphan its group.

* fix(claude): run the structured Claude child under the POSIX provider supervisor

A close now stops Claude with a SIGTERM through the supervisor instead of letting
stdin end finish the turn, and Orca's death stops it through the supervisor.

* test(claude): pin the supervised stop against a real Claude CLI, opt-in

* test(claude): a requested stop reads interrupted through the frames the supervised SIGTERM makes Claude emit

* test(claude): show what the real CLI did when it never ran the tool

* test(claude): Orca's death now reaps Claude's own tool through its SIGTERM

* refactor(claude): take supervision from the spawn spec so the close ladder cannot disagree with the spawn

createProviderSpawnSpec now reports whether it wrapped the provider in the supervisor, and the Claude spawner reads that instead of repeating the platform check. The close ladder's SIGTERM follows the process actually spawned.

* fix(native-chat): derive quit's chat-eviction bound from the longest supervised provider close

Quit's child-eviction phase was a hand-picked 8 s. It is now the sink drain plus the longest
supervised close over Claude and Codex plus a named 1 s margin, so a provider close that grows
widens it instead of silently outrunning it. A close's tree-kill fallback stays outside the bound:
once main exits, the supervisor stops its provider group on owner death, which a new test now
proves for a clean owner exit, and next launch's recovery settles the lease.
2026-09-28 20:49:32 -07:00
Brennan Benson 2dc2693953 fix(native-chat): a turn a proven crash cut short reads interrupted (#23456)
* fix(native-chat): a turn a proven crash cut short reads interrupted, ending when it was last seen working

* fix(native-chat): end a probe-proven turn at the last row the journal wrote live

A revised item keeps its first sighting's timestamp, so a long command or a streamed reply read as ending when it started. The reducer now tracks the latest live row the same way it tracks all activity.

* test(native-chat): give the unexpected-exit fake journal its live-activity read

* refactor(native-chat): read the journal's live bound only for a probe-proven death

* fix(native-chat): mark what a journal open settles for a gone host as crash reconciliation

A crashed host's working roster is retired when the journal reopens. That row was
written live, so a probe-proven turn ended at the relaunch and counted the downtime.

* fix(native-chat): bound a probe-proven death with the last time Orca proved the owner alive

A crash mid-tool left the turn ending at the tool call's start, because Claude writes nothing
while a Bash call runs. The death evidence now records the lease's last renewal before the
death as lastProvenAliveAt, and the turn ends at the later of that and the last live row,
capped at the probe.

Parking a lease in recovery no longer stamps lastRenewedAt, since nothing proved the owner
alive then; a child that outlived Orca would otherwise have its turn count the downtime.

* test(native-chat): a failed acquisition parked in recovery keeps its last proof of life, and the timing read goes through the display selector

* refactor(native-chat): move the submission dispatch folds out of the journal reducer

Main grew the reducer to its line limit, so the live-activity bound tipped it over. The dispatch
row and echo-acceptance folds move unchanged into their own module.

* test(native-chat): a send or reader that opens a crashed chat settles a proven death interrupted

Main's open-time settle test still asserted the old rule, where only a watched exit proved a death.

* docs(native-chat): say which proofs of death record a last proof of life

* fix(native-chat): a proof of life bounds only the owner that wrote the turn

A start after a crash that spawned a child and then failed without proving it gone parks that
child for recovery; when recovery finds it gone, the proof of death carries its proof of life,
which is after the crash. The older turn then ended there and counted the downtime.

The journal now derives the fence of its newest live writer, and the lease that holds the proof
names the owner it released by the fence it moved to. The last renewal counts only when that
move was one step past the writer of the turn.

* test(native-chat): a crash with a send in doubt still ends at the last proof of life

The reopen settles that send at the new fence, so the owner check must read the fence of live rows only.

* fix(native-chat): a proof of death judges only the turn its own owner wrote

The settle read the record's latest proof of death for whatever turn a gone generation left
running. After a crash, a start that reserved a new fence cleared the relaunch's proof, and if
it then failed, its own child's death (a watched exit at the failure, or a probe finding the
child it left for recovery gone) ended the older turn an hour after the crash.

Every proof of death now records ownerFence, the fence of the owner or reservation it is about;
a fence names exactly one owner. The journal derives the fence each item was created at, and a
running turn is interrupted only by a proof naming its own owner; otherwise it is unverifiable.
Evidence older builds wrote keeps their rule. This replaces the derived one-step fence check.

* test(native-chat): every writer of a watched exit names the owner it released

A watched exit that names no owner reads the older rule, so the settle alone cannot tell a
dropped field; the writers are pinned directly, including past a recovery floor.

* test(native-chat): give the fake journal's cast its safety rationale

* fix(native-chat): a proof of death written after a chat opened revises the turn it left unverifiable

On desktop the chat on screen at relaunch opens before the startup reconcile has probed its owner,
so the open settles the cut-off turn unverifiable. When the reconcile then records the proof, it
re-runs the same settle for every open conversation, which revises that owner's unverifiable turns
to interrupted with the proof's end. Any later open re-runs it too, so a failed write converges.
Only upward, only for a proof that names the turn's own owner.

* test(native-chat): a chat read before the reconcile reads unverifiable, then interrupted

Covers the reconcile revising an open chat to the last renewal (27 s) and a subscriber being sent
both states, a start after the crash whose running turn the queued revision leaves alone, a failed
revision write converging at the next open, a proof about another owner or from an older build
never revising, and a second settle writing nothing.

* fix(native-chat): revise an open chat's turn wherever a proof of death is written

The store tells its listeners, once committed, of each record a transaction gave a new proof of
death, so every writer (the startup reconcile, a recovery that stopped a child which outlived
Orca, a failed start, a watched exit) triggers the same serialized resettle for a chat already
open. The reconcile's own callback is gone. Quit stops listening first, and a queued resettle is
drained with the starts.

* test(native-chat): a failed exit settlement is retried in place once the exit is recorded

Recording a watched exit now queues the same settle an open runs, so the turn converges without
waiting for the chat to be reopened. The reopen and read-after-restart cases now refuse that retry
too, so they still pin the open's own settle.

* test(native-chat): tests that pin a send settling a failed exit refuse the in-place retry too

The exit's release now queues the same settle, so two tests named for the send's settle refuse that
retry as well; the comments that said nothing retries it now say what does.

* fix(native-chat): name the explanation row by the death it explains, so a retried settle adds no second row

* fix(native-chat): end a crashed turn at its owner's provider output, never at a later send

A send accepted into a crashed chat before the proof of death wrote a submission row live, and the journal-wide last-live-row bound counted it, so the revised turn ended at the send and counted Orca's downtime. The bound is now the last row the owner's provider child wrote, per writer fence: submissions, dispatch rows and crash reconciliation are Orca's or the user's, and a newer owner's work says nothing about the dead one.

* fix(native-chat): end a crashed turn at its last proof of life, never at a timeline row

The end of a probe-proven death was the later of the last renewal and the last live timeline row. A send accepted into a crashed chat before the proof writes a row live, so the revised turn ended at the send and counted Orca's downtime. Rows cannot tell the agent's output from Orca's or the user's, so the end is now the last renewal alone, never after the probe, and never before the turn began. The journal's live-activity bound and the reopen's recovered marker, which existed only for it, are gone.

* test(native-chat): drop a stale reference to the removed live-activity bound
2026-09-28 20:48:34 -07:00
Brennan Benson b283a09688 fix(native-chat): stop flashing "still starting" on every chat launch (#23666)
* fix(native-chat): stop flashing "still starting" on every chat launch

Every structured chat passes through a short startup phase, and the pane
showed "<agent> is still starting…" for all of it, so a normal launch
flashed the notice for a fraction of a second. The notice now goes through
a keyed delayed status: it appears only once startup outlasts a grace
period, stays up for a minimum time once shown, and resets per session.

* fix(native-chat): reset startup notice for each provider child
2026-09-28 20:47:51 -07:00
Brennan Benson 94b5a7d256 fix(native-chat): only Codex's turn completion ends a Codex turn (#23682)
* fix(native-chat): the sink queue keeps a settlement's first batch, as the journal does

The journal applies a lifecycle batch's settlement id once and skips any later
batch with the same id. The deferred sink queue coalesced the same key the other
way: a second batch replaced the first while it was still queued. So which
record survived depended on whether the first had drained yet.

A lifecycle batch now keeps the queued operation with its key, and a later one
is accepted and dropped, which is what the journal does once the first is
written.

* fix(native-chat): only turn/completed ends a Codex turn

Codex follows every turn-ending `error` (willRetry=false) with a failed
`turn/completed` for the same turn, 0-32 ms later. That was captured from the
real app-server on 0.141.0 and 0.158.0 across eight failure scenarios, and it is
how Codex builds a failed turn: it records the error as the turn's last error,
records any pending input, and then derives `failed` from that error when it
completes the turn.

The translator ended the turn twice: once on the error, and again on the
completion, with a guard to make the first end final. Ending on the error threw
away what only the completion carries: Codex's duration, and the completion's
receipt time. It also forgot the turn before Codex recorded the turn's pending
input.

Now the error is only the row the user reads, inside the still-open turn, and
`turn/completed` is the turn's only live end. A process exit between the two is
the existing exit sweep's observed end, recorded as interrupted.

A failed completion is stored as completed with outcome failure, live and on
restore alike. Only `interrupted` maps to the interrupted state.

The first-end-final guard is gone. Codex sends one completion per turn, the only
redelivery Orca has is the retry of a refused frame (which changes nothing), and
the settlement id already keeps the first record in the queue and the journal.

* refactor(codex): delete the unreachable oversized-notification settlement

The translator settled a streamed item when the transport rejected its
notification as oversized. Nothing can produce that frame. The Codex stdio
reader frames with `maxLineBytes: Number.POSITIVE_INFINITY`
(codex-app-server-record-reader.ts), which it has done since the app-server
records were uncapped. With an infinite limit the framer never reports
`line-too-long`: no line, pending suffix or paused queue can exceed it. So the
dispatcher never emits `frame:oversized-notification`, and the arm that settles
it never runs.

The arm, its helper module and its test go. In place of the test, the
connection test now proves the reason: a notification past the old 16 MiB wire
limit arrives whole, and no oversized frame is reported.

* test(codex): replace the captured ids in the turn-endings fixture with synthetic ones

The replay reads ids only to group frames, so the real thread, turn and
response ids from the capture account carry nothing the test needs. The
fixture moves beside the Codex tests that read it.

* test(codex): use a neutral made-up status as the unknown-status example

'cancelled' read as a stop being recorded as a completion.

* test(codex): a restored turn with a status Orca cannot place ends with no verdict

Codex's history carries the same status field as the live completion, so the
restore path is pinned to the same mapping: completed, and no outcome.
2026-09-28 20:47:45 -07:00
Brennan Benson a5ce8251e3 Agent launches carry the surface that started them (#23697)
* feat(agent-launch): every launch carries the surface that started it

The host now attributes every agent it builds to the surface that asked for
it, resolving a missing or unrecognized surface to 'unknown' in one place
instead of silently skipping it. The CLI names itself on worktree.create and
orchestration workers name themselves host-side.

* fix(agent-launch): attribute the agent a startup-draft create launches

The host builds a third kind of agent launch: a worktree.create with a
startupDraft and no startupAgent, where the host picks the agent itself.
It carried no launch record at all and ignored the caller's launchSource.
Route it through the same resolver as the other two builders, and derive
the startupAgent terminal record only from the resolver so no prebuilt
record can stand in for it.

* fix(agent-launch): attribute the agent a host-built agent session launches

terminal.createAgentSession builds a fresh agent's launch on the host, like the
other startup builders, but spawned it with no launch record, so those launches
were never counted. Record them through the same resolver; the request names no
surface, so they count as unknown.

* test(agent-launch): require an attribution decision for every host-built agent startup
2026-09-28 20:33:59 -07:00
Jinjing 85ad292930 fix(status-bar): re-measure collapsing levels when only the collapsed width moves (#23773) 2026-09-28 20:23:28 -07:00
Jinwoo Hong a134d1259e feat(mobile): tell users when a newer app binary is installable (#23755)
* feat(mobile): tell users when a newer app binary is installable

With OTA page updates, store releases get rare and users stop looking.
The shell now asks the channel that installed it. Android sideload
reads GitHub's mobile-android-v* tag refs and proves the release has
an APK. iOS reads the App Store lookup. A home card above Desktops,
dismissible per version, and Settings rows surface the result. The
check runs on the desktop updater's cadence: cold start, foreground
once 24 h have passed, and a 1 h retry after a failure.

The releases atom feed was not used because it lists only the 10
newest releases, which are all desktop builds, so it never carries a
mobile tag.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): parse update replies with zod schemas

The anti-slop gate refuses Reflect.get on dynamic input. The GitHub
refs, the release, the App Store lookup and the stored update record
are now parsed into named schemas before they are read.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say why the Android update source reads tag refs

Record why the Android source reads tag refs. The releases atom feed
and /releases?per_page=100 are both newest-first windows that desktop
releases fill. Either would silently report "current" once a run of
desktop builds pushes the newest mobile release out.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): load update state once and apply review rulings

Every check and dismissal now awaits one shared store load. This
replaces the merge that guessed whether a check had landed during the
load. A manual check that fails while the store loads therefore keeps
its 1 h retry instead of re-checking at once.

- checking is derived from the in-flight check.
- start() uses a per-start flag, so a StrictMode double start applies
  one load.
- A check that finishes after stop() writes nothing.
- A corrupt stored update record loses only itself.
- Tag refs are parsed with a single schema.
- The runtime wiring is folded into one file, and the card moves to
  home/.
- The recorded App Store fixture is oxfmt-formatted, with the same
  parsed value.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the update timer armed across a stop and restart

A check that stayed in flight across stop and restart returned
'failed' without rescheduling. The restart skipped arming because a
check was in flight, which left a live checker with no timer until
the next foreground. The stop counter is removed. schedule() already
arms nothing while no start is active, and saving a real result after
a stop is harmless.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore: retrigger CI after the RPC recording repin (#23757) landed on main

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): trim the update checker and Settings rows

- The load sets prefs and the due time only. start() re-arms the
  schedule after it.
- The Settings result hides through one effect keyed on the result.
- onUpdate receives the URL.
- The version row is bound once.
- The retry and timeout constants are no longer exported.
- The unused AppUpdateChecker type is deleted.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): run the update check when its timer fires

The armed timer is the due time. Re-checking the wall clock when the
timer fired meant a clock stepped back skipped the check and re-armed
nothing. The due-time guard now applies only on foreground.

The binary version still comes from expoConfig.version. SDK 55
removed Constants.nativeAppVersion, so the no-expo-updates invariant
is now named in the comment.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): recover from a future check time and use Apple's page URL

If the device clock was ahead when a check ran and was corrected
later, the stored check time is in the future. Cold starts then armed
a timer for the whole skew, and foreground never came due. The stored
state now reads as never checked in that case.

The iOS link is the lookup's trackViewUrl instead of a URL built from
trackId.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(mobile): fit the future-check-time comment in the print width

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 22:28:50 -04:00
Neil ec9f35e2ee perf(ci): plan the unit shards before the static-analysis gate instead of behind it (#23743)
A caller's `needs` gate the whole called workflow, so while the plan job lived in
unit-tests.yml it could not start until static analysis and typecheck had both
finished and passed -- and the shard matrix then waited on it. The two hops were
serial when they did not need to be: planning reads the checkout, a git diff
against HEAD^1, the import graph and the checked-in timing baseline in
config/scripts/ci-shard-timings.json, and consumes nothing that static analysis,
typecheck or the native-cache primer produce.

Planning moves to its own reusable workflow so pr.yml can run it against
code_paths alone, overlapping it with the gate. Measured across 99 runs, the
shard matrix is created a median 93s earlier (p25 47s, p90 241s, never later).
Planning stays a required predecessor of the shards, so an empty assignment
cannot expand the matrix.

The gate itself is deliberately left in place. It fires on 22% of runs, and the
shard queue wait knees hard above ~9 concurrent ARM jobs -- 4s median below that
against 218s at 15-19 -- so admitting 8 doomed shards per failed run would cost
more in queue pressure than it returns in latency.

Cost is one 37s ubuntu-latest job, which does not touch the ARM pool the shards
contend for.

A planning failure still fails the PR: the shards are skipped, and verify's
check_job requires success whenever the classifier says tests should run, so it
reports `test: expected success, got skipped`.
2026-09-28 19:22:10 -07:00
Jinwoo Hong 28942eed5a test(mobile): repin the RPC recording corpus to main after #22762 (#23757)
#22762 squash-merged as 29c7d5d983 with the corpus pinned to its branch
commit 03995ae29d, which the squash left unreachable from main. Repin
baseline to main's tip and re-record the full corpus in place; only the
baseline header moves on every golden.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 21:16:37 -04:00
Neil ccdb324b63 Add CodeBuddy as a built-in coding agent (#23740)
* feat(agents): integrate CodeBuddy launch, status and session history

* docs: record CodeBuddy lifecycle verification

* fix(codebuddy): backfill scoped history and negotiate remote resume

* test(cli): include CodeBuddy in known search agents
2026-09-28 18:11:25 -07:00
Neil da48d98040 fix: bound combined diff editors and scope chat style invalidation (#23725)
* fix(editor): bound offscreen combined diff rendering by height

* perf: scope native chat relational styles to their ancestor
2026-09-28 17:57:08 -07:00
Neil d606be3ade test: wait for remote terminal grid convergence after reveal (#23738)
* test: wait for revealed remote PTY grid convergence

* test: retain reveal diagnostics on geometry failure
2026-09-28 17:45:09 -07:00
Neil bd5dca4406 fix(editor): preserve combined diff scroll on line focus (#23735) 2026-09-28 17:14:03 -07:00
Neil 8bb78f0ddd test(browser): create probe body before starting frame requests (#23730) 2026-09-28 17:07:07 -07:00
Neil 64dbe87de6 test(terminal): wait for decoy panes before host parking (#23729) 2026-09-28 16:54:02 -07:00
Brennan Benson 4b622e1b13 fix(runtime): reopen the quiet-foreground tui-idle lane for agents with no other rest signal (#23598)
* fix(runtime): reopen the quiet-foreground tui-idle lane for agents with no other rest signal

A tui-idle wait could never settle on a pane running amp, goose, crush, kimi,
qwen-code, rovo, auggie and other agents whose titles Orca cannot classify:
the quiet-foreground lane was closed for every launched agent, and it was the
only lane those agents could reach, so worker start failed at agent_readiness
after 60s.

Model each agent's rest signal, derived from the tables that already encode
it (synthetic ready titles, the title classifier, the DSH hook and Muse ready
screen lanes). The lane stays closed where a stronger signal will arrive and
reopens for agents with none. On a reopened lane, silence counts only after
the TUI has painted: an agent that has painted nothing is still booting.

Linear: STA-7440

* fix(runtime): count only the command's own output as an agent's paint on the tui-idle foreground lane

The after-paint lane accepted any output, and the shell's prompt and echoed launch
command always land before the agent starts, so a silently booting agent could still
settle and lose its first prompt. The runtime now reads the shell integration's
command-start marker and requires visible output after it; panes whose shell emits
no marker keep the any-output rule.

Also skip the foreground-process read while the pane cannot settle, and register the
new title-classifier call site in the pane agent identity inventory.

* fix(runtime): classify Freebuff's rest signal and skip the backward marker scan on chunks without one

Main added the Freebuff agent after this branch point, so the full per-agent rest-signal table no longer matched on the merge ref. Freebuff derives `none`: its screen reports a first-party `done`, which tui-idle trusts only for DSH, so the quiet-foreground lane is its only one.

The command-start scan ran a backward search over every PTY chunk; a forward check first cuts that to the cost of a plain substring test on chunks with no marker.

* fix(runtime): classify Qoder's rest signal after merging main

Main added Qoder with its own readiness branch returning a boolean quiet-foreground
flag; map it to the lane type and classify Qoder by its ready screen so the full
rest-signal table and lane-agreement check stay exhaustive. Say what `none`
actually means: no stronger lane tui-idle trusts, not no hooks at all.

* refactor(runtime): track command paint with the shared OSC 133 scanner

The command-paint tracker had its own split-unsafe 133;C parser. Reuse the
chunk-boundary-safe scanner, which now reports where in the chunk the marker
ended, so a marker split across reads is still found. Correct the unmarked-launch
list: bash and zsh mark typed launches after the echo.

* fix(runtime): drop command-paint state on an output gap or a new process

A dropped chunk can cut a command-start marker in half, and the scanner's
carry then completes it on unrelated output after the gap, leaving the
pane waiting for a paint that already happened. Reset it with the other
cross-chunk carries.

* fix(terminal): keep the command-start offset out of renderer lifecycle callbacks
2026-09-28 16:48:07 -07:00
Neil 2aed2cf64a fix(terminal): reveal splits while the source pane binds (#23692)
* test(e2e): observe passive terminal restoration before activation

* fix(terminal): reveal persisted splits during source binding publication

* test(terminal): handle nullable persisted layout roots
2026-09-28 16:34:36 -07:00
Neil d0db35c18c fix(terminal): preserve typing while a remote pane reattaches (#23701)
* fix(terminal): retain typing while a parked remote pane reattaches

* test: persist restored remote terminal screenshots

* fix(remote): buffer recovery reconnect input

* fix(remote): retain input across restored pane attach

* fix(remote): flush attach input after subscription

* fix(remote): flush reattach input after attach readiness

* fix(remote): stop buffering after reattach readiness

* test(remote): trace parked reattach input lifecycle

* test(remote): forward paired client lifecycle diagnostics

* fix(remote): preserve restored typing before connect starts

* chore(i18n): refresh runtime required catalog

* fix(i18n): ship compact agent runtime label

* fix(i18n): merge required label into existing sidebar catalog
2026-09-28 16:12:39 -07:00
Brennan Benson d68eee3757 fix(runtime): retire an exited terminal before its stream end (#23492)
* fix(runtime): retire an exited terminal before its stream end

An exit's durable retirement became asynchronous, so onPtyExit released
the terminal stream before the retirement landed. A paired client answers
a stream end by re-activating its pane; that activation still found the
exited leaf, materialized it under the same session id, and registerPty
dropped the pending retirement. The exited split pane came back as a
fresh shell.

The exit now stages the retirement into the in-memory session and
publishes it synchronously, then notifies exit listeners, and only then
makes it durable. A failed durable write is logged and left in memory for
the next profile write instead of being rolled back, since the process is
gone either way. This removes the pending-retirement latch and its
post-await incarnation fence: there is no longer a window for them to
guard.

* test(runtime): a failed exit retirement still reaches disk

Pins the no-rollback contract through a real Store and SQLite authority:
when the retirement's own durable write fails, the in-memory retirement
is carried by the next unrelated profile write, and by the app-quit
flush when no other write happens. The delayed authority fixture can now
fail its next write, and the acknowledged-retirement fixture reads the
database a relaunch would load and models the quit flush.

* test(runtime): a stream end observes the exit retirement already published

The re-activation check alone passes with the listener ordering reverted,
because activation awaits before its lookup. Record the session binding
and publication count at the moment the exit listener fires so the
ordering itself is pinned.

* fix(runtime): an exit cleanup fault still ends the terminal stream

* perf(runtime): exits retired together share one durable write

* test(runtime): a refused staging write still retires the pane and ends the stream

* refactor(runtime): describe exit retirement as staged, not durably accepted

The retirement result is staged in memory before any write, and the removable-surface comment and the replacement-admission test name still described the old publish-after-durable rule.
2026-09-28 16:12:26 -07:00
Brennan Benson 2af897d7ea fix(codex): the provider supervisor outlives its provider group when stopped (#23466)
* fix(codex): the provider supervisor outlives its provider group when stopped

A signalled supervisor forwards the signal to the provider group, escalates to
SIGKILL after the grace, and exits only once the group is gone, so recovery's
proof that the recorded pid is dead also proves the provider is. It refuses to
spawn when its parent is already not the owner named in its spec, and watches
that owner rather than whichever parent it first saw. The grace is a spec
field. Recovery's SIGTERM stage now outlasts the supervisor's own stop, since a
SIGKILL that lands first cannot be handled and leaves the group running.

* fix(codex): a closed owner pipe no longer ends the supervisor before its provider group

When Orca dies, the supervisor's stdout pipe has no reader. Provider output in the
window before the parent-death watch fired raised an unhandled EPIPE that exited the
supervisor with the provider group still running.

* fix(codex): bound the supervisor grace so recovery's SIGTERM stage always covers it

Recovery sized its SIGTERM stage from the default grace, so a launch with a
longer grace would be SIGKILLed mid-stop and orphan its group with no test
noticing. The spec now refuses any grace above one exported maximum, and
recovery derives its SIGTERM stage from that maximum.

* fix(codex): every supervisor stop asks the provider with SIGTERM first

Owner death, stdin end after the grace, and a signal to the supervisor now all
take one path: SIGTERM the provider group, SIGKILL it after the grace, and exit
only once it is gone. The signal handlers are registered before the provider
is spawned, so a stop that lands in the spawn window still reaps it. The
longest stop grows to two graces plus the reap wait, and both recovery's
SIGTERM stage and the connection's graceful close now wait that long before
forcing, since forcing the supervisor sooner can orphan its group.

* fix(codex): give the provider 1 s after stdin end and 3 s after SIGTERM to flush before SIGKILL

The supervisor's stop was stdin end, 1.25 s, SIGTERM, 1.25 s, SIGKILL. Codex
now gets 3 s after SIGTERM to flush its state. The two graces are separate
constants, the longest stop they derive becomes 5.5 s, and a test keeps it
inside quit's 8 s child-eviction bound.

* test(codex): count eviction's pre-stop drain in the quit budget test

Eviction drains the sink for up to 1 s before it stops the child, inside the same 8 s bound.
2026-09-28 16:04:04 -07:00
Kelvin AmoabaandBrennan Benson 64569a8183 fix(ssh): don't overwrite remote agent config after a failed read (#22644)
* fix(ssh): don't overwrite remote agent config after a failed read

A flaky read was treated as an empty file, wiping the user's config.

Fixes #22638

* test(runtime): model remote missing-config reads as relay ENOENT errors

The runtime harness stubbed isENOENT as code-only, and the remote Codex
startup specs rejected with a generic error that only passed while any
read failure seeded an empty config. Use the real isENOENT and the
message-only shape the relay actually delivers.

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-09-28 15:53:11 -07:00
Brennan Benson 9b11b1d594 fix(agent-hooks): compare status rows structurally instead of serializing both (#23585)
Status-row change detection stringified two full IPC payloads on every
status write, including an up-to-8 KB lastAssistantMessage re-posted on
every OpenCode streamed part. Compare the same published field set with
the existing structural-equality helper, with a same-reference fast path.

Linear: STA-7432
2026-09-28 15:44:59 -07:00
Brennan Benson e7940121eb fix(tab-group): measure fallback pane geometry once per tab group, only while visible (#23592)
* fix(tab-group): measure fallback pane geometry once per tab group, only while visible

* refactor(tab-group): derive the shared resize listener's lifetime from the source map

The map already drops empty groups, so a separate counter was a second copy that could disagree.
2026-09-28 15:43:09 -07:00
Brennan Benson e8e144bf3c fix(native-chat): a subagent's words are presented as that subagent's, never the parent's (#23605)
* fix(native-chat): a subagent's words are presented as that subagent's, never the parent's

The journal already names the agent that produced every row, but the transcript
projection dropped it, so a subagent's prose rendered as the parent's reply, its
tool calls folded into the parent's runs, and a settled turn could fold down to
a subagent's words as its only visible answer.

The transcript message now keeps the row's producer. The fold keeps each agent's
calls in that agent's own run, a turn's answer is the session's own agent's last
prose, and a subagent's row names the subagent on desktop, mobile and a worker's
transcript text.

* test(native-chat): give the window fixture's slot the attribution field it now carries

* fix(mobile): read the subagent label the row is given, and pin the caption

* fix(native-chat): keep interleaved agents in order and each agent's own run live

Review follow-ups:
- the fold is main's adjacency fold plus one condition: a row never folds into
  another agent's run, so an agent's later call stays below its subagent's work
  instead of jumping back into its earlier row
- each agent has its own live frontier, so a parent still inside its spawn call
  reads as running while its subagent works below it
- mobile names no one on a row whose only content is hidden behind its settled turn
- a pending question from a subagent keeps its producer
- worker reads serve only the producing agent's id, bounded like the roster key
  that names it, and drop the provenance fields
- the single-message worker formatter is private, so no caller can drop names
2026-09-28 15:28:01 -07:00
Brennan Benson c5330d0d52 fix(native-chat): stop killing processes that only inherited a chat's spawn tag (#23460)
* fix(native-chat): stop signalling processes that only inherited a spawn token

A spawn token is an environment variable, so every descendant of a provider child
carries it. The Linux-only startup scan treated any carrier no lease claimed as a lost
provider child and sent it SIGTERM, which also hit editors, tmux servers and nested
Orca processes the agent had started. Remove that scan's killing consumer; the token
scan stays for the reservation probe, and recorded owners are still stopped by
identity during recovery.

* fix(codex): remove the token-scan kill path from app-server teardown

Every descendant inherits the spawn token, so killing each pid that carries it can
reach processes the agent started that are not the provider. Production never
injected this path; teardown always uses the process-group and descendant-snapshot
proof. Drop it, its deps, and the now-unused spawn-token argument.
2026-09-28 15:25:24 -07:00
Brennan Benson 2ca4ecbc61 feat(orchestration): let a structured chat run orchestration as itself (#22568)
* feat(orchestration): inject the Orca session id into structured children and let the CLI act as it

Every structured session's child (native Claude, native Codex, and the terminal
view) carries ORCA_AGENT_SESSION_ID and reaches the Orca CLI. The CLI sends the id
in the orchestration envelope; when present it is the caller, and a caller flag
naming anyone else is refused before any request. The id is stripped from
inherited PTY env and from the SSH host-CLI passthrough, and crosses into WSL so
the host can refuse the cross-host claim.

* test(orchestration): pin session id injection for native Claude, native Codex, the terminal view, WSL, PTY inheritance and SSH

* test(orchestration): pin one caller precedence rule across every CLI verb that names its caller

Adds the per-verb table (flagless acts as the session; a conflicting --from or
--terminal is refused before any request; the session's own spellings are
accepted), the enumerated guess population with its positive control, the
structured worker's own handle, the identity-less refusal for an older child,
the unchanged terminal agent, and the envelope. dispatch-show's --from only fills
preview text, so it passes through unfenced and a session's flagless preview
names the address the real dispatch writes.

* refactor(orchestration): keep the identity-less marker reader to the marker; the id is checked first

* test(orchestration): pin that a host refusal of the session surfaces verbatim from the CLI

* fix(orchestration): keep the identity-less marker beside the id for CLIs that predate it

A CLI older than the id, reached through a global install when a shell rc resets
PATH, would otherwise guess a sibling's terminal in a chat that no longer carries
the marker. It refuses on the marker instead; a current CLI checks the id first,
so the marker never makes a session with an id identity-less.

* fix(orchestration): refuse a conflicting --from on gate-list and task-list scoped by --run

A --run listing needs no caller, so both handlers skipped the resolver and a
--from naming another actor was dropped silently under a session. The conflict
check now runs on that branch too; terminal callers are unchanged.

* fix(orchestration): name this app's CLI by absolute path for a structured session's login shells

A provider can run each command in a login shell: Codex runs zsh -lc, and the
profile rebuilds PATH, putting a global install (possibly an older Orca) ahead of
the directory Orca prepended. ORCA_CLI_COMMAND, which an agent resolves the CLI
from first, is now the absolute launcher in that directory (the native launcher
on Windows), so no shell's startup files can swap it. The PATH prepend stays for
shells that read no profile. Found by the live coordinator run of the next PR.

* test(orchestration): pin a structured worker's CLI command as this app's absolute launcher

* test(orchestration): run the zsh login-shell arm in the real-shell lane that installs zsh

The ordinary Linux unit lane has no /bin/zsh, so the zsh arm failed there with
ENOENT. It moves to a live-shell file registered in the shell-contracts lane; the
bash arm keeps running in every lane. The lane guard's detector now also sees a
zsh spawned through the ProcessSpec program field, which is how this test
escaped it.

* fix(orchestration): omit a structured child's CLI command when no launcher resolves, and pin its instance

A bare `orca` fallback named GNOME's screen reader on packaged Linux, and an inherited value named
another app's CLI. The builder now deletes any inherited value, sets the absolute launcher only when
one resolved, and pins ORCA_USER_DATA_PATH so a current CLI dials the instance that minted the id.
Renames the marker reader to hasStructuredSessionMarker and records why the terminal view carries
the id without the marker.

* fix(terminal): name this app's CLI launcher by absolute path in every local terminal

ORCA_CLI_COMMAND meant three things by lane: an absolute launcher for a structured session, a bare
name for WSL, and nothing for any other terminal, so a structured session's terminal view lost it.
Local terminals now get the same absolute launcher the structured lane gets; WSL keeps its guest
command name, and a terminal whose launcher does not resolve still gets none.

* feat(cli): hand a command to the session's own CLI when another Orca CLI was invoked

A login shell can reorder PATH behind a global install, and an agent or its helper script can run
bare `orca`, so the binary that answered depended on the agent following instructions. Orca's
packaged launchers and bare-orca shims now export ORCA_CLI_SELF (outermost wins). At the CLI entry,
when it names a different launcher than ORCA_CLI_COMMAND, the command re-runs once through the named
launcher with ORCA_CLI_REEXEC=1 and exits with its status; both variables are consumed so no child
inherits them. Dev launchers export no self on purpose, WSL and SSH names never qualify, and a
launcher that cannot start leaves the command to run here. The Windows launcher no longer rewrites
ORCA_CLI_COMMAND; the legacy ask protocol normalizes its resume command itself.

* refactor(orchestration): declare which flag names the caller on each spec and refuse at the CLI entry

Each handler hand-classified its --from/--terminal as the caller or a target, and the refusal of a
conflicting caller flag ran inside the caller resolver plus two standalone calls for --run listings,
so a new verb that read its flag raw would pass a sibling's handle to a pre-session host. Specs now
declare identityFlagRoles, the CLI entry refuses a conflicting caller flag once from the spec, the
resolver only applies the id-wins rule, and a test fails any orchestration verb that accepts --from
or --terminal without classifying it.

* perf(cli): keep the session caller check off the actor codec's module graph

The check runs at the CLI entry for every command, and the actor codec pulls zod through the session
record. Compare the session's own spellings as plain strings instead.

* refactor(cli): spell a session's address from the one prefix constant, off the codec's module graph

The Orca session address prefix moves to a leaf module with no imports, re-exported by
the address codec, so the CLI entry check derives `session:<id>` from that constant
instead of re-typing it and still stays off the codec's zod graph. Prose and test names
say caller or Orca session id, not actor.

* refactor(orchestration): drop the session id's terminal-view spawn now that the handoff is gone

The terminal handoff was removed, so no terminal is ever a structured session:
- delete the terminal-view identity env and its WSL passthrough, and their tests;
- strip the session caller keys from every terminal's env unconditionally;
- the CLI's own-address spelling moves beside the injected id in src/shared, with
  a test pinning it to the address the host's party resolver gives that session.

* fix(terminal): run the Codex launch preflight through the CLI the terminal names

Packaged Linux names the userData shim in ORCA_CLI_COMMAND, while the preflight
ran the bundled launcher behind it. The CLI saw a different launcher and handed
the preflight off to the shim, booting Electron twice before every codex launch.

* revert(terminal): keep terminals on main's ORCA_CLI_COMMAND and Codex preflight

Only a structured session needs an absolute ORCA_CLI_COMMAND; local terminals go back to
naming none (WSL keeps its guest command), and the Codex launch preflight goes back to the
bundled launcher. The CLI handoff is scoped to sessions, so a terminal's preflight can no
longer be handed off and start Electron twice.

This reverts commit d2cefb6c03 and commit dd2853a5a9.

* fix(cli): hand off to the session's CLI only inside a structured session

The handoff ran whenever an Orca launcher's ORCA_CLI_SELF differed from an absolute
ORCA_CLI_COMMAND, so any process with both - a terminal, a script - ran another install's CLI
instead of the one invoked: a beta's --version lied, and an AppImage command from a terminal
that outlived its Orca failed. It now requires the injected session id, the identity it exists
to deliver. The launcher variables are still consumed in every process.

* fix(cli): name the packaged Windows command after the handoff decision

The launcher stopped writing orca/orca-ide over ORCA_CLI_COMMAND so the handoff could see a
session's absolute launcher, which also changed what every Windows terminal's CLI read. The CLI
entry now applies the launcher's rule itself once the handoff is decided, so terminals and the
legacy ask resume command see exactly what they saw before, and the resume-command reader
goes back to its original form.

* refactor(cli): decide the session handoff from the CLI's own entry, not a launcher export

Every packaged launcher, shim and dispatcher exported ORCA_CLI_SELF so the CLI could tell which
launcher ran it, and compared that with the session's ORCA_CLI_COMMAND. Two launchers of the same
app are different files, so a session that reached its own app through a global orca-ide on Linux
still handed off and started Electron twice, and the export rode artifacts every terminal uses.

A structured session now also names the JS entry its launcher runs (ORCA_SESSION_CLI_ENTRY), and
the CLI compares its own argv entry with it: any launcher of the same app stays, another install
hands off. The launcher scripts, Linux shim and dispatcher go back to main; the Windows launcher
keeps only leaving ORCA_CLI_COMMAND for the CLI to name after the handoff decision.

* refactor(cli): drop the session CLI handoff; the pinned instance and injected id already bind any current CLI

Every current Orca CLI dials the instance ORCA_USER_DATA_PATH names and sends the injected
session id in the orchestration envelope, so a bare `orca` that reaches another install's
current CLI already acts as the session. An older CLI has no handoff code and refuses on the
marker. The handoff only lined up versions between two current CLIs, and comparing two
separately derived paths kept misfiring (an AppImage's mount against its registered
extraction started the CLI twice on every call).

Removes the re-exec, ORCA_SESSION_CLI_ENTRY and ORCA_CLI_REEXEC, and the CLI-side Windows
command naming; the packaged Windows launcher rewrites ORCA_CLI_COMMAND again, as on main,
inside its own process only. resolveHostCliEntryPath goes back to the SSH passthrough.

* test(orchestration): say why the registered worker case pins the handle, now that every session's env is populated
2026-09-28 15:19:44 -07:00
Neil a880c885a6 test(browser): check grab scope through responsive chrome (#23709) 2026-09-28 14:59:29 -07:00
Brennan Benson 153d3fd3fa feat(native-chat): Codex sessions write their subagents into the host status store (#22553)
* refactor(native-chat): the Codex acquire names its turn-boundary methods as a set

Behavior-neutral: the same two methods stamp receipt time. Keeps the file
under the size limit once the child-work sink lands.

* feat(native-chat): Codex sessions write their subagents into the host status store

A Codex child thread and each persistent command become host child records,
fed through the same delivery, ingest and reducer the Claude lane uses. The
child's own turn decides it: turn start is live, turn completion settles it
with the outcome Codex reports, and a follow-up turn reopens the same record
as a new run. Its open tool call, last message, usage and waiting-on-user flag
come from its own thread's frames. A parent turn ending settles nothing.

* fix(native-chat): close a Codex child's tool call by its item id alone

A completion frame need not restate the tool it ran, so reading the tool name
before closing left the call open and the record naming a finished tool.

* test(native-chat): pin the Codex child-work evidence and every hop to the host's records

Child turn start/end/follow-up, open tool call, last message, usage, waiting,
the persistent command a child owns and its monitoring display, a primary
turn end settling nothing, and session end. End to end through the real
adapter: evidence after the journal and the legacy republish, and the parent
state the records imply equals today's at every frame of a scripted session.
Through the production runtime: a Codex session's child work reaches the
status sink under its own address, and a provider exit ends it there.

* test(native-chat): a Codex child's new run never inherits the last run's open call

* test(native-chat): a Codex session with no child-work sink holds no evidence

* test(native-chat): deliver a Codex child's announcement twice, as Codex does, before counting edges

* refactor(native-chat): hand the Codex producer's pending edge over directly

* fix(native-chat): name every Codex turn state in the outcome map; type the runtime test's fake opener

* fix(native-chat): a Codex child's turn ends on the error that ends it, or on its thread closing

Codex can end a child's turn with no turn/completed: an error it will not
retry is that turn's own end (the verdict the transcript already settles the
same turn on), and a closed thread ran its last turn. The executions, the one
owner of child turn state, now end the turn on both, so the strip drops the
child and its record settles (failed, or unknown for a close) together,
instead of reading working for the life of the session. A systemError status
is not an ending: Codex raises it for errors that leave the turn running.

A child fact whose frame names no turn now belongs to the turn the child is
running, instead of counting for every run.

* test(native-chat): a Codex child's turn ending by fatal error or thread close settles strip and record together

* test(native-chat): the Codex parity script reads a waiting child through the shared fold's waiting arm

* test(native-chat): a Codex child row's journal attempt is its record's generation

The journal numbers a Codex child's runs by the turns it observed on the
child's thread; the host record numbers them by the runs its evidence
opened. Both are keyed by the child's own turn id, so they must agree run for
run, including when Codex reports the child's first turn before the spawn
that announces it.

* test(native-chat): a Codex session's end settles its live children and keeps the ended ones

The host no longer erases a session's children when its provider goes away: a
child still running settles with an outcome nobody reported, and a child that
had already ended keeps what it said. The producer tests now expect exactly
that, from the close path and from an unexpected exit.

* fix(native-chat): a Codex subagent's shell is its open tool until the process exits

Codex runs every agent shell through unified exec, so every subagent shell
arrives with the source the persistent-command tracker keys on. The producer
skipped those items, so a working subagent never named its shell, and an
approved command (started on the approval path, completed from unified exec)
stayed its open tool until the turn ended. The tracker still records the
process separately, so a command that outlives the turn reads as monitoring.

* fix(native-chat): a Codex shell becomes a subagent's own work only once it outlives its turn

Codex runs every agent shell through unified exec and never says when one is
left running, so the producer turned every shell, even a millisecond `rg`, into
a command record the moment it started. Each settled into the session's pool
of 32 settled records, so a busy turn evicted a finished subagent's record
(its outcome row would vanish) and listed dozens of finished shells beside it.

A command now becomes a record at the first turn boundary of the thread that
launched it while its process still runs: until then it is the agent's open
call. A shell that exits within its turn never becomes a record.

* refactor(native-chat): child records keep every settled child and can be removed outright

Settled child records now stay until the host drops the session's row; the
32-record trim is gone. A producer can say work stopped with nothing to
report, and its record (and the handles it answered to) goes instead of
settling. Evidence stays host-internal: the producer and the store share
one process.

* fix(native-chat): a Codex command is live work from its start until its process stops

The command tracker is now the one owner of a Codex command's lifetime. It
admits every command whatever `source` Codex tags it with (the approval
path starts one as `agent`), and ends it when its process exits, when its
thread closes (Codex stops the processes first, so no exit ever arrives),
or when the session ends. The producer mirrors that one-to-one: a live
record from the start, removed when the command stops, never settled.

This removes the turn-boundary rule: a command that was only recorded at
its turn's end left the parent reading done for one publish when the main
agent's turn ended with a shell still running. The parity script now
checks the parent at every journal write, not only at frame end.

* fix(native-chat): a Codex command whose approval its turn abandoned never ran

Codex starts an approval's command item before it asks, and when the turn
ends with the question unanswered (the user stops at the approval), it drops
the question and never completes the item. The command tracker admitted that
start as a running process, so the strip kept a phantom command row and the
session row read working until the session ended.

The prompt registry, which owns which approvals are still unanswered, reports
the command approvals a turn ended without; the tracker ends those commands
with the frame that ended the turn. An answered approval keeps its command.

* test(native-chat): start the Codex child-work runtime test without the removed hold

Main no longer has host.hold: creating the session starts its child, and
nothing a viewer does keeps it running. The test attaches and asserts the
one child that attach started, then drives it as before.
2026-09-28 14:56:48 -07:00
OrcaWinandm4air 813aff8f8a fix(opencode): stop OpenCode 2 loading a stale plugin from the retired shared hooks dir (#23500)
* fix(opencode): stop OpenCode 2 loading a stale plugin from the retired shared hooks dir

Before 1.4.209 Orca pointed OPENCODE_CONFIG_DIR at <userData>/opencode-hooks/shared
and wrote a server()-only status plugin there. 1.4.209 moved the plugin to OpenCode's
global config dir and 1.4.210 added the v2 setup() export, but nothing rewrote the
old file. Shells, daemon-persisted panes and OpenCode 2 background services that
still carry that OPENCODE_CONFIG_DIR load only that dir under OpenCode 2 (it replaces
the global dir), so the v2 loader rejects the stale plugin with "Plugin must export a
default definition with an id and an effect or setup function" and pane status dies.

- Refresh the plugin in the retired shared dir (only when it already exists and its
  content differs) so OpenCode processes started later from old shells load the dual
  v1/v2 export. Runs on OpenCode pane spawns and on any spawn that inherits the
  retired dir, even with agent status hooks off.
- Drop an inherited OPENCODE_CONFIG_DIR / ORCA_OPENCODE_* marker that points at the
  retired dir when building a new pane env, so new panes use global discovery.

Limitation: an OpenCode 2 background service already running from an old pane keeps
its cached copy of the stale module even after the file is rewritten (verified with
opencode2 v2.0.18). It must be restarted (`opencode service restart`); a restart from
a new Orca pane then picks up the global config because the env is stripped.

* fix(opencode): harden legacy plugin repair and inherited config cleanup

* fix(opencode): preserve daemon-owned user config during legacy cleanup

* fix(opencode): sanitize inherited sources and repair unseen legacy copies

* test(opencode): update shared PTY mocks for legacy repair

* test(opencode): annotate shared repair mock signature

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-28 14:55:42 -07:00
Neil b482b4d3b4 test: measure pointer gestures on the isolated visible display (#23678)
* test(claude): expect typed cancellation in queued-send settlements

* fix(ci): respect disabled terminal links and await browser recovery

* test(e2e): give legacy close client a profile authority

* test(e2e): account for frame pacing in pointer latency budgets

* test: compare pointer timing on isolated visible display
2026-09-28 14:34:53 -07:00
Neil e87772b3a4 test: retire a dormant worker through host-owned status (#23686)
* test(e2e): await renderer recovery after worker exit

* test(e2e): publish worker recovery through authenticated hooks

* test: keep retired background worker dormant before activation
2026-09-28 14:34:49 -07:00
Neil 8c61a5df1f fix(windows): require signed release binaries and identify CLI launcher (#23680) 2026-09-28 14:34:08 -07:00
Brennan Benson 29c7d5d983 fix(mobile): start AI-button agents through agent.launch, never a bare shell (#22762)
* fix(mobile): start AI-button agents through agent.launch, never a bare shell

"Fix checks with AI", "Resolve conflicts with AI", commit-failure recovery and
diff review's "New Agent Session" created a terminal with no agent and typed the
multi-line prompt into the shell, so each line ran as a shell command.

They now call agent.launchReplay into the existing workspace with the prompt;
the host picks chat or terminal from the user's default and delivers the prompt.
Hosts without the launch capabilities get the buttons disabled with update copy.

The agent comes from the desktop's own resolution (moved to src/shared). The
replay loop and capability read are shared with the workspace-create launch.

* test(mobile): repin bridged-parity tallies for the AI-button launch goldens

The corpus goes from 787 to 790 goldens: five shell-path goldens are removed and
eight agent.launch ones added; one lands in identical and two in
result-absent-settlement.

* test(mobile): re-record goldens for AI-button launches through agent.launch

Repinned baseline to 514ab7f868 and re-recorded all goldens. Against the branch
point: 781 header-only (baseline on all; adapterSha256 on the 48 goldens whose
adapter module changed; scenarioSha256 on 3), one body moved
(pr-triage-launch: createTerminal + terminal.send becomes agent.launchReplay),
eight added (the new launch outcomes and their reply matrices) and five deleted
(the shell-path scenarios and their matrices).

* fix(mobile): show review notes' agent launch progress and failures, once

"New Agent Session" left the sheet open with no progress for the whole launch
(up to a minute while a terminal agent readies), so a second tap started a
second agent, and a launch that never started or could not be confirmed
rejected an unobserved promise, showing nothing. The sheet now closes on tap,
the review screen says "Starting an agent...", one launch runs at a time, and
every outcome lands in the review screen's status line.

Marking notes sent now reads the screen state when the launch settles, so a
note written during the wait is not dropped by the whole-list save.

* test(mobile): re-record goldens for review notes' launch outcome on the review screen

Repins the corpus to 6f3018576d. One golden body moves:
review-create-agent-refused now fulfils with "Workspace not found" in the
review screen's status line and the sheet closed, where it previously
rejected an unobserved promise and left the status line empty. The other
789 goldens move only their baseline header.

* test(agent-status): drop the retired PR-triage terminal send from the identity inventory

The phone's AI buttons no longer create a terminal and send the prompt into it
(`createTerminalAndSendPrompt` is gone); the host's agent launch delivers it.
There is no terminal action consumer left in that file to pin.

* fix(runtime): publish saved source-control launch recipes to paired clients

settings.get is an allowlist and omitted sourceControlAi, so the phone never
saw an agent saved globally for "Fix checks", "Resolve conflicts" or commit
recovery and always fell back to the default agent. The host now publishes
the launch actions' recipes (agent, prompt template, agent args), normalized
so legacy saved defaults are already migrated. A new optional reply field:
older clients ignore it, and a client talking to an older host sees none and
keeps using the default agent.

* fix(mobile): ask to update Orca only when the host answered without agent launch

An unread or failed status read settles with no capabilities, which the AI
buttons read as an old host and showed "Update Orca on your computer". The
update copy now needs a status the host actually returned; an unread one keeps
the buttons disabled without blaming the desktop's version.

* fix(mobile): send an AI button's saved agent arguments with its launch

The desktop's direct launches for "Fix checks", "Resolve conflicts" and commit
recovery pass the action's saved agent arguments to agent.launch; the phone
honoured the saved agent but dropped its arguments. It now sends them the same
way: absent when none are saved, so the host keeps the user's configured
defaults. A host that predates the field ignores it.

* test(mobile): repin the RPC recording corpus after the launch recipe and availability fixes

Repins to fe85e0346f. All 790 goldens move only their baseline header: no
scenario saves agent arguments or reads an unreadable status, so no recorded
behaviour changes.

* fix(mobile): say the host status is unreadable instead of nothing when it is

With the update copy now reserved for a host that answered without agent
launch, an unread status left the AI buttons disabled with no explanation.
They now say "Could not read this host's status. Go back and reopen it.", the
words the mobile web shell already uses for the same failure; leaving the host
re-reads its status.

* fix(mobile): wrap an AI button's prompt in the action's saved template

The desktop renders every source-control launch's prompt through the action's
saved template (buildSourceControlRecoveryAgentCommandInput); the phone sent
its built-in prompt as is. Now that the host publishes the recipes, the phone
renders through the same shared function, refuses an empty result as the
desktop does, and offers the rendered text when it could not be delivered.
Review notes have no recipe and are unchanged.

* test(mobile): re-record goldens for the templated AI-button prompt

Repins to 7fd1555d20. One golden body moves: pr-triage-prompt-not-delivered
now carries the prompt as sent (rendered through the action's template) on its
prompt-not-sent result, which is what Copy prompt offers. The other 789
goldens move only their baseline header.

* fix(mobile): re-read a host status that failed while the connection stayed up

A status.get that timed out or was cut over settled the host's gates closed
and was never asked again until the connection state changed, so the phone's
AI buttons stayed disabled behind "Could not read this host's status" on a
link that was working. The gate still settles closed at once, so a failed
read never holds the host screen, but it now re-asks in the background with
the same backoff the runtime capability probe uses, and opens once a status
lands. A reply this app cannot decode is not re-asked.

* test(mobile): repin the RPC recording corpus after the host status re-read

The status gate change moves no recorded behavior: every golden's body is
unchanged and only its baseline header moves to the new pin.

* fix(mobile): show a launch's host warning as a note, not an error

A launch that went ahead can carry a host warning (a structured chat ignores saved agent
arguments, including the '' a template-only save writes). The AI buttons rendered it in the red
error line beside a success haptic. The notice now carries it separately as secondary text, and
review notes keep saying they were sent. Commit recovery also takes the synchronous in-flight lock
the PR triage buttons use, so two taps before a re-render start one agent.

* chore(mobile): record the host status re-read timer for React Doctor

The status re-read arms one retry timer from inside its read and clears it in the effect's
cleanup. React Doctor reports that self-rescheduling shape even in its minimal form, which failed
both changed-lines gates. Suppressed the same way as the session startup timers.

* fix(mobile): say review notes are waiting for the desktop instead of doing nothing

With no live connection, New Agent Session threw from a handler whose promise the sheet drops, so
the tap did nothing visible while the button stayed enabled (proven capabilities survive a drop).
It now closes the sheet and shows "Waiting for desktop..." as the other AI buttons do.

* fix(mobile): stop sending an AI button's saved agent arguments

Whether saved arguments apply depends on the route and shell the host settles after the request
(a chat ignores them and warns; malformed ones fail after admission), and the desktop sends them
only when they apply. The phone cannot know that, so it now leaves them out and the agent's default
arguments apply, as before this series. The saved agent and prompt template still apply.

* fix(mobile): mark review notes sent through the latest save

The sent marks after an agent launch went through the save callback captured at tap time, whose
rollback restores the screen from that moment, so a failed save could drop notes written during
the launch. It now uses the latest render's save, as it already did for the screen state.

* refactor(mobile): own the host status re-read outside the effect

The re-read loop lived inside the effect body, so React Doctor could not see its cleanup and
needed an inline suppression plus a config allowlist entry. The loop is now a plain function that
returns its stop handle, and the effect returns that handle, the same shape every caller of the
runtime capability probe uses. Both suppressions are removed; behaviour is unchanged.

* test(mobile): record the host descriptor from a background status re-read

Pins that the status read records the host descriptor when a re-read succeeds after a failed first
read, not only on the first answer.

* fix(mobile): show a PR AI launch notice only under the button that launched it

Fix checks and Resolve conflicts shared one error, warning and undelivered prompt, so a Fix checks
launch whose prompt was not sent also offered "Copy prompt" under Resolve conflicts, copying the
fix-checks prompt. Notices are now kept per button. The host availability notice stays under each
disabled button, since it explains why that button cannot be tapped.

* fix(mobile): say the host status is being retried instead of asking to reopen it

The host status gate now re-reads a failed status in the background, so "Go back and reopen it"
asked the user for a step that is no longer needed. The review sheet hint uses the same words.
The mobile web shell keeps its own copy.

* refactor(mobile): run the host status gate on the shared status probe

The gate had its own copy of the status probe's retry loop (same delays, same cutover and backoff
split, same stop on an undecodable status). The probe now takes an optional callback for each
failed attempt, which the gate uses to settle closed on the first failure, and the duplicate loop
and its now-unused reader are removed. Existing probe callers are unchanged.

* test(mobile): repin the RPC recording corpus after merging main

Re-records every golden against the merge commit and drops the three goldens whose
scenarios this branch removed, which the merge had restored from main.

* test(mobile): re-record the RPC goldens on the merge with main

Conflicted goldens were seeded from main and re-recorded against the merged
tree; every value either side recorded survives except main's terminal.send in
the PR triage launch, which this branch removes. Drops three goldens main still
had for scenarios this branch deleted.

* feat(mobile): confirm an AI button's agent started, naming the workspace

Fix checks, Resolve conflicts and commit recovery now show "Agent started in
<workspace>" under the button once the host started the agent with its prompt,
so a tap is no longer silent. The workspace label comes from the Source Control
panel and falls back to the branch.

* test(mobile): repin the recording baseline to the success-confirmation commit (header-only)

* fix(mobile): name the workspace in the diff review's AI-button confirmation

The diff review screen mounted the PR sidebar without a workspace label, so
"Agent started in ..." under Fix checks and Resolve conflicts named the branch
while the screen header named the workspace. The sidebar now requires the label
so no screen can drop it, and the diff review passes the one its header shows.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit, the last commit to touch a fenced
path. Against this branch before the merge, only header fields move:
baseline on every golden, and adapterSha256 on the 14 review-action goldens
whose adapter main now drives through the review sheet state. No recorded
body changed.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit, the last commit to touch a fenced
path. Against this branch before the merge only the baseline header moves,
on every golden; no recorded body changed.

* test(mobile): give the send-sheet stacking test the review controller's host status inputs

The merge with main brought in #22951's stacking test, which builds the review
controller without the host capability and status inputs this branch made
required, so the mobile test typecheck ratchet failed.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit 03995ae29d, the last commit to touch a
fenced path. Against this branch before the merge only headers move: baseline
on every golden, and adapterSha256 on the 14 goldens recorded through the
terminal adapter main changed in #23080. No recorded body changed, and the
merged corpus differs from main exactly as this branch did before.
2026-09-28 14:28:09 -07:00
Brennan Benson dbc4e21d9c fix(sidebar): move agent child disclosure to the right (#23577)
* fix(sidebar): move agent child disclosure to the timestamp slot

* fix(sidebar): align agent disclosure with summary caret
2026-09-28 14:27:29 -07:00
Neil 1a0ff42eb1 test: finish sidebar lazy imports before teardown (#23694) 2026-09-28 14:16:40 -07:00
Neil 0ceb3fa2af test(e2e): give wheel probes a running TUI fixture (#23691) 2026-09-28 14:12:24 -07:00
Neil 0f52bb8be5 perf(ci): use four ARM test workers and remove repeated compilation (#23685)
* ci: benchmark per-job Node compile caching on full unit shards

* ci: measure unit shards with three and four workers

* ci: benchmark localization extraction CLI patch

* perf(build): reuse identical relay bundles across platforms

* ci: compare Vitest 4 and 5 on complete ARM shards

* perf(ci): upgrade localization extraction to skip irrelevant syntax walks

* perf(ci): use all four ARM cores and remove benchmark workflows

* ci: preserve failures while capturing unit source revision

* fix(ci): preserve commented and escaped localization calls

* ci: remove corrected localization benchmark harness
2026-09-28 14:05:30 -07:00