Commit Graph
1053 Commits
Author SHA1 Message Date
Neil 76d480f808 Fix unsafe test fixtures and the Bun version pin (#26051)
* Keep test interruption signals within owned processes

* Pin Bun and add optional unit runner shutdown diagnostics

* Unblock CI lint without changing session host runtime

* Avoid duplicating runtime import-check dependency bundles

* Leave unit runner diagnostics disabled by default

* test: keep runner incident follow-up focused on durable guards

* test: apply transcript replacements as authoritative snapshots
2026-10-06 21:42:02 -07:00
Neil d3e1494674 test: bound memory used by the runtime Electron audit (#26049) 2026-10-06 20:21:34 -07:00
Brennan Benson 0bdcaf36ed fix(claude): start a Claude chat with its saved options and send the first message at once (#25152)
* fix(claude): end a Claude start that never answers initialize after 120 s

* Read the Claude startup deadline inside startup; fix a stale test comment

* fix(claude): start a Claude chat on its initialize answer, not on a frame only a SessionStart hook sends

Startup waited for system/init or a SessionStart hook frame as well as the initialize answer.
Before the first turn only a SessionStart hook sends one, and Orca adds that hook only through
its optional status hooks, so with them off the first message was held forever. Startup now
lands on the initialize answer; a start frame already seen is still checked, and one naming
another session ends a started session. The deadline drops to 90 s so it fires inside the
host's 120 s start wait.

* fix(claude): time the Claude start by silence, and fail it at once on another session's frame

Claude answers initialize only after its SessionStart hooks finish, so a total-time deadline
would fail every start behind a slow hook. Each start frame now restarts the clock. A frame
naming another session fails a start still waiting on initialize at once, as before.

* test(claude): a real Claude chat starts and answers with every hook disabled

* Say what the start-frame re-arm covers, and check only start frames in the hook-less real test

* fix(native-chat): a Claude chat starts with its saved options and takes its first message at once

Saved model, effort, Fast and permission mode are passed as launch options, checked against the
account's cached model catalog, instead of restored by control requests after initialize. With
nothing left to restore, the host no longer holds a message until the CLI answers initialize, and
the 90 s startup deadline is gone. A Stop on a start that never answers ends the child and settles
what it was handed as stopped. A failed result that repeats the turn's own API error reply writes
no second row.

* fix(native-chat): a host stop of a Claude start fails the message it was handed, with one row

With no start-hold the delivery loop no longer sees a host stop of a start it waited on. The
child's end now rejects what it handed over with the host-stopped words and writes the one row,
as an exit of its own would; an idle start the host stops still goes quietly.

* test(native-chat): a Claude chat's first message is written before initialize answers

Rewrites the tests that encoded the start-hold, the startup deadline and the option restore to
the new contract, and adds: saved options at launch (catalog checks, bypass, fresh-session Fast),
a message written before initialize answers (adapter and runtime), Stop on a start that never
answers (stopped, child closed, nothing working), and an API error said once.

* revert(native-chat): keep a failed Claude turn's error row

The shared turn fold already shows a failed turn's error once after it settles, as an error;
dropping the row left the CLI's synthetic reply looking like something Claude said.

* fix(native-chat): pass saved Claude options unchecked, heal a retired model on the CLI's word, and never leave an unrun message in doubt

- Saved model, effort and Fast are launched as picked; only values no Claude can parse are left
  out. The pre-spawn cache check is gone.
- A saved Fast on for a new conversation is applied once the settings readback shows no
  per-session opt-in (dropped when there is one, or when the model is listed without Fast), with
  nothing waiting on it; the record keeps the pick.
- Under an Agent Permissions bypass, a saved narrower mode launches with the allow flag so bypass
  stays reachable.
- A turn whose reply is the CLI's model_not_found for the launched model drops that model from
  the record; the launch's own row for it is kept out of the account model cache.
- A child that ends before it answered initialize, for any reason, settles every message it was
  handed as not sent (cancelled for a Stop).
- A launched effort the CLI reports only as `applied.effort` is confirmed from there.
- The untimed-initialize comment is back to main's text.
- A real-CLI test for a message written before initialize answers, under saved options.

* fix(native-chat): type the close's ended event and the start-exit test fixtures

The close's ended event is typed as the adapter event so its optional startupUnanswered spread
fits exactOptionalPropertyTypes; two tests guard the fixture's optional generation, and the
hung-start fixture records initialize on the fake connection it holds.

* fix(native-chat): a Claude model heal keeps a later pick, a refused Fast is dropped, no allow flag

- `options-skipped` carries the retired value; the record drops it only while it still holds it.
- `started` carries the values a heal retired, and the record does not take them back from the
  CLI's report of the same value.
- A saved Fast on a new conversation is applied before `started`: a refusal drops it and records
  it skipped, as main's refused restore did; silence keeps it wanted and unconfirmed.
- A saved narrower mode under an Agent Permissions bypass launches without any bypass flag again:
  the allow flag is one older CLIs reject at start. Kept as a known limit.
- The real-CLI test asserts the message was written before initialize answered.
- The fake reports a launch effort only under `applied`, and a misplaced doc comment moves back.

* fix(native-chat): a new Claude chat reports started before its saved Fast is applied

The Fast apply on a new conversation now runs after `started`, so a Stop interrupts a running
first turn and an option write is not refused while the round trip is out. A refusal drops the
pick through `options-skipped`, in order after `started`; silence keeps it unconfirmed. The
launch's unreachable skipped-model branch is gone.

* fix(native-chat): a healed Claude chat goes back to the default model live; comments match the no-hold design

When the CLI says the launched model does not exist, the live child is also put back on the CLI's
own default (set_model with no model, fire-and-forget), so later messages in the same chat run; a
user's pick sent after it wins, and a refused or unanswered reset only logs. Comments that still
described the start-hold or the option restore now describe the launch options and the
handed-over, never-echoed rule.

* fix(native-chat): a message handed to a Claude start that never answered is kept as main keeps an unsent one

A child that ended before it answered initialize ran nothing it was handed, the same fact as a
send accepted and never handed over. Its end now settles those sends exactly as the chat settles
a queued send for that end: a quit keeps a person's message as a held card (restart words), a
close keeps it as a held card (closed words), a person's Stop withdraws it as cancelled, and a host
stop fails the start with one row.

* fix(native-chat): a quit during a Claude start that never answered offers no resume for the message it keeps as a card

The restart snapshot now reads the same never-answered fact the exit does, so a message handed
to such a start counts as queued work, not as work to resume. The retired-model reset comment
names the default it really applies.

* refactor(native-chat): the saved permission-mode launch helpers live with the spawn options that use them

Keeps claude-structured-launch-resolution.ts within max-lines once merged with main, and names
the hung-start test envelope's field type.

* fix(native-chat): a Claude chat's saved options take precedence over the agent Arguments' own flags

Main now passes the saved agent Arguments to the Claude child, and the SDK writes them after its
own options. An Arguments --model or --effort therefore reached the CLI as a second flag after the
chat's saved pick (a commander CLI keeps the last), and a saved Fast's launch settings replaced an
Arguments --settings file outright. The saved model and effort now stand in for the Arguments'
flags, and a saved Fast beside an Arguments --settings is applied by the start instead of at
launch.

* test(native-chat): the hand-built Claude session in the options test carries fastModeAtStart

* test(native-chat): the queued rig's start spy carries a named Mock type

An unannotated vi.fn() inferred @vitest/spy's internal Procedure, which CI's typecheck cannot
name in the factories' inferred return types (TS2883).

* fix(ci): run the PR's SQLite-backed tests in the Node runtime project

Lists the hung-start Stop test, renames the send-during-startup entry from its old name, and
carries main's own two entries from #26010 so the boundary test passes before the next merge.
2026-10-06 19:19:47 -07:00
Brennan Benson b3b6c5dc13 Give native chat names one source for tabs, sidebar and AI Vault (list and search) (#25986)
* Give native chat names one renderer source and drop Vault's name repair copies

The host's saved conversation name now rides the structured session status
feed, which already exists per host, is keyed by the durable session id, and
keeps a closed chat's summary. Tab strip, sidebar rows and AI Vault (list and
search) read it through one hook and one display order (tab alias, saved name,
host label). Vault no longer copies names into its cached results, so the
projection, recovery and pending-title modules and their tab-snapshot lanes are
removed. Indexed search hits now carry the native owner and saved name from the
host that indexed them.

* Type the sidebar name test fixture without an assertion

* Keep Vault search working when the chat host will not install

Naming and owning search hits is bookkeeping: if the native chat host fails to
install, return the plain hits instead of failing the search. The runtime RPC
only installs the host for clients that will receive the owners.

* Publish chat names to the feed independently of the tab retitle

A failed feed publication no longer skips retitling the open tab. The publish
now lives in the naming deps, where a test covers it.

* Note why the status feed must keep closed chats' summaries

* Bound names and owner ids that come from a paired host

Drop a published chat name the record store would refuse, and cap a search
hit's owner workspace id at the same length the list row and record use.

* Let native chat search hits from a paired host open their chat

A paired host's search hits carry no resume command, so the row disabled
every open action even for a native chat it can open through its owner, as
its list row does.

* Ignore workspace ids that name object members in tab lookups

A paired host's row or search hit could carry a workspace id such as
"constructor", which read an Object.prototype member as a tab list and broke
the render. The shared tab index now reads only own workspace entries.
2026-10-06 19:04:04 -07:00
Jinwoo Hong 825d7bd5a9 test(vitest): run agent-launch-instant-tab in the SQLite runtime project (#26028)
#25430 added a test that opens a real agent-session record store, but not to the
SQLite runtime list, so vitest-sqlite-runtime-boundary fails on main.
2026-10-06 21:33:40 -04:00
Neil 3fb72d135d Run Node event-loop measurement after ordinary test suites (#26015) 2026-10-06 18:08:50 -07:00
Neil 37ff3873a0 Run combined localization catalog verification on Bun (#25999) 2026-10-06 18:03:47 -07:00
Neil f0520851ab Keep mobile restore SQLite fixture in the Node test runtime 2026-10-06 17:46:39 -07:00
Neil 1040f673b1 Keep orchestration SQLite fixtures in the Node test runtime 2026-10-06 17:46:39 -07:00
Neil 3308ff8b26 Bound YAML merge conversion and close SQLite routing review gaps (#25998)
* Close test runtime and YAML merge review gaps

* Bound YAML conversion inside explicitly tagged pairs

* Make completion notification fixture cadence deterministic
2026-10-06 17:31:56 -07:00
Jinwoo Hong faed899cd3 fix(ci): run three new SQLite-backed tests in the Node runtime project (#26010)
#25888 and #25766 added tests that import the orchestration database or the
structured session runtime without registering them in the Node runtime list,
so vitest-sqlite-runtime-boundary fails on main and every PR.
2026-10-06 19:48:41 -04:00
Neil d8c871a1f0 Speed up unit tests with cross-runtime duration scheduling (#25967)
* Schedule unit tests across runtimes by measured duration

* Keep new SQLite fixture suites on Node after updating main

* Inject scheduling timings instead of mocking the module

* Keep agent-session database lifecycle contracts on Node
2026-10-06 15:33:40 -07:00
Jinwoo HongandClaude 72b84118e0 test(e2e): terminal layout parity check against main (#25681)
* test(e2e): add terminal layout parity check for topology refactor PRs

Runs fixed terminal-layout journeys in the real app on two builds, captures
the renderer topology and the saved workspace session, normalizes volatile
values and fails on any difference not declared for a named bug.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(e2e): record quit exit status and report paths main does not reproduce

Co-Authored-By: Claude <noreply@anthropic.com>

* test(e2e): invoke pnpm correctly under corepack and silence the typeless-module warning

Co-Authored-By: Claude <noreply@anthropic.com>

* test(e2e): accept pnpm's forwarded -- in the layout parity runner

Co-Authored-By: Claude <noreply@anthropic.com>

* test(e2e): close parity panes through the user's chord and treat absent maps as empty

Driving PaneManager.closePane directly left main to learn of the close from the
PTY exit, which raced quit; the keyboard path commits the close in main first.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(e2e): build parity checkouts before running, forward -g, and add a drag-out scenario

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-06 18:26:49 -04:00
Brennan Benson 6aa12c30ef Name native Claude and Codex chats after their first message (#25724)
* feat: generate chat names through configured text agents

* Project structured chat names across stored tabs and session rows

* Name native chats from their first live message

* Restore the journal test provider handle import

* Fix first-message naming and live Vault title updates

* Preserve Unicode characters in bounded chat naming prompts

* Read chat naming settings only when a turn needs them

* fix(chat): store only generated conversation names

* fix(chat): keep ordinary labels across unnamed chat surfaces

* fix(chat): preserve naming after fast first turns

* Preserve native chat names across command-first sends and Vault lifecycle

* Route native chat SQLite contracts through the existing Node test pool

* feat(settings): add chat naming controls to Chat page

* Preserve chat naming drafts and configure Custom commands through host settings

* Scope synthetic command output to its journal thread in naming test

* Keep naming test fixtures within their typed project boundaries

* Resolve the real Vault hook directly in its integration test

* Publish chat naming save refs after render commits

* fix: keep Japanese chat naming copy stable during repair

* Update test contracts for the naming integration

* Keep journal fixture reads on the host clock

* Give Chat names a separate settings section
2026-10-06 14:42:59 -07:00
Jinjing f7c542c7a3 Keep profile saving alive after a stalled main loop (#25318)
* Keep profile saving alive after a stalled main loop

After a long main-loop stall (overnight sleep, dark wakes), the profile
writer's overdue 30s timeout could run before an acknowledgement that was
already queued, permanently retiring the writer until restart. Terminal
creation then failed because pane bindings could not be saved.

- Writer deadlines measure lateness on the monotonic clock and grant a
  bounded fresh window when the callback is overdue or the system reports
  suspended; resume re-arms without spending grace. Applies to
  initialization, every command, and the close/exit wait.
- The "Saving stopped" alert is parented to a visible main window (never a
  parentless synchronous macOS alert), deferred until shown, deduplicated,
  and says whether the latest change is unconfirmed.
- Timeouts, grace, writer faults, and alert presentation leave sanitized
  durable breadcrumbs.

* Fix profile writer timeout and shutdown races

* Run profile writer stall regression on Linux and Windows

* Keep Electron probes out of headless runtime qualification
2026-10-06 14:13:39 -07:00
Neil 83cf7cf5e2 perf(ci): compile release JavaScript once for all packaging hosts (#25828)
* perf(ci): share release JavaScript across packaging hosts

* fix(ci): verify the projected web entry in release archives

* fix(ci): use the Windows system archive tool for release bundles

* test(ci): retain stylesheet evidence in release build comparisons

* test(ci): verify release parity across native color rounding

* test(ci): normalize manifest asset references without changing import order

* test(ci): compare portable outputs across Windows text and color formatting

* test(ci): preserve module identity across dependent asset hashes

* fix(ci): keep SVG build inputs identical across release hosts

* fix(ci): stabilize compiler inputs and projected web bindings

* fix(ci): retain vendor minification in projected web output

* test(ci): normalize platform-specific pnpm manifest source paths

* fix(packaging): exclude shared build staging from application files

* test: align thinking-state fixtures with the current source shape

* test(mobile): reuse message fixtures within the line limit
2026-10-06 13:22:04 -07:00
Jinwoo Hong ac8ea9f958 fix(codex): write Orca's hook into ~/.codex only when something changed (#25743)
* fix(codex): write Orca's hook in ~/.codex only when something changed

One reconcile replaces the per-launch writer and its background approval
session. It runs at app start after PATH hydration, when the setting turns on,
once per native pane spawn, and (with a bounded 3 s wait) on Codex launches and
resumes. Orca's entry lives alone in a matcherless group, appended last unless
already in place; other copies are removed and the shifted user approvals move
verbatim under both key spellings. The approval, with Codex's own hash, goes in
first. The routing gate closes only for a hooks.json with a bad shape.

* fix(codex): report hook status for the home the next pane uses

Status reads ~/.codex (both key spellings) when launches use it, or the CLI
has no selection, and the selected managed home otherwise. It explains: update
Codex, Codex not found, not asked yet, approved by Orca but not yet confirmed
by Codex, a hooks.json shape that moves panes to Orca's own home, and inline
config.toml approvals Orca cannot add to.

* test(codex): pin the ~/.codex approval against a real Codex, and run it on those files

The contract now also writes Orca's entry into a throwaway ~/.codex and checks
that Codex lists it trusted and enabled, turns it back on over a /hooks
switch-off, lists it for review after a user inserts a hook ahead until the next
check, keeps an inline config.toml loadable, and shows no review in a real TUI
start. The real-binary CI job now runs when the ~/.codex writer changes.

* test(ci): keep the scope test under its line cap; assert the ~/.codex paths beside the contract job

* test(codex): type the hoisted test holders instead of asserting

* test(codex): cover the matcher rule, an older build's event, and the spawn trigger

Adds tests that failed against mutants which survived the first pass: an
entry alone in a matcher group, a current copy's approval in an event left to
an older build, one run per spawn, a spawn riding a running reconcile, and the
native-only spawn trigger. The opt-out's Codex-hash cleanup now runs once,
after the sweep, instead of twice.

* docs(codex): name the ~/.codex reconcile in the legacy sweep's lane comment

* fix(codex): approve ~/.codex with Orca's own hash while Codex's answer is pending

The reconcile waits at most 0.5 s for Codex's hash. If the lookup is still
running, an entry already in place keeps its approval and nothing is written;
a missing entry or approval gets Orca's computed hash, approval first, inside a
launch's 3 s wait, as main's did. When the lookup lands, the reconcile runs
again and Codex's hash replaces it.

* refactor(codex): one start for the hook lookup and the ~/.codex reconcile

startCodexHooks replaces the two start functions and takes the PATH wait that
the reconcile request's after option carried. A spawn reconciles unless one is
running, without a scheduled flag, launches call one reconcileCodexHooksForLaunch
on the shared withTimeout, and the reconcile alone reads the hooks setting. The
startup test now runs the ready phase instead of matching its source text.

* refactor(codex): plan ~/.codex with the shared Orca-hook pruner and approval reader

The planner prunes with removeManagedCommands, as the opt-out does (so a hook
that runs Orca's script through its args goes too), and returns a prune or
settle union. While Codex has not answered, the kept approval comes from the
shared per-event reader. The reconcile result is its outcome alone, a
concurrent save spends a pass of the one bound, the opt-out removes Orca's
approvals once, the approval-first writer takes the hooks path, and
getRealHomeConfigTomlPath gives way to getSystemCodexConfigTomlPath.

* fix(codex): ~/.codex keeps only an approval holding a hash Orca's entry may carry

* refactor(codex): name the ~/.codex hooks-file check for what it reports

* test(codex): the ~/.codex entry tests know Codex's earlier hashes, as the app's lookup does

* refactor(codex): the ~/.codex reconcile uses the lookup's answer names

* test(codex): ~/.codex status tests keep a Codex on PATH unless one is missing

* refactor(codex): one stopgap for both homes; the ~/.codex pass finds its own home, hash and user data

Also passes every approval to the approval-first writer, which already skips the ones in place.

* refactor(codex): the reconcile start owns the warm-up; no rerun-on-answer flag

The lookup start only records the hydrated PATH; one catch, one request type, and the app config is kept as given.

* refactor(codex): the ~/.codex pass returns nothing; tests read the files

Also makes the ~/.codex opt-out take Codex's hashes, as every production caller passes them.

* refactor(codex): the ~/.codex approval cleanup takes Codex's hashes

* test(startup): check the Codex hook start waits for PATH by behavior, not identity

* fix(codex): a status-hook problem never moves the system default off ~/.codex

* chore(ci): run the real-Codex contract when the ~/.codex reconcile changes

* docs(codex): the ~/.codex reconcile's refused branch covers every definitive refusal

* test(codex): start the failure-memo lookup with the reconcile's PATH-only start

* fix(codex): ~/.codex gets nothing while Codex is missing, a pending answer uses the saved one, a conversion waits for a run that writes, and status and the reconcile read one home and one Orca-hash rule

* test(codex): ~/.codex and managed-home status tests name their home; ~/.codex rewrites the backslash key Codex on Windows reads

* test(codex): the Windows upgrade test reads status for the managed home it installs

* test(codex): the real-TUI contract trusts its workdir by its real path, and its userData exists

* fix(codex): keep Orca's approval at every slot its entry holds in ~/.codex
2026-10-06 16:12:23 -04:00
Brennan Benson ab6389045e fix(native-chat): start Windows chats without reading process creation times (#25718)
* fix(native-chat): start Windows chats without reading process creation times

Native chat on Windows refused to start ("Orca can't run this agent in a chat
here") whenever the process-table addon could not report process creation
times. Chat never needed them; only the bookkeeping around stopping the agent
did.

Windows now follows the common pattern: Stop ends the agent's tree with
`taskkill /T /F` on the child Orca still holds, and reports the tree gone only
when taskkill exits 0. A saved pid is never signalled after a restart.

- Remove the Claude and Codex location gate and the Codex launch refusal.
- A start time that cannot be read is recorded as unknown instead of refusing
  the session; recovery already releases such an owner without signalling it.
- Delete Claude's Windows creation-time descendant snapshot and its verifier.
- Codex's Windows teardown reports the real taskkill outcome.
- The renderer no longer waits on the capability flag; the host keeps
  publishing it for older clients (temporary).

* fix(native-chat): treat a Windows Claude exit after stdin end as a proven close

On Windows an idle Claude leaves on its own once its stdin ends, so every Stop
and close read as unproven: live background work settled as unknown and the
trace logged a close that "did not finish cleanly". As with the Codex close,
that exit is now the close and Orca makes no claim about processes Claude
started; a forced close still rests on taskkill's own report.

Also drop the Settings clause about Windows needing process start times, and
fix comments and the tracked process-enumeration doc that still described the
removed start-time gate and descendant snapshot.

* fix(native-chat): renew a held child's lease without a PID probe

An owner recorded without a process start time could never renew: the renewer
re-proves every live owner by PID identity, and with no start time and no
spawn-token echo that probe is indeterminate. The lease's last renewal then
stayed at the spawn, so a turn cut by an Orca crash was dated to its own start,
and every tick logged a failed renewal and split the batch into one store
transaction per chat.

The runtime now renews a lease for a child it still holds at the record's
fence and whose exit it has not received, recorded as a `held-child` match.
Receipt of the exit ends that proof before the exit is settled, and a restart
holds no child, so a dead owner's lease still expires and restart adjudication
is unchanged. Records this runtime does not hold keep the PID probe.

Also correct the identity probe's comment about shipped addons and note that
recovery's stop ladder is POSIX-only.

* fix(native-chat): keep a failed Windows taskkill unproven across a retried close

A Windows close counts Claude leaving on its own after its stdin ends as the
close. That shortcut also caught a retried close whose first attempt forced a
taskkill that failed: once the root exited, the retry returned true and the
failure read as a proven close. The tree reaper now records whether a reap
ever reached the live root, here, on an earlier close or from a transport
failure, and the shortcut applies only when none did; otherwise taskkill's
verdict stands and no new taskkill runs against the exited root.

Pin the platform on the existing tests that assume the POSIX close, and say
what `exit-proven` means on Windows in the acquisition-failure docs.

* fix(native-chat): derive a held child's liveness from the adapter's own handle

Lease renewal trusted a held child until the host settled its exit, and the
host hears of an exit late: Claude runs its close ladder and a store write
first, unexpected exits wait on one delivery chain shared by every chat, and a
Claude close that cannot prove its tree publishes nothing at all. A dead root
could keep renewing through that window, so a crash in it dated the cut turn
late. Renewal now asks the adapter, which owns the process handle and sees the
exit first: a child is held only while it is on record at the lease's fence
and its adapter still runs that exact acquisition with no root exit seen. An
adapter that cannot answer falls back to the PID probe. The stored
exit-received mark is gone.

The held-child read is now required by the runtime state and the renewer, and
a host-level test drives renewal through the real host wiring.
2026-10-06 12:09:58 -07:00
Jinwoo Hong 8d049b594d fix(codex): approve Orca's hook in managed Codex homes with Codex's own hash (#25742)
* feat(codex): ask Codex for its hash of Orca's hook in a throwaway home, cross-checked by position and path

* feat(codex): cache Codex's hook hashes per binary and version, asked one at a time and only by the app

* feat(codex): write a hook approval before its entry, and take back only its own on failure

* feat(codex): approve Orca's hook in managed Codex homes with Codex's own hash, written first

Managed homes (the shared mirror and per-account homes) no longer run a
background approval session. Status reads the home's files against Codex's
answer, and turning hooks off recognizes every saved version's hashes.

* feat(codex): managed homes approve Orca's hook with Codex's hash; drop their background approval

The previous commit carried only the managed resume's wait; this one holds
the managed install it relies on. Managed homes (the shared mirror and
per-account homes) write Codex's hash before the entry, fall back to their
own approvals when the answer is late, and strip Orca's entry only when
Codex itself answered with nothing to approve. Status reads the home's files
against Codex's answer, and turning hooks off recognizes every saved
version's hashes.

* feat(codex): only an Orca-launched Codex waits up to 3 s for the hook hash; warm it after PATH hydration

* feat(cli): name the file each agent hook status reports on

* test(codex): real-Codex contract for the derived hook hash in managed homes, on both pins and latest

* test(codex): type the hook-hash test fixtures and drop a duplicate import

* fix(codex): give a Codex launch its own install run instead of joining a plain terminal's

* test(codex): a user hook's approval stays put in an event Codex does not list

* test(codex): cover late answers, first-install mirroring, stale approvals and opt-out re-asking

* test(codex): a long managed home gets the daemon guard on its first install

* fix(codex): until Codex answers, approve a managed home's hook with Orca's own hash, as main did

A late, temporary or missing answer with no earlier approval in the home now
writes main's self-computed approval instead of leaving the hook out. Codex's
answer replaces it at the next install, a definitive answer (no hooks/list,
a refused cross-check, 0.128) never uses it, and status says the approval
is Orca's until Codex confirms it.

* test(codex): status flags an unapproved entry while Codex has not answered

* fix(codex): managed stopgap fills each missing event

Until Codex answers, a managed home kept only the events it had already
approved and dropped Orca's entry from the rest. Each event now keeps the
home's approval, else gets Orca's own hash, as main wrote every event. One
reader of the approval at Orca's entry serves the stopgap and status.

* refactor(codex): one Codex answer type, one in-process answer map, a disk-only memo

- One answer type with a kind (hashes, refused, pending) replaces two types
  and the three-field decoding at each caller.
- The lookup keeps one in-process answer per binary path, replacing the
  process memo, the global latest answer and the transient-failure map;
  status now reads the answer for the codex on PATH, not the last one asked.
- The memo file keeps Codex's refusals per version, like its hashes.
- Derivation takes the version it is given; one 30 s version-probe timeout.
- The launch wait reuses withTimeout, and launch prep passes launchesCodex
  down instead of a wait in milliseconds.
- Turning hooks off no longer forgets Codex's answer.
- Tests mock the derivation instead of a test-only resolver in production.

* chore(codex): list the approval reader for the CLI build; fold two identical scope checks

* fix(codex): count an approval at Orca's key only when it holds a hash Orca's entry may carry

* refactor(codex): one append for hook trust tables

* refactor(codex): the lookup keeps no entry for a missing Codex, and status checks the binary's fingerprint

Also names the lookup functions for the answer they return.

* refactor(codex): one stopgap reader for the managed home; the refused branch reads its own status

* refactor(codex): drop defaults and exports only tests relied on

* fix(codex): a failed ask of Codex stays pending instead of refusing its version

* chore(ci): run the real-Codex contract when the approval reader changes

* fix(codex): only a scratch home Codex loaded can refuse; the memo takes any hash and writes only on change

* fix(codex): hooks turned off during a launch's wait win, Off re-keys mirrored user approvals, and one rule says which hashes are Orca's

* fix(codex): an approval counts only under every key spelling Orca writes, as Codex on Windows reads only the backslash one
2026-10-06 14:15:35 -04:00
Brennan Benson 468e4e1167 fix(native-chat): a prompt card owns the chat input until its answer lands (terminal-backed chat, desktop and phone) (#25761)
* fix(native-chat): an answerable prompt card owns the chat input until its answer lands

* fix(mobile): a terminal chat's composer waits while its prompt card is up

* test(native-chat): type the prompt card fixtures without casts

* fix(native-chat): scope replies to acknowledged prompt occurrences

* fix(native-chat): preserve answer ordering and verified delivery

* test(native-chat): keep mock RPC client inside test boundary

* test(native-chat): place mock fixtures in the test-only scope

* Keep runtime comments within the module size limit

* test: preserve prompt delivery coverage in desktop CI

* Treat an older host's accepted write as delivered

A newer desktop or phone talking to a host that predates the write
settlement field read every accepted reply as "unconfirmed". Prompt cards
never dismissed, the phone showed "Response unconfirmed" on every tap and
ordinary chat messages were held as "Delivery unconfirmed".

The reader now uses writeSettlement when present and otherwise keeps the
host's whole-write accepted/refused verdict, exactly as before this branch.
Only prompt answers ask for provider settlement; ordinary callers
(follow-up delivery, paste drafts, option commands, composer sends) are
back on the original contract, so the legacy-handoff error class, its
flag, the sequence-only send helper and the mobile handoff hook are gone.

* Keep terminal-pane Escape on the plain accepted write

Every pane's Escape/Ctrl+C goes through pty:writeAccepted. This branch had
switched that IPC to wait for provider settlement, which dropped the
"remount this pane" signal for a daemon session awaiting recovery and could
stall later keystrokes behind a slow daemon acknowledgment.

pty:writeAccepted is back to its original local-only, synchronous write.
Prompt answers opt into settlement with requireWriteSettlement on the same
channel, and a settled refusal while the daemon recovers now sends the same
remount signal. Ordinary verified sends regain their original fallback write.

* Report a partly accepted local paste as unconfirmed

A settled local write split into chunks returned plain false when a later
chunk was refused after earlier ones were accepted. Callers read false as
"nothing was written", so chat showed "Message not sent" with a prefix
already in the agent's input. It now reports the write as unconfirmed,
the same verdict the paired host gives for a partial write.

* Hide the chat composer under a prompt card instead of unmounting it

When an approval or question card took the input region, the composer
unmounted. A message still waiting for its Enter was cancelled and its
bubble deleted after the draft had already been cleared, so the message
vanished without a notice; composer history was also wiped each time.

The composer now stays mounted but hidden while a card owns input, so its
state survives. A send that has not submitted yet is still stopped (its
Enter would answer the card), but its bubble stays with "Message not sent"
so the text is not lost. The composer ref is detached while hidden, so
root typing, paste and reveal focus never reach it.

* Keep an answered prompt hidden after the chat view remounts

The "answered" dismissal lived in component state. Toggling chat to
terminal and back, a PTY reconnect, or leaving the phone session and coming
back while the approved tool was still running brought the answered
approval back, and it then took over the input again.

Desktop now keeps the answered occurrence per pane outside the view;
phone keeps it per chat tab outside the controller. Both still retire it
when the pane observes the prompt clear or change, desktop also when the
tab retires, and both maps are size-bounded.

* Update the prompt-reply reliability gate for the review fixes

Older hosts' accepted answers now dismiss like acknowledged ones, the
composer stays mounted under a card, and answered prompts survive a view
remount. The gate's invariant, oracle, assertion list, new test files and
the two latest evidence runs now describe that contract.

* Let users hide a prompt card, keep Escape from denying, and gate only Send on the phone

The chat input could stay locked behind a card the host never closes (for
example after a Deny typed in the terminal), and Escape on a focused
approval card denied the tool even when the user meant to close a picker.

- A Hide control (chevron) on terminal approval and question cards, desktop
  and phone, hides that prompt occurrence and gives the input back. It writes
  nothing to the agent and uses the same per-occurrence dismissal as an
  acknowledged answer, so a new occurrence shows the card again.
- On desktop, Escape on a card now does the same Hide instead of Deny, and a
  card that appears while the user is typing no longer takes focus.
- On the phone, a card blocks only Send: typing, dictation and image attach
  keep working on the draft. The placeholder is back to the normal one.

* Fix two comments that still called older-host replies unconfirmed

Since an older host's accepted write now counts as delivered, the
requireWriteSettlement comment and the reliability gate's oracle said the
opposite of the code. Both now describe the current rule.

* Collapse prompt cards to a strip instead of hiding them, and close the round-2 gaps

Hide removed a card completely, so nothing on screen said a prompt was still
waiting, and several edges let the chat type into a live prompt.

- Collapse (the header chevron, or Escape on desktop) folds the card to a
  one-line strip above the composer; the strip's chevron expands it back.
  Collapsing writes nothing, frees the composer, and is disabled while an
  answer is still being written. Each pane or tab keeps the occurrence as
  answered or collapsed, so a remount restores the same view.
- Questions now carry the host wait's start like approvals, so an identical
  question in a new wait shows again (desktop and phone). A transcript-only
  prompt, which has no wait start, is dropped when the view stops observing
  it, and a transcript still loading no longer clears a dismissal.
- Desktop: while a card owns the input, the hidden composer cannot send or
  interrupt even if it still has keyboard focus, and the card takes focus in
  the same commit. A send the card retires no longer types Ctrl+U under it.
- Phone: an Ask hides the heuristic card read from the same waiting status,
  and the dismissal store is scoped by host, worktree and tab.

* Keep a collapsed card's partial answer, and scope its focus to its own pane

Collapsing a question card unmounted it, so expanding it again lost the
chosen step, selections and typed "Other" text; Escape typed in that text
field collapsed the card. A card arriving while the user typed in another
surface (sidebar, notes, a browser URL bar) also took the keyboard.

- The collapsed card now stays mounted but hidden (and inert on desktop)
  under its strip, on desktop and phone, so a partial answer survives
  collapse and expand. Escape inside the card's text field no longer
  collapses it. The question card shows the same focus ring as the approval
  card.
- A card takes focus only from inside its own pane (its hidden composer) or
  from the page body, never from a text field elsewhere.
- Desktop and phone share one dismissal store in src/shared, bounded by the
  existing scope-cache helper, which moves to src/shared with it.
- The card send imports the verified helper from its own module, and the
  phone files are split so each name matches its contents (header action,
  strip, lane selector).

* Return focus to the composer after a prompt card collapses

Since a collapsed card stays mounted, Escape or the chevron left keyboard
focus inside the now hidden, inert card. The composer's reveal-focus took
that as focus already in the pane and stood down, then the browser dropped
focus to the page body, so typed keys went nowhere.

Reveal-focus now treats focus inside a hidden or inert subtree as not in the
pane and focuses the composer. On the phone, collapsing a card also
dismisses the keyboard so a hidden reply field does not keep it.

* Keep the question card's collapse chevron beside its Cancel button

The question card header spread its three items with justify-between, which
put the new chevron in the middle of the header. The title now takes the
free space, as in the approval card, so the chevron sits next to Cancel at
the right edge.

* Run the prompt tests on the merged main

Main now runs Vitest under Bun, which resolves a long data: URL import as a
package name, so the SSH delivery test loads its bundled mobile module from a
temp file instead. The phone prompt harnesses mock the live line that main's
view now renders, and add Platform, which main's text-selection helper reads,
the same way main's own view tests do.
2026-10-06 10:57:58 -07:00
Neil 3ec38b8c6d Run Vitest on Bun with Node runtime contracts (#25840)
* Run Vitest on Bun while preserving Node runtime contracts

* Preserve runtime timing provenance and keep the Bun pin in config

* Scope builtin compatibility mocks to test-only lint exceptions

* Give capture retention fixtures distinct filesystem timestamps

* Await the copy button success state in the React fixture

* Bound Node test worker shutdown and tighten migration fixtures
2026-10-06 03:18:06 -07:00
Neil 13ea35973c Stop expensive checks when an unmerged PR closes (#25829)
* Cancel active checks when an unmerged PR closes

* Register owned-branch cancellation qualification

* Keep temporary cancellation qualification outside the review diff
2026-10-06 01:03:24 -07:00
Neil f6f96db6be Build SSH hostile-host Linux slots independently (#25821) 2026-10-06 01:02:53 -07:00
Brennan Benson eaaae0196f feat(native-chat): record fresh sessions after failed restoration (#25747)
* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore

A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.

Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.

* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row

Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.

* test(native-chat): prove replacement rows survive downgrade and re-upgrade
2026-10-05 23:22:24 -07:00
Neil f9c8cd4fc3 Reduce redundant test coverage and unnecessary CI waits (#25806) 2026-10-05 23:19:18 -07:00
Brennan Benson 020cebeff6 Add standalone Agent Client Protocol client layer (#24990)
* Add standalone ACP protocol client and session runtime

* Protect ACP transport teardown from late stream errors

* Retire incoming ACP request ids before publishing responses

* Narrow ACP configuration requests and transport message types

* Remove redundant ACP request handler return unions

* Keep ACP waits caller-owned and preserve protocol extensions

* Preserve open ACP decisions through prompt completion

* Generate open ACP enums and check the generated schema offline

A newer or vendor enum value (tool kind, tool status, option kind, stop
reason) no longer fails the whole message: generated enums accept the known
literals plus any other string, typed so callers can still narrow on the
known ones. The generated header now records the pinned input digests, the
generator digest and a body hash, so `verify:acp-protocol` catches a stale or
hand-edited file without network access; it runs in lint and the PR workflow.

* Land the ACP runtime contract the agent adapters use

- Deliver notifications other than session/update through
  onExtensionNotification, in arrival order with session updates.
- Accept _meta on prompt, setMode, setModel, setConfigOption and cancel.
- cancel() always sends session/cancel once the session runs, since the
  agent can be in a turn it began itself; only a successful send is shared,
  so a failed write is retried.
- Cancel aborts each open agent request's signal and lets its handler send
  its own answer; -32800 only when the handler rejects.
- Permission requests validate only the session, tool call id and options;
  unreadable fields are dropped with a diagnostic, and any answer Orca
  cannot send is `cancelled` instead of a JSON-RPC error. Agent-started
  turns may ask; whether to show it is the caller's decision.
- AcpAgentError marks the agent's own errors; AcpInvalidResponseError keeps
  the raw answer and validation issues for answers Orca could not read.
- Lines over the size limit are classified by prefix (shared with the Codex
  reader): the owed request fails, an oversized agent request is answered
  with an error, and an unattributable response closes the connection.

* Answer every agent request after an ACP cancel

A cancel that lands before a permission handler starts now still runs the
permission path, so the agent gets the `cancelled` outcome rather than a
request-cancelled error. A handler that ignores the abort no longer leaves
the agent waiting: once the abort has run through, any request still
unanswered gets request-cancelled. Handlers that answer on abort keep their
own reply.

Also renames a lint-rejected helper parameter, replaces a Reflect.apply in a
test, and stops the permission diagnostic from firing with an empty list.

* Let each ACP request handler own its answer after a cancel

Removes the next-event-loop-turn fallback that answered request-cancelled
for any handler still silent after a cancel. It raced answers that were
still being saved (an approval mid-journal-write reached the agent as an
error) and made the outcome depend on event-loop timing. The handler that
owns an agent request now always sends its answer, or throws for
request-cancelled; a request it never answers ends when the connection
closes. A permission whose handler had not started still answers
`cancelled`.

* Register the ACP schema verify step in the PR preflight phase test

* feat(acp): a steer's cancel asks once and never ends the agent

The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.

* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns

Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.

* test(ratchet): require src/main/acp now that this PR lands it
2026-10-05 22:25:40 -07:00
Neil ca4e239861 Remove low-value test inventories and duplicate fuzz oracles (#25791) 2026-10-05 22:23:48 -07:00
Brennan Benson 6c693edf40 Show native chat tool calls as plain sentences (#25654)
* feat(native-chat): add chat-scoped color tokens

* feat(native-chat): soften transcript and composer appearance

* fix(native-chat): refine code spacing and faint text styling

* fix(native-chat): wrap prose links at word boundaries

* test(native-chat): refresh background task strip snapshots

* feat(native-chat): show tool calls as plain sentences

* fix(native-chat): make tool sentences reflect call state

* fix(native-chat): clarify failed commands and subagent sentences

* test(native-chat): exercise command disclosure with real results

* fix(native-chat): preserve command inputs and localize failure rows

* fix(native-chat): retain complete padded command input

* fix(native-chat): read named tool input fields safely

* Refresh native chat tool rows when the UI language changes
2026-10-05 22:05:13 -07:00
Jinjing 059e81a106 chore(i18n): use 智能体 for Chinese Agent copy (#25767)
Simplified Chinese rendered the Agent concept as 代理, which collides with
代理 = proxy. Standardize on 智能体 for Agent (and 子智能体 for subagent),
while keeping 代理 for proxy senses: HTTP/network proxy, SSH Proxy Command,
reverse proxy, and browser user agent.

- Converted 131 zh catalog values (incl. 子代理 -> 子智能体); 26 proxy /
  user-agent values left as 代理.
- locale-phrase-fixes.mjs: 客服人员/代理商/座席 -> 智能体, 代理 -> 智能体 when
  the English names an agent (guard excludes "user agent"); removed the old
  智能体 -> 代理 rule so the pipeline no longer reverts it.
- Updated value/key/search/macos-tcc overrides to 智能体; proxy keyword and
  proxy override entries unchanged.
- Updated the two policy tests that pinned the old 代理 output.

Gates: catalog verify, coverage --check, extraction, runtime-catalog, and
the locale vitest suites (296 tests) pass.
2026-10-05 20:19:43 -07:00
Neil d73efccc7d Reduce localization audit and relay setup work in CI (#25665) 2026-10-05 19:37:11 -07:00
Brennan Benson b01df814e6 test(ratchet): require src/main/provider-process now that it has landed (#25710) 2026-10-05 16:38:02 -07:00
Kelvin Amoabaandmmarabel 1168e0f8c8 fix(ssh): let a placed worktree seed while its host is in conflict (#23213)
One host tab on a folder this client lacks marked every worktree on the host unverifiable, so reopening an emptied one never got a terminal.

Fixes #22015

Co-authored-by: mmarabel <166927047+mmarabel@users.noreply.github.com>
2026-10-05 16:22:17 -07:00
Brennan Benson d791421568 Extract provider process supervision and stream reading from Codex (#24989)
* Move provider process supervision and stream reading out of Codex

* Preserve teardown behavior with checked mock types after move

* Apply provider launch environment and caller teardown labels

* Give provider child env one owner and gate Codex contract on the shared reader

resolveProviderChildEnv is now the only place that overlays and strips a
provider's environment; the spawn spec and the request-scoped Codex session
both call it. supervisedPosixLaunch only accepts a launch without env fields,
so an override can no longer be silently ignored there. Edits to the shared
stream reader or the env rule now run the real-binary Codex contract job.
2026-10-05 14:49:17 -07:00
Brennan Benson 7208be6902 test(native-chat): cover paired runtime launch compatibility (#25072)
* test(native-chat): cover paired runtime launch compatibility

* test(native-chat): use production prompt response fingerprints

* test(native-chat): pass journal items to the paired-runtime outbox hook

Main made journalItems a required outbox input; mirror production by
passing the read state's items.

* test(native-chat): pin paired-runtime routing and capability questions

- Move the paired-runtime chat test into its own file and render the real
  session hook against a paired server that advertises today's capabilities
  while this machine advertises none, so Stop, prompt cancel, repeated Stop
  and question answers fail if their capability question goes to the wrong
  runtime. Add a local and a paired session sharing one id.
- Cross-version: state that the file pins the capability strings a
  v1.4.219 server really advertises, use the desktop's real handshake list,
  and keep the current server only as the control for the refusal.
- agent-launch-routing: give the released-server case a capable client so
  it is refused by the server gate, not the client one.

* test(native-chat): route the paired launch suite and cover the outline read

- Start the cross-version job when the launch route or the client list the
  desktop sends a paired host changes, since the paired launch suite runs
  them.
- Check the rail's conversation outline is read from the paired server, and
  narrow the hook test's header to what it covers.
2026-10-05 14:48:52 -07:00
Brennan Benson 508419f11e test: check structured-chat code for Electron imports even before the runtime loads it (#24988)
* test: keep structured chat free of Electron imports

* test: widen runtime Electron ratchet to structured chat

* test: check whole structured-chat directories for Electron imports

Cover src/main/{native-chat,claude,codex}, src/shared and every structured-*
or agent-session-* file under src/main/runtime, so new files are checked by
default. A lane that exists today now fails loudly if it goes missing; only
the not-yet-landed acp/ and provider-process/ may be absent.

Move the test-only file rule into one classifier shared with the
localization audit, so -test-<thing> helpers and test doubles no longer
enter the production gate.

* test: run the Electron-import CLI path in tests and retire may-be-absent lanes

Export main() so the real-tree tests exercise the entry list CI uses,
check every required lane for the missing-directory error, and fail once
acp/ or provider-process/ exists while still allowed to be absent.
2026-10-05 14:48:33 -07:00
Brennan Benson e817b0e237 refactor(native-chat): keep the provider resume handle opaque to shared code (#24991)
* refactor(native-chat): keep the provider resume handle opaque to shared code

Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).

Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.

The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).

No user-visible change.

* fix(native-chat): derive journal-row provider handles from the journal identity

The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.

* fix(native-chat): refuse a stored provider handle written in both forms

A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.

* test(native-chat): use opaque handle in queued rejection fixture

* test(native-chat): share one Codex journal identity in the integration suite

Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.

* refactor(agent-session): name the handle's adapter state resumeCursor

Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.

State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.
2026-10-05 14:48:12 -07:00
Brennan Benson 761d63a4e5 feat(agent-launch): keep long prompts off the launch line and paste them after readiness (step 1 of 7) (#24257)
* feat(agent-launch): host-side prompt delivery for agent.launch

The host's agent.launch typed any launch prompt into the shell as part of the
launch command. A long or multi-line prompt then ran line by line in the
shell, and an agent that never showed readiness or crashed at startup had
nothing guarding where its text went.

agent.launch now carries a prompt on the typed line only when the line stays
one line, control-free and at most 512 bytes; otherwise the agent starts
clean and the host pastes the prompt once the agent's own ready signal fires
(bracketed paste plus its composer marker or a quiet render, read only after
the shell's last hand-off, never while the pane's own shell is proven in
front), with main's draft-paste bytes and an Enter 50 ms later. Orchestration
worker starts wait on tui-idle as before. A replay-safe launch admits and
claims its ledger row in one write, Qwen Code gets a second Enter, the
desktop and phone share one launch-refusal classifier, and hosts advertise
agent.launch.prompt-carry.v1.

Split out of #23748, which moves the desktop source-control buttons onto
this path.

* fix(agent-launch): keep a short-lined multi-line prompt on a local zsh launch line, as main did

#24257 moved every multi-line or over-512-byte prompt off the typed launch line and pasted it after readiness. The phone's AI buttons and review notes, whose multi-line prompts main typed whole into zsh, then reached Claude 0.5-3 s later and their RPC reply waited for the paste.

The host now names the shell a local macOS or Linux line is typed into, the way the spawn picks it, and a multi-line prompt rides a zsh line when every line is at most 512 bytes and the whole line at most 8 KB. A real-zsh test types such a line through Orca's own ready barrier and startup write, including when a slow user config makes the write land early. Elsewhere the measured unsafe cases keep the paste: bash 3.2 runs multi-line lines piecemeal, fish drops an early multi-line write, and any shell loses a line over 1 KB written early.

* test(agent-launch): keep the real-zsh launch-line test out of the Windows lane's gate scan

The Windows lane registration check read `const ZSH_PATH = process.platform === 'win32'` (the head of a multi-line ternary) as a Windows-true flag, so `describe.skipIf(!ZSH_PATH)` looked like a Windows-only suite. The file is POSIX-only; the zsh lookup is now a function.

* refactor(protocol): move the agent.launch capabilities into their own module

Main's protocol-version.ts sits at the 300-line cap, so the prompt-carry capability pushed it over. The four agent.launch capabilities and their doc move to agent-launch-runtime-capability.ts, re-exported by name and spread into RUNTIME_CAPABILITIES at the same position; the advertised lists and every export are unchanged.

* refactor(protocol): import the agent.launch capabilities from their own module

`export *` from protocol-version left the four names undefined under the mobile recording loader, which resolves a relative import through a Proxy with no own keys, so 37 phone recordings lost agent.launchReplay. Importers now name agent-launch-runtime-capability directly; protocol-version only spreads its list.

* refactor(agent-launch): drop the unshipped viewMode field and trusted local caller id

Both were inert in step 1 and existed only for step 2. agent.launch will become a
public plugin API, so every wire field is permanent once shipped; a top-level
viewMode reads as "choose terminal vs chat", which the host decides. Step 2
introduces placement and view intent under a placement object instead.

* fix(agent-launch): read Codex's provisional startup from the rule files' hold anchor

Main (#24375) moved Codex's provisional-header check into codex.json's
provisional_startup hold anchor and deleted codex-terminal-readiness.ts, so the
launch readiness hold now asks showsHoldAnchor, as main's own settled check does.

* fix(agent-launch): hold rule-file name titles to quiet for a launch, and census the zsh fixture

Main (#24375) answers a name-only title from each agent's rule file ahead of the
sustained-title lane, so gemini.json's name_title settled a launch readiness wait
on the shell's auto-title while Gemini was still booting. A launch now asks quiet
of every weak idle verdict, as that lane did.

Main's readiness census requires a recorder for every runtime fixture; the zsh
prompt recording is a non-agent control. Gemini's synthetic baseline is
regenerated for this PR's stated change: a bare gemini title is no longer its
rest mark, so name-only rows settle weak, and a fresh working or blocked status
is no longer overridden.

* fix(agent-launch): paste a launch prompt only when the launched agent is proven in front

A launch pasted its prompt unless a shell was proven in the terminal's
foreground, so any read that could not prove one let the prompt through. After
an agent exited at startup, its shell turned bracketed paste on at the next
prompt, readiness fired on it, and the prompt was typed into the shell:

- macOS: a pane runs its shell under login, so the process-group fence's root
  was never the shell's group and never proved it; the cached foreground name
  could also still name the exited process.
- Windows Git Bash and WSL: the shell-alone-in-its-job check never answers.

Now one fresh read of the terminal's foreground decides: agent, shell or
unknown. Only 'agent' lets a write through (paste, Enter, second Enter, reused
panes too); 'shell' still drops a ready signal. A Windows host never proves the
agent, so there the launch line carries the prompt at any size, as on main.

* test(agent-launch): cover the Windows QA stub, a grok override that exits at once

* fix(agent-launch): keep the local socket alive while a prompted launch waits for its agent

A launch with a prompt now waits up to 60 s for the terminal agent to be
ready before it writes the prompt, and reports not-delivered when the agent
never is. The local runtime socket closes a connection idle for 30 s unless
the request is a long poll, so a launch whose agent exited at startup lost its
reply and the caller saw 'runtime closed the connection' instead of
not-delivered. Classify a prompted agent.launch and agent.launchReplay as a
long poll, as orchestration.workerStart already is for the same wait.

* refactor(agent-launch): narrow the launch params by 'in' instead of a cast

* fix(agent-launch): find a launched agent behind a wrapper that leads its process group

A tcsh or nu launch line runs the agent from /bin/sh '<script>', and a
wrapper script that does not exec its agent does the same: the wrapper leads
the terminal's foreground process group and the agent is a member of it. The
fresh foreground read names the group's leader, sh, so a prompted launch was
refused or pasted late (M4Air tcsh: 2 of 4 not delivered, 2 pasted ~9 s late).

Before that read, take the host's process-group observation as positive proof
when it names the launched agent among the foreground group's members and is
younger than a ready signal's quiet window. It never proves a shell.

* fix(agent-launch): judge the foreground-group proof by when its capture began, not how long ps took

The age the host stamps on a process-group observation runs from the start
of its whole-machine ps, so on a loaded Mac a capture begun after the read
was asked for still read as older than 1 s and the proof was dropped. Count
an observation whose capture began after the read was asked for, less the
window a shared capture is reused across.

* test(agent-launch): keep the crash-guard live test out of the Windows lane's gate scan

The Windows-lane registration scan read the const assigned from a platform
check as a Windows-only gate, though the suite runs everywhere but Windows;
find zsh in a function instead, as the real-zsh typed-line test does.

Under load the fresh foreground scan can fail to answer, which lets the
shell's prompt settle readiness (2 of 4 paired runs). The guard still refuses
that write, so assert the refused write, the property that must always hold.

* perf(agent-launch): read a local pane's foreground from its own terminal, not the whole process table

The foreground read that gates every launch paste ran the daemon's
inspectProcess capture and then a fresh scan, each a whole-machine ps; the
fresh one also waits for any capture already running before it starts its own.
Measured here at load 5: 1.2 s a read (M4Air QA: 3.4-5.0 s, and worker starts
17.6-32 s against main's 9-12 s at load 25-84).

On a local macOS or Linux host, take the pane's root pid from the provider's
session inventory and run one ps limited to that pane's terminal. Its
foreground process group decides: the launched agent or any non-shell member
is the agent (a wrapper that did not exec its agent leads the group), a group
of shells alone is the shell. Same pane, same verdict: 2.7 ms a read. SSH hosts
keep the relay's observation and name.

* test(mobile): re-measure the web app's script sweep after agent.launch's capabilities moved out of protocol-version

The mobile web bundle check failed at 124 assets against a ceiling of 123. Main already sat
exactly on that ceiling: its sweep table read 69 scripts at 16 routes while the tree builds 73,
the whole margin of 4. This branch imports the agent.launch capabilities from their own module,
so protocol-version is no longer pulled into the root layout and four other routes. That moves
which routes share which modules, and the Qoder capability module, imported by protocol-version
and the AI-vault resume path, no longer shares an importer set with anything, so it gets a chunk
of its own: 74 scripts.

The fence says to re-derive the bound rather than raise it, so the sweep is re-measured on this
head (every prefix of the sorted route list). The worst route now adds 10 scripts (session), not
9, which moves the pinned shell crossing from 32 to 30 routes; main re-measured on its own lands
on the same crossing.

* fix(agent-launch): a worker's brief needs its agent found in front, and Grok's start answers on its composer

A paired-server worker start whose agent exited at startup typed its brief into the server's
shell, which ran it: the idle wait can settle on a shell back at its prompt, and the brief was
written with no foreground read. Both worker-start paths now check before each brief write, as a
launch prompt is checked: on a host that can find the agent in front it must be there; on one that
cannot (Windows) a shell proven in front still refuses, and anything else writes as before.

A Grok worker start waited ~10 s more than main: its only rest signal is its bare name, which a
launch holds to quiet output, and Grok animates its logo for ten seconds after its composer glyph.
A worker start for an agent whose rest signal is its bare name and whose composer draws a marker
(Grok, DSH, mimo-code) now also answers on that marker, whichever comes first.
2026-10-05 11:26:52 -07:00
OrcaWinandOrca Worker a68ee67cf9 fix(codex): stop prompting a Codex restart when only its home changed (#25421)
* fix(codex): stop prompting a Codex restart when only its home changed

A Codex terminal that outlived the move to ~/.codex was blocked behind "This
Codex session is using an outdated configuration" until the user restarted it,
which starts a fresh codex and ends the conversation. That terminal keeps
working on Orca's old Codex home, which is refreshed from ~/.codex, so the
change did not need the user's answer; the prompt also never appeared for
Codex in Git Bash terminals, whose process tree Orca cannot see through.

The restart prompt now covers only an account switch, which the user caused.
The pane record keeps its home route, which the retained-home refresh and the
shared-server check still read; only the prompt and the custom-home downgrade
that existed to keep it from guessing are gone.

* refactor(codex): drop what the home-route prompt left behind

- Record a pane's own CODEX_HOME override as-is. The filter that kept only
  overrides Orca could re-derive existed for the removed route comparison;
  its one remaining reader, the shared-server check, returns null for the
  shared-home route those panes record either way.
- Treat custom-home like shared-home in resolveCodexPaneHome: the removed
  downgrade only wrote it without an override, so it never named a home.
- Inline the restart notice key; its route/account prefix was its only job.
- Delete the two pane-local override tests, including a POSIX-only one that
  still expected custom-home and would have failed on Linux and macOS CI.
- Cover the account recheck branches the deleted route-recheck file was the
  last to exercise: main reporting a pane current, and main not answering.

* refactor(codex): retire custom-home and the last re-check helpers

- Read an older build's custom-home record as shared-home, which it always
  was, and drop custom-home from the route type and every check.
- Delete shellStartupCodexHomeOverrideMatches and its comparison helper;
  nothing re-checks a recorded override any more.
- Build the pane launch record in one return: the route is always set, and
  only a resumed launch differs, in how it picks the account.
- Drop the restart dialog's notice key; its focus effect now depends on the
  pane and both account labels directly.
- Give the account recheck test a typed window stub, so CI's type-assertion
  gate passes, and remove timing entries for deleted test files.

* fix(codex): drop the removed dialog strings main added to es.json

* refactor(codex): stop recording a pane's custom CODEX_HOME

The pane record kept a pane's CODEX_HOME override so the removed route
prompt could re-check it later. Its one other reader, the shared-server
check, only looked at it for a real-home pane, and a custom CODEX_HOME
always routes a pane to Orca's mirror, so it was never used.

Drop both record fields, their validators, equality checks and spawn
plumbing. getCustomCodexHomeOverrideForLaunch folds into the existing
hasCustomCodexHomeOverrideForLaunch, and real-home resolves to ~/.codex.
Records from older builds still parse; the parser keeps only known keys.

The fish test's decoy no longer sets CODEX_HOME, so the boolean check
still fails if the launch env XDG_CONFIG_HOME is ignored.

* test(codex): drop setup the yes/no CODEX_HOME checks no longer need

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
2026-10-05 11:06:40 -07:00
Neil 67fc708b3b Group CSV editor modules and tests in their own folder (#25476) 2026-10-05 01:16:20 -07:00
Neil 1bec53ceb2 Reduce repeated CI setup and overlap mobile typechecks (#25359)
* Measure remaining CI import, diagnostic and checkout savings

* Qualify remaining CI candidates on hosted runners

* Qualify independent mobile typecheck overlap on Actions

* Keep explicit RPC test registries from loading unused methods

* Qualify complete RPC registry cohort and mobile cancellation

* Promote measured CI setup and typecheck savings

* Recognize the shared RPC test guard in lint policy

* Align the mobile barrier contract with independent typechecks
2026-10-04 19:58:35 -07:00
822fc5bed4 Add repository OpenCode permission defaults (#25326)
* test(config): reproduce rejected repository OpenCode config

* Add repository OpenCode permissions and allow its reviewed root config

* fix: preserve sensitive OpenCode confirmation prompts

---------

Co-authored-by: Orca campaign recovery <campaign-recovery@example.invalid>
Co-authored-by: Orca OpenCode Campaign <opencode-campaign@local.invalid>
2026-10-04 19:34:47 -07:00
keiandsetodeve cecb62158a fix(ui): restore IME Enter protection in workspace details (#24099)
Restore IME Enter protection in workspace details by reusing the existing composition tracker. Reset Notes ownership at textarea detachment and preserve sizing behavior. Repair isolated native test-window delivery without changing the production foreground policy or original native input assertions.

Fixes #24097

Related contributor history: #10711, #11067, #13128, #13282.
Original implementation and macOS recordings: @setodeve, commit b30f095.
Verified on required stock Linux X11/Wayland checks and independent frozen-source review.

Co-authored-by: setodeve <keinick11@outlook.com>
2026-10-04 16:21:04 -07:00
Jinwoo Hong 1978469fd2 fix(mobile): keep the working rings turning on the OTA page (#25299)
* fix(mobile): keep the working rings turning on the OTA page

Animated.loop starts a native loop whenever the timing asks for the native
driver; the web has none, so the JS fallback ran one turn and froze at 360deg.
Ask for the native driver only off the web.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the native driver on native spinners

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): share the working ring rotation between both rings

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-10-04 19:12:00 -04:00
Neil 0c761a7610 Admit short required auxiliary checks after PR preflight (#25317) 2026-10-04 16:01:22 -07:00
Jinwoo Hong baa56fd10d Simplify the phone-control and phone-size terminal dialogs (#25307)
* Redesign the phone-control and phone-size terminal dialogs

Drop the eyebrow label and circled icon, shorten the copy so it no longer
restates the buttons, and give each state one primary action with a quieter
"all" action. Collapse moves out of the button row into a Minimize icon in
the corner. Behavior is unchanged.

* Point the phone settings copy at the renamed Restore button; drop dead ko overrides

The phone app and desktop update independently, so name only "Restore",
which matches both the old and new desktop banner labels.
2026-10-04 18:59:50 -04:00
Neil 8e8efb1947 Reduce avoidable work in PR checks and SSH test setup (#25309) 2026-10-04 15:26:53 -07:00
Neil b32462f246 Replace patched JSON parser with stream-json (#25202)
* Replace patched JSON parser with stream-json

* Isolate dependencies for historical server compatibility builds
2026-10-04 13:05:30 -07:00
Neil d77c57022e Verify shared preflight selection and record full unit timings (#25239)
* Strengthen shared preflight contracts and record unit timing results

* Record rejected shard-weight holdouts
2026-10-04 06:24:30 -07:00
d9173ffbdb Keep Orca CLI first after shell startup (#25130)
* Restore the owning Orca CLI path after shell profiles

* Use a literal marker for the Bash lookup regression

* Preserve plain panes and initialize zsh after prompt hook replacement

* Preserve user line-editor dispatchers during deferred startup

* fix: retain CLI startup when global Zsh replaces prompt hooks

* test: replay global Zsh hook replacement after host startup

* test: isolate controlled Zsh widgets from distro keyboard setup

* fix(shell): preserve user hooks during deferred zsh initialization

* Keep completed Zsh startup hooks retired when the wrapper is sourced again

---------

Co-authored-by: Codex <codex@openai.com>
Co-authored-by: Orca maintenance <orca-maintenance@users.noreply.github.com>
Co-authored-by: Orca campaign <orca-campaign@local.invalid>
2026-10-04 02:55:59 -07:00
Neil c4e8735f45 Share PR preflight setup to reduce runner demand (#25150)
* Share PR static analysis and compiler runner

* Preserve evidence document final newline for concurrent merges

* Keep readiness reuse contracts aligned with the physical preflight gate
2026-10-04 02:05:05 -07:00