Commit Graph
75 Commits
Author SHA1 Message Date
Brennan Benson 444e0b1cf9 fix(codex): recognise Codex's quoted spellings in config.toml, and repair Orca's duplicates (#22592) (#23958)
* fix(codex): recognise Codex's quoted project-trust spellings in config.toml (#22592)

Codex's settings screen writes project trust as ["projects"."/p"] and
"trust_level" = "trusted". Orca's matchers only knew the bare spelling, so a
trust write appended a second [projects."/p"] table (or a second trust_level
line) and every codex command then failed with "duplicate key". The config
mirror kept both spellings in Orca-managed homes for the same reason.

- Project table headers are now read through the existing TOML key-path
  parser, so bare, quoted, literal-quoted, mixed and spaced spellings are the
  same table for trust writes and the managed-home mirror/dedupe.
- trust_level is found by decoded key, in both the trust writer and the
  mirror's trust reader, and an existing key is rewritten, never duplicated.
- On the next trust write, a table older Orca appended (exactly
  [projects."<p>"] holding only trust_level = "trusted") that duplicates the
  user's table, or the bare line it inserted under a quoted "trust_level", is
  removed; the user's table wins and the atomic writer keeps config.toml.bak.
  Any other duplicate, or a repair that would still leave one, leaves the file
  untouched and logs once.

* build(cli): list the new Codex trust modules in the CLI project

* fix(codex): recognise Codex's quoted hooks.state spellings and repair Orca's copies (#22592)

Codex writes hook trust as ["hooks"."state"."<key>"] (and the parent as
["hooks"."state"]). Orca's hook-trust writer, parent-table check and mirror
only knew the bare spelling, so a hook-trust write appended a bare copy and
the file failed to parse with "Cannot declare ... twice".

- The hooks.state header, parent-table and mirror checks now use the TOML
  key-path parser, like project tables.
- The duplicate repair now also removes Orca's own hooks.state tables (an
  exact [hooks.state."<k>"] with only enabled + trusted_hash, or an empty
  [hooks.state]) that repeat a table in another spelling, and runs on hook
  trust writes too, so a file with both project and hook duplicates is fully
  repaired. The Orca-shaped copy is removed whichever order the two tables are
  in, only when exactly one other table (the user's) remains; anything else is
  left untouched and logged once.

* fix(codex): carry plain-Codex plugin and project hook trust into Orca's Codex homes (#22592)

Codex keeps hook trust in $CODEX_HOME/config.toml under hooks.state, keyed
by the hook's source. Plugin keys (`id@mkt:path`) and project keys
(`<repo>/.codex/...`) are the same in every home, but the mirror dropped
every hooks.state table from ~/.codex, so Codex inside Orca asked users to
re-trust plugin and project hooks they had already trusted in plain Codex.

- classifyHookTrustKey splits keys into home-scoped (the home's own
  hooks.json/config.toml, re-keyed by install as before) and shared.
- The mirror now carries shared hook trust from ~/.codex in every spelling.
  A key the managed home already holds keeps the managed copy, a key
  repeated in ~/.codex is carried once, and the parent [hooks.state] table
  is never copied, so the result never declares a table twice.
- mergeSystemCodexConfigIntoRuntime moves to codex-config-mirror-merge.ts
  to keep codex-config-mirror.ts under the line limit.
- Tests cover plugin/project carry in each spelling, user-hook keys staying
  out, repeated launches, managed-copy precedence, Windows key spellings,
  parent tables, and user-hook trust re-keying (trusted_hash and enabled)
  from every ~/.codex spelling.

* fix(codex): carry session_end and interrupt hook trust into Orca's Codex homes (#22592)

The shared-trust classifier parsed hook keys with Orca's own trust-key
parser, which only knows the ten events Orca installs hooks for. Keys for
Codex's session_end and interrupt events did not parse, so their plugin and
project trust was treated as home-scoped and left out of the managed home.

The classifier now reads the source path from Codex's key shape
`{source}:{event}:{group}:{handler}` for any event label. A key without
that shape is still never carried. Tests cover both events for plugin and
project keys in both spellings, user-layer keys for both events, and five
unattributable key shapes.
2026-09-30 00:09:26 -07:00
Neil ccdb324b63 Add CodeBuddy as a built-in coding agent (#23740)
* feat(agents): integrate CodeBuddy launch, status and session history

* docs: record CodeBuddy lifecycle verification

* fix(codebuddy): backfill scoped history and negotiate remote resume

* test(cli): include CodeBuddy in known search agents
2026-09-28 18:11:25 -07:00
Brennan Benson a84bd1c4fd fix(claude): write only the hook events and statusLine the user's Claude accepts (#23614)
* refactor(claude): name the Claude version module after the hook events it gates

Pure move of claude-session-end-hook-capability.ts and its tests; the next
commit turns its one-event SessionEnd floor into a per-event version table.

* fix(claude): write only the hook events the resolved Claude knows

Claude 1.0.81 through 2.1.100 validate settings.json `hooks` against a
closed event enum and discard the whole file on one unknown name, so
Orca's install made Claude <= 2.1.77 silently ignore the user's env,
permissions and hooks. Each managed event now carries the first Claude
release that knows it (pinned to per-release enums read from the npm
packages), and install, status and the SSH/WSL relay installer write
only the events the resolved Claude accepts. An unresolved version gets
the set every tabled Claude knows; a downgrade removes only Orca's own
entry for an event the older Claude would reject.

* refactor(claude): move the managed Claude hook events into their own module

hook-settings.ts is at its line limit; the event list and its version gate
move out whole so the next change has room.

* fix(claude): an unresolved Claude version never removes Orca's hook entries

A failed or timed-out version probe is no evidence of an old Claude, so it
must not strip StopFailure, PermissionRequest and the other newer events a
version-aware install wrote. With the version unknown, install adds only
the set every tabled Claude knows and leaves every other entry exactly as
it is; only a known version that lacks an event retires Orca's entry.

* fix(claude): gate the core hook events on the Claude release that added them

Claude validates hooks against a closed event list from 1.0.23, not 1.0.81.
The table treated SessionStart, UserPromptSubmit, Stop, SubagentStop,
PreToolUse and PostToolUse as known by every resolved version, so a Claude
from 1.0.23 to 1.0.61 was still sent names it rejects, and it dropped the
whole settings file. Pin each to its first release from the packed enums and
keep the unresolved-version set as its own policy.

* fix(claude): write Orca's statusLine only for a Claude that knows it

Claude 1.0.49 through 1.0.66 also reject any unknown top-level settings
key, and statusLine joined that schema only in 1.0.64. Orca wrote its
statusLine for every Claude, so 1.0.49 to 1.0.63 still dropped the whole
settings file even with the event gate. Gate statusLine on 1.0.64, pinned
by the packed schemas; a known older Claude has Orca's own statusLine
removed along with the opt-out marker, so an upgrade re-adds it. An
unresolved version is now assumed to be 1.0.64, which knows the same core
events and keeps the statusLine install it had before.

* test(claude): a user statusLine opt-out survives a downgrade and upgrade

Retiring Orca's statusLine for a Claude older than 1.0.64 forgets the
install marker only when Orca's own statusLine was removed. Pin that, so
a user who deleted Orca's statusLine is not opted back in by an upgrade.

* test(claude): check the whole written settings file against each strict schema

Claude 1.0.49 through 1.0.66 discard the whole settings file over any
top-level key their schema lacks. The fixture recorded only whether each
release knew statusLine, so a new top-level key Orca wrote would pass every
test. Record each release's top-level keys instead (statusLine is derived
from them), add the hook enums for every packed release in that window,
and check that a real install and a downgrade write only keys and events
each strict release accepts.
2026-09-28 12:04:59 -07:00
8b410b4893 feat: add first-class Qoder CLI support (#23581)
feat: add first-class Qoder CLI support

Integrate Qoder launch, identity, canonical hook status, trust and resume.
Verify with captured Qoder 1.1.64 transcripts and hidden Electron sidebar checks.

Builds on and cross-reviews #7502, #8611, #9655, #12910, #13311 and #15291.

Co-authored-by: dalveytech-vincent <vincent@dalveytech.com>
Co-authored-by: Eridanus117 <45489268+Eridanus117@users.noreply.github.com>
Co-authored-by: xingqingzzp-gif <xingqingzzp-gif@users.noreply.github.com>
Co-authored-by: jyang2004 <jyang2004@users.noreply.github.com>
Co-authored-by: yunqian <yunqian@alibaba-inc.com>
Co-authored-by: huzhening.hzn <huzhening.hzn@alibaba-inc.com>
2026-09-28 02:59:50 -07:00
Neil 45f3512a33 feat(agents): add first-class DeepSeek Harness (dsh) support (#22468)
* feat(agents): add first-class DeepSeek Harness (dsh) support

Register DSH as a supervised Orca agent: catalog entry and detection for its
dsh-tui profile, status/question hooks through DeepSeek's own Claude-Code hook
bridge, composer-ready prompt delivery, session resume, headless Source Control
AI, and title identity that no longer collides with Gemini's.

* fix(dsh): reach Orca through DSH's credential scrub and stop reading its title as Gemini

DSH runs command hooks through its own shell executor, which drops every env var whose
name contains KEY, TOKEN, SECRET or PASSWORD — taking ORCA_PANE_KEY and
ORCA_AGENT_LAUNCH_TOKEN with it, so every hook exited without posting. Mirror both onto
scrub-safe aliases at spawn and restore them at the top of the DSH hook script.

Its title collided too: DSH rests on the same glyph Gemini works on, so a resting DSH
pane was relabelled Gemini CLI and reported working forever. Defer both the Gemini
classifier and the title status detector on DSH's whale, in the base module both copies
of that classifier read.

* test(mobile): repin the session-route closure for the DSH agent icon

* fix(dsh): address review — never splice user rows, cover remote panes, keep the diff off argv

- findManagedDshPatchRegion paired an orphan start marker with a later block's end, so a
  truncated write made install/remove delete the user's own rows. Pair each end with the
  nearest preceding start; regression test fails without the fix.
- The relay PTY env builder never applied the scrub-safe aliases, so remote DSH status
  silently never appeared even with the remote hook installed.
- Source Control AI sent the whole diff on argv; send it over stdin with DSH's '-' marker.
- dsh-tui/dst already chose the interactive profile, so a workspace folder named 'web' or
  'plugin' no longer marks a live agent pane non-interactive.
- Isolate USERPROFILE as well as HOME so a Windows run cannot edit the real home.
- Drop the duplicate README badge and revert an incidental doc reformat.

* refactor(dsh): share the managed-hooks reader and tighten the new modules

Reuse before reimplementing: readManagedDshHookEvents was a near-verbatim copy of Muse's,
with byte-identical private helpers. Both now call one readManagedHookEventsFromJson.

Also: one readTextOrAbsent instead of two spellings of the same read (dropping an
existsSync TOCTOU), one status() builder instead of four inline literals, rmSync(force)
instead of exists-then-unlink, and a redundant empty-string guard before JSON.parse.
The patch-file transforms lose their index juggling for a predicate plus a filter.

* fix(dsh): refuse a flow-style patch file, keep its mode, and stop the relay inheriting a pane

- applyManagedDshPatch matched only an exact `[]`, so `[] # keep empty` or a non-empty
  flow sequence got a block entry appended after it — invalid YAML that would leave DSH
  unable to load the user's own patch layer either. It now strips the token from an empty
  sequence (keeping a trailing comment) and returns null for a non-empty one; install
  reports that and changes nothing.
- The patch rewrite dropped an owner-only file to the umask default (CWE-732); pass
  preserveMode.
- The relay PTY env never dropped inherited pane identity the way the local and daemon
  builders do, so a spawn that specified none could inherit the relay's own and every
  agent's hook would report against that pane.

* fix(dsh): keep the flow-style refusal in every status read, and scope the mode test to POSIX

A refused patch file carries no managed region, so getStatus() fell through to a bare
not_installed with detail null — the actionable 'rewrite it as a block sequence' message
only ever reached the one-shot install() return. Export the predicate and check it first,
behind one shared message constant.

The owner-only mode assertion cannot hold on Windows, where chmod only toggles the
read-only attribute and mode & 0o777 reads 0o666 for any writable file.

* docs(readme): restore the DeepSeek Harness badge lost in the rebase

* test(mobile): repin the session-route closure to the measured 4221

Measured, not derived: 4220 without the DSH icon entry, 4221 with it. Two of the three
modules above main's 4218 pin are not this change's — they arrived with the mobile work
after #22570 and were never repinned; the changelog records that split explicitly.

* fix(dsh): settle tui-idle on the agent's own hook, so supervised workers see it ready

Reported by a tester on the adhoc build: `terminal wait --for tui-idle` ran to its 90s
timeout against an already-ready DSH composer, so a supervised worker never sees the agent
as ready.

Every existing tier reads the title, and DSH deliberately carries no title status: its rest
prefix is Gemini's working glyph, so the detector reports none. A fresh first-party `done`
is better evidence than any title anyway — it is the agent's own account of its own turn,
and normalizeDshEvent drops subagent events, so it is the lead's. Scoped to DSH: for agents
whose hooks report child turns, a mid-turn `done` is the #6011 class this file prevents.

* test(daemon): record the DSH transcript's true-colour I2 divergences

Adding the dsh-tui capture to __fixtures__ enrolled it in the serialize replay sweep, where
it reports 10 I2 divergences and failed the unlisted-transcript default of 0.

Every one is the same shape — visible-grid row=0, a 24-bit background the round trip does
not restore to default — which is DSH's whale intro painting whole rows of true colour.
Verified as an upstream limitation rather than a regression by replaying against the
previous build (build-serialize-addon-at-ref.mjs --ref origin/main): I1 and I3 both hold.

* fix(dsh): return the new tui-idle verdict from the first-party done lane

Main refactored isTuiIdleSatisfied into evaluateTuiIdle, which returns a verdict rather
than a boolean. The DSH lane still returned `true`; it is tier-1 positive evidence, so it
returns READY_STRONG like the title/body lane above it. Re-verified the regression test
still fails without the lane.

* test(relay): pin the scrub-safe pane-identity aliases on the relay spawn path

The relay builds a remote pane's env itself, so the alias mirroring there had no
test: removing the call left every suite green while remote DSH status silently
vanished. Both cases fail without it.

* docs(dsh): point the hook service at the integration reference

The reference doc had no inbound link from anywhere in the repo.
2026-09-27 22:44:18 -07:00
459410a63b Preserve Hermes YAML configuration during hook installation (#23324)
Preserve supported Hermes YAML values and comments while installing or removing Orca hooks.

Adapted from manthis and Pr1p proposals #22366 and #20632.

Co-authored-by: Maxime AUBURTIN <m@hellomax.io>
Co-authored-by: Chen <zwq19980411@gmail.com>
2026-09-26 22:13:24 -07:00
OrcaWinandm4air 8416e8de10 refactor(persistence): retire ordinary JSON profile writes (#23202)
* refactor(persistence): retire ordinary JSON profile writes

Require SQLite for writable profiles and keep import, compatibility export, and recovery in a documented legacy-json boundary.

* fix(cli): preserve dynamic profile imports in release output

* test(persistence): exercise SQL races and verify packaged CLI imports

* test(persistence): consolidate shared fixture imports

* test(persistence): close SQLite fixtures before cleanup and await launcher output

* test(automations): use SQLite fixtures for dispatch fencing and skip coalescing

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 12:56:32 -07:00
OrcaWinandm4air 3a081abf71 fix(persistence): reclaim Windows profile locks after PID reuse (#23122)
* fix(persistence): identify reused Windows profile-owner processes

* fix(persistence): preserve absent-owner recovery without native registry

* ci: build Windows registry before native profile identity checks

* fix(cli): include native profile-owner dependencies in typecheck

* test: register native profile owner test in Windows PR lane

* test: use resilient Windows profile-owner cleanup

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 02:26:10 -07:00
OrcaWin 6fc3cdcad6 Bundle Bun for headless Orca and profile persistence (#22635)
Bundle a pinned, verified Bun runtime for headless Orca so existing Node launch commands can hand off before opening a profile. Keep desktop execution on Electron.

Add the Bun SQLite adapter and terminal backend, bounded shutdown, process inspection and cross-platform artifact qualification. Keep future managed SSH deployment separate from current production launch paths.
2026-09-25 22:49:06 -07:00
OrcaWin 82412dab8b Persist profile state in SQLite with background writes (#22612)
Migrate profile state to SQLite and move writes and backups into a background worker. Acknowledge terminal, SSH and automation changes only after durable saves. Preserve JSON import, recovery, rollback and compatibility exports.

Validate migration, worker failures, maintenance, cross-profile moves and terminal lifetime races with unit, integration and end-to-end coverage.
2026-09-25 22:47:33 -07:00
Neilandmmarabel 85ac14e9c2 fix(codex): retain runtime MCP entries without losing revocation (#22426)
* Retain runtime-only MCP entries

Adapted from the investigation and proposal by @mmarabel.

Co-authored-by: mmarabel <mmarabel@users.noreply.github.com>

* fix(codex): respect inline and dotted canonical MCP ownership

* fix(codex): retain canonical MCP removal across upgrades

* fix(types): include MCP ownership in CLI project

* Keep unrelated main test formatting unchanged

---------

Co-authored-by: mmarabel <mmarabel@users.noreply.github.com>
2026-09-25 21:37:54 -07:00
8846987c99 feat(rate-limits): add Cursor usage tracking (#22633)
* feat(rate-limits): add Cursor usage tracking

## ELI5

If you use Cursor, Orca now shows how much of your monthly Cursor plan you
have used, next to the Claude, Codex and Grok meters, and in Settings →
Accounts. It reads the sign-in Cursor already saved on this computer and never
changes it.

## What changed

Cursor becomes a rate-limit provider like Grok: a status-bar meter (default-on,
with its own toggle), a row in the usage roster, and a Settings → Accounts
section naming the signed-in account.

The credential is read from whichever of three stores has it, first match wins,
all read-only:

- the macOS login keychain item `cursor-access-token` / `cursor-user`, which is
  where `cursor-agent` 2026.06+ keeps the session;
- `~/.cursor/auth.json` and its platform variants, used by older CLIs;
- the Cursor IDE's `state.vscdb` (`cursorAuth/accessToken`), for people who
  never run the CLI.

The keychain entry is the one current CLIs use, and reading only `auth.json`
finds nothing on an up-to-date macOS install. A locked keychain cannot mask a
readable `auth.json`, and a locked `state.vscdb` cannot mask either.
`~/.cursor/cli-config.json` supplies the account's email and display name; it
never holds a token.

Usage comes from the dashboard route the Cursor web dashboard itself reads,
because Cursor documents no individual-user usage API — every documented API is
team- or Enterprise-scoped. Per Cursor's pricing docs an individual plan has two
pools, Cursor Models and Other Models, both resetting with the billing cycle,
plus optional on-demand spend; each becomes a named bucket. The headline
percentage prefers `used / limit` over the sibling percentage fields, which are
pre-rounded for the dashboard's own copy. Because the route is undocumented the
mapping is defensive: an unrecognised payload resolves to `unavailable` and
hides the bar rather than publishing a zero that reads as "no usage".

Orca never runs `cursor-agent login` and never writes, refreshes or rotates a
Cursor credential. An expired token short-circuits to an actionable
"run cursor-agent login" instead of spending a request that can only 401 — not a
rare case, since `cursor-agent status` still reports `isAuthenticated: true`
against a token that expired months ago.

## Why this shape

Six open PRs implement this feature and none reads the keychain, so each finds
nothing for a large share of users; this takes the auth layer further and keeps
what those PRs verified live. The bar is not gated on `cursor-agent` being on
PATH, unlike other CLI providers, because an IDE-only session is real usage with
no CLI to detect.

`readKeychainPassword` moved out of the Claude keychain reader into
`src/main/macos-keychain/generic-password.ts` so both providers share one
`security(1)` wrapper. It is a byte-for-byte relocation, so Claude's credential
path is unchanged; the two child_process allowlists move the entry with it and
neither ratchet count changes.

Co-authored-by: Preschian Febryantara <preschian@users.noreply.github.com>
Co-authored-by: Qwesdy <qwezdi@proton.me>
Co-authored-by: ivo922 <github.concur614@passmail.net>
Co-authored-by: Mihail Vratchanski <mivrkiki@gmail.com>
Co-authored-by: Tauri-EPO <enrico.pin@gmail.com>
Co-authored-by: Raajik <44516546+Raajik@users.noreply.github.com>

* test(rate-limits): name the JWT helper's segment type in the Cursor tests

The anti-slop gate rejects a bare `object` parameter; the fixtures build a
claims record, so say that.

* fix(rate-limits): render Cursor's pools and keep its plan total visible

Review of the first commit found the meter effectively blank for a healthy
account, which the screenshots missed because the only Cursor session on hand
had expired and never reached the success path.

- The verbose status-bar segment filtered buckets through an allowlist written
  for Gemini's experimental models, so both Cursor pools were dropped and the
  fallback needed a `session` window Cursor never reports. A signed-in account
  rendered an icon and no number. The allowlist now admits Cursor's pools, and
  the fallback accepts a monthly window.
- `getWindowSections` dropped `monthly` whenever buckets existed. Cursor puts
  the plan total there and its sub-pools in buckets, so a plan at 92% showed as
  50% in the roster, the tooltip, and the tightest-usage pick.
- A plan reporting `enabled: false` still published its 0% pools, painting a
  healthy meter for a pool the account does not own and skipping the
  request-quota fallback.
- `redirect: 'error'` turned the dashboard's bounce to /login into a generic
  network failure, hiding the actionable sign-in message.
- A busy `state.vscdb` (the IDE holds it open) surfaced as a provider error,
  which would pin an alert bar on Cursor IDE users who never set Cursor up in
  Orca. It falls through to "no credential" instead.
- Refreshing the Accounts section read the keychain twice for one update.

* fix(rate-limits): pin the platform in the Cursor keychain tests

Review caught three cases that assumed macOS: the keychain source is behind an
explicit `process.platform` check, so on the Linux CI runner the mocked read was
never reached and the tests read the CLI file instead. They now set the platform
they mean, and two new cases assert the off-macOS fall-through.

Also track the credentials reference doc (docs/** is ignored by default, so a
new reference needs its own allowlist entry) and give the visibility fixtures
their own provider id instead of Grok's.

* fix(rate-limits): prefer a live Cursor session and report a failed refresh

Review round two, from CodeRabbit and Pullfrog.

- Credential precedence returned the first token that parsed, so an expired
  keychain token in front of a fresh Cursor IDE session reported "sign-in
  expired" on every poll while a usable session sat one source below. A live
  session now wins; the expired one is returned only when nothing live exists,
  so the actionable message still appears in that case.
- The usage schema took `.optional()` where the route sends `null` for an absent
  sub-object, so one null pool failed the parse for the whole body and threw
  away valid pools and the billing cycle with it.
- Cursor usage could survive an account switch: a failed refresh for account B
  kept account A's figures beside B's name in Accounts. The snapshot now carries
  a hashed account fingerprint, and a known-and-changed identity clears the
  previous reading. A refresh that names no account still keeps its own.
- The Accounts section rendered nothing at all when a signed-in account's fetch
  failed, and could repaint an older account when two status reads overlapped.
  It now states the failure — beside the numbers when a stale snapshot remains —
  and ignores superseded reads.
- A web client claimed "not signed in" for a host it cannot read, contradicting
  the meter beside it; it now says the detail is host-only.
- Signed-out copy named `cursor-agent login` as the only way in, though an IDE
  sign-in works just as well.
- The census comment ended at 4219 after the pacer squash without naming the two
  modules #22616 added; recorded them, re-measured on a clean origin/main.
- Narrowed the docs claim: Cursor documents all-plan APIs, but no individual
  usage endpoint.

* fix(i18n): localize the web client's Cursor host-only notice

It reaches the Accounts pane like any other string, so the coverage gate is
right to want it in the catalog rather than allowlisted.

* fix(rate-limits): name the Cursor account on failed refreshes, and ship the reworded copy

Review round three. Both findings say an earlier fix did not actually take.

- The account-switch guard reads `authProvenance` off the fresh result, but the
  fetcher stamped it only on success and network failures. The `stale-token`,
  429, 5xx and parse results omitted it, and so did the expired-session branch —
  so a switch whose first refresh failed, which is precisely the case the guard
  exists for, still rendered the previous account's figures under the new name.
  Every failure holding a readable session now names its account; a missing or
  unreadable credential still names none. The service test also fed a result
  shape the fetcher never produces, so it proved nothing; it now uses the real
  stale-token shape, and the fetcher test asserts provenance across 401/429/5xx
  and expiry.
- The reworded signed-out copy never rendered: a present catalog value beats the
  `translate()` fallback, and `sync:localization-catalog` only adds missing keys
  rather than updating changed defaults. Updated both strings in en.json, which
  also prunes them from the runtime-required catalog now that they match.

* docs: keep the Cursor credentials reference out of the tree

Its content lives in the PR description instead; docs/** stays ignored rather
than gaining an allowlist entry for this branch.

* test(mobile): drop the census note main no longer pins

main removed `SESSION_ROUTE_MODULES` and re-pinned this lane on a different
count, so the paragraph this branch added documents a number series that is
gone. The branch touches nothing in this file now.

---------

Co-authored-by: Preschian Febryantara <preschian@users.noreply.github.com>
Co-authored-by: Qwesdy <qwezdi@proton.me>
Co-authored-by: ivo922 <github.concur614@passmail.net>
Co-authored-by: Mihail Vratchanski <mivrkiki@gmail.com>
Co-authored-by: Tauri-EPO <enrico.pin@gmail.com>
Co-authored-by: Raajik <44516546+Raajik@users.noreply.github.com>
2026-09-25 18:54:11 -07:00
c220d92c03 fix(codex): Codex 0.157+ starts in Orca-managed homes instead of failing with SUN_LEN (#22878)
* fix(codex): turn off Codex daemon auto-start in homes whose socket path exceeds sun_path

Codex >= 0.157 auto-starts a background app-server daemon and connects to
<CODEX_HOME>/app-server-control/app-server-control.sock. Orca's managed homes
under userData make that path longer than sun_path (104 bytes on macOS, 108 on
Linux/Windows), so every interactive codex in an Orca terminal failed with
'path must be shorter than SUN_LEN'. The config mirror now writes a marked
[features] daemon_auto_start = false into only those homes, removes it when the
home fits, and never promotes it into ~/.codex.

* fix(codex): address review of the daemon socket guard

- A runtime config.toml holding only Orca's daemon override no longer reads as a
  config-sync stall, so users without ~/.codex/config.toml get no false
  "missing" warning in the accounts pane.
- The legacy shared-home refresh re-applies the guard, so retained pre-rollout
  panes keep daemon auto-start off after a system-default launch.
- Warn once when an inline `features = {...}` or `[[features]]` blocks the
  override instead of failing silently.
- Rename the upsert's TUI-specific internals now that it serves any table.

* fix(codex): apply the daemon socket guard even when the settings mirror stalls

When the settings write-back or mirror refused (unreadable baseline, failed
write to ~/.codex, unreadable source), the whole pass returned before the
daemon guard was applied. A home whose config.toml predates the guard then
kept failing with SUN_LEN on every launch for as long as the stall lasted.
The guard now lands on those paths too; the mirror itself is unchanged.

* fix(codex): guard managed account homes when ~/.codex/config.toml is missing

* test(codex): keep reset-credit ownership checks scoped to the retry, not service construction

* test(codex): build the account mirror test without a type cast

* fix(codex): keep blocking WSL ownership checks off the no-config guard pass

Guarding account homes with no ~/.codex/config.toml ran the WSL ownership
check, a synchronous wsl.exe call per account, at startup before the window
opens and on every account switch. WSL homes are guarded by WSL launch prep,
so that pass now covers host homes only.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-09-25 14:29:08 -07:00
Neil 90801e2deb feat(agents): add first-class ZCode harness (#22464)
* feat(agents): add first-class ZCode harness

Add ZCode (Z.ai's `zcode` CLI) as a supervised Orca agent: managed lifecycle
hooks on local, SSH and Windows hosts; status, question and approval reporting;
synthetic status titles; session resume; orchestration worker launch options;
and desktop + mobile agent-picker registration.

Written against the newly open-sourced `zai-org/ZCode` (agent CLI 0.16.9), not
against a remembered screen:

- ZCode's hook runner writes a Claude-compatible stdin alias set, so it routes
  through the existing Claude-compatible vendor path while keeping its own
  identity in the sidebar.
- `PermissionRequest` fires only once the approval card is on screen and racing
  the user's answer, so it is proof the pane is blocked, not an auto-approval.
- ZCode's clarification tool is literally `AskUserQuestion` with Claude's
  questions/options shape, so Orca's question card renders it unchanged.
- ZCode's `hooks.enabled` defaults to false, which is why configured hooks were
  reported as never firing; the installer sets it.
- ZCode renames its own process to `zcode-cli`, so the expected foreground
  process cannot be the launch command or dispatch refuses the pane.
- ZCode emits no OSC title in any state and repaints its ASCII banner forever,
  so readiness comes from Orca's synthetic hook title and launch drafts wait on
  the composer box rather than on a quiet render window.

Three files crossed their max-lines limit, so each is split along a real seam:
command-line entrypoint parsing out of agent process recognition, skill
classification out of skill root discovery, and registry coverage out of the
remote hook installer tests.

Refs #10564

* fix(zcode): drop the session-option catalog and pin the orchestration contract

ZCode's CLI exposes no `--model` flag at all, and the session-option launch path
refuses to apply any option until a model id is chosen. A catalog therefore could
not deliver `--mode` per worker, and would have accepted `--model` only to drop
it silently. Take opencode's position instead: no catalog, so `worker-start
--model` is refused with a clear message and ZCode launches with the model from
its own config. `--mode` stays reachable through agent args, which is also how
the yolo default is applied.

Add a contract test covering the parts that make ZCode a usable worker:
dispatchable foreground process, stdin prompt delivery, the prompt staying out
of the launch command, and the composer-gated draft paste.

* refactor(zcode): reuse shared helpers and cut the harness down

No behaviour change; every ZCode test still passes.

- Use installer-utils' own `hookDefinitionHasManagedCommand` instead of
  re-walking a hook definition by hand, which also drops a local string reader.
- Share one `readZCodeEventMap` instead of keeping the same narrowing in both
  hook-settings and hook-config-json.
- Collapse five identical error returns into one `zcodeHookError` builder, and
  return early from the status branches instead of assigning through `let`.
- Split the event-to-status decision out of `normalizeZCodeEvent` into a pure
  `readZCodeTurn`, so the normalizer reads as decide-then-build and stops
  computing the tool name for events that never look at it.
- Take a script file name in `readManagedZCodeHookEvents` like its siblings,
  which removes a `Parameters<typeof …>` indirection at the call site.
- Drop the unused `ZCodeHookEvent` export and inline a single-use path helper.
- Correct a stale comment: ZCode's loader is a strict `JSON.parse`, so the
  in-place edit preserves key order and indentation, not comments.

* fix(zcode): address review — keep unmanaged event keys, correct comment, de-dupe README

- `removeZCodeManagedHooks` deleted any event key whose list ended up empty, so an
  unrelated `"Notification": []` the user wrote was removed as collateral whenever a
  managed hook elsewhere made the write happen. Only touch an event Orca actually
  owned something in; covered by a new regression test.
- The `isNewTurnEvent` comment claimed UserPromptSubmit was ZCode's only turn
  boundary while the expression below it also returned true for SessionStart. Say
  what the code does: SessionStart lands the idle boundary, UserPromptSubmit is the
  turn boundary (the Codex/Claude shape).
- ZCode appeared twice in the README's single agent-badge block; keep the
  local-icon entry the link checker validates and drop the favicon duplicate.

* docs(zcode): call out that the desktop bundle's CLI cannot open a session

From live testing on #22464: pointing `zcode` at the desktop app's bundled
`glm/zcode.cjs` installs Orca's hooks fine but then fails with
`Cannot find package '@zcode/tui'`, so the pane never opens a session. The
symptom reads as a broken harness when the CLI simply has no TUI. Say which
build to use and how to check before reporting a problem.

Reported-by: JWu527
2026-09-25 02:17:51 -07:00
NeilandAdrien De oliveira ebed0964a2 feat(agents): add first-class Muse Code harness (#22216)
* feat(agents): add first-class Muse Code harness

Add Muse as a supervised Orca agent across desktop, mobile, session history, source control, local hooks, SSH, WSL, and native Windows. Preserve user settings, support Muse 1.3 hook environment allowlists, and recognize versioned foreground processes. Include question, waiting, completion, resume, and readiness coverage.

Co-authored-by: homesh-dev <300847526+homesh-dev@users.noreply.github.com>

Co-authored-by: jeffhuen <32542276+jeffhuen@users.noreply.github.com>

Co-authored-by: John Cusack <johncusackccm@gmail.com>

Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com>

* test(agents): cover Muse remote hook registration

* test(agents): cover Muse hook and source-control contracts

* test(agents): exclude Muse hook metadata from script mode check

* test(agents): keep Muse skill picker coverage stable

* test(ai-vault): include Muse in every-agent fixture

* test(mobile): repin Muse agent icon closure

* fix(muse): detect questions and approvals from structured Muse signals

Muse 1.3 fires no hook for request_user_input, so a pending question left
the pane "working". Its internal reminder subagents also post hooks with
their own session ids (even after Stop), which surfaced "tool failed" rows
and flipped finished panes back to working.

- Read pending questions from Muse's session log
  (user_input_prompt_requested/settled) via the existing transcript poll,
  now generalized from Codex subagents to Muse on main and relay.
- Drop child-session hooks (SubagentStart ids, or turn_id === session_id).
- Treat Notification permission_prompt as the approval wait; PermissionRequest
  also fires for auto-approved calls, so it only caches the approval card.
- Ignore Notification copy as the prompt; poll replays are not new prompts
  or turn boundaries.
- Allowlist USERPROFILE so Windows cmd AutoRun doesn't fail every hook.

* perf(muse): parse only question events from the session log

Most Muse session-log lines are large model/tool records. Filter raw lines
by the user_input_prompt_ marker before JSON.parse via an optional
readJsonlCursor line filter.

* fix(muse): unwrap batched log records and scope questions to the live turn

Review follow-ups: question events inside retained_frame batches were
skipped, and a question left open by a crash or interrupt stayed pending
for the pane's life. Share the history scanner's retained_frame unwrapper,
and only report a pending question whose run_id matches the hook turn_id.

* refactor(muse): drop type assertion in retained_frame unwrap

* fix(agent-hooks): satisfy exhaustive-switch lint in transcript poll policy

---------

Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com>
2026-09-22 19:13:11 -07:00
Brennan Benson 33149fcde5 fix(claude): install SessionEnd for capable versions (#20530) 2026-09-13 21:59:57 -07:00
Brennan Benson e944e76537 fix(grok): stop replayed Claude/Cursor hooks reporting Grok panes as Claude (#20507)
* fix(grok): stop replayed Claude/Cursor hooks reporting Grok panes as Claude

Grok's hook discovery reads ~/.claude/settings.json (and the Cursor equivalent)
for vendor compatibility, and that is on by default. So inside every Grok pane
Orca's managed Claude hook fires in addition to Orca's managed Grok hook, and
both POST the same Grok envelope. The Claude-routed copy lands last and wins, so
the pane's agent type is resolved from the POST route as "claude" and no
Grok-specific normalization runs for it.

Guard the managed Claude and Cursor scripts on GROK_HOOK_EVENT, which Grok's hook
runner stamps into every hook subprocess it spawns — including replayed vendor
configs — after any user-supplied environment, so a hook cannot spoof it. This
mirrors the existing DEVIN_PROJECT_DIR guard in the same script, which solves the
identical problem for another agent that imports Claude hooks.

Placement is load-bearing: the guard sits after the stdin capture, so Grok's
writer never blocks, and before both the spool write and the HTTP POST, so a
replayed event cannot leave a spool entry that replays later. The Windows
variants jump to the stdin-drain label rather than exiting, because abandoning
stdin there hangs the writer.

The guard is scoped to agent === 'claude'; OpenClaude reuses ClaudeHookService
with its own settings file, which Grok does not replay, so it is unaffected.

Verified live against Grok 1.0.25 in a dev instance: the pane's reported agent
type goes from "claude" to "grok" on every turn-end, including the hidden
follow-up turns Grok runs when background work finishes.

The guard pushed hook-service.ts past the 300-line cap, so the script builder
moves to a sibling hook-script.ts. That mirrors the existing split under
src/main/cursor/, where the service owns install/status and the script module
owns script text.

* fix(agent-hooks): preserve Windows background worker stdin contract
2026-09-13 15:55:41 -07:00
20ab995065 fix(codex): reconcile marketplace and plugin tables through the config mirror (#20150)
Scalar promotion omitted the marketplace and plugin tables, and the mirror rebuilt
ordinary config from canonical while only trust sections survived, so a managed-account
registration and refreshed provider metadata were both destroyed at the same boundary.

Registrations now reconcile through one baseline-aware pass before the canonical->runtime
copy: a runtime-only table is promoted, a table the canonical config removed since the
last mirror stays removed, canonical wins on an identity change, marketplace refresh
metadata is promoted only for a strictly newer valid timestamp with its paired revision,
and a plugin `enabled` toggle promotes only when the runtime alone changed it.

The settings baseline gains an optional `registrations` map at version 3. Absent means
never mirrored, which makes the v2 upgrade lossless; an older build rejects version 3 and
rebuilds, so downgrade is a safe degrade.

Verified end to end against a real codex-cli binary, which wrote the registration into a
managed home and read the promoted result back: `No marketplace plugins found.` becomes
`ponytail@ponytail  installed, enabled`.

Fixes #10489
Fixes #11770

Co-authored-by: BsTiger <96857444+Bongseop-Kim@users.noreply.github.com>
Co-authored-by: Rod Boev <rod.boev@gmail.com>
2026-09-11 21:22:10 -07:00
Neilandmaoking 438e0f4f5a fix(hooks): stop orphaned managed markers from consuming user TOML (#20148)
A managed block missing its end marker was treated as Orca-owned through EOF,
so uninstall/reinstall deleted appended user tables. The same shape existed a
second time in the Codex legacy profile cleanup.

Ownership is now two separate claims: a marker pair proves extent, and a
provider that can recognize its own emitted tables owns them wherever they
sit. An orphaned marker owns only its own line. Recognition uses the same
test for remove, install and status, so a table Orca cannot see is never one
it leaves running.

Co-authored-by: maoking <secretxierluo@gmail.com>

Fixes #18861
2026-09-11 21:22:06 -07:00
a899f92402 feat(windows): enable structured Codex chat on native Windows (#18519)
* feat(native-chat): enable Windows structured sessions

* fix(codex): prove native Windows process identity

* style(codex): format Windows session seam

* fix Windows structured Codex admission

* fix(windows): reprobe missing process identity capability

* fix(windows): decide folder-workspace WSL routing before the click

Review found pathUsesWslUnc exported but unused, and the folder composer
hardcoding worktreeUsesWslPath:false. Together those meant a folder picked
under a \\wsl.localhost\ parent routed to structured chat, then got refused
by the host and fell back AFTER the click -- which defeats the lane's own
design goal that create cannot fail after the click.

The group's parentPath is in scope at submit and the workspace is created
under it, so the parent decides WSL-ness pre-click. Wires pathUsesWslUnc
there and adds tests for the helper, including the unhydrated-store case
that previously threw.

* fix(windows): collapse the gate derivation to one call, restoring max-lines

CI static analysis failed: launch-agent-in-new-tab.ts crossed the 300-line
oxlint ceiling. Adding a max-lines disable is forbidden, so the two gate
derivations collapse into one readWindowsStructuredGateInputs() call --
a store-backed site now adds one line and one import name instead of two.
Better shape anyway: one derivation entry point rather than two reads a
call site must remember to pair.

* fix(windows): engage the legacy fallback when the host THROWS a refusal

Review found a P1 this merge composes: neither parent could reach it. At the
lane head the only structured entry was launch-agent-in-new-tab (full
store-backed WSL check); on main all win32 was refused. The merge enables
win32 in creation flows that pass no projectRuntime, so a WSL folder
workspace, a WSL-configured repo, or a repair-required runtime now routes
structured -- and the host refuses correctly, but by THROWING rather than
returning {ok:false, refusal}.

Callers engage their legacy-terminal fallback on the refusal CLASS, so an
unmapped throw arrives as a generic RPC rejection: no fallback, empty
workspace, error toast, prompt stranded in the launch outbox. Pre-merge the
same action opened a legacy terminal agent.

Map the host's thrown definitive refusals onto the refusal class at the
launch boundary, so every creation flow -- present and future -- degrades to
the legacy terminal instead of stranding. Narrow predicate: unrelated
failures (ECONNRESET, empty message, non-Error) still propagate untouched.

Ablation-proven: removing the mapping reddens the fallback test.

* fix(windows): teach the mobile RPC double the status probe the lane added

CI's first-ever run on this lane caught a pre-existing lane defect. The lane
changed status.get to resolve through
runtime.getStatusAfterWindowsProcessStartTimeProbe(), but never taught the
mobile-surface runtime double about it, so status.get failed for mobile
clients with "not a function". The lane's own test list did not include this
file and the lane had zero CI, so nothing ever ran it.

The real runtime always implements the method; the double omitted it.

* chore: merge current main and regenerate the localization runtime catalog

CI static analysis failed on a stale en-runtime-required.json: main added
onboarding integration-capability keys, and the generated catalog is checked
against the PR MERGE result, not the branch alone -- so it read clean locally
while failing in CI. Merging current main (90780acb85) and regenerating.

Gates after the merge: pnpm tc 0, oxlint 0, changed-code quality 0/56,
7 gate/lane test files 69 tests green.

* fix: route structured launches by execution host platform

* fix: recover paired structured session mirror on host swap

* Revert "fix: recover paired structured session mirror on host swap"

This reverts commit 81bfca0007.

* Revert "fix: route structured launches by execution host platform"

This reverts commit 47abbd354a.

* fix(windows): refuse structured chat in a paired web client

Reverts the two review-loop commits (restoring a tree byte-identical to the
validated head) and closes the hole they were aiming at, without their cost.

A paired web client's `platform` describes the browser's machine, not the host
that will run the agent, so the Windows gate cannot be evaluated there. Before
this, a browser on macOS driving a Windows runtime read "not win32", skipped the
creation-time proof entirely and allowed structured chat — fail-OPEN, the
dangerous direction, bypassing the guarantee this lane is built on.

`isWebClient` is a required input like the other gate fields, so the compiler
enumerated all seven call sites. Refusal is synchronous and fail-closed: no
async round-trip, no null window, no cache to invalidate — unlike keying on an
asynchronously-fetched host platform, which would have made every desktop
launch wait on a round-trip to fix a paired-web-only hole.

Paired web therefore gets the legacy chat until the host publishes eligibility
itself; that is the proper fix and belongs in its own PR.

Ablation-proven: removing the guard reddens both refusal tests; the
desktop-unaffected test is a preservation check and passes either way.
Gates: tc 0, oxlint 0.

Known open: repos-onboarding-folder-startup.test.ts fails on this branch and
passes on plain main — under investigation, NOT caused by this commit.

* test(onboarding): mock the web-client check the store path now reaches

The web-client refusal added `isWebClientLocation()` to the launch-route
inputs, which this suite's store path reaches while adding the FIRST folder.
The suite stubs `window` as `{ api }` with no `location`, so the function
cleared its `typeof window === 'undefined'` guard and then threw on
`window.location.pathname`.

That threw inside addNonGitFolder's own catch, so folder-1 never activated;
folder-2 then returned early (a project already existed) before reaching the
call at all, leaving exactly one activation with no startup seed.

Test artifact, not a product defect: a real renderer always has
`window.location`, so the seeding path is intact for users. Mocking the module
is the convention 7 other suites already use, and keeps product code free of
defensive branches that only exist to satisfy a stub.

Ablation-proven: removing the mock reproduces the original failure exactly.

* fix(renderer): make the web-client check total over a partial window

isWebClientLocation() guarded `typeof window === 'undefined'` and then assumed
`window.location` existed. A window stubbed without a location cleared the
guard and threw on `.pathname`.

That matters because this branch put the call on the launch-routing path,
where the throw is swallowed by the caller's catch and silently becomes a
FAILED LAUNCH rather than a visible error. CI caught it as 9 failures in
launch-work-item-direct.test.ts.

I previously "fixed" this by mocking the module in the one suite I knew about.
That was whack-a-mole against an unbounded set, and it missed this one. The
defect is the partial-window assumption, so fix it there: the mock is removed
from the onboarding suite and both suites now pass on the hardening alone.

Ablation-proven: reverting to the unguarded form reddens 11 tests across the
new unit suite and launch-work-item-direct.

Gates: tc 0, oxlint 0, changed-code quality 0/58.

* Move Codex's Windows structured-chat eligibility onto the host createSupport probe

The renderer no longer decides Codex win32 eligibility: launchStructuredAgentSession
probes agentSession.createSupport for both providers, the host answers via
supportsCodexStructuredLocation (process start-time proof + WSL refusal), and the
create path re-checks live. Deletes the client-side windows gate module and its
routing inputs (windowsProcessStartTime, worktreeUsesWslPath, isWebClient, platform)
from six call sites. Splits killCodexAppServerProcessTree out of
codex-app-server-session to hold the max-lines ceiling without a disable.

* fix(ci): keep pnpm lockfile stable

* test(windows): align foreground snapshot flags

* Restore main's pane-snapshot flag contract

Main asks for CreationTime on both projections; this branch's hot-path
isolation went away with the async probe it served.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
Co-authored-by: Merge Sim <sim@local>
Co-authored-by: Merge Sim <merge@localhost>
2026-09-07 09:18:38 -07:00
Brennan BensonandMerge Sim 298571ad9f fix(codex): uncap app-server stdio records (#18590)
Co-authored-by: Merge Sim <sim@local>
2026-09-06 11:49:42 -07:00
OrcaWinandOrca Worker b6ca8dad99 fix(hooks): register the Claude hook script directly on Windows (#18875) (#18905)
* fix(hooks): register the Claude hook script directly on Windows (#18875)

The Windows Claude Code lifecycle hook was registered as
`powershell.exe -NoProfile -EncodedCommand <...>` whose entire decoded payload
was a `Test-Path` and a call to `~/.orca/agent-hooks/claude-hook.cmd`. Every
hook event paid a full PowerShell start-up to reach a script that exits at its
first `ORCA_PANE_KEY` guard, so sessions outside Orca paid it to do nothing.

Register the script path itself instead, with `|| echo {}` for the
neutral-JSON-when-missing contract (#14818). Measured on Windows 11, invoked as
Claude Code invokes it (`printf payload | bash -c -l "<command>"`):

  idle (n=12)          baseline 177ms | before 471ms | after 213ms
  10-way conc (n=40)            --    | before 656ms | after 296ms
  p95 under load                --    | before 696ms | after 337ms

It also drops an interpreter from the chain the hook's timeout kill must tear
down. Killing the hook does not kill its PowerShell grandchild, which still
holds the stdout handle the agent reads to EOF -- measured, EOF arrived 352ms
AFTER the kill, when the orphan exited by itself. msys2 creates children
suspended and resumes them after, so a kill landing in that window strands one
that never exits and EOF never comes; that is the reported frozen session.

The encoded launcher stays as the fallback for profile paths the shells cannot
carry bare (space, `%`, `^`, `&`, non-ASCII) and for hosts where Git Bash is not
resolvable, because PowerShell 5.1 rejects `||`. Every other agent's hook is
untouched, as is the remote/SSH path.

Not adopted from the report: `cmd.exe /d /c <path>` (MSYS rewrites the `/c`
under Git Bash -- measured, the invocation fails), and raising the 10s timeout
(the orphan survives the kill regardless; the fast path puts the hook 30x under
the budget so the kill effectively stops firing).

* fix(build): list the new hook launcher modules in the CLI tsconfig project

config/tsconfig.cli.json enumerates its files explicitly, so the two new
imports reached by src/main/claude/hook-settings.ts failed tc:cli with TS6307.
src/main/git-bash.ts pulls in only node:fs, node:path and a shared constant,
so it adds nothing heavy to the CLI project.

* fix(hooks): address review of the direct Windows Claude hook launcher

- Make the Windows hook suites host-independent. A box with a cmd.exe AutoRun
  (HKCU\...\Command Processor\AutoRun) failed them at HEAD too: the tests
  redirect USERPROFILE, the AutoRun target vanishes, and MSYS spawns a .cmd
  without /d so AutoRun runs and lands on the hook's stderr. Seed an empty
  target, including under the deliberately-absent profile.
- Note in managed-hook-stdin-lifecycle why the "missing managed script" case no
  longer exercises the fallback for the direct shape (it carries an absolute
  path, so a redirected profile changes nothing); that path is covered live in
  windows-direct-cmd-hook-command.test.ts.
- Keep the direct shape off UNC profiles: WINDOWS_CMD_SAFE_PATH admits them, but
  //server/share/... is not a command cmd.exe reliably starts.
- Correct the comments: `|| echo {}` also fires when cmd.exe itself exits
  non-zero (failing AutoRun), printing {} twice. The encoded launcher exited 1
  on that same box, so neither shape is clean there.
- Test the contract that replaced runtime %USERPROFILE% resolution (STA-3348): a
  stale absolute path reports not_installed and is rewritten on install.
- Record the standing unmeasured assumption in windows-edr-posture.md: `||` does
  not parse in Windows PowerShell 5.1, so a compat consumer that hosts hook
  strings there would fail closed. Measure before widening to another agent.
- Trim the launcher comments per AGENTS.md; the numbers live in the doc.

* test(win32): register the new Windows-gated hook test in the CI lane

win32-test-lane-registration guards against exactly this: a Windows-gated file
that self-skips on ubuntu and reports success, so it runs on no machine. The new
windows-direct-cmd-hook-command.test.ts needs both entries — WINDOWS_PACKAGE_TESTS
decides whether package_windows runs for a diff, and the workflow argv decides
whether the file runs once that job started.

* test(win32): remove the hook temp tree through the retrying helper

windows-lane-tree-removal-boundary scans exactly the specs in the Windows CI
lane, so registering windows-direct-cmd-hook-command.test.ts subjected it to the
rule: cmd.exe and bash have just exited in that tree, and a raw recursive rm
throws EPERM on Windows while their handles drain, turning a green spec into a
lane failure. Use removeTreeSync, which carries the repo's maxRetries policy.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
2026-09-05 17:50:33 -07:00
Neil f37d2fec97 fix(linux): land the reviewed Linux packaging stack on main (#18100)
* fix(linux): give the CLI one entrypoint by extracting the AppImage once

* refactor(linux): trim AppImage CLI registration seams

* test(cli): assert registration lock serialization

* fix(linux): fence AppImage terminal shim mounts

* fix(linux): accept extracted AppImage runtimes with APPDIR only

* docs(linux): make headless AppImage extraction runnable

* refactor(linux): import bundled launcher directly

* fix(linux): reclaim superseded AppImage payloads and packaged symlinks

Pruning removed 3215 of 3216 files from a superseded generation and always
stranded resources/app.asar, leaking ~105 MB per version update. Electron's
asar shim reports a *.asar file as a directory, so the recursive remove tried
to rmdir a real file and failed with ENOTEMPTY; the .catch(() => {}) hid it.
Reproduced end to end on Ubuntu 24.04: 519M -> 623M across one update, and
519M again once the payload is actually reclaimed.

removeExtractedAppImagePayload holds process.noAsar for the removal, counted
so overlapping removals cannot hand the shim back early, and the prune site
now warns with the path instead of swallowing the rejection. All three
removal sites use it -- staging cleanup and displaced roots leaked the same
way.

Also reclaim symlinks left by a packaged deb/rpm install, which the
extracted-cache-only rule turned into a hard conflict on a deb -> AppImage
migration, and name the remedy in the conflict error.

* fix(linux): bound the CLI registration lock wait

`retries: 1000` caps the attempt count, not elapsed time, so at up to 1s per
attempt an IPC-driven registration could hang ~16 minutes against a wedged
holder with no feedback.

A legitimate holder is bounded by the extraction timeout, so wait that plus
slack and then fail with a message naming the lock file, rather than hanging.
`maxRetryTime` is forwarded verbatim to the `retry` package by proper-lockfile.

* fix(linux): stop re-extracting the AppImage on inode metadata churn

The extracted-payload cache key hashed ctime alongside dev/ino/size/mtime.
ctime moves on any inode metadata write -- `chmod +x`, which every AppImage
user is told to run, plus `chown`, an ACL or SELinux relabel, and a backup
restore -- none of which alter a byte of the payload.

Measured on Ubuntu 24.04: `chmod +x` leaves dev, ino, size and mtime
identical and moves ctime alone, so the key changed and the next launch paid
a full ~519 MB re-extraction and a multi-second stall to rebuild a payload it
already had, then pruned the old generation.

Key on content identity instead. An in-place content change moves mtime and
almost always size; a replacement moves the inode. The existing
replace-in-place test still passes.

* fix(linux): stop CLI commands from falling through to Chromium startup

* refactor(cli): remove redundant command membership check

* test(cli): cover command-named project selectors

* fix(cli): redirect the open-url command before startup

* test(linux): cover AUR serve wrapper flags

* fix(linux): tighten CLI launch detection

* fix(linux): respect CLI flag value boundaries

* fix(linux): strip injected Chromium switches from CLI args

* fix(linux): report a missing display instead of dying in uv_close

* refactor(linux): read display locks without a preflight race

* fix(linux): preserve unverified external displays

* chore: format reliability gate manifest

* test(packaging): split runtime resource checks

* fix(linux): fail serve when no display is available

* fix(linux): do not treat a lockless X socket as a dead display

An X server writes its lock beside its socket and both survive a crash
(verified against Xvfb under SIGKILL), so a socket with no lock was never
left by a crashed server. It is an endpoint published from elsewhere: a
container bind-mounting only /tmp/.X11-unix, WSLg, or a foreign PID
namespace. Declaring those dead made the desktop gate exit(1) on displays
that work, with no workaround, and the serve gate refuse to start.

Liveness now splits by ownership. A foreign DISPLAY trusts a lockless
socket; Orca's own :99 does not, because removeStaleDisplayArtifacts
unlinks the lock before the socket and so manufactures that state itself --
adopting it would resurrect the orphan-socket bug and stop the cleanup from
self-healing. The stale-lock rejection is unchanged.

Also correct four doc statements this behaviour falsified.

* fix(linux): fail closed when a stale socket blocks the Xvfb rebind

Readiness only checked that /tmp/.X11-unix/X99 exists. A stale socket we
could not unlink still exists after our own Xvfb refused to bind, so Orca set
DISPLAY to a dead server and Chromium died in Ozone init.

Measured on Ubuntu 24.04 against the pre-fix build: with a leftover :99
socket and no lock, serve exits 139 (SIGSEGV), the socket inode is unchanged
before and after, and no lock is recreated -- it neither cleaned up nor
respawned. To a user that is a crash, not a misconfiguration.

This is reachable in the documented topology, where orca-xvfb.service has no
User= and runs as root while serve runs as User=orca: /tmp is sticky, so the
orca uid cannot unlink a root-owned socket, rmSync fails, and Xvfb exits with
the display already active.

Readiness now requires the display to actually be live -- our socket plus a
lock naming a running process -- so the same state reports an unusable
display and exits 1 with the existing diagnosis.

* fix(linux): recognise abstract X sockets and inherited Wayland fds

Two display setups this gate could not prove were refused outright, and on the
desktop path that is app.exit(1) with no workaround.

An X server may bind only the abstract namespace (`@/tmp/.X11-unix/X0`), which
leaves no filesystem socket to stat. Abstract addresses are kernel-owned and
vanish the moment the owner exits, so an entry in /proc/net/unix is proof of a
live server -- no lock file needed and no stale entry possible. Verified on
Ubuntu 24.04, where 139 such addresses were present.

WAYLAND_SOCKET is an already-connected fd handed over by the compositor, so
there is no path to stat and WAYLAND_DISPLAY may be unset entirely. Its
presence is the display.

Both are consulted only after the filesystem-socket check fails, so no
existing verdict changes.

* fix(linux): never treat Orca's own display number as a foreign endpoint

Recognising a lockless X socket as live is correct for an endpoint published
from elsewhere -- a container bind mount, WSLg -- because an X server writes
its lock beside its socket and both survive a crash. It is wrong for
VIRTUAL_DISPLAY_NUMBER, because Orca's own teardown unlinks the lock before
the socket and so manufactures that exact state.

The managed branch was already strict, but a caller that sets DISPLAY=:99
explicitly takes the foreign path and skipped it, accepting a dead display
left by Orca's own interrupted cleanup. Route the managed number through the
strict probe on both paths.

Found by an adversarial audit of the asymmetry introduced earlier in this
branch; the documented systemd topology is unaffected because its Xvfb writes
a real lock.

* test(linux): add a packaged-artifact contract for the CLI launch paths

* test(linux): avoid buffered serve readiness detection

* test(linux): signal AppImage serve owner directly

* test(linux): tolerate readiness timeout boundary

* test(linux): add startup margin to shutdown oracle

* ci(linux): give package contracts timeout headroom

* fix(ci): route all Linux packaging contract changes

* test(linux): poll shutdown readiness without tail leaks

* test(linux): bound shutdown cleanup grace

* test(linux): assert on CLI output, not the harness's own control lines

run-cli-case.sh echoes `RESULT status=N case=<name>`, and the two cases named
*-skills asserted `expectOutput: 'skills'`. That substring was satisfied by
the case name in the harness's own line, so 2 of 8 cases asserted nothing
about the command -- gutting `skills` entirely would still have gone green.

Control lines are now excluded before matching, and both cases assert the
rendered help header, which only real help output produces. Verified on an
Ubuntu 24.04 host: 8/8 still pass against a stack-tip AppImage.

Also register the gate in reliability-gates.jsonc, which #15085 added a CI
Docker gate without. Red/green is recorded from a stock release AppImage
failing 4 of 8, three of them at status 133 (SIGTRAP).

* fix(linux): require static AppImage runtimes (#17319)

* test(linux): reject a wrong-architecture native binary at packaging time

Cross-building the arm64 slice on an x64 host silently packed an x86-64
`pty.node` -- the rebuild logged "Forcing native rebuild for linux-arm64" and
shipped the host's binary anyway. Every gate here inspects symbol versions,
which are perfectly valid on the wrong architecture, so nothing noticed.

Observed on a Raspberry Pi 5: the packaged app loaded, then failed with
"Failed to load native module: pty.node", and the launch contract reported
3 of 8 cases crashed rather than naming the cause. Swapping in the aarch64
`pty.node` took the same build to 8/8.

Compare ELF `e_machine` against the slice being packaged and fail with the
offending path. Checked before the glibc pass, because a wrong-architecture
binary's symbol versions are valid but meaningless and would send the reader
down the wrong path.

Release CI builds arm64 on a native runner, so this guards local and future
cross-builds rather than a shipped artifact.

* test(linux): judge per-arch vendored binaries against their own path

The first CI run of the architecture gate failed the x64 package job on
`@parcel/watcher-linux-arm64-glibc/watcher.node`. That binary is arm64 on
purpose: the package ships every architecture and its loader picks the match,
so its presence in an x64 build is correct.

Judge a binary against the architecture its own path names, falling back to
the slice when the path names none. That keeps the case this gate exists for
-- `bin/linux-arm64-*/node-pty.node` holding an x86-64 binary, which is what
shipped to a Raspberry Pi 5 -- while letting multi-arch dependencies through.

Dry-run over the real dependency tree flags nothing for either target arch.

* fix(linux): move deb/rpm update installation outside Orca (#17318)

* fix(linux): complete deb/rpm package metadata

* fix(linux): preserve CLI link during package upgrades

* docs(linux): document local RPM build prerequisites

* fix(linux): move deb/rpm update installation outside Orca

* fix(updater): preserve Linux recovery across stale events

* fix(updater): fence stale downloaded events by active target

* fix(updater): preserve active Linux package recovery

* test(linux): keep workflow order assertion in scope

* test(updater): assert stale recovery stays silent

* fix(updater): preserve Linux package recovery after checks

* refactor(updater): keep Linux marker message with status

* fix(linux): describe the right manual update path for deb/rpm hosts

A remote host installed from .deb or .rpm now reports
manual-service-update-required, and the guidance told the operator to
"update through the service manager that starts this server" -- which is
correct for unsupported-headless-serve but wrong for a package install,
where nothing about the remedy involves the service manager.

Say both, keyed on how the host was installed.

* docs(linux): document orcad update restart safety

* docs(linux): scope restart census omissions

* docs(linux): use absolute service CLI launcher

* fix(serve): validate in-process serve options before startup (#17683)

* fix(linux): stop offering updates a distro-managed install cannot apply (#17918)

Closes #17702.

The resources/package-type marker is authoritative but never checked against
the host, so any repackager that unpacks Orca's .deb -- AUR, Nix, a container
rebuild -- inherits `deb` verbatim. Install feasibility was then computed
after a ~165 MB download, so those users got check -> download -> a card
promising an install command -> a dead end.

Validate the marker against the host: a deb/rpm marker with no matching
package manager in the trusted directories means a package manager owns this
install. This reuses the exact lists and resolver that
buildLinuxPackageInstallCommand already loops over, so a false positive is
impossible by construction -- any host flagged here would have failed with
no-package-manager after the download anyway. The gate only moves that
verdict earlier. Verified across Debian 12, Ubuntu 24.04, Arch, Fedora 40 and
openSUSE Leap: no false positive on a real deb host, correct on every
repackaging host.

The release is still reported, because the user does want to know 1.4.194
exists and to update through their distro; only the download path is closed.
`externallyManaged` is an additive optional field on the existing `available`
status, so older paired clients decode it unchanged. downloadUpdate() refuses
authoritatively, since main owns this verdict rather than the card, and
unwinds any pinned-build state first -- a Linux pinned jump resolves to
'release', and stranding isPinnedBuildActive would silently kill every
background check for the rest of the process.

Note the fix the issue suggests cannot work: electron-updater builds a
PacmanUpdater whose doDownloadUpdate looks for a .pacman asset Orca does not
publish, then dereferences undefined.

* style(cli): restore prettier wrapping on install error copy

* test(linux): re-pin the child-process ratchets and the batch-shim allowlist after the merge
2026-09-02 03:08:01 -07:00
Neil 20a12a6a46 perf(codex): share one launch-prep hook install across a spawn burst (#17669)
* perf(codex): share one launch-prep hook install across a spawn burst

Codex launch prep runs a full managed-hook install on every local PTY
spawn, and both install lanes serialize globally per Codex home. Opening
a multi-pane worktree therefore paid N full installs back to back, and a
resumed Codex pane prepares twice. Concurrent spawns for the same runtime
home now share one run; the promise is dropped as soon as it settles, so
the next launch still re-reads hooks.json and the user's trust state.

Also split the `host_env` spawn-timing phase, which spanned the entire
Codex preamble and pinned that cost on the env builder that ran last.

* refactor(codex): unify the two hook-install single-flight lanes

Both the WSL and launch-prep lanes now share one generic in-flight helper
instead of duplicating the map bookkeeping. Also routes the WSL launch-prep
install through the serialized variant, which closes the same per-spawn
serialization gap on WSL that the native lane just got.

* refactor: extract the shared in-flight run dedupe

The codex hook service and the GitHub conflict-summary cache had grown
near-identical private copies of the same single-flight helper. Both now
use one module, which also keeps the hook service clear of the 300-line
budget. The shared copy keeps the identity check on clear so a late settle
cannot evict a newer entry for the same key.
2026-08-31 15:16:38 -07:00
Neil 1369821bad Split Codex hook service responsibilities (#17260)
* Split speech session lifecycle

* Split terminal output scheduler pipeline

* Split mobile browser pane modules

* Prune resolved max-lines suppressions

* Split pane tree equalization logic

* Extract mobile troubleshoot screen styles

* Split external automation manager

* Split main window service attachments

* Split hosted review creation checks

* Split automation dispatch event handling

* Split settings navigation metadata

* Split daemon initialization lifecycle

* Split GitLab item dialog

* Split relay dispatcher layers

* Split mobile host screen

* Retarget mobile view settings source test

* Split runtime file client layers

* Split ports panel layers

* Split runtime environments pane layers

* Split local PTY provider responsibilities

* Split CDP bridge responsibilities

* Split relay Git handler responsibilities

* Track moved relay Git fetch audit

* Split Linear item drawer responsibilities

* Split telemetry event schema responsibilities

* Split resource usage status responsibilities

* Split remote terminal multiplexer responsibilities

* Split Git worktree responsibilities

* Split Codex hook service responsibilities

* Keep mirrored hook trust type private

* Fix F3-speech for #17123

* Fix F1-cycle for #17131

* Fix F4-navtest for #17157

* Fix F2-allowlist for #17161
2026-08-29 20:21:14 -07:00
Mark XianandBrennan Benson cc384c5a3d fix(agent-hooks): post posix payloads as json (#11292)
* fix(agent-hooks): post posix payloads as json

* fix(agent-hooks): mark header merged envelopes

* docs(agent-hooks): describe header merge envelope

* fix(agent-hooks): encode posix metadata headers

* test(agent-hooks): update WSL JSON hook assertions

* fix(agent-hooks): negotiate raw JSON transport

* fix(agent-hooks): preserve packed metadata in POSIX shells

* test(agent-hooks): include hook envelope in relay boundary inventory

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-27 12:32:29 -07:00
Brennan Benson f352e3e27d fix(cursor): emit Cursor-contract JSON from managed hooks
Merge rebased conflict repair after exact-head tests, typecheck, lint, format, and all required GitHub checks passed.
2026-08-27 00:25:31 -07:00
Neil 26721bd632 fix(codex): stop blocking the main thread on trust grants (#16441) (#16594)
* fix(codex): stop blocking the main thread on trust grants (#16441)

Codex hook trust was granted by blocking the Electron main thread on
`spawnSync` of a bundled ELECTRON_RUN_AS_NODE entry for the whole
app-server deadline: 15s native, 35s WSL, ~45s on the real-home path
(rebase inspect + repair + grant). Cold start and every Codex pane
launch showed "Not Responding"; the reported event-loop gap was
15,049 ms.

The subprocess only ever existed to donate an event loop to a
deliberately blocked parent — `runCodexHookTrustGrantSession` was
already the real async implementation. Make the callers async and the
fork is unnecessary, so the bridge, the forked entry and its envelope
are deleted along with their build/knip/tsconfig registrations. The CLI
`agent hooks prepare-codex` handler is already async, so it awaits the
in-process session and saves a process spawn per managed-home shell.

`resolveCodexTrustGrantHost` is async too; the WSL identity probe moves
from `execFileSync` to `runProcess`, dropping that file from the
child-process import allowlist. Status reads keep a synchronous
native-only stamp path.

Two invariants that held only because the lane blocked:

- Overlapping capability probes were impossible by construction.
  `GitCapabilityCache`'s dedupe engine is extracted to a shared
  `CapabilityProbeCache` and `CodexAppServerCapabilityCache` now
  inherits it, so concurrent launches against a cold host share one
  app-server session instead of one each.
- Two grants on one `config.toml` could not interleave capture and
  restore. A reentrant per-file lane now serializes the whole install
  sequence (managed, WSL runtime, real-home ensure, legacy sweep) and
  the grant and rebase inside it.

Cold-start work moves off the critical path: retained-home
reconciliation (N sequential sessions) is fire-and-forget behind the
daemon provider, and the startup real-home ensure chains into managed
hook reconciliation instead of blocking app init.

Every preserved semantic is unchanged: never throws, the
ORCA_DISABLE_CODEX_TRUST_RPC kill switch, ledger hits, backfill-pending
and cooldown fallbacks, config rollback on every failure path,
pre-grant self-computed trust removal, the verify-failure taxonomy,
diagnostics and telemetry.

* fix(codex): widen the trust-config lane to every config.toml writer

Review follow-ups on #16441's async trust grant:

- `markCodexProjectTrusted` now runs inside the runtime+system config.toml
  lanes, so a project-trust write can no longer land inside a hook grant's
  capture->restore window and be silently reverted. Its callers await it.
- `install`/`refreshRuntimeUserHooks`/`remove` hold the system config.toml
  lane as well as the runtime one — they promote approvals into
  ~/.codex/config.toml and mirror it back. Lock order is runtime-before-system
  everywhere.
- The real-home ensure chain resumes after a rejection instead of returning
  the same rejected promise to every later pane launch, and resolving the real
  home is now inside the module's never-throws boundary.
- `buildSpawnEnv` awaits inside a cancelable pending-spawn registration, so
  shutdown during the (now long) env build stops the PTY from launching.
  `prepareLocalPtySpawn` generalizes into `awaitCancelableLocalPtySpawn`.
- CapabilityProbeCache drops the test-only `nowMs` passthrough; its probe
  backstop comment now describes what it actually guards.
- Preflight is a plain async function; the trust dispatch in orca-runtime
  collapses into one `markWorkspaceTrustedForAgent`.

* test(codex): exercise the trust-config lane under real concurrency

The async grant makes two pane launches overlap for the first time. These
drive the real modules end to end on real files: a rollback swallowing a
sibling's grant, a markCodexProjectTrusted write landing inside a capture
-> restore window, shared capability-probe dedupe on a cold host, the
host-scoped transient cooldown, and reentrancy from inside an installer.

Each was verified to fail against a deliberately broken implementation
(lane removed, dedupe disabled, cooldown made global, reentrancy pass-
through disabled).

* test(codex): stop hook-service suites spawning the developer's real codex

The forked grant bundle never existed under vitest, so the RPC lane was
unreachable in tests on main. Running it in-process makes these suites
spawn a real `codex app-server` when one is installed: 38 spawns and two
failures in hook-service-runtime-trust-repair on a machine with codex,
green in CI where there is none. Stand in for the missing binary so both
environments exercise the same fallback lane.

* docs(codex): scope the trust-RPC kill switch comment to what it actually gates

The comment read as though the flag forces the fallback lane everywhere. It
gates the managed grant only: the real-home rebase still runs its own
inspect/repair app-server sessions when Orca's insertion shifts a user's hook
positions, and never reads the flag.

Verified by exercise, not by reading — with the flag set, both
inspect-user-hook-trust and repair-user-hook-trust still ran. Pre-existing:
main has no check there either, it just blocked the main thread while doing it.

Widening the flag to cover the rebase is a follow-up; this only stops the
comment promising something the constant does not do.
2026-08-26 16:44:55 -07:00
Brennan BensonandSiddiqui Qamar 5a59bc5bc4 fix(grok): stop Orca's Grok hooks from costing anything outside Orca (#16666)
* fix(grok): stop Orca's Grok hooks from costing anything outside Orca

Orca registers Grok agent-status hooks in the global $GROK_HOME/hooks. Grok
loads that directory on every session, so a Grok run that Orca did not launch
still paid for the hook on every event, and Orca rewrote the file even after a
user had emptied it to opt out (#15518).

The registered POSIX command now guards on ORCA_PANE_KEY before doing anything.
That variable is part of the pane identity Orca injects into terminals it
launches, and unlike the port and token it never comes from the endpoint file,
so it is present exactly when the session belongs to Orca. A standalone session
short-circuits without spawning a shell for the managed script at all. The same
guard is applied to the remote install, because a remote host runs standalone
Grok sessions too.

PreToolUse is no longer registered. It is a blocking hook, so Orca sat on the
critical path of every tool call and doubled the per-tool spawns, for a
transition PostToolUse already reports.

Windows cannot use the guard: the command there must be a single spawnable
token, so it is a bare script path with no shell to evaluate a test. For that
case the hooks are removed when Orca quits -- locally, on WSL guests, and on
connected SSH hosts -- and reinstalled on the next launch. A config the user has
emptied is left alone on startup; turning the setting back on in Settings is an
explicit and later choice, so that path reinstalls.

Removal is careful about what it is deleting. It strips only Orca's own entries,
keeps user-authored ones, and deletes the file only when no hook entries remain
-- keying that off the whole object would leave a stray non-hook key behind, and
the emptied-config check would then read that remnant as a deliberate opt-out
and never reinstall. A config the user has symlinked into a dotfiles repo is
written through rather than unlinked, and is exempt from the emptied-config
check for the same reason: after a quit it is a file Orca emptied, not one the
user did.

Writes go through temp+rename. Grok refuses to build a sandbox profile for a
hook JSON with more than one hard link, so publishing by hard link would fail
any session that started during the write.

Install and removal on remote hosts now read the platform from the same field.
They did not, so a Windows remote whose bridge env was incomplete had hooks
installed and never removed.

Co-authored-by: Siddiqui Qamar <137684575+siddqamar@users.noreply.github.com>

* fix(grok): preserve hook state outside Orca

---------

Co-authored-by: Siddiqui Qamar <137684575+siddqamar@users.noreply.github.com>
2026-08-26 12:48:52 -07:00
Neil 48e63c015f refactor agent config and auth services (#16195)
* refactor: split agent config and auth services

* chore: repoint wsl and global-fetch guards at split module paths

* fix: restore merge-base Claude CLI error propagation

Drop the secret-redaction rewriting added to Claude CLI error paths in the
refactor: spawn errors again reject with the original Error (preserving
.code/.errno/.syscall/.stack) and command output/auth-status logs are no
longer rewritten.
2026-08-24 23:15:01 -07:00
Brennan Benson 0b80a773a4 fix(codex): stop overwriting and deleting Codex files that were merely unreadable (STA-4737) (#15287)
* fix(codex): stop overwriting and deleting Codex files that were merely unreadable (STA-4737)

Three modules shared by the host and WSL Codex lanes decided a file was absent
from a read that had only failed, and then wrote over it or removed it.

- `codex-config-mirror`: `existsSync` on the RUNTIME config.toml returned false
  for a locked file exactly as for an absent one, so the mirror took the
  "seed a fresh runtime config" branch and replaced the user's config wholesale.
- `config-settings-promotion`: an unreadable ~/.codex/config.toml counted as
  having no promoted settings, and the write path then rebuilt the user's
  canonical Codex config from Orca's runtime copy.
- `codex-home-paths`: both delete branches in `linkSystemCodexResource` remove
  Orca's mirrored copy because the system resource "is not there". `existsSync`
  and `systemResourceIsRegularFile`'s `catch { return false }` both reported
  that for a source nobody could read, so one denied read on ~/.codex/AGENTS.md
  removed the managed copy on the next launch.

`src/shared/definitive-filesystem-absence.ts` now owns the one errno allowlist —
ENOENT and ENOTDIR, with every other code including unrecognised ones treated as
indeterminate — and `host-codex-managed-home-ownership.ts` drops its private
copy rather than letting the two drift. `codex-path-observation.ts` builds the
three-valued observation on top of it.

The resource sync's two `existsSync`/`statSync` probes collapse into one
resolved stat, which answers reachability and regular-file-ness together and
closes the window between them.

`config-settings-promotion.ts` crossed its max-lines budget, so the write-target
resolution moves to its own module rather than taking a lint exemption.

Deliberately not here: the hook-service trust writes that run after a refused
mirror, and the promotion write target's own classification, which is
unreachable because it always resolves to the same file the read above already
refused. Both are noted in comments rather than half-built.

* fix(codex): preserve resource copies on indeterminate reads
2026-08-18 14:10:32 -07:00
Brennan Benson 8ea5dd80c3 fix(antigravity): install a PreToolUse status hook without deciding tool permissions (#14701)
* fix(antigravity): install a PreToolUse status hook without deciding tool permissions

Antigravity is the only supported agent with no pre-tool signal, so its panes
show a bare "Working" spinner for the whole tool call instead of the live
"Working - <tool>(<input>)" readout every other agent gets.

The consumer side already handles it — extractAntigravityToolFields and
normalizeAntigravityEvent parse PreToolUse (including the `waiting` state for
ask_question/ask_permission) and are covered by tests. Only the installer was
missing the event.

PreToolUse was installed originally and removed in a480e6b7 (#2501) because the
observational `{}` response violates Antigravity's schema and is read as a deny
(#2426). Re-add it with the one documented decision that defers to the user's
permission config rather than overriding it:

- `ask` — "Prompts the user, but respects 'Always Allow' settings."
- `allow` — "Automatically allows the tool execution." (would silently
  auto-approve every tool call Orca observes)

The hooks.json guard for a missing managed script previously drained stdin and
printed nothing, which is exactly the #2426 deny. Gate events now emit their
response from the guard too, so a swept ~/.orca cannot brick tools on a host
whose ~/.gemini/config/hooks.json survives.

Fixes #12898

* fix(antigravity): list posix-hook-command.ts in the CLI typecheck project

The CLI project enumerates its main-side files explicitly, so the extracted
module needs an entry alongside its sibling runtime-home-hook-command.ts.
2026-08-18 00:44:56 -07:00
OrcaWinandOrcaWin 02ba70a847 fix(agent-hooks): make the Windows managed hook survive Claude-hooks-compat consumers (#14825)
* fix(agent-hooks): make the Windows managed hook survive Claude-hooks-compat consumers

`~/.claude/settings.json` is not read only by Claude Code. Third-party
Claude-hooks-compat layers (cursor-agent, Devin) import the same file and
reimplement hook execution, so Orca's entry has to survive consumers that
support strictly less than the documented schema. Three separate defects
came from assuming otherwise.

1. The entry depended on `args`, which a compat consumer ignores.
   `args` is valid Claude Code syntax, but cursor-agent spawns `command`
   alone -- so `conhost.exe` ran bare, which opens an interactive console
   that never closes. Hook payloads were typed into those stranded shells
   (#14815). The entry is now one self-contained `command` string that
   depends on nothing optional.

2. `conhost.exe --headless` never relayed anything. It implements the
   ConPTY server protocol, not a generic no-window wrapper: it does not
   wait for the hosted process and relays neither exit code nor stdout.
   Measured directly -- `conhost --headless cmd /c "echo X& exit /b 42"`
   yields empty stdout and no exit code, while the replacement returns
   both and waits. So every hook was fire-and-forget, and whatever it
   printed was discarded. Replaced with `-WindowStyle Hidden`, which
   suppresses the window and keeps wait/exit-code/stdout intact.

3. The hook never wrote anything to stdout. Guards exited silently and
   curl's output went to nul. Claude Code documents empty stdout as "no
   decision", but cursor-agent treats PreToolUse as a permission gate,
   fails to parse empty stdout as JSON, and blocks the tool call -- so
   every shell command in every cursor-agent session on Windows failed
   (#14818). The script now writes `{}` first, on both the Windows and
   POSIX branches, which is documented to be identical to writing nothing
   for real Claude Code. Gemini and Antigravity already did this.

Defects 2 and 3 are causally linked: `{}` cannot reach any consumer while
conhost is swallowing stdout, so neither fix works without the other.

Also fixed while establishing the contract:

- The launcher's own missing-script fallback returned empty stdout,
  reproducing #14818 whenever `~/.orca` was cleaned or an install was
  half-finished. It now emits `{}` too.
- PowerShell serializes progress records to stderr as CLIXML when stderr
  is redirected; a consumer merging stderr into stdout would see those
  bytes before the JSON. Every encoded payload now silences progress.
- `runtime-home-hook-command.ts` built its own launcher without window
  suppression -- exactly the drift #14815 asks to prevent. All launcher
  construction now goes through `windows-powershell-hook-launcher.ts`, so
  the switch list cannot be present in one installer and missing in
  another.
- Renamed `usesWindowsHeadlessHook` to `usesWindowsPowerShellLauncher`;
  nothing is headless anymore, and the flag selects a launcher.

Testing: the new regression test asserts the effect a consumer observes
-- it runs the exact `command` string from settings.json through both
cmd.exe and Git Bash, across the guard-exit, reached-curl, and
missing-script paths, and parses stdout. Verified it fails when
`conhost --headless` is reintroduced. The previous tests all asserted
installer intent, which is why they passed through all three defects.

* fix(agent-hooks): close hook launcher review gaps

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-16 20:48:26 -07:00
Neil bc28107864 refactor(hooks,relay): split agent hook services and relay under the max-lines budget (#14725)
The four agent hook services, the main hooks module, and the two relay modules
each carried a file-level `eslint-disable max-lines` and ran 365-628 counted
lines against a 300-line budget. AGENTS.md calls for splitting rather than
suppressing, and config/max-lines-baseline.txt is a shrink-only ratchet, so this
removes all seven suppressions and prunes their entries (341 -> 334).

Pure move, no behavior change. Each hook service splits into its managed script
source, its config/bundle serialization, and its remote-install path, keeping the
per-agent integrations independent: copilot, amp, antigravity and hermes each
retain their own getManagedScript rather than sharing one, because each emits a
different script body for a different agent. Merging them by name would have
been a behavior change, not a refactor.

For antigravity the suppression's stated rationale -- that local install, Windows
wrapper generation, status cleanup, and SSH remote install must share one event
list and managed-command matcher so stale-hook cleanup cannot drift by platform
-- is now enforced structurally instead: both install paths call
buildInstalledConfig + createAntigravityManagedCommandMatcher over the single
ANTIGRAVITY_EVENTS catalog, with the graph a strict DAG.

Also registers the six new antigravity/ and copilot/ modules in
config/tsconfig.cli.json. That project uses a curated `include` list rather than
a glob, so an unlisted module fails `tsc -p config/tsconfig.tc.cli.json` with
TS6307 even though the entire unit suite passes.

Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(remaining failures are pre-existing load flakes in untouched files, green when
re-run serially), no new runtime import cycles, and no lint suppression added.
2026-08-15 18:25:37 -07:00
Brennan Benson 537864a248 Fix Codex hook trust before manual shell launches (#14326)
* fix codex hook trust before shell launch

* fix packaged cli preflight dependency

* fix codex shell preflight safety

* fix Codex shell preflight settings and startup safety
2026-08-13 17:02:28 -07:00
OrcaWinandBrennan Benson f226fcfc4b fix(claude): make managed hook paths portable (STA-3348) (#13442)
* fix(claude): make managed hook paths portable

* perf(claude): keep portable hooks shell-native

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-11 14:03:16 -07:00
OrcaWin c96ded8dfd fix(startup): restore Windows PATH before shell changes (#13792) 2026-08-11 11:50:21 -07:00
Brennan Benson 2ee43bfc0d fix(agent-hooks): refresh existing Orca launchers when agent CLIs are unavailable (#13378)
* fix(agent-hooks): refresh existing shared hook scripts when the CLI is no longer detected

A CLI that falls off PATH (moved npm prefix, relocated shim) keeps its user-wide
config invoking Orca's launcher script under ~/.orca/agent-hooks, but the
presence gate skips install() with no removal — freezing the script at whatever
Orca generated last. Anyone in that state kept the pre-#11568 more.com-leaking
.cmd forever, because no launcher script is ever deleted and Windows startup
deliberately skips shell PATH hydration.

Reconcile before gating: every existing shared launcher/statusline script is
rewritten to the current template on each install pass. Creating scripts stays
behind the presence gate — an existing file is proof of a prior install; a
missing one means the gate did its job. Amp and Hermes are deliberately absent:
they write provider-native plugin code with its own install lifecycle, not
shared launchers.

- refreshManagedScriptIfPresent() in installer-utils (no-op unless the file exists)
- refreshManagedScripts() on the 11 launcher-writing services (openclaude via
  the shared Claude class)
- reconcile pass in installManagedAgentHooks before presence detection,
  filtered by the agents option, best-effort per agent
- coverage gate: a launcher written to ~/.orca/agent-hooks without a matching
  refresher entry fails the suite, in both directions

* perf(agent-hooks): refresh launchers off the main thread

* test(agent-hooks): keep refresh mode assertion POSIX-only
2026-08-10 16:34:15 -07:00
Brennan Benson f0443c326a fix(codex): recover interrupted state DB backfills (#12617)
* fix(codex): recover interrupted state DB backfills

* fix(codex): detect mixed-case backfill timeout

* fix(codex): harden backfill recovery review findings

* fix(codex): keep process identity retries safe
2026-08-07 14:06:29 -07:00
Wooseong KimandJinwoo-H f057cbc85f fix(serve): recognize CLI-form serve args on the Electron process (#12818)
* fix(serve): recognize CLI-form serve args on the Electron process

When the binary is launched as `… serve --port …` without the CLI rewrite
that injects `--serve`, normalize argv so isServeMode, headless GPU flags,
and serve option parsing all engage.

Preserves existing `--serve*` flag behavior for the CLI-spawned path.

Fixes #12677

* fix(serve): treat only CLI subcommand position as serve

Parse bare `serve` as the first positional token after flags/values so an
option value named `serve` cannot enable headless mode.

Addresses CodeRabbit on #12818.

* fix(serve): keep CLI redirects ahead of the serve argv rewrite

Rewriting argv before maybeRedirectAppImageCliLaunch replaced the `serve`
positional with `--serve`, so the redirect's command-name lookup saw a port
number and bailed — dropping AppImage serve launches out of the CLI path.

Also translate `--port=6768` (the CLI accepts it, getServeOptions only reads
the next token) and the mixed `--serve --port` form, so a security-shaped flag
like `--no-pairing` can no longer read as accepted while pairing stays on.
Map lookups replace `in` on object literals, which turned a stray `serve
toString` positional into a function spliced onto argv.

* fix(serve): close the CLI-form serve gaps found in review

second-instance: shouldActivateDesktopForSecondInstance matched only `--serve`,
so a duplicate `<binary> serve --port …` — the ExecStart shape documented in
docs/reference/headless-linux-server.md — promoted the live headless server to a
desktop window, un-fixing #11935 on exactly the launch shape this PR legitimizes.

findServeSubcommandIndex consumed a flag's value unconditionally while the
rewrite consumed it only when the next token was not flag-shaped. The two could
disagree and swallow the `serve` token, leaving `--serve` uninjected: #12677
again in a new shape (`--port --port serve`, `--port -- serve`). Both scans now
share one definition of value consumption.

`<binary> serve --help` / `serve help` bound a network-exposed runtime server
with pairing on and printed nothing; the AppImage redirect already routes those
three tokens to the CLI, so refuse them here too.

`--no-pairing=false` translated to `--serve-no-pairing` with the value dropped,
disabling pairing for an operator who asked for the opposite. The CLI reads its
serve booleans as `flags.get(name) === true`, so a boolean is now translated only
in its bare form and the `=` form rides through as the CLI treats it.

Tests: spec-derived parity between src/cli/specs/serve.ts and the rewrite,
covering both ends of the contract (serveOrcaApp and getServeOptions); a
source-text lock on the index.ts redirect/rewrite ordering, which reverted
silently green before; an exhaustive self-consistency property test; and the
real GUI launch argv shapes that must never enter serve mode.

---------

Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
2026-08-06 23:56:34 -07:00
KyouandOrcaWin 74ac7049ec fix(windows): make managed grok-hook.cmd safe when GROK_HOME is unset (#11782)
* fix(windows): make managed grok-hook.cmd safe when GROK_HOME is unset

Fixes #9358 and #9941.

cmd.exe expands %VAR:~n,m% at parse time. When GROK_HOME is unset (default
outside Orca terminals), the generated length/trailing-backslash guards
became a syntax error and every Grok hook event failed with exit 255.

- Skip substring work when GROK_HOME is undefined (if defined + goto)
- Replace if "%x:~-1%"=="\" (itself a quote-parser bug) with findstr
- Extract Windows script builder; add template + spawn tests

* fix(windows): harden grok-hook GROK_HOME guards and tests

Address review on #11782:
- Inject grokHome via buildWindowsAgentHookPostCommand extra form lines
  (no fragile string replace of the shared payload line)
- Spawn tests delete GROK_HOME and keep PORT/TOKEN/PANE_KEY set so the
  GROK_HOME path actually runs before curl

* fix(windows): cover Grok hook home boundaries

---------

Co-authored-by: OrcaWin <alpha-eng@stably.ai>
2026-08-05 13:31:26 -07:00
Jinwoo HongandOrcaWin 8f7692aa12 Fix packaged skills CLI runtime ownership (#11627)
* fix(cli): make packaged skills runtime self-contained

* fix(cli): address packaged skills review feedback

* ci(cli): smoke packaged skills on Windows

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-30 18:27:16 -07:00
650dd48ec9 feat(cli): add orca account add / account list for headless hosts (Claude + Codex) (#9177)
* feat(cli): add `orca account add` / `account list` for headless hosts

The desktop "Add account" UI is disabled when the renderer drives a remote
runtime (isRemoteAccountScope === kind:'environment'), so a headless server
reached from a remote desktop/web client has no way to register managed
Claude accounts. Add a host-local CLI path that reuses the existing capture
logic:

- ClaudeAccountService.addAccountFromConfigDir(): register a managed account by
  capturing credentials from an already-authenticated CLAUDE_CONFIG_DIR instead
  of spawning the interactive browser login (extracted persist/rollback helpers
  shared with the existing add flow)
- RPC accounts.addClaudeFromConfigDir, bridged via OrcaRuntime; rejected for
  mobile device tokens (host-local only)
- `orca account add` runs `claude login` in the user's own terminal into a temp
  CLAUDE_CONFIG_DIR, then registers it via the local runtime; `orca account list`
  lists managed accounts

Switching (select) already works from a remote client; only adding was blocked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(cli): support Codex in `orca account add` / `account list`

Mirror the Claude headless-account CLI for Codex:

- CodexAccountService.addAccountFromHome(): register a managed Codex account by
  importing auth.json from an already-authenticated CODEX_HOME, reusing a shared
  persist helper extracted from doAddAccount (no interactive login spawned here)
- RPC accounts.addCodexFromHome + OrcaRuntime.addCodexAccountFromHome bridge,
  rejected for mobile device tokens (host-local only)
- `orca account add --agent claude|codex` (default claude); `orca account list`
  now renders both Claude and Codex managed-account blocks

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: cover headless account-add capture paths (Claude + Codex)

- ClaudeAccountService.addAccountFromConfigDir: registers a managed account by
  capturing an authenticated CLAUDE_CONFIG_DIR; rejects and rolls back when the
  dir has no .credentials.json
- CodexAccountService.addAccountFromHome: imports auth.json from an
  authenticated CODEX_HOME into a managed account; rejects when auth.json is
  missing

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: address CodeRabbit review on headless account-add flows

- CLI login spawn uses a shell on Windows so `.cmd` agent shims resolve without
  ENOENT (args are fixed literals, no injection risk)
- Claude capture skips the `.credentials.json` precheck on macOS, where creds
  live in the Keychain and captureAuthFromConfigDir reads them
- Claude add rollback is best-effort: a failed rematerialization no longer skips
  managed-auth cleanup or masks the original add error
- Codex persist restores the prior account/selection if a post-write sync or
  rate-limit refresh fails, so a failure can't leave a dangling managed account
- Codex sync passes the account's selection target (correct runtime for WSL)
- Add JSDoc to the new public service methods and CLI functions

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(cli): harden headless account capture

* fix(cli): correct account command flag surface and interrupt cleanup

- `account` commands no longer accept or advertise the browser `--page`
  flag; `supportsBrowserPageFlag` allow-listed them by omission, so
  `orca account list --page x` was silently accepted and `--help`
  rendered a browser-only option
- account specs declare GLOBAL_FLAGS, so `--help`/`--json` render in the
  Options block like every other command
- `--agent` on `account add` documents the account provider instead of
  the terminal TUI-agent meaning inherited from the shared flag table
- a SIGINT/SIGTERM during the interactive login now removes the temp
  login dir (and restores the macOS Keychain item) before exiting 130;
  Node terminates without unwinding `finally`, which stranded live OAuth
  credentials on disk

* perf(cli): stop `account list` forcing a provider usage refresh

`accounts.list` awaited refreshAccountsForMobile(), which runs
fetchAll({ force: true }) — bypassing both the poll throttle and the
per-provider Retry-After gate — then O(N) serial per-account round
trips. `orca account list` renders only emails and the active ids, so
all of that work was discarded. The RPC now takes `refreshUsage`
(default true, so mobile and web keep the forced lane) and the CLI opts
out. Older hosts declare `params: null` and ignore the field, so a newer
CLI degrades to the previous behavior rather than failing.

Also documents on `account list` that `--environment` does not retarget
it, matching the host-local behavior of shouldIgnoreRemoteSelection.

* fix(cli): survive repeated and hangup signals during account add

withInterruptCleanup latched cleanup behind a boolean, so a second signal
got an already-resolved promise and its process.exit fired while the first
cleanup was still inside a Keychain call (3s each) — the temp dir's OAuth
credentials and the swapped macOS Keychain item both survived. Memoize the
cleanup promise so every signal awaits the same run, and register with
`on` instead of `once` so a second Ctrl-C cannot fall through to Node's
terminate-immediately default mid-cleanup.

Handle SIGHUP too. This flow exists for headless/SSH hosts, where the most
likely interrupt is the connection dropping, which hangs up the login's
terminal and previously ran no cleanup at all.

Warn when the interrupt lands after sign-in completed: the runtime finishes
the add independently of this process, so exiting 130 silently would tell
the user it was cancelled when the account may exist.

Reject a valueless `--agent`; the parser turns it into boolean true, which
silently ran a full OAuth login for Claude when the user asked for another
provider.

Also lock two behaviors the refactor changed but left uncovered: a WSL Codex
add must sync the WSL runtime lane rather than the default host lane, and
rename the account-spec help test to describe the Options block it actually
asserts rather than the usage string it never reads.

* fix(build): bundle the main modules the account CLI imports

electron-vite cleans out/main and emits only its declared entries, and
`build:desktop` runs it after `build:cli`, so the tsc-emitted copies of
`claude-accounts/keychain`, `codex-cli/command` and `win32-utils` were
deleted before packaging. Both `orca account add` and `orca account list`
then died at require time with "Cannot find module
'../../main/claude-accounts/keychain'" — reproduced against a real
`--serve` host. `agent-hooks/managed-agent-hook-controls` already carried
an entry for exactly this reason; these three were missing.

Adds a parity test so any future CLI import of a `src/main` module fails
in CI rather than at a user's shell after packaging.

* test: cover the desktop add-path behavior this PR changes

Both changes ride in the persist/rollback helpers the existing GUI add
flow shares with the new headless path, and neither had coverage:

- Claude: rollbackAddAccount now guards forceMaterializeCurrentSelection-
  ForRollback, so a rejecting rematerialization no longer replaces the
  real add error nor skips safeRemoveManagedAuth. Asserts the original
  error surfaces and the throwaway auth dir is gone.
- Codex: the desktop add now passes the account's selection target to
  syncForCurrentSelection, matching reauthenticate and select. Asserts
  the host target alongside the existing WSL assertion.

Both fail when the corresponding change is reverted.

* fix(cli): close the remaining account-add interrupt and preflight gaps

The round-1 interrupt fix detached the signal handlers before running the
finally-path cleanup, so the very window it was meant to protect — the two
serial 3s `security` calls plus rmSync on the success/error path — was
still covered only by Node's terminate-immediately default. Both review
lanes reproduced it independently. Await cleanup first, detach in a nested
finally, and stop a cleanup failure from replacing the error that actually
explains why the add failed.

Do not burn the interactive login when the runtime is unreachable. The
RuntimeClient is lazily constructed and the first call was the registration
RPC itself, so "Requires the Orca runtime to be running" was discovered
only after the user completed a full OAuth round trip. Preflight with the
now-cheap `accounts.list { refreshUsage: false }`.

Reject `--environment` / `--pairing-code` on `account add`.
shouldIgnoreRemoteSelection pins account commands to the local runtime, so
`orca account add --environment homelab` silently registered the account on
the laptop instead of the headless host it names.

Survive a daemon that cannot spawn `claude`. `allowFailure` is honored in
onClose but not onError, and unlike the GUI flow nothing has run `claude` in
the daemon before this point — so a launchd/systemd daemon with a minimal
PATH hard-failed an add the user had already signed in for, even though
identity resolves fine from the config dir's oauthAccount.

Also align the `--agent` help description with the global flag column.

* fix(cli): reject runtime selectors on `account list` too

`orca account list --environment homelab` was accepted and silently
listed the LOCAL machine's accounts, because shouldIgnoreRemoteSelection
pins account commands to the local runtime. Documenting that in --help
does not reach someone who already typed the flag, and answering with the
wrong host's accounts is the specific wrong answer they would act on.

`account add` already errors; this makes the new command group internally
consistent. The other groups in shouldIgnoreRemoteSelection keep their
existing silent-ignore behavior — changing those is not this PR's job.

* test: harden account-add signal tests and cover cleanup failure

- Identify the handler under test by set difference instead of
  `process.listeners(sig).at(-1)`. Vitest installs its own once-wrapped
  SIGINT teardown, so the positional lookup could grab the wrong listener;
  the helper also asserts exactly one new listener was added.
- Mock rmSync while keeping the real implementation by default, so the
  temp-dir assertions elsewhere stay honest.
- Cover that a cleanup failure in the `finally` does not replace the error
  explaining why the add failed. Fails when that guard is removed.

Completes the review loop's final round; the loop died on an API error
before it could commit this, and its `import()` type annotation would
have failed oxlint.

* fix(cli): harden interactive account add

* test(cli): make account cancellation coverage portable

* fix(cli): preserve merged skills runtime modules

---------

Co-authored-by: Dominik <marketing@gavaplast.sk>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-07-30 12:50:07 -07:00
Sebastián Castañoandscastanoh21 676ef7fab8 feat(cli): add orca skills install and orca skills update for headless skill setup (#9201)
Adds `orca skills install` and `orca skills update` so skills can be set up without the GUI — SSH hosts, containers, CI. Previously `orca skills` had only `list` and `get`, so there was no headless path.

**Agent targeting is scoped explicitly rather than delegated to detection.** The `skills` CLI decides which agents to install into, and with `-y` and zero detected agents it takes `targetAgents = validAgents` — all ~75. That is not a corner case for a headless CLI: a fresh SSH box or container with no agent installed is the normal starting state. Measured on a bare host, the unscoped command created **52 top-level agent directories and 54 junctions** (one real payload in `~/.agents/skills`, the rest links) on Windows, and 52/53 on macOS.

The CLI now passes `--agent` derived from Orca's own detection, mapped to the `skills` key namespace, plus `universal`. Supplying `--agent` makes `runAdd` use it directly and never call `detectInstalledAgents()`, so the fan-out branch is unreachable. On a bare host it now refuses with `No coding agent detected on this host` and exit 1, creating nothing. Same command with scoping: **1 directory, 0 junctions.**

`universal` alone would under-install — Claude Code is not in that set, and 19 of 28 mapped keys write agent-private homes `universal` never touches. `--agent '*'` is the bug itself. The mapping is hedged three ways: `null` for any agent whose key could not be confirmed, `satisfies Record<TuiAgent, …>` so a new Orca agent is a compile error, and a test pinning every mapped key against the CLI's own valid list.

Fixed during review — two holes that each restored the full fan-out through a different door:
- `--agent ','` trimmed to nothing, which skipped the refusal *and* emitted no `--agent`.
- `--agent -y` passed an emptiness check, and the vendor CLI silently drops `-`-leading values, re-emptying its list.

The real invariant is argument *shape*, not emptiness, and it is now enforced at the choke point in `buildAgentFeatureSkillInstallArgs`, so no caller can emit `-y` without a usable target. `*` remains allowed — asking for every agent explicitly is a choice, not an accident. Verified with 51 hostile inputs through the built binary, each recorded argv replayed through the vendor's own parser.

Also fixed: the `ORCA_CLI_CWD` refusal now runs before target resolution (it was quoting the wrong host's agent list), and `--dry-run` is refused in a forwarded shell rather than printing a command naming the wrong machine.

Validated on a real Windows host across PowerShell 7, PowerShell 5.1, cmd.exe and Git Bash: `.cmd` shims route through `cmd.exe` and `.exe` shims spawn directly (proved with instrumented shims, not inferred), the ENOENT path produces an actionable error rather than a silent failure, and `skills update` genuinely restores a corrupted skill byte-for-byte.

Known, not addressed here — both upstream behaviours this only forwards: a partial install failure exits 0, and "no installed skills found" exits 0. Both are invisible to the headless callers this feature exists for.

Co-authored-by: scastanoh21 <scastanoh21@gmail.com>
2026-07-30 11:20:29 -07:00
Brennan Benson f8b553b7d5 fix(agent-hooks): skip unavailable agent homes (#11442)
* fix(agent-hooks): skip unavailable agent homes

* refactor(agent-hooks): separate Pi and OMP home fix

* test(agent-hooks): update merged protocol harnesses

* fix(agent-hooks): avoid redundant reconciliation

* fix(agent-hooks): harden reconciliation and detection

* test(agent-hooks): cover settings reconciliation

* fix(agent-hooks): hydrate PATH for paired clients
2026-07-29 20:19:18 -07:00
Brennan Benson 9c5d827d6a fix(codex): keep history, restarts, and account identity across an account switch (#10770)
Fixes #10757. Switching Codex accounts broke three ways, all rooted in the
self-contained per-account CODEX_HOME from #9501.

HISTORY DISAPPEARED. Codex's own /resume picker only lists rollouts under the
launch CODEX_HOME, and nothing bridged history into a per-account home — only
the AI Vault's discovery scan knew about the other homes. Every other
Orca-visible home's rollouts are now hardlinked in, on selection and again at
launch, so one physical log is listed everywhere.

THE RESTART PANEL STUCK. A queued restart was only drained by a mounted
TerminalPane, but the prompt covered every stale pane in the worktree including
parked and cold-deferred tabs. Requesting a restart now answers the prompt
immediately while the pane keeps its pending restart, and a pane drains it when
its reconnected PTY binds.

PANES STAYED ON THE OLD ACCOUNT. CODEX_HOME is fixed in a shell's environment at
spawn and the daemon keeps those shells alive across app restarts, while the
restart notices are renderer state and are discarded. Each PTY's launch account
is now recorded on disk and compared against the current selection at startup.

Also merged in: #10802 (a dismissed notice no longer kills the pane's keyboard),
#10803 (the sweep arms on real PTY binds, and launcher Codex panes are no longer
filtered out by Windows deepest-process reporting), #10804 (a resume-pinned pane
now says which account it is on), #10870 (the restart card no longer parks focus
on its destructive Restart button), #10853 (the retry ladder is widened past the
Windows worst case).

Six independent reviews found real defects in every original PR, several of them
dead-keyboard bugs and three introduced by the fix for another defect in the same
loop. Live QA on macOS covered every PR; Windows was validated three times.

WINDOWS: pass 1 found two defects that made the stale-account fix a no-op there
(the sweep fired before any PTY was bound and never retried; launcher panes were
filtered out). Pass 3 at the merged head: the prompt appears on its own after a
restart — warm ~3.7-4.2s, cold ~21s needing rung 4, so #10853's widening was
load-bearing rather than precautionary; an ordinary sentence typed into a healthy
pane while another pane's card is up reaches that pane and kills nothing; a pane
running vim after exiting Codex gets no card, still none 45s later. auth.json
byte-identical across every pass.

KNOWN GAPS, stated rather than implied: #10804 is unverified on Windows
(auto-resume could not be manufactured there); cross-volume Windows is untested
and expected to yield no bridged history (EXDEV, and Codex ignores symlinked
rollouts); a cold-parked pane never binds so the sweep never covers it; the
subagent-deepest launcher shape could not be reproduced on Windows, so that
branch is fixture-verified only; WSL passed isolation but the resume mechanism is
host-lane only. A host-account switch also marks and mutes live SSH remote panes
— confirmed pre-existing on main by two independent QA runs — tracked separately
in #10992. Related pre-existing defect filed as #10863.
2026-07-27 17:53:41 -07:00
Brennan Benson c3526cc19d feat(codex): surface a stalled config sync instead of failing silently (#10449)
* feat(codex): surface a stalled config sync instead of failing silently

Why: the mirror keeps serving the last synced settings when ~/.codex/config.toml
is missing, blank, or unreadable. That is the right call for data safety, but it
is invisible — a downed WSL distro or an unhydrated cloud-synced home leaves
"Orca ignores my config edits" with no log line and no UI to diagnose.

Status is derived on demand from the same predicates the mirror uses, so the two
cannot disagree. The stall is logged once per episode rather than on every launch
and quota poll, and the Codex account section names the file and what to do.

* fix(codex): latch an unreadable source and stop over-claiming recovery

An unreadable source throws out of the mirror, so reporting only on the success
path left that stall latch-less: it logged the raw failure on every launch and
quota poll while its reason never reached the surfaced status. Report from the
catch path too.

The clear message also claimed the source was "readable again", which is false
when the stall ended because the runtime config was removed rather than because
the source came back.

Restoring console.warn now happens in afterEach — an inline mockRestore is
skipped by a failing assertion, and the leaked spy made every later case in the
block fail spuriously.

* fix(codex): latch the stall promotion hits first, and scope it to the host

Review round 1 findings:

- The unreadable-source latch still never fired in the steady state. Once a
  baseline exists, promotion reads the source before the mirror does, so it
  throws first and `!promotionPlan` returned before any reporting — logging a
  reasonless failure every launch and quota poll, which is exactly what the
  previous commit claimed to fix. Report from that branch too. The test only
  passed because its fixture had no baseline; it now seeds one first and fails
  without the fix.
- The banner named the host's ~/.codex while a WSL or per-account runtime was
  selected, whose real source is a different file entirely. Gate it to the host
  scope, matching how the sign-in warning is already gated.
- Three new translate keys were missing from the locale catalogs, failing the
  localization gate in `pnpm lint`.
- The registrar mock was never asserted, so deleting the registration left the
  suite green.
- `codexConfigSyncStatus` hung off the `agentHooks` namespace despite having
  nothing to do with agent hooks; moved to its own `codexConfigSync.status`
  while it is still a four-file change.

* fix(codex): report sync health for the home the selection actually mirrors

Review round 2:

- The status resolved the shared runtime home, but the system default now runs
  Codex directly against ~/.codex and managed accounts get their own home. So a
  stalled per-account mirror showed no banner at all, while a stale shared home
  could warn about a config the active lane never reads. Resolve the mirrored
  home from the current selection, and report synced when the lane has no mirror
  to fall behind.
- The round-1 report on the promotion failure path could clear the latch on a
  pass where no mirror ran, claiming a recovery that never happened and
  silencing every later pass. Only ever latch a stall there; leave clearing to
  the path that actually mirrored.

* fix(codex): refetch sync status when the active Codex account changes

Review round 3:

- Resolving the status per selection made the fetch account-dependent, but the
  effect was not keyed on the active account. Switching accounts left the banner
  describing the previous one — and switching INTO a stalled account showed
  nothing at all, which is the silence this change exists to remove.
- Pin the home resolution itself: it had no direct test, and its shared-home
  path was a hand-copied literal that could drift from the real helper and
  silence the banner with every other test still green.
- Narrow the handler's dependency to the one method it calls, which also drops
  an `as unknown as` cast from its test.
- Skip the chmod-based test on Windows, where a read-only directory does not
  block writes so the scenario cannot be constructed; matches the convention
  already used in config-settings-promotion.test.ts.

* chore(codex): restore the handler docstring and isolate the resolver suite

Round 4 returned clean; these are its two non-blocking nits.

Narrowing the handler param left its JSDoc stranded above the new type, so the
function had no hover doc. The resolver suite also read the developer's real
CODEX_HOME and shell rc, so anyone exporting one would see it fail locally.
2026-07-24 18:30:54 -07:00
kin001andBrennan Benson 6a72c8f120 fix(codex): preserve runtime config when system source is missing (#9127)
* fix(codex): preserve runtime config without system source

* fix(codex): retain baseline when mirror is skipped

* refactor(codex): extract deprecated hook-flag normalization

Why: codex-config-mirror.ts sat at the 300-line cap, so the missing-source
guard could not land without a max-lines disable.

* fix(codex): bootstrap a baseline when the mirror is skipped

Why: a runtime home seeded outside the mirror (WSL, per-account) never got a
baseline while the source was missing, so promotion stayed inert and silently
reverted the in-Codex change once the source returned.

* fix(codex): stop a synthesized source config from wiping runtime settings

Two routes still reached the #9073 data loss after the missing-source guard:

- Promotion runs before the guard and, with no ~/.codex/config.toml, created
  one holding only the promoted keys. The next mirror treated that skeleton as
  authoritative and deleted every other runtime setting. It needs no missing
  file: `codex mcp add` inside an Orca-launched Codex plus /model was enough to
  drop the MCP server for good. Promotion now seeds a brand-new system config
  from the runtime's ordinary settings, so the mirror round-trips them.
- A 0-byte source (half-written, or an unhydrated cloud-synced home) still read
  as an authoritative empty config and advanced the baseline, making the loss
  unrecoverable. A blank source is now treated like a missing one.

Moves the TOML section model out of codex-config-mirror.ts so promotion can
share it without a cycle.

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-07-24 13:38:10 -07:00
NeilandOrca aab112933e Revert "fix(memory): bound OOM-prone accumulators (#10179)" (#10255)
Co-authored-by: Orca <help@stably.ai>
2026-07-23 18:35:31 -07:00
Brennan Benson ef985ed800 fix(codex): re-land TUI settings promotion with anchored baseline (#10213)
* Reapply "Preserve Codex [tui] settings across managed CODEX_HOME remirrors (#9475)" (#10085)

This reverts commit 09756dfaff.

* fix(codex): anchor settings baseline migration
2026-07-23 16:08:46 -07:00