diff --git a/.github/workflows/pr.yml b/.github/workflows/pr.yml
index ed37cefb434..ca1651325c9 100644
--- a/.github/workflows/pr.yml
+++ b/.github/workflows/pr.yml
@@ -816,6 +816,7 @@ jobs:
src/main/cli/wsl-cli-powershell-boundary.test.ts
src/main/cursor/hook-service.test.ts
src/main/orca-profiles/profile-index-store.test.ts
+ src/main/startup/windows-install-dir-acl-repair.win32.test.ts
src/main/runtime/repo-worktree-admin-fingerprint.test.ts
src/main/runtime/worktree-scan-admin-fingerprint-gate.test.ts
src/shared/secure-file-fsync-flags.test.ts
diff --git a/AGENTS.md b/AGENTS.md
index 6915b246bcc..8b0156ba6b1 100644
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -53,18 +53,6 @@ Orca targets macOS, Linux, and Windows. Keep all platform-dependent behavior beh
- **WSL commands**: build argv with `buildWslExecArgs` (always `--exec` — under `--`, `wsl.exe` expands `$name` in every argument and silently rewrites the script), and fence anything whose stdout you parse with `buildWslCapturedLoginShellCommand`, because the interactive login shell prints the distro banner to stdout. See [`docs/reference/wsl-command-execution.md`](./docs/reference/wsl-command-execution.md).
- **Linux native modules**: keep the glibc floor at Ubuntu 20.04 / glibc 2.31. A module compiled from source on a newer runner can reference symbol versions absent on the floor and crash the app on startup. See [`docs/reference/linux-glibc-compatibility.md`](./docs/reference/linux-glibc-compatibility.md); packaging fails if a bundled native binary needs newer glibc.
-## Localization (i18n)
-
-All user-facing copy is localized. `src/renderer/src/i18n/locales/en.json` is the source of truth; the `zh`, `ja`, `ko`, and `es` catalogs mirror its keys. Strings reach the UI through `translate('auto.', 'English fallback')` — never hardcode display text.
-
-When you touch user-facing copy, keep all five catalogs in sync:
-
-- **New strings** — wrap them in `translate(...)` with an English fallback, run `pnpm sync:localization-catalog` to register the keys in `en.json` and add placeholders to every other locale, then `pnpm bootstrap:-catalog` (e.g. `bootstrap:ja-catalog`) to translate the placeholders.
-- **Reworded strings** — changing the value of an existing key updates only `en.json`. The other locales keep the key with its old translation, and **the lint checks will not catch this**: `verify:localization-catalog` enforces key _parity_, not translation _freshness_. Update the same key in `zh/ja/ko/es` by hand, or re-translate it via `bootstrap:-catalog`.
-- **Removed strings** — delete the key from _every_ locale; the parity check rejects a key that exists in one catalog but not another.
-
-Before pushing copy changes, run `pnpm verify:localization-catalog` and `pnpm verify:localization-coverage` (both also run in `pnpm lint`).
-
## SSH Use Case
All changes must consider the SSH use case. Don't assume local-only execution. Before changing anything that reports on, stops, or lists remote work, follow [`docs/reference/ssh-execution-boundary.md`](./docs/reference/ssh-execution-boundary.md): the execution host owns everything that touches execution, and loss of contact is never evidence of process death — the verdict vocabulary is `live` / `unverifiable` / `exited`, with no synonyms.
diff --git a/README.md b/README.md
index 805c94295cc..2ae59035da8 100644
--- a/README.md
+++ b/README.md
@@ -12,7 +12,7 @@
@@ -238,9 +238,9 @@ Pair with your desktop app to monitor and steer your agents from your phone.
- **Discord:** Join the community on **[Discord](https://discord.gg/fzjDKHxv8Q)**.
- **Twitter / X:** Follow **[@orca_build](https://x.com/orca_build)** for updates and announcements.
-- **WeChat:** Scan to join the Orca community WeChat group 8.
+- **WeChat:** Scan to join the Orca community WeChat group 8. Group 8 may be full; if so, scan the Group 9 QR code instead.
-
+
- **Feedback & Ideas:** We ship fast. Missing something? [Request a new feature](https://github.com/stablyai/orca/issues).
- **Privacy:** See the [privacy & telemetry docs](https://www.onorca.dev/docs/telemetry) for what anonymous usage data Orca collects and how to opt out.
diff --git a/cloud/apps/relay-ops/src/incident-monitor.test.ts b/cloud/apps/relay-ops/src/incident-monitor.test.ts
index ea5ad55b645..4e1da9fab26 100644
--- a/cloud/apps/relay-ops/src/incident-monitor.test.ts
+++ b/cloud/apps/relay-ops/src/incident-monitor.test.ts
@@ -111,27 +111,41 @@ describe('incident monitor evaluator', () => {
})
})
- it('freezes when postgres retries exceed the recalibrated ceiling', () => {
- const sample = healthySample()
- sample.sources['relay-logs']!.signals['relay.postgres_retries'] =
- signal(INCIDENT_MONITOR_THRESHOLDS.relayPostgresRetries + 1)
- expect(evaluateIncidentSample(sample, startedAt)).toMatchObject({
+ // Why: the global relay_cells lock made retries a steady-state rate (24 h p99
+ // 1320/5min on 2026-09-04); the bar fences only unbounded growth beyond that.
+ it('tolerates the measured healthy retry rate and freezes above the bar', () => {
+ const healthy = healthySample()
+ healthy.sources['relay-logs']!.signals['relay.postgres_retries'] = signal(1504)
+ expect(evaluateIncidentSample(healthy, startedAt).status).toBe('green')
+
+ const incident = healthySample()
+ incident.sources['relay-logs']!.signals['relay.postgres_retries'] = signal(2001)
+ expect(evaluateIncidentSample(incident, startedAt)).toMatchObject({
status: 'freeze',
failures: [
- expect.objectContaining({ signal: 'relay.postgres_retries', threshold: 300 })
+ expect.objectContaining({ signal: 'relay.postgres_retries', threshold: 2000 })
]
})
})
- // Why: sweeps no longer reach the retry wrapper, so any exhaustion left in this
- // counter is a request path that terminally failed. It must still freeze.
- it('freezes on a single exhausted request-path transaction', () => {
- const sample = healthySample()
- sample.sources['relay-logs']!.signals['relay.postgres_retry_exhausted'] = signal(1)
- expect(evaluateIncidentSample(sample, startedAt)).toMatchObject({
+ // Why: since #18521 the request path fails fast on the cell-inventory lock, so
+ // exhaustion is a steady contention rate (post-#18521 p90 147/5min, max 220),
+ // not an anomaly. The bar bounds it below the 2026-08-23 incident peak of 467.
+ it('tolerates the measured healthy exhaustion rate and freezes above the bar', () => {
+ const healthy = healthySample()
+ healthy.sources['relay-logs']!.signals['relay.postgres_retry_exhausted'] = signal(220)
+ expect(evaluateIncidentSample(healthy, startedAt).status).toBe('green')
+
+ const atLimit = healthySample()
+ atLimit.sources['relay-logs']!.signals['relay.postgres_retry_exhausted'] = signal(300)
+ expect(evaluateIncidentSample(atLimit, startedAt).status).toBe('green')
+
+ const incident = healthySample()
+ incident.sources['relay-logs']!.signals['relay.postgres_retry_exhausted'] = signal(301)
+ expect(evaluateIncidentSample(incident, startedAt)).toMatchObject({
status: 'freeze',
failures: [
- expect.objectContaining({ signal: 'relay.postgres_retry_exhausted', threshold: 0 })
+ expect.objectContaining({ signal: 'relay.postgres_retry_exhausted', threshold: 300 })
]
})
})
diff --git a/cloud/apps/relay-ops/src/incident-monitor.ts b/cloud/apps/relay-ops/src/incident-monitor.ts
index a1936f5df6c..a121568d918 100644
--- a/cloud/apps/relay-ops/src/incident-monitor.ts
+++ b/cloud/apps/relay-ops/src/incident-monitor.ts
@@ -15,8 +15,8 @@ export const INCIDENT_MONITOR_THRESHOLDS = {
// Why: healthy latest-sum backends idle near 100 but spike to 216 in 1-minute
// bursts (~10 min/day exceeded the old bar of 160 on 2026-08-26, freezing a
// pre-drain gate on baseline noise). 250 clears measured healthy peaks while
- // firing well before the verified 400-connection ceiling; pool-wait and
- // exhausted-retry signals keep their strict thresholds.
+ // firing well before the verified 400-connection ceiling; the retry signals
+ // below discriminate incident-class contention.
cloudSqlBackends: 250,
// Bound the observed recovery load; deadlocks remain zero-tolerance.
cloudSqlLockWaits: 20,
@@ -32,27 +32,35 @@ export const INCIDENT_MONITOR_THRESHOLDS = {
relayPoolWaiting: 800,
relayPoolWaitMs: 2_500,
// Why: successful lock retries are the contention machinery working, not harm.
- // Healthy 2026-08-26 baseline bursts to 234/5min (26% of windows crossed the old
- // bar of 20, set unmeasured at the monitor's 2026-07-28 birth); the 2026-08-23
- // incident ran ~2,200-3,000/5min. 300 clears healthy bursts with ~10x incident
- // margin; relayPostgresRetryExhausted below stays at zero tolerance, so any
- // transaction that terminally fails still freezes the gate.
- relayPostgresRetries: 300,
- // Why: this bar stays at zero. Cell-inventory contention reaches the retry
- // wrapper from exactly two kinds of caller, and neither is a sweep tick that
- // can shrug the failure off:
- // - request paths, which take a wait bounded at CELL_INVENTORY_LOCK_TIMEOUT_MS
- // (assignment, control activation, activity, admin drain/evacuate/supersede);
- // - sweep-reachable code that a request also enters, which keeps the pool
- // lock_timeout so it cannot fail faster than before this change: the
- // completeEvacuation site that waits, reconcileReservationAccounting, and
- // placement re-entered from evacuateDeadCells.
- // Sweep-only sites take the inventory NOWAIT, so their contention becomes
- // database_lock_unavailable, which is not a retryable abort and never reaches
- // this counter. The relay's cell-inventory-lock census test holds that split.
- // Splitting the metric by the phase label PR #423 put on the log payload would
- // need a labelled log-based metric, which this signal's counter does not carry.
- relayPostgresRetryExhausted: 0,
+ // Recalibrated 2026-09-04 from 300, which was set 2026-08-26 when healthy bursts
+ // reached 234/5min. The global relay_cells FOR UPDATE lock has since become the
+ // fleet's steady state: measured fleet-wide (director + cells, summed per five
+ // minutes) 2026-09-03T05Z..2026-09-04T05Z p50 430 / p90 924 / p99 1320 / max
+ // 1504, with 55% of windows over 300 and only 22% of 15-minute gates clean, so
+ // the bar blocked the very cell roll that carries the 500 ms lock wait (#18521)
+ // and the beginProof crash guard to the cells. The 2026-08-23 lock incident on
+ // this same metric peaked at 1510 in one window and 646 in the next, so it is
+ // not separable from today's contention by retries alone; it is caught by
+ // relayPostgresRetryExhausted (467 at the peak vs a 300 bar), director
+ // concurrency, and the pool bars. 2000 passes every healthy 15-minute window
+ // measured in the last 24 h and still fences unbounded growth. Re-tighten once
+ // the fleet is on the 500 ms lock wait and the baseline is re-measured.
+ relayPostgresRetries: 2000,
+ // Why: 300 per five minutes, recalibrated 2026-09-04 from a bar of zero that no
+ // production window has cleared since #18521 shipped to the director. That
+ // change cut the request-path cell-inventory wait from the 1 s pool lock_timeout
+ // to 500 ms, so a contended waiter now fails fast (one /v1/assign 503 with
+ // Retry-After, which the client retries) instead of succeeding slowly, and the
+ // exhaustion count became a steady-state contention rate rather than an
+ // anomaly. Measured fleet-wide (director + cells) per five minutes over
+ // 2026-09-03T03Z..2026-09-04T02Z: every one of 236 windows was non-zero;
+ // quiet hours p50 2 / max 36; pre-#18521 daytime p50 10 / p90 25 / max 87;
+ // post-#18521 p50 42 / p90 147 / max 220. The 2026-08-23 lock incident peaked
+ // at 467. 300 clears every measured healthy window and still sits below the
+ // incident shape; retries above fence only unbounded growth.
+ // User-facing /v1/assign 503 share did not move with #18521 (13.9% old image
+ // vs 12.3% new, same evening), so exhaustion is not a proxy for user harm.
+ relayPostgresRetryExhausted: 300,
// Why: public admission is a per-instance semaphore, so fleet assignment capacity is
// concurrency x instances. A floor of 1 let the 2026-08-04 collapse from five instances
// to two pass unnoticed, which is the exact failure this monitor exists to catch. Keep in
diff --git a/cloud/apps/relay/src/assignment-connection-headroom-postgres.test.ts b/cloud/apps/relay/src/assignment-connection-headroom-postgres.test.ts
index 6ac9521c3d6..80a74a47eeb 100644
--- a/cloud/apps/relay/src/assignment-connection-headroom-postgres.test.ts
+++ b/cloud/apps/relay/src/assignment-connection-headroom-postgres.test.ts
@@ -44,6 +44,12 @@ describePostgres('PostgreSQL assignment connection headroom', () => {
`DELETE FROM relay_assignments
WHERE user_id LIKE 'connection-headroom-postgres-%'`
)
+ // A snapshot left by an aborted run rejects the replayed watermark
+ // with stale_connection_snapshot.
+ await database.query(
+ `DELETE FROM relay_cell_connection_snapshots WHERE cell_id = ?`,
+ [cell.id]
+ )
await database.query(
`DELETE FROM relay_cell_connection_runtime WHERE cell_id = ?`,
[cell.id]
diff --git a/cloud/apps/relay/src/assignment-control-supersession-postgres.test.ts b/cloud/apps/relay/src/assignment-control-supersession-postgres.test.ts
index 10193b78cc6..cf8819686b5 100644
--- a/cloud/apps/relay/src/assignment-control-supersession-postgres.test.ts
+++ b/cloud/apps/relay/src/assignment-control-supersession-postgres.test.ts
@@ -38,6 +38,10 @@ describePostgres('PostgreSQL control supersession', () => {
[identity.userId]
)
await database.query(`DELETE FROM relay_assignments WHERE user_id = ?`, [identity.userId])
+ // A snapshot left by an aborted run rejects the replayed watermark with stale_connection_snapshot.
+ await database.query(`DELETE FROM relay_cell_connection_snapshots WHERE cell_id = ?`, [
+ cell.id
+ ])
await database.query(`DELETE FROM relay_cell_connection_runtime WHERE cell_id = ?`, [cell.id])
await database.query(`DELETE FROM relay_cell_connection_limits WHERE cell_id = ?`, [cell.id])
await database.query(`DELETE FROM relay_cell_runtime WHERE cell_id = ?`, [cell.id])
diff --git a/cloud/apps/relay/src/assignment-store.ts b/cloud/apps/relay/src/assignment-store.ts
index 96d66eb7dc2..9d240e304a7 100644
--- a/cloud/apps/relay/src/assignment-store.ts
+++ b/cloud/apps/relay/src/assignment-store.ts
@@ -337,8 +337,8 @@ export type CellInventoryLockMode =
// Never queue: the caller handles database_lock_unavailable and moves on.
| 'nowait'
// A sweep can enter here, so keep the pool default. Failing sooner would turn
- // ordinary contention into a 55P03 the retry wrapper reports as terminal, and
- // one terminal failure freezes the incident gate.
+ // ordinary contention into a 55P03 the retry wrapper reports as terminal, which
+ // spends the incident gate's bounded exhausted-retry budget (300 per 5 min).
| 'pool-default'
// Why: stranded detection (issue #225) needs a grant old enough that a real
// attach would have registered (the 90s activity lease covers dial +
@@ -3202,8 +3202,7 @@ export class RelayAssignmentStore {
)
const requestDelta = ACTIVITY_REQUEST_UNITS[kind] * (after - before)
if (requestDelta !== 0) {
- await this.lockCellInventory(transaction, 'request')
- await this.adjustCellReservation(transaction, text(row, 'cell_id'), requestDelta)
+ await this.adjustCellReservationAtomically(transaction, text(row, 'cell_id'), requestDelta)
}
})
})
@@ -3263,9 +3262,12 @@ export class RelayAssignmentStore {
}
const units = ACTIVITY_REQUEST_UNITS[input.kind]
if (existing) {
- await this.lockCellInventory(transaction, 'request')
+ // Why: a client-chosen activity id can move between cells, so lock the
+ // one or two rows this path touches in cell_id order, the same order
+ // placement takes the inventory in, and no cycle can form.
+ await this.lockCellRows(transaction, [text(existing, 'cell_id'), input.cellId])
await this.removeActivityLease(transaction, identity, existing, now)
- await this.adjustCellReservation(transaction, input.cellId, units)
+ await this.adjustCellReservationAtomically(transaction, input.cellId, units)
}
await this.adjustActivityCount(transaction, identity, input.kind, 1, expiresAt, now)
await transaction.query(
@@ -3580,8 +3582,7 @@ export class RelayAssignmentStore {
)
await this.touchAssignment(transaction, identity, expiresAt, now)
} else {
- await this.lockCellInventory(transaction, 'request')
- await this.adjustCellReservation(transaction, input.cellId, 1)
+ await this.adjustCellReservationAtomically(transaction, input.cellId, 1)
await this.adjustActivityCount(transaction, identity, 'control', 1, expiresAt, now)
await transaction.query(
`INSERT INTO relay_assignment_activity_leases
@@ -6954,6 +6955,19 @@ export class RelayAssignmentStore {
return rows
}
+ // Per-connection paths touch one or two cells. Locking exactly those rows,
+ // in the same ascending order the inventory lock uses (ORDER BY fixes the
+ // row-lock order), keeps them off the fleet-wide lock without a cycle.
+ private async lockCellRows(database: RelayDatabase, cellIds: string[]): Promise {
+ const distinct = [...new Set(cellIds)]
+ return await database.queryLocked(
+ `SELECT * FROM relay_cells WHERE cell_id IN (${distinct.map(() => '?').join(', ')})
+ ORDER BY cell_id ASC`,
+ distinct,
+ { lockTimeoutMs: CELL_INVENTORY_LOCK_TIMEOUT_MS }
+ )
+ }
+
private async lockGeneralCellInventory(
database: RelayDatabase,
mode: CellInventoryLockMode
@@ -7590,7 +7604,10 @@ export class RelayAssignmentStore {
) {
throw new Error('activity_lease_shape_mismatch')
}
- const cells = await this.lockCellInventory(database, 'request')
+ // Why: this recomputes one cell's reservation from its leases, so only that
+ // row needs to be held; the 23-row inventory lock here serialised every
+ // desktop control rebind in the fleet behind every other one.
+ const cellRow = (await this.lockCellRows(database, [cellId]))[0]
await database.query(
`DELETE FROM relay_assignment_activity_leases
WHERE user_id = ? AND relay_host_id = ? AND activity_kind = 'control'
@@ -7611,7 +7628,6 @@ export class RelayAssignmentStore {
[cellId]
)
)[0]!
- const cellRow = cells.find((cell) => text(cell, 'cell_id') === cellId)
const cellUnits = integer(cellUnitsRow, 'request_units')
if (!cellRow) throw new Error('assigned_cell_missing')
if (cellUnits > integer(cellRow, 'capacity_requests')) {
diff --git a/cloud/apps/relay/src/cell-inventory-lock-census.test.ts b/cloud/apps/relay/src/cell-inventory-lock-census.test.ts
index b2b5b65684c..a26e15f8e1d 100644
--- a/cloud/apps/relay/src/cell-inventory-lock-census.test.ts
+++ b/cloud/apps/relay/src/cell-inventory-lock-census.test.ts
@@ -25,10 +25,12 @@ const CENSUS: CensusEntry[] = [
{ method: 'assignOnce', mode: 'nowait', reach: 'both' },
{ method: 'assignOnce', mode: 'nowait', reach: 'both' },
{ method: 'refreshDrainMigrationLeasesOnce', mode: 'request', reach: 'request' },
- // Reachable from neither: changeActivity has no production callers, only tests.
- { method: 'changeActivity', mode: 'request', reach: 'orphan' },
- { method: 'acquireActivity', mode: 'request', reach: 'request' },
- { method: 'activateControl', mode: 'request', reach: 'request' },
+ // changeActivity, acquireActivity, activateControl and
+ // removeSupersededSameCellControls no longer take the inventory: they lock
+ // only the one or two cell rows they touch, in cell_id order (lockCellRows),
+ // so they cannot cycle with placement's ordered inventory lock, and the
+ // 23-row lock there had serialised every reconnect in the fleet behind every
+ // other one.
{ method: 'startEvacuation', mode: 'request', reach: 'request' },
{ method: 'completeEvacuationFromDeadSourceOnce', mode: 'request', reach: 'request' },
{ method: 'completeEvacuationFromDeadSourceOnce', mode: 'nowait', reach: 'request' },
@@ -48,8 +50,31 @@ const CENSUS: CensusEntry[] = [
{ method: 'releaseExpiredActivityLeases', mode: 'nowait', reach: 'sweep' },
{ method: 'releaseExpiredActivity', mode: 'nowait', reach: 'sweep' },
{ method: 'reconcileReservationAccounting', mode: 'pool-default', reach: 'both' },
- { method: 'leastLoadedCell', mode: 'pool-default', reach: 'both' },
- { method: 'removeSupersededSameCellControls', mode: 'request', reach: 'request' }
+ { method: 'leastLoadedCell', mode: 'pool-default', reach: 'both' }
+]
+
+// Every inline `FROM relay_cells ... FOR UPDATE` outside the named lock helpers,
+// in source order: whole-table locks in reconciliation and sticky placement,
+// and single-row locks for a cell the method is already scoped to (heartbeat,
+// fence, drain generation, configuration, or a reservation adjust that runs
+// under a lock its caller already holds). A new inline lock fails the census
+// below until it is listed here; per-connection paths that touch more than one
+// cell go through lockCellRows so the order is fixed.
+const NAMED_LOCK_HELPERS = ['lockCellInventory', 'lockGeneralCellInventory', 'lockCellRows']
+
+const INLINE_CELL_LOCK_SITES = [
+ 'reconcileCellsWithOptions',
+ 'assignStickyOnce',
+ 'recordCellHeartbeat',
+ 'attestCellFence',
+ 'adoptLegacyCellFence',
+ 'commitLegacyCellFenceAdoption',
+ 'prepareCellFenceAttempt',
+ 'attestCellFenceAttempt',
+ 'attestCellFenceAttempt',
+ 'configureCell',
+ 'assertDrainCellGeneration',
+ 'adjustCellReservation'
]
// The background sweeps, and nothing else. A method reachable from one of these
@@ -151,6 +176,42 @@ describe('cell inventory lock call-site census', () => {
)
})
+ // Why: the census only sees lockCellInventory calls, so a hand-written
+ // `relay_cells ... FOR UPDATE` would escape classification entirely.
+ it('routes every relay_cells row lock through a named lock helper', () => {
+ const lines = storeSource()
+ const rawSites: string[] = []
+ // Whole statements, not a fixed window: a wide column list or a raw
+ // FOR UPDATE inside query() must not slip past.
+ const source = lines.join('\n')
+ const bounds: { name: string; start: number }[] = []
+ lines.forEach((line, index) => {
+ const declaration = DECLARATION.exec(line)
+ if (declaration) bounds.push({ name: declaration[1]!, start: index })
+ })
+ const methodAt = (offset: number): string => {
+ const lineIndex = source.slice(0, offset).split('\n').length - 1
+ let name = ''
+ for (const bound of bounds) if (bound.start <= lineIndex) name = bound.name
+ return name
+ }
+ const tick = String.fromCharCode(96)
+ const statementCall = new RegExp(
+ '\\.(queryLocked|query)\\(\\s*' + tick + '([^' + tick + ']*)' + tick,
+ 'g'
+ )
+ for (const call of source.matchAll(statementCall)) {
+ const statement = call[2]!
+ if (!/\bFROM\s+relay_cells\b/.test(statement)) continue
+ const locks = call[1] === 'queryLocked' || /\bFOR\s+UPDATE\b/.test(statement)
+ if (!locks) continue
+ const method = methodAt(call.index)
+ if (NAMED_LOCK_HELPERS.includes(method)) continue
+ rawSites.push(method)
+ }
+ expect(rawSites).toEqual(INLINE_CELL_LOCK_SITES)
+ })
+
it('leaves no call site taking the inventory without naming a mode', () => {
const source = readFileSync(new URL('./assignment-store.ts', import.meta.url), 'utf8')
const unclassified = source
@@ -170,8 +231,8 @@ describe('cell inventory lock call-site census', () => {
})
// Why: this is the whole point of the classification. A shorter wait on a
- // sweep-reachable site turns contention into a terminal transaction failure,
- // and relayPostgresRetryExhausted freezes the incident gate at zero.
+ // sweep-reachable site turns contention into a terminal transaction failure
+ // that counts against the incident gate's relayPostgresRetryExhausted bar.
// Why: the hold distribution is what the 500ms bound will be tuned against, so
// a mode that stops asking for it goes unmeasured in exactly the lane that
// matters. Nothing else in the suite reads the pool-default branch.
diff --git a/cloud/apps/relay/src/cell-inventory-lock-contention.test.ts b/cloud/apps/relay/src/cell-inventory-lock-contention.test.ts
index 783d8a62e8b..23ec001c573 100644
--- a/cloud/apps/relay/src/cell-inventory-lock-contention.test.ts
+++ b/cloud/apps/relay/src/cell-inventory-lock-contention.test.ts
@@ -300,8 +300,8 @@ describe('bounded cell-inventory lock wait', () => {
})
})
-// Why: the incident monitor freezes at zero exhausted transactions. A sweep that
-// steps aside must not spend the retry budget or report a terminal failure.
+// Why: exhausted transactions count against the incident monitor's bounded bar.
+// A sweep that steps aside must not spend the retry budget or report a terminal failure.
describe('sweep lock skips stay off the transaction retry counters', () => {
it('reports neither a retry nor an exhaustion when NOWAIT finds the lock held', async () => {
const database = await openFakePostgres()
diff --git a/cloud/apps/relay/src/control-rebind-inventory-lock-postgres.test.ts b/cloud/apps/relay/src/control-rebind-inventory-lock-postgres.test.ts
new file mode 100644
index 00000000000..e990ac1ed1a
--- /dev/null
+++ b/cloud/apps/relay/src/control-rebind-inventory-lock-postgres.test.ts
@@ -0,0 +1,260 @@
+import { afterAll, beforeAll, describe, expect, it } from 'vitest'
+import { RelayAssignmentStore } from './assignment-store.js'
+import { openRelayDatabase, type RelayDatabase } from './database.js'
+
+const databaseUrl = process.env.ORCA_RELAY_TEST_POSTGRES_URL
+const describePostgres = databaseUrl ? describe : describe.skip
+
+// Three cells: the inventory lock covers more than the rows a move touches, and
+// a high-to-low move exposes any lock taken out of cell_id order.
+const cells = [
+ {
+ id: 'rebind-inventory-postgres-a',
+ url: 'https://rebind-inventory-postgres-a.example.com',
+ capacityRequests: 1_000,
+ connectionHardCap: 600 as const,
+ connectionUnobservedBound: 50
+ },
+ {
+ id: 'rebind-inventory-postgres-b',
+ url: 'https://rebind-inventory-postgres-b.example.com',
+ capacityRequests: 1_000,
+ connectionHardCap: 600 as const,
+ connectionUnobservedBound: 50
+ },
+ {
+ id: 'rebind-inventory-postgres-c',
+ url: 'https://rebind-inventory-postgres-c.example.com',
+ capacityRequests: 1_000,
+ connectionHardCap: 600 as const,
+ connectionUnobservedBound: 50
+ }
+]
+const identity = { userId: 'rebind-inventory-postgres-user', relayHostId: 'rebindinvhost001' }
+
+function heartbeat(cell: (typeof cells)[number]) {
+ return {
+ cellId: cell.id,
+ cellUrl: cell.url,
+ cellIncarnation: '11111111-1111-4111-8111-111111111111',
+ startedAt: 50,
+ ready: true,
+ observedRequests: 0,
+ totalConnections: 0,
+ inFlightConnections: 0,
+ reservedConnectionUnits: 0,
+ enforcedConnectionUnits: 0,
+ connectionInclusionWatermark: 1,
+ connectionHardCap: 600 as const,
+ connectionUnobservedBound: 50
+ }
+}
+
+// Why: every desktop control rebind used to take the fleet-wide relay_cells
+// FOR UPDATE lock, so a rebind on one cell queued behind whatever held any
+// other cell's row, until COMMIT (55P03 at the request bound). A rebind only
+// touches its own cell row, so it must proceed while another cell's row is
+// held elsewhere.
+describePostgres('PostgreSQL control rebind under a held cell row', () => {
+ const databases: RelayDatabase[] = []
+
+ beforeAll(async () => {
+ databases.push(
+ await openRelayDatabase({ databaseUrl, dataDir: '' }),
+ await openRelayDatabase({ databaseUrl, dataDir: '' })
+ )
+ })
+
+ async function removeTestRows(database: RelayDatabase): Promise {
+ await database.query(
+ `DELETE FROM relay_control_connection_reservations WHERE user_id = ?`,
+ [identity.userId]
+ )
+ for (const table of [
+ 'relay_assignment_activity_leases',
+ 'relay_post_drain_migration_pins',
+ 'relay_assignment_migration_incarnations',
+ 'relay_assignment_migrations',
+ 'relay_assignments'
+ ]) {
+ await database.query(`DELETE FROM ${table} WHERE user_id = ?`, [identity.userId])
+ }
+ for (const cell of cells) {
+ for (const table of [
+ 'relay_cell_connection_snapshots',
+ 'relay_cell_connection_runtime',
+ 'relay_cell_connection_limits',
+ 'relay_cell_runtime',
+ 'relay_cells'
+ ]) {
+ await database.query(`DELETE FROM ${table} WHERE cell_id = ?`, [cell.id])
+ }
+ }
+ }
+
+ afterAll(async () => {
+ if (databases[0]) await removeTestRows(databases[0])
+ for (const connection of databases) await connection.close()
+ })
+
+ it("rebinds and supersedes a control while another cell's row is held", async () => {
+ // A prior aborted run leaves connection snapshots that reject a replayed watermark.
+ await removeTestRows(databases[0]!)
+ const store = new RelayAssignmentStore(databases[0]!, () => 100)
+ await store.reconcileCells(cells)
+ for (const cell of cells) await store.recordCellHeartbeat(heartbeat(cell))
+ // Pin the host to cell A so placement is deterministic.
+ await store.setCellEnabled(cells[1]!.id, false)
+ await store.setCellEnabled(cells[2]!.id, false)
+ const assignment = await store.assign(identity)
+ expect(assignment.cellId).toBe(cells[0]!.id)
+ await store.setCellEnabled(cells[1]!.id, true)
+ await store.setCellEnabled(cells[2]!.id, true)
+ await store.activateControl(identity, {
+ cellId: cells[0]!.id,
+ assignmentEpoch: assignment.assignmentEpoch,
+ generation: 1,
+ connectionInclusionWatermark: 10
+ })
+
+ // Hold only cell B's row on a second connection, the way a rebind on B
+ // does, for longer than the request-path lock bound.
+ let releaseInventory!: () => void
+ const inventoryReleased = new Promise((resolve) => {
+ releaseInventory = resolve
+ })
+ let inventoryHeld!: () => void
+ const inventoryHeldPromise = new Promise((resolve) => {
+ inventoryHeld = resolve
+ })
+ const holder = databases[1]!.transaction(async (transaction) => {
+ await transaction.queryLocked(`SELECT * FROM relay_cells WHERE cell_id = ?`, [cells[1]!.id])
+ inventoryHeld()
+ await inventoryReleased
+ })
+ await inventoryHeldPromise
+
+ // A generation-2 rebind on cell A supersedes generation 1. It must not
+ // wait on cell B's row.
+ const startedAt = Date.now()
+ const blockedStatement = async (): Promise => {
+ const rows = await databases[1]!.query(
+ `SELECT left(query, 160) AS q FROM pg_stat_activity
+ WHERE datname = current_database() AND wait_event_type = 'Lock'`
+ )
+ return rows.map((row) => String(row.q)).join(' | ')
+ }
+ const timeout = new Promise((_, reject) =>
+ setTimeout(
+ () =>
+ void blockedStatement().then((statement) =>
+ reject(new Error(`rebind on cell A blocked behind cell B's row: ${statement}`))
+ ),
+ 2_000
+ )
+ )
+ const rebound = await Promise.race([
+ store.activateControl(identity, {
+ cellId: cells[0]!.id,
+ assignmentEpoch: assignment.assignmentEpoch,
+ generation: 2,
+ connectionInclusionWatermark: 11
+ }),
+ timeout
+ ])
+ const elapsedMs = Date.now() - startedAt
+ releaseInventory()
+ await holder
+
+ expect(rebound).toBe(`control:${cells[0]!.id}:2`)
+ expect(elapsedMs).toBeLessThan(2_000)
+ const controls = await databases[0]!.query(
+ `SELECT activity_id FROM relay_assignment_activity_leases
+ WHERE user_id = ? AND activity_kind = 'control' ORDER BY activity_id`,
+ [identity.userId]
+ )
+ expect(controls).toEqual([{ activity_id: `control:${cells[0]!.id}:2` }])
+ const reserved = await databases[0]!.query(
+ `SELECT reserved_requests FROM relay_cells WHERE cell_id = ?`,
+ [cells[0]!.id]
+ )
+ expect(Number(reserved[0]!.reserved_requests)).toBe(1)
+ }, 15_000)
+
+ // Why: a phone's activity id is client-chosen and can follow the host across
+ // a migration, so acquireActivity may touch two cell rows. Moving from the
+ // higher cell to the lower one is where an unordered lock cycles with
+ // placement's ascending inventory lock (reproduced live before this fix).
+ it('moves an activity from a higher cell to a lower one in cell_id order', async () => {
+ await removeTestRows(databases[0]!)
+ const [cellA, cellB, cellC] = cells as [typeof cells[0], typeof cells[0], typeof cells[0]]
+ const store = new RelayAssignmentStore(databases[0]!, () => 100)
+ await store.reconcileCells(cells)
+ for (const cell of cells) await store.recordCellHeartbeat(heartbeat(cell))
+ await store.setCellEnabled(cellA.id, false)
+ await store.setCellEnabled(cellB.id, false)
+ const assignment = await store.assign(identity)
+ expect(assignment.cellId).toBe(cellC.id)
+ await store.setCellEnabled(cellA.id, true)
+ await store.setCellEnabled(cellB.id, true)
+ const activityId = 'splice:rebind-inventory-postgres'
+ await store.acquireActivity(identity, { activityId, kind: 'splice', cellId: cellC.id })
+ // The migration makes B authoritative; the lease still sits on C.
+ const migration = await store.startEvacuation(identity, cellB.id)
+ expect(migration.targetCellId).toBe(cellB.id)
+
+ // Hold B elsewhere. An ordered move locks B first and queues here holding
+ // nothing else. Locking C first (the old lease's row, as an unordered move
+ // does) or the whole inventory (which takes A) shows up as a held row.
+ let releaseRow!: () => void
+ const rowReleased = new Promise((resolve) => {
+ releaseRow = resolve
+ })
+ let rowHeld!: () => void
+ const rowHeldPromise = new Promise((resolve) => {
+ rowHeld = resolve
+ })
+ const heldWhileMoverWaits: string[] = []
+ const holder = databases[1]!.transaction(async (transaction) => {
+ await transaction.queryLocked(`SELECT * FROM relay_cells WHERE cell_id = ?`, [cellB.id])
+ rowHeld()
+ await rowReleased
+ for (const cell of [cellA, cellC]) {
+ try {
+ await transaction.queryLocked(`SELECT * FROM relay_cells WHERE cell_id = ?`, [cell.id], {
+ failIfUnavailable: true
+ })
+ } catch {
+ heldWhileMoverWaits.push(cell.id)
+ }
+ }
+ })
+ await rowHeldPromise
+ const move = store.acquireActivity(identity, { activityId, kind: 'splice', cellId: cellB.id })
+ let moved = false
+ void move.then(() => {
+ moved = true
+ })
+ await new Promise((resolve) => setTimeout(resolve, 250))
+ expect(moved).toBe(false)
+ releaseRow()
+ await holder
+ await move
+ expect(heldWhileMoverWaits).toEqual([])
+
+ const reservations = await databases[0]!.query(
+ `SELECT cell_id, reserved_requests FROM relay_cells
+ WHERE cell_id IN (?, ?, ?) ORDER BY cell_id ASC`,
+ [cellA.id, cellB.id, cellC.id]
+ )
+ const reserved = reservations.map((row) => [String(row.cell_id), Number(row.reserved_requests)])
+ expect(reserved).toEqual([
+ [cellA.id, 0],
+ // Migration grant plus the moved splice, as in the SQLite origin-scoped
+ // reservation case: the lock change did not alter accounting.
+ [cellB.id, 6],
+ // The sticky grant stays on the source until the migration completes.
+ [cellC.id, 1]
+ ])
+ }, 15_000)
+})
diff --git a/cloud/apps/relay/src/host-close-reason-memory.test.ts b/cloud/apps/relay/src/host-close-reason-memory.test.ts
new file mode 100644
index 00000000000..2985e6f1a1d
--- /dev/null
+++ b/cloud/apps/relay/src/host-close-reason-memory.test.ts
@@ -0,0 +1,82 @@
+import { ASSIGNMENT_LIMITS, RELAY_HOST_CLOSE_REASON } from '@orca-cloud/relay-contract'
+import { describe, expect, it } from 'vitest'
+import { HostCloseReasonMemory } from './host-close-reason-memory.js'
+
+function memoryAt(clock: { now: number }): HostCloseReasonMemory {
+ return new HostCloseReasonMemory(() => clock.now)
+}
+
+describe('HostCloseReasonMemory', () => {
+ it('remembers only reasons it knows', () => {
+ const clock = { now: 1_000 }
+ const memory = memoryAt(clock)
+
+ memory.record('a', RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+ memory.record('b', 'quitting')
+ memory.record('c', Buffer.alloc(0))
+ memory.record('d', undefined)
+
+ expect(memory.read('a')).toBe(RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+ expect(memory.read('b')).toBeNull()
+ expect(memory.read('c')).toBeNull()
+ expect(memory.read('d')).toBeNull()
+ })
+
+ it('accepts the reason as the Buffer a ws close delivers', () => {
+ const clock = { now: 1_000 }
+ const memory = memoryAt(clock)
+
+ memory.record('a', Buffer.from(RELAY_HOST_CLOSE_REASON.SIGNED_OUT))
+
+ expect(memory.read('a')).toBe(RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+ })
+
+ it('expires an entry once its host may have been rebalanced away', () => {
+ const clock = { now: 1_000 }
+ const memory = memoryAt(clock)
+ memory.record('a', RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+
+ clock.now += ASSIGNMENT_LIMITS.dormantTtlMs - 1
+ expect(memory.read('a')).toBe(RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+
+ clock.now += 1
+ expect(memory.read('a')).toBeNull()
+ expect(memory.size()).toBe(0)
+ })
+
+ it('forgets on demand', () => {
+ const clock = { now: 1_000 }
+ const memory = memoryAt(clock)
+ memory.record('a', RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+
+ memory.forget('a')
+
+ expect(memory.read('a')).toBeNull()
+ })
+
+ it('drops the oldest survivors rather than growing without bound', () => {
+ const clock = { now: 1_000 }
+ const memory = memoryAt(clock)
+ for (let index = 0; index < 50_050; index++) {
+ memory.record(`host-${index}`, RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+ }
+
+ expect(memory.size()).toBe(50_000)
+ expect(memory.read('host-0')).toBeNull()
+ expect(memory.read('host-50049')).toBe(RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+ })
+
+ it('re-recording refreshes recency so a live host is not evicted first', () => {
+ const clock = { now: 1_000 }
+ const memory = memoryAt(clock)
+ memory.record('a', RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+ memory.record('b', RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+ memory.record('a', RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+
+ expect([...['a', 'b'].map((key) => memory.read(key))]).toEqual([
+ RELAY_HOST_CLOSE_REASON.SIGNED_OUT,
+ RELAY_HOST_CLOSE_REASON.SIGNED_OUT
+ ])
+ expect(memory.size()).toBe(2)
+ })
+})
diff --git a/cloud/apps/relay/src/host-close-reason-memory.ts b/cloud/apps/relay/src/host-close-reason-memory.ts
new file mode 100644
index 00000000000..ed01aacd666
--- /dev/null
+++ b/cloud/apps/relay/src/host-close-reason-memory.ts
@@ -0,0 +1,72 @@
+import {
+ ASSIGNMENT_LIMITS,
+ relayHostCloseReasonFrom,
+ type RelayHostCloseReason
+} from '@orca-cloud/relay-contract'
+
+// Retention matches the dormant assignment TTL: past it the host may have been
+// rebalanced onto another cell, so this cell is no longer the one a phone asks.
+const RETENTION_MS = ASSIGNMENT_LIMITS.dormantTtlMs
+// A fleet-wide auth outage signs out every host at once; the cap bounds that
+// burst well above any single cell's host count without becoming a leak.
+const MAX_ENTRIES = 50_000
+
+// Why in-memory and not Postgres: a phone reaches the cell its host's assignment
+// row already names, which is the same cell that watched the control socket
+// close. Losing this on a cell restart degrades to the pre-existing generic
+// verdict, so the failure mode is the old behaviour rather than a wrong one.
+export class HostCloseReasonMemory {
+ private readonly entries = new Map()
+
+ constructor(private readonly now: () => number = Date.now) {}
+
+ // Silently ignores anything that is not a known reason, which is every close
+ // from a host that predates the field and every abrupt 1006.
+ record(key: string, reason: unknown): void {
+ const parsed = relayHostCloseReasonFrom(reason)
+ if (!parsed) {
+ return
+ }
+ this.entries.delete(key)
+ this.entries.set(key, { reason: parsed, expiresAt: this.now() + RETENTION_MS })
+ this.evict()
+ }
+
+ forget(key: string): void {
+ this.entries.delete(key)
+ }
+
+ read(key: string): RelayHostCloseReason | null {
+ const entry = this.entries.get(key)
+ if (!entry) {
+ return null
+ }
+ if (entry.expiresAt <= this.now()) {
+ this.entries.delete(key)
+ return null
+ }
+ return entry.reason
+ }
+
+ size(): number {
+ return this.entries.size
+ }
+
+ private evict(): void {
+ const now = this.now()
+ for (const [key, entry] of this.entries) {
+ if (entry.expiresAt > now) {
+ break
+ }
+ this.entries.delete(key)
+ }
+ // Insertion order is recency order (record deletes before setting), so the
+ // head is always the oldest survivor.
+ for (const key of this.entries.keys()) {
+ if (this.entries.size <= MAX_ENTRIES) {
+ break
+ }
+ this.entries.delete(key)
+ }
+ }
+}
diff --git a/cloud/apps/relay/src/host-session-registry.ts b/cloud/apps/relay/src/host-session-registry.ts
index 11b7d1de030..5c53041e7ff 100644
--- a/cloud/apps/relay/src/host-session-registry.ts
+++ b/cloud/apps/relay/src/host-session-registry.ts
@@ -14,7 +14,8 @@ import {
HostHelloSchema,
InviteCreateSchema,
RELAY_PROTOCOL_LIMITS,
- RELAY_CLOSE_CODE
+ RELAY_CLOSE_CODE,
+ type RelayHostCloseReason
} from '@orca-cloud/relay-contract'
import nacl from 'tweetnacl'
import type WebSocket from 'ws'
@@ -25,6 +26,7 @@ import {
RelayCredentialStore,
type CredentialReservation
} from './credential-store.js'
+import { HostCloseReasonMemory } from './host-close-reason-memory.js'
import { relayHostLogDigest } from './relay-host-log-digest.js'
import type { RelayTokenClaims } from './relay-token-verifier.js'
import type { RelayRuntimeObserver } from './relay-observability.js'
@@ -130,6 +132,10 @@ const ACTIVATION_QUEUE_WAIT_MS = 30_000
export class HostSessionRegistry {
private readonly sessions = new Map()
private readonly activationQueues = new Map>()
+ // Why it outlives `sessions`: the orphan grace deletes the session within 30s,
+ // but a signed-out desktop never comes back, so the phone that asks minutes
+ // later would otherwise find nothing to explain its rejection with.
+ private readonly hostCloseReasons = new HostCloseReasonMemory(() => this.now())
private draining = false
constructor(
@@ -175,7 +181,8 @@ export class HostSessionRegistry {
return
}
this.observer.recordAuth(true)
- const session = this.sessions.get(this.key(reservation.userId, hostId))
+ const sessionKey = this.key(reservation.userId, hostId)
+ const session = this.sessions.get(sessionKey)
if (
!session ||
session.state !== 'active' ||
@@ -184,7 +191,13 @@ export class HostSessionRegistry {
) {
capacityReservation?.release()
await this.store.failReservation(reservation)
- this.rejectClient(socket, RELAY_CLOSE_CODE.HOST_OFFLINE)
+ // The only rejection that can name a cause: the host is genuinely absent.
+ // The attach-deadline 4404 below fires while control is still connected.
+ this.rejectClient(
+ socket,
+ RELAY_CLOSE_CODE.HOST_OFFLINE,
+ this.hostCloseReasons.read(sessionKey)
+ )
return
}
if (session.activeConnIds.size + session.pendingConns.size >= 8) {
@@ -793,7 +806,10 @@ export class HostSessionRegistry {
regionalDrainTimer: null,
regionalDrainExpiresAt: null
}
- this.sessions.set(this.key(identity.sub, identity.relayHostId), session)
+ const sessionKey = this.key(identity.sub, identity.relayHostId)
+ // A host that proved itself again is not signed out, whatever it said last.
+ this.hostCloseReasons.forget(sessionKey)
+ this.sessions.set(sessionKey, session)
this.wireActiveControl(session)
this.sendHelloAck(session)
}
@@ -813,6 +829,11 @@ export class HostSessionRegistry {
})
socket.once('close', (code, reason) => {
this.observer.recordControlClose?.(code)
+ // Guarded on identity: a predecessor retired by a rebind must not stamp a
+ // cause onto the live session that replaced it.
+ if (session.socket === socket) {
+ this.hostCloseReasons.record(this.key(session.identity.sub, session.relayHostId), reason)
+ }
// One line per control close makes reconnect churners attributable by
// host digest without exposing the raw relay host id.
console.warn(
@@ -1187,9 +1208,16 @@ export class HostSessionRegistry {
if (session.socket) send(session.socket, 'control-error', { ...(reqId ? { reqId } : {}), code })
}
- private rejectClient(socket: WebSocket, code: number): void {
+ // hostCloseReason rides the WebSocket close reason, never relay-hello: every
+ // shipped phone parses relay-hello with a strict schema that rejects an
+ // unknown key, and none of them read the close reason at all.
+ private rejectClient(
+ socket: WebSocket,
+ code: number,
+ hostCloseReason?: RelayHostCloseReason | null
+ ): void {
send(socket, 'relay-hello', { ok: false, code })
- closeRelayWebSocket(socket, code, 'relay connection rejected')
+ closeRelayWebSocket(socket, code, hostCloseReason ?? 'relay connection rejected')
}
private releaseControlActivity(session: HostSession): void {
diff --git a/cloud/apps/relay/src/host-signed-out-rejection.test.ts b/cloud/apps/relay/src/host-signed-out-rejection.test.ts
new file mode 100644
index 00000000000..0f8542c960f
--- /dev/null
+++ b/cloud/apps/relay/src/host-signed-out-rejection.test.ts
@@ -0,0 +1,206 @@
+import { EventEmitter } from 'node:events'
+import {
+ CONTROL_CONTINUITY_LIMITS,
+ RELAY_CLOSE_CODE,
+ RELAY_HOST_CLOSE_REASON
+} from '@orca-cloud/relay-contract'
+import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
+import type WebSocket from 'ws'
+import type { RelayAssignmentStore } from './assignment-store.js'
+import type { RelayConfig } from './config.js'
+import type { RelayCredentialStore } from './credential-store.js'
+import { HostSessionRegistry } from './host-session-registry.js'
+import type { RelayRuntimeObserver } from './relay-observability.js'
+import type { RelayTokenClaims } from './relay-token-verifier.js'
+import { ProcessQueuedByteBudget } from './splice-forwarder.js'
+
+class FakeSocket extends EventEmitter {
+ readonly OPEN = 1
+ readonly CLOSED = 3
+ readyState = this.OPEN
+ readonly send = vi.fn()
+ readonly close = vi.fn((code?: number, reason?: string) => {
+ this.readyState = this.CLOSED
+ this.emit('close', code, Buffer.from(reason ?? ''))
+ })
+ readonly terminate = vi.fn(() => {
+ this.readyState = this.CLOSED
+ this.emit('close', 1006, Buffer.alloc(0))
+ })
+}
+
+const config = {
+ port: 8080,
+ publicUrl: 'https://relay-c3.example.com',
+ cellUrl: 'https://relay-c3.example.com',
+ authIssuer: 'https://auth.example.com',
+ authAudience: 'orca-relay',
+ jwksUrl: 'https://auth.example.com/jwks',
+ assignmentSigningKey: new Uint8Array(32),
+ role: 'cell',
+ cellId: 'production-gce-c3',
+ cells: []
+} as unknown as RelayConfig
+
+const identity = {
+ sub: 'user-1',
+ prof: 'profile-1',
+ org: 'org-1',
+ relayHostId: 'AbCdEf0123_-xyZ9'
+} as unknown as RelayTokenClaims
+
+const reservation = {
+ userId: identity.sub,
+ relayHostId: identity.relayHostId,
+ credentialKind: 'resume',
+ relayDeviceId: 'device-1',
+ leaseExpiresAt: Date.now() + 60_000
+}
+
+function createRegistry() {
+ const store = {
+ resolveResume: vi.fn().mockResolvedValue({ userId: identity.sub }),
+ reserveCredential: vi.fn().mockResolvedValue(reservation),
+ failReservation: vi.fn().mockResolvedValue(undefined)
+ }
+ const assignments = {
+ activateControl: vi.fn().mockResolvedValue('control:production-gce-c3:1'),
+ markMigrationTargetRegistered: vi.fn().mockResolvedValue(undefined),
+ resolve: vi.fn().mockResolvedValue({ cellId: config.cellId }),
+ acquireActivity: vi.fn().mockResolvedValue(undefined),
+ renewControlActivity: vi.fn().mockResolvedValue(undefined),
+ releaseActivity: vi.fn().mockResolvedValue(true)
+ } as unknown as RelayAssignmentStore
+ const observer = {
+ recordAuth: vi.fn(),
+ recordForwardedBytes: vi.fn(),
+ recordHttp: vi.fn(),
+ recordReconnect: vi.fn(),
+ recordSql: vi.fn(),
+ recordControlClose: vi.fn(),
+ recordSpliceClose: vi.fn()
+ } satisfies RelayRuntimeObserver
+ const registry = new HostSessionRegistry(
+ config,
+ vi.fn(),
+ store as unknown as RelayCredentialStore,
+ assignments,
+ new ProcessQueuedByteBudget(),
+ observer
+ )
+ const activate = (socket: WebSocket, generation: number): Promise =>
+ (
+ registry as unknown as {
+ activate: (
+ socket: WebSocket,
+ identity: RelayTokenClaims,
+ existing: null,
+ generation: number,
+ rebind: boolean,
+ assignmentEpoch: number,
+ appVersion: string
+ ) => Promise
+ }
+ ).activate(socket, identity, null, generation, false, 1, '1.4.173')
+ return { registry, activate }
+}
+
+async function dialPhone(registry: HostSessionRegistry): Promise {
+ const phone = new FakeSocket()
+ await registry.acceptClient(phone as unknown as WebSocket, identity.relayHostId, 'credential')
+ return phone
+}
+
+// The 4404 hello body is unchanged: every shipped phone parses it with a strict
+// schema, so the cause has to ride the close frame instead.
+const HOST_OFFLINE_HELLO = JSON.stringify({
+ type: 'relay-hello',
+ ok: false,
+ code: RELAY_CLOSE_CODE.HOST_OFFLINE
+})
+
+describe('host sign-out reason on phone rejection', () => {
+ beforeEach(() => vi.useFakeTimers())
+ afterEach(() => {
+ vi.clearAllTimers()
+ vi.useRealTimers()
+ })
+
+ it('names the sign-out to a phone that arrives after the host is gone', async () => {
+ const { registry, activate } = createRegistry()
+ const control = new FakeSocket()
+ await activate(control as unknown as WebSocket, 1)
+
+ control.close(1000, RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+ vi.advanceTimersByTime(CONTROL_CONTINUITY_LIMITS.orphanGraceMs + 1)
+
+ const phone = await dialPhone(registry)
+ expect(phone.send).toHaveBeenCalledWith(HOST_OFFLINE_HELLO)
+ expect(phone.close).toHaveBeenCalledWith(
+ RELAY_CLOSE_CODE.HOST_OFFLINE,
+ RELAY_HOST_CLOSE_REASON.SIGNED_OUT
+ )
+ })
+
+ it('says nothing when the host died without naming a cause', async () => {
+ const { registry, activate } = createRegistry()
+ const control = new FakeSocket()
+ await activate(control as unknown as WebSocket, 1)
+
+ control.terminate()
+ vi.advanceTimersByTime(CONTROL_CONTINUITY_LIMITS.orphanGraceMs + 1)
+
+ const phone = await dialPhone(registry)
+ expect(phone.close).toHaveBeenCalledWith(
+ RELAY_CLOSE_CODE.HOST_OFFLINE,
+ 'relay connection rejected'
+ )
+ })
+
+ it('ignores a close reason the host invented', async () => {
+ const { registry, activate } = createRegistry()
+ const control = new FakeSocket()
+ await activate(control as unknown as WebSocket, 1)
+
+ control.close(1000, 'signed-out-ish')
+ vi.advanceTimersByTime(CONTROL_CONTINUITY_LIMITS.orphanGraceMs + 1)
+
+ const phone = await dialPhone(registry)
+ expect(phone.close).toHaveBeenCalledWith(
+ RELAY_CLOSE_CODE.HOST_OFFLINE,
+ 'relay connection rejected'
+ )
+ })
+
+ it('forgets the sign-out once the host proves itself again', async () => {
+ const { registry, activate } = createRegistry()
+ const control = new FakeSocket()
+ await activate(control as unknown as WebSocket, 1)
+ control.close(1000, RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+ vi.advanceTimersByTime(CONTROL_CONTINUITY_LIMITS.orphanGraceMs + 1)
+
+ const reconnected = new FakeSocket()
+ await activate(reconnected as unknown as WebSocket, 2)
+ // Drop it abruptly, as a network death would, so only the stale memory
+ // could still name a cause.
+ reconnected.terminate()
+ vi.advanceTimersByTime(CONTROL_CONTINUITY_LIMITS.orphanGraceMs + 1)
+
+ const phone = await dialPhone(registry)
+ expect(phone.close).toHaveBeenCalledWith(
+ RELAY_CLOSE_CODE.HOST_OFFLINE,
+ 'relay connection rejected'
+ )
+ })
+
+ // A live host is present: the 4404 there is an attach deadline, not absence.
+ it('never names a cause while the host control is connected', async () => {
+ const { registry, activate } = createRegistry()
+ const control = new FakeSocket()
+ await activate(control as unknown as WebSocket, 1)
+
+ const phone = await dialPhone(registry)
+ expect(phone.close).not.toHaveBeenCalled()
+ expect(control.send).toHaveBeenCalledWith(expect.stringContaining('"type":"conn-open"'))
+ })
+})
diff --git a/cloud/dev/fixtures/terraform-root-partition/families.json b/cloud/dev/fixtures/terraform-root-partition/families.json
index 9664c1eb299..3c57b2ea7cb 100644
--- a/cloud/dev/fixtures/terraform-root-partition/families.json
+++ b/cloud/dev/fixtures/terraform-root-partition/families.json
@@ -133,7 +133,10 @@
"google_logging_metric.relay_snapshot",
"google_monitoring_alert_policy.relay_assignment_5xx",
"google_monitoring_alert_policy.relay_assignment_edge_429",
+ "google_monitoring_alert_policy.relay_cloud_nat_port_drops",
"google_monitoring_alert_policy.relay_cloud_sql_backends",
+ "google_monitoring_alert_policy.relay_cloud_sql_checkpoint_loop",
+ "google_monitoring_alert_policy.relay_cloud_sql_disk",
"google_monitoring_alert_policy.relay_custom",
"google_monitoring_alert_policy.relay_gce_connection_headroom",
"google_monitoring_alert_policy.relay_postgres_retry_exhausted",
diff --git a/cloud/docs/relay-incident-monitor.md b/cloud/docs/relay-incident-monitor.md
index 3e5fc04836e..870c95dd413 100644
--- a/cloud/docs/relay-incident-monitor.md
+++ b/cloud/docs/relay-incident-monitor.md
@@ -99,8 +99,8 @@ durably marked consumed before mutation and cannot authorize another run.
| Cloud SQL deadlocks | over 0 |
| Relay pool waiters | over 800 |
| Relay pool wait | over 2,500 ms |
-| PostgreSQL retries in five minutes | over 300 |
-| Exhausted PostgreSQL retries | over 0 |
+| PostgreSQL retries in five minutes | over 2,000 |
+| Exhausted PostgreSQL retries in five minutes | over 300 |
| Director instances | outside 5–6 |
| Director CPU or memory | over 80% |
| Director concurrency | over 64 |
@@ -132,15 +132,41 @@ heartbeats, and matching live admission.
10 minutes over the old bar of 160 — enough to freeze roughly one in ten
15-minute pre-drain gates on baseline noise. 250 clears measured healthy
peaks and still fires well before the verified 400-connection ceiling;
- pool waiters, pool wait latency, and exhausted retries keep their strict
- thresholds.
+ pool waiters and pool wait latency keep their strict thresholds.
- Recalibrated the PostgreSQL-retry freeze from 20 to 300 per five minutes
(2026-08-26). Basis, measured from
`jsonPayload.event="orca_relay_postgres_transaction_retry"` in production
logs: healthy-day bursts reach 234/5min with zero exhausted retries and 26%
of five-minute windows over 20, while the 2026-08-23 lock-contention
- incident ran roughly 2,200–3,000/5min. Exhausted retries stay at zero
- tolerance.
+ incident ran roughly 2,200–3,000/5min by raw log-line count (the gate's
+ own `orca_relay_postgres_retries` metric read 1,510 for that window; see the
+ 2026-09-04 entry).
+- Recalibrated the PostgreSQL-retry freeze from 300 to 2,000 per five minutes
+ (2026-09-04). Basis: the global `relay_cells FOR UPDATE` lock made
+ successful retries a steady-state rate. Measured fleet-wide (director +
+ cells, summed per five minutes from the `orca_relay_postgres_retries`
+ log metric) over 2026-09-03T05Z..2026-09-04T05Z: p50 430 / p90 924 /
+ p99 1,320 / max 1,504; 55% of windows over 300; only 22% of 15-minute gates
+ clean at 300 versus 100% at 2,000. Three read-only dry-runs on 2026-09-04
+ froze on this bar (runs 33836470590, 33838698725) or on a genuine six-cell
+ crash storm (33837160275), blocking the same-cap roll that carries #18521
+ and the `beginProof` crash guard to the 23 cells. The 2026-08-23 incident
+ on this metric peaked at 1,510 then 646, so retries alone no longer
+ separate it from today's baseline; the exhausted-retry bar (incident peak
+ 467 vs bar 300), director concurrency, and the pool bars carry that role.
+ Re-tighten after the fleet is on the 500 ms lock wait.
+- Recalibrated the exhausted-PostgreSQL-retry freeze from 0 to 300 per five
+ minutes (2026-09-04). Basis: #18521 cut the request-path cell-inventory
+ lock wait from the 1 s pool `lock_timeout` to 500 ms, so contended waiters
+ now fail fast (one `/v1/assign` 503 with `Retry-After`) instead of
+ succeeding slowly, and `orca_relay_postgres_transaction_exhausted` became
+ a steady contention rate. Measured fleet-wide per five minutes over
+ 2026-09-03T03Z..2026-09-04T02Z: 236 of 236 windows non-zero; quiet hours
+ p50 2 / max 36; pre-#18521 daytime p50 10 / p90 25 / max 87; post-#18521
+ p50 42 / p90 147 / max 220; the 2026-08-23 incident peaked at 467. Every
+ pre-drain dry-run since the director deploy froze at minute one on this
+ bar, which blocked the cell roll that carries the same fix to the 23 GCE
+ cells. `/v1/assign` 503 share was unchanged by #18521 (13.9% vs 12.3%).
- Added a fail-closed state machine with latched threshold freezes,
generation-scoped checkpoint boundaries, continuity-reset evidence, cadence
accounting, restart-gap recovery, and the 15-minute pre-drain gate.
diff --git a/cloud/infra/terraform/relay-gce-foundation.tf b/cloud/infra/terraform/relay-gce-foundation.tf
index aab3b4579eb..a8d64b3fcea 100644
--- a/cloud/infra/terraform/relay-gce-foundation.tf
+++ b/cloud/infra/terraform/relay-gce-foundation.tf
@@ -42,6 +42,12 @@ resource "google_compute_router_nat" "relay_gce" {
router = google_compute_router.relay_gce[0].name
nat_ip_allocate_option = "AUTO_ONLY"
source_subnetwork_ip_ranges_to_nat = "LIST_OF_SUBNETWORKS"
+ # Cells reach Cloud SQL's public IP through this NAT. The static default of 64 ports per VM
+ # filled during the 2026-09-04 incident and every cell's proxy dial timed out at once.
+ enable_dynamic_port_allocation = true
+ enable_endpoint_independent_mapping = false
+ min_ports_per_vm = 64
+ max_ports_per_vm = 4096
subnetwork {
name = google_compute_subnetwork.relay_gce[0].id
@@ -85,6 +91,12 @@ resource "google_compute_router_nat" "relay_gce_additional" {
router = google_compute_router.relay_gce_additional[each.key].name
nat_ip_allocate_option = "AUTO_ONLY"
source_subnetwork_ip_ranges_to_nat = "LIST_OF_SUBNETWORKS"
+ # Cells reach Cloud SQL's public IP through this NAT. The static default of 64 ports per VM
+ # filled during the 2026-09-04 incident and every cell's proxy dial timed out at once.
+ enable_dynamic_port_allocation = true
+ enable_endpoint_independent_mapping = false
+ min_ports_per_vm = 64
+ max_ports_per_vm = 4096
subnetwork {
name = google_compute_subnetwork.relay_gce_additional[each.key].id
diff --git a/cloud/infra/terraform/relay-observability.tf b/cloud/infra/terraform/relay-observability.tf
index a8fa4276eb2..5467555ccba 100644
--- a/cloud/infra/terraform/relay-observability.tf
+++ b/cloud/infra/terraform/relay-observability.tf
@@ -37,6 +37,10 @@ locals {
description = "Relay PostgreSQL transactions that exhausted bounded retry."
filter = "((resource.type=\"cloud_run_revision\" AND (${local.relay_service_log_filter})) OR resource.type=\"gce_instance\") AND jsonPayload.event=\"orca_relay_postgres_transaction_exhausted\""
}
+ cloud_sql_wal_checkpoint = {
+ description = "Cloud SQL checkpoints triggered by WAL volume instead of the timed schedule; a sustained run is the fsync loop that stalled every relay process at once on 2026-09-04."
+ filter = "resource.type=\"cloudsql_database\" AND resource.labels.database_id=\"${var.project_id}:${local.relay_database_instance_name}\" AND textPayload:\"checkpoint starting: wal\""
+ }
}
relay_runtime_metrics = {
@@ -523,3 +527,109 @@ resource "google_monitoring_alert_policy" "relay_cloud_sql_backends" {
mime_type = "text/markdown"
}
}
+
+resource "google_monitoring_alert_policy" "relay_cloud_sql_checkpoint_loop" {
+ project = var.project_id
+ display_name = "Orca Relay: Cloud SQL checkpoint loop"
+ combiner = "OR"
+ enabled = true
+ notification_channels = var.relay_alert_notification_channels
+
+ conditions {
+ display_name = "WAL-triggered checkpoints above 3 in 5 minutes"
+
+ condition_threshold {
+ filter = "resource.type=\"cloudsql_database\" AND metric.type=\"logging.googleapis.com/user/orca_relay_cloud_sql_wal_checkpoint\""
+ comparison = "COMPARISON_GT"
+ threshold_value = 3
+ duration = "300s"
+
+ aggregations {
+ alignment_period = "300s"
+ per_series_aligner = "ALIGN_SUM"
+ cross_series_reducer = "REDUCE_SUM"
+ }
+
+ trigger {
+ count = 1
+ }
+ }
+ }
+
+ documentation {
+ content = "Healthy operation is one timed checkpoint every 5 minutes. Repeated `checkpoint starting: wal` lines mean WAL is outrunning `max_wal_size` and every checkpoint fsync stalls all relay SQL for seconds. Check `checkpoint complete` sync= times and disk write throughput against the PD-SSD ceiling; the fix is disk size and `max_wal_size` in the Terraform root that owns the instance (orca-cloud `infra/terraform-foundation`)."
+ mime_type = "text/markdown"
+ }
+
+ depends_on = [google_logging_metric.relay_incident]
+}
+
+resource "google_monitoring_alert_policy" "relay_cloud_sql_disk" {
+ project = var.project_id
+ display_name = "Orca Relay: Cloud SQL disk utilization"
+ combiner = "OR"
+ enabled = true
+ notification_channels = var.relay_alert_notification_channels
+
+ conditions {
+ display_name = "Cloud SQL disk above 70%"
+
+ condition_threshold {
+ filter = "resource.type=\"cloudsql_database\" AND resource.label.\"database_id\"=\"${var.project_id}:${local.relay_database_instance_name}\" AND metric.type=\"cloudsql.googleapis.com/database/disk/utilization\""
+ comparison = "COMPARISON_GT"
+ threshold_value = 0.7
+ duration = "600s"
+
+ aggregations {
+ alignment_period = "300s"
+ per_series_aligner = "ALIGN_MAX"
+ }
+
+ trigger {
+ count = 1
+ }
+ }
+ }
+
+ documentation {
+ content = "The shared auth/relay Cloud SQL disk is filling. `refresh_tokens` is the largest table and grows without pruning; grow the disk (IOPS scale with size) before it reaches the WAL checkpoint loop, and prune revoked token rows."
+ mime_type = "text/markdown"
+ }
+}
+
+resource "google_monitoring_alert_policy" "relay_cloud_nat_port_drops" {
+ count = local.relay_gce_configured ? 1 : 0
+
+ project = var.project_id
+ display_name = "Orca Relay: Cloud NAT port exhaustion"
+ combiner = "OR"
+ enabled = true
+ notification_channels = var.relay_alert_notification_channels
+
+ conditions {
+ display_name = "NAT packets dropped for lack of ports"
+
+ condition_threshold {
+ filter = "resource.type=\"nat_gateway\" AND resource.label.\"gateway_name\"=monitoring.regex.full_match(\"${local.relay_gce_name}(-.*)?\") AND metric.type=\"router.googleapis.com/nat/dropped_sent_packets_count\" AND metric.label.\"reason\"=\"OUT_OF_RESOURCES\""
+ comparison = "COMPARISON_GT"
+ threshold_value = 0
+ duration = "120s"
+
+ aggregations {
+ alignment_period = "60s"
+ per_series_aligner = "ALIGN_SUM"
+ cross_series_reducer = "REDUCE_SUM"
+ group_by_fields = ["resource.label.\"gateway_name\""]
+ }
+
+ trigger {
+ count = 1
+ }
+ }
+ }
+
+ documentation {
+ content = "Relay cells reach Cloud SQL's public IP through this NAT. Port exhaustion makes every cell's Cloud SQL Auth Proxy dial time out at once, which reads as a fleet-wide SQL stall with a healthy database. Check `nat/port_usage` per VM and raise `max_ports_per_vm` in `relay-gce-foundation.tf`, or move the database to a private IP."
+ mime_type = "text/markdown"
+ }
+}
diff --git a/cloud/packages/relay-contract/src/host-close-reason.ts b/cloud/packages/relay-contract/src/host-close-reason.ts
new file mode 100644
index 00000000000..3a5abde3f00
--- /dev/null
+++ b/cloud/packages/relay-contract/src/host-close-reason.ts
@@ -0,0 +1,18 @@
+// Mirror of src/shared/relay-host-close-reason.ts in the Orca app repo half.
+// A host control socket may close with one of these as its WebSocket close
+// reason; the cell records it so a later phone rejection can name the cause.
+// Anything else (including the empty reason of an abrupt 1006) means "unknown",
+// which is what every peer that predates this file sends.
+export const RELAY_HOST_CLOSE_REASON = {
+ SIGNED_OUT: 'signed-out'
+} as const
+
+export type RelayHostCloseReason =
+ (typeof RELAY_HOST_CLOSE_REASON)[keyof typeof RELAY_HOST_CLOSE_REASON]
+
+const REASONS: readonly string[] = Object.values(RELAY_HOST_CLOSE_REASON)
+
+export function relayHostCloseReasonFrom(value: unknown): RelayHostCloseReason | null {
+ const text = typeof value === 'string' ? value : (value?.toString() ?? '')
+ return REASONS.includes(text) ? (text as RelayHostCloseReason) : null
+}
diff --git a/cloud/packages/relay-contract/src/index.ts b/cloud/packages/relay-contract/src/index.ts
index 2a7d7d0feda..aab3b53b5f3 100644
--- a/cloud/packages/relay-contract/src/index.ts
+++ b/cloud/packages/relay-contract/src/index.ts
@@ -5,6 +5,7 @@ export * from './control-messages.js'
export * from './control-continuity.js'
export * from './credential-messages.js'
export * from './director-messages.js'
+export * from './host-close-reason.js'
export * from './host-proof-transcript.js'
export * from './persistence-invariants.js'
export * from './protocol-limits.js'
diff --git a/config/electron-builder.config.cjs b/config/electron-builder.config.cjs
index 06d41bad344..ebf4d275678 100644
--- a/config/electron-builder.config.cjs
+++ b/config/electron-builder.config.cjs
@@ -90,7 +90,19 @@ const bundledPluginResources = {
// from package directories where pnpm's symlink farm is absent. Copy the exact
// runtime dependency closure to Resources/node_modules so bare require() calls
// do not fall through to a developer checkout's node_modules.
-const commonExtraResources = [relayExtraResource, bundledPluginResources, skillFreshnessResources]
+// Why the single file rather than the package root: app.asar carries no node_modules, so main's
+// lazy require in deferred-emoji-shortcode-dataset.ts resolves only out of Resources/node_modules,
+// but emojibase-data is 49 MB of locale datasets and worktree naming reads exactly this 166 KB file.
+const emojiShortcodeDatasetResource = {
+ from: 'node_modules/emojibase-data/en/shortcodes/emojibase.json',
+ to: 'node_modules/emojibase-data/en/shortcodes/emojibase.json'
+}
+const commonExtraResources = [
+ relayExtraResource,
+ bundledPluginResources,
+ skillFreshnessResources,
+ emojiShortcodeDatasetResource
+]
// Why: native speech addons must be real files outside app.asar; copy only the
// package matching the artifact target instead of every optional variant.
const macSpeechNativeResource = {
diff --git a/config/patches/node-pty@1.1.0.patch b/config/patches/node-pty@1.1.0.patch
index ff474f7d95e..8f5045b932a 100644
--- a/config/patches/node-pty@1.1.0.patch
+++ b/config/patches/node-pty@1.1.0.patch
@@ -603,7 +603,7 @@ index 7b4b9e1f990fbf95b51528bb56dc9717f5b87532..2ae787c5bd4f3eba470584dc658a01a5
}
#endif
diff --git a/src/win/conpty.cc b/src/win/conpty.cc
-index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a6a4082ce 100644
+index 7b286d3d644c26141df516929703aa6e129df4b2..4aed260dd68e6a171dcfd349e9a7c5c97209248e 100644
--- a/src/win/conpty.cc
+++ b/src/win/conpty.cc
@@ -18,6 +18,7 @@
@@ -614,7 +614,7 @@ index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a
#include
#include
#include
-@@ -44,12 +45,29 @@ struct pty_baton {
+@@ -44,12 +45,39 @@ struct pty_baton {
HANDLE hOut;
HPCON hpc;
@@ -630,22 +630,32 @@ index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a
+ // refused to create or assign one (an outer job without breakaway rights),
+ // in which case callers fall back to their pre-job behaviour.
+ HANDLE hJob = nullptr;
++
++ // Orca: teardown needs BOTH the shell's death and an explicit kill() before
++ // the baton can be freed, so each side records that it has run. Whichever
++ // arrives second frees it. Freeing on the shell's death alone -- what this
++ // file did before -- destroyed the only record of `hpc` while
++ // ClosePseudoConsole was still owed, which is why a self-exiting shell
++ // leaked its pseudoconsole and the console host it reaps (#18601 / F24).
++ bool shellExited = false;
++ bool consoleClosed = false;
pty_baton(int _id, HANDLE _hIn, HANDLE _hOut, HPCON _hpc) : id(_id), hIn(_hIn), hOut(_hOut), hpc(_hpc) {};
};
static std::vector> ptyHandles;
-+// Orca: guards the job accessors below against the exit watcher thread. It does
-+// NOT make the whole table safe -- PtyResize/PtyClear/PtyKill read it unlocked,
-+// as they always have -- but it closes the window this patch opened, where the
-+// watcher can close hShell/hJob and free the baton between a lookup and its use.
++// Orca: guards the job accessors below, and PtyKill, against the exit watcher
++// thread. It does NOT make the whole table safe -- PtyResize and PtyClear still
++// read it unlocked, as they always have -- but it closes the window this patch
++// opened, where the watcher can close hShell/hJob and free the baton between a
++// lookup and its use.
+// Handle VALUES are recycled aggressively, so an unguarded read could pass the
+// shell-pid check against an unrelated process and terminate the wrong job.
+static std::mutex ptyJobMutex;
static volatile LONG ptyCounter;
static pty_baton* get_pty_baton(int id) {
-@@ -102,8 +120,27 @@ void SetupExitCallback(Napi::Env env, Napi::Function cb, pty_baton* baton) {
+@@ -102,8 +130,31 @@ void SetupExitCallback(Napi::Env env, Napi::Function cb, pty_baton* baton) {
// Get process exit code.
GetExitCodeProcess(baton->hShell, (LPDWORD)(&exit_event->exit_code));
// Clean up handles
@@ -665,9 +675,13 @@ index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a
+ // Why inside the lock: erasing frees the baton the job accessors hold a
+ // pointer to. Note remove_pty_baton must not be an assert() argument --
+ // NDEBUG would compile the call away and leak every baton.
-+ const bool removed = remove_pty_baton(baton->id);
-+ assert(removed);
-+ (void)removed;
++ baton->shellExited = true;
++ if (baton->consoleClosed) {
++ const bool removed = remove_pty_baton(baton->id);
++ assert(removed);
++ (void)removed;
++ }
++ // Else PtyKill has not run yet and still owns hpc. It frees the baton.
+ }
+ // Why the lock ends here: BlockingCall below waits on the JS thread, and the
+ // JS thread can be waiting on ptyJobMutex inside PtyTerminateJob. Holding
@@ -675,7 +689,7 @@ index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a
auto status = tsfn.BlockingCall(exit_event, callback); // In main thread
switch (status) {
-@@ -409,6 +446,15 @@ static Napi::Value PtyConnect(const Napi::CallbackInfo& info) {
+@@ -409,6 +460,15 @@ static Napi::Value PtyConnect(const Napi::CallbackInfo& info) {
throw errorWithCode(info, "UpdateProcThreadAttribute failed");
}
@@ -691,7 +705,7 @@ index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a
PROCESS_INFORMATION piClient{};
fSuccess = !!CreateProcessW(
nullptr,
-@@ -416,7 +462,10 @@ static Napi::Value PtyConnect(const Napi::CallbackInfo& info) {
+@@ -416,7 +476,10 @@ static Napi::Value PtyConnect(const Napi::CallbackInfo& info) {
nullptr, // lpProcessAttributes
nullptr, // lpThreadAttributes
false, // bInheritHandles VERY IMPORTANT that this is false
@@ -703,7 +717,7 @@ index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a
envArg, // lpEnvironment
mutableCwd.get(), // lpCurrentDirectory
&siEx.StartupInfo, // lpStartupInfo
-@@ -426,8 +475,47 @@ static Napi::Value PtyConnect(const Napi::CallbackInfo& info) {
+@@ -426,8 +489,47 @@ static Napi::Value PtyConnect(const Napi::CallbackInfo& info) {
throw errorWithCode(info, "Cannot create process");
}
@@ -753,7 +767,7 @@ index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a
if (useConptyDll && fLoadedDll)
{
PFNRELEASEPSEUDOCONSOLE const pfnReleasePseudoConsole = (PFNRELEASEPSEUDOCONSOLE)GetProcAddress(
-@@ -440,6 +528,8 @@ static Napi::Value PtyConnect(const Napi::CallbackInfo& info) {
+@@ -440,6 +542,8 @@ static Napi::Value PtyConnect(const Napi::CallbackInfo& info) {
// Update handle
handle->hShell = piClient.hProcess;
@@ -762,7 +776,91 @@ index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a
// Close the thread handle to avoid resource leak
CloseHandle(piClient.hThread);
-@@ -567,6 +657,143 @@ static Napi::Value PtyKill(const Napi::CallbackInfo& info) {
+@@ -544,29 +648,215 @@ static Napi::Value PtyKill(const Napi::CallbackInfo& info) {
+ int id = info[0].As().Int32Value();
+ const bool useConptyDll = info[1].As().Value();
+
+- const pty_baton* handle = get_pty_baton(id);
++ // Orca: resolve the DLL BEFORE touching any baton state, for the same reason
++ // PtyConnect does it before creating anything. LoadConptyDll throws when
++ // conpty.dll is missing, and a throw after consoleClosed was set would strand
++ // the pseudoconsole permanently: the retry would find the work already
++ // claimed and do nothing. Only the useConptyDll path can throw here; the
++ // other returns kernel32.
++ HANDLE hLibrary = LoadConptyDll(info, useConptyDll);
++ PFNCLOSEPSEUDOCONSOLE pfnClosePseudoConsole = nullptr;
++ if (hLibrary != nullptr) {
++ pfnClosePseudoConsole = (PFNCLOSEPSEUDOCONSOLE)GetProcAddress(
++ (HMODULE)hLibrary,
++ useConptyDll ? "ConptyClosePseudoConsole" : "ClosePseudoConsole");
++ }
+
+- if (handle != nullptr) {
+- HANDLE hLibrary = LoadConptyDll(info, useConptyDll);
+- bool fLoadedDll = hLibrary != nullptr;
+- if (fLoadedDll)
+- {
+- PFNCLOSEPSEUDOCONSOLE const pfnClosePseudoConsole = (PFNCLOSEPSEUDOCONSOLE)GetProcAddress(
+- (HMODULE)hLibrary,
+- useConptyDll ? "ConptyClosePseudoConsole" : "ClosePseudoConsole");
+- if (pfnClosePseudoConsole)
+- {
+- pfnClosePseudoConsole(handle->hpc);
++ // Orca: the baton now outlives the shell, so this runs on a self-exited pty
++ // too -- that is the whole point. Take what we need under the lock: the
++ // watcher thread nulls hShell the moment the shell dies, and TerminateProcess
++ // on a handle it just closed is an invalid-handle operation. Duplicating
++ // rather than reordering keeps upstream's close-then-terminate sequence.
++ HPCON hpc = nullptr;
++ HANDLE hShellDup = nullptr;
++ bool owed = false;
++ {
++ std::lock_guard guard(ptyJobMutex);
++ pty_baton* handle = get_pty_baton(id);
++ // Why the consoleClosed check: a second kill() would otherwise close the
++ // same pseudoconsole twice. Upstream relied on the baton being gone.
++ if (handle != nullptr && !handle->consoleClosed) {
++ hpc = handle->hpc;
++ owed = true;
++ handle->consoleClosed = true;
++ // Null hShell means a self-exited pty, where there is nothing to kill.
++ if (useConptyDll && handle->hShell != nullptr) {
++ if (!DuplicateHandle(GetCurrentProcess(), handle->hShell, GetCurrentProcess(),
++ &hShellDup, 0, FALSE, DUPLICATE_SAME_ACCESS)) {
++ // Why terminate here instead of skipping: a failed duplication leaves
++ // hShellDup null, which is indistinguishable from the self-exit case,
++ // and skipping would leave the shell RUNNING after its pane closed --
++ // a worse outcome than the leak this all exists to fix. hShell is
++ // valid under this lock and TerminateProcess does not block, so the
++ // only cost is that this rare path kills before the console closes.
++ hShellDup = nullptr;
++ TerminateProcess(handle->hShell, 1);
++ }
++ }
++ if (handle->shellExited) {
++ const bool removed = remove_pty_baton(id);
++ assert(removed);
++ (void)removed;
+ }
++ // Else the shell is still running and the watcher frees the baton.
+ }
+- if (useConptyDll) {
+- TerminateProcess(handle->hShell, 1);
++ }
++
++ // Why outside the lock: ClosePseudoConsole blocks until the conout side has
++ // drained, and the watcher must be able to take the lock while it does.
++ if (owed) {
++ if (pfnClosePseudoConsole)
++ {
++ pfnClosePseudoConsole(hpc);
++ }
++ if (hShellDup != nullptr) {
++ TerminateProcess(hShellDup, 1);
++ CloseHandle(hShellDup);
+ }
+ }
+
return env.Undefined();
}
@@ -808,9 +906,11 @@ index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a
+ * Orca: the pids still alive in this pty's tree, straight from the kernel.
+ *
+ * Descendant liveness for a tree that is still tracked, including children that
-+ * detached from the console. Once the shell exits the baton is gone, so this
-+ * returns null rather than an empty list -- null means "no answer", never
-+ * "they died". Also returns null when no job was assigned.
++ * detached from the console. Once the shell exits the watcher nulls hJob, which
++ * ownsShell rejects, so this returns null rather than an empty list -- null
++ * means "no answer", never "they died". (The baton itself now outlives the
++ * shell, until kill() runs; hJob is what makes the answer null.) Also returns
++ * null when no job was assigned.
+ *
+ * Does not include the ConPTY console host: CreatePseudoConsole spawns it
+ * before this job exists, so it is not a member and ClosePseudoConsole is what
@@ -906,7 +1006,7 @@ index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a
/**
* Init
*/
-@@ -577,6 +804,9 @@ Napi::Object init(Napi::Env env, Napi::Object exports) {
+@@ -577,6 +867,9 @@ Napi::Object init(Napi::Env env, Napi::Object exports) {
exports.Set("resize", Napi::Function::New(env, PtyResize));
exports.Set("clear", Napi::Function::New(env, PtyClear));
exports.Set("kill", Napi::Function::New(env, PtyKill));
@@ -917,7 +1017,7 @@ index 7b286d3d644c26141df516929703aa6e129df4b2..ec6bf3932c65b89c013ff133dc6bf46a
};
diff --git a/lib/windowsPtyAgent.js b/lib/windowsPtyAgent.js
-index a358ffb..fb3a96f 100644
+index a358ffb177357e177661033c1b092f9c9d0e5f5a..26c2a4c58799ce649f5113131e4c52f7ed2d87ad 100644
--- a/lib/windowsPtyAgent.js
+++ b/lib/windowsPtyAgent.js
@@ -136,6 +136,9 @@ var WindowsPtyAgent = /** @class */ (function () {
@@ -930,6 +1030,20 @@ index a358ffb..fb3a96f 100644
this._outSocket.readable = false;
this._getConsoleProcessList().then(function (consoleProcessList) {
consoleProcessList.forEach(function (pid) {
+@@ -154,9 +157,10 @@ var WindowsPtyAgent = /** @class */ (function () {
+ // Close the input write handle to signal the end of session.
+ this._inSocket.destroy();
+ this._ptyNative.kill(this._pty, this._useConptyDll);
+- this._outSocket.on('data', function () {
+- _this._conoutSocketWorker.dispose();
+- });
++ // Orca: dispose unconditionally, as the non-DLL branch above does.
++ // Waiting for another 'data' event leaks the conout worker on every
++ // self-exiting shell, because no more data ever arrives (F24).
++ this._conoutSocketWorker.dispose();
+ }
+ }
+ else {
diff --git a/lib/windowsTerminal.js b/lib/windowsTerminal.js
index 3c38f89..e20b3e6 100644
--- a/lib/windowsTerminal.js
@@ -1015,7 +1129,7 @@ index 3c38f89..e20b3e6 100644
\ No newline at end of file
+//# sourceMappingURL=windowsTerminal.js.map
diff --git a/src/windowsPtyAgent.ts b/src/windowsPtyAgent.ts
-index d705444..ce611b8 100644
+index d7054449516f0c9a62af351c2caa17331206d530..0c28a32e2e1db2b3f208ddde8443cd4e67bb1ad6 100644
--- a/src/windowsPtyAgent.ts
+++ b/src/windowsPtyAgent.ts
@@ -143,6 +143,9 @@ export class WindowsPtyAgent {
@@ -1028,6 +1142,20 @@ index d705444..ce611b8 100644
this._outSocket.readable = false;
this._getConsoleProcessList().then(consoleProcessList => {
consoleProcessList.forEach((pid: number) => {
+@@ -159,9 +162,10 @@ export class WindowsPtyAgent {
+ // Close the input write handle to signal the end of session.
+ this._inSocket.destroy();
+ (this._ptyNative as IConptyNative).kill(this._pty, this._useConptyDll);
+- this._outSocket.on('data', () => {
+- this._conoutSocketWorker.dispose();
+- });
++ // Orca: dispose unconditionally, as the non-DLL branch above does.
++ // Waiting for another 'data' event leaks the conout worker on every
++ // self-exiting shell, because no more data ever arrives (F24).
++ this._conoutSocketWorker.dispose();
+ }
+ } else {
+ // Because pty.kill closes the handle, it will kill most processes by itself.
diff --git a/src/windowsTerminal.ts b/src/windowsTerminal.ts
index 13f6c6d..eda63c8 100644
--- a/src/windowsTerminal.ts
diff --git a/config/relay-assets/node-pty-1.1.0-windows-pty-teardown-patch.cjs b/config/relay-assets/node-pty-1.1.0-windows-pty-teardown-patch.cjs
new file mode 100644
index 00000000000..dd26784ee46
--- /dev/null
+++ b/config/relay-assets/node-pty-1.1.0-windows-pty-teardown-patch.cjs
@@ -0,0 +1,158 @@
+const { createHash } = require('node:crypto')
+const { readFileSync, renameSync, rmSync, writeFileSync } = require('node:fs')
+const { join, resolve } = require('node:path')
+
+/**
+ * Release the ConPTY teardown handles a relay's npm-installed node-pty never releases.
+ *
+ * Two files, and the ORDER of one of the edits is the whole fix.
+ *
+ * `windowsPtyAgent.js` -- `kill()` flips `readable` on both sockets and destroys neither.
+ * `_cleanUpProcess` destroys `_outSocket`, so the conout handle comes back; nothing ever destroys
+ * `_inSocket`, and it wraps a real Windows named-pipe handle from `fs.openSync(term.conin, 'w')`.
+ * Every terminal leaks one File handle for the life of the host process.
+ *
+ * The obvious fix -- and the one the desktop patch ships -- releases it at the TOP of the branch,
+ * before `_getConsoleProcessList()` forks and before the native kill. That is measurably worse than
+ * leaving the leak alone: teardown aborts partway, the forked console-list agent is never reaped,
+ * and both pipe handles stay alive instead of one. This asset releases it at the END of the branch
+ * instead, after the fork and the kill have already happened.
+ *
+ * Measured on a Windows SSH host, 20 spawn/kill cycles, handles bucketed by NT object type
+ * (identical numbers standalone and through a real relay):
+ *
+ * published node-pty File +1/terminal, Process flat
+ * desktop patch placement File +2/terminal, Process +1/terminal <-- 3x WORSE
+ * released last (here) File flat, Process flat
+ *
+ * `windowsTerminal.js` carries the desktop's error-listener hunks verbatim. The conin listener is
+ * what keeps a pipe error retiring one terminal instead of the host -- its own comment names the
+ * failure mode: "Without a listener, Node promotes errors such as write EAGAIN to uncaughtException".
+ * It is not what fixes the leak (adding it changed nothing on its own), but it is the guard that
+ * makes destroying conin safe at all.
+ *
+ * Why this ships as a relay asset rather than only in config/patches/node-pty@1.1.0.patch: pnpm
+ * patches do not cross the SSH boundary -- a relay host runs the tree `npm install` put there.
+ *
+ * DELIBERATE DIVERGENCE FROM THE DESKTOP: the desktop patch has the early placement and therefore
+ * the +2 File / +1 Process regression, measured against its exact installed tree. Correcting it
+ * there is a separate change with its own verification, so the two trees differ on this one hunk on
+ * purpose, and the test pins that so a future "sync the patches" does not copy the bug back.
+ *
+ * NOT ADDRESSED, AND A SEPARATE DEFECT THAT IS STILL OPEN: a terminal that exits on its own is
+ * still torn down through `kill()` -- both hosts call `destroy()` on natural exit and
+ * `WindowsTerminal.destroy()` is `kill()` -- but the shell is already gone by then, and the
+ * ordering this patch relies on does not hold. Measured over 20 self-exit cycles with that
+ * `destroy()` issued: published +3 File/+1 Process per terminal, desktop-patched +2/+1, this tree
+ * +2/+1. So this patch does not close it and the desktop patch does not either. It is reachable
+ * for every Windows user, local and relay, on every terminal closed by typing `exit`.
+ */
+
+const EXPECTED_NODE_PTY_VERSION = '1.1.0'
+
+/** Each entry is one published file, its patched form, and the edits between them. */
+const PATCH_TARGETS = [
+ {
+ relativePath: ['lib', 'windowsPtyAgent.js'],
+ originalSha256: '8636d16b38266112204061a22b135734177c242837982fd3a4055be726efa64a',
+ patchedSha256: '1e23ef480569e73706e3ab4f5482c7e553c76f51414ae8e7b0bdcc2fd75f7280',
+ replacements: [
+ [
+ ' this._ptyNative.kill(this._pty, this._useConptyDll);\n this._conoutSocketWorker.dispose();\n',
+ ' this._ptyNative.kill(this._pty, this._useConptyDll);\n this._conoutSocketWorker.dispose();\n // Orca: released AFTER the console-list fork and the native kill, not before them.\n // Destroying conin first aborts teardown partway -- measured on a Windows SSH relay\n // as +2 File and +1 Process handles per terminal, against +1 File unpatched.\n this._inSocket.destroy();\n'
+ ]
+ ]
+ },
+ {
+ relativePath: ['lib', 'windowsTerminal.js'],
+ originalSha256: 'c3a65716f53fed0135a8a633373d5f9c2ab092544d651f27ef0a67096dd3bcd9',
+ patchedSha256: '8247ecd69be8b18257050fb026b290024612c5ffc6d492ff1d46f81e613be2cf',
+ replacements: [
+ [
+ ' _this._agent = new windowsPtyAgent_1.WindowsPtyAgent(file, args, parsedEnv, cwd, _this._cols, _this._rows, false, opt.useConpty, opt.useConptyDll, opt.conptyInheritCursor);\n _this._socket = _this._agent.outSocket;\n // Not available until `ready` event emitted.\n _this._pid = _this._agent.innerPid;',
+ " _this._agent = new windowsPtyAgent_1.WindowsPtyAgent(file, args, parsedEnv, cwd, _this._cols, _this._rows, false, opt.useConpty, opt.useConptyDll, opt.conptyInheritCursor);\n _this._socket = _this._agent.outSocket;\n // Attach before readiness so a broken ConPTY output pipe cannot be unhandled.\n _this._socket.on('error', function (err) {\n var code = err && err.code;\n // PTY output can report EPIPE before `_close()` wins the race.\n _this._close();\n if (code === 'EPIPE' || code === 'ERR_STREAM_PUSH_AFTER_EOF' || code === 'ERR_STREAM_DESTROYED') {\n return;\n }\n // EIO, happens when someone closes our child process: the only process\n // in the terminal.\n // node < 0.6.14: errno 5\n // node >= 0.6.14: read EIO\n if (typeof code === 'string') {\n if (~code.indexOf('errno 5') || ~code.indexOf('EIO'))\n return;\n }\n // Throw anything else.\n if (_this.listeners('error').length < 2) {\n throw err;\n }\n });\n // Not available until `ready` event emitted.\n _this._pid = _this._agent.innerPid;"
+ ],
+ [
+ " }\n });\n // Shutdown if `error` event is emitted.\n _this._socket.on('error', function (err) {\n // Close terminal session.\n _this._close();\n // EIO, happens when someone closes our child process: the only process\n // in the terminal.\n // node < 0.6.14: errno 5\n // node >= 0.6.14: read EIO\n if (err.code) {\n if (~err.code.indexOf('errno 5') || ~err.code.indexOf('EIO'))\n return;\n }\n // Throw anything else.\n if (_this.listeners('error').length < 2) {\n throw err;\n }\n });\n // Cleanup after the socket is closed.\n _this._socket.on('close', function () {",
+ " }\n });\n // Cleanup after the socket is closed.\n _this._socket.on('close', function () {"
+ ],
+ [
+ ' _this._readable = true;\n _this._writable = true;\n _this._forwardEvents();\n return _this;',
+ " _this._readable = true;\n _this._writable = true;\n // A ConPTY input-pipe error must retire only this terminal. Without a listener, Node promotes\n // errors such as write EAGAIN to uncaughtException and kills every PTY in the daemon.\n _this._agent.inSocket.on('error', function () {\n if (!_this._writable) {\n return;\n }\n _this._close();\n try {\n _this._agent.kill();\n }\n catch (_a) {\n // The failing pipe may have raced process exit; the terminal is already unwritable.\n }\n });\n _this._forwardEvents();\n return _this;"
+ ],
+ [
+ 'exports.WindowsTerminal = WindowsTerminal;\n//# sourceMappingURL=windowsTerminal.js.map',
+ 'exports.WindowsTerminal = WindowsTerminal;\n//# sourceMappingURL=windowsTerminal.js.map\n'
+ ]
+ ]
+ }
+]
+
+function inspectTarget(relayDir, target) {
+ const nodePtyDir = resolve(relayDir, 'node_modules', 'node-pty')
+ const packageJson = JSON.parse(readFileSync(join(nodePtyDir, 'package.json'), 'utf8'))
+ if (packageJson.version !== EXPECTED_NODE_PTY_VERSION) {
+ throw new Error(
+ `Refusing to patch node-pty ${packageJson.version}; expected ${EXPECTED_NODE_PTY_VERSION}`
+ )
+ }
+ const filePath = join(nodePtyDir, ...target.relativePath)
+ return { filePath, source: readFileSync(filePath, 'utf8') }
+}
+
+function assertPatchedNodePtyWindowsTeardown(relayDir = process.cwd()) {
+ for (const target of PATCH_TARGETS) {
+ const inspected = inspectTarget(relayDir, target)
+ if (sourceSha256(inspected.source) !== target.patchedSha256) {
+ throw new Error(
+ `node-pty ConPTY teardown release is not installed in ${target.relativePath.join('/')}`
+ )
+ }
+ }
+}
+
+function patchNodePtyWindowsTeardown(relayDir = process.cwd()) {
+ for (const target of PATCH_TARGETS) {
+ const inspected = inspectTarget(relayDir, target)
+ const sourceHash = sourceSha256(inspected.source)
+ if (sourceHash === target.patchedSha256) {
+ continue
+ }
+ if (sourceHash !== target.originalSha256) {
+ throw new Error(
+ `Refusing to patch unexpected node-pty source in ${target.relativePath.join('/')}`
+ )
+ }
+ let patchedSource = inspected.source
+ for (const [from, to] of target.replacements) {
+ // Why the count check: an anchor that matched twice would patch the wrong site silently, and
+ // the hash below would then reject a tree this script had already rewritten.
+ if (patchedSource.split(from).length - 1 !== 1) {
+ throw new Error(`Refusing to patch ${target.relativePath.join('/')}; anchor is not unique`)
+ }
+ patchedSource = patchedSource.replace(from, to)
+ }
+ const temporaryPath = `${inspected.filePath}.orca-patch-${process.pid}`
+ // Why: a terminated remote install must leave either known source version recoverable on reconnect.
+ try {
+ writeFileSync(temporaryPath, patchedSource)
+ renameSync(temporaryPath, inspected.filePath)
+ } finally {
+ rmSync(temporaryPath, { force: true })
+ }
+ }
+ assertPatchedNodePtyWindowsTeardown(relayDir)
+}
+
+function sourceSha256(source) {
+ return createHash('sha256').update(source).digest('hex')
+}
+
+if (require.main === module) {
+ patchNodePtyWindowsTeardown()
+}
+
+module.exports = {
+ assertPatchedNodePtyWindowsTeardown,
+ patchNodePtyWindowsTeardown
+}
diff --git a/config/scripts/build-relay.mjs b/config/scripts/build-relay.mjs
index 289c7a957bd..4d408712f97 100644
--- a/config/scripts/build-relay.mjs
+++ b/config/scripts/build-relay.mjs
@@ -57,6 +57,13 @@ const NODE_PTY_CONSOLE_LIST_PATCH_SOURCE = join(
'relay-assets',
NODE_PTY_CONSOLE_LIST_PATCH_FILENAME
)
+const NODE_PTY_WINDOWS_TEARDOWN_PATCH_FILENAME = 'node-pty-1.1.0-windows-pty-teardown-patch.cjs'
+const NODE_PTY_WINDOWS_TEARDOWN_PATCH_SOURCE = join(
+ ROOT,
+ 'config',
+ 'relay-assets',
+ NODE_PTY_WINDOWS_TEARDOWN_PATCH_FILENAME
+)
const NODE_PTY_MASTER_CLOEXEC_PATCH_FILENAME = 'node-pty-1.1.0-master-cloexec-patch.cjs'
const NODE_PTY_MASTER_CLOEXEC_PATCH_SOURCE = join(
ROOT,
@@ -132,6 +139,10 @@ for (const platform of RELAY_BUILD_PLATFORMS) {
NODE_PTY_CONSOLE_LIST_PATCH_SOURCE,
join(outDir, NODE_PTY_CONSOLE_LIST_PATCH_FILENAME)
)
+ copyFileSync(
+ NODE_PTY_WINDOWS_TEARDOWN_PATCH_SOURCE,
+ join(outDir, NODE_PTY_WINDOWS_TEARDOWN_PATCH_FILENAME)
+ )
}
copyFileSync(
NODE_PTY_MASTER_CLOEXEC_PATCH_SOURCE,
diff --git a/config/scripts/electron-builder-runtime-resources.test.mjs b/config/scripts/electron-builder-runtime-resources.test.mjs
index d2407776fa7..453d5702cb0 100644
--- a/config/scripts/electron-builder-runtime-resources.test.mjs
+++ b/config/scripts/electron-builder-runtime-resources.test.mjs
@@ -1,14 +1,18 @@
+import { readFileSync, readdirSync } from 'node:fs'
import { cp, mkdir, mkdtemp, readFile, readdir, rm, stat, writeFile } from 'node:fs/promises'
import { createRequire } from 'node:module'
import { tmpdir } from 'node:os'
-import { join } from 'node:path'
+import { dirname, join, relative, resolve } from 'node:path'
import { describe, expect, it } from 'vitest'
const require = createRequire(import.meta.url)
+const projectRoot = resolve(import.meta.dirname, '..', '..')
const electronBuilderConfig = require('../electron-builder.config.cjs')
const {
createPackagedRuntimeNodeModuleResources,
findAsarEntry,
+ isPackagedExternalSpecifier,
+ packageNameFromSpecifier,
prunePackagedNodePty,
prunePackagedParcelWatcher,
prunePackagedSherpaOnnx,
@@ -306,3 +310,91 @@ describe('packaged runtime resources', () => {
}
)
})
+
+// Why source-anchored: the bundler renames a createRequire()'d require, so
+// verifyPackagedMainRuntimeDeps' `require("x")` scan cannot see these specifiers — packaging
+// stays green while the packaged app throws MODULE_NOT_FOUND the first time the path runs.
+function collectLazyRequireSpecifiers(directory, found = new Map()) {
+ for (const entry of readdirSync(directory, { withFileTypes: true })) {
+ const entryPath = join(directory, entry.name)
+ if (entry.isDirectory()) {
+ collectLazyRequireSpecifiers(entryPath, found)
+ continue
+ }
+ if (!entry.isFile() || !entry.name.endsWith('.ts') || entry.name.includes('.test.')) {
+ continue
+ }
+ const source = readFileSync(entryPath, 'utf8')
+ if (!source.includes('createRequire(')) {
+ continue
+ }
+ for (const match of source.matchAll(/\brequire[A-Za-z0-9_]*\(\s*'([^']+)'\s*\)/g)) {
+ if (isPackagedExternalSpecifier(match[1])) {
+ found.set(match[1], relative(projectRoot, entryPath).replaceAll('\\', '/'))
+ }
+ }
+ }
+ return found
+}
+
+function packagedResourceDestinations(platform) {
+ return new Set(
+ (electronBuilderConfig[platform].extraResources ?? []).map((resource) =>
+ String(resource.to).replaceAll('\\', '/')
+ )
+ )
+}
+
+describe('lazily required packages reach Resources/node_modules', () => {
+ it('copies every createRequire specifier main uses into the packaged resource plan', () => {
+ const specifiers = collectLazyRequireSpecifiers(join(projectRoot, 'src', 'main'))
+ expect(specifiers.size).toBeGreaterThan(0)
+
+ const destinations = {
+ win: packagedResourceDestinations('win'),
+ mac: packagedResourceDestinations('mac'),
+ linux: packagedResourceDestinations('linux')
+ }
+ for (const [specifier, source] of specifiers) {
+ const packageName = packageNameFromSpecifier(specifier)
+ const covered = (platform) =>
+ destinations[platform].has(`node_modules/${packageName}`) ||
+ destinations[platform].has(`node_modules/${specifier}`)
+ // Windows carries the full closure, so an uncovered specifier is uncovered everywhere.
+ expect(
+ covered('win'),
+ `${source} lazily requires '${specifier}', but nothing copies it to Resources/node_modules`
+ ).toBe(true)
+ if (covered('mac') && covered('linux')) {
+ continue
+ }
+ // Only the Windows-native loaders may be absent from the mac/linux plans.
+ expect(source, `'${specifier}' is packaged for Windows only`).toContain('windows')
+ }
+ })
+
+ it('resolves the copied emoji dataset the way the packaged main bundle does', async () => {
+ const resourcesDir = await mkdtemp(join(tmpdir(), 'orca-lazy-require-'))
+ try {
+ const datasetPath = 'node_modules/emojibase-data/en/shortcodes/emojibase.json'
+ const entry = electronBuilderConfig.mac.extraResources.find(
+ (resource) => String(resource.to) === datasetPath
+ )
+ expect(entry).toBeDefined()
+ const destination = join(resourcesDir, ...datasetPath.split('/'))
+ await mkdir(dirname(destination), { recursive: true })
+ await cp(join(projectRoot, ...String(entry.from).split('/')), destination)
+
+ // app.asar's parent is Resources, so main's bare require walks into Resources/node_modules.
+ const packagedMainDir = join(resourcesDir, 'app.asar', 'out', 'main')
+ await mkdir(packagedMainDir, { recursive: true })
+ const probe = join(packagedMainDir, 'probe.cjs')
+ await writeFile(probe, 'module.exports = require', 'utf8')
+
+ const dataset = require(probe)('emojibase-data/en/shortcodes/emojibase.json')
+ expect(Object.keys(dataset).length).toBeGreaterThan(1000)
+ } finally {
+ await rm(resourcesDir, { recursive: true, force: true })
+ }
+ })
+})
diff --git a/config/scripts/locale-ko-key-overrides.json b/config/scripts/locale-ko-key-overrides.json
index f368ecc3cbc..bf5f62d1fa5 100644
--- a/config/scripts/locale-ko-key-overrides.json
+++ b/config/scripts/locale-ko-key-overrides.json
@@ -492,7 +492,7 @@
"ko": "agent CLI를 찾지 못했습니다. 하나를 설치하거나 설정에서 기본 agent를 선택하세요."
},
"auto.components.Terminal.7958465754": {
- "ko": "실행 중인 프로세스가 있는 로컬 terminals이 있습니다. 그래도 창을 닫으시겠습니까?"
+ "ko": "실행 중인 프로세스가 있는 terminals이 있습니다. 그래도 창을 닫으시겠습니까?"
},
"auto.components.Terminal.cdc9ac4b2d": {
"ko": "편집기"
diff --git a/config/scripts/node-pty-windows-pty-teardown-patch.test.mjs b/config/scripts/node-pty-windows-pty-teardown-patch.test.mjs
new file mode 100644
index 00000000000..64fb1b056b8
--- /dev/null
+++ b/config/scripts/node-pty-windows-pty-teardown-patch.test.mjs
@@ -0,0 +1,213 @@
+// The relay's copy of the ConPTY teardown release, and the guard that keeps it in lockstep with the
+// desktop's own node-pty patch. pnpm patches do not cross the SSH boundary, so a relay runs the tree
+// `npm install` put there; the desktop had this fix and the relay did not, and every terminal on a
+// Windows SSH host leaked one File handle for the life of the relay process.
+//
+// The ORDER of the conin release is the fix. Releasing it at the top of the branch -- what the
+// desktop patch does -- was measured at 3x WORSE than shipping nothing (File +2/terminal and a new
+// Process +1/terminal); releasing it after the console-list fork and the native kill is flat.
+import { createRequire } from 'node:module'
+import { existsSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs'
+import { join, resolve } from 'node:path'
+import { afterEach, describe, expect, it } from 'vitest'
+
+const require = createRequire(import.meta.url)
+const {
+ assertPatchedNodePtyWindowsTeardown,
+ patchNodePtyWindowsTeardown
+} = require('../relay-assets/node-pty-1.1.0-windows-pty-teardown-patch.cjs')
+const projectDir = resolve(import.meta.dirname, '..', '..')
+const cleanupDirs = []
+
+const PATCHED_FILES = ['windowsPtyAgent.js', 'windowsTerminal.js']
+
+/** The hunks config/patches/node-pty@1.1.0.patch adds to the installed desktop tree. */
+const DESKTOP_HUNKS = {
+ 'windowsPtyAgent.js': [
+ [
+ [
+ ' this._inSocket.readable = false;',
+ ' // The non-DLL path previously only flipped `readable`, leaving the',
+ ' // conin PipeWrap alive until the host exited (#947).',
+ ' this._inSocket.destroy();',
+ ' this._outSocket.readable = false;',
+ ''
+ ].join('\n'),
+ [
+ ' this._inSocket.readable = false;',
+ ' this._outSocket.readable = false;',
+ ''
+ ].join('\n')
+ ],
+ // The useConptyDll branch, which only the DESKTOP runs -- the relay takes the
+ // non-DLL branch above, where the dispose is already unconditional. Listed here
+ // so un-applying still yields published; the relay asset needs no counterpart.
+ [
+ [
+ ' // Orca: dispose unconditionally, as the non-DLL branch above does.',
+ " // Waiting for another 'data' event leaks the conout worker on every",
+ ' // self-exiting shell, because no more data ever arrives (F24).',
+ ' this._conoutSocketWorker.dispose();',
+ ''
+ ].join('\n'),
+ [
+ " this._outSocket.on('data', function () {",
+ ' _this._conoutSocketWorker.dispose();',
+ ' });',
+ ''
+ ].join('\n')
+ ]
+ ],
+ 'windowsTerminal.js': [
+ [
+ ' // Attach before readiness so a broken ConPTY output pipe cannot be unhandled.',
+ null
+ ],
+ [' // A ConPTY input-pipe error must retire only this terminal.', null]
+ ]
+}
+
+function desktopPath(file) {
+ return join(projectDir, 'node_modules', 'node-pty', 'lib', file)
+}
+
+afterEach(() => {
+ for (const dir of cleanupDirs.splice(0)) {
+ rmSync(dir, { recursive: true, force: true })
+ }
+})
+
+describe('Windows SSH relay node-pty ConPTY teardown patch', () => {
+ // Why reconstruct rather than vendor upstream: the installed tree IS the published file plus the
+ // desktop's hunks, so un-applying them yields upstream exactly -- and pinning that against this
+ // asset's own hashes is what fails loudly if either side of the pair moves.
+ it('takes the desktop error listeners verbatim', () => {
+ const fixture = writeNodePtyFixture('1.1.0')
+ patchNodePtyWindowsTeardown(fixture.root)
+
+ expect(readFileSync(join(fixture.libDir, 'windowsTerminal.js'), 'utf8')).toBe(
+ readFileSync(desktopPath('windowsTerminal.js'), 'utf8')
+ )
+ })
+
+ // The one hunk that must NOT match the desktop, and the reason is measured, not stylistic:
+ // releasing conin before `_getConsoleProcessList()` forks aborts teardown partway.
+ it('releases conin after the console-list fork, not before it like the desktop patch', () => {
+ const fixture = writeNodePtyFixture('1.1.0')
+ patchNodePtyWindowsTeardown(fixture.root)
+ const patched = readFileSync(join(fixture.libDir, 'windowsPtyAgent.js'), 'utf8')
+
+ const branch = patched.slice(
+ patched.indexOf('if (!this._useConptyDll) {'),
+ patched.indexOf('else {', patched.indexOf('if (!this._useConptyDll) {'))
+ )
+ expect(branch).toContain('this._inSocket.destroy();')
+ expect(branch.indexOf('this._inSocket.destroy();')).toBeGreaterThan(
+ branch.indexOf('this._conoutSocketWorker.dispose();')
+ )
+ expect(branch.indexOf('this._inSocket.destroy();')).toBeGreaterThan(
+ branch.indexOf('this._getConsoleProcessList()')
+ )
+ // Pinned so a future "sync the relay asset to config/patches" cannot copy the regression back.
+ expect(patched).not.toBe(readFileSync(desktopPath('windowsPtyAgent.js'), 'utf8'))
+ })
+
+ it('installs and verifies idempotently', () => {
+ const fixture = writeNodePtyFixture('1.1.0')
+
+ patchNodePtyWindowsTeardown(fixture.root)
+ const once = PATCHED_FILES.map((file) => readFileSync(join(fixture.libDir, file), 'utf8'))
+ for (const file of PATCHED_FILES) {
+ expect(existsSync(`${join(fixture.libDir, file)}.orca-patch-${process.pid}`)).toBe(false)
+ }
+ expect(() => assertPatchedNodePtyWindowsTeardown(fixture.root)).not.toThrow()
+
+ patchNodePtyWindowsTeardown(fixture.root)
+ expect(PATCHED_FILES.map((file) => readFileSync(join(fixture.libDir, file), 'utf8'))).toEqual(
+ once
+ )
+ })
+
+ it('refuses a different package version or unexpected source', () => {
+ const wrongVersion = writeNodePtyFixture('1.2.0-beta.11')
+ expect(() => patchNodePtyWindowsTeardown(wrongVersion.root)).toThrow('expected 1.1.0')
+
+ for (const file of PATCHED_FILES) {
+ const drifted = writeNodePtyFixture('1.1.0')
+ const path = join(drifted.libDir, file)
+ writeFileSync(path, `${readFileSync(path, 'utf8')}\n// drift`)
+ expect(() => patchNodePtyWindowsTeardown(drifted.root)).toThrow('unexpected node-pty')
+ }
+ })
+
+ it('refuses a half-applied tree, so one file cannot pass for both', () => {
+ for (const file of PATCHED_FILES) {
+ const partial = writeNodePtyFixture('1.1.0')
+ const fixture = writeNodePtyFixture('1.1.0')
+ patchNodePtyWindowsTeardown(fixture.root)
+ writeFileSync(join(partial.libDir, file), readFileSync(join(fixture.libDir, file), 'utf8'))
+ expect(() => assertPatchedNodePtyWindowsTeardown(partial.root)).toThrow('is not installed')
+ }
+ })
+})
+
+/** A published node-pty tree, rebuilt by un-applying the desktop hunks from the installed one. */
+function writeNodePtyFixture(version) {
+ const root = mkdtempSync(join(projectDir, '.node-pty-teardown-patch-test-'))
+ cleanupDirs.push(root)
+ const libDir = join(root, 'node_modules', 'node-pty', 'lib')
+ mkdirSync(libDir, { recursive: true })
+ writeFileSync(join(root, 'node_modules', 'node-pty', 'package.json'), JSON.stringify({ version }))
+ for (const file of PATCHED_FILES) {
+ const desktop = readFileSync(desktopPath(file), 'utf8')
+ for (const [marker] of DESKTOP_HUNKS[file]) {
+ expect(desktop).toContain(marker)
+ }
+ writeFileSync(join(libDir, file), unapplyDesktopHunks(file, desktop))
+ }
+ return { root, libDir }
+}
+
+/**
+ * Reverse of the published-to-desktop transform.
+ *
+ * `windowsTerminal.js` is taken verbatim from the desktop, so the asset's own replacement table is
+ * the transform and reversing it is exact. `windowsPtyAgent.js` deliberately diverges, so its
+ * published form is rebuilt from the desktop hunk instead -- which is also what makes this file the
+ * place that notices if the desktop hunk itself ever moves.
+ */
+function unapplyDesktopHunks(file, desktop) {
+ if (file === 'windowsPtyAgent.js') {
+ let published = desktop
+ for (const [patched, original] of DESKTOP_HUNKS[file]) {
+ expect(published.split(patched).length - 1).toBe(1)
+ published = published.replace(patched, original)
+ }
+ return published
+ }
+ const asset = readFileSync(
+ join(projectDir, 'config', 'relay-assets', 'node-pty-1.1.0-windows-pty-teardown-patch.cjs'),
+ 'utf8'
+ )
+ const { PATCH_TARGETS } = loadPatchTargets(asset)
+ const target = PATCH_TARGETS.find((entry) => entry.relativePath.at(-1) === file)
+ expect(target).toBeDefined()
+ let published = desktop
+ for (const [from, to] of target.replacements.toReversed()) {
+ expect(published.split(to).length - 1).toBe(1)
+ published = published.replace(to, from)
+ }
+ return published
+}
+
+function loadPatchTargets(assetSource) {
+ const module = { exports: {} }
+ const factory = new Function(
+ 'module',
+ 'exports',
+ 'require',
+ `${assetSource}\nmodule.exports.PATCH_TARGETS = PATCH_TARGETS`
+ )
+ factory(module, module.exports, require)
+ return module.exports
+}
diff --git a/config/scripts/pr-code-change-scope.mjs b/config/scripts/pr-code-change-scope.mjs
index 7111531c35c..f5a73f6239f 100644
--- a/config/scripts/pr-code-change-scope.mjs
+++ b/config/scripts/pr-code-change-scope.mjs
@@ -228,6 +228,7 @@ const WINDOWS_PACKAGE_TESTS = [
'src/main/cli/wsl-cli-powershell-boundary.test.ts',
'src/main/cursor/hook-service.test.ts',
'src/main/orca-profiles/profile-index-store.test.ts',
+ 'src/main/startup/windows-install-dir-acl-repair.win32.test.ts',
'src/main/runtime/repo-worktree-admin-fingerprint.test.ts',
'src/main/runtime/worktree-scan-admin-fingerprint-gate.test.ts',
'src/shared/secure-file-fsync-flags.test.ts',
diff --git a/config/scripts/skill-description-length.test.mjs b/config/scripts/skill-description-length.test.mjs
new file mode 100644
index 00000000000..e7a9db79541
--- /dev/null
+++ b/config/scripts/skill-description-length.test.mjs
@@ -0,0 +1,39 @@
+import { readdirSync, readFileSync } from 'node:fs'
+import { join, resolve } from 'node:path'
+import { describe, expect, it } from 'vitest'
+import { parse } from 'yaml'
+
+const skillsDir = resolve(import.meta.dirname, '../../skills')
+// Why: the Agent Skills spec caps `description` at 1024 chars and conforming installers
+// reject the whole skill (#17935); the frontmatter is what the installer parses, so check it.
+const MAX_DESCRIPTION_LENGTH = 1024
+
+function readDescription(skillName) {
+ const skillMarkdown = readFileSync(join(skillsDir, skillName, 'SKILL.md'), 'utf8')
+ const frontmatter = /^---\r?\n([\s\S]*?)\r?\n---\r?\n/u.exec(skillMarkdown)?.[1]
+
+ expect(frontmatter, `${skillName}: missing frontmatter`).toBeDefined()
+
+ return parse(frontmatter ?? '').description
+}
+
+describe('bundled skill descriptions', () => {
+ const skillNames = readdirSync(skillsDir, { withFileTypes: true })
+ .filter((entry) => entry.isDirectory())
+ .map((entry) => entry.name)
+
+ it('discovers the bundled skills', () => {
+ expect(skillNames).toContain('orchestration')
+ })
+
+ it.each(skillNames)('%s keeps description within the Agent Skills spec limit', (name) => {
+ const description = readDescription(name)
+
+ expect(typeof description, `${name}: description must be a string`).toBe('string')
+ expect(description.trim().length, `${name}: description is empty`).toBeGreaterThan(0)
+ expect(
+ description.length,
+ `${name}: description is ${description.length} chars`
+ ).toBeLessThanOrEqual(MAX_DESCRIPTION_LENGTH)
+ })
+})
diff --git a/config/tsconfig.tc.web.json b/config/tsconfig.tc.web.json
index 56253527c69..2caf2149f73 100644
--- a/config/tsconfig.tc.web.json
+++ b/config/tsconfig.tc.web.json
@@ -19,6 +19,7 @@
"../src/preload/usage-provider-api.ts",
"../src/shared/**/*",
"../src/main/gitlab/mappers.ts",
+ "../src/main/ipc/deferred-emoji-shortcode-dataset.ts",
"../src/main/ipc/worktree-branch-name.ts",
"../src/main/ipc/worktree-logic.ts",
"../src/main/ipc/worktree-display-name.ts",
diff --git a/docs/assets/readme-downloads.svg b/docs/assets/readme-downloads.svg
index ef8ebb61bb4..c240c965fee 100644
--- a/docs/assets/readme-downloads.svg
+++ b/docs/assets/readme-downloads.svg
@@ -1,5 +1,5 @@
-
@@ -243,9 +243,9 @@ Associez-la à l'app de bureau pour surveiller et piloter vos agents depuis votr
- **Discord :** Rejoignez la communauté sur **[Discord](https://discord.gg/fzjDKHxv8Q)**.
- **Twitter / X :** Suivez **[@orca_build](https://x.com/orca_build)** pour les news et annonces.
-- **WeChat :** Scannez pour rejoindre le groupe WeChat 8 de la communauté Orca.
+- **WeChat :** Scannez pour rejoindre le groupe WeChat 8 de la communauté Orca. Le groupe 8 est peut-être complet ; dans ce cas, scannez plutôt le QR code du groupe 9.
-
+
- **Feedback & idées :** On ship vite. Il manque quelque chose ? [Demandez une feature](https://github.com/stablyai/orca/issues).
- **Confidentialité :** Voir la [doc confidentialité & télémétrie](https://www.onorca.dev/docs/telemetry) pour ce qu'Orca collecte en anonyme et comment désactiver la télémétrie.
diff --git a/docs/readme/README.ja.md b/docs/readme/README.ja.md
index a256f5239c4..cce2032a67c 100644
--- a/docs/readme/README.ja.md
+++ b/docs/readme/README.ja.md
@@ -12,7 +12,7 @@
@@ -238,9 +238,9 @@ yay -S stably-orca-bin
- **Discord:** **[Discord](https://discord.gg/fzjDKHxv8Q)** 커뮤니티에 참여하세요.
- **Twitter / X:** 업데이트와 공지는 **[@orca_build](https://x.com/orca_build)** 를 팔로우하세요.
-- **WeChat:** QR 코드를 스캔해 Orca 커뮤니티 WeChat 그룹 8에 참여하세요.
+- **WeChat:** QR 코드를 스캔해 Orca 커뮤니티 WeChat 그룹 8에 참여하세요. 그룹 8이 가득 찼을 수 있으니, 그런 경우 그룹 9 QR 코드를 스캔하세요.
-
+
- **피드백과 아이디어:** 우리는 빠르게 출시합니다. 필요한 기능이 있나요? [새 기능을 요청](https://github.com/stablyai/orca/issues)하세요.
- **개인정보 보호:** Orca가 수집하는 익명 사용 데이터와 수집 거부 방법은 [개인정보 및 텔레메트리 문서](https://www.onorca.dev/docs/telemetry)를 참고하세요.
diff --git a/docs/readme/README.pt.md b/docs/readme/README.pt.md
index 667cebb685f..86d998a4e5f 100644
--- a/docs/readme/README.pt.md
+++ b/docs/readme/README.pt.md
@@ -12,7 +12,7 @@
- AI-оркестратор для розробників рівня 100x.
- Запускайте Codex, Claude Code, OpenCode або Pi паралельно — кожен у власному worktree, усі під контролем в одному місці.
-
-
-### Супутній мобільний застосунок
-
-Стежте за агентами та керуйте ними з телефону — отримуйте сповіщення про завершення роботи агента та надсилайте подальші вказівки, де б ви не були.
-
-[App Store для iOS](https://apps.apple.com/us/app/orca-ide/id6766130217) · [TestFlight](https://testflight.apple.com/join/YjeGMQBA) · [Android APK 0.0.44](https://github.com/stablyai/orca/releases/download/mobile-android-v0.0.44/app-release.apk) · [Документація →](https://www.onorca.dev/docs/mobile)
-
-
-
-
-
-
-
-
-
-### Паралельні worktree
-
-Надішліть один промпт одразу п’ятьом агентам, кожен із яких працюватиме у власному ізольованому git worktree, — порівняйте результати та виконайте злиття найкращого з них.
-
-[Документація →](https://www.onorca.dev/docs/model/worktrees)
-
-
-
-
-
-
-
-
-
-### Розділені термінали
-
-Термінали рівня Ghostty з рендерингом на WebGL, необмеженою кількістю розділень і буфером прокручування, який зберігається після перезапуску.
-
-[Документація →](https://www.onorca.dev/docs/terminal)
-
-
-
-
-
-
-
-
-
-### Режим дизайну
-
-Клацніть на будь-якому елементі інтерфейсу у справжньому вікні Chromium, щоб надіслати його HTML, CSS і обрізаний скриншот прямо в промпт агента.
-
-[Документація →](https://www.onorca.dev/docs/browser/design-mode)
-
-
-
-
-
-
-
-
-
-### GitHub і Linear, нативно
-
-Переглядайте PR, issue та дошки проєктів прямо в застосунку — відкривайте worktree з будь-якої задачі та рев'юйте без перемикання контексту.
-
-[Документація →](https://www.onorca.dev/docs/review/linear)
-
-
-
-
-
-
-
-
-
-### SSH worktree
-
-Запускайте агентів на потужній віддаленій машині з повноцінним редагуванням файлів, git і терміналами — з автоперепідключенням і прокиданням портів.
-
-[Документація →](https://www.onorca.dev/docs/ssh)
-
-
-
-
-
-
-
-
-
-### Анотуйте diff-и агентів
-
-Залишайте коментарі на будь-якому рядку diff-у й надсилайте їх агенту — рев'юйте, редагуйте та комітьте, не виходячи з Orca.
-
-[Документація →](https://www.onorca.dev/docs/review/annotate-ai-diff)
-
-
-
-
-
-
-
-
-
-### Перетягуйте файли агентам
-
-Редактор на базі VS Code з автозбереженням усюди — перетягуйте файли чи зображення прямо в промпт агента.
-
-[Документація →](https://www.onorca.dev/docs/editing/file-explorer)
-
-
-
-
-
-
-
-
-
-### Orca CLI
-
-Агенти теж керують Orca — автоматизуйте будь-який робочий процес командами `orca worktree create`, `snapshot`, `click` і `fill`.
-
-[Документація →](https://www.onorca.dev/docs/cli/overview)
-
-
-
-
-
-
-
-
-**Також у комплекті:**
-
-- **[Швидкий пошук](https://www.onorca.dev/docs/model/quick-open)** — Шукайте серед worktree, файлів, агентів, команд і контексту репозиторію, не відриваючись від роботи.
-- **[Перемикач акаунтів і відстеження використання](https://www.onorca.dev/docs/agents/usage-tracking)** — Стежте за використанням Claude і Codex та скиданням лімітів, перемикайте акаунти на льоту без повторного входу.
-- **[Розширені перегляди репозиторію](https://www.onorca.dev/docs/editing/markdown)** — Переглядайте Markdown, зображення, PDF та документацію репозиторію прямо в робочому просторі.
-- **[Computer Use](https://www.onorca.dev/docs/cli/computer-use)** — Дозвольте агентам керувати десктопними застосунками та видимим інтерфейсом, коли робочий процес потребує реальної взаємодії.
-- **[Сповіщення та статус непрочитаного](https://www.onorca.dev/docs/notifications)** — Дізнавайтеся, коли агент завершив роботу або потребує уваги, і позначайте треди як непрочитані, щоб повернутися пізніше.
-- **І багато іншого** — ми випускаємо оновлення щодня, тож цей список завжди відстає. Справжній перелік можливостей — це [changelog](https://github.com/stablyai/orca/releases).
-
----
-
-## Підтримувані агенти
-
-Працює з **будь-яким CLI-агентом** — якщо він запускається в терміналі, він запуститься і в Orca.
-
-
-
----
-
-## Встановлення
-
-### Десктоп — macOS, Windows, Linux
-
-- **[Завантажити з onOrca.dev](https://onorca.dev/download)**
-- Або завантажте білд напряму: [macOS Apple Silicon](https://github.com/stablyai/orca/releases/latest/download/orca-macos-arm64.dmg) · [macOS Intel](https://github.com/stablyai/orca/releases/latest/download/orca-macos-x64.dmg) · [Windows (.exe)](https://github.com/stablyai/orca/releases/latest/download/orca-windows-setup.exe) · [Linux AppImage](https://github.com/stablyai/orca/releases/latest/download/orca-linux.AppImage) · [Усі білди](https://github.com/stablyai/orca/releases/latest)
-- Запускаєте `orca serve` на headless Linux-сервері? Дивіться [посібник із headless Linux-сервера](../reference/headless-linux-server.md).
-
-_Або через пакетний менеджер:_
-
-```bash
-# macOS (Homebrew)
-brew install --cask stablyai/orca/orca
-
-# Arch Linux (AUR) — або stably-orca-git для збірки з джерела
-yay -S stably-orca-bin
-```
-
-### Супутній мобільний застосунок — iOS, Android
-
-Під’єднайте мобільний застосунок до десктопного, щоб стежити за агентами та керувати ними з телефону.
-
-- **iOS:** [Завантажити з App Store](https://apps.apple.com/us/app/orca-ide/id6766130217) або [приєднатися до TestFlight](https://testflight.apple.com/join/YjeGMQBA)
-- **Android:** [Завантажити APK 0.0.44](https://github.com/stablyai/orca/releases/download/mobile-android-v0.0.44/app-release.apk) · [Інструкція зі встановлення](https://www.onorca.dev/docs/android-apk)
-
----
-
-## Спільнота та підтримка
-
-- **Discord:** Приєднуйтеся до спільноти в **[Discord](https://discord.gg/fzjDKHxv8Q)**.
-- **Twitter / X:** Стежте за **[@orca_build](https://x.com/orca_build)**, щоб бути в курсі оновлень і анонсів.
-- **WeChat:** Відскануйте QR-код, щоб приєднатися до групи № 7 спільноти Orca у WeChat. Якщо вона заповнена, приєднайтеся до групи № 8.
-
-
-
-
-- **Зворотний зв'язок та ідеї:** Ми випускаємо оновлення швидко. Чогось бракує? [Запропонуйте нову функцію](https://github.com/stablyai/orca/issues).
-- **Конфіденційність:** Перегляньте [документацію про конфіденційність і телеметрію](https://www.onorca.dev/docs/telemetry), щоб дізнатися, які анонімні дані про використання збирає Orca і як від цього відмовитися.
-- **Підтримайте нас:** Поставте [зірку](https://github.com/stablyai/orca) цьому репозиторію, щоб стежити за нашими щоденними релізами.
-
----
-
-## Розробка
-
-Хочете зробити внесок або запустити проєкт локально? Перегляньте наш посібник [CONTRIBUTING.md](../../.github/CONTRIBUTING.md).
-
-
-
-
-
-
-
-
-
-## Підписані білди
-Підписання коду для Windows надано за підтримки [SignPath.io](https://signpath.io), сертифікат надано [SignPath Foundation](https://signpath.org).
-
-## Ліцензія
-
-Orca — безкоштовний проєкт із відкритим кодом за ліцензією [MIT](../../LICENSE).
diff --git a/docs/readme/README.zh-CN.md b/docs/readme/README.zh-CN.md
index 160f5b05611..10f47e20fe6 100644
--- a/docs/readme/README.zh-CN.md
+++ b/docs/readme/README.zh-CN.md
@@ -12,7 +12,7 @@
@@ -235,9 +235,9 @@ yay -S stably-orca-bin
- **Discord:** 加入 **[Discord](https://discord.gg/fzjDKHxv8Q)** 社区。
- **Twitter / X:** 关注 **[@orca_build](https://x.com/orca_build)** 获取更新和公告。
-- **微信:** 扫码加入 Orca 社区微信第 8 群。
+- **微信:** 扫码加入 Orca 社区微信第 8 群。第 8 群可能已满,如遇这种情况请扫描第 9 群二维码。
-
+
- **反馈与想法:** 我们发布很快。缺少什么功能?[提交功能请求](https://github.com/stablyai/orca/issues)。
- **隐私:** 查看[隐私与遥测文档](https://www.onorca.dev/docs/telemetry),了解 Orca 收集哪些匿名使用数据以及如何退出。
diff --git a/mobile/src/home/MobileHomeHostList.tsx b/mobile/src/home/MobileHomeHostList.tsx
index 3907df16f03..41d1f07156d 100644
--- a/mobile/src/home/MobileHomeHostList.tsx
+++ b/mobile/src/home/MobileHomeHostList.tsx
@@ -19,6 +19,7 @@ type MobileHomeHostListProps = {
hostAttempts: Record
hostLastConnected: Record
hostPairingRejected: Record
+ hostSignedOut: Record
hostPaths: Record
hostPendingPaths: Record
hosts: HostCatalogEntry[]
@@ -40,6 +41,7 @@ export function MobileHomeHostList(props: MobileHomeHostListProps) {
hostAttempts={props.hostAttempts}
hostLastConnected={props.hostLastConnected}
hostPairingRejected={props.hostPairingRejected}
+ hostSignedOut={props.hostSignedOut}
hostPaths={props.hostPaths}
hostPendingPaths={props.hostPendingPaths}
hostStates={props.hostStates}
@@ -54,6 +56,7 @@ export function MobileHomeHostList(props: MobileHomeHostListProps) {
props.hostAttempts,
props.hostLastConnected,
props.hostPairingRejected,
+ props.hostSignedOut,
props.hostPaths,
props.hostPendingPaths,
props.hostStates,
@@ -91,6 +94,7 @@ type MobileHomeHostRowProps = Pick<
| 'hostAttempts'
| 'hostLastConnected'
| 'hostPairingRejected'
+ | 'hostSignedOut'
| 'hostPaths'
| 'hostPendingPaths'
| 'hostStates'
@@ -113,7 +117,8 @@ const MobileHomeHostRow = memo(function MobileHomeHostRow(props: MobileHomeHostR
lastConnectedAt: props.hostLastConnected[item.id] ?? null,
endpoint: item.endpoint,
pendingPath: props.hostPendingPaths[item.id] ?? null,
- pairingRejected: props.hostPairingRejected[item.id] ?? false
+ pairingRejected: props.hostPairingRejected[item.id] ?? false,
+ hostSignedOut: props.hostSignedOut[item.id] ?? false
})
const open = useCallback(() => onOpen(item), [item, onOpen])
const longPress = useCallback(() => onLongPress(item), [item, onLongPress])
diff --git a/mobile/src/home/MobileHomeScreen.tsx b/mobile/src/home/MobileHomeScreen.tsx
index 83cf3de4b8a..7772f5152c6 100644
--- a/mobile/src/home/MobileHomeScreen.tsx
+++ b/mobile/src/home/MobileHomeScreen.tsx
@@ -143,6 +143,7 @@ export function MobileHomeScreen() {
hostAttempts={data.hostAttempts}
hostLastConnected={data.hostLastConnected}
hostPairingRejected={data.hostPairingRejected}
+ hostSignedOut={data.hostSignedOut}
hostPaths={data.hostPaths}
hostPendingPaths={data.hostPendingPaths}
hosts={data.sortedHostCatalog}
diff --git a/mobile/src/home/home-host-connection-projection.ts b/mobile/src/home/home-host-connection-projection.ts
index f8fdfd4bcdf..9a49f6186c4 100644
--- a/mobile/src/home/home-host-connection-projection.ts
+++ b/mobile/src/home/home-host-connection-projection.ts
@@ -5,12 +5,14 @@ export type HomeHostConnectionProjectionEntry = {
path: MobileConnectionPath
pendingPath: MobileConnectionPath | null
pairingRejected: boolean
+ hostSignedOut: boolean
}
export type HomeHostConnectionProjection = {
hostPaths: Record
hostPendingPaths: Record
hostPairingRejected: Record
+ hostSignedOut: Record
}
/** Build all host lookup maps while reading each connection entry once. */
@@ -22,16 +24,19 @@ export function projectHomeHostConnections(
const hostPaths = Object.create(null) as Record
const hostPendingPaths = Object.create(null) as Record
const hostPairingRejected = Object.create(null) as Record
+ const hostSignedOut = Object.create(null) as Record
- for (const { hostId, path, pendingPath, pairingRejected } of entries) {
+ for (const { hostId, path, pendingPath, pairingRejected, hostSignedOut: signedOut } of entries) {
hostPaths[hostId] = path
hostPendingPaths[hostId] = pendingPath
hostPairingRejected[hostId] = pairingRejected
+ hostSignedOut[hostId] = signedOut
}
Object.setPrototypeOf(hostPaths, Object.prototype)
Object.setPrototypeOf(hostPendingPaths, Object.prototype)
Object.setPrototypeOf(hostPairingRejected, Object.prototype)
+ Object.setPrototypeOf(hostSignedOut, Object.prototype)
- return { hostPaths, hostPendingPaths, hostPairingRejected }
+ return { hostPaths, hostPendingPaths, hostPairingRejected, hostSignedOut }
}
diff --git a/mobile/src/home/use-mobile-home-data.ts b/mobile/src/home/use-mobile-home-data.ts
index c77d024158e..6b28354a86d 100644
--- a/mobile/src/home/use-mobile-home-data.ts
+++ b/mobile/src/home/use-mobile-home-data.ts
@@ -179,6 +179,7 @@ export function useMobileHomeData() {
connectedHosts,
hostCatalog,
hostPairingRejected: hostConnectionProjection.hostPairingRejected,
+ hostSignedOut: hostConnectionProjection.hostSignedOut,
hostPaths: hostConnectionProjection.hostPaths,
hostPendingPaths: hostConnectionProjection.hostPendingPaths,
primaryHost,
diff --git a/mobile/src/session/mobile-image-attachment.test.ts b/mobile/src/session/mobile-image-attachment.test.ts
index 9d725d5fe60..eead9691303 100644
--- a/mobile/src/session/mobile-image-attachment.test.ts
+++ b/mobile/src/session/mobile-image-attachment.test.ts
@@ -50,7 +50,9 @@ describe('attachMobileImageToTerminal', () => {
const sendCall = client.calls.find((c) => c.method === 'terminal.send')
expect(sendCall?.params).toEqual({
terminal: 'term-1',
- text: '\x1b[200~/tmp/orca-attach.png\x1b[201~',
+ // Trailing space: the user types on this same line next, so a bare
+ // `…\x1b[201~` would arrive as `…pngadd` (STA-4847).
+ text: '\x1b[200~/tmp/orca-attach.png\x1b[201~ ',
enter: false,
client: { id: 'device-9', type: 'mobile' }
})
diff --git a/mobile/src/session/mobile-image-attachment.ts b/mobile/src/session/mobile-image-attachment.ts
index 567be99a9a1..9cb7d60aa8e 100644
--- a/mobile/src/session/mobile-image-attachment.ts
+++ b/mobile/src/session/mobile-image-attachment.ts
@@ -1,4 +1,5 @@
import type { RpcClient } from '../transport/rpc-client'
+import { separateImagePasteFromFollowingText } from '../../../src/shared/image-paste-following-text'
import {
buildMobileImagePastePayload,
saveMobileClipboardImageAsTempFile
@@ -47,7 +48,10 @@ export async function attachMobileImageToTerminal(
})
// Why: a generated image path is terminal image injection, so it's always
// bracketed (matching desktop paste) regardless of terminal mode.
- const payload = buildMobileImagePastePayload(imagePath)
+ // Always separated: attach-then-type is the whole interaction here, so the user's
+ // next keystroke would otherwise glue onto the path (`…pngadd`). Unlike native
+ // chat there is no batch to look ahead in, and a trailing space is inert.
+ const payload = separateImagePasteFromFollowingText(buildMobileImagePastePayload(imagePath), true)
if (beforeTerminalSend && !(await beforeTerminalSend(terminal))) {
return false
}
diff --git a/mobile/src/session/mobile-native-chat-image-send.test.ts b/mobile/src/session/mobile-native-chat-image-send.test.ts
index cf1c59adf0f..a41b3fca3f1 100644
--- a/mobile/src/session/mobile-native-chat-image-send.test.ts
+++ b/mobile/src/session/mobile-native-chat-image-send.test.ts
@@ -38,7 +38,8 @@ describe('pasteMobileNativeChatImagePaths', () => {
client,
terminal: 'term-1',
deviceToken: 'device-9',
- imagePaths: ['/tmp/a.png', '/tmp/b.png', '/tmp/c.png']
+ imagePaths: ['/tmp/a.png', '/tmp/b.png', '/tmp/c.png'],
+ followedByText: true
})
expect(ok).toBe(true)
@@ -55,7 +56,7 @@ describe('pasteMobileNativeChatImagePaths', () => {
})
expect(client.calls[1]?.params.text).toBe('\x1b[200~/tmp/a.png\x1b[201~')
expect(client.calls[2]?.params.text).toBe('\x1b[200~/tmp/b.png\x1b[201~')
- expect(client.calls[3]?.params.text).toBe('\x1b[200~/tmp/c.png\x1b[201~')
+ expect(client.calls[3]?.params.text).toBe('\x1b[200~/tmp/c.png\x1b[201~ ')
})
it('stops and reports failure as soon as a paste is rejected', async () => {
@@ -66,7 +67,8 @@ describe('pasteMobileNativeChatImagePaths', () => {
client,
terminal: 'term-1',
deviceToken: null,
- imagePaths: ['/tmp/a.png', '/tmp/b.png']
+ imagePaths: ['/tmp/a.png', '/tmp/b.png'],
+ followedByText: true
})
expect(ok).toBe(false)
@@ -94,7 +96,8 @@ describe('pasteMobileNativeChatImagePaths', () => {
client,
terminal: 'term-1',
deviceToken: null,
- imagePaths: ['/tmp/a.png', '/tmp/b.png']
+ imagePaths: ['/tmp/a.png', '/tmp/b.png'],
+ followedByText: true
})
expect(ok).toBe(false)
@@ -121,6 +124,7 @@ describe('clearing a parked multi-line launch draft before the image paste', ()
terminal: 'term-1',
deviceToken: null,
imagePaths: ['/tmp/a.png'],
+ followedByText: true,
clearInput
})
@@ -137,6 +141,7 @@ describe('clearing a parked multi-line launch draft before the image paste', ()
terminal: 'term-1',
deviceToken: null,
imagePaths: ['/tmp/a.png', '/tmp/b.png'],
+ followedByText: true,
clearInput
})
@@ -151,9 +156,27 @@ describe('clearing a parked multi-line launch draft before the image paste', ()
client,
terminal: 'term-1',
deviceToken: null,
- imagePaths: ['/tmp/a.png']
+ imagePaths: ['/tmp/a.png'],
+ followedByText: true
})
expect(client.calls[0]?.params.text).toBe('\x15')
})
+
+ it('keeps image writes byte-clean when no text or submit follows', async () => {
+ const client = clientWithResponses([sendResult(true), sendResult(true), sendResult(true)])
+
+ await pasteMobileNativeChatImagePaths({
+ client,
+ terminal: 'term-1',
+ deviceToken: null,
+ imagePaths: ['/tmp/a.png', '/tmp/b.png'],
+ followedByText: false
+ })
+
+ expect(client.calls.slice(1).map((call) => call.params.text)).toEqual([
+ '\x1b[200~/tmp/a.png\x1b[201~',
+ '\x1b[200~/tmp/b.png\x1b[201~'
+ ])
+ })
})
diff --git a/mobile/src/session/mobile-native-chat-image-send.ts b/mobile/src/session/mobile-native-chat-image-send.ts
index adb996b7612..9f3c555062f 100644
--- a/mobile/src/session/mobile-native-chat-image-send.ts
+++ b/mobile/src/session/mobile-native-chat-image-send.ts
@@ -1,4 +1,5 @@
import type { RpcClient } from '../transport/rpc-client'
+import { imagePasteWritesFollowedByText } from '../../../src/shared/image-paste-following-text'
import { buildMobileImagePastePayload } from './mobile-clipboard-image'
import {
MOBILE_NATIVE_CHAT_MIN_WRITE_TIMEOUT_MS,
@@ -23,6 +24,7 @@ type PasteImagesArgs = {
readonly terminal: string
readonly deviceToken: string | null
readonly imagePaths: readonly string[]
+ readonly followedByText: boolean
/** Budget shared with the rest of the user action (the text body that follows, or
* the send this is healing for). Omit to open a fresh one for this paste alone. */
readonly deadline?: number
@@ -42,6 +44,7 @@ export async function pasteMobileNativeChatImagePaths({
terminal,
deviceToken,
imagePaths,
+ followedByText,
deadline: sharedDeadline,
clearInput
}: PasteImagesArgs): Promise {
@@ -55,7 +58,7 @@ export async function pasteMobileNativeChatImagePaths({
const deadline = sharedDeadline ?? openMobileNativeChatSendBudget()
for (const text of [
clearInput ?? MOBILE_NATIVE_CHAT_CLEAR_UNSUBMITTED_INPUT,
- ...imagePaths.map(buildMobileImagePastePayload)
+ ...imagePasteWritesFollowedByText(imagePaths.map(buildMobileImagePastePayload), followedByText)
]) {
const remainingMs = deadline - Date.now()
// Why: the budget is the whole sequence's — starting a write it can't fund would
diff --git a/mobile/src/session/mobile-native-chat-stale-input.ts b/mobile/src/session/mobile-native-chat-stale-input.ts
index 18d84d2e503..cdc21067683 100644
--- a/mobile/src/session/mobile-native-chat-stale-input.ts
+++ b/mobile/src/session/mobile-native-chat-stale-input.ts
@@ -55,6 +55,7 @@ export async function healMobileNativeChatStaleInput(args: {
terminal: args.terminal,
deviceToken: args.deviceToken,
imagePaths: [],
+ followedByText: false,
...(args.deadline === undefined ? {} : { deadline: args.deadline })
})
} catch {
diff --git a/mobile/src/session/use-mobile-native-chat-image-attachments.test.ts b/mobile/src/session/use-mobile-native-chat-image-attachments.test.ts
index cb8522bae7a..022bb8c3953 100644
--- a/mobile/src/session/use-mobile-native-chat-image-attachments.test.ts
+++ b/mobile/src/session/use-mobile-native-chat-image-attachments.test.ts
@@ -182,9 +182,12 @@ describe('useMobileNativeChatImageAttachments', () => {
expect(sendCalls).toHaveLength(2)
expect(sendCalls[0]?.params).toMatchObject({ text: '\x15', enter: false })
expect(sendCalls[1]?.params).toMatchObject({
- text: '\x1b[200~/tmp/a.png\x1b[201~',
+ text: '\x1b[200~/tmp/a.png\x1b[201~ ',
enter: false
})
+ const combined = String(sendCalls[1]?.params.text ?? '') + 'look at this'
+ expect(combined).toContain('.png\x1b[201~ look')
+ expect(combined).not.toContain('.png\x1b[201~look')
// Clear, then paste, then settle, then the text send — in that order.
expect(order).toEqual(['clear', 'paste', 'settle', 'text:look at this'])
// The local preview URI rides along so the sent bubble shows the photo.
@@ -267,7 +270,10 @@ describe('useMobileNativeChatImageAttachments', () => {
}
})
- it('routes an attachments-only send through baseSend with empty text so the echo still shows the photo', async () => {
+ it.each([
+ ['empty', ''],
+ ['whitespace-only', ' ']
+ ])('routes an attachments-only send through baseSend with %s text', async (_label, text) => {
pick.mockResolvedValue([{ base64: 'AAAA', uri: 'file:///a.jpg' }])
const client = makeClient([
methodNotFound('start'),
@@ -283,16 +289,17 @@ describe('useMobileNativeChatImageAttachments', () => {
})
let accepted = false
await act(async () => {
- accepted = await hook!.sendNativeChat('')
+ accepted = await hook!.sendNativeChat(text)
})
expect(accepted).toBe(true)
- // Empty text still goes through baseSend (which submits the bare Enter) so the
+ // Attachment-only text still goes through baseSend (which submits Enter) so the
// optimistic echo carries the preview URI.
- expect(baseSend).toHaveBeenCalledWith('', ['file:///a.jpg'], expect.any(Number))
+ expect(baseSend).toHaveBeenCalledWith(text, ['file:///a.jpg'], expect.any(Number))
const sendCalls = client.calls.filter((c) => c.method === 'terminal.send')
// Only the clear + image paste hit the wire here; baseSend owns the submit.
expect(sendCalls).toHaveLength(2)
+ expect(sendCalls[1]?.params.text).toBe('\x1b[200~/tmp/a.png\x1b[201~')
expect(hook!.attachments).toEqual([])
})
diff --git a/mobile/src/session/use-mobile-native-chat-image-attachments.ts b/mobile/src/session/use-mobile-native-chat-image-attachments.ts
index 77d839adf34..c36e30f44d7 100644
--- a/mobile/src/session/use-mobile-native-chat-image-attachments.ts
+++ b/mobile/src/session/use-mobile-native-chat-image-attachments.ts
@@ -240,6 +240,7 @@ export function useMobileNativeChatImageAttachments({
terminal: handle,
deviceToken: deviceTokenRef.current,
imagePaths: pendingImages.map((attachment) => attachment.path),
+ followedByText: text.trim().length > 0,
deadline,
...(seededLaunchDraft
? { clearInput: buildAgentTuiClearInputForText(seededLaunchDraft) }
diff --git a/mobile/src/transport/client-context-connection-metrics.ts b/mobile/src/transport/client-context-connection-metrics.ts
index 1e2ff6f7919..1fe72687bd2 100644
--- a/mobile/src/transport/client-context-connection-metrics.ts
+++ b/mobile/src/transport/client-context-connection-metrics.ts
@@ -30,14 +30,16 @@ export function useConnectionPathStatus(hostId: string | undefined): {
export function useRelayRecoveryStatus(hostId: string | undefined): {
pendingPath: MobileConnectionPath | null
pairingRejected: boolean
+ hostSignedOut: boolean
} {
return useHostMetric(
hostId,
(context, id) => ({
pendingPath: context.getPendingPath(id),
- pairingRejected: context.isPairingRejected(id)
+ pairingRejected: context.isPairingRejected(id),
+ hostSignedOut: context.isHostSignedOut(id)
}),
- { pendingPath: null, pairingRejected: false }
+ { pendingPath: null, pairingRejected: false, hostSignedOut: false }
)
}
diff --git a/mobile/src/transport/client-context.test.ts b/mobile/src/transport/client-context.test.ts
index 4a7d5d7b0e8..56b227bdff7 100644
--- a/mobile/src/transport/client-context.test.ts
+++ b/mobile/src/transport/client-context.test.ts
@@ -533,12 +533,12 @@ describe('useAllHostClients', () => {
await Promise.resolve()
})
act(() => client.emitPendingPath('relay'))
- expect(status).toEqual({ pendingPath: 'relay', pairingRejected: false })
+ expect(status).toEqual({ pendingPath: 'relay', pairingRejected: false, hostSignedOut: false })
// Why: the desktop refusing the credential is a status-only change — no
// transport state moves, so only the connection-path signal can carry it.
act(() => client.emitPairingRejected(true))
- expect(status).toEqual({ pendingPath: 'relay', pairingRejected: true })
+ expect(status).toEqual({ pendingPath: 'relay', pairingRejected: true, hostSignedOut: false })
act(() => renderer.unmount())
})
diff --git a/mobile/src/transport/connection-health.ts b/mobile/src/transport/connection-health.ts
index 858b13a8b24..1a9e282047f 100644
--- a/mobile/src/transport/connection-health.ts
+++ b/mobile/src/transport/connection-health.ts
@@ -29,6 +29,10 @@ const STALE_SINCE_LAST_CONNECT_MS = 60_000
// instead of leaving the user staring at a generic "Can't connect".
const TAILSCALE_HINT = 'check Tailscale'
+// No hint field: the remedy is the label, and appending "— check Tailscale" to
+// it would be wrong advice for a desktop that is reachable but signed out.
+const SIGNED_OUT_LABEL = 'Desktop signed out — sign in to Orca on your desktop to reconnect'
+
export type ConnectionVerdict =
| { kind: 'normal'; label: string }
| { kind: 'warning'; label: string; hint?: string } // "Can't connect"
@@ -54,6 +58,10 @@ export function classifyConnection(args: {
// The desktop has repeatedly refused this device's relay credential — retrying
// cannot fix it, so it outranks any "still connecting" reading (STA-4681).
pairingRejected?: boolean
+ // The relay says the desktop's last control close named its own Orca Cloud
+ // sign-out. Retrying is still correct and still happens on the same cadence,
+ // but only the desktop's owner can end it, so the label has to say so.
+ hostSignedOut?: boolean
nowMs?: number
}): ConnectionVerdict {
const { state, reconnectAttempts, lastConnectedAt } = args
@@ -70,6 +78,17 @@ export function classifyConnection(args: {
return { kind: 'normal', label: 'Connected' }
}
+ // Ahead of the attempt thresholds: this is evidence, not an inference from a
+ // failure streak, and waiting twelve dials to show it wastes the whole point.
+ // Below auth-failed because a revoked pairing cannot be fixed by signing in.
+ if (args.hostSignedOut) {
+ return {
+ kind: 'unreachable',
+ label: SIGNED_OUT_LABEL,
+ reason: lastConnectedAt == null ? 'never-connected' : 'stale'
+ }
+ }
+
// A disconnected pending path can survive a cleared retry timer during a
// lifecycle race. Only narrate Relay while dialing or after a retry has
// recorded progress; otherwise the idle transport must read Disconnected.
diff --git a/mobile/src/transport/host-client-context-state.ts b/mobile/src/transport/host-client-context-state.ts
index 859e5d847f9..db4c7908b41 100644
--- a/mobile/src/transport/host-client-context-state.ts
+++ b/mobile/src/transport/host-client-context-state.ts
@@ -89,10 +89,16 @@ export function createHostClientSelectors(
getPendingPath: (hostId: string): MobileConnectionPath | null =>
clientPendingPath(entries.get(hostId)?.client),
isPairingRejected: (hostId: string): boolean =>
- clientPairingRejected(entries.get(hostId)?.client)
+ clientPairingRejected(entries.get(hostId)?.client),
+ isHostSignedOut: (hostId: string): boolean => clientHostSignedOut(entries.get(hostId)?.client)
}
}
+export function clientHostSignedOut(client: RpcClient | undefined): boolean {
+ const logical = client as Partial | undefined
+ return logical?.isHostSignedOut?.() ?? false
+}
+
export function clientPairingRejected(client: RpcClient | undefined): boolean {
const logical = client as Partial | undefined
return logical?.isPairingRejected?.() ?? false
diff --git a/mobile/src/transport/logical-client-connection-path.ts b/mobile/src/transport/logical-client-connection-path.ts
index 55c6d1d28f3..b7f02b40b88 100644
--- a/mobile/src/transport/logical-client-connection-path.ts
+++ b/mobile/src/transport/logical-client-connection-path.ts
@@ -5,6 +5,7 @@ export class LogicalClientConnectionPath {
private recovery: MobileConnectionPath | null = null
private recoveryAttempt = 0
private pairingRejected = false
+ private hostSignedOut = false
private readonly listeners = new Set<() => void>()
constructor(private readonly isConnected: () => boolean) {}
@@ -35,12 +36,23 @@ export class LogicalClientConnectionPath {
})
}
+ isHostSignedOut(): boolean {
+ return this.hostSignedOut
+ }
+
+ setHostSignedOut(signedOut: boolean): void {
+ this.update(() => {
+ this.hostSignedOut = signedOut
+ })
+ }
+
clearAfterConnected(): void {
this.migration = null
this.recovery = null
this.recoveryAttempt = 0
// Why: an authenticated session is the desktop accepting this device.
this.pairingRejected = false
+ this.hostSignedOut = false
}
setRecovery(path: MobileConnectionPath | null, attempt?: number): void {
@@ -69,11 +81,13 @@ export class LogicalClientConnectionPath {
const previousPath = this.pending()
const previousAttempt = this.reconnectAttempt(0)
const previousRejected = this.pairingRejected
+ const previousSignedOut = this.hostSignedOut
apply()
if (
previousPath === this.pending() &&
previousAttempt === this.reconnectAttempt(0) &&
- previousRejected === this.pairingRejected
+ previousRejected === this.pairingRejected &&
+ previousSignedOut === this.hostSignedOut
) {
return
}
diff --git a/mobile/src/transport/mobile-endpoint-lifecycle.ts b/mobile/src/transport/mobile-endpoint-lifecycle.ts
index 8ee8df6948e..7ec5f28b945 100644
--- a/mobile/src/transport/mobile-endpoint-lifecycle.ts
+++ b/mobile/src/transport/mobile-endpoint-lifecycle.ts
@@ -86,7 +86,7 @@ function createSupervisor(
): MobileEndpointSupervisor {
return new MobileEndpointSupervisor(logical, host, {
openDirect: (endpoint) => connect(endpoint, host.deviceToken, host.publicKeyB64, { onLog }),
- openRelay: (relay, credential, confirmReqId) =>
+ openRelay: (relay, credential, confirmReqId, onHostCloseReason) =>
connectMobileRelayRpcSession({
relay,
resumeToken: credential.token,
@@ -94,6 +94,7 @@ function createSupervisor(
resumeConfirmReqId: confirmReqId,
deviceToken: host.deviceToken,
desktopPublicKeyB64: host.publicKeyB64,
+ onHostCloseReason,
onLog
}),
resolveRelay: resolveMobileRelayEndpoint,
diff --git a/mobile/src/transport/mobile-endpoint-supervisor-contract.ts b/mobile/src/transport/mobile-endpoint-supervisor-contract.ts
index 0098de6e079..2a784fd8895 100644
--- a/mobile/src/transport/mobile-endpoint-supervisor-contract.ts
+++ b/mobile/src/transport/mobile-endpoint-supervisor-contract.ts
@@ -1,4 +1,5 @@
import type { MobileRelayEndpoint } from '../../../src/shared/mobile-relay-credential-contract'
+import type { RelayHostCloseReason } from '../../../src/shared/relay-host-close-reason'
import type { MobileRelayCredentialBundle } from './mobile-relay-credential-bundle'
import type { MobileRelayRpcSession } from './mobile-relay-rpc-session'
import type { resolveMobileRelayEndpoint } from './mobile-relay-resume-director'
@@ -10,7 +11,8 @@ export type MobileEndpointSupervisorDependencies = {
openRelay: (
relay: MobileRelayEndpoint,
credential: { token: string; version: number },
- confirmReqId: string
+ confirmReqId: string,
+ onHostCloseReason?: (reason: RelayHostCloseReason) => void
) => MobileRelayRpcSession
resolveRelay: typeof resolveMobileRelayEndpoint
readBundle: (hostId: string) => Promise
diff --git a/mobile/src/transport/mobile-endpoint-supervisor-test-fakes.ts b/mobile/src/transport/mobile-endpoint-supervisor-test-fakes.ts
index 1dc1473d9db..80f4438c160 100644
--- a/mobile/src/transport/mobile-endpoint-supervisor-test-fakes.ts
+++ b/mobile/src/transport/mobile-endpoint-supervisor-test-fakes.ts
@@ -134,10 +134,22 @@ export class FakeLogicalClient extends FakeSession implements StableLogicalRpcCl
}
})
isPairingRejected = () => this.pairingRejected
+ private hostSignedOut = false
+ setHostSignedOut = vi.fn((signedOut: boolean) => {
+ if (this.hostSignedOut === signedOut) {
+ return
+ }
+ this.hostSignedOut = signedOut
+ for (const listener of this.pathListeners) {
+ listener()
+ }
+ })
+ isHostSignedOut = () => this.hostSignedOut
// Mirrors LogicalClientConnectionPath.clearAfterConnected.
publishState(state: ConnectionState): void {
if (state === 'connected') {
this.pairingRejected = false
+ this.hostSignedOut = false
}
super.publishState(state)
}
diff --git a/mobile/src/transport/mobile-endpoint-supervisor.test.ts b/mobile/src/transport/mobile-endpoint-supervisor.test.ts
index aeb9cddef63..10ef892a479 100644
--- a/mobile/src/transport/mobile-endpoint-supervisor.test.ts
+++ b/mobile/src/transport/mobile-endpoint-supervisor.test.ts
@@ -185,7 +185,12 @@ describe('mobile endpoint supervisor', () => {
await supervisor.start()
expect(deps.resolveRelay).toHaveBeenCalledOnce()
- expect(openRelay).toHaveBeenLastCalledWith(resolved, expect.any(Object), expect.any(String))
+ expect(openRelay).toHaveBeenLastCalledWith(
+ resolved,
+ expect.any(Object),
+ expect.any(String),
+ expect.any(Function)
+ )
expect(deps.saveHost).toHaveBeenCalledWith(
expect.objectContaining({ relay: resolved, endpoint: host.endpoint })
)
@@ -556,7 +561,8 @@ describe('mobile endpoint supervisor', () => {
expect(openRelay).toHaveBeenLastCalledWith(
relay,
expect.objectContaining({ version: 3 }),
- expect.any(String)
+ expect.any(String),
+ expect.any(Function)
)
supervisor.stop()
})
@@ -603,7 +609,8 @@ describe('mobile endpoint supervisor', () => {
expect(openRelay).toHaveBeenLastCalledWith(
relay,
expect.objectContaining({ version: 3 }),
- expect.any(String)
+ expect.any(String),
+ expect.any(Function)
)
supervisor.stop()
})
diff --git a/mobile/src/transport/mobile-relay-e2ee-link.ts b/mobile/src/transport/mobile-relay-e2ee-link.ts
index 7743deb23dd..9b1f7a9a355 100644
--- a/mobile/src/transport/mobile-relay-e2ee-link.ts
+++ b/mobile/src/transport/mobile-relay-e2ee-link.ts
@@ -2,6 +2,10 @@ import {
RelayPhoneHelloSchema,
type RelayPhoneHello
} from '../../../src/shared/mobile-relay-phone-protocol'
+import {
+ relayHostCloseReasonFrom,
+ type RelayHostCloseReason
+} from '../../../src/shared/relay-host-close-reason'
import { MobileE2EEV2ClientSession } from './mobile-e2ee-v2-client-session'
import { MobileE2EEV2PhysicalChannel } from './mobile-e2ee-v2-physical-channel'
import { websocketPayloadToUint8 } from './websocket-payload-bytes'
@@ -26,6 +30,12 @@ type MobileRelayE2eeLinkOptions = {
onText: (plaintext: string) => void
onBinary: (plaintext: Uint8Array) => void
onHello?: (hello: Extract) => void
+ // The cell's account of why the desktop is absent, read off the close frame.
+ // Reported separately from onError because a rejection is delivered as both a
+ // relay-hello and a close, and which one the runtime dispatches first is not
+ // ordered — only the close carries the reason, and it must not be lost to
+ // that race.
+ onHostCloseReason?: (reason: RelayHostCloseReason) => void
// Fired once relay-auth is on the wire: from here the cell owns the wait.
onOpen?: () => void
onError: (error: Error) => void
@@ -129,6 +139,11 @@ export class MobileRelayE2eeLink {
clearTimeout(this.transportErrorTimer)
this.transportErrorTimer = null
}
+ // Ahead of fail(), which no-ops once the hello already reported this close.
+ const hostCloseReason = relayHostCloseReasonFrom(event.reason)
+ if (hostCloseReason) {
+ this.options.onHostCloseReason?.(hostCloseReason)
+ }
this.fail(new RelayOuterError(event.code || 1006))
}
}
diff --git a/mobile/src/transport/mobile-relay-rpc-session.ts b/mobile/src/transport/mobile-relay-rpc-session.ts
index 947a1d23ce8..203a0329192 100644
--- a/mobile/src/transport/mobile-relay-rpc-session.ts
+++ b/mobile/src/transport/mobile-relay-rpc-session.ts
@@ -13,6 +13,7 @@ import { RelayDialStageTracker, type RelayDialStageSource } from './relay-dial-s
import { RelayPendingRequests } from './relay-pending-requests'
import { RpcSessionLivenessWatchdog } from './rpc-session-liveness-watchdog'
import { settleMobileRuntimeCapabilities } from './mobile-runtime-capability-negotiation'
+import type { RelayHostCloseReason } from '../../../src/shared/relay-host-close-reason'
import type { RpcClient } from './rpc-client'
import type { ConnectionLogSink, ConnectionState, RpcResponse } from './types'
@@ -40,6 +41,7 @@ export function connectMobileRelayRpcSession(args: {
desktopPublicKeyB64: string
requestTimeoutMs?: number
createSocket?: (url: string) => WebSocket
+ onHostCloseReason?: (reason: RelayHostCloseReason) => void
onLog?: ConnectionLogSink
}): MobileRelayRpcSession {
const requestTimeoutMs = args.requestTimeoutMs ?? 30_000
@@ -69,6 +71,7 @@ export function connectMobileRelayRpcSession(args: {
deviceToken: args.deviceToken,
desktopPublicKeyB64: args.desktopPublicKeyB64,
createSocket: args.createSocket,
+ onHostCloseReason: args.onHostCloseReason,
onOpen: () => dialStage.advance('awaiting-hello'),
onHello: (hello) => {
if (
diff --git a/mobile/src/transport/mobile-relay-runtime-failover.test.ts b/mobile/src/transport/mobile-relay-runtime-failover.test.ts
index 01f4d45feb0..ce7cca3fd9f 100644
--- a/mobile/src/transport/mobile-relay-runtime-failover.test.ts
+++ b/mobile/src/transport/mobile-relay-runtime-failover.test.ts
@@ -145,10 +145,22 @@ class FakeLogicalClient extends FakeSession implements StableLogicalRpcClient {
}
})
isPairingRejected = () => this.pairingRejected
+ private hostSignedOut = false
+ setHostSignedOut = vi.fn((signedOut: boolean) => {
+ if (this.hostSignedOut === signedOut) {
+ return
+ }
+ this.hostSignedOut = signedOut
+ for (const listener of this.pathListeners) {
+ listener()
+ }
+ })
+ isHostSignedOut = () => this.hostSignedOut
// Mirrors LogicalClientConnectionPath.clearAfterConnected.
publishState(state: ConnectionState): void {
if (state === 'connected') {
this.pairingRejected = false
+ this.hostSignedOut = false
}
super.publishState(state)
}
@@ -264,7 +276,8 @@ describe('relay runtime recovery without direct connectivity', () => {
expect(openRelay).toHaveBeenLastCalledWith(
relay,
expect.objectContaining({ version: 3 }),
- expect.any(String)
+ expect.any(String),
+ expect.any(Function)
)
expect(logical.getActivePath()).toBe('relay')
supervisor.stop()
@@ -353,7 +366,8 @@ describe('relay runtime recovery without direct connectivity', () => {
expect(deps.openRelay).toHaveBeenLastCalledWith(
relay,
expect.objectContaining({ version: 2 }),
- expect.any(String)
+ expect.any(String),
+ expect.any(Function)
)
expect(logical.getActivePath()).toBe('relay')
supervisor.stop()
@@ -382,7 +396,8 @@ describe('relay runtime recovery without direct connectivity', () => {
expect(openRelay).toHaveBeenLastCalledWith(
relay,
expect.objectContaining({ version: 1 }),
- expect.any(String)
+ expect.any(String),
+ expect.any(Function)
)
expect(logical.getActivePath()).toBe('relay')
supervisor.stop()
diff --git a/mobile/src/transport/mobile-relay-session-establisher.ts b/mobile/src/transport/mobile-relay-session-establisher.ts
index 7a8ce372155..9a04ae44137 100644
--- a/mobile/src/transport/mobile-relay-session-establisher.ts
+++ b/mobile/src/transport/mobile-relay-session-establisher.ts
@@ -10,6 +10,7 @@ import type { MobileRelayCredentialBundle } from './mobile-relay-credential-bund
import type { RelayReconnectController } from './mobile-relay-reconnect-controller'
import type { StableLogicalRpcClient } from './stable-logical-rpc-client'
import type { MobileRelayEndpoint } from '../../../src/shared/mobile-relay-credential-contract'
+import { RELAY_HOST_CLOSE_REASON } from '../../../src/shared/relay-host-close-reason'
import type { HostProfile } from './types'
type EstablishResult = { ok: true } | { ok: false; error: Error }
@@ -100,7 +101,16 @@ export class MobileRelaySessionEstablisher {
const session = args.openRelay(
relay,
credential,
- `confirm-${encodeBase64Url(args.randomBytes(16))}`
+ `confirm-${encodeBase64Url(args.randomBytes(16))}`,
+ // Latched on the logical client, not on the dial result: the close that
+ // carries the reason can land after this dial has already reported its
+ // failure. Clearing is clearAfterConnected's job, so any path that
+ // reaches connected retires it.
+ (reason) => {
+ if (reason === RELAY_HOST_CLOSE_REASON.SIGNED_OUT) {
+ args.logical.setHostSignedOut(true)
+ }
+ }
)
try {
// Why: backgrounding or a direct winner withdraws this dial before cutover.
diff --git a/mobile/src/transport/relay-host-signed-out-supervisor.test.ts b/mobile/src/transport/relay-host-signed-out-supervisor.test.ts
new file mode 100644
index 00000000000..cc7a9ac6c35
--- /dev/null
+++ b/mobile/src/transport/relay-host-signed-out-supervisor.test.ts
@@ -0,0 +1,60 @@
+import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
+import { MOBILE_RELAY_CLOSE_CODE } from '../../../src/shared/mobile-relay-close-codes'
+import { RELAY_HOST_CLOSE_REASON } from '../../../src/shared/relay-host-close-reason'
+import { RelayOuterError } from './mobile-relay-e2ee-link'
+import {
+ dependencies,
+ FakeLogicalClient,
+ FakeRelaySession,
+ host
+} from './mobile-endpoint-supervisor-test-fakes'
+import { MobileEndpointSupervisor } from './mobile-endpoint-supervisor'
+
+vi.mock('react-native', () => ({ Platform: { OS: 'ios' } }))
+vi.mock('expo-secure-store', () => ({ WHEN_UNLOCKED_THIS_DEVICE_ONLY: 'when-unlocked' }))
+vi.mock('expo-crypto', () => ({ getRandomBytes: (length: number) => new Uint8Array(length) }))
+
+// The reason travels from the cell's close frame to the screens. This covers
+// the production wiring between them: the supervisor's own openRelay callback.
+describe('a signed-out desktop reaches the phone verdict', () => {
+ beforeEach(() => {
+ vi.useFakeTimers()
+ vi.setSystemTime(new Date('2026-07-13T12:00:00Z'))
+ })
+ afterEach(() => vi.useRealTimers())
+
+ function supervisorOver(closeReason: string | null) {
+ const logical = new FakeLogicalClient('disconnected', 'lan')
+ const deps = dependencies({
+ openDirect: vi.fn(() => new FakeRelaySession('disconnected')),
+ openRelay: vi.fn((_relay, _credential, _confirmReqId, onHostCloseReason) => {
+ if (closeReason) {
+ onHostCloseReason?.(closeReason as never)
+ }
+ return new FakeRelaySession(
+ 'disconnected',
+ new RelayOuterError(MOBILE_RELAY_CLOSE_CODE.HOST_OFFLINE)
+ )
+ })
+ })
+ return { logical, supervisor: new MobileEndpointSupervisor(logical, host, deps) }
+ }
+
+ it('latches the sign-out the cell reported', async () => {
+ const { logical, supervisor } = supervisorOver(RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+
+ await supervisor.start()
+ await vi.waitFor(() => expect(logical.isHostSignedOut()).toBe(true))
+
+ supervisor.stop()
+ })
+
+ it('stays quiet for an ordinary host-offline rejection', async () => {
+ const { logical, supervisor } = supervisorOver(null)
+
+ await supervisor.start()
+
+ expect(logical.isHostSignedOut()).toBe(false)
+ supervisor.stop()
+ })
+})
diff --git a/mobile/src/transport/relay-host-signed-out-verdict.test.ts b/mobile/src/transport/relay-host-signed-out-verdict.test.ts
new file mode 100644
index 00000000000..2607b922b58
--- /dev/null
+++ b/mobile/src/transport/relay-host-signed-out-verdict.test.ts
@@ -0,0 +1,216 @@
+import { describe, expect, it, vi } from 'vitest'
+
+vi.mock('./mobile-e2ee-v2-client-session', () => ({
+ MobileE2EEV2ClientSession: { create: () => ({}) }
+}))
+
+vi.mock('./mobile-e2ee-v2-physical-channel', () => ({
+ MobileE2EEAuthenticationError: class extends Error {},
+ MobileE2EEV2PhysicalChannel: class {
+ start = vi.fn()
+ handleMessage = vi.fn(async () => {})
+ sendText = vi.fn(() => true)
+ sendBinary = vi.fn(() => true)
+ dispose = vi.fn()
+ }
+}))
+
+import { RELAY_HOST_CLOSE_REASON } from '../../../src/shared/relay-host-close-reason'
+import { MOBILE_RELAY_CLOSE_CODE } from '../../../src/shared/mobile-relay-close-codes'
+import { classifyConnection, verdictDisplayLabel } from './connection-health'
+import { MobileRelayE2eeLink, RelayOuterError } from './mobile-relay-e2ee-link'
+import { LogicalClientConnectionPath } from './logical-client-connection-path'
+import { RelayReconnectController } from './mobile-relay-reconnect-controller'
+
+const SIGNED_OUT_LABEL = 'Desktop signed out — sign in to Orca on your desktop to reconnect'
+
+class FakeSocket {
+ static readonly OPEN = 1
+ readonly OPEN = FakeSocket.OPEN
+ readyState = FakeSocket.OPEN
+ bufferedAmount = 0
+ onopen: (() => void) | null = null
+ onmessage: ((event: { data: unknown }) => void) | null = null
+ onerror: (() => void) | null = null
+ onclose: ((event: { code: number; reason: string }) => void) | null = null
+ send = vi.fn()
+ close = vi.fn()
+}
+
+function linkOver(
+ socket: FakeSocket,
+ onHostCloseReason: (reason: string) => void,
+ onError: (error: Error) => void
+): MobileRelayE2eeLink {
+ return new MobileRelayE2eeLink({
+ endpoint: { cellUrl: 'https://relay-c1.onorca.dev', relayHostId: 'AbCdEf0123_-xyZ9' },
+ credential: 'credential',
+ expectedCredentialKind: 'resume',
+ deviceToken: 'device-token',
+ desktopPublicKeyB64: 'desktop-key',
+ onAuthenticated: vi.fn(),
+ onText: vi.fn(),
+ onBinary: vi.fn(),
+ onHostCloseReason,
+ onError,
+ createSocket: () => socket as unknown as WebSocket
+ })
+}
+
+describe('relay close reason on the phone', () => {
+ it('reports the cell close reason and still fails with 4404', () => {
+ const socket = new FakeSocket()
+ const onHostCloseReason = vi.fn()
+ const onError = vi.fn()
+ linkOver(socket, onHostCloseReason, onError)
+
+ socket.onclose?.({
+ code: MOBILE_RELAY_CLOSE_CODE.HOST_OFFLINE,
+ reason: RELAY_HOST_CLOSE_REASON.SIGNED_OUT
+ })
+
+ expect(onHostCloseReason).toHaveBeenCalledWith(RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+ expect(onError).toHaveBeenCalledWith(new RelayOuterError(MOBILE_RELAY_CLOSE_CODE.HOST_OFFLINE))
+ })
+
+ // An old cell sends its constant, and every other close sends nothing.
+ it('reports nothing for a reason it does not know', () => {
+ const socket = new FakeSocket()
+ const onHostCloseReason = vi.fn()
+ linkOver(socket, onHostCloseReason, vi.fn())
+
+ socket.onclose?.({
+ code: MOBILE_RELAY_CLOSE_CODE.HOST_OFFLINE,
+ reason: 'relay connection rejected'
+ })
+
+ expect(onHostCloseReason).not.toHaveBeenCalled()
+ })
+
+ // The rejection arrives as a relay-hello AND a close, in an unordered pair.
+ // Whichever lands first, the reason must survive.
+ it('still reports the reason when the hello already failed the link', async () => {
+ const socket = new FakeSocket()
+ const onHostCloseReason = vi.fn()
+ linkOver(socket, onHostCloseReason, vi.fn())
+
+ socket.onmessage?.({
+ data: JSON.stringify({
+ type: 'relay-hello',
+ ok: false,
+ code: MOBILE_RELAY_CLOSE_CODE.HOST_OFFLINE
+ })
+ })
+ await Promise.resolve()
+ await Promise.resolve()
+ socket.onclose?.({
+ code: MOBILE_RELAY_CLOSE_CODE.HOST_OFFLINE,
+ reason: RELAY_HOST_CLOSE_REASON.SIGNED_OUT
+ })
+
+ expect(onHostCloseReason).toHaveBeenCalledWith(RELAY_HOST_CLOSE_REASON.SIGNED_OUT)
+ })
+})
+
+describe('the signed-out signal on the logical client', () => {
+ it('publishes on change and retires when any path reaches connected', () => {
+ const path = new LogicalClientConnectionPath(() => false)
+ const changes = vi.fn()
+ path.subscribe(changes)
+
+ path.setHostSignedOut(true)
+ path.setHostSignedOut(true)
+ expect(path.isHostSignedOut()).toBe(true)
+ expect(changes).toHaveBeenCalledTimes(1)
+
+ path.clearAfterConnected()
+ expect(path.isHostSignedOut()).toBe(false)
+ })
+})
+
+describe('RelayReconnectController cadence', () => {
+ // The reason changes no recovery decision; 4404 keeps the host-offline
+ // backoff it has always had, so a phone on this build retries exactly as
+ // often as one that never hears the reason.
+ it('keeps the host-offline retry delay for a 4404', () => {
+ const delays: number[] = []
+ const controller = new RelayReconnectController(
+ {
+ now: () => 0,
+ randomBytes: () => new Uint8Array([0, 0]),
+ setTimer: ((callback: () => void, delay: number) => {
+ delays.push(delay)
+ return 1 as unknown as ReturnType
+ }) as unknown as typeof setTimeout,
+ clearTimer: (() => {}) as unknown as typeof clearTimeout
+ },
+ vi.fn()
+ )
+
+ controller.registerFailure(new RelayOuterError(MOBILE_RELAY_CLOSE_CODE.HOST_OFFLINE))
+
+ // hostOfflineDelayMs' 5s floor, not the 250ms transport-backoff floor.
+ expect(delays.at(-1)).toBe(5_000)
+ })
+})
+
+describe('classifyConnection with a signed-out desktop', () => {
+ const base = { reconnectAttempts: 0, lastConnectedAt: null, hostSignedOut: true }
+
+ it('says so from the first failed dial instead of "Connecting via Relay…"', () => {
+ const verdict = classifyConnection({
+ ...base,
+ state: 'connecting',
+ pendingPath: 'relay'
+ })
+
+ expect(verdict).toEqual({
+ kind: 'unreachable',
+ label: SIGNED_OUT_LABEL,
+ reason: 'never-connected'
+ })
+ expect(verdictDisplayLabel(verdict)).toBe(SIGNED_OUT_LABEL)
+ })
+
+ it('replaces "Can\'t reach desktop" on the direct path too', () => {
+ expect(
+ classifyConnection({ ...base, state: 'reconnecting', reconnectAttempts: 20 }).label
+ ).toBe(SIGNED_OUT_LABEL)
+ })
+
+ it('reads as stale once this session had been connected', () => {
+ expect(
+ classifyConnection({ ...base, state: 'reconnecting', lastConnectedAt: 1, nowMs: 2 }).reason
+ ).toBe('stale')
+ })
+
+ // A Tailscale endpoint cannot make "sign in on your desktop" better advice.
+ it('never appends the Tailscale hint', () => {
+ expect(
+ classifyConnection({ ...base, state: 'reconnecting', endpoint: '100.64.0.1' })
+ ).not.toHaveProperty('hint')
+ })
+
+ it('never outranks a connected session', () => {
+ expect(classifyConnection({ ...base, state: 'connected' }).label).toBe('Connected')
+ })
+
+ // Re-pairing, not signing in, is the remedy when the pairing itself is dead.
+ it('never outranks a revoked pairing', () => {
+ expect(classifyConnection({ ...base, state: 'reconnecting', pairingRejected: true }).kind).toBe(
+ 'auth-failed'
+ )
+ })
+
+ it('leaves every other verdict alone when the desktop is not signed out', () => {
+ expect(
+ classifyConnection({
+ state: 'connecting',
+ reconnectAttempts: 0,
+ lastConnectedAt: null,
+ pendingPath: 'relay',
+ hostSignedOut: false
+ }).label
+ ).toBe('Connecting via Relay…')
+ })
+})
diff --git a/mobile/src/transport/rpc-client-context-contract.ts b/mobile/src/transport/rpc-client-context-contract.ts
index 54e25973c7f..65262a6fc15 100644
--- a/mobile/src/transport/rpc-client-context-contract.ts
+++ b/mobile/src/transport/rpc-client-context-contract.ts
@@ -24,6 +24,7 @@ export type RpcClientContextValue = {
getActivePath: (hostId: string) => MobileConnectionPath
getPendingPath: (hostId: string) => MobileConnectionPath | null
isPairingRejected: (hostId: string) => boolean
+ isHostSignedOut: (hostId: string) => boolean
subscribeHostState: (hostId: string, listener: (state: ConnectionState) => void) => () => void
getAllClients: () => { hostId: string; client: RpcClient }[]
subscribeAllHosts: (listener: () => void) => () => void
diff --git a/mobile/src/transport/stable-logical-rpc-client.ts b/mobile/src/transport/stable-logical-rpc-client.ts
index d1f1701aed9..fb514128382 100644
--- a/mobile/src/transport/stable-logical-rpc-client.ts
+++ b/mobile/src/transport/stable-logical-rpc-client.ts
@@ -55,6 +55,9 @@ export type StableLogicalRpcClient = RpcClient & {
// Latched when the desktop has repeatedly refused this device's relay credential.
setPairingRejected(rejected: boolean): void
isPairingRejected(): boolean
+ // Latched when the relay named the desktop's own sign-out as the reason it is absent.
+ setHostSignedOut(signedOut: boolean): void
+ isHostSignedOut(): boolean
// Recovery attempts share this signal so status-only changes rerender.
onConnectionPathChange(listener: () => void): () => void
getGeneration(): number
@@ -282,6 +285,8 @@ export function createStableLogicalRpcClient(
setRecoveryAttempt: (attempt) => connectionPath.setRecoveryAttempt(attempt),
setPairingRejected: (rejected) => connectionPath.setPairingRejected(rejected),
isPairingRejected: () => connectionPath.isPairingRejected(),
+ setHostSignedOut: (signedOut) => connectionPath.setHostSignedOut(signedOut),
+ isHostSignedOut: () => connectionPath.isHostSignedOut(),
onConnectionPathChange: (listener) => connectionPath.subscribe(listener),
getGeneration: () => generation
}
diff --git a/mobile/src/transport/use-all-host-clients.ts b/mobile/src/transport/use-all-host-clients.ts
index 03ac5890015..70c709d965f 100644
--- a/mobile/src/transport/use-all-host-clients.ts
+++ b/mobile/src/transport/use-all-host-clients.ts
@@ -138,6 +138,7 @@ export function useAllHostClients(hostIds: string[], options?: UseAllHostClients
path: MobileConnectionPath
pendingPath: MobileConnectionPath | null
pairingRejected: boolean
+ hostSignedOut: boolean
}>((hostId) => {
const client = clientsByHostId.get(hostId)
return client
@@ -148,7 +149,8 @@ export function useAllHostClients(hostIds: string[], options?: UseAllHostClients
state: ctx.getState(hostId),
path: ctx.getActivePath(hostId),
pendingPath: ctx.getPendingPath(hostId),
- pairingRejected: ctx.isPairingRejected(hostId)
+ pairingRejected: ctx.isPairingRejected(hostId),
+ hostSignedOut: ctx.isHostSignedOut(hostId)
}
]
: []
diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml
index 0b72bb101ba..6b59d23e026 100644
--- a/pnpm-lock.yaml
+++ b/pnpm-lock.yaml
@@ -116,7 +116,7 @@ patchedDependencies:
'@xterm/addon-webgl@0.20.0-beta.299': 94687e89a0115e6e6aa102837f986debdc029c091527ee5eb4a4e17ceaf9473e
'@xterm/xterm@6.1.0-beta.303': 98756bcedc402bcdb7c6ab7b015d2e59cd18e97b03a2c06a27e95bb3ba429d9d
lint-staged@16.4.0: 7333b3837f80a7fbd045964db6d76ba4fc118e49134bdbabb00585b6b7b60673
- node-pty@1.1.0: e262847f57a1d4d3f2287a843822f7dcf3c9d8655892b07a69eba464e1317eaa
+ node-pty@1.1.0: 7cc9d45f3d2c38f142490d0805e75db55f0eef5174ad41c4b52abc5fbe079ad1
importers:
@@ -160,7 +160,7 @@ importers:
version: 3.3.1
node-pty:
specifier: ^1.1.0
- version: 1.1.0(patch_hash=e262847f57a1d4d3f2287a843822f7dcf3c9d8655892b07a69eba464e1317eaa)
+ version: 1.1.0(patch_hash=7cc9d45f3d2c38f142490d0805e75db55f0eef5174ad41c4b52abc5fbe079ad1)
posthog-node:
specifier: ^5.33.3
version: 5.33.3
@@ -12285,7 +12285,7 @@ snapshots:
node-int64@0.4.0: {}
- node-pty@1.1.0(patch_hash=e262847f57a1d4d3f2287a843822f7dcf3c9d8655892b07a69eba464e1317eaa):
+ node-pty@1.1.0(patch_hash=7cc9d45f3d2c38f142490d0805e75db55f0eef5174ad41c4b52abc5fbe079ad1):
dependencies:
node-addon-api: 7.1.1
diff --git a/resources/skills/current-manifest.json b/resources/skills/current-manifest.json
index c8b204f00c5..fdb54016a8f 100644
--- a/resources/skills/current-manifest.json
+++ b/resources/skills/current-manifest.json
@@ -131,17 +131,17 @@
"name": "orchestration",
"sourcePath": "skills/orchestration",
"releaseRevision": 29,
- "packageDigest": "7a386ce558ba54abe02b4a0de5d71fe3d63c944ef0888ddde130c729b37f7cc8",
- "gitTreeSha": "4199ec6988801dd491631706cba62631b4bed8fb",
+ "packageDigest": "689e31d84256aded123c801eaa87413474943a9a30d96bff9a19d0a321aefb54",
+ "gitTreeSha": "902cc33dd65730b32ac234dd0ae7166d75498b46",
"files": [
{
"path": "SKILL.md",
- "size": 4451,
+ "size": 4398,
"executable": false,
"classification": "text",
- "exactSha256": "a7e3350f037698ebbce36b2818d8383ec6e96c2a4825537caab39a83b4fb4b7f",
- "textNormalizedSha256": "a7e3350f037698ebbce36b2818d8383ec6e96c2a4825537caab39a83b4fb4b7f",
- "identitySha256": "a7e3350f037698ebbce36b2818d8383ec6e96c2a4825537caab39a83b4fb4b7f"
+ "exactSha256": "19ffdc1fe0d2c97dae845e8d636edb16781453ce2ec26f65a323c492ef90da18",
+ "textNormalizedSha256": "19ffdc1fe0d2c97dae845e8d636edb16781453ce2ec26f65a323c492ef90da18",
+ "identitySha256": "19ffdc1fe0d2c97dae845e8d636edb16781453ce2ec26f65a323c492ef90da18"
}
]
}
diff --git a/resources/skills/snapshot-registry.json b/resources/skills/snapshot-registry.json
index 5842b48858a..5b3412a497b 100644
--- a/resources/skills/snapshot-registry.json
+++ b/resources/skills/snapshot-registry.json
@@ -1046,17 +1046,17 @@
},
{
"releaseRevision": 29,
- "packageDigest": "7a386ce558ba54abe02b4a0de5d71fe3d63c944ef0888ddde130c729b37f7cc8",
- "gitTreeSha": "4199ec6988801dd491631706cba62631b4bed8fb",
+ "packageDigest": "689e31d84256aded123c801eaa87413474943a9a30d96bff9a19d0a321aefb54",
+ "gitTreeSha": "902cc33dd65730b32ac234dd0ae7166d75498b46",
"files": [
{
"path": "SKILL.md",
- "size": 4451,
+ "size": 4398,
"executable": false,
"classification": "text",
- "exactSha256": "a7e3350f037698ebbce36b2818d8383ec6e96c2a4825537caab39a83b4fb4b7f",
- "textNormalizedSha256": "a7e3350f037698ebbce36b2818d8383ec6e96c2a4825537caab39a83b4fb4b7f",
- "identitySha256": "a7e3350f037698ebbce36b2818d8383ec6e96c2a4825537caab39a83b4fb4b7f"
+ "exactSha256": "19ffdc1fe0d2c97dae845e8d636edb16781453ce2ec26f65a323c492ef90da18",
+ "textNormalizedSha256": "19ffdc1fe0d2c97dae845e8d636edb16781453ce2ec26f65a323c492ef90da18",
+ "identitySha256": "19ffdc1fe0d2c97dae845e8d636edb16781453ce2ec26f65a323c492ef90da18"
}
]
}
diff --git a/skill-guides/orchestration.md b/skill-guides/orchestration.md
index b9b7ca79442..0878532c447 100644
--- a/skill-guides/orchestration.md
+++ b/skill-guides/orchestration.md
@@ -3,18 +3,18 @@ name: orchestration
description: >-
Use Orca orchestration for structured multi-agent coordination: threaded
messages, blocking ask/reply flows, task dispatch, worker_done/escalation
- waits, task DAGs, decision gates, coordinator loops, or decomposing work
- across agents. Use `orca-cli` instead for full ownership handoffs, including
- requests phrased as "hand off", "handoff", "handover", "give this to another
- agent", or "another worktree" when the user did not explicitly ask to
- supervise, monitor, wait for results, or coordinate a DAG. Use `orca-cli` for
- terminal control, lightweight terminal prompts, shell commands, Orca
- worktree management, reading or waiting on terminals, and automation of the
- browser embedded inside Orca. Use Computer Use for external browser windows,
- webviews, Orca app UI, or desktop UI outside Orca's embedded browser only when
- the task requires OS/window-level control such as focus, menus, dialogs,
- coordinates, or screenshots. Use `orca-cli` for Orca's embedded pages and a
- page-automation tool such as Playwright or CDP for external pages.
+ waits, task DAGs, decision gates, or coordinator loops. Use `orca-cli`
+ instead for full ownership handoffs, including requests phrased as "hand
+ off", "handoff", "handover", "give this to another agent", or "another
+ worktree" when the user did not explicitly ask to supervise, monitor, wait
+ for results, or coordinate a DAG. Use `orca-cli` for terminal control,
+ lightweight terminal prompts, shell commands, Orca worktree management,
+ reading or waiting on terminals, and the Orca embedded browser. Use Computer
+ Use for external browser windows, webviews, Orca app UI, or desktop UI
+ outside Orca's embedded browser only when the task requires OS/window-level
+ control such as focus, menus, dialogs, coordinates, or screenshots. Use
+ `orca-cli` for Orca's embedded pages and a page-automation tool such as
+ Playwright or CDP for external pages.
---
# Orca Inter-Agent Orchestration
diff --git a/skills/orchestration/SKILL.md b/skills/orchestration/SKILL.md
index 85a0ff8c4b0..5725a8f5512 100644
--- a/skills/orchestration/SKILL.md
+++ b/skills/orchestration/SKILL.md
@@ -3,18 +3,18 @@ name: orchestration
description: >-
Use Orca orchestration for structured multi-agent coordination: threaded
messages, blocking ask/reply flows, task dispatch, worker_done/escalation
- waits, task DAGs, decision gates, coordinator loops, or decomposing work
- across agents. Use `orca-cli` instead for full ownership handoffs, including
- requests phrased as "hand off", "handoff", "handover", "give this to another
- agent", or "another worktree" when the user did not explicitly ask to
- supervise, monitor, wait for results, or coordinate a DAG. Use `orca-cli` for
- terminal control, lightweight terminal prompts, shell commands, Orca
- worktree management, reading or waiting on terminals, and automation of the
- browser embedded inside Orca. Use Computer Use for external browser windows,
- webviews, Orca app UI, or desktop UI outside Orca's embedded browser only when
- the task requires OS/window-level control such as focus, menus, dialogs,
- coordinates, or screenshots. Use `orca-cli` for Orca's embedded pages and a
- page-automation tool such as Playwright or CDP for external pages.
+ waits, task DAGs, decision gates, or coordinator loops. Use `orca-cli`
+ instead for full ownership handoffs, including requests phrased as "hand
+ off", "handoff", "handover", "give this to another agent", or "another
+ worktree" when the user did not explicitly ask to supervise, monitor, wait
+ for results, or coordinate a DAG. Use `orca-cli` for terminal control,
+ lightweight terminal prompts, shell commands, Orca worktree management,
+ reading or waiting on terminals, and the Orca embedded browser. Use Computer
+ Use for external browser windows, webviews, Orca app UI, or desktop UI
+ outside Orca's embedded browser only when the task requires OS/window-level
+ control such as focus, menus, dialogs, coordinates, or screenshots. Use
+ `orca-cli` for Orca's embedded pages and a page-automation tool such as
+ Playwright or CDP for external pages.
---
# Orca Orchestration
diff --git a/src/cli/bundled-skill-guides.ts b/src/cli/bundled-skill-guides.ts
index 66625fc9a1f..06aae7bdd7d 100644
--- a/src/cli/bundled-skill-guides.ts
+++ b/src/cli/bundled-skill-guides.ts
@@ -30,7 +30,7 @@ const ORCA_LINEAR_MARKDOWN = "---\nname: orca-linear\ndescription: >-\n Use Orc
const ORCA_PER_WORKSPACE_ENV_MARKDOWN = "---\nname: orca-per-workspace-env\ndescription: >-\n Set up, review, debug, or validate Orca per-workspace environment recipes —\n on-demand, disposable runtimes (cloud sandboxes, VMs, or local) created fresh\n for each workspace. Covers first-time setup (provider prerequisites, the\n reusable base snapshot, the coding-agent auth snapshot, credentials, and\n state), not just the per-workspace lifecycle scripts. Use to stand up\n per-workspace environments, fix an `environmentRecipes` entry in `orca.yaml`, scaffold\n provider lifecycle scripts, or resolve an `orca vm recipe doctor` failure.\n---\n\n# Per-Workspace Environments\n\nHelp a user stand up and maintain a repo-owned per-workspace environment recipe end to end. Each\nworkspace gets its own on-demand, disposable runtime (a cloud sandbox, a VM, or a local one),\ncreated fresh and torn down after.\n\nOrca is a **thin wrapper**: you guide, detect, and scaffold; you never own the user's cloud account,\nbilling, images, or credentials.\n\n- **You DO:** sequence the setup, detect what's detectable (provider CLI present/logged-in? recipe\n present? `doctor` passing?), scaffold provider-templated scripts the user fills in, drive the slow\n snapshot/auth phases with the user, and always show the next action.\n- **You DO NOT:** create accounts, choose plans/regions, invent org/project/scope ids, store or print\n secrets, or run anything that spends money without an explicit user OK.\n\nFirst-time setup has **four phases before the per-workspace recipe runs** — easy to miss, so walk\nthem in order:\n\n1. **Prerequisites** — cloud account, provider CLI, scope/project, plan limits, git token (§2).\n2. **Base snapshot** — reusable image: tools + repo + headless build, snapshotted once (§3).\n3. **Agent-auth snapshot** — boot the base, run interactive device-auth, re-snapshot (§4).\n4. **State** — thread snapshot id / scope / project / port between phases via a state file (§6).\n\nThen the **per-workspace contract** (create/suspend/resume/destroy) runs fast (§8).\n\n**The one branch that shapes everything — connection mode:** **Orca-server** (`create` runs `orca serve`\nin the env and emits a `pairingCode`; §7c/§7f) vs **SSH** (`create` runs no server and emits a\n`connection.type:\"ssh\"` block Orca dials into; §7g/§7h). Settle this first — it changes the `create`\noutput shape and half the templates.\n\nKeep Orca's checkout behavior unchanged by default: omit `checkoutMode`, emit schema version 1, and\nlet Orca create a linked worktree. Only use `checkoutMode: provisioned-root` when the user explicitly\nwants one ephemeral machine to clone the finished workspace itself. This niche mode currently requires\ndirect SSH, an ordinary non-bare/non-sparse primary checkout at `projectRoot`, and schema version 2.\n\n**Quick-start (happy path):** interview the user (connection mode Orca-server vs SSH, provider, agent CLI,\ngit auth — §1.2) + read the provider's CLI docs → scaffold `scripts/orca-vm/` from §7 → run the\nbase-snapshot script, then the auth script (you invoke these by hand; not via `orca.yaml`) → wire\n`environmentRecipes` in `orca.yaml` → `orca vm recipe doctor --json` (free) → then the `--provision`\nself-test loop (§9) until it passes.\n\n---\n\n## 1. Setup workflow\n\nDrive these with the user. **[CHECKPOINT]** steps need explicit confirmation — they spend money, take\na long time, or need the user at the keyboard. Never create an Orca workspace or commit unless asked.\n\n1. **Inspect the repo** for an existing `environmentRecipes` entry, `scripts/orca-vm/`, a state file, or setup\n notes. If a working recipe exists, jump to Doctor (§9) instead of rebuilding.\n2. **Interview the user up front** — gather these choices and confirm them back before scaffolding\n anything. Don't pick for them (§11); don't guess.\n - **Connection mode:** how Orca attaches to the environment — an **Orca server** (the VM runs\n `orca serve` and Orca pairs over its pairing URL; worked example §7f) or **SSH** (Orca connects to\n the host over SSH; §7g). This decides the recipe's connection shape, so settle it first.\n - **Checkout ownership:** do not ask by default. Only when the user requires the environment to\n create the exact final checkout, confirm `provisioned-root` and direct SSH; otherwise omit it.\n - **Provider:** Vercel Sandbox, Fly, Modal, an existing SSH host, … For non-obvious providers, also\n ask scope/project/region and plan limits (§2). Then **read that provider's CLI/SDK docs** (or\n ` --help`) before scaffolding — you need its exact create/exec/snapshot/remove verbs.\n If a provider advertises `ssh`, verify whether it exposes a real dialable SSH target\n (host/port/user/key or proxy command) or only a provider-mediated interactive shell; Orca SSH mode\n needs the former.\n - **Coding-agent CLI + account:** which agent runs in the VM (`codex`, `claude`, …) and that the user\n has an account for it — it gets logged in during the Phase-3 auth snapshot (§4).\n - **Git auth:** the token source for cloning a private repo (`GH_TOKEN`/`GITHUB_TOKEN` or `gh auth\ntoken`; §5).\n3. **Check prerequisites (§2)** — detect the provider CLI + auth and confirm the items above are in\n place before any paid step.\n4. **Scaffold scripts + state file** from §7 (worked Vercel example: §7f; SSH host: §7g; Docker SSH:\n §7h; Windows: §7i), filling in the provider's real commands. Make them executable.\n5. **[CHECKPOINT] Build the base snapshot (§3)** — paid, slow.\n6. **[CHECKPOINT] Authenticate the agent (§4)** — interactive; the user follows a URL/code. **You cannot\n drive this step** — you run commands non-interactively, so there's no TTY for `docker exec -it` /\n `ssh -t` to prompt against. The **user** runs the Phase-3 login in their own terminal (or via the\n Claude Code harness bang-prefix — `! `, with the required space after `!`); you scaffold and drive\n the non-interactive phases around it. After kicking it off, **ask the user to report back once the login\n finishes** — you can't observe it completing, and you need that confirmation before resuming the\n non-interactive steps (base/auth commit, doctor, provision).\n7. **Wire the recipe** so `orca.yaml` points create/suspend/resume/destroy at the scripts (§8). The\n workspace composer reads `environmentRecipes` from the project's primary checkout of `orca.yaml`, **not** from\n a feature branch or worktree. So a recipe added only on a branch won't appear as a \"Run on\" option\n until that `orca.yaml` change is committed and merged to the project's primary branch. Tell the user\n this up front: `doctor`/`--provision` validate the scripts from the working copy on any branch, but\n creating a workspace from the recipe in the picker needs it on primary.\n8. **Dry-run doctor** — `orca vm recipe doctor --repo-path --json` (free, static; §9).\n Fix every failure before going live.\n9. **[CHECKPOINT] Live self-test** — get the user's OK once, then run\n `orca vm recipe doctor --provision --json` as a loop: it runs create → validates →\n destroys, and on failure returns a full transcript. Read it, fix the scripts, and re-run yourself until\n it passes (§9). Spends cloud money; the one approval covers the loop.\n10. **[CHECKPOINT] Optional workspace test** — only if asked: create a workspace via the picker, then\n verify sleep/wake/delete.\n\n---\n\n## 2. Phase 1 — Prerequisites\n\nThe user's responsibility; verify what's verifiable, ask for the rest, invent nothing. State which\nitems you verified vs. which the user asserted.\n\n- **Connection mode** (Orca server vs SSH) confirmed with the user — see §1 step 2; it shapes the recipe.\n- **Cloud account + plan** that allows sandboxes/VMs. Ask.\n- **Provider CLI installed + authenticated** — detect (`command -v `), check auth (e.g.\n `vercel whoami`). If missing, point at the provider's docs; don't log them in.\n- **Scope / project / region** the sandboxes live under. Ask; flows into every script via state.\n- **Plan / timeout / RAM caps.** Record them — e.g. Vercel Hobby caps sandbox timeout at **45m**,\n which limits both the base build and per-workspace runtime (see §10).\n- **Git token for private repos** (`GH_TOKEN`/`GITHUB_TOKEN`, or the provider's git auth; can fall back\n to `gh auth token`). See §5.\n- **Coding-agent CLI choice** (`codex`, `claude`…) and that the user has an account — it gets\n authenticated into the VM in Phase 3.\n\n---\n\n## 3. Phase 2 — Base snapshot (the reusable image)\n\nBuild **once**, snapshot, and every workspace boots from it in seconds instead of rebuilding.\nProvisioning + building takes a while (often ~20–30 min), so it runs behind a checkpoint. The script\nshape is §7a; key points:\n\n- Build the **headless Electron main only** (not the renderer) so it fits in plan RAM.\n- Use the VM image's package manager (`apt`/`dnf`/`apk`, per the base distro — not the provider brand).\n- Clone with the git token via `GIT_ASKPASS` (§5).\n- **Trap errors and remove the half-built sandbox** so a crash doesn't leave a paid resource running.\n- **Never snapshot a machine on which the Orca runtime has already run.** The first `orca serve` creates\n the runtime's user-data dir, and everything in it gets baked into the image and shared by every VM\n booted from it: the pairing keypair and device-token registry (`orca-devices.json`,\n `orca-e2ee-keypair.json`), `agent-session-authority.key`, and the build box's logs, terminal history\n and orchestration db. Confirmed: two VMs from one such snapshot emitted **identical `deviceToken` and\n `pairedDeviceId`**. Snapshot **before** the runtime has ever run, or delete the resolved user-data\n directory first: `orca_user_data_path=\"${ORCA_USER_DATA_PATH:-${XDG_CONFIG_HOME:-$HOME/.config}/orca}\"; rm -rf -- \"$orca_user_data_path\"`.\n This matches Orca's Linux precedence for custom and default paths; deleting a named file list will\n drift as Orca adds state.\n- Snapshot the stopped sandbox, parse the snapshot id, and write it + scope/project/port/repo to state.\n\n---\n\n## 4. Phase 3 — Agent-auth snapshot (interactive)\n\nThe base snapshot has the agent CLI installed but **not logged in**, and per-workspace VMs are\nephemeral — so authenticate once and bake it into a second snapshot layer. Script shape is §7b:\n\n1. Boot a sandbox from the base `snapshotId` (from state).\n2. Run the agent's login **interactively** (`--interactive --tty`); the user completes the URL/code in\n their browser. On a **headless VM this must be the device-auth flow** (e.g. `codex login --device-auth`),\n **not** plain `codex login`: the default OAuth login starts a loopback callback server on a container\n port the host browser can't reach, so it hangs. Device-auth instead prints a URL + code the user opens\n on the **host**.\n3. Verify login; **refuse to snapshot an unauthenticated VM.** Prefer the status command's **exit code**\n (most agent CLIs exit non-zero when unauthenticated). If you grep instead, agent status often goes to\n **stderr** (e.g. `codex login status` prints \"Logged in using ChatGPT\" there), so **fold stderr first**\n (`... 2>&1 | grep …`) and match the agent's **exact success line** — never `grep -qi 'logged in'`, which\n also matches \"**not** logged in\" and would commit an unauthenticated image.\n4. Re-snapshot, parse the new id, and overwrite `snapshotId` in state to the authenticated image\n (recording `authSourceSnapshotId`). Remove the auth sandbox.\n\n**You can't drive step 2 yourself** (you run commands non-interactively — no TTY). The **user** runs it in\ntheir own terminal, or via the Claude Code harness bang-prefix (`! `, with the required space after\n`!`). You scaffold/boot the sandbox and run steps 3–4, but **you cannot observe the interactive login\nfinishing** — so **ask the user to tell you when it's done** before you verify and re-snapshot.\n\nThis layer inherits §3's rule: if you started `orca serve` on the base or auth sandbox to smoke-test it,\ndelete the runtime's user-data dir (`~/.config/orca` on Linux) before re-snapshotting, or every workspace\nbooted from this image shares one pairing identity and one `agent-session-authority.key`.\n\nIf the agent's credentials are short-lived, warn that the snapshot may need periodic re-auth (§10).\n\nFor disposable runtimes, do **not** treat a host agent config directory (for example `~/.codex`) as the\nauth snapshot by bind-mounting or copying it wholesale. Agent homes often contain sqlite state, hook\napproval state, caches, logs, and host-specific env/config. Instead, authenticate/configure the agent\ninside the disposable runtime and snapshot/commit that runtime layer.\n\n---\n\n## 5. Credentials\n\n- **Never** commit secrets or put them in `userData`, recipe JSON, comments, docs, or the state file.\n- **Git token:** read from env (`GH_TOKEN`/`GITHUB_TOKEN`), falling back to `gh auth token`. Pass to the\n VM only via the provider's ephemeral `--env`. Inside the VM, use a `GIT_ASKPASS` helper with\n `x-access-token` (not the token in the clone URL) and `GIT_TERMINAL_PROMPT=0` so a missing token fails\n fast instead of hanging. When you write the helper from inside `bash -lc` under `set -u`, escape the\n positional arg and the token (`\\$1`, `\\$GH_TOKEN`) so they land **literally** and resolve at git-runtime\n — an unescaped `$1` aborts with \"unbound variable\", and a literal `$GH_TOKEN` keeps the real token out of\n the written file. `rm -f` the helper after the clone/fetch.\n- **Provider auth:** rely on the provider CLI's logged-in session, not checked-in keys.\n- **Agent auth:** lives in the authenticated snapshot (Phase 3) — never a file you write or commit.\n- State holds only **non-secret** wiring (snapshot ids, scope, project, port, repo url/ref).\n\n---\n\n## 6. State file\n\nA repo-local JSON file (e.g. `scripts/orca-vm/-state.json`) threads non-secret values between\nphases. Each script resolves values as **env var → state → built-in fallback**, and merges its outputs\nback. Phase 2 writes the base `snapshotId`; Phase 3 overwrites it with the authenticated snapshot;\nper-workspace `create` boots from `snapshotId`.\n\n```json\n{\n \"baseName\": \"orca-base\",\n \"snapshotId\": \"snap_authenticated_image_id\",\n \"authSourceSnapshotId\": \"snap_base_image_id\",\n \"scope\": \"\",\n \"project\": \"\",\n \"port\": 7331,\n \"repoUrl\": \"https://host/org/repo.git\",\n \"repoRef\": \"main\",\n \"projectRoot\": \"/abs/path/on/remote/repo\"\n}\n```\n\n---\n\n## 7. Script templates (provider-agnostic shapes)\n\nScaffold under `scripts/orca-vm/`. These are **shapes** — fill in the provider's real commands. All\nreserve stdout for the final JSON and log progress to stderr. Include a shared `json_value ` /\n`env_value ` reader (env → state → fallback) in each.\n\n**Where each script runs:**\n\n- **Local-side** (`create`/`suspend`/`resume`/`destroy` + the base-snapshot/auth scripts the user\n invokes) runs **on the user's desktop**, so it must run on their OS. macOS/Linux: `#!/usr/bin/env\nbash`, `set -euo pipefail`, quoted paths. **Windows:** a bare `.sh` won't run — scaffold `.ps1`/`.cmd`\n or require WSL/Git-Bash and point `orca.yaml` at the right launcher.\n- **Remote-side** (commands you `exec` _inside_ the Linux VM) always runs in the VM's Linux shell, so\n bash is fine there regardless of the user's OS.\n\n### 7a. Base-snapshot (`-base-snapshot.sh`) — Phase 2\n\n```bash\n#!/usr/bin/env bash\nset -euo pipefail\n# resolve base_name/repo_url/repo_ref/project_root/port/scope/project/timeout (env→state→fallback)\n# resolve gh token: GH_TOKEN | GITHUB_TOKEN | `gh auth token`\n# 1. provision a sandbox (timeout/vcpus/published port/snapshot retention); trap: remove on error\n# 2. remote exec (long timeout): install pkgs + gh + corepack/pnpm + agent CLI;\n# clone with GIT_ASKPASS(token); write headless main-only build config;\n# dev setup; pnpm install; build CLI; build headless electron main; smoke-check tools\n# 3. snapshot stopped sandbox; parse snapshot id (fail if unparseable)\n# 4. merge { baseName, snapshotId, projectRoot, repoUrl, repoRef, port, scope, project } into state\n# print only the state JSON to stdout\n```\n\nWorked Vercel commands for this phase are in §7f. You run this script by hand (not via `orca.yaml`),\nafter exporting the first-run inputs the state file doesn't have yet — e.g. provider scope/project, the\nrepo URL/ref, and a git token (`GH_TOKEN`); later runs read them back from state.\n\n### 7b. Auth (`-base-auth.sh`) — Phase 3\n\n```bash\n#!/usr/bin/env bash\nset -euo pipefail\n# read source snapshot from state.snapshotId (fail if absent); auth_name=\"${base_name}-auth\"\n# 1. boot sandbox from source snapshot; trap: remove on error\n# 2. INTERACTIVE/TTY remote exec: agent login — user completes URL/code. Headless VM: MUST use the\n# device-auth flow (e.g. `codex login --device-auth`) — plain OAuth login binds a loopback callback\n# port the host can't reach and hangs. User runs this themselves (you have no interactive TTY); ask\n# them to report back when it's done before continuing.\n# 3. verify login, then refuse to snapshot if not logged in. Prefer the status command's EXIT CODE (most\n# agent CLIs exit non-zero when unauthenticated) over string-matching. If you must grep, fold stderr\n# first (`status 2>&1 | grep …` — many agents print the success line there) and match the agent's exact\n# success line; never `grep -qi 'logged in'`, which also matches \"not logged in\". Codex example: §7f.\n# 4. snapshot; parse new id\n# 5. merge { snapshotId:, authSourceSnapshotId: } into state; remove auth sandbox\n# print only the state JSON to stdout\n```\n\n### 7c. Create (`-create.sh`) — per workspace\n\n```bash\n#!/usr/bin/env bash\nset -euo pipefail\n# read authenticated snapshotId/scope/project/port/repo*/project_root (env→state→fallback)\n# fail clearly if snapshotId is missing (point back to Phases 2–3)\n# name = orca-${ORCA_RECIPE_ID}-${ORCA_VM_INSTANCE_ID} (sanitized, length-capped)\n# 1. boot sandbox from snapshotId with a published port; capture the public URL → pairing address\n# (an externally reachable wss:// URL); trap: remove sandbox on error\n# 2. remote exec: ensure repo at desired commit; rebuild only if commit changed (cache marker)\n# 3. remote exec: start orca serve in the background and read the recipe JSON it writes (see below)\n# 4. print serve's JSON to stdout, optionally enriched with userData:\n# { schemaVersion:1, pairingCode, projectRoot, userData:{ provider, resourceId:name, snapshotId } }\n```\n\n**The exact `orca serve` invocation and its output (verified — do not improvise the flags).** Inside the\nVM, run:\n\n```bash\norca serve \\\n --port \"$PORT\" \\\n --project-root \"$ABS_REPO_PATH_ON_REMOTE\" \\\n --pairing-address \"$EXTERNAL_WSS_URL\" \\\n --recipe-json\n```\n\n**Binary name:** in a VM built from source (the Phase-2 flow), run it as `pnpm exec orca-dev serve …`\nfrom the repo root — `orca-dev` is the in-repo entrypoint and is what the §7f example uses. Plain\n`orca serve …` is the same command when the built CLI is installed on the VM's PATH. The flags/output\nare identical either way.\n\nThere is **no `--host` flag**. `--project-root` must be an absolute directory on the remote. With\n`--recipe-json` the server **stays running** and prints exactly this single object to **stdout**, then\nkeeps serving:\n\n```json\n{\n \"schemaVersion\": 1,\n \"pairingCode\": \"\",\n \"projectRoot\": \"\"\n}\n```\n\n`pairingCode` is the pairing URL, already pointing at whatever you passed as `--pairing-address` — so set\n`--pairing-address` to the externally reachable address and **pass `pairingCode` through unchanged; never\nhand-rewrite it**. Because serve runs in the foreground and doesn't exit, redirect its stdout to a file\nand poll until that file parses as JSON (and bail if the process dies — dump its stderr log). Your\n`create` script then prints that JSON (optionally merging `userData`). Concrete pattern: §7f.\n\n### 7d. Suspend / resume / destroy — per workspace\n\n```bash\n#!/usr/bin/env bash\nset -euo pipefail\npayload=\"$(cat)\" # Orca passes lifecycle JSON on stdin\nresource_id=\"$(node -e 'const d=JSON.parse(process.argv[1]); process.stdout.write(d.recipeResult?.userData?.resourceId ?? \"\")' \"$payload\")\"\n[ -n \"$resource_id\" ] || { echo \"No resource id in lifecycle payload\" >&2; exit 1; }\n# suspend: provider suspend \"$resource_id\"\n# resume: provider resume \"$resource_id\"; then RE-EMIT fresh recipe JSON (pairing may change)\n# destroy: provider remove \"$resource_id\" (or set destroy: none in orca.yaml)\n```\n\n### 7e. State file — scaffold with scope/project/repo filled in and snapshot ids empty (§6).\n\n### 7f. Worked example — Vercel Sandbox (all three phases)\n\nA real, working shape (the Vercel surface is a CLI: `vercel sandbox create|exec|snapshot|remove`). Adapt\nnames; verify flags against `vercel sandbox --help` for the user's CLI version before relying on them.\nThese ground §7a (base snapshot) and §7b (auth), which are otherwise generic skeletons.\n\n**Phase 2 — base snapshot (§7a):** provision → install tools + clone + headless build → snapshot.\n\n```bash\n# provision a fresh build sandbox (retain a couple of snapshots); trap-remove on error\nvercel sandbox create --name \"$base\" --runtime node24 --timeout 30m --vcpus 4 --publish-port \"$port\" \\\n --snapshot-expiration 30d --keep-last-snapshots 2 \"${vercel_args[@]}\" >&2\n# remote build (long timeout): install pkgs+gh+pnpm+agent CLI, clone with GIT_ASKPASS (write the helper\n# with LITERAL \\$1/\\$GH_TOKEN so they resolve at git-runtime, not write-time — see §5/§7f create — then\n# `rm -f /tmp/askpass.sh`), write the headless main-only build config (drop the renderer), dev setup,\n# build CLI + headless main, smoke-check\nvercel sandbox exec \"$base\" \"${vercel_args[@]}\" --timeout 25m --env \"GH_TOKEN=$gh_token\" … -- bash -lc '…build…' >&2\n# snapshot the STOPPED sandbox and parse the id from CLI output (fail if unparseable)\nout=\"$(vercel sandbox snapshot \"$base\" --stop --expiration 30d \"${vercel_args[@]}\" 2>&1)\"; printf '%s\\n' \"$out\" >&2\nsnapshot_id=\"$(printf '%s\\n' \"$out\" | sed -nE 's/.*(snap_[A-Za-z0-9]+).*/\\1/p' | tail -1)\"\n# merge { baseName, snapshotId, scope, project, port, repoUrl, repoRef, projectRoot } into state; print state JSON\n```\n\n**Phase 3 — agent-auth snapshot (§7b):** boot the base, log the agent in interactively, re-snapshot.\n(`codex` below is an example — substitute the user's chosen agent's login/status verbs, e.g. `claude`.)\n\n```bash\nvercel sandbox create --name \"$auth\" --snapshot \"$snapshot_id\" --timeout 30m --publish-port \"$port\" \"${vercel_args[@]}\" >&2\n# INTERACTIVE — the USER runs this in their own terminal (you have no interactive TTY) and completes the\n# URL/code on the HOST. --device-auth is MANDATORY on a headless VM: plain `codex login` binds a loopback\n# callback port the host browser can't reach and hangs. Ask the user to report back when login finishes.\nvercel sandbox exec --interactive --tty \"$auth\" \"${vercel_args[@]}\" -- bash -lc 'codex login --device-auth'\n# refuse to snapshot an unauthenticated VM — fold stderr, match codex's exact success line (§4)\nvercel sandbox exec \"$auth\" \"${vercel_args[@]}\" --timeout 30s -- bash -lc 'codex login status 2>&1' | grep -Eqi 'Logged in using ChatGPT|Logged in via device' \\\n || { echo \"agent not logged in; not snapshotting\" >&2; exit 1; }\nout=\"$(vercel sandbox snapshot \"$auth\" --stop --expiration 30d \"${vercel_args[@]}\" 2>&1)\"; printf '%s\\n' \"$out\" >&2\nnew_id=\"$(printf '%s\\n' \"$out\" | sed -nE 's/.*(snap_[A-Za-z0-9]+).*/\\1/p' | tail -1)\"\n# overwrite state.snapshotId = new_id, record authSourceSnapshotId = snapshot_id; remove the auth sandbox\n```\n\n**Per-workspace `create`** (the fast path):\n\n```bash\n#!/usr/bin/env bash\nset -euo pipefail\n# resolve from env→state→fallback: snapshot_id, scope, project, port, repo_url, repo_ref, project_root\nvercel_args=(); [ -n \"$scope\" ] && vercel_args+=(--scope \"$scope\"); [ -n \"$project\" ] && vercel_args+=(--project \"$project\")\n[ -n \"$snapshot_id\" ] || { echo \"snapshotId missing — run Phases 2–3 first\" >&2; exit 1; }\ngh_token=\"${GH_TOKEN:-${GITHUB_TOKEN:-$(command -v gh >/dev/null 2>&1 && gh auth token 2>/dev/null || true)}}\"\nrecipe_id=\"${ORCA_RECIPE_ID:-vercel-sandbox}\"\nrecipe_id=\"${recipe_id//./-}\" # Vercel names forbid dots.\ninstance_id=\"${ORCA_VM_INSTANCE_ID:-$(date +%s)}\"\nmax_recipe_id_length=$((128 - ${#instance_id} - 6)) # Preserve the unique instance suffix.\n[ \"$max_recipe_id_length\" -gt 0 ] || { echo \"ORCA_VM_INSTANCE_ID is too long for a Vercel sandbox name\" >&2; exit 1; }\nname=\"orca-${recipe_id:0:max_recipe_id_length}-${instance_id}\"\n\n# Arm cleanup BEFORE create so a failing create can't leak a half-built paid sandbox.\ncleanup_on_error() { [ \"$?\" -ne 0 ] && vercel sandbox remove \"$name\" \"${vercel_args[@]}\" >/dev/null 2>&1 || true; }\ntrap cleanup_on_error EXIT\n\n# 1. boot from the authenticated snapshot, publish the serve port\ncreate_output=\"$(vercel sandbox create --name \"$name\" --snapshot \"$snapshot_id\" \\\n --timeout 30m --publish-port \"$port\" \"${vercel_args[@]}\" 2>&1)\"; printf '%s\\n' \"$create_output\" >&2\n# Vercel prints the published https URL; derive the external wss:// pairing address from it\npublic_url=\"$(printf '%s\\n' \"$create_output\" | sed -nE 's#.*(https://[^[:space:]]+\\.vercel\\.run).*#\\1#p' | head -1)\"\n[ -n \"$public_url\" ] || { echo \"no published URL in create output\" >&2; exit 1; }\npairing_ws=\"${public_url/https:\\/\\//wss://}\"\n\n# 2. (remote) ensure the repo is at the right commit; rebuild only if the commit changed (cache marker)\nvercel sandbox exec \"$name\" \"${vercel_args[@]}\" --timeout 20m \\\n --env \"GH_TOKEN=$gh_token\" --env \"ORCA_PROJECT_ROOT=$project_root\" \\\n --env \"ORCA_REPO_URL=$repo_url\" --env \"ORCA_REPO_REF=$repo_ref\" \\\n -- bash -lc 'set -euo pipefail; cd \"$ORCA_PROJECT_ROOT\"; \\\n # Re-establish git auth for the private-repo fetch (why + full rationale: §5); else it hangs on a prompt.\n # Load-bearing escaping: \\$1 and \\$GH_TOKEN must land LITERALLY and resolve at git-runtime. Test after\n # any edit here — reformatting the nested printf/node quoting silently breaks the fetch or leaks the token.\n if [ -n \"${GH_TOKEN:-}\" ]; then \\\n printf \"%s\\n\" \"#!/usr/bin/env bash\" \"case \\\"\\$1\\\" in *Username*) echo x-access-token;; *Password*) echo \\\"\\$GH_TOKEN\\\";; esac\" > /tmp/askpass.sh; \\\n chmod 700 /tmp/askpass.sh; export GIT_ASKPASS=/tmp/askpass.sh GIT_TERMINAL_PROMPT=0; fi; \\\n git fetch origin \"$ORCA_REPO_REF\"; \\\n git checkout -B \"$ORCA_REPO_REF\" FETCH_HEAD; \\\n rm -f /tmp/askpass.sh; \\\n c=\"$(git rev-parse HEAD)\"; [ -f .orca-built ] && [ \"$(cat .orca-built)\" = \"$c\" ] || { \\\n pnpm install --prefer-offline && pnpm run build:cli && \\\n node config/scripts/run-electron-vite-build.mjs --config config/electron-vite.vm-serve.config.ts && \\\n printf \"%s\" \"$c\" > .orca-built; }' >&2\n\n# 3. (remote) start orca serve in the background, writing recipe JSON to a file; poll until it parses\nrecipe_json=\"$(vercel sandbox exec \"$name\" \"${vercel_args[@]}\" --timeout 60s \\\n --env \"ORCA_PORT=$port\" --env \"ORCA_PROJECT_ROOT=$project_root\" --env \"ORCA_PAIRING_ADDRESS=$pairing_ws\" \\\n -- bash -lc 'set -euo pipefail; cd \"$ORCA_PROJECT_ROOT\"; rm -f /tmp/orca-recipe.json /tmp/orca-serve.log; \\\n nohup pnpm exec orca-dev serve --port \"$ORCA_PORT\" --project-root \"$ORCA_PROJECT_ROOT\" \\\n --pairing-address \"$ORCA_PAIRING_ADDRESS\" --recipe-json >/tmp/orca-recipe.json 2>/tmp/orca-serve.log /dev/null 2>&1 && { cat /tmp/orca-recipe.json; exit 0; }; \\\n kill -0 \"$pid\" 2>/dev/null || { cat /tmp/orca-serve.log >&2; exit 1; }; sleep 0.25; \\\n done; cat /tmp/orca-serve.log >&2; echo \"serve recipe JSON timed out\" >&2; exit 1')\"\n\n# 4. print serve's JSON enriched with userData (single object on stdout)\nnode -e 'const p=JSON.parse(process.argv[1]); console.log(JSON.stringify({...p, schemaVersion:1,\n userData:{...p.userData, provider:\"vercel-sandbox\", resourceId:process.argv[2], snapshotId:process.argv[3]}}))' \\\n \"$recipe_json\" \"$name\" \"$snapshot_id\"\ntrap - EXIT\n```\n\n`suspend`/`resume`/`destroy` use `vercel sandbox stop|...|remove \"$resource_id\"` reading\n`userData.resourceId` from stdin (§7d). This is the **Orca-server** connection mode (the recipe emits a\npairing URL). If the user chose **SSH** in the §1 interview, use §7g instead.\n\n### 7g. Worked example — existing SSH host (SSH connection mode)\n\nSSH mode is **fundamentally different from §7c/§7f**, not a relabeling of them:\n\n- **`create` does NOT run `orca serve` and does NOT emit a `pairingCode`.** Orca itself connects to the\n host over its SSH relay, brings up the git + filesystem providers, and imports the repo. The script's\n only job is to make the host ready and **print SSH connection details** Orca will dial.\n- The result uses a `connection` block with `type: \"ssh\"` and a `target`, **not** the flat\n `pairingCode`/`projectRoot` shape. Exact shape (Orca rejects anything else):\n\n```json\n{\n \"schemaVersion\": 1,\n \"connection\": {\n \"type\": \"ssh\",\n \"projectRoot\": \"/abs/path/to/repo/on/host\",\n \"target\": {\n \"label\": \"my-box\",\n \"host\": \"192.0.2.10\",\n \"port\": 22,\n \"username\": \"ubuntu\",\n \"identityFile\": \"~/.ssh/id_ed25519\",\n \"jumpHost\": \"bastion.example.com\",\n \"proxyCommand\": \"cloudflared access ssh --hostname %h\",\n \"relayGracePeriodSeconds\": 0,\n \"portForwards\": []\n }\n }\n}\n```\n\n`label`, `host`, `port`, `username` are required; the rest are optional — omit any you don't need.\n\nFor an explicitly requested one-VM-per-workspace checkout, the create script must read\n`ORCA_RECIPE_RESULT_SCHEMA_VERSION`, `ORCA_REPO_URL`, `ORCA_REPO_REF`, `ORCA_REPO_REF_HEAD`, and\n`ORCA_REPO_BRANCH`. Use `ORCA_REPO_REF` to fetch the selected source, but create\n`ORCA_REPO_BRANCH` at the exact `ORCA_REPO_REF_HEAD` commit; resolving the symbolic ref again can race\nwith an upstream update. `ORCA_REPO_URL` and `ORCA_REPO_REF` are a matched fetch pair, including when\nthe desktop source uses multiple remotes. Return that primary checkout at `projectRoot` and emit the\nsame SSH result with:\n\n```bash\n[ -n \"${ORCA_REPO_REF_HEAD:-}\" ] || { echo \"missing pinned source commit\" >&2; exit 1; }\ngit fetch origin \"$ORCA_REPO_REF\"\ngit cat-file -e \"${ORCA_REPO_REF_HEAD}^{commit}\"\ngit checkout -B \"$ORCA_REPO_BRANCH\" \"$ORCA_REPO_REF_HEAD\"\n```\n\n```json\n{\n \"schemaVersion\": 2,\n \"checkoutMode\": \"provisioned-root\",\n \"connection\": {\n \"type\": \"ssh\",\n \"projectRoot\": \"/abs/repo\",\n \"target\": { \"label\": \"my-box\", \"host\": \"192.0.2.10\", \"port\": 22, \"username\": \"ubuntu\" }\n }\n}\n```\n\nFail if the requested schema is not `2`; do not silently fall back to the ordinary recipe shape.\n\n**Networking → which `target` fields to set** (how _your desktop_ reaches the box — there is no\n`orca serve` URL in SSH mode):\n\n- Public IP / DNS, or a Tailscale/VPN address → `host`; SSH port → `port` (usually 22).\n- Key auth → `identityFile` (add `identitiesOnly: true` if the agent has many keys).\n- Through a bastion → `jumpHost` (a `user@host` ProxyJump) **or** a full `proxyCommand` (e.g. an access\n proxy). Use one, not both.\n- A service port the workspace needs → add entries to `portForwards`.\n- `relayGracePeriodSeconds` (optional): how long Orca keeps the SSH relay alive after the workspace\n detaches before tearing it down; `0` = tear down immediately. Leave it off unless the user wants a\n reconnect grace window.\n\n**Toolchain & agent auth on a persistent (no-snapshot) host — do this ONCE, by hand, before wiring the\nrecipe** (there's no base image to bake; the host _is_ the base). Run the §7f Phase-2 install steps and\nthe §7f Phase-3 ` login --device-auth` **directly over SSH on the host** (interactive, e.g.\n`ssh -t user@host ' login --device-auth'`). After that the host stays ready across workspaces.\n\n```bash\n#!/usr/bin/env bash\nset -euo pipefail\n# resolve from env→state→fallback (default unset optionals to \"\"): ssh_username, host,\n# ssh_port (default 22), identity_file, jump_host, proxy_command, project_root, repo_url, repo_ref\n: \"${identity_file:=}\"; : \"${jump_host:=}\"; : \"${proxy_command:=}\" # avoid set -u aborts on optionals\ngh_token=\"${GH_TOKEN:-${GITHUB_TOKEN:-$(command -v gh >/dev/null 2>&1 && gh auth token 2>/dev/null || true)}}\"\nssh_target=\"${ssh_username}@${host}\"\nssh_opts=(-p \"$ssh_port\"); [ -n \"$identity_file\" ] && ssh_opts+=(-i \"$identity_file\")\n# Why: a fresh host's key isn't in known_hosts; a StrictHostKeyChecking prompt would HANG a\n# non-interactive create. Pre-add the key (or set the option) so it can't block.\nssh-keyscan -p \"$ssh_port\" \"$host\" >> \"$HOME/.ssh/known_hosts\" 2>/dev/null || true\n\n# 1. ensure the repo is present and at the right commit on the host (NO orca serve here)\nssh \"${ssh_opts[@]}\" \"$ssh_target\" \\\n \"GH_TOKEN='$gh_token' GIT_TERMINAL_PROMPT=0 bash -lc '\n set -euo pipefail\n [ -d \\\"$project_root/.git\\\" ] || git clone \\\"$repo_url\\\" \\\"$project_root\\\"\n cd \\\"$project_root\\\" && git fetch origin \\\"$repo_ref\\\" && git checkout -B \\\"$repo_ref\\\" FETCH_HEAD\n '\" >&2\n\n# 2. print the SSH connection block (NO pairingCode, NO orca serve). host/port/username tell Orca's\n# relay how to dial in; identityFile/jumpHost/proxyCommand/portForwards are emitted when set.\nnode -e 'const [host,port,user,idf,jh,pc,root]=process.argv.slice(1);\n const target={ label:\"per-workspace-host\", host, port:Number(port), username:user };\n if(idf) target.identityFile=idf; if(jh) target.jumpHost=jh; if(pc) target.proxyCommand=pc;\n // add target.portForwards=[...] here if the workspace needs forwarded service ports\n console.log(JSON.stringify({ schemaVersion:1, connection:{ type:\"ssh\", projectRoot:root, target } }))' \\\n \"$host\" \"$ssh_port\" \"$ssh_username\" \"$identity_file\" \"$jump_host\" \"$proxy_command\" \"$project_root\"\n```\n\n`suspend`/`resume`/`destroy`: on a persistent host there's usually nothing to tear down — set\n`destroy: none` and omit suspend/resume. (Orca still disconnects/reconnects its own SSH relay on\nsleep/wake/delete — that's separate from these scripts.)\n\nIf the SSH host is instead an **ephemeral/snapshot-capable VM** (your hypervisor, or a cloud VM with\nimage support), keep the §7f Phase-2/3 base-image model for provisioning, but still emit the\n`connection.type:\"ssh\"` block above instead of starting `orca serve`.\n\n### 7h. Worked example — local Docker SSH (SSH connection mode)\n\nLocal Docker can model an ephemeral SSH VM without cloud cost: build a base image with `sshd`, tools,\nrepo prerequisites, and the agent CLI; run an **interactive auth container** once; then `docker commit`\nthat container as the authenticated image used by per-workspace `create`.\n\nKey points:\n\n- Publish container SSH to a random localhost port (`-p 127.0.0.1::22`) and emit\n `connection.type:\"ssh\"` with `host:\"127.0.0.1\"`, that port, `username`, `identityFile`, and\n `identitiesOnly:true`.\n- Generate a repo-local SSH key if needed, but gitignore the private/public key files.\n- **Bake SSH host keys into the base image** (`ssh-keygen -A` at **build** time; at runtime only generate\n if absent). Ephemeral containers all present the **same** host key, so `known_hosts` on `127.0.0.1`\n doesn't churn as the published port rotates across workspaces (otherwise every container's freshly\n generated key collides on `localhost` and trips host-key-changed warnings).\n- The auth image is the Docker equivalent of Phase 3: the **user** runs the agent login **inside** the\n container (you can't drive it — you have no interactive TTY), configures proxy env/config, approves\n hooks, and you commit once they report it's done. On a headless container use the **device-auth** flow\n (§4). Verify login before committing — exit code, or fold stderr and match the exact success line (§4).\n- Do not bind-mount or copy the host's full agent home into the image. Let each container have writable\n agent state; only the committed auth image should carry reusable authenticated state.\n- If committing from an interactive shell, force the runtime entrypoint back to `sshd`:\n `docker commit --change='ENTRYPOINT [\"/usr/local/bin/orca-docker-ssh-entrypoint\"]' …`.\n- `destroy` should read `recipeResult.userData.resourceId` and run `docker rm -f \"$resource_id\"`.\n\nValidation before wiring/live use:\n\n```bash\ndocker image inspect \"$auth_image\" --format '{{json .Config.Entrypoint}}'\ndocker run -d --name \"$name\" -p 127.0.0.1::22 -e \"ORCA_SSH_PUBLIC_KEY=$pubkey\" \"$auth_image\"\ndocker ps -a --filter \"name=$name\"\ndocker logs \"$name\"\nssh -i \"$key\" -p \"$port\" -o IdentitiesOnly=yes user@127.0.0.1 'codex --version'\n```\n\nIf the container exits immediately, inspect logs before the cleanup trap removes it; a committed\ninteractive image with `ENTRYPOINT [\"bash\"]` is a common cause.\n\nAlso confirm the **host key is stable** across containers: the SSH `ssh -i … 127.0.0.1` dial should not\ntrigger a host-key-changed warning when a second container reuses the port. If it does, the host keys\nweren't baked into the base image (see the `ssh-keygen -A` point above).\n\n### 7i. Windows local-side scripts\n\nThe local-side scripts run on the user's desktop. On **Windows**, a bare `.sh` won't execute. Either\nrequire WSL/Git-Bash (and point `orca.yaml` at e.g. `bash ./scripts/orca-vm/.sh` via a `.cmd`\nlauncher), or scaffold PowerShell equivalents. Minimal PowerShell shape:\n\n```powershell\n#requires -Version 5\n$ErrorActionPreference = 'Stop'\n# resolve env→state→fallback; run the provider CLI / ssh the same way;\n# capture provider output; build the result object for the chosen mode and write ONE line of JSON to stdout.\n# Orca-server mode: @{ schemaVersion=1; pairingCode=$pairingCode; projectRoot=$projectRoot; userData=@{...} }\n# SSH mode: @{ schemaVersion=1; connection=@{ type=\"ssh\"; projectRoot=$projectRoot;\n# target=@{ label=$label; host=$host; port=$port; username=$user } } } (see §7g/§7h)\n($result | ConvertTo-Json -Compress -Depth 6)\n# progress/errors → Write-Error / the error stream, never stdout.\n```\n\nThe remote-side commands you run _inside_ the Linux VM stay bash regardless of the desktop OS.\n\n---\n\n## 8. Per-workspace recipe contract (the fast path)\n\nOnce the authenticated snapshot exists, this runs on every workspace create. Define recipes in\n`orca.yaml`:\n\n```yaml\nenvironmentRecipes:\n - id: cloud-sandbox\n name: Cloud Sandbox\n create: ./scripts/orca-vm/cloud-sandbox-create.sh\n suspend: ./scripts/orca-vm/cloud-sandbox-suspend.sh\n resume: ./scripts/orca-vm/cloud-sandbox-resume.sh\n destroy: ./scripts/orca-vm/cloud-sandbox-destroy.sh\n```\n\n`create` runs **locally from the repo root** and prints **one** JSON object to stdout. Its shape depends\non the connection mode chosen in §1:\n\n**Orca-server mode** — boot the env, start `orca serve` in it, and print serve's result:\n\n```json\n{\n \"schemaVersion\": 1,\n \"pairingCode\": \"orca-pairing-code-or-url\",\n \"projectRoot\": \"/absolute/path/to/repo/on/remote\",\n \"userData\": { \"provider\": \"example\", \"resourceId\": \"provider-resource-id\" }\n}\n```\n\nHere `pairingCode` (from `orca serve --recipe-json`) and `projectRoot` are required; `schemaVersion` (`1`)\nand `userData` are optional.\n\n**SSH mode** — do **not** run `orca serve`; print the `connection.type:\"ssh\"` block instead (full shape +\nworked script in §7g). `pairingCode` is **not** used in SSH mode.\n\n**Optional provisioned root** — only for direct SSH and only when explicitly requested. Add\n`checkoutMode: provisioned-root` to the recipe, require `ORCA_RECIPE_RESULT_SCHEMA_VERSION=2`, create\nthe requested `ORCA_REPO_BRANCH` at the pinned `ORCA_REPO_REF_HEAD` commit (use `ORCA_REPO_REF` only\nto fetch that commit) at the returned `projectRoot`, and emit schema version 2 with\n`checkoutMode: \"provisioned-root\"`. All recipes without this field retain the schema-v1 behavior above.\n\nLifecycle hooks (all run locally):\n\n- `create`: required. Prints recipe result JSON.\n- `suspend`: optional. Sleep; reads lifecycle payload on stdin.\n- `resume`: optional. Wake; reads payload on stdin and **prints fresh recipe JSON** (pairing may change).\n- `destroy`: optional unless `destroy: none`. Delete/cleanup; reads payload on stdin.\n\nStart Orca remotely with `orca serve --port \"$PORT\" --project-root \"$ABS_ROOT\" --pairing-address\n\"$EXTERNAL_WSS_URL\" --recipe-json` (exact flags + output in §7c). Set `--pairing-address` to the\nexternally reachable address so the emitted `pairingCode` is reachable; tunneling/port mapping is the\nscript's job.\n\nBackward compatibility: `command`→`create`, `cleanup`→`destroy`, `cleanup: none`→`destroy: none`.\nPrefer the lifecycle names.\n\n---\n\n## 9. Doctor and validation\n\nValidate in two stages — the cheap dry run first, then the live self-test.\n\n### Dry run (free, non-destructive) — always do this first\n\n`orca vm recipe doctor --repo-path --json` validates **static wiring only** — it does\n**not** boot anything. It checks: local-host execution (v1), repo path, recipe id exists,\ncreate/destroy/suspend/resume command paths resolve, suspend/resume are paired, and each script is\nexecutable (POSIX exec bit; skipped on Windows). Fix every failure here before spending any cloud money.\n\n### Live self-test (`--provision`) — diagnose and iterate yourself\n\n`orca vm recipe doctor --repo-path --provision --json` actually runs the recipe end\nto end: it executes `create`, validates the returned recipe JSON, then runs `destroy` to **tear the\nenvironment back down** (so the test leaves nothing running, as long as `destroy` works). It spends real\ncloud money, so get the user's OK **once** before starting — that one approval covers the whole loop\nbelow; do not re-ask before each run.\n\nOn failure, the JSON result includes a `provisionTranscript` with the **complete** captured output of\neach stage so you can self-diagnose without asking the user to relay logs:\n\n```json\n{\n \"ok\": false,\n \"checks\": [{ \"id\": \"recipe.provision\", \"status\": \"fail\", \"message\": \"…\" }],\n \"provisionTranscript\": {\n \"provision\": { \"exitCode\": 0, \"signal\": null, \"stdout\": \"…\", \"stderr\": \"…\", \"parseError\": \"…\" },\n \"destroy\": { \"exitCode\": 0, \"signal\": null, \"stdout\": \"…\", \"stderr\": \"…\" }\n }\n}\n```\n\n**Run it as a loop:** read `provisionTranscript.provision.stderr` / `.stdout` / `.parseError` (and\n`destroy.*`), fix the script, and re-run `--provision` until `ok` is `true` — iterating on your own\nrather than waiting for the user to paste errors. Common reads: a non-empty `stderr` with `exitCode 0`\nplus a `parseError` means `create` ran but printed something other than the single recipe-result JSON on\nstdout (often a stray `echo` — route it to stderr, see §10); a non-zero `exitCode` is a provider/script\nfailure described in `stderr`. Each stream is redacted and capped (head+tail) — large logs keep both the\nsetup context and the failure.\n\nThe self-test cannot see provider-side truth beyond what the scripts print, so still confirm: state has a\npopulated **authenticated** `snapshotId` (Phases 2–3 done), and `destroy` is implemented/tested (or\nexplicitly `none` — in which case the self-test won't tear down, so clean up manually).\n\nFor SSH recipes, also smoke-test the exact emitted target before declaring success: dial the host/port\nwith the identity/proxy settings, run `pwd`, verify the repo path, check the agent binary, and confirm\n`destroy` removes the provider resource/container. For Docker, inspect the auth image entrypoint and do a\nstartup-only `docker run` before the full clone/install path.\n\n---\n\n## 10. Failure modes\n\n- **Build exceeds plan timeout (e.g. Hobby 45m).** Use enough vCPUs and a timeout covering the build;\n else split work or use a higher plan. The cap also limits per-workspace runtime — surface it.\n- **Build exceeds plan RAM.** Build the **headless main only** (drop the renderer) — the biggest fitter.\n- **Private-repo clone hangs/fails.** Wrong/missing token. Use `GIT_ASKPASS` + `GIT_TERMINAL_PROMPT=0`\n so it fails fast instead of prompting.\n- **`GIT_ASKPASS` helper aborts the clone with \"`$1: unbound variable`\".** The `printf`/heredoc that writes\n the helper inside `bash -lc` under `set -u` expanded `$1`/`$GH_TOKEN` at **write** time. Escape them\n (`\\$1`, `\\$GH_TOKEN`) so they land literally and resolve at git-runtime; this also keeps the real token\n out of the file. `rm -f` the helper afterward (§5, §7f).\n- **Agent verified as \"not logged in\" despite a good login.** `codex login status` (and similar) print\n \"Logged in …\" to **stderr**; an stdout-only `grep` misses it. Prefer the status **exit code**; if you\n grep, fold stderr first (`status 2>&1 | grep …`) and match the exact success line — not `grep -qi\n'logged in'`, which also matches \"not logged in\".\n- **Headless agent login hangs.** Plain OAuth `login` starts a loopback callback server on a VM/container\n port the host browser can't reach. Use the **device-auth** flow (`login --device-auth`) — it prints a\n URL + code the user opens on the host.\n- **`known_hosts` host-key churn on local Docker.** Each ephemeral container regenerating its SSH host key\n collides on `127.0.0.1` as the published port rotates. Bake host keys into the base image at build time\n (`ssh-keygen -A`; runtime generates only if absent) so all containers share one stable key (§7h).\n- **Snapshot expired/evicted.** If `create` hits an unknown snapshot id, rerun Phases 2–3 and update\n `snapshotId`.\n- **Agent auth didn't persist.** Confirm `snapshotId` points at the **authenticated** snapshot; re-run\n Phase 3. Warn that short-lived tokens may need periodic re-auth.\n- **Agent auth copied from the host breaks.** Do not bind-mount/copy a full host agent home; sqlite\n files can be unwritable or host-specific, hooks may need approval again, and config may reference\n local-only env vars. Authenticate inside the runtime and snapshot/commit that layer.\n- **Docker auth image exits immediately.** Inspect `docker image inspect … .Config.Entrypoint` and\n `docker logs`. If the image was committed from an interactive shell, reset the entrypoint to the SSH\n entrypoint during `docker commit`.\n- **Leaked paid resource.** Every long script must trap errors and remove the sandbox it created.\n- **`create` emits non-JSON on stdout.** A stray `echo` corrupts the result — stdout is for the final\n JSON only; everything else to stderr. The `--provision` self-test surfaces this as `exitCode 0` + a\n `parseError` with the offending stdout in `provisionTranscript` (§9).\n\n---\n\n## 11. Boundaries\n\n- Don't create accounts, choose plans/regions, or invent scope/project/org/image/billing ids.\n- Don't invent or store credentials; no secrets in `userData`, state, comments, docs, or commits.\n- Don't run paid/long phases (base snapshot, auth, live test) without an explicit OK.\n- Don't hide provider errors behind generic messages — preserve actionable stderr.\n- Don't make Orca own provider lifecycle beyond invoking the configured scripts.\n- Don't commit or create an Orca workspace unless asked.\n"
// oxfmt-ignore
-const ORCHESTRATION_MARKDOWN = "---\nname: orchestration\ndescription: >-\n Use Orca orchestration for structured multi-agent coordination: threaded\n messages, blocking ask/reply flows, task dispatch, worker_done/escalation\n waits, task DAGs, decision gates, coordinator loops, or decomposing work\n across agents. Use `orca-cli` instead for full ownership handoffs, including\n requests phrased as \"hand off\", \"handoff\", \"handover\", \"give this to another\n agent\", or \"another worktree\" when the user did not explicitly ask to\n supervise, monitor, wait for results, or coordinate a DAG. Use `orca-cli` for\n terminal control, lightweight terminal prompts, shell commands, Orca\n worktree management, reading or waiting on terminals, and automation of the\n browser embedded inside Orca. Use Computer Use for external browser windows,\n webviews, Orca app UI, or desktop UI outside Orca's embedded browser only when\n the task requires OS/window-level control such as focus, menus, dialogs,\n coordinates, or screenshots. Use `orca-cli` for Orca's embedded pages and a\n page-automation tool such as Playwright or CDP for external pages.\n---\n\n# Orca Inter-Agent Orchestration\n\nOrchestration is Orca's structured coordination layer for agent messages, task ownership, dispatch state, and worker completion tracking.\n\nUse this skill when coordination state matters. For lightweight terminal prompts or basic worktree/terminal/built-in-browser control, use `orca-cli`.\n\n## Tool Boundary\n\nIf a task says to use Orca orchestration, the coordinator must create or bind a Run, create the Task with `orca orchestration task-create`, then attach the worker with either the preferred `orca orchestration worker-start` composition or the low-level `orca orchestration dispatch --inject` path.\n\nDo not substitute non-Orca subagent tools, generic agent-spawn APIs, or chat-only parallel worker features. Those may create useful workers, but they do not create Orca task/dispatch provenance, injected lifecycle preambles, `worker_done` authority, or decision gates.\n\nBefore claiming a worker was orchestrated, verify the task/dispatch exists:\n\n```bash\norca orchestration task-list --json\norca orchestration dispatch-show --task --json\n```\n\nIf the work was accidentally run outside Orca orchestration, say so plainly. To repair provenance, rerun or revalidate the needed work through a fresh Orca terminal plus injected dispatch; do not retroactively describe the external worker as orchestrated.\n\n## When To Use\n\n- Send/reply/ask between agent terminals with persistent messages.\n- Dispatch structured tasks to workers and wait for `worker_done` or `escalation`.\n- Track task DAGs with dependencies.\n- Run coordinator loops or decision gates.\n\nDo not use orchestration merely because the user says \"hand off\", \"handoff\", \"handover\", \"give this to another agent\", or asks for another worktree/agent/model/effort. Those are full ownership transfers unless the user explicitly asks to supervise, monitor, wait for worker completion/results, coordinate a DAG, use decision gates, or keep a blocking ask/reply loop.\n\n## Preconditions\n\n- `orca status --json` should show a running runtime.\n- `orca` must be on PATH (`orca-ide` on Linux).\n- The orchestration experimental feature must be enabled in Settings > Experimental.\n- `orca orchestration` commands are RPC calls to the running Orca runtime.\n\n## Contract Migration\n\nOrca adopts a live pre-update orchestration assignment into an ordinary Run. Adoption preserves the existing agent process, PTY/session, terminal handle, tab/leaf/pane, worktree or folder workspace, Task, and Dispatch; it never restarts or replaces the worker. The retired scheduler is not revived, and a newly created attempt uses the current grammar.\n\nTreat the authority label on injected or formatted messages as definitive:\n\n- `[LEGACY COMPATIBILITY]` is live and attested. Run only the exact supported command printed with the message, using the same CLI executable and arguments that the original prompt supplied.\n- `[LEGACY RECOVERY REPLAY — MAY HAVE BEEN SEEN]` is one bounded, at-least-once cutover replay. Process it idempotently and acknowledge it only through the exact displayed guidance.\n- `[LEGACY READ-ONLY]` is inspection-only. It has no reply, acknowledgment, or lifecycle action.\n- An unlabeled current message uses the current guide and current grammar.\n\nAn explicitly selected current Run, attested current Run binding, current Dispatch, or federated attachment takes precedence over legacy fallback. A retained adoption record alone never turns a current command into a legacy call.\n\nDatabase provenance, an old-looking terminal, or a legacy Run ID does not prove mutation authority. If the runtime cannot prove liveness, principal ownership, capability, or the exact legacy contract, it degrades to read-only inspection and must not fall back to local execution. Exact recovery may restore the already-live PTY once in its original inactive background tab. It must not spawn, write, signal, stop, switch, focus, split, or inject a terminal. Loss of lifecycle authority does not invalidate the existing assignment, process, or filesystem work.\n\nCompatibility retries have narrow guarantees. A pending ask, a reply, a final Dispatch settlement, and a consuming check have durable recovery identities. A-era heartbeat and escalation calls remain at-least-once across a manual A-to-B retry because identical later signals may be intentional. If an A-era ask may already have been answered, run the exact non-consuming recovery check printed by the runtime first; after its answer is printed and acknowledged, a new invocation with the same question creates a new question. Never guess among multiple identical question threads.\n\nWhen a compatibility or recovery command returns structured next-step arguments, run those exact arguments with the same CLI executable. The arguments intentionally omit the executable name so the guidance works with `orca`, `orca-ide`, `orca-dev`, or another configured Orca CLI command. Do not translate the command from memory, broaden its recipient, or retry it as a current mutation unless the returned guidance explicitly says to.\n\nOn packaged Windows, a legacy ask uses a two-step commit/resume protocol. The initial command durably commits the question, prints its exact `ask --resume ` command, and exits with launcher status `75`; it does not wait for the answer. Run that exact resume command after the launcher or update boundary. Resume is idempotent and read-oriented: it waits for the already-committed question and does not create another one. For a WSL process that received compatibility proof at launch, use the printed executable `orca-ide` WSL resume command so the same distro and packaged launcher authority are preserved; do not substitute a PATH-resolved local CLI. Older WSL processes that never received the hidden launch token remain lifecycle read-only after the update, even while their terminal and filesystem work continue.\n\nLegacy inspection remains available without consuming mail:\n\n```bash\norca orchestration run-list --json\n# run_legacy_local is an empty audit tombstone after adoption.\norca orchestration run-show --id run_legacy_local --json\n# In run-list, find the ordinary Run whose objective is:\n# \"Recovered orchestration work from a contract update\"\norca orchestration run-show --id --json\norca orchestration task-list --run --json\norca orchestration inbox --full --json\norca orchestration check --terminal --peek --format --json\norca terminal read --terminal --json\norca terminal wait --terminal --for tui-idle --timeout-ms 60000 --json\n```\n\nIf the original coordinator is unavailable or cannot prove its retained authority, a current coordinator may explicitly take over the adopted Run from its own live agent terminal:\n\n```bash\norca orchestration run-use --id --takeover-legacy --json\norca orchestration check --run --json\n```\n\nTakeover fences only the old coordinator, binds the current one, and moves pending worker mail into current Run Delivery. It is bound to the authenticated invoking terminal; `--from` cannot name another coordinator. Live legacy workers keep their original Tasks, Dispatches, processes, filesystems, and old prompt commands; their later questions, escalations, and completion reports route to the current coordinator. Do not use takeover while the original coordinator is still actively coordinating, because its later lifecycle mutations are rejected.\n\nDo not launch a replacement editor merely because the desktop app or runtime was updated. If adoption cannot prove continuing authority, keep the original worker as the only editor until it reaches a stable handoff point, then use a new current Dispatch in a conflict-free placement for any remaining work.\n\n## Ownership\n\nNew orchestration messages and tasks belong to one explicitly bound Run. A Run is only a durable namespace and coordinator inbox; it never schedules or places workers. Lifecycle authority comes from the active Dispatch, and terminal handles remain routing metadata rather than durable identity. Send `worker_done` and `heartbeat` from the worker's own terminal; Orca routes them to that Dispatch's Run.\n\nClassify inherited context before sending lifecycle messages:\n\n- Coordinated subtask: a live coordinator owns the DAG and waits on this dispatch. Follow the preamble exactly, including `worker_done`, heartbeat/status, `ask`, and `escalation`.\n- Full handoff means ownership transfer, not supervised dispatch. The original actor is not monitoring a DAG, so do not create lifecycle obligations unless the user explicitly asks you to supervise.\n- Classify requests containing \"hand off\", \"handoff\", \"handover\", \"give this to another agent\", \"give this to another worktree\", \"another agent\", or \"another worktree\" as full handoffs by default, even when the user names a custom model or reasoning effort.\n- Use supervised orchestration only when the user explicitly asks you to \"supervise\", \"monitor\", \"wait\", \"track completion\", \"wait for worker_done\", return results, coordinate a DAG, use a decision gate, or manage ask/reply flow.\n- Do not use `orca orchestration dispatch --inject` for full handoffs. It injects a coordinator preamble that tells the worker to send `worker_done`, heartbeat, and `ask` messages, then end its turn under the original terminal's dispatch lifecycle.\n- Do not run `orca orchestration task-create`, `orca orchestration dispatch --inject`, or `orca orchestration check --wait` for full handoffs. Do not peek at terminal output after prompt delivery to monitor progress.\n- A review-only `worker_done` reports findings; it does not authorize coordinator file edits. After a review-only completion, synthesize findings, ask a decision gate if ownership is unclear, and dispatch or hand off fixes unless the user explicitly asked the coordinator to own fixes.\n- If the user's plan names a next owner agent (for example, \"then use opencode to create a PR\"), post-review corrections and PR prep belong to that named owner. The coordinator routes, synthesizes, asks decision gates when needed, and supervises; the named owner edits files and creates the PR.\n\nIf unclear, inspect orchestration state before sending lifecycle messages:\n\n```bash\norca orchestration task-list --json\norca terminal list --json\n# If inherited context includes a task id:\norca orchestration dispatch-show --task --json\n```\n\n## Messaging\n\n```bash\norca orchestration send --subject [--to ] [--from ] [--body ] [--type ] [--priority ] [--thread-id ] [--payload ] [--json]\norca orchestration check [--terminal ] [--ack ] [--peek|--all] [--types ] [--format] [--wait] [--timeout-ms ] [--json]\norca orchestration reply --id --body [--from ] [--json]\norca orchestration ask (--question |--resume ) [--options ] [--timeout-ms ] [--from ] [--json]\norca orchestration inbox [--limit ] [--json]\n```\n\nRules:\n\n- Omit `--from` unless impersonating another terminal; Orca auto-resolves it from the current terminal.\n- A coordinator `check` returns the bound Run's oldest FIFO Delivery (up to 50 messages) and replays that exact batch until `--ack `. Process every message before acknowledging; `check --ack --wait` acknowledges, checks, and waits in one operation.\n- Use `--peek` and `--all` only for read-only history/debugging. Type filters decide when a waiter wakes; the returned actionable Delivery is still the oldest full batch.\n- Use `dispatch:` for coordinator guidance to one supervised worker. Orca routes that stable address locally or through the connected-server relay; do not substitute a remote terminal handle.\n- Terminal handles remain appropriate for low-level pre-Dispatch messaging. Prefer `agentTerminalHandle` from the create response, fall back to `startupTerminal.handle` for older runtimes, then re-resolve with `orca terminal list --worktree ... --json` if missing or stale. Continue with the replacement handle only; never dual-send to old and new handles.\n- `terminal list --json` omits `visualLayouts` because handle recovery does not need topology. Add `--include-visual-layouts` only for explicit tab and pane inspection.\n- `orca orchestration check --peek --format --json` returns locally formatted unread mail without consuming it; it never writes to terminal input or remotely wakes another terminal. Use `orchestration dispatch --inject` to deliver a tracked task, or `terminal send` when an existing agent needs a free-form prompt.\n- While supervising workers manually, use `check --wait --types worker_done,escalation,question --timeout-ms ` instead of sleep/poll loops. Process the whole Delivery, reply to `question` messages with `orca orchestration reply --id --body --json`, then acknowledge and keep waiting.\n- `check --json` prints exactly one JSON document on stdout. While `--wait` blocks it also prints keepalive lines (`{\"_keepalive\":true,...}`) to stderr so you can tell the process is alive; those are never on stdout. Do not merge the streams before a parser — `check --wait --json 2>&1 | ` fails with \"Extra data: line 2\". Pipe stdout only.\n- Treat a `check --wait` timeout or `{count:0}` as a checkpoint, not a worker failure. Long coding tasks routinely run 15-60 minutes; keep using rolling waits unless you receive `worker_done`/`escalation`, the terminal exits or disappears, or the user explicitly asks you to stop.\n- Heartbeats and visible terminal activity mean the worker is alive, not done. Do not stop, close, kill, or restart a worker just because it has not produced a completion message yet.\n- Use `ask` when a worker needs a blocking answer from the coordinator; it defaults to the active Dispatch's Run. Timeout or disconnect leaves the question pending, so resume by its original message ID instead of asking again.\n- `check --wait` returns one bounded Delivery, not every future completion. Process every message, acknowledge it, then keep waiting until every expected Dispatch settles.\n- Group addresses include `@all`, `@idle`, `@claude`, `@codex`, `@opencode`, `@gemini`, `@droid`, `@grok`, `@cursor`, and `@worktree:`.\n- Message types include `status`, `dispatch`, `worker_done`, `merge_ready`, `escalation`, `handoff`, `question`, `decision_gate` (legacy/gates), and `heartbeat`.\n- Use group addresses only for messages that are genuinely useful to many terminals, such as `status` broadcasts or intentional fan-out questions. Do not send dispatch lifecycle messages to groups.\n- `worker_done` belongs to the active Dispatch and defaults to its Run mailbox; never target a group.\n- A valid `worker_done` for the active `taskId` + `dispatchId` marks the task and dispatch completed automatically. Do not follow it with `task-update --status completed`; reserve manual updates for explicit recovery or overrides.\n- `heartbeat` is also Dispatch-scoped. Include both IDs and omit `--to` so Orca uses the owning Run; use `status` for broad progress updates.\n\n## Tasks And Dispatch\n\nA Run is the namespace/inbox, a Task is the work item, and a Dispatch assigns one Task attempt to a terminal. Create or bind a Run once before the common loop.\n\n```bash\norca orchestration run-create --objective --json\norca orchestration task-create --spec [--deps ] [--parent ] [--json]\norca orchestration task-list [--status ] [--ready] [--brief] [--json]\norca orchestration task-update --id --status [--result ] [--json]\norca orchestration dispatch --task --to [--from ] [--inject] [--json]\norca orchestration dispatch-show --task [--json]\n```\n\nTask statuses: `pending`, `ready`, `dispatched`, `completed`, `failed`, `blocked`.\n\nDispatch rules:\n\n- `--inject` sends the task spec plus preamble into a recognized agent CLI so it can report `worker_done`.\n- If the target is a bare shell, omit `--inject`, dispatch for tracking if needed, then send the prompt manually with `orca terminal send --terminal --text --enter --json`.\n- After 3 consecutive failures on one task, the dispatch context circuit-breaks and the task is marked failed.\n- Use `task-list --brief --json` for coordinator sweeps; it collapses whitespace and caps each echoed spec at 160 characters (`spec_truncated` marks shortened rows). Omit `--brief` when the full spec is required, or when an older CLI rejects it as an unknown flag.\n\n## How deep workers can nest\n\nA dispatched worker normally cannot dispatch sub-workers. Attempting it fails with\n`nested_worker_depth_exceeded` and a message telling the worker to complete the task\nitself. Do that — do not try to route around it.\n\nThe limit is a number, not an on/off switch. `Settings -> Orchestration -> Nested worker depth`\nsets how many generations are allowed:\n\n- `1` (default): a coordinator dispatches workers; those workers do not dispatch.\n- `2`: workers may dispatch one further generation.\n\nDepth is counted from the terminal that issues the command, not from the Run. Creating a\nnew Run does not reset it — a worker that runs `run-create` then `worker-start` is still a\nworker, and still counted. This is the part that changed: the old behaviour rejected\nsub-dispatch only because a worker's terminal was not bound to a Run, so creating a Run was\nenough to slip past it.\n\nTwo limits worth knowing:\n\n- **It is a guardrail, not a security boundary.** A caller that declares another terminal's\n handle while its own launch evidence is unverifiable (an ordinary restored terminal, for\n example) can be counted as that terminal instead. Orca does not treat workers as hostile.\n- **It applies while a Dispatch is active.** After `worker_done`, or after a coordinator\n settles the task, the terminal is no longer a worker and is counted as a root again. The\n process may still be alive; that is the documented boundary, not an accident.\n\n## Preferred Supervised Worker Loop\n\nUse `worker-start` for the normal supervised path. It composes the existing worktree, terminal, readiness, and dispatch primitives while returning exact created/reused effects. Agents still choose placement and concurrency; Orca does not schedule workers or infer conflicts.\n\nCreate the Run and every independent Task first, then start all independent workers before waiting:\n\n```bash\norca orchestration run-create --objective \"\" --json\norca orchestration task-create --spec \"\" --json\norca orchestration task-create --spec \"\" --json\norca orchestration worker-start --task --worktree current --agent codex --json\norca orchestration worker-start --task --worktree current --agent claude --json\n```\n\n`current` and exact existing worktrees create a fresh agent terminal and do not rerun setup. Reuse an existing agent only with `--terminal `.\n\nFor a per-invocation Claude, Codex, or Cursor launch, pass an opaque provider model id with `--model`; add `--effort` only when that agent/model supports the level. These options apply only to fresh agent terminals, override general agent default arguments, and are reported under `launch.requested` and `launch.effective` in the receipt:\n\n```bash\norca orchestration worker-start --task --worktree current --agent claude --model opus --effort high --json\n```\n\n`--effort` requires `--model`, and neither option can combine with `--terminal`. A connected worker server must advertise launch-preference support before Orca forwards either option.\n\nFor a new worktree, setup runs by default and agent-first creation reuses the returned startup agent terminal:\n\n```bash\norca orchestration worker-start --task --worktree new-child --name --agent codex --setup run --json\n# Independent/top-level:\norca orchestration worker-start --task --worktree new-top-level --name --agent codex --setup run --json\n```\n\nSetup normally starts alongside the agent. Only a repository explicitly configured with `wait-for-setup` delays agent launch until setup succeeds. Use `--setup skip` or `--setup inherit` only for a concrete reason.\n\nRead the returned receipt before continuing: `ready` plus setup `running` is normal for start-immediately, while wait-for-setup returns setup `succeeded` before accepting task input. A failed or unknown start exits nonzero; inspect its `stage`, `effects`, and `residualResources` instead of guessing or automatically retrying. A wait-for-setup timeout can honestly leave setup `running`, which is not proof of failure.\n\nTo run the worker on another connected Orca server, add `--on `. The Run and Tasks remain authoritative on the current server; later commands route by Dispatch ID, so never repeat `--on`:\n\n```bash\n# Mac Run home -> Windows worker (the reverse is identical from a Windows Run home)\norca orchestration worker-start --task --on windows --worktree new-top-level --repo --name --agent codex --setup run --json\norca orchestration worker-show --dispatch --json\norca orchestration worker-read --dispatch --limit 50 --json\norca orchestration send --to dispatch: --subject \"Follow-up\" --body \"\" --json\n```\n\nRemote `current` and `new-child` are intentionally invalid because those words are ambiguous across servers. Use an exact discovered remote worktree selector or `new-top-level` with an explicit remote repo selector.\n\nThe follow-up is structured inbox mail, not prompt injection. The worker's next\n`orchestration check` receives it even when the Dispatch is on another connected Orca server.\n\n`worker-read` defaults to `--source auto`: Orca returns the exact hook-reported Codex, Claude, OpenClaude, or Grok transcript when it can prove the worker session, otherwise it returns bounded terminal output with `source: \"terminal\"` and a typed `fallbackReason`. Continue with the returned top-level `cursor`; it stays pinned to that exact source. If Orca reports `source_changed`, start a fresh read without the old cursor. Never supply or guess a provider session ID or transcript path.\n\nWait until every expected Dispatch settles, not for a fixed number of batches:\n\n```bash\norca orchestration check --wait --types worker_done,escalation,question --timeout-ms 900000 --json\n# Process every message. For each accepted worker_done that is not immediately reused:\norca orchestration worker-release --dispatch --json\n# Acknowledge only after every message and required release decision is handled:\norca orchestration check --ack --wait --types worker_done,escalation,question --timeout-ms 900000 --json\n```\n\nAfter processing each accepted `worker_done`, choose the terminal's next owner before you acknowledge the Delivery or wait again. If the same exact agent has an immediate follow-up Task, read the `worker.agent_terminal_handle` field of `worker-show --dispatch --json`, then run `orca orchestration worker-start --task --terminal --json` so Orca transfers cleanup ownership to the new Dispatch. Otherwise run `orca orchestration worker-release --dispatch --json`.\n\nRun `worker-release` after both succeeded and failed `worker_done` reports unless the user explicitly asked to keep that worker live. Release is post-completion cleanup, not cancellation: Orca first preserves inspectable output, then closes only the exact agent terminal owned by that settled Dispatch. Reused or pre-existing terminals, setup terminals, coordinators, active workers, user-taken-over terminals, and identities Orca cannot prove are retained. If the user explicitly asks to keep the live terminal for debugging, record that exception with `orca orchestration worker-retain --dispatch --json` instead of silently skipping cleanup. When the user is finished, the same Dispatch can be passed to `worker-release`, which clears the requested retention and releases the terminal.\n\nDo not release a worker because of a timeout, TUI idle state, heartbeat, status, question, escalation, or rejected/stale `worker_done`. If release returns `release_pending` or `release_unknown`, do not substitute `terminal close`; follow the exact recovery action in the receipt. A replayed Delivery may repeat `worker-release` safely.\n\nWorkers report exactly once using the IDs and capability injected by Orca; they do not supply Run/server/terminal identity:\n\n```bash\norca orchestration send --type worker_done --subject \"\" --body \"\" --task-id --dispatch-id --outcome succeeded --files-modified \"path/a,path/b\" --json\n# On failure, use --outcome failed; never encode failure only in prose.\n```\n\nA worker question defaults to its owning Run. Timeout leaves it pending:\n\n```bash\norca orchestration ask --question \"\" --options \"yes,no\" --timeout-ms 600000 --json\norca orchestration ask --resume --timeout-ms 600000 --json\n# Coordinator:\norca orchestration reply --id --body \"\" --json\n```\n\nRecovery is conditional, never a fixed destructive sequence:\n\n- The response was lost and named no Dispatch: run `orca orchestration request-show --request --json` first. It is read-only. `completed` means the mutation already took effect. `pending` means the original mutation is still running or Orca restarted before recording its outcome. For either state, replaying the original command with `--retry-request ` reuses the same operation identity so Orca can replay, join, or safely recover it without starting a separate duplicate. `absent` means this runtime holds no receipt under your caller identity and is not proof that nothing happened; inspect the affected state before deciding whether to retry.\n- `worker-show --dispatch ` says `ready`: keep waiting or read bounded output.\n- It proves `failed` or `stopped`: start a replacement with `worker-start --task --retry-of ` plus an explicit `--on`/`--worktree` and `--agent`/`--terminal` choice. Retry does not silently inherit placement.\n- It remains `outcome_unknown`: either `worker-stop --dispatch ` and inspect again, or explicitly `worker-abandon --dispatch ` while accepting that resources may still be live. Abandon performs no remote, process, or filesystem action.\n- `worker-stop` closes only the exact supervised agent terminal. It never deletes the worktree, setup terminal, configured tabs, or unrelated processes.\n\nLow-level `worktree create`, `terminal create`, and `dispatch --inject` remain valid recipes for custom argv or topology that `worker-start` does not express.\n\n`dispatch --inject` deliberately keeps an operator-started terminal unsupervised: it never creates a `worker_dispatches` row and `worker-stop`/`worker-abandon` never close that process. The dispatch context is still authoritative, so `worker-show`, `worker-read`, and `worker-list` report it as `unsupervised`; settled `worker-retain` and `worker-release` report `retained` with `no_owned_resource` and take no process action. Use `worker-start --terminal ` when supervision and worker lifecycle state are required.\n\n## Gates And Legacy Inspection\n\n```bash\norca orchestration gate-create --task --question [--options ] [--json]\norca orchestration gate-resolve --id --resolution [--json]\norca orchestration gate-list [--task ] [--status ] [--json]\n```\n\nUse `ask` for worker-to-coordinator questions; it creates a `question` message that the coordinator answers with `reply`. Use `gate-create` only for coordinator-managed task DAG decisions, not for answering a worker's `ask`.\n\n`coordinator-start`, `coordinator-stop`, `run`, and `run-stop` are retired scheduler commands. They perform no effects and return the current-skill recovery action. They are not aliases for lightweight Run creation or binding.\n\nRecovery only: `orca orchestration reset --tasks|--messages|--all --json` clears the selected local orchestration database state. Do not run it during active coordination unless explicitly abandoning that state.\n\n## Full Handoffs\n\nFor full ownership transfer, use non-lifecycle terminal/worktree commands and then stop monitoring unless the user asks for supervision.\n\nTreat these as full handoff requests by default: \"hand off\", \"handoff\", \"handover\", \"give this to another agent\", \"give this to another worktree\", \"send this to another agent\", \"another agent\", \"another worktree\", or \"launch another agent to own this.\" Custom model or reasoning effort words such as `gpt-5.5`, `high`, or `xhigh` do not make the handoff supervised.\n\nSupervised orchestration remains available only when the user explicitly asks for supervision or coordination: \"supervise\", \"monitor\", \"wait for worker_done\", \"wait for results\", \"track completion\", \"DAG\", \"decision gate\", \"ask/reply\", or \"coordinate workers.\"\n\nDo not run `orca orchestration task-create`, `orca orchestration dispatch --inject`, or `orca orchestration check --wait` for full handoffs. `task-create` is also forbidden because it records coordinator-owned tracking state; if a task row is needed, the user asked for supervised orchestration. Do not create a `taskId`/`dispatchId`, inject a lifecycle preamble, wait for completion, or read the worker terminal after prompt delivery except to avoid losing the initial prompt.\n\nNew top-level worktree handoff:\n\n```bash\norca worktree create --name --no-parent --agent codex --prompt \"\" --setup run --json\n```\n\nBefore creating a new worktree from an active feature branch, decide and state whether the desired Orca lineage is child or top-level. Use child worktree lineage only when the new work is conceptually stacked under or dependent on the active worktree. For independent repo-wide fixes, standalone feature work, or unrelated follow-up tasks, create a top-level worktree with `--no-parent`.\n\nExisting terminal handoff:\n\n```bash\norca terminal send --terminal --text \"\" --enter --json\n```\n\nCustom Codex model/effort handoff:\n\n`orca worktree create --agent codex --prompt ...` launches the known Codex agent but does not accept Codex-specific `--model` or `-c model_reasoning_effort=...` arguments. When the user asks for a specific Codex model or effort, create the independent worktree first, launch Codex with the requested command in that worktree, wait only for TUI readiness if prompt delivery would otherwise race startup, send the prompt, and stop.\n\nThe two-step custom-argv path cannot enforce a repository's explicit `wait-for-setup` startup policy because the later `terminal create` is not the startup owned by `worktree create`. Use it only when the repository starts agents immediately. If the repository requires `wait-for-setup`, use an agent-first configured launcher that can preserve sequencing, or stop and ask rather than silently bypassing the policy.\n\nNote: when no repo default-terminal configuration supplies a primary terminal, bare create opens a fallback shell before `terminal create` adds the agent. Configured default tabs are materialized instead and may run real commands. Prefer `--agent` whenever custom argv is not required. With the two-step path, target only the agent handle; close a prior terminal only after `terminal list` or `terminal show` confirms it is an unused shell.\n\nUse the exact full `::` worktree id returned by `orca worktree create --json`; a bare repo id cannot target the new worktree.\n\n```bash\norca worktree create --name --no-parent --setup run --json\norca terminal create --worktree id: --title --command 'codex --model gpt-5.5 -c model_reasoning_effort=\"xhigh\"' --json\norca terminal wait --terminal --for tui-idle --timeout-ms 60000 --json\norca terminal send --terminal --text \"\" --enter --json\n```\n\nWait only for `tui-idle` when needed to avoid losing the prompt. Do not monitor task completion.\n\n`--no-parent` only controls Orca lineage; it does not choose the Git base. If the work should start from the repo default base, omit `--base-branch` so Orca uses that default, or explicitly pass the repo default base (`origin/main`, `origin/master`, or the `orca repo show --repo --json` value); never base it on the current feature branch unless the user explicitly asks for stacked work or \"branch from current\". Put current-branch context in the prompt instead.\n\n## Worker Terminals\n\nChoose the worker location before creating a terminal. `Fresh worker` means a fresh agent session, not a new git worktree. For parallel work, create one fresh agent terminal per worker in the same required worktree, falling back to the active worktree when none is named. If the task says current worktree only, depends on uncommitted files/artifacts, or must validate/PR the current branch, keep every worker in the active worktree:\n\n```bash\norca terminal create --worktree active --title --command \"codex\" --json\norca terminal wait --terminal --for tui-idle --timeout-ms 60000 --json\norca orchestration dispatch --task --to --inject --json\n```\n\nReuse an idle agent in the required worktree only if the prompt allows reuse; otherwise create a fresh terminal there. Create a new worktree only when the user explicitly requests one or a concrete checkout or filesystem conflict makes sharing unsafe or impossible; if the user did not request it, state that conflict before running `worktree create`. Independent tasks, parallel execution, convenience, or a preference for separate checkouts are not isolation requirements.\n\nWhen a new worktree is allowed, use child lineage for isolated work that is stacked under or dependent on the active worktree, and use `--no-parent` when it is not stacked. Decide the Git base separately: `--no-parent` makes the worktree top-level in Orca, while omitted `--base-branch` uses the repo default base.\n\nFor every new worktree, pass `--setup run` so any configured repository setup hook runs. This does not mean waiting for setup before agent launch: preserve the repository's startup policy, whose default starts setup and the agent side by side. Use `--setup skip` or `--setup inherit` only when there is a concrete task-specific reason, and state that reason before creating the worktree. This rule does not rerun setup for current or existing worktrees.\n\n```bash\norca worktree create --name --agent codex --setup run --json\n# or: --agent claude | omp | pi | grok | ...\n# Read from agentTerminalHandle, falling back to startupTerminal.handle.\norca terminal wait --terminal --for tui-idle --timeout-ms 60000 --json\norca orchestration dispatch --task --to --inject --json\n```\n\nFor new-worktree workers, read the id and `agentTerminalHandle` from `worktree create`, falling back to `startupTerminal.handle` for older runtimes. Use that as the sole worker handle when present; otherwise use `terminal list` to resolve the agent handle. Omit `--repo` only inside an Orca-managed worktree; otherwise pass `--repo `.\n\n**For an allowed new worktree, use agent-first:** `--agent` reveals the new worktree and launches the selected agent **in its first terminal**, without adding a separate fallback shell for that worker. Pass `--setup run`; repo setup and default-terminal settings may add intentional tabs or splits. Do **not** run bare `worktree create` and then `terminal create --command ` for the same worker when agent-first create is available: without configured default tabs, that two-step path leaves a fallback shell + agent pair. Only use it when custom agent argv is required (for example Codex model/effort flags) or when an older CLI rejects `--agent`; if you must, message only the agent handle. Configured default tabs are intentional surfaces, so close a prior terminal only after `terminal list` or `terminal show` confirms it is an unused shell. Do not run `worktree create` when the task must stay in the current worktree.\n\nUse `orca worktree create --prompt ...` or `orca terminal send ...` for full handoffs or untracked/lightweight prompts. Those paths do not attach `taskId`/`dispatchId`; the worker should not send lifecycle messages unless the prompt supplies a live orchestration preamble.\n\nSidebar lineage and orchestration lifecycle are related but not identical. A same-worktree worker may appear as a peer under that worktree in the sidebar while remaining a child dispatch in orchestration state; only an actual child worktree creates visible parent/child worktree lineage.\n\nOther terminal commands coordinators often need:\n\n```bash\norca terminal list [--worktree ] [--include-visual-layouts] [--json]\norca terminal create [--worktree ] [--title ] [--command ] [--json]\norca terminal split --terminal [--direction horizontal|vertical] [--command ] [--json]\norca terminal wait --terminal --for tui-idle --timeout-ms --json\norca terminal read --terminal --json\norca terminal send --terminal --text --enter --json\n```\n\nIf an older CLI rejects `worktree create --agent`, create the worktree normally, then run `orca terminal create --worktree --command \"codex\" --json` or `--command \"claude\"`.\n\nWait for `tui-idle` before dispatching. Always pass `--timeout-ms`; real coding tasks can take 15-60 minutes. During supervision, use rolling `check --wait` windows. If a window returns no matching message, inspect `task-list`, `terminal read`, or `terminal wait --for tui-idle` as a liveness checkpoint; if the terminal is still working or producing activity, keep waiting instead of retrying the task.\n\n## Agent Guidance\n\n- Workers with a valid live preamble must send `worker_done` exactly once from their own terminal with an explicit `--outcome succeeded` or `--outcome failed`:\n `orca orchestration send --type worker_done --subject \"\" --body \"<3-sentence summary: what you did, what you found, what's left>\" --task-id --dispatch-id --outcome succeeded --files-modified \"path/a\" --report-path \"\" --json`\n- A failed outcome is still a terminal report, but Orca records both the Dispatch and Task as failed. Never encode failure only in the subject/body.\n- After sending `worker_done`, end that dispatched turn and idle at the agent prompt. Do not autonomously start more work, poll, or attempt to close the terminal yourself. A direct user instruction takes precedence and starts ordinary user-owned work: follow it without coordinator approval or a fresh Dispatch, never refuse it because of worker/coordinator roles, and do not reuse the settled Dispatch's lifecycle IDs. A coordinator-supervised follow-up still arrives with a fresh preamble + TASK block.\n- For long tasks, send heartbeat/status only when the preamble asks for it, including both IDs:\n `orca orchestration send --type heartbeat --subject \"alive\" --payload '{\"taskId\":\"\",\"dispatchId\":\"\",\"phase\":\"implementing\"}' --json`\n- If blocked before completion, use `ask`; use `escalation` only when ownership is valid and the coordinator must intervene.\n- Treat preambles inherited through terminal history or full handoffs as stale unless the current prompt explicitly keeps that coordinator in the loop.\n- Coordinators must account for every settled worker terminal before waiting again or ending the turn: immediately reuse the exact worker for a new Dispatch, explicitly retain it at the user's request with `worker-retain`, or run `worker-release`. Do not leave a completed worker live merely to inspect output; released workers remain readable through `worker-read`.\n- Coordinators should use `task-list --ready` as external memory, dispatch parallel waves, and avoid dependency chains deeper than 3-4 steps.\n\n## Example\n\n```bash\norca terminal create --worktree active --title login-css-worker --command \"claude\" --json\norca terminal wait --terminal --for tui-idle --timeout-ms 60000 --json\norca orchestration task-create --spec \"Fix the login button CSS\" --json\norca orchestration dispatch --task --to --inject --json\norca orchestration check --wait --types worker_done,escalation,question --timeout-ms 900000 --json\n```\n\n## Next Action\n\nCoordinator: confirm `orca status --json`, create or bind a Run, inspect `task-list`/`dispatch-show` if inheriting state, then use the explicit supervised loop (`task-create` -> `worker-start` -> `check --wait`). Use low-level terminal creation plus `dispatch --inject` only when the composed start does not express the needed topology. After every accepted `worker_done`, either transfer the exact terminal to an immediate follow-up Dispatch or run `worker-release` before the next wait.\n\nWorker: if the current prompt contains a live dispatch preamble, do the task, use `ask` for blocking questions, and send `worker_done` once with the required payload. If the preamble is stale or absent, do not send lifecycle messages; inspect state or treat the prompt as an ordinary handoff.\n"
+const ORCHESTRATION_MARKDOWN = "---\nname: orchestration\ndescription: >-\n Use Orca orchestration for structured multi-agent coordination: threaded\n messages, blocking ask/reply flows, task dispatch, worker_done/escalation\n waits, task DAGs, decision gates, or coordinator loops. Use `orca-cli`\n instead for full ownership handoffs, including requests phrased as \"hand\n off\", \"handoff\", \"handover\", \"give this to another agent\", or \"another\n worktree\" when the user did not explicitly ask to supervise, monitor, wait\n for results, or coordinate a DAG. Use `orca-cli` for terminal control,\n lightweight terminal prompts, shell commands, Orca worktree management,\n reading or waiting on terminals, and the Orca embedded browser. Use Computer\n Use for external browser windows, webviews, Orca app UI, or desktop UI\n outside Orca's embedded browser only when the task requires OS/window-level\n control such as focus, menus, dialogs, coordinates, or screenshots. Use\n `orca-cli` for Orca's embedded pages and a page-automation tool such as\n Playwright or CDP for external pages.\n---\n\n# Orca Inter-Agent Orchestration\n\nOrchestration is Orca's structured coordination layer for agent messages, task ownership, dispatch state, and worker completion tracking.\n\nUse this skill when coordination state matters. For lightweight terminal prompts or basic worktree/terminal/built-in-browser control, use `orca-cli`.\n\n## Tool Boundary\n\nIf a task says to use Orca orchestration, the coordinator must create or bind a Run, create the Task with `orca orchestration task-create`, then attach the worker with either the preferred `orca orchestration worker-start` composition or the low-level `orca orchestration dispatch --inject` path.\n\nDo not substitute non-Orca subagent tools, generic agent-spawn APIs, or chat-only parallel worker features. Those may create useful workers, but they do not create Orca task/dispatch provenance, injected lifecycle preambles, `worker_done` authority, or decision gates.\n\nBefore claiming a worker was orchestrated, verify the task/dispatch exists:\n\n```bash\norca orchestration task-list --json\norca orchestration dispatch-show --task --json\n```\n\nIf the work was accidentally run outside Orca orchestration, say so plainly. To repair provenance, rerun or revalidate the needed work through a fresh Orca terminal plus injected dispatch; do not retroactively describe the external worker as orchestrated.\n\n## When To Use\n\n- Send/reply/ask between agent terminals with persistent messages.\n- Dispatch structured tasks to workers and wait for `worker_done` or `escalation`.\n- Track task DAGs with dependencies.\n- Run coordinator loops or decision gates.\n\nDo not use orchestration merely because the user says \"hand off\", \"handoff\", \"handover\", \"give this to another agent\", or asks for another worktree/agent/model/effort. Those are full ownership transfers unless the user explicitly asks to supervise, monitor, wait for worker completion/results, coordinate a DAG, use decision gates, or keep a blocking ask/reply loop.\n\n## Preconditions\n\n- `orca status --json` should show a running runtime.\n- `orca` must be on PATH (`orca-ide` on Linux).\n- The orchestration experimental feature must be enabled in Settings > Experimental.\n- `orca orchestration` commands are RPC calls to the running Orca runtime.\n\n## Contract Migration\n\nOrca adopts a live pre-update orchestration assignment into an ordinary Run. Adoption preserves the existing agent process, PTY/session, terminal handle, tab/leaf/pane, worktree or folder workspace, Task, and Dispatch; it never restarts or replaces the worker. The retired scheduler is not revived, and a newly created attempt uses the current grammar.\n\nTreat the authority label on injected or formatted messages as definitive:\n\n- `[LEGACY COMPATIBILITY]` is live and attested. Run only the exact supported command printed with the message, using the same CLI executable and arguments that the original prompt supplied.\n- `[LEGACY RECOVERY REPLAY — MAY HAVE BEEN SEEN]` is one bounded, at-least-once cutover replay. Process it idempotently and acknowledge it only through the exact displayed guidance.\n- `[LEGACY READ-ONLY]` is inspection-only. It has no reply, acknowledgment, or lifecycle action.\n- An unlabeled current message uses the current guide and current grammar.\n\nAn explicitly selected current Run, attested current Run binding, current Dispatch, or federated attachment takes precedence over legacy fallback. A retained adoption record alone never turns a current command into a legacy call.\n\nDatabase provenance, an old-looking terminal, or a legacy Run ID does not prove mutation authority. If the runtime cannot prove liveness, principal ownership, capability, or the exact legacy contract, it degrades to read-only inspection and must not fall back to local execution. Exact recovery may restore the already-live PTY once in its original inactive background tab. It must not spawn, write, signal, stop, switch, focus, split, or inject a terminal. Loss of lifecycle authority does not invalidate the existing assignment, process, or filesystem work.\n\nCompatibility retries have narrow guarantees. A pending ask, a reply, a final Dispatch settlement, and a consuming check have durable recovery identities. A-era heartbeat and escalation calls remain at-least-once across a manual A-to-B retry because identical later signals may be intentional. If an A-era ask may already have been answered, run the exact non-consuming recovery check printed by the runtime first; after its answer is printed and acknowledged, a new invocation with the same question creates a new question. Never guess among multiple identical question threads.\n\nWhen a compatibility or recovery command returns structured next-step arguments, run those exact arguments with the same CLI executable. The arguments intentionally omit the executable name so the guidance works with `orca`, `orca-ide`, `orca-dev`, or another configured Orca CLI command. Do not translate the command from memory, broaden its recipient, or retry it as a current mutation unless the returned guidance explicitly says to.\n\nOn packaged Windows, a legacy ask uses a two-step commit/resume protocol. The initial command durably commits the question, prints its exact `ask --resume ` command, and exits with launcher status `75`; it does not wait for the answer. Run that exact resume command after the launcher or update boundary. Resume is idempotent and read-oriented: it waits for the already-committed question and does not create another one. For a WSL process that received compatibility proof at launch, use the printed executable `orca-ide` WSL resume command so the same distro and packaged launcher authority are preserved; do not substitute a PATH-resolved local CLI. Older WSL processes that never received the hidden launch token remain lifecycle read-only after the update, even while their terminal and filesystem work continue.\n\nLegacy inspection remains available without consuming mail:\n\n```bash\norca orchestration run-list --json\n# run_legacy_local is an empty audit tombstone after adoption.\norca orchestration run-show --id run_legacy_local --json\n# In run-list, find the ordinary Run whose objective is:\n# \"Recovered orchestration work from a contract update\"\norca orchestration run-show --id --json\norca orchestration task-list --run --json\norca orchestration inbox --full --json\norca orchestration check --terminal --peek --format --json\norca terminal read --terminal --json\norca terminal wait --terminal --for tui-idle --timeout-ms 60000 --json\n```\n\nIf the original coordinator is unavailable or cannot prove its retained authority, a current coordinator may explicitly take over the adopted Run from its own live agent terminal:\n\n```bash\norca orchestration run-use --id --takeover-legacy --json\norca orchestration check --run --json\n```\n\nTakeover fences only the old coordinator, binds the current one, and moves pending worker mail into current Run Delivery. It is bound to the authenticated invoking terminal; `--from` cannot name another coordinator. Live legacy workers keep their original Tasks, Dispatches, processes, filesystems, and old prompt commands; their later questions, escalations, and completion reports route to the current coordinator. Do not use takeover while the original coordinator is still actively coordinating, because its later lifecycle mutations are rejected.\n\nDo not launch a replacement editor merely because the desktop app or runtime was updated. If adoption cannot prove continuing authority, keep the original worker as the only editor until it reaches a stable handoff point, then use a new current Dispatch in a conflict-free placement for any remaining work.\n\n## Ownership\n\nNew orchestration messages and tasks belong to one explicitly bound Run. A Run is only a durable namespace and coordinator inbox; it never schedules or places workers. Lifecycle authority comes from the active Dispatch, and terminal handles remain routing metadata rather than durable identity. Send `worker_done` and `heartbeat` from the worker's own terminal; Orca routes them to that Dispatch's Run.\n\nClassify inherited context before sending lifecycle messages:\n\n- Coordinated subtask: a live coordinator owns the DAG and waits on this dispatch. Follow the preamble exactly, including `worker_done`, heartbeat/status, `ask`, and `escalation`.\n- Full handoff means ownership transfer, not supervised dispatch. The original actor is not monitoring a DAG, so do not create lifecycle obligations unless the user explicitly asks you to supervise.\n- Classify requests containing \"hand off\", \"handoff\", \"handover\", \"give this to another agent\", \"give this to another worktree\", \"another agent\", or \"another worktree\" as full handoffs by default, even when the user names a custom model or reasoning effort.\n- Use supervised orchestration only when the user explicitly asks you to \"supervise\", \"monitor\", \"wait\", \"track completion\", \"wait for worker_done\", return results, coordinate a DAG, use a decision gate, or manage ask/reply flow.\n- Do not use `orca orchestration dispatch --inject` for full handoffs. It injects a coordinator preamble that tells the worker to send `worker_done`, heartbeat, and `ask` messages, then end its turn under the original terminal's dispatch lifecycle.\n- Do not run `orca orchestration task-create`, `orca orchestration dispatch --inject`, or `orca orchestration check --wait` for full handoffs. Do not peek at terminal output after prompt delivery to monitor progress.\n- A review-only `worker_done` reports findings; it does not authorize coordinator file edits. After a review-only completion, synthesize findings, ask a decision gate if ownership is unclear, and dispatch or hand off fixes unless the user explicitly asked the coordinator to own fixes.\n- If the user's plan names a next owner agent (for example, \"then use opencode to create a PR\"), post-review corrections and PR prep belong to that named owner. The coordinator routes, synthesizes, asks decision gates when needed, and supervises; the named owner edits files and creates the PR.\n\nIf unclear, inspect orchestration state before sending lifecycle messages:\n\n```bash\norca orchestration task-list --json\norca terminal list --json\n# If inherited context includes a task id:\norca orchestration dispatch-show --task --json\n```\n\n## Messaging\n\n```bash\norca orchestration send --subject [--to ] [--from ] [--body ] [--type ] [--priority ] [--thread-id ] [--payload ] [--json]\norca orchestration check [--terminal ] [--ack ] [--peek|--all] [--types ] [--format] [--wait] [--timeout-ms ] [--json]\norca orchestration reply --id --body [--from ] [--json]\norca orchestration ask (--question |--resume ) [--options ] [--timeout-ms ] [--from ] [--json]\norca orchestration inbox [--limit ] [--json]\n```\n\nRules:\n\n- Omit `--from` unless impersonating another terminal; Orca auto-resolves it from the current terminal.\n- A coordinator `check` returns the bound Run's oldest FIFO Delivery (up to 50 messages) and replays that exact batch until `--ack `. Process every message before acknowledging; `check --ack --wait` acknowledges, checks, and waits in one operation.\n- Use `--peek` and `--all` only for read-only history/debugging. Type filters decide when a waiter wakes; the returned actionable Delivery is still the oldest full batch.\n- Use `dispatch:` for coordinator guidance to one supervised worker. Orca routes that stable address locally or through the connected-server relay; do not substitute a remote terminal handle.\n- Terminal handles remain appropriate for low-level pre-Dispatch messaging. Prefer `agentTerminalHandle` from the create response, fall back to `startupTerminal.handle` for older runtimes, then re-resolve with `orca terminal list --worktree ... --json` if missing or stale. Continue with the replacement handle only; never dual-send to old and new handles.\n- `terminal list --json` omits `visualLayouts` because handle recovery does not need topology. Add `--include-visual-layouts` only for explicit tab and pane inspection.\n- `orca orchestration check --peek --format --json` returns locally formatted unread mail without consuming it; it never writes to terminal input or remotely wakes another terminal. Use `orchestration dispatch --inject` to deliver a tracked task, or `terminal send` when an existing agent needs a free-form prompt.\n- While supervising workers manually, use `check --wait --types worker_done,escalation,question --timeout-ms ` instead of sleep/poll loops. Process the whole Delivery, reply to `question` messages with `orca orchestration reply --id --body --json`, then acknowledge and keep waiting.\n- `check --json` prints exactly one JSON document on stdout. While `--wait` blocks it also prints keepalive lines (`{\"_keepalive\":true,...}`) to stderr so you can tell the process is alive; those are never on stdout. Do not merge the streams before a parser — `check --wait --json 2>&1 | ` fails with \"Extra data: line 2\". Pipe stdout only.\n- Treat a `check --wait` timeout or `{count:0}` as a checkpoint, not a worker failure. Long coding tasks routinely run 15-60 minutes; keep using rolling waits unless you receive `worker_done`/`escalation`, the terminal exits or disappears, or the user explicitly asks you to stop.\n- Heartbeats and visible terminal activity mean the worker is alive, not done. Do not stop, close, kill, or restart a worker just because it has not produced a completion message yet.\n- Use `ask` when a worker needs a blocking answer from the coordinator; it defaults to the active Dispatch's Run. Timeout or disconnect leaves the question pending, so resume by its original message ID instead of asking again.\n- `check --wait` returns one bounded Delivery, not every future completion. Process every message, acknowledge it, then keep waiting until every expected Dispatch settles.\n- Group addresses include `@all`, `@idle`, `@claude`, `@codex`, `@opencode`, `@gemini`, `@droid`, `@grok`, `@cursor`, and `@worktree:`.\n- Message types include `status`, `dispatch`, `worker_done`, `merge_ready`, `escalation`, `handoff`, `question`, `decision_gate` (legacy/gates), and `heartbeat`.\n- Use group addresses only for messages that are genuinely useful to many terminals, such as `status` broadcasts or intentional fan-out questions. Do not send dispatch lifecycle messages to groups.\n- `worker_done` belongs to the active Dispatch and defaults to its Run mailbox; never target a group.\n- A valid `worker_done` for the active `taskId` + `dispatchId` marks the task and dispatch completed automatically. Do not follow it with `task-update --status completed`; reserve manual updates for explicit recovery or overrides.\n- `heartbeat` is also Dispatch-scoped. Include both IDs and omit `--to` so Orca uses the owning Run; use `status` for broad progress updates.\n\n## Tasks And Dispatch\n\nA Run is the namespace/inbox, a Task is the work item, and a Dispatch assigns one Task attempt to a terminal. Create or bind a Run once before the common loop.\n\n```bash\norca orchestration run-create --objective --json\norca orchestration task-create --spec [--deps ] [--parent ] [--json]\norca orchestration task-list [--status ] [--ready] [--brief] [--json]\norca orchestration task-update --id --status [--result ] [--json]\norca orchestration dispatch --task --to [--from ] [--inject] [--json]\norca orchestration dispatch-show --task [--json]\n```\n\nTask statuses: `pending`, `ready`, `dispatched`, `completed`, `failed`, `blocked`.\n\nDispatch rules:\n\n- `--inject` sends the task spec plus preamble into a recognized agent CLI so it can report `worker_done`.\n- If the target is a bare shell, omit `--inject`, dispatch for tracking if needed, then send the prompt manually with `orca terminal send --terminal --text --enter --json`.\n- After 3 consecutive failures on one task, the dispatch context circuit-breaks and the task is marked failed.\n- Use `task-list --brief --json` for coordinator sweeps; it collapses whitespace and caps each echoed spec at 160 characters (`spec_truncated` marks shortened rows). Omit `--brief` when the full spec is required, or when an older CLI rejects it as an unknown flag.\n\n## How deep workers can nest\n\nA dispatched worker normally cannot dispatch sub-workers. Attempting it fails with\n`nested_worker_depth_exceeded` and a message telling the worker to complete the task\nitself. Do that — do not try to route around it.\n\nThe limit is a number, not an on/off switch. `Settings -> Orchestration -> Nested worker depth`\nsets how many generations are allowed:\n\n- `1` (default): a coordinator dispatches workers; those workers do not dispatch.\n- `2`: workers may dispatch one further generation.\n\nDepth is counted from the terminal that issues the command, not from the Run. Creating a\nnew Run does not reset it — a worker that runs `run-create` then `worker-start` is still a\nworker, and still counted. This is the part that changed: the old behaviour rejected\nsub-dispatch only because a worker's terminal was not bound to a Run, so creating a Run was\nenough to slip past it.\n\nTwo limits worth knowing:\n\n- **It is a guardrail, not a security boundary.** A caller that declares another terminal's\n handle while its own launch evidence is unverifiable (an ordinary restored terminal, for\n example) can be counted as that terminal instead. Orca does not treat workers as hostile.\n- **It applies while a Dispatch is active.** After `worker_done`, or after a coordinator\n settles the task, the terminal is no longer a worker and is counted as a root again. The\n process may still be alive; that is the documented boundary, not an accident.\n\n## Preferred Supervised Worker Loop\n\nUse `worker-start` for the normal supervised path. It composes the existing worktree, terminal, readiness, and dispatch primitives while returning exact created/reused effects. Agents still choose placement and concurrency; Orca does not schedule workers or infer conflicts.\n\nCreate the Run and every independent Task first, then start all independent workers before waiting:\n\n```bash\norca orchestration run-create --objective \"\" --json\norca orchestration task-create --spec \"\" --json\norca orchestration task-create --spec \"\" --json\norca orchestration worker-start --task --worktree current --agent codex --json\norca orchestration worker-start --task --worktree current --agent claude --json\n```\n\n`current` and exact existing worktrees create a fresh agent terminal and do not rerun setup. Reuse an existing agent only with `--terminal `.\n\nFor a per-invocation Claude, Codex, or Cursor launch, pass an opaque provider model id with `--model`; add `--effort` only when that agent/model supports the level. These options apply only to fresh agent terminals, override general agent default arguments, and are reported under `launch.requested` and `launch.effective` in the receipt:\n\n```bash\norca orchestration worker-start --task --worktree current --agent claude --model opus --effort high --json\n```\n\n`--effort` requires `--model`, and neither option can combine with `--terminal`. A connected worker server must advertise launch-preference support before Orca forwards either option.\n\nFor a new worktree, setup runs by default and agent-first creation reuses the returned startup agent terminal:\n\n```bash\norca orchestration worker-start --task