Commit Graph
71 Commits
Author SHA1 Message Date
Jinwoo-H 630bc6642c feat(cli): add skills get --reference and --references selectors
An action gate names one reference, but --full was the only way to reach it and
returned the kernel plus every reference. The selector serves one document, and
the orchestration guide now teaches it with --full as the older-CLI fallback.
2026-09-05 00:21:24 -04:00
Jinwoo-H 0032074fd7 skills: keep only the orchestration guide rewrite in this PR
The seven non-orchestration guide rewrites, the shared stub fragment, and their guards move to a separate PR. The orca-cli cross-guide pin follows the worktree-selector rule into the orchestration placement reference.
2026-09-04 18:17:37 -04:00
Jinwoo-H 5bcc79923c skills: spell every invocation ORCA, restore the auth-home rule, restate dense sentences plainly 2026-09-04 17:53:28 -04:00
Jinwoo-H 3ebda3f0b7 skills: reconcile integrated guards, name the emulator_no_active recovery 2026-09-04 17:36:25 -04:00
Jinwoo-H 3753fac07e Merge branch 'skills-fix-env' into skills-optimization
# Conflicts:
#	config/scripts/generate-bundled-skill-guides.test.mjs
#	skill-stubs/orca-per-workspace-env.md
2026-09-04 17:34:30 -04:00
Jinwoo-H fe7d20f569 Merge branch 'skills-fix-linear-emu' into skills-optimization
# Conflicts:
#	config/scripts/generate-bundled-skill-guides.test.mjs
2026-09-04 17:34:01 -04:00
Jinwoo-H e3df3bbfe4 Merge branch 'skills-fix-cli' into skills-optimization
# Conflicts:
#	config/scripts/generate-bundled-skill-guides.test.mjs
2026-09-04 17:33:42 -04:00
Jinwoo-H a66bd9dc5a docs(skills): split per-workspace-env guide into a kernel plus references
Apply the Change findings from the skill review.

Correctness:
- relayGracePeriodSeconds now states that 0 is unbounded (the relay stays up
  until explicitly terminated) and names the accepted domain, 0 or 60..604800
  from EphemeralVmRecipeSshTargetSchema. Removed from the SSH exemplar so the
  omitted default and a written value cannot disagree.
- The SSH exemplar carries only the required fields. jumpHost/proxyCommand
  exclusivity is stated once, where the choice is made.
- The provisioned-root snippet fetches "$ORCA_REPO_URL", the remote Orca
  resolved the base ref against, not origin.
- portForwards entries name localPort/remoteHost/remotePort and optional label
  against the strict SavedPortForwardSchema.
- The free doctor gate is clear only with no fail and no warn; buildDoctorResult
  leaves ok true with warnings.
- The worked auth check uses the status command's exit code by default via a
  sentinel, since a provider CLI need not propagate a remote exit code. The grep
  recipe stays as the named fallback and matches a shell variable, so no
  pipefail/SIGPIPE hazard remains.

Structure:
- One outcome spine with a joined done bar, and one autonomy envelope replacing
  the four money restatements, the five checkpoint rationales, and Boundaries.
- Quick-start deleted; the numbered sequence is the single copy.
- Kernel keeps what must fire without a read; the four worked examples and the
  failure modes move to skill-guides/orca-per-workspace-env/references/.
- Description replaced with the review's proposal; the ORCA placeholder rule is
  stated once and the Claude Code bang prefix is fenced as a harness adapter.

The generator test's reference assertions are generalized per guide, and the
Vercel name-building pins are repointed to the file that now carries them.
2026-09-04 17:29:44 -04:00
Jinwoo-H 47e5e2ba68 docs(skills): correct linear and emulator skill guides
Emulator: drop the camera-injection verb (it does not exist in
src/cli/specs/emulator.ts), drop iOS permissions (the iOS backend
declares permissions: false so the bridge throws emulator_unsupported),
fix the Android permissions positional order and remove the nonexistent
list op, and delete the stale 'visual pane in development' status and
the 'once the attach/active flow lands' qualifier. State the wrapped-verb
and backend-capability conditions once, add an outcome spine, and drop
the ASCII diagrams and identity prose.

Linear: delete the 27-line Common Commands mirror of --help, replace the
verb-keyed unconfirmed-write rules with the payload-keyed condition that
covers every write verb, add a skill-level done bar for all five
branches, and move examples onto the ORCA placeholder.

All four descriptions rewritten off body-owned detail.
2026-09-04 17:25:06 -04:00
Jinwoo-H 47e3bf4ac1 docs(skills): give orca-cli and computer-use an outcome spine and conditional references
Applies the Change findings from the orca-cli / computer-use skill review.

orca-cli guide:
- Delete the duplicated mobile-emulator tail and the second `Next Action`; the
  emulator now routes through one conditional-references row.
- State the handoff done bar once, and state the `terminal wait` gate with its
  failure direction beside the recipe: `terminal wait` prints an ordinary result
  envelope on timeout and signals the unsatisfied wait only through the exit
  code, so an agent could send a brief into a half-started TUI.
- State the one-agent-handle invariant once in Worktrees; the Terminals copy and
  the two restatements are gone.
- Drop the executable-resolution ladder the discovery stub owns and keep one
  placeholder rule, in the shape `skill-guides/orchestration.md` uses.
- Move the reconstructible command catalogs (browser, automations, artifact and
  skill publishing) behind `skills get orca-cli --full`; the routing paragraph,
  the untrusted-page-content rule and the artifact publish gate stay inline.
- 424 always-loaded lines to 260.

computer-use guide: promote the verification vocabulary to a done bar, and drop
its copy of the resolver ladder.

Repointed the generator's resolver-phrase assertions for these two guides to the
stub projections that now own the ladder, and generalized the bundled-reference
assertions beyond orchestration.
2026-09-04 17:23:11 -04:00
Jinwoo-H 610baee9d4 skills(orchestration): state the outcome spine's consumer and one positive-proof condition
Applies the Change findings from the orchestration skill review.

- Outcome names the next consumer (the requesting user) and the turn-end report
  contract, and states the positive-proof condition once beside Safe failure.
- The three wait/release gates cite that condition instead of carrying divergent
  case lists.
- Drops the kernel's third copy of the worker_done command; the runtime preamble
  owns it at the point of use and worker-contract.md owns it behind the gate.
- Drops messaging-and-gates.md's duplicate coordinator delivery loop block; the
  paragraph below it already states the --terminal delta as a condition.
- worker-start gets an ordered failure hatch naming failedStage and
  residualResources.
- Conditional references describes what --full returns instead of promising
  selective loading; the worker-contract gate row states the condition the kernel
  does not decide.
- Moves the worktree-selector form to placement-and-remote.md, where an exact
  selector is consumed, and aligns worker-contract.md on the kernel's
  capability-first question-TUI wording.

Pins move with the facts: the worker_done flag spellings to worker-contract.md's
recipe case, the worktree-id form to the placement reference.
2026-09-04 17:20:34 -04:00
Jinwoo-H f1e6c0d9e3 fix(orchestration): report a paneless caller before the settled-Attempt fence
The superseded-Attempt fence ran ahead of the stable_pane_required guard, so a
terminal that had only lost its pane binding was told it had been re-attached and
to stop, and lost the run-use recovery data. The fence message and the worker
contract now also cover the no-successor case they already fired on.
2026-09-04 15:59:49 -04:00
Jinwoo-H e443ed63ae Merge branch 'fix-final-surface' into integrate-final
# Conflicts:
#	src/cli/bundled-skill-guides.ts
2026-09-04 15:44:27 -04:00
Jinwoo-H 4a09ba3ad7 docs(orchestration): say an empty check never means you were replaced
The worker contract calls an empty check a checkpoint rather than a failure, so
a fenced worker had no sentence telling it that consumer_fenced, not silence, is
how it learns the Task moved. Pinned next to the existing consumer_fenced pin.
2026-09-04 15:41:29 -04:00
Jinwoo-H 2b21ce5ce2 docs(orchestration): enumerate remote workers on the kernel stall path
A coordinator that never loads a reference read unverifiable for every
--on <environment> worker. Fits the existing wrap, so the 202-line budget
is unchanged.
2026-09-04 15:34:26 -04:00
Jinwoo-H 8e28b7de61 docs(orchestration): stop the kernel from looping on an informational nextAction
For a healthy in-progress worker the projection returns nextAction
{kind:'inspect', argv:['orchestration','worker-show','--dispatch',<same id>]}
with attention.requiresAction false, and the kernel told the coordinator to
follow the literal argv. Raises the kernel line budget 200 -> 202: the
paragraph had zero slack and no existing guidance was worth cutting.
2026-09-04 15:30:09 -04:00
Jinwoo-H a4f416f5b4 docs(orchestration): state that --types is a wake condition, not a batch filter
check --types without --wait returned an unmatched type. wakeTypes is an
existence probe in getOrCreateMailboxDelivery; the batch query that
follows has no type predicate, so a Delivery is never filtered. Behavior
is correct; only the guide and --help were silent.
2026-09-04 15:25:13 -04:00
Jinwoo-H cbe2d378fb docs(orchestration): name --include-remote on the guides' enumeration paths
The stall path told coordinators to enumerate with plain worker-list, but
a worker started --on <environment> reads unverifiable without
--include-remote. Also names page.nextCursor for fleets past 100 rows.
2026-09-04 15:24:31 -04:00
Jinwoo-H d571ab154f docs(orchestration): tell a fenced worker to stop instead of retrying check 2026-09-04 14:09:48 -04:00
Jinwoo-H 40aa364df7 docs(orchestration): correct the projection.attention path and restore routing phrases 2026-09-04 04:04:52 -04:00
Jinwoo-H 3252d28d71 docs(orchestration): keep the id:<newFullWorktreeId> placeholder the cross-skill guidance test pins 2026-09-04 03:45:59 -04:00
Jinwoo-H 5ab25102e5 docs(orchestration): restore the routing triggers to the skill description
The branch's rewrite dropped main's verbatim phrases ("hand off", "handoff",
"handover", "give this to another agent", "another worktree", threaded
messages, worker_done/escalation waits, decision gates, reading or waiting on
terminals) — the only text a model sees when choosing this skill. Restored in
both the kernel frontmatter and the identical stub, still shorter than main's,
and pinned by a routing test.
2026-09-04 03:23:41 -04:00
Jinwoo-H e484d6a1a7 docs(cli): name the prompt stages the runtime actually reports
The guide still documented queued_pending_turn and submission_observed, and the
send-submit repro gated on the removed submission_observed literal, so its
unsubmitted assertion passed vacuously. Both now use input_accepted then
turn_started.
2026-09-04 03:22:02 -04:00
Jinwoo-H 31c3f74f79 fix(cli): publish terminal send delivery warnings in --json
The four prompt-delivery warnings were built inside the text formatter, so
--json callers (every agent) saw none of them. The receipt now carries a
warnings array built from the same function, and the swallowed-Enter warning is
ordered ahead of the unsupported-observation arm so an agent provider always
gets the recovery command; a plain shell keeps its cannot-report-delivery text.
2026-09-04 03:16:54 -04:00
Jinwoo-H c78f40fdd0 docs(orchestration): require positive evidence of exit before ending a wait
The kernel's stall exit fired on "not live", which includes every unverifiable
arm — all of which are absence — contradicting the safety floor and the recovery
table. It now names the positive signals (exited liveness, the worker's own
observation of exit, a final agent turn with no worker_done) and states that
unverifiable never authorizes stop, abandon, retry, or release. Also corrects the
projection.* field paths the worker-list row actually nests.
2026-09-04 03:14:51 -04:00
Jinwoo-H dc6e0f6e47 docs(orchestration): give the kernel loop an exit condition and the missing commands
D1: name the two liveness layers (worker-list projection.liveness is the fleet
verdict, worker-show observation.status is PTY-only) and give the supervised
loop a bounded stall procedure instead of an unbounded wait.
D2: put worker-list, attention, requiresAction and nextAction in the loop and in
completion accounting.
D4: document request-show / --retry-request / terminal send --wait-submit.
D5: check names its caller with --terminal, never --from.
D7: give a dispatched worker a concrete follow-up read cadence.
D10: document the real folder-workspace route (project setup-existing-folder).
worker-start --spec is now the canonical loop's default.
Guidance pins are contracts via squash() instead of reflow-fragile prose.
2026-09-04 01:56:59 -04:00
Jinwoo-H 07fa28a3ae Merge origin/main into orchestration-v3
Runtime and rpc hunks from the pre-split monolith still need rehoming into
main's split modules; terminal.ts max-lines follow-up pending.
2026-09-04 00:49:43 -04:00
Jinwoo Hong 573537ecd4 feat(cli): make terminal close the canonical workspace teardown (#18073)
* fix(runtime): recover stale session owners and await retirement

* fix(runtime): preserve session hydration and smoke compatibility

* test(runtime): cover empty and unindexed session owners

* feat(cli): make terminal close the canonical workspace teardown

* fix(preload): align ssh termination result type

* test(runtime): assert folder hydration owner

* fix(runtime): fence legacy terminal stop by worktree host

* fix(preload): reconcile ssh result import with main

* fix(runtime): keep same-id sibling hosts out of workspace close

The stale-owner fallback in the session controller re-routed any worktree whose
catalog partition had no tabs to whichever other partition held tabs. Only
`runtime:` environment ids rotate across relay restarts; `repoId::path` legitimately
repeats across hosts, so an SSH workspace close could retire the local copy's
tabs and resume records, or flip owners mid-close and strand the SSH PTY.

Restrict the fallback to runtime hosts, and pin the session partition once per
workspace close so record clearing targets the partition that owned the tabs.

* test(runtime): give the cross-host close fixture a real resume record

* fix(preload): take main's ssh-bridge import order so the merge stays duplicate-free
2026-09-03 03:58:45 -04:00
Jinwoo Hong b44ef1e59d fix(skills): narrow computer-use discovery boundary (#17736)
* fix(skills): narrow computer-use discovery boundary

* chore: remove merge-formatting noise

* fix(skills): name browser page automation surfaces
2026-08-31 18:57:52 -04:00
Jinwoo-H df91d88bd1 fix(orchestration): close CI durability regressions 2026-08-31 15:50:58 -04:00
Jinwoo-H e92d7812d9 feat(orchestration): make multi-agent workflows durable 2026-08-31 15:48:28 -04:00
NeilandBrennan Benson fbe94ceff6 fix: close readiness gaps found by merged-change audit (#17159)
* fix(ssh): fence stale kills and retired pane replay

* fix(ssh): support cancellable interactive authentication

* fix(ssh): await remote catalog before snapshot adoption

* fix(pty): contain Windows ConPTY input failures

* fix(power): avoid redundant macOS display blocking

* perf(editor): narrow markdown override subscriptions

* fix(quick-open): close directory handles after reads

* refactor(linux): remove unused proc socket scanner

* fix(usage): apply flat Sonnet 4.6 pricing

* ci: prime Node next native test cache

* docs(skills): resolve snapshot cleanup data path

* fix(ssh): recover install locks after host reboot

* test(ssh): recognize boot-aware install locks

* test(ssh): prove previous-boot lock recovery live

* test(wire): pin pre-metadata release coverage

* fix(terminal): preserve remote tab ownership through recovery races

* test(runtime): fence replaced terminal handles in agent guard

* fix(ssh): preserve remote snapshot authority across polls

* fix(pty): contain late ConPTY output EPIPE

* test(pty): register Windows exit watcher before kill

* fix: close SSH and tab readiness race gaps

* fix(tabs): retain headless order and placeholder titles

* fix(build): avoid parallel electron-vite config race

* test(windows): avoid MSYS temp path rewriting

* test(windows): avoid killing exited PTY

* fix(pty): avoid late ConPTY input teardown race

* fix(terminal): sync reconnect error ownership after commit

* fix(runtime): use canonical worktree identity comparison

* test(ssh): assert complete cold-hydration baseline

* test(windows): invoke quoted retention fixture via PowerShell

* test(windows): read ConPTY grid through mode con

* fix(terminal): publish PTY replacements atomically

* fix(terminal): infer stale identity on reattach

* fix(terminal): fence stale pane PTY callbacks

* fix(terminal): fence stale pane binds after rebind

* fix(terminal): reject stale pane transport callbacks

* fix(terminal): fence mirrored reattach spawn callbacks

* fix(terminal): replace stale pane PTYs on remount

* fix(ci): size the Windows launcher-compile test budget from measurement

`native-smoke (windows-latest)` fails ~4.5% of runs on
`preserves a multiline argument through the compiled remote launcher`
with "Test timed out in 15000ms" — on unrelated PRs, for reasons that
have nothing to do with them. Across 176 sampled attempts it is the only
red that job produced, and it hit seven different PRs in two days:
#16900, #16904, #16915, #16955 (twice), #16979, #17014, #17085.

The test is six process creations: powershell.exe forks csc.exe, then
the freshly compiled orca.exe forks node.exe, twice. Hosted Windows
runners periodically slow process creation down, and this test amplifies
that far harder than anything else in the job. Comparing the 80 attempts
where it ran under 3s against the 12 where it ran over 12s, its own
median goes 2198ms -> 15917ms (7.2x) while the same file's
powershell-only test moves 556 -> 686ms (1.2x), the cmd.exe and Git Bash
process tests in the neighbouring file move 1.4x, and the other 35 files
put together move 1.5x.

Measured across those 176 attempts: 1881ms to 35438ms, p50 4264ms,
correlation +0.881 with the job's total Vitest duration. 8 of 176 (4.5%)
exceeded the 15s cap; 2 of 176 (1.1%) also exceeded the shared 30s
testTimeout, so deleting the override and inheriting the config is not
enough on its own. 60s clears all 176 with 1.7x headroom on the worst.

This is slow, not hung. Every body here is synchronous spawnSync, so
Vitest cannot interrupt one — the timer fires only after the body
returns and the reported duration is real elapsed time. That is why a
failure reads `× ... 22464ms` under `Test timed out in 15000ms`. The
work finished; the stopwatch was short. Seven reruns at one identical
head measured 2053 / 4680 / 5551 / 8732 / 13506 / 14868 / 21937ms — the
last of those would have been red on code that had not changed.

The 15s came from #8897, which raised this test off Vitest's built-in 5s
default because the job then ran bare `pnpm vitest run`. #8909 landed
3h27m later and pointed the job at config/vitest.config.ts, which is the
real fix for that. The constant stayed behind and has been the binding
budget ever since.

* fix(terminal): fence stale remount reattach ownership

* fix(terminal): reconcile mounted pane identity after replacement

* fix(terminal): fence stale reattach fallback ownership

* fix(terminal): fence deferred SSH reattach ownership

* fix(terminal): fence stale split pane ownership callbacks

* fix(terminal): keep stale spawns from consuming startup

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-31 08:17:40 -07:00
Brennan Benson cc6b600e21 Fix orchestration CLI recovery, settled-Dispatch mail, and guide defects (#16919)
* Fix orchestration CLI recovery, settled-Dispatch mail, and guide defects

Five reported orchestration CLI defects, verified individually before fixing.
Two were real code defects, one was a docs error, one was correct as-is, and
one was correct on both ends except for its recovery wording.

- Mail addressed to a settled `dispatch:<id>` was accepted and silently dropped.
  Local sends bypassed the settlement check the federated branch already had, so
  the caller was told success for a delivery no worker would ever read. Reject
  with `dispatch_inactive` and name the Run mailbox to use instead.

- A lost mutation response offered no read-only way to ask whether it took
  effect. `--retry-request` does dedupe correctly, but the recovery guidance
  emitted a query command only when the payload carried a dispatch id, which is
  exactly what a lost response lacks. Add read-only
  `orca orchestration request-show --request <id>` over the durable receipt
  ledger, and always emit a read-only step before the keyed retry.

- The bundled `orca-cli` guide documented `check --unread --inject`, a flag the
  parser rejects. Correct it to `--format` and add a ratchet that runs every
  orchestration invocation in the bundled guides through the real CLI parser.

- `check --json` is one stdout document and its keepalives are stderr-only; the
  reported `Extra data: line 2` came from merging the streams. Document the
  contract rather than changing the wire.

- A rejected lifecycle message is loud on both ends already, but the rejection
  never named the flag that supplies the missing capability. Name it.

* Harden orchestration mutation recovery guidance
2026-08-30 18:12:58 -07:00
Neil 2b0ee06205 docs(env-recipes): warn that snapshotting a started runtime bakes its identity (#17001)
Snapshotting a VM on which `orca serve` has already run captures the
runtime's user-data dir into the image. Every VM booted from that image
then shares one pairing identity and one agent-session-authority key,
which defeats the per-device token design.

Confirmed by booting two VMs from one such snapshot: both emitted
identical deviceToken and pairedDeviceId.

Adds the rule to the base-snapshot section and repeats it for the
agent-auth layer, which is the likelier place to start the runtime by
hand while smoke-testing. Says to delete the whole user-data dir rather
than a named file list, since that list drifts as Orca adds state.
2026-08-28 02:37:55 -07:00
Jinjing c4b39295c1 style: format codebase (#16935)
* style: format codebase

* style: format codebase

* refactor: extract skill install dialog footer and content

Extract footer and content sections from SkillInstallDialog and
SkillInstallManagementDialog into separate components for improved
maintainability and clarity of component responsibilities.
2026-08-28 00:59:21 -07:00
Brennan Benson 9135b6f004 feat(orchestration): surface nested worker depth and propagate it across hosts (#16669)
* feat(orchestration): surface nested worker depth and propagate it across hosts

Builds on the depth enforcement in the previous commit, which shipped with the
setting reachable only by editing settings.json and with workers never told they
could nest.

Adds the Settings -> Agents control (a 1/2/3 select rather than a free-form
number, which bounds the value without inventing a numeric input primitive). The
key stays absent from the SettingsUpdate RPC schema, matching agentSkillSharingEnabled:
settings.update is reachable from the CLI, so an RPC-writable depth would let a
worker raise its own cap.

Adds a SUB-DISPATCH block to the dispatch preamble, emitted only when the worker
actually has budget left. A worker told it "usually cannot" delegate still tries and
then reports the refusal as a blocker, so the section is omitted entirely rather
than softened.

Propagates depth to federated worker hosts. Previously the home side computed and
stored a depth the remote host never received, so a remote attachment always read
as depth 1. That is correct at the default cap and wrong as soon as the cap is
raised — precisely when someone starts relying on nesting. The field is optional,
so an older Run home simply omits it and the attachment's NOT NULL DEFAULT 1 keeps
the fail-closed behaviour. Enforcement still runs on the executing host against
that host's own cap, consistent with the SSH execution boundary.

* fix(orchestration): close nested depth readiness gaps

* fix(settings): defer nested depth translations

* fix(orchestration): drop federated depth keys that main already landed

The enforcement PR's review pass added the same federated depth propagation
before it merged, so replaying this branch onto main produced duplicate object
keys. Keep main's versions -- its schema entry validates an integer >= 1 rather
than any finite number.

* fix(settings): label nested worker depth select

* fix(settings): move nested depth to orchestration

* fix(settings): refine nested depth placement
2026-08-26 16:16:05 -07:00
Brennan Benson 8a07bbd8cf fix(orchestration): enforce nested worker depth instead of an accidental fence (#16668)
* fix(orchestration): enforce nested worker depth instead of an accidental fence

Orca documented that "dispatched workers cannot spawn their own sub-workers
(worker-start is coordinator-fenced)". No such check existed. What existed was a
single Run-binding check in the workerStart RPC: a worker's terminal is not bound
to a Run, so worker-start happened to fail. The rule was emergent, asserted by no
test, and written in no doc — and it leaked. A worker could run-create its own
Run, task-create, and worker-start: now bound, the check passed.

Replace it with a real, configurable depth cap.

Depth is derived from the caller's own active Dispatch rather than from Run
binding, which is what dissolves the run-create bypass: creating a Run does not
stop you being a worker. Enforcement lives in a single dispatch-row writer that
owns all three INSERTs that mint a live worker — the generic claim, the supervised
worker-start path (including every retry), and the remote attachment. Two of those
were missed by earlier drafts of this change, so `creator` and `maxDepth` are
required parameters: a new spawn path cannot compile without deciding, and a
boundary test refuses the SQL anywhere else.

Schema v30 adds depth to dispatch_contexts and remote_dispatch_attachments,
NOT NULL DEFAULT 1 and backfilled to 1 so an unstamped or pre-upgrade row fails
closed rather than reading as a root coordinator. The attachment pane indexes
widen to the five states in which a remote worker may still be running:
loss of contact is not evidence of process death, so an unverifiable worker still
counts as a nesting parent.

Also adds the caller-evidence assertion that workerStart was the only Run-scoped
verb to skip, so a declared --from cannot name another terminal's pane and inherit
its depth.

Default is 1, so behaviour is unchanged unless the new setting is raised. Two
limitations are deliberate and documented rather than papered over: this is a
guardrail and not a security boundary, since a caller whose launch evidence is
unverifiable (any ordinary restored terminal) can declare another handle; and it
is enforced at supervised dispatch creation, so a settled worker whose process is
still alive counts as a root again.

* fix(orchestration): share caller resolution and pin worker gaps

* refactor(orchestration): make the caller resolver's pane contract explicit

Overloads so requireStablePane callers get a non-null string instead of casting,
and rename the attestation opt-out to say what it means: the caller asserts it
itself. A flag called assertEvidence:false reads as "attestation optional",
which is the hole this helper exists to close.

* fix(orchestration): propagate dispatch depth to federated workers

* chore(cli): refresh bundled orchestration guide
2026-08-26 13:22:09 -07:00
Jinwoo HongandJinwoo-H a9781a4118 STA-4150: client-hosted remote browser (consolidated) (#15448)
Co-authored-by: Jinwoo-H <jinwoo@stably.ai>
2026-08-25 15:36:51 -07:00
Brennan Benson 3fca1d1648 fix(linear): unbound list-issues by default, surface truncation, bind cursor workspace (#15824)
Fixes STA-5076.

list-issues capped at 50 by default and hard-clamped at 250, with hasMore buried
under result.meta and no stderr warning for --json, so a page that stopped early
read as a complete answer. Omitting --limit now walks Linear's pages until they
run out (meta.limit is null), and --limit <n> is the only cap, paging past
Linear's 250-per-request maximum to reach it. result.truncated sits next to
result.issues and is set only when a cap actually held results back; human output
prints "truncated: showing N".

The read still has to fit the CLI's 60s RPC budget, so a 20s wall-clock deadline
and a 200-page ceiling stop the walk early and report truncated with a
continuation cursor rather than failing the command.

Also:
- issued --cursor values bind the resolved workspace, so call -> nextCursor ->
  call works without --workspace; raw Linear cursors still need one and now carry
  nextSteps
- issued cursors whose payload smuggles back `all` or an empty workspace are
  rejected at decode, since either would widen the read past the bound workspace
- JSON issue rows carry priorityLabel (none/urgent/high/medium/low), matching
  orca linear priority set
- truncated and priorityLabel are optional on the wire, so a host that predates
  either is not read as "complete"; readers fall back to meta.hasMore
- the truncation line prints the rows actually rendered, so a remote result with
  no meta.returned cannot print "showing undefined"
2026-08-21 14:28:55 -07:00
Brennan Benson 2a760e310b fix(computer): report unasserted accessibility actions (#15028)
* fix(computer): report unasserted accessibility actions

* fix(computer): fail closed on missing action metadata

* Fix merged tab search test fixture
2026-08-18 11:29:26 -07:00
Brennan Benson fc8b92e507 docs(computer): explain screenshot file requirements (#15054)
* docs(computer): clarify screenshot output requirements

* fix(cli): do not advertise an unshipped --probe flag

The capabilities help line referenced --probe, which does not exist yet;
it ships in a later change. Advertising it here would be false until then.

* fix(cli): align computer-use screenshot guidance

* docs(computer): document inline screenshot fallback

* docs(computer): keep screenshot summary accurate

* docs(computer): keep screenshot guidance general
2026-08-18 01:18:56 -07:00
Jinwoo Hong 0bedeea642 fix(orchestration): expose unsupervised dispatch lanes (#15105) 2026-08-17 13:53:26 -07:00
Jinwoo Hong fa9b20cb41 feat(skills): reland private bundle sharing safely (#14934) 2026-08-16 13:45:54 -07:00
Jinjing 763b1febeb Revert "feat(skills): add private bundle sharing (#14401)" (#14913)
This reverts commit 757fae28d7.
2026-08-16 10:39:57 -07:00
Jinwoo HongandE2E Test 757fae28d7 feat(skills): add private bundle sharing (#14401)
Co-authored-by: E2E Test <e2e@test.local>
2026-08-16 02:36:18 -07:00
Brennan Benson 66dfdc456f feat(computer-use): support macOS middle click and stop the silent left-click fallback (#14721)
* feat(computer-use): support macOS middle click and gate the AX click path

`--mouse-button middle` already validated end-to-end through the CLI, the
zod schema, and the provider validator, and both the Windows and Linux
providers honored it. Only the macOS provider rejected it outright with
"middle-click is not yet supported", so the flag was a dead end on the one
platform that has no fallback.

Two changes:

- Add `.middle` to the macOS button mapping. macOS has no dedicated middle
  event family, so it rides `otherMouseDown`/`otherMouseUp` with the button
  number carried by `mouseButton: .center`; that constructor argument is
  honored for exactly the `otherMouse*` types, so no extra field write is
  needed.
- Validate the requested button before the accessibility fast path, and skip
  that path for buttons it cannot express. Previously the raw string was read
  unvalidated, and `performClickAction` only special-cased `right`, so
  `click --mouse-button middle --element-index N` (no modifiers, count 1) fell
  through to `AXPress` — a left click — and reported success with
  `path: "accessibility"`. Any unrecognized button string did the same. This
  matches guards the Windows and Linux providers already had.

The button enum moves into `OrcaComputerUseMacOSCore` so it is unit-testable;
`main.swift` keeps only the CoreGraphics mapping.

Also documents `--mouse-button` in the computer-use skill guide, which never
mentioned the flag, so agents on Windows and Linux had no way to discover it.

* test(computer-use): cover macOS middle click in the real-desktop e2e suite

* test(computer-use): prove macOS middle-click delivery
2026-08-15 00:41:45 -07:00
Jinwoo Hong 500b72d8ef fix(vm): harden provisioned root ownership and cleanup (#14477)
* fix(vm): verify provisioned root ownership

* test(vm): retry transient removal menu

* test(vm): stabilize provisioned root teardown

* fix(vm): clarify recipe-owned cleanup

* fix(vm): pin provisioned root source commit

* fix(vm): make runtime cleanup user-cancellable
2026-08-14 19:04:55 -04:00
Jinwoo Hong 77b37d85e2 feat(vm): create workspaces from provisioned SSH roots (#14359)
* feat(vm): use recipe-provisioned SSH roots

* fix(vm): preserve ordinary create failure timing

* test(vm): prepare provisioned root SSH fixture

* ci(vm): enable SSH setup for provisioned root E2E
2026-08-13 16:23:22 -07:00
erishandJinwoo-H 1f4b731f7c fix(skills): use exported recipe id in environment guide (#14280)
* fix(skills): use exported recipe id in environment guide

* fix(skills): keep recipe-derived Vercel names valid

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
2026-08-13 14:18:57 -07:00
Brennan Benson cbca291aa7 fix(orchestration): preserve direct user authority after worker_done (#14192)
* fix(orchestration): preserve direct user authority

* test(orchestration): assert settled dispatch boundaries
2026-08-13 12:00:04 -07:00