Commit Graph
847 Commits
Author SHA1 Message Date
Neil f69052e113 Reuse qualified Windows server builds and dependency verification records (#24448) 2026-10-01 16:19:40 -07:00
OrcaWinandm4air 43d9b43d3f feat(ssh): remote orcad stop by request file and journaled decommission (#16741 T6-4) (#24449)
Clients stop an orcad that advertises health.stopRequests through its slot-local request file and keep SIGTERM for older builds. Decommission runs through the activation journal and fence: it refuses while the terminal census is live or uncounted, stops the instance with an instance-bound managed request, cancels a stop orcad never acted on, and deactivates the record only on proven exit. orcad gains --cancel-managed-stop and an exclusive per-transaction decision file so a cancel can never race a dispatched stop. POSIX-only and inert: no production caller.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 13:32:25 -07:00
OrcaWinandm4air b093d3ab20 feat(orcad): supervisable server: stop requests, managed stop receipts and a lifetime that keeps its lock on failed teardown (#16741 T6-3) (#24433)
orcad stops through slot-local and instance-bound request files, so a reused PID is never signalled. A managed stop is proven by its completion command and recorded as a receipt. Optional daemon retirement is best effort: an idle daemon retires, while a busy or unverifiable one stays up with its admission fence released. Runtime teardown runs in reverse order and keeps the instance lock and profile admission when any writer fails to stop. Browser discovery no longer delays readiness. Legacy worker recovery and watcher children are drained before the final flush. Headless terminal close no longer waits on a renderer tab that does not exist. No production deployment.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 12:45:34 -07:00
Neil 197ea3a3b3 Free PR CI capacity by avoiding repeated setup and real-time test waits (#24355)
* Reduce repeated PR setup and transcript timing waits; add hosted comparisons

* Align parallelism contract with Node-only external rebuild toolchain

* Record hosted coverage and launch package, store, and cancellation comparisons

* Apply hosted Windows setup savings and remove measured test waits

* Keep measured PR package gains and remove completed comparison jobs

* Report measured test counts with precise units
2026-10-01 11:51:43 -07:00
github-actions[bot] 198fe72066 Update README downloads badge 2026-10-01 18:35:36 +00:00
Brennan Benson c6cfcc034e refactor(native-chat): structured chat failures always reach the diagnostics log (#24312)
* refactor(native-chat): give the structured chat host one required logger

The structured chat runtime took an optional onError callback that the
desktop never passed, so a late dispatch settlement, an unanswered-dispatch
release, a journal event-sink write and a provider lifecycle delivery that
failed were dropped with no trace. Other host failures went to scattered
console.warn calls, which reach nothing in a packaged desktop build.

The runtime and host now take one required logger (warn/error with a scope
and fields). The production logger writes each entry as a failed span to
<userData>/logs/main.trace.ndjson, which the diagnostic bundle collects, and
to the console (stderr under a supervised headless host). The runtime and the
host wrap it so a logger that throws never fails what it reports, and the
install refuses without one. Sites that deliberately kept a recovery-capsule
error out of the log still log no error object.

* refactor(native-chat): hand the chat host's collaborators the logger, and give orcad its trace file

The delivery loop, idle sweep, queued-message drain, lease renewer, event
sink, conversation map and provider start/exit settlement each took an
internal error callback that the host mapped onto the logger. They now take
the logger itself and log under their own scope. The event sink keeps one
onFailed hook, which decides whether to stop the provider, not whether to
report. The dead-generation settlement returns its failure so each caller
logs it under its own scope.

orcad now installs the desktop's local trace sink under its own data root, so
a headless host's chat failures reach <data-root>/logs/main.trace.ndjson as
well as stderr.

Also passes the logger in the test fixtures the first commit missed, which
tc:node caught.

* fix(native-chat): keep repeated chat failures from flooding the trace file, and record their causes

- The production structured-chat logger writes a repeated failure (same level, scope, session,
  message and error text) once per 5 minutes, carrying how many repeats it swallowed; the
  tracked set is capped at 256.
- Trace entries now carry the error's code (and SQLite errcode) and up to three causes by name and
  message.
- A chat read whose conversation will not open is logged through the host's logger
  (open-for-read), and so are the runtime's chat-tab bookkeeping failures that already hold the
  host.
- orcad writes its own orcad.trace.ndjson, closes it after every quit handler, and flushes it on
  process exit; a trace file that cannot be opened leaves tracing off instead of stopping the app
  or orcad.
- Tests: the desktop wiring test proves the logger reaches the trace sink, and the privacy tests
  read every level the logger received.

* fix(native-chat): log a created chat's tab-publication and launch-prompt failures through the host's logger

* fix(native-chat): key a repeated chat failure on everything its entry writes

The repeat suppression keyed on the message and the error's text, so two refusals with the same
code but different causes, a plain error and a refusal of one code, or two object-valued errors
shared a key and the second was swallowed for five minutes. The key is now the entry's whole
written content (fields, code, errcode, refusal reason, cause chain, a stable rendering of a
non-error value) plus the error's name and message; a refusal's reason is also written.

* test(native-chat): pin that an error's name keeps two repeated failures apart

* test(native-chat): build the refusal in the repeat-key test as the wire does

* fix(native-chat): read an error's code and a refusal's reason by narrowing, not Reflect.get
2026-10-01 11:21:50 -07:00
OrcaWinandm4air 6d1a97ef98 fix(ssh): launch the Windows relay outside sshd's job so standard users work (#24224)
* fix(ssh): launch the Windows relay outside sshd's job without WMI

Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js
gains a one-shot launcher mode that starts the detached relay with
CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard
user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a
relay without the addon, and a refusal there is named. The Windows SSH-host
lanes drop their WMI grant and assert the breakaway route and adoption.

* fix(ssh): find runtime holds without WMI on a standard-user Windows host

The store GC read held runtimes through Get-CimInstance Win32_Process, which
WMI refuses to a standard user's SSH logon, so the pass kept every runtime.
On a refusal it now reads this account's own process image paths through
Get-Process.

* build(relay): ship the Windows relay launcher addon in every desktop package

macOS and Linux packages carried Windows relays without windows-process-tree.node,
so a legacy-runtime relay they uploaded to a Windows SSH host could not launch
outside sshd's job and fell back to WMI, which a standard user is refused.

A reusable Windows job now compiles the x64 and arm64 addons once and uploads
them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds
download them before build:release and require both arches. Staging now rejects
a binary with the wrong PE machine, the ReadProcessMemory import, or no
spawnOutsideJob export, so a stale pre-launcher build cannot ship.

* ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change

The staging and gyp-rebuild scripts decide which windows-process-tree addon the
relay ships, so a change to either must re-prove the Windows host cells.

* test(ci): find the mac orcad-template download by artifact name

The release mac job now also downloads the relay Windows process-tree addons, so
the first download-artifact step is no longer the template's.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 05:32:09 -07:00
14d4bb2e2a fix(ssh): Windows hosts without Add-Type staging; runtime-store GC on Windows (#24149)
* fix(ssh): collect the pinned-Node runtime store on Windows hosts

Windows SSH hosts now run runtime-store GC instead of skipping it: one
PowerShell inventory reads .runtime-ref-node-<sha> and .runtime-node refs from
every version dir, and one Get-CimInstance Win32_Process query filtered on an
image path under runtimes\ adds process holds (never by image name; a failed
query keeps everything). Stale upload stages are swept with the same rule as
POSIX. Promotion and the post-upload hold check now take the store lock on
Windows too, and the lock's own commands run unwrapped there.

Windows relay version-dir liveness now honours .relay-pid (design D5): a live
PID answers ALIVE before any pipe is touched, a dead one (ESRCH) plus refusing
pipes is exited, anything else is unverifiable. The runtime probe adopts a
pinned node.exe an earlier vault reader left without a .verified marker after
running it.

* fix(ssh): Windows stage fencing and vault runtime go through the verified node.exe

Upload-stage file identity on Windows no longer compiles an Add-Type P/Invoke
helper when the relay runs on Orca's verified pinned node.exe: the stage
commands run a fixed fs.lstatSync(..., {bigint:true}) script through it. It
prints the legacy helper's vol:high:low lowercase hex, and identity files are
compared after normalising hex spelling, so old and new clients recover each
other's stages. Host-Node relays keep the legacy helper; the choice is
documented in windows-edr-posture.md.

The Windows OpenCode vault reader now installs the pinned runtime through
ensureRemoteOrcadNodeRuntime (official zip, host-side extraction, .verified,
store lock) instead of uploading a client-extracted node.exe, and the relay dir
gains a .runtime-ref-node-<sha> so store GC keeps the runtime the vault uses.

* test(ssh): run the Windows stage-identity and store-GC tests on the Windows lane

The legacy/node.exe identity compatibility test and the Win32_Process hold path
were gated to win32 but no CI lane ran them. Add both files to the Windows
package lane and a real running-node.exe hold test.

* test(ssh): tear down Windows-lane temp trees through removeTreeSync

* test(ssh): grant the store lock to the Windows OpenCode runtime setup test

The Windows promote now runs under runtimes/.store-lock, so the mocked host
must answer the lock's CreateNew step.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-01 03:25:42 -07:00
Neil bd90da7a5b ci: share PR planning setup and reuse the static native cache (#24329) 2026-10-01 02:46:18 -07:00
OrcaWinandm4air 53fd2dea0b feat(ssh): relay runtime fallback ladder, telemetry and host runtime setting (#24133)
* feat(ssh): complete the relay runtime fallback ladder (D6 rungs B slot, C, D)

Rung C runs the relay on the host's Node >= 18 with Orca's prebuilt N-API
addons and no npm (addon-only probe mode). Rung B is a data-driven slot chosen
only when a compat runtime is listed. Rung D fails the connect with a
classified reason carried as a TerminalUnavailableCause. The ladder steps
down only on classified refusals; unanswered probes throw. The rung decision
is persisted per host keyed by (glibc, runtime hash, Orca major), and
ssh_remote_runtime_resolved reports it once per host per session.

* feat(settings): SSH host runtime choice (Auto | Orca-managed Node | Host Node)

* docs(telemetry): describe ssh_remote_runtime_resolved

* fix(ssh): let a passing rung C disprove a remembered noexec; allow glibc-less compat runtimes

A remembered rung A noexec was re-persisted even after rung C self-tested addons from the same
~/.orca-remote tree, so rung A stayed skipped until the key changed. Rung B's evaluator also could
never match a musl compat runtime.

* test(ssh): import node:fs once in the host-node addon test

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 02:08:26 -07:00
OrcaWinandm4air ddd4927a0b build(orcad): server node-pty slots at glibc 2.28, plus a glibc 2.17 compat slot (#24134)
* build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot

Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in
manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25
through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps
its Ubuntu 20.04 (2.31) default.

Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS
pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside
SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014
with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in
DT_NEEDED, and smokes it under the glibc-217 Node.

* refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 01:10:57 -07:00
OrcaWinandm4air 6593d7d194 feat(orcad): run orcad on the pinned Node instead of Bun (#24110)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release

Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

* fix(runtime): reject a pinned archive that belongs to another target

* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol

D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.

* feat(orcad): select pinned-Node slots by a .runtime-node marker

D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.

* fix(runtime): load the Node pin without the typeless-module warning

check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.

* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout

* feat(orcad): 8-slot node-pty prebuilds against the pinned Node headers at N-API 8

- build-orcad-prebuilds.mjs adds win32-x64/arm64 (conpty.node, the vendored
  conpty.dll/OpenConsole.exe, upstream's N-API conpty_console_list.node), compiles
  in a scratch copy against the hash-verified pinned headers (node.lib pinned per
  Windows arch) with NAPI_VERSION=8, rejects post-8 node_api_* imports, and writes
  a schema 2 manifest with per-file sha256, N-API level and the glibc need.
- --require-slots [slots] verifies files against hashes; --smoke loads the slot
  under the pinned Node and spawns a PTY; --print-slot names the host slot.
- The slot installer gates on N-API, libc, arch, glibc and file hashes instead of
  the exact NODE_MODULE_VERSION, and installs nested files (conpty/).
- bun-profile-tests.yml builds, verifies and smokes each runner's slot.

* fix(orcad): scope node-pty's glibc .symver pins to glibc on musl prebuild slots

musl's unversioned libc cannot satisfy openpty@GLIBC_* references at link
time, so the Alpine slot compile would fail. Pin the staged pty.cc guard to
__GLIBC__ and assert both musl transforms against the installed patch.

* feat(orcad): run orcad on the pinned Node instead of Bun

A packaged orcad slot now references the pinned Node 24.21.0 by its
executableSha256 (`.runtime-node`, `.server-target`) instead of carrying
bun-runtime, and ships node-pty from the slot's prebuild, only its own
ripgrep, and no Windows Bun PTY gate. The runtime lives beside the slots
at runtimes/node-<sha>/bin/node (node.exe on Windows, upstream name).

- build:orcad (build-orcad-node.mjs) builds the host slot's prebuild when
  missing and places the pinned runtime; the template is schema 3 with
  per-target files.
- handoffToBundledOrcad() resolves the slot's runtime reference and checks
  process.versions.node against the pin; a host Node >= 18 still hands off.
  Startup preflight keys on running as that runtime; callers expect 'node'.
- orcad and its daemon use node-pty (ConPTY + windows-pty-job on Windows);
  the Bun PTY sources, gate entry and canUseBunPty branches are removed.
- SSH deploy uploads the official archive once per pin, extracts and
  hash-checks it on the host, and self-tests it before publishing. Bun
  slots stay launchable for rollback; Node slots never use host Node.
- The runtime materializer is generic over pinned assets; the Bun wrapper
  remains only for the OpenCode vault reader (design Phase 2).
- Cross-runtime test: a profile DB written by Bun 1.4.2 (WAL left by
  SIGKILL) opens and backs up under the pinned Node, and the reverse.

No daemon PROTOCOL_VERSION change (design D7.1 R3).

* docs(ci): name the headless lanes after the pinned Node, drop Bun shard timings

Design D10: ci-demand-rollout.md and ci-runner-efficiency.md follow the
bun-profile-tests.yml -> node-server-tests.yml rename; shard timings drop the
deleted Bun PTY tests and follow the renamed ones.

* chore(ci): count the runtime archive download as a runtime launcher path

* fix(orcad): pin the macOS C++ standard for node-pty prebuilds

The official Node headers' config.gypi sets clang: 0, so common.gypi skips its
gnu++20 xcode_settings and Apple clang 15 (macos-14 runners) compiles
node-addon-api as C++98.

* fix(orcad): resolve the preflight's slot through realpath, as the handoff does

A symlinked orcad.js handed off to its real slot's pinned Node, but the
startup and profile preflights read the symlink's directory, found no
runtime marker there, and silently skipped the readiness check.

* refactor(ssh): drop materializeCachedNodeRuntime, which nothing calls

Deploys upload the verified official archive (design D5); no client path
needs an extracted Node executable cached by digest.

* test(orcad): gate the Bun-to-Node upgrade and Node-to-Bun rollback with live terminals

Design D7.1 R1/R3/R4 and D7.2. The last Bun orcad and this checkout's Node
slot are installed side by side under ~/.orca-remote, launched and stopped
with the client's own deploy commands, and share one data root. Each
direction proves the incoming orcad adopts the outgoing runtime's daemon
(same PID, same shell, output continues), opens its profile database and
backs it up with its own shipped worker, and that GC keeps the slot the
live daemon was forked from.

The node-server Linux lanes provide Bun 1.4.2 and build that Bun orcad
from main, and run with --cross-runtime. --artifact and --cross-runtime
now make their tests fail on a missing input instead of skipping.

* ci(node-server): pin node:24.21.0-alpine by its multi-arch index digest

* test(ssh): name the runtime archive fixture after its role

* test(node-server): load node-pty from the packaged slot in artifact runs

The node-server lane installs dependencies without building node-pty, and
Linux has no upstream prebuild, so the real-PTY failed-I/O teardown test
(picked up by the pty-subprocess selector) could not load pty.node. In
--artifact runs, alias node-pty to out/orcad's shipped slot so the test
exercises the addon orcad actually runs under the pinned Node.

* fix(orcad): let the Windows profile preflight exit after its PTY probe

On Windows, node-pty keeps the conout worker thread and pseudoconsole alive
until kill(), even after the shell exits. The PTY health probe never killed a
cleanly exited probe, so the packaged preflight printed its readiness line
and then hung until the build's 30s timeout, reported with an empty stderr.

- The probe kills its PTY on Windows after exit and uses the bundled ConPTY
  the daemon spawns with.
- The preflight exits once stdout is flushed; its owner reads to EOF.
- Preflight failures now report code, signal, timeout, stdout and stderr.

* test(node-server): load the slot's node-pty in the real-PTY test, not by alias

A vite alias redirected only ESM imports of node-pty; windows-pty-job and
local-pty-utils resolve it through require, so Windows loaded two conpty.node
copies and the Git Bash job-membership proof read an empty job. The failed-I/O
teardown test now loads node-pty through a fixture that picks the packaged slot
in artifact lanes.

The pty-subprocess selector was a prefix that also pulled in its POSIX-host
sibling unit tests, which pr.yml runs and which were never qualified on
Windows. Select the directory plus the two sibling files that belong here.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 00:39:00 -07:00
github-actions[bot] d4ae6a904b Update README downloads badge 2026-10-01 06:46:31 +00:00
Jinwoo Hong daf63e659c fix(runtime): read Antigravity, Cline and Prime Agent readiness from the live screen (#24222)
* fix(runtime): decide Antigravity readiness from the live screen

agy paints its composer with cursor addressing, so the line-folded wait
text misses the 1.2.14 accept-edits and plan composers and an ended turn,
while the grid keeps the bare `>` caret painted mid-turn and behind the
/model picker. Read the screen's bottom rows instead: rule, caret, rule,
`? for shortcuts`. A clocked pane is held to quiescence (tier 1b) because
the submit repaint reads ready for a moment; a clockless restored pane
settles from the screen alone. When a trustworthy screen exists it
decides, so the name-only title lane no longer settles an open picker.

Retires the visible-read probe's Antigravity branch: the probe now runs
the shared screen rule for any screen-ruled agent without an output
clock, and keeps its generic empty-pane read for everyone else.

Adds twelve agy 1.2.14 recordings and a replay suite shared by
screen-ruled agents. STA-8741.

* fix(runtime): decide Cline readiness from the live screen

Cline paints its composer box with cursor addressing on the alternate
screen, so no text rule saw it and worker-start timed out at
agent_readiness (#23268). Read the box off the grid: rule, an empty
composer with one of the captured placeholders, rule, the Plan/Act row
and the auto-approve row, with no braille spinner above it.

A streaming reply repaints the same box once its spinner has scrolled
away, so Cline is tier 1b only: a clocked pane waits for quiet and a
clockless one never settles from the screen. The screen now decides for
a Cline pane, which shuts the quiet-process lane that would have settled
its unworded tool-approval prompt and the Cline Desktop promo.

readLiveTerminalScreenLines now returns raw rows: the read projection
blanks a composer it takes for a draft, and it takes Cline's
placeholder for one, so a typed draft and an empty composer looked the
same.

Adds nine cline 3.0.66 recordings (macOS) and the 3.0.65 Windows capture
from #23269. STA-8741.

* fix(runtime): decide Prime Agent readiness from the live screen

Prime redraws its composer on the alternate screen, so the text tail
never showed a settled prompt and tui-idle timed out (#22153). Read the
grid instead: a bare `>` directly over the `<- manage` footer, with no
braille status row (`Writing - 6s`) above it. The footer and caret alone
stay painted for a whole turn.

Replayed chunk by chunk, Prime erases that status row before redrawing
it, and on first launch paints the idle composer just before the
trace-sharing question covers it. Both keep repainting, so a clocked
pane is held to quiescence (tier 1b); a clockless restored pane settles
from the screen alone.

Adds nine prime-agent 0.9.8 recordings (isolated HOME, OpenRouter) and
the two 0.9.5 captures from #22154. STA-8741.

* refactor(runtime): drop Cline-only readiness branches

Cline now follows the same pattern as Antigravity and Prime: a screen
rule plus table entries.

- Drop MID_TURN_COMPOSER_AGENTS. onPtyData stamps lastOutputAt on every
  chunk, so a re-attached streaming pane has an output clock from its
  first byte; the exception only guarded a pane that printed nothing
  since attach. A clockless Cline pane now settles from its screen like
  the other two.
- Drop the 'ready-body' rest-signal entries for all three agents. The
  rest signal is read only by quietForegroundLane, and a readable screen
  already shuts that lane and the title lane (isReadinessDecidedByScreen),
  so the entries only removed the quiet-process fallback for a pane with
  no trustworthy grid. The census now checks that screen-shut instead.
- Drop the Cline rule's auto-approve row check; no recorded verdict
  depends on it.

Kept: raw rows from readLiveTerminalScreenLines. Every frame of every
codex-* and qoder-* capture at 120x40, 80x24 and 100x32 gives the same
isKnownReadyPromptBody (with and without a clock) and
isQuietReadyScreenBody verdict through both readers.

Serializer known-failures for the new captures are pre-existing
serializer behaviour, not this branch: row-0 cells restore with a
true-colour background where the source has the default (the DSH
class), and Prime's cursor restores at column 119 instead of the pending
wrap at 120 (the qoder class). STA-8741.

* fix(runtime): trust a screen rule only on the PTY's own grid

Review findings on the screen-ruled readiness (STA-8741):

- A grid out of step with the PTY garbles cursor-addressed chrome, and a
  model resize does not make the TUI repaint. readLiveTerminalScreenLines
  now returns null unless the emulator's grid matches the PTY's reported
  size and was never reflowed without a repaint (a re-attach that learned
  the real size late), so the pre-existing lanes decide there instead of
  timing out.
- The visible-read probe reads the draft-blanking projection, which
  turns Cline's `❯ Ask anything...` into a bare `❯`. It now restores the
  blanked composer row before the rule reads it; `terminal read --screen`
  output is unchanged.
- The quiet lane no longer ORs the text rules over a trustworthy screen
  that refused; without one, tier 1 already ran them. No recorded
  verdict changes.

Tests: ready recordings on a mismatched and on a reflowed grid settle
through the old lanes; the restored-pane probe runs every ready
recording through the real projection; the rest-signal census checks
the lane verdict with and without a screen.

* test(runtime): trim STA-8741 recordings to the screens they prove

* refactor(runtime): one screen verdict for every screen-ruled lane

readScreenRuledReady, readScreenRuledQuietReady and isReadinessDecidedByScreen
each re-derived the same thing: the agent's rule applied to a trustworthy live
screen. They collapse into readScreenRuledVerdict (true / false / null), which
tier 1, tier 1b and the lane gate read.

This also makes a refusal final in tier 1: a clockless pane whose trustworthy
screen refused fell through to the text rules, so retained ready text could
settle over an open picker (Greptile review). The quiet tier already refused
there; now both do.

The tier-1b agent set derives the screen-ruled agents from the rule table
instead of listing them again, and the lane test that repeated the census
case is dropped.

* refactor(runtime): let the visible-read probe read its own output clock

The probe's clock was captured at start and threaded through the wait
dependencies as a one-off parameter. The probe now reads it from the live
record when its screen read returns, which is also the fresher answer.

* fix(runtime): trust a reflowed grid again once a PTY resize repaints it

The reattach-reflow flag was never cleared, so a pane stayed on the old lanes
for the rest of its life even after a real resize made the TUI repaint
(Greptile review). The record now keeps the reflowed grid, and a PTY resize
off that grid clears it; an echo of the same size sends no SIGWINCH and keeps
it.

Tests: the reflow case in every screen-ruled suite now includes a same-size
echo, and an Antigravity recording only the screen reads ready settles after a
resize and repaint.

* refactor(runtime): keep screen-rule trust and raw rows to screen-ruled agents

Two shared changes reached agents this PR does not target: the live
screen reader returned raw rows, and it refused a grid that did not
match the PTY. Both now live in readScreenRuledLines, which only the
screen-ruled agents read (screenReader picks it from the rule table);
readLiveTerminalScreenLines is main's again. The probe keeps main's
Antigravity-banner trigger, so a Codex or unknown pane is probed exactly
as before.

Proof: the non-screen-ruled suites give identical pass sets on this
branch and its base (1,781 tests), and replaying every other recording
frame by frame through the readiness and blocked verdicts, for its
agent and for an unknown pane, gives identical results (93 pairs). A new
test keeps a Codex pane reading its screen when the PTY reports another
grid; it fails if the trust check moves back into the shared reader.
2026-10-01 00:02:39 -04:00
Brennan Benson a4606ccae3 fix(cli): orca file open no longer moves your view unless you pass --focus (#24244)
* docs(cli): file open/diff/open-changed say they switch the user's view and are for user requests only

Refs #9944

* fix(cli): file open/diff/open-changed leave the user's view alone unless --focus

`orca file open`, `file diff` and `file open-changed` always switched the
desktop to the target worktree, selected the tab and revealed it in the
sidebar. An agent skill that opens its answer pulled the user out of whatever
they were typing in (#9944), and a phone opening a file moved the desktop too.

The commands now add the tab in its worktree without changing anything on
screen, including when that worktree is the one being viewed: the new tab is
added to the tab bar but the active tab, tab type and focus stay put. In a
worktree the user is not viewing, the tab becomes that worktree's selection so
it is in front when they go there. `--focus` keeps today's behavior.

files.open / files.openDiff take an optional `navigation` target (the existing
RUNTIME_NAVIGATION_TARGETS vocabulary); the CLI sends 'all' for --focus, like
`worktree create --activate`, and nothing otherwise. The renderer moves the
host view only when the target reaches the host; a missing field (phones,
older CLIs) leaves it still. Editor opens for a worktree other than the
on-screen one no longer write the global activeFileId/activeTabType.

Refs #9944

* test(cli): justify the window and runtime stubs in the file-open notification test

* fix(cli): keep phone file opens switching the desktop; the CLI asks for 'caller'

Phone opens send no `navigation` field, and the phone's diff-review "Open in
session" relies on the desktop selecting the diff it opened. A missing field
now keeps the original switch exactly; the CLI says what it wants instead:
'caller' (no host move) by default and 'all' for --focus. Older CLIs, which
send nothing, keep switching as they always have.

Refs #9944

* fix(cli): background file opens select the tab without counting as a visit

A CLI open into a worktree the user is not viewing selected the new tab with
the same activation a user click uses, which stamps lastFocusedAt and the
group's recency list. The worktree jump palette sorts recent tabs by that
time, so every agent `orca file open` into another worktree jumped to the top
of the user's recent tabs.

Editor opens now take a selection mode: 'focus' (default, unchanged),
'background' (select within its worktree without recording focus or recency)
and 'none' (add only). createUnifiedTab and activateTab gain recordFocus:false
for the background case.

Also: tests for reopening an already-open file or diff without --focus, a
comment that file opens move only the host window ('all' acts as 'host'),
root help lines back under 100 columns, and an accurate remote test title.

Refs #9944

* fix(tabs): a background-selected tab still joins its group's tab history

recordFocus:false skipped both the focus-time stamp and the group's
recentTabIds append while still making the tab the group's active tab. Ctrl+Tab
looks the active tab up in that history, so after a background CLI open it
did nothing (or went to the wrong tab) once the user switched to that
worktree, and hydrate kept the broken history across a restart.

Only the focus-time stamp is skipped now; the jump palette's recent rows sort
by that alone, so the palette fix stands.

Refs #9944

* fix(cli): file open/diff/open-changed --focus help says it brings the user to the file

The three commands borrowed the shared --focus line written for terminal
create ("Reveal the created terminal session in Orca"). They now use the
per-command flag help table; terminal create's line is unchanged.

Refs #9944
2026-09-30 20:46:47 -07:00
Brennan Benson 0b79720c2e feat(native-chat): the chat strip and the sidebar read the host's child records (#22614)
* feat(native-chat): publish the host's child records to the status summary and the chat strip

The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.

* feat(native-chat): the chat strip reads the host's child records with its parent's verdict

The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.

Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.

* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open

- The view decoder ignores unknown keys, degrades unknown kinds, states,
  outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
  roster of finished children and never the views themselves; a stop-only
  reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.

* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered

* test(native-chat): type the switch tests' mocks instead of asserting them

* test: remote clients advertise reading child views

* docs(agent-status): the structured row folds the store's child records

* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary

The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.

* refactor(native-chat): the status summary's broadcast equality gets its own module

The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.

* fix(native-chat): command admission reads the strip's child records

A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.

Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.

* refactor(native-chat): command admission takes only what it reads of a turn

* fix(native-chat): the session list drops a session's children when the store does

A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.

The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.

* test(native-chat): write the Codex frame script's parent row out step by step

Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.

* fix(native-chat): the idle sweep and the restart snapshot read the host's child records

The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.

The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.

* test(native-chat): the child-record tests follow the merged command lifecycle

A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.

Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.

* refactor(native-chat): the status feed's journal projection cache gets its own module

The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.

* test(native-chat): the admission test's compaction resolves with a real outcome

Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.

* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished

The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.

This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.

* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source

`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.

A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.

* test(native-chat): the switch test passes the startup child key main's status bar takes

* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own

Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.

Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.

* fix(native-chat): a background Stop reaches the tasks the child records show

The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.

The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.

* fix(native-chat): one rule for a finished child that still owns live work, at any depth

The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.

* fix(native-chat): an older client sees a Codex child's shell as it did before views

Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.

* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives

The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.

* fix(native-chat): the strip channel forgets a closed conversation's roster

It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.

* docs(native-chat): rewrap the retention comment

* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays

Two lifecycle gaps from the round-1 fixes.

A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.

A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.

Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.

* fix(native-chat): the strip keeps one empty list for a roster that omits one

A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.

* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent

The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.

* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once

A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.

The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.

Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.

* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader

CI on dbd2439cd4 was red in three places:
- first-work-branch-rename and the agentSession.subscribeStatus RPC test feed the status feed a
  journal whose snapshot lists items only. The projection reads the user's newest accepted send
  from `snapshot.submissions`; it now tolerates their absence, as the status projection beside
  it already did.
- the cross-version downgrade test still passed `backgroundTasks` to the teardown's working
  marker, which now takes `childWork`.
- an e2e unit test still gave the status feed the removed `readBackgroundTasks` dependency
  (harmless at run time, a type error in the tests/ project).

* fix(native-chat): the chat strip lists running children only, by the sidebar's rule, and hides when none runs

A finished subagent's result is already in the transcript ("Ran N subagents ·
completed"), so the strip is for work that runs. It now lists exactly what the
sidebar lists, by one predicate (a running child, or a finished one whose own
shell still runs, which reads monitoring), and the host sends no roster once none
runs, so the strip hides.

Gone with it: the 100-row budget and the running-then-newest-finished
selection, the re-homing of a child whose owner the budget cut, and the RPC
gate's rule for a roster of finished rows only (no such roster exists now).
Older clients still get their derived task list, running work only.

Finished records still stay in the host's store until the user's next accepted
message: they refuse a late frame of their run, let a task's own ending replace
an acknowledged Stop's, and keep a running shell's owner. Dating that retention
by when the user wrote the message only kept finished rows visible longer, so it
is removed.

* fix(native-chat): the strip shows running work only from any host, and hides after a released session's last child

- A new app paired with an older host no longer shows that host's finished task
  rows: the strip lists running work only, whatever host sent it, and hides when
  an older host's roster has only finished rows left.
- A test for the path that hides the strip when a session's last running child
  settles after the provider let go of the session (Claude's release path): the
  channel sends `null` though no provider answers for the session any more.
- A test comment still described the strip keeping finished children.
2026-09-30 18:23:24 -07:00
Brennan BensonandClaude 24edf0f64b fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm (#23467)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm

When the provider gave no verdict, the structured status projection now derives
one from the newest turn's lifecycle: interrupted -> interruption, unverifiable ->
unconfirmed. Nothing new is journaled, the completion feed stays provider-only, and
the legacy interrupted flag stays a user stop only. Every verdict reader handles
both arms explicitly.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* fix(native-chat): a folded turn a crash cut off reads Interrupted after N

The settled-turn timing now carries the turn's verdict, derived by the same
agentTurnVerdict the status row uses. A turn that ended interrupted with no
provider verdict heads its fold 'Interrupted after N'; a user's stop keeps
'Worked for N'.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* fix(native-chat): the user's close of a chat records the turn it cuts short as their cancellation

The expected-close settle writes outcome cancellation when the user aimed the
stop at this chat: a Stop while the agent starts, the chat's tab closed (the
agentSession.close RPC, or session.tabs.close with a user reason), or /clear.
A quit, an idle eviction, a worktree teardown or an orchestration stop leaves the
turn with no verdict, so it still reads Interrupted.

* test(native-chat): a Claude turn a newer send superseded reads Interrupted

The supersede fires for any send Orca dispatched, the user's or another agent's,
and nothing at that site records the sender, so the turn keeps no verdict and
folds as Interrupted after N.

* fix(native-chat): the user's close records cancellation on the turn the provider settled on its way out

The Codex and Claude adapters settle their open turn as interrupted, with no
verdict, while the host stops them. The expected-close settle then found no
running turn, so a user's close of a mid-turn chat read Interrupted. The stop
now reads the running turns before it reaches the provider and records the
user's cancellation on each one it cut short, unless the provider gave a
verdict of its own. The close-verdict test's adapter now settles its turn on
close the way the real adapters do.

* test(native-chat): update the close and settled-turn expectations for the host-observed verdict

agentSession.close now passes the user's word to the host, and a settled
interrupted turn with no provider verdict carries `interruption`. Also merge a
duplicate import the code-quality gate rejects.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): the user's close cancels a turn whose start landed as the provider stopped

The close read which turns were running before the stop. A turn whose start was
still in flight (a send echo not yet journaled) was absent from that read, so the
provider's verdict-less settle on the way out left it Interrupted. The close now
reads which turns were already over instead, and records the user's cancellation
on every other turn the stop left running or interrupted with no verdict. A turn
cut off earlier, or finished during the stop, keeps its end.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

* refactor(native-chat): the adapter settles the turn a stop cuts with the stop's typed cause

The host hands its stop's cause to the adapter's close. Each adapter settles its own
open turn on 'ended' through one mapping, turnVerdictForChildEnd, and the host's
dead-generation fallback uses the same mapping for any turn no adapter settled. A
user's close or stop of this chat is their cancellation; a quit, eviction, teardown
or an exit the adapter saw first is news.

Deletes the snapshot-and-diff reconstruction (endedTurnItemIds,
userStoppedTurnRevisions, settleUserStoppedTurns) and the requestedByUser flag.
host.close now takes a required cause.

* test(native-chat): a user's close drops the chat's status row like an eviction

* test(native-chat): expect the eviction cause on the host closes of idle release, worker stop and worker discard

The stop's typed cause now travels into host.close and the adapter's close, so these
three non-user closes assert the 'evict' they pass.

* refactor(native-chat): every stop names its cause, so none defaults to the user's cancellation

stopStructuredAgentSessionAgentUnderSerialize defaulted its ending to 'user-stop', which now
settles the cut turn as the user's cancellation. Every caller already passes a cause; the
parameter is now required, and a type-level test fails to compile if the default returns.

* test(native-chat): pin who a chat's session.tabs.close speaks for, older clients' reasonless close included

The mapping lived inline in a type-unchecked file, and only the explicit user reason had a test: an older client's reasonless close, or a lifecycle echo read as the user's, stayed green. It is now one exhaustive, type-checked function with a case per reason.

* style(mobile): draw the unconfirmed dot in the theme's status amber, not an inline hex

* docs(agent-status): the main agent's outcome also carries the host-observed end, interruption or unconfirmed

* fix(native-chat): a chat the user closed while its agent started is not a failed start

A still-starting child the user's close cut counted as a failed start, since only 'user-stop' was excluded: a start-failure row, and queued messages rejected as a provider failure. Whether an ending fails its start is now one exhaustive switch, shared by the failed-start read and the delivery loop's handover: a user's stop or close never does; an exit, a failed attach and the host's own stops still do.

* fix(native-chat): a chat the user closed closes its queued messages, and starts no agent for them

After 876b6989f1 a user's close of a still-starting chat went on like a Stop, so when the close did not complete the delivery loop started a new agent for the message queued behind it. A child's end now has three dispositions, not a failed-start boolean: a user's Stop lets the queue go on, the user's close closes what was queued before it, and any other end fails it. The close is the one a completed close does (the provider-closed rejection, no verdict), applied at the top of each delivery step and ordered against the close so a later send still goes on.

* fix(activity): a crash-cut turn draws the interrupted glyph; only a user's Stop keeps the done check

The Activity page drew every Interrupted row with the done check, which #2569 chose for a user's Stop. With a crash now reading Interrupted, that put a green check on a turn nobody asked to stop. The row's glyph is now an exhaustive switch over the verdict: a cancellation keeps the done check, and an interruption draws the existing interrupted dot. An unconfirmed end already drew its own glyph.

* fix(activity): the Interrupted group header draws the done check only when every row is a user's Stop

A user's Stop and a crash share the Interrupted group, and its header took its newest row's glyph, so a Stop newer than a crash put a green check over the crash. The header is now folded over the group's rows: the done check only when every row draws it, the interrupted dot otherwise.

* fix(native-chat): a failed close of what the user closed starts no agent for it

Rejecting the messages a user's close left queued swallowed a journal write failure, so the delivery step went on to start an agent for a message in a chat the user closed. The rejection now reports whether it landed, and a step whose rejection failed stops instead; the next wake re-derives and retries it. Also pins that the ordering against the close holds only within its epoch, since a later epoch's sequences restart.

* test(native-chat): the idle sweep's stop is an eviction, so its close carries that cause

Main's idle sweep now stops an idle agent through the conversation lifetime, which this branch gives the 'evict' cause; its expectations name it.

* fix(native-chat): a retried stop keeps the cause of the stop it finishes

A user's Stop or close whose wind-down failed after the child was proven gone was finished by the idle sweep as an eviction, so the turn it cut read Interrupted. The owed wind-down now carries its stop's cause, and a retry with no child settles with it.

* test(native-chat): the idle sweep's close of a retrying Claude chat carries the eviction cause

Main's new test expected the adapter close with the session id alone; every stop now names its cause, and the idle sweep's is 'evict'.

* fix(status): a user's Stop marks done on the tab and sidebar; red Interrupted is only a turn cut short by something else

The tab, the worktree card and the sidebar rows drew a Stop with the same red dot as a crash. The
verdict mark now maps a cancellation to done, still saying "Interrupted by user" in the row text,
and the mobile mirror follows. The Activity page keeps grouping a Stop under Interrupted with the
done check, as before.

* test(cross-version): a new phone reads a user's Stop as done; an old phone still draws it interrupted

* fix(native-chat): a user's Stop inside a live Claude chat reads as their cancellation

Stopping a running Claude turn interrupts it and keeps the session, so the turn's end comes from
the CLI's result frame. Claude CLIs before 2.1.91 send that frame with no terminal_reason, and later
ones may still omit it, so the user's own Stop was recorded as a failure with an error row.

Orca now records the stop on the open turn when it sends the interrupt. An error result for that
turn reads as the user's cancellation whatever reason the CLI gives. The stop belongs to that one
turn, so it cannot reach the next, and it is withdrawn when the CLI refuses the interrupt.

* docs(agent-status): a user's stop marks done; name the tab close cause by its type

The reference still said a stop marks a row interrupted and ranks between live work and an
unconfirmed end. A cancellation now marks done, and only a turn cut short by something else ranks
as interrupted. The runtime's tab close restated the close cause's union; it now uses the type.

* test(native-chat): a proven crash reads as an interruption on the status feed and in the chat

A crash the relaunch proves now settles its turn interrupted, and the status feed works the verdict
out from that record, so the restart test expects interruption for a proven crash and unconfirmed
for one it cannot prove, never a cancellation. A chat read before the proof lands reports
unconfirmed, then interruption and a folded "Interrupted after 27s" once the proof revises it.

* fix(status): a user's Stop reads Interrupted, and a turn anything else cut short reads Failed

The verdict mark now maps a cancellation, the user's own Stop, to interrupted, and an interruption,
a turn cut short by a crash or a killed agent, to failed, the same as a failure, which outranks live
subagent work. An unconfirmed end is unchanged. This applies to every agent, in a terminal or a chat,
on the tab, the sidebar rows and worktree card, the dashboard row, Cmd+J and the phone. A Stop is
not news, so the rollups rank it below an unconfirmed end, and notifications word an interruption
"failed". Recording is unchanged.

* fix(activity): group a user's Stop under Interrupted and a crash with failures

A user's Stop draws the interrupted glyph and sits alone in Interrupted, and a turn anything else cut
short sits in Failed, titled "Agent failed". Every row in a status group now draws the group's own
glyph, so the header is the group's status and the rule that folded a Stop's done check into the
header is gone. Interrupted ranks below an unconfirmed end, as in the sidebar.

* fix(status): draw a user's Stop in the muted tone, not the fault red

The interrupted dot, which now means only a user's Stop, draws in the muted foreground token on the
agent rows, the sidebar card and the phone. Red stays for a failure or a turn cut short by anything
else, and green for a finish.

* fix(native-chat): fold a stopped turn as "Interrupted after N" and a failed or crash-cut one as "Failed after N"

The settled turn header now follows the verdict mark: a user's Stop reads "Interrupted after N", and
a failure or a turn anything else cut short reads "Failed after N", under the new key
components.native-chat.status.failedAfter in all six catalogs and the boot catalog. Desktop and
phone share the one description, so they agree.

* docs(agent-status): describe the Interrupted and Failed marks

The reference and the phone's turn bar still described a user's stop as done and a crash as
interrupted. A fault now reads failed, a user's stop reads interrupted in the muted tone, and the
rollups rank an unconfirmed end above a stop.

* test(status): a crash the relaunch recovers marks failed

The recovery test still expected a recovered interruption to mark interrupted; it now marks failed,
as a failure does. Formatting only elsewhere.

* fix(native-chat): record a turn a newer request replaced as superseded, and show it Interrupted

A Claude turn that a newer send replaced before its result arrived was recorded as interrupted with
no verdict, which reads as a turn cut short by something else, now "Failed". It is now recorded with
its own outcome, `superseded`, where the replacement is detected. That outcome names no sender, so a
dispatch from another agent is never recorded as the user's Stop, and it sets no legacy flag.

Every reader handles it in an exhaustive switch: it draws the muted Interrupted mark with the plain
text "Interrupted", folds as "Interrupted after N", and attention demotes it with a Stop, through
the renamed agentTurnEndedOnRequest. Older builds read an arm they do not know as no verdict, which
is what this turn carried before, so their rows keep reading done; the cross-version suites pin an
older desktop's journal and status readers and an older phone.

* refactor(status): name the attention predicate for a turn ended on purpose

agentTurnEndedOnRequest becomes agentTurnEndedOnPurpose: a user's Stop or a newer request's
replacement, never a fault. The Claude turn-end comment no longer says a replaced turn carries no
verdict.

* test(native-chat): a turn cut off by a restart or by quitting Orca reads Failed after N

On main a restart-cut turn shows the done tick. Pin the chat's turn bar and
the tab's mark for both cuts, through the recovery settlement and the quit's
child-end mapping, and pin the quit's turn bar through the host's own quit.

* style(native-chat): format the superseded turn-bar expectations

* fix(claude): a Stop that names no turn is the user's stop of the open turn

The chat's Stop button names no turn. Claude's conversation Stop recorded the
user's stop only for a named turn, so an older CLI's error result after that
Stop read Failed. It now records it on the open turn through the same intent,
dropped when Claude refuses the interrupt and never carried to the next turn.

* fix(native-chat): keep the attach context's publishStatus required

The lifetime context type makes publishStatus optional, so the attach context
that spreads it no longer satisfied its own type once its duplicate
publishStatus went. The host's lifetime context is now inferred, and checked
with satisfies, so the spread carries the member it always sets.

* fix(activity): rank a user's Stop below live work in the status grouping

The Activity page's status grouping put the Interrupted group (a user's Stop, or a
turn a newer request replaced) above Working and Monitoring, so a Stop still sorted
like news there while the sidebar, worktree card and Cmd+J rank it below live work.
It now follows live work and stays above Done; Failed and Couldn't confirm keep
their places above live work.

* docs(agent-status): say which turn outcomes the journal records and which are derived

The journal now records superseded as well as the provider's verdict and a stop;
interruption and unconfirmed are derived on read. The resume row no longer claims
interrupted renders red.

* test(native-chat): the idle sweep's held-send rest closes with the evict cause, like its siblings

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-30 15:02:44 -07:00
github-actions[bot] 8cc28545a2 Update README downloads badge 2026-09-30 12:44:27 +00:00
Neil d2dbe2c385 fix(windows): replace the managed CLI launcher with a native one (#24094)
* docs(security): add the antivirus clearance path for future releases

Every AV false positive here has been handled one vendor and one shipped
version at a time. Document the programs that clear future releases instead --
signer and product enrollment rather than per-build sample submission -- and add
a script that reports an RC's current detection state by hash, so a verdict is
found before users meet it in an issue report.

Hash lookup only by default; --upload transmits the artifact and stays manual.

* fix(windows): replace the managed CLI launcher with a native one

resources\bin\orca.exe was a csc-compiled MSIL assembly: a small, freshly
compiled .NET image in a user-writable directory that mutates environment
variables and proxies a child process. That is the shape .NET dropper
heuristics are trained on, and every verdict against it named the family --
MSILHeracles from two vendors, Wacatac!ml from a third. Signing the file does
not change its shape, so signing never cleared it.

Rebuild it in Rust. Same resolution, same environment contract, same argv
passthrough that keeps newline-bearing orchestration bodies intact (#8374), and
the child still inherits our environment block rather than an explicit map, so
a block carrying both PATH and Path survives (#12046). The PE now carries
publisher, version, icon and an asInvoker manifest from build.rs.

Refs #23383

* ci(windows): install the Rust toolchain before building the CLI launcher

The hosted runners happen to ship cargo, but a real Windows dev box does not --
verified on our own Windows QA host, where cargo and rustc were both absent.
Relying on the image means a future image change fails deep inside
electron-builder's native hook instead of at an obvious step.
2026-09-30 02:49:03 -07:00
d68a5be13a fix(claude): run Windows hooks without shell operators (#23944)
* fix(claude): run Windows hooks without shell operators

Keep neutral replies inside the managed entry and payload scripts, repair missing files from managed registrations, and stop using Git Bash discovery to guess Claude's hook shell.

Co-authored-by: latte271 <junghyeyun27@gmail.com>
Co-authored-by: Bing.Z <zzb@gxsmjx.com>

* fix(claude): keep the Windows hook refresh async and its scripts after uninstall

- List windows-hook-files.ts in the CLI project so typecheck passes.
- Refresh the entry/payload pair only from a surviving entry, with an async
  existence check, so startup refresh stays off the main thread on Windows.
- Keep both scripts on uninstall like every other agent; a Claude session
  still holding old settings keeps answering instead of erroring per event.
- A payload that exists but cannot start falls through to the neutral reply,
  and the missing-payload branch exits early for background jobs.
- Update the EDR posture reference for the operator-free command.

* test(claude): run the Windows hook host legs for real

The live Windows host legs never ran: runProcessSync cannot take a string
stdin (it forces encoding 'buffer'), so every leg threw before starting a
host. Use async runProcess, pass PATHEXT (without it Windows PowerShell 5.1
prints nothing and exits 0 for a .cmd path), and name the host in each
assertion. Drop the POSIX pwsh leg: its drive-mapping shim proved nothing
about Windows, and the Windows legs cover both PowerShell hosts.

---------

Co-authored-by: latte271 <junghyeyun27@gmail.com>
Co-authored-by: Bing.Z <zzb@gxsmjx.com>
2026-09-30 01:07:09 -07:00
Jinjing f33f3093cb docs(wechat): point community QR code at group 11 (#24050)
Group 10 is full; swap the README QR code and copy (all locales) to the new group 11 invite.
2026-09-30 00:11:49 -07:00
Brennan Benson 21d4ae9448 feat(agents): pre-trust the folder wherever Orca starts an agent (#23744)
* feat(claude): pre-trust worktrees Orca creates

Claude Code asks "Do you trust this folder?" on first launch in any folder it
has not seen, which blocks unattended launches in worktrees Orca itself made.
Orca now records where a worktree's content came from when it creates it, and
before each Claude launch writes Claude's own folder-trust entry for that
worktree's root (never the main checkout) when the new setting is on and the
content is the user's repository. Forks, bare commits, folder workspaces and
external checkouts keep Claude's prompt. The write takes Claude's lock, never
creates or breaks the file, runs on the SSH host itself, and is revoked when
the worktree is removed or the setting is turned off.

Launches that already pass --dangerously-skip-permissions also skip the trust
prompt for that one process only, via CLAUDE_CODE_SANDBOXED=1 on the command.

* fix(claude): parse the relay trust request with a schema and ship its search keys

* fix(claude): never write a WSL guest's trust into the Windows config

A WSL worktree's Claude reads the guest's own config. Two paths still wrote
its trust into the Windows host's ~/.claude.json instead: the Claude auth prep's
fallback (runtime 'wsl' but the host config dir, when the WSL home cannot be
resolved), which wrote a Linux-path key the removal revoke can never delete;
and the Agent Teams leader, which passes no auth or distro and wrote a UNC key.
Require the guest's own config dir, and treat any WSL worktree path as guest-only.

* revert(claude): drop the skip-permissions trust shortcut

Pre-trust stays limited to worktrees Orca creates from the user's own
repository. The per-launch CLAUDE_CODE_SANDBOXED prefix skipped Claude's
trust question in every folder for launches carrying the skip-permissions
flag (Orca's default Claude args), including the user's own folders and fork
PR worktrees, and the setting could not turn it off. Remove the prefix, the
inherited-variable strip that existed only for it, and the Agent Teams
leader-to-teammate propagation; restore the tests that pinned the prefixed
launch string.

* fix(claude): revoke SSH trust in the config file the grant used

At spawn the relay resolves Claude's config from the launch env, which carries
a CLAUDE_CONFIG_DIR set in Orca's Claude default env. claudeTrust.converge had
only the relay's own process env, so removing the worktree or turning the
setting off revoked in the default file and the grant outlived the worktree.
Send the config-file keys with the request, as the local revoke already uses.

* i18n(settings): translate the Claude worktree trust setting

* fix(settings): say Claude trust applies when Orca starts Claude

The description said Claude skips its trust prompt in any worktree Orca
created. Trust is written only when Orca itself starts Claude there, so a
`claude` typed by hand in a fresh worktree still asks. Say that, and bring
the es/fr/ja/ko/zh translations in line with the new text.

* fix(worktrees): treat a base on an Orca-added fork remote as fork content

A worktree based on a named ref was always stamped as the repository's own
content, so picking the fork remote Orca adds for a pull request (or a local
branch tracking it) as the base made a fork's code eligible for Claude trust.
At create time, read the repo's `remote.<name>.orca-created` markers and each
branch's tracked remote in one `git config` call; a base on such a remote is
stamped as a fork's content, and a read failure is not vouched for. Remotes
the user added, such as `upstream`, stay first-party.

* perf(claude): revoke worktree trust once per config file, not per worktree

Turning "Trust worktrees Orca creates for Claude" off read and parsed the
whole Claude config once per Orca worktree on the main process. Group the
revocations by config file locally and by SSH connection, and make
claudeTrust.converge take a batch of requests.

* feat(settings): one agent-wide "trust the folder" setting in Settings > Agents

Replace the Claude-only worktree trust toggle with a single setting,
agentWorkspaceTrustEnabled (on by default; unreleased, so no migration).
The row says what it does for every agent: agents Orca starts skip their
"trust this folder?" prompt in that worktree or folder, turning it off
stops new trust while existing trust stays, and while it is off unattended
launches (orchestration workers, automations, the phone) stop at the
agent's trust question until someone answers.

Translations for es/fr/ja/ko/zh. Also restores the `awaitingUnnamed` chat
catalog keys an earlier merge of main dropped from this branch.

* feat(agent-trust): pre-trust the workspace for every preset agent at PTY spawn

Every Orca-started agent PTY passes through one of the two spawn builders
with its declared launchAgent, which survives setup-script wrapping. The
builders now call one hook that, for a fresh launch (never a reattach or
restored pane) with the setting on, applies the agent's trust preset to
the worktree, folder workspace or main checkout it starts in.

- One dispatcher, applyAgentWorkspaceTrust(preset, workspacePath, launch
  context), carries what a writer needs: the final spawn env, the Claude
  managed-account auth prep, the WSL distro and the SSH connection.
- Claude joins the presets on both the claude and claude-agent-teams
  entries. Its writer stays grant-only in claude-folder-trust-file.ts:
  the file Claude reads (CLAUDE_CONFIG_DIR / custom-OAuth suffix / legacy
  .config.json / a WSL guest's own file), Claude's <file>.lock never
  broken and taken only when a write is due, atomic temp+rename keeping
  mode and symlinks, never creating the file, NFC + realpath keys.
- SSH Claude launches forward the optional claudeFolderTrust spawn field
  so the relay grants with its own spawn env; old relays ignore it and
  Claude asks. Other presets keep the SFTP writer. A WSL launch never
  writes the Windows home: non-Claude presets skip it, Claude writes the
  guest's file or nothing.
- Codex keeps the 20 s deadline its shared config lane needs; every other
  preset gets 1.5 s. A miss means the agent asks; trust bookkeeping never
  fails or blocks a launch.

Removes the Claude-only machinery this replaces: the eligibility/host/
lifecycle/spawn modules, the persisted creation content-origin field and
its classification, revoke-on-removal, the setting-off sweep, the
claudeTrust.converge relay method and the Agent Teams leader special case
(the leader pane now spawns through the hook with the claude preset).
The agent config types move to tui-agent-config-types.ts so the config
table stays under the line budget.

* refactor(agent-trust): delete the pre-spawn trust writes the spawn hook replaces

The spawn hook is now the only owner of agent folder trust, so remove
every other writer:

- the agentTrust:markTrusted IPC channel, its preload bridge and types,
  and all renderer callers (agent-trust-preflight and its callers in the
  background session, work-item direct launch, session continuation,
  worktree creation, folder workspace composer and session fork);
- the main pre-spawn sites: the createdWithAgent preflight in
  worktree-remote.ts, markLocalWorktreeTrusted/markRemoteWorktreeTrusted
  and the runtime's markWorkspaceTrustedForAgent family with the
  markTrusted ports of the runtime create flows;
- Codex's own launch-prep and resume-prep trust writes.

Each of those launches reaches a spawn builder with launchAgent set, so
the hook covers it. This also fixes a live gap: the worktree-remote.ts
copy of the preset switch omitted Antigravity, so an agy agent started
from a desktop worktree create still asked; the single dispatcher covers
it. Trust is also written on the host the PTY actually spawns on, which
removes the #11163 class of writing the wrong host's config.

* test(agent-trust): type the spawn-builder trust fixtures and prove the spawn waits for trust

The builder test passed untyped args (a string launchAgent) and cast its deps,
which failed tc:node. It now builds both spawn states from a fully typed deps
fixture and a typed restored pane, with no casts.

Adds a case that holds the trust write pending and checks the builder does not
finish until it settles, the ordering the deleted renderer and launch-prep
tests used to cover.

* fix(agent-trust): give SSH trust writes the 20 s deadline again

The dispatcher gave every non-Codex preset a 1.5 s budget, including the SSH
writers for Cursor, Copilot and Qoder, which make several round trips over the
link. Before this PR those writes had 20 s (desktop) or no limit (runtime), so
on a slow link an unattended SSH worker would now stop at the agent's trust
question. SSH writes get the 20 s deadline back; local non-Codex writers keep
the short budget, and Codex keeps 20 s.

The relay's Claude grant keeps the short budget: it writes the relay host's own
disk and does not cross the link.

* fix(agent-trust): never pre-trust a home folder or a filesystem root

A folder workspace can be the user's home folder or a disk root. Claude and
Copilot let a trusted folder cover every folder under it, so pre-trusting one
of those would silently trust everything on the machine for those agents.

One check, isHomeOrFilesystemRoot, now refuses them for every preset: the
dispatcher checks roots and this machine's homes (including the spawn env's
HOME and a cached WSL guest home), the SSH writer checks the remote home it
already resolves, and the relay checks its own home. The agent then asks, as
it would without Orca.

* refactor(agent-trust): drop the Codex launch plumbing that only carried trust

The spawn hook replaced the trust writes in Codex launch prep and resume prep,
which left the fields that fed them unread: CodexHomeLaunchContext.workspacePath
and .launchAgent, the resume prep's workspacePath, and the structured Codex
launch input's workspacePath (plus the extra target lookup that produced it).
Remove them and their plumbing; unavailableManagedHomePath stays.

Also removes test stubs of runtime trust methods this PR deleted, whose
not-called assertions could no longer fail, and two comments that still
described the old trust preflight.

* chore(reliability-gates): point the trust gate at the spawn-time trust tests

The agent-session trust gate still listed three test files this PR deleted
(the renderer preflight, the Codex launch-prep deadline and the e2e trust
completion suites), so check:reliability-gates, which runs in PR CI and in
pnpm lint, failed on missing files. Its invariant also described the deleted
IPC handler and pre-spawn writers.

The gate now covers what replaced them: the spawn builders holding the spawn
until trust settles, the fresh-launch and setting gates, the per-preset
deadlines, and the home and root refusal.

* fix(agent-trust): skip the relay Claude grant for a WSL shell

On a Windows SSH host whose pane shell is wsl.exe, Claude runs inside the WSL
guest and reads the guest's config. The relay still granted trust in the
Windows host's own .claude.json, writing the Windows home for a WSL launch,
which the local path never does. The relay now skips the grant there, so that
Claude asks, as a local WSL launch does when Orca cannot reach the guest file.

* perf(agent-trust): only agent launches wait on the trust hook

Both spawn builders awaited the trust hook on every spawn, including plain
shells, reattaches and agents without a preset. Awaiting even a resolved
promise adds microtask ticks ahead of the pane-spawn reservation check, and
this handler already keeps non-Codex spawns off an await because an extra tick
reorders those reservation races.

The hook now returns null when there is nothing to write, and the builders
await only a real trust write.

* test(agent-trust): keep the home and root cases off any real Claude config

The home and root cases ran the real Claude writer with the test process's
env, so a regression in the guard would have written trust for the home
folder and / into whatever Claude config that env named. They now point
CLAUDE_CONFIG_DIR at a folder that does not exist, and the writer never
creates a config.

* test(runtime): drop needless casts from the launch-host test

The renamed launch-host test kept three `as never` casts on launch options
that already match launchAgentTerminal's parameter type. The changed-lines
casting gate reads the renamed file as new and failed on them.

* test(agent-trust): type the Claude grant mock with the real writer's signature

The mock took an unknown target, so installing the real writer as its
implementation would not typecheck under strict function types.

* fix(codex): drop the launch context the trust move left unread in local spawn env

* fix(agent-trust): queue Claude grants per config file so a launch burst keeps them all

Concurrent grants in one process retried Claude's file lock in lockstep, so each
retry round admitted about one winner. Starting 12 Claude agents at once left 6
of them at the trust question with nothing logged. Grants for one config file now
queue in-process; only Claude's own writes contend for the lock. The relay shares
the writer, so bursts of SSH launches are covered too.

* fix(agent-trust): never pre-trust a folder above a home either

The guard refused only an exact home or a filesystem root. A folder workspace at
/Users, /home or C:\Users was still pre-trusted, and Claude walks up parent folders
for a non-git folder, so every non-git folder in the user's home became trusted.
The guard now also refuses any folder that contains a home, on every host, and is
renamed to say what it decides.

* perf(agent-trust): skip the SSH round trips for Antigravity, which has no remote writer

Every Antigravity launch over SSH now reaches the remote trust writer, which
resolved the remote home and realpath'd the workspace over the link before
writing nothing (the known remote gap). That delayed each launch by two SSH round
trips, and up to the 20 s deadline on a stalled link. It now returns first.

* fix(settings): keep the hidden folder trust row out of web-client settings search

The paired web client hides the host-only "Trust the folder" row, but settings
search still listed it, so searching "trust" opened the Agents pane with no
matching row. Its search entry is now filtered the same way as Agent Awake.

* fix(agent-trust): a failing breadth guard skips trust instead of failing the spawn

* docs(qoder): New Tab now pre-trusts through the agent-wide spawn hook

* fix(agent-trust): never pre-trust a home reached through a symlink

Every trust writer stores the workspace's resolved path, but the breadth
guard compared only the path as given. A folder workspace that is a
symlink to the home folder (or a real home picked while HOME names a
symlinked one, as on distros that link /home to /var/home) passed the
guard, and Claude, Copilot and Cursor then trusted the home itself.

The local dispatcher and the relay now compare given and resolved forms
of both the workspace and each home. The SSH writer resolves the remote
home alongside the workspace, in parallel, so it adds no round trip.
Local non-Claude WSL launches still skip before any filesystem call.

* fix(relay): a failing breadth guard skips Claude trust instead of failing the SSH spawn

The relay ran its home/root guard and homedir() before its catch, so a
throw there rejected the relay's terminal spawn. Same fix as the main
dispatcher's: the whole grant, guard included, is best-effort.

* fix(agent-trust): guard the path each writer stores, not the path Orca was asked to trust

The breadth guard checked the launch's workspace while each writer stored a
transformed path, so every new transformation opened a hole. Codex stores a
linked worktree's main checkout: with a git repo rooted at the home, a Codex
launch in one of its worktrees wrote trust for the whole home.

One relay-safe host module now computes the stored path (Codex's main-checkout
hop, then given and resolved forms of it and of each home), refuses a root, a
home or a folder above one, and only then writes. Main uses it for local and
WSL launches and the relay for Claude. An unknown home writes nothing, and the
WSL home cache is keyed case-insensitively by distro.

* fix(ssh): the relay writes every preset's trust on the SSH host itself

Codex, Cursor, Copilot and Qoder trust over SSH was written from the desktop
over SFTP: four or five round trips per launch, so it needed a 20 s deadline
that outlasted the 8 s draft paste, the 10 s phone wait and the 15 s web-client
create. It also skipped Claude's atomic rename, ignored CODEX_HOME, and stored
the worktree where local Codex stores the main checkout.

The unreleased `claudeFolderTrust` spawn field becomes `agentWorkspaceTrust`,
sent for every preset. The relay derives the preset from the `launchAgent` it
already receives and runs the same host writer main uses, on its own disk,
within 1.5 s and with no extra round trip. Antigravity still returns early on
the relay (its writer is unverified on SSH hosts), a WSL shell still skips,
and any throw means the agent asks.

Deleted: the SFTP preset writer, the remote Qoder writer, the SSH deadline
clause and the desktop-side SSH root pre-check.

* test(e2e): keep CLAUDE_CONFIG_DIR out of isolated Electron launches

The spawn hook now writes Claude folder trust into the config
CLAUDE_CONFIG_DIR names, so an e2e run started from a shell that sets it
could add trust entries to the developer's real Claude config. Also drops
a stale comment that still named Codex launch prep as the trust owner.

* fix(agent-trust): guard Claude's resolve() form of the workspace too

Claude's writer stores both resolve(path) and the realpath. The breadth
guard compared only the given path and its realpath, so a workspace
path that does not exist and climbs back with `..` (for example
<home>/missing/..) passed the guard while Claude stored a key for the
home itself. The guard now also compares resolve(path), so it sees
every form a writer stores.

* test(relay): pty.spawn writes agent trust before the agent's process starts

Nothing exercised the relay handler's call into the trust writer, so
removing that call, or no longer awaiting it, left every suite green
while SSH launches silently stopped pre-trusting. The new case holds the
trust call pending and checks the spawn waits for it, and that the call
gets the request, the declared agent and the final spawn env.

The reliability gate lists the new suite and records the resolve() form
the breadth guard now compares.

* fix(agent-trust): refuse a home only for agents that inherit trust from it

The home and root refusal applied to every preset, so Codex, Cursor and
Antigravity started asking in a home folder workspace, where they did not
before. Only Claude, Copilot and Qoder let trust on a folder cover the
folders below it; Codex matches its start folder or that folder's repo
root, Antigravity the exact folder, and Cursor itself never inherits from
a home, a folder above one or a shallow path. The refusal now reads a
per-preset table in the host module, so the local and relay writers share
the rule.

* fix(agent-trust): trust Codex at the folder it starts in, as before

Before this PR, Codex launch prep trusted the spawn's start folder. The
spawn hook trusted only the workspace root and skipped terminals with no
workspace, so Codex began asking in a floating terminal and in a subfolder
of a non-git folder workspace: its lookup checks the start folder, then
that folder's repo root, and a plain folder above it is neither. The hook
now passes the resolved start folder for presets marked as keyed by it
(Codex only), falling back to the workspace root.

* fix(agent-trust): pre-trust a structured Codex chat's folder, as before

Before this PR, creating a structured (native) Codex chat pre-wrote Codex
trust for its folder through launch preparation. The PR removed that write
and routed trust through the PTY spawn builders, which a structured chat
never passes. Codex's app-server trusts the folder itself only when the
chat's permissions can write it, so a read-only chat started running
untrusted and ignored the project's .codex config. Creating the chat now
calls the same dispatcher, behind the same setting, before launch prep.

* fix(settings): plainer folder trust setting text
2026-09-29 16:03:28 -07:00
github-actions[bot] a5c7dd9671 Update README downloads badge 2026-09-29 18:34:50 +00:00
OrcaWinandm4air e1362ada4c fix(terminal): stop inline-image decoders exhausting the renderer's wasm memory budget (#23499)
V8 reserves an 8 GiB guard region per wasm memory inside its 1 TiB sandbox,
so an Electron renderer can hold only ~124 live wasm memories regardless of
free RAM. @xterm/addon-image instantiated a SIXEL decoder per terminal at
activation (and kept IIP decoders after the first image), so ~120+ terminals
exhausted the budget: new panes raised 'WebAssembly.instantiate(): Out of
memory' rejections, and the next Kitty/IIP image threw 'WebAssembly.Memory():
could not allocate memory' out of the parser, permanently wedging that
terminal's write queue.

The addon-image source patch now borrows SIXEL decoders from a shared pool
only while a sequence is open (color registers stay on the terminal), drops
IIP decoders after each image, and turns a failed decoder allocation into a
dropped image instead of a parser throw. Bundles regenerated with
regenerate-xterm-patches.mjs --write.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-29 01:20:37 -07:00
Neil f8f656ca19 perf(ci): spend fewer concurrency slots per pull request (#23810)
A concurrency slot is charged per job, not per core, and the account's cap is
the scarce resource: standard runner minutes are free and unlimited on a public
repository. Two paths spent slots that bought nothing.

The unit matrix ran eight fixed shards averaging 6.5 minutes each, 3384
job-slots a day and 68% of all slot demand, while the arm pool queued 10.5
minutes at p95 — the queue was the oversharding. Five shards run the same work
in ~10.5 minutes each for three fewer slots per run.

Bun profile persistence escalated to all six platforms on `config/`,
`resources/` and `.github/` wholesale, which took 36.5% of the last 1100
commits through the full matrix where a platform-flavoured predicate takes 19%.
A pull request now qualifies one platform unless the change is platform-
flavoured, and the push to main re-qualifies all six, so an unescalated miss
surfaces minutes after merge rather than at the next cron. Missing changed-file
evidence and an unavailable dependency graph still fail closed to all six.
2026-09-29 00:13:33 -07:00
github-actions[bot] 5560e534ff Update README downloads badge 2026-09-29 06:47:59 +00:00
Neil ccdb324b63 Add CodeBuddy as a built-in coding agent (#23740)
* feat(agents): integrate CodeBuddy launch, status and session history

* docs: record CodeBuddy lifecycle verification

* fix(codebuddy): backfill scoped history and negotiate remote resume

* test(cli): include CodeBuddy in known search agents
2026-09-28 18:11:25 -07:00
Jinwoo Hong a9195eedfa docs: update GitHub star history chart (#23661) 2026-09-28 14:03:01 -04:00
Jinwoo Hong 9c1f9b514e docs(readme): remove the TestFlight link (#23660) 2026-09-28 14:00:21 -04:00
github-actions[bot] aedb9305cd Update README downloads badge 2026-09-28 12:43:36 +00:00
8b410b4893 feat: add first-class Qoder CLI support (#23581)
feat: add first-class Qoder CLI support

Integrate Qoder launch, identity, canonical hook status, trust and resume.
Verify with captured Qoder 1.1.64 transcripts and hidden Electron sidebar checks.

Builds on and cross-reviews #7502, #8611, #9655, #12910, #13311 and #15291.

Co-authored-by: dalveytech-vincent <vincent@dalveytech.com>
Co-authored-by: Eridanus117 <45489268+Eridanus117@users.noreply.github.com>
Co-authored-by: xingqingzzp-gif <xingqingzzp-gif@users.noreply.github.com>
Co-authored-by: jyang2004 <jyang2004@users.noreply.github.com>
Co-authored-by: yunqian <yunqian@alibaba-inc.com>
Co-authored-by: huzhening.hzn <huzhening.hzn@alibaba-inc.com>
2026-09-28 02:59:50 -07:00
400e4e7957 feat(agents): add Freebuff launch and sidebar status support (#23567)
Add Freebuff launch support and execution-host status reporting for the sidebar, including running, question, blocked, and settled states. Validate against captured CLI transcripts and real rendered sidebar evidence.

Cross-referenced community implementations #17065, #20839, and the Freebuff portion of #18790. Preserve their agent/catalog/mobile/documentation coverage and add canonical status publication and regression tests.

Co-authored-by: Harkaran Brar <18134082+harkaranbrar7@users.noreply.github.com>
Co-authored-by: Prarambha369 <98906077+Prarambha369@users.noreply.github.com>
Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>
2026-09-28 02:32:41 -07:00
Neil 9179b93ebf ci: reduce repeated runner work and validate affected-test selection (#23540)
* ci: stage heavy checks and measure affected-test selection

* fix(ci): exercise the real Git boundary in unit selection planning

* Harden review cancellation and CI demand reporting
2026-09-27 23:25:20 -07:00
Neil 45f3512a33 feat(agents): add first-class DeepSeek Harness (dsh) support (#22468)
* feat(agents): add first-class DeepSeek Harness (dsh) support

Register DSH as a supervised Orca agent: catalog entry and detection for its
dsh-tui profile, status/question hooks through DeepSeek's own Claude-Code hook
bridge, composer-ready prompt delivery, session resume, headless Source Control
AI, and title identity that no longer collides with Gemini's.

* fix(dsh): reach Orca through DSH's credential scrub and stop reading its title as Gemini

DSH runs command hooks through its own shell executor, which drops every env var whose
name contains KEY, TOKEN, SECRET or PASSWORD — taking ORCA_PANE_KEY and
ORCA_AGENT_LAUNCH_TOKEN with it, so every hook exited without posting. Mirror both onto
scrub-safe aliases at spawn and restore them at the top of the DSH hook script.

Its title collided too: DSH rests on the same glyph Gemini works on, so a resting DSH
pane was relabelled Gemini CLI and reported working forever. Defer both the Gemini
classifier and the title status detector on DSH's whale, in the base module both copies
of that classifier read.

* test(mobile): repin the session-route closure for the DSH agent icon

* fix(dsh): address review — never splice user rows, cover remote panes, keep the diff off argv

- findManagedDshPatchRegion paired an orphan start marker with a later block's end, so a
  truncated write made install/remove delete the user's own rows. Pair each end with the
  nearest preceding start; regression test fails without the fix.
- The relay PTY env builder never applied the scrub-safe aliases, so remote DSH status
  silently never appeared even with the remote hook installed.
- Source Control AI sent the whole diff on argv; send it over stdin with DSH's '-' marker.
- dsh-tui/dst already chose the interactive profile, so a workspace folder named 'web' or
  'plugin' no longer marks a live agent pane non-interactive.
- Isolate USERPROFILE as well as HOME so a Windows run cannot edit the real home.
- Drop the duplicate README badge and revert an incidental doc reformat.

* refactor(dsh): share the managed-hooks reader and tighten the new modules

Reuse before reimplementing: readManagedDshHookEvents was a near-verbatim copy of Muse's,
with byte-identical private helpers. Both now call one readManagedHookEventsFromJson.

Also: one readTextOrAbsent instead of two spellings of the same read (dropping an
existsSync TOCTOU), one status() builder instead of four inline literals, rmSync(force)
instead of exists-then-unlink, and a redundant empty-string guard before JSON.parse.
The patch-file transforms lose their index juggling for a predicate plus a filter.

* fix(dsh): refuse a flow-style patch file, keep its mode, and stop the relay inheriting a pane

- applyManagedDshPatch matched only an exact `[]`, so `[] # keep empty` or a non-empty
  flow sequence got a block entry appended after it — invalid YAML that would leave DSH
  unable to load the user's own patch layer either. It now strips the token from an empty
  sequence (keeping a trailing comment) and returns null for a non-empty one; install
  reports that and changes nothing.
- The patch rewrite dropped an owner-only file to the umask default (CWE-732); pass
  preserveMode.
- The relay PTY env never dropped inherited pane identity the way the local and daemon
  builders do, so a spawn that specified none could inherit the relay's own and every
  agent's hook would report against that pane.

* fix(dsh): keep the flow-style refusal in every status read, and scope the mode test to POSIX

A refused patch file carries no managed region, so getStatus() fell through to a bare
not_installed with detail null — the actionable 'rewrite it as a block sequence' message
only ever reached the one-shot install() return. Export the predicate and check it first,
behind one shared message constant.

The owner-only mode assertion cannot hold on Windows, where chmod only toggles the
read-only attribute and mode & 0o777 reads 0o666 for any writable file.

* docs(readme): restore the DeepSeek Harness badge lost in the rebase

* test(mobile): repin the session-route closure to the measured 4221

Measured, not derived: 4220 without the DSH icon entry, 4221 with it. Two of the three
modules above main's 4218 pin are not this change's — they arrived with the mobile work
after #22570 and were never repinned; the changelog records that split explicitly.

* fix(dsh): settle tui-idle on the agent's own hook, so supervised workers see it ready

Reported by a tester on the adhoc build: `terminal wait --for tui-idle` ran to its 90s
timeout against an already-ready DSH composer, so a supervised worker never sees the agent
as ready.

Every existing tier reads the title, and DSH deliberately carries no title status: its rest
prefix is Gemini's working glyph, so the detector reports none. A fresh first-party `done`
is better evidence than any title anyway — it is the agent's own account of its own turn,
and normalizeDshEvent drops subagent events, so it is the lead's. Scoped to DSH: for agents
whose hooks report child turns, a mid-turn `done` is the #6011 class this file prevents.

* test(daemon): record the DSH transcript's true-colour I2 divergences

Adding the dsh-tui capture to __fixtures__ enrolled it in the serialize replay sweep, where
it reports 10 I2 divergences and failed the unlisted-transcript default of 0.

Every one is the same shape — visible-grid row=0, a 24-bit background the round trip does
not restore to default — which is DSH's whale intro painting whole rows of true colour.
Verified as an upstream limitation rather than a regression by replaying against the
previous build (build-serialize-addon-at-ref.mjs --ref origin/main): I1 and I3 both hold.

* fix(dsh): return the new tui-idle verdict from the first-party done lane

Main refactored isTuiIdleSatisfied into evaluateTuiIdle, which returns a verdict rather
than a boolean. The DSH lane still returned `true`; it is tier-1 positive evidence, so it
returns READY_STRONG like the title/body lane above it. Re-verified the regression test
still fails without the lane.

* test(relay): pin the scrub-safe pane-identity aliases on the relay spawn path

The relay builds a remote pane's env itself, so the alias mirroring there had no
test: removing the call left every suite green while remote DSH status silently
vanished. Both cases fail without it.

* docs(dsh): point the hook service at the integration reference

The reference doc had no inbound link from anywhere in the repo.
2026-09-27 22:44:18 -07:00
Brennan BensonandClaude 85067494a1 fix(native-chat): a request that failed reads as failed (#22944)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 22:23:49 -07:00
Mr. ZengandNeil ef987e42d7 feat(analytics): persist local usage session identities (#18759)
Co-authored-by: Neil <neil@stably.ai>
2026-09-27 14:29:19 -07:00
github-actions[bot] 6c75837750 Update README downloads badge 2026-09-27 18:30:40 +00:00
OrcaWinandm4air 27b823f934 ci: compile the E2E CLI once for all consumers (#23384)
* ci: share compiled CLI output across E2E consumers

* ci: preserve CLI setup and old-ref fallback for shared artifacts

* docs: record shared E2E CLI benchmark evidence

* docs: include final CLI reuse timing range

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 02:02:29 -07:00
OrcaWinandm4air c15f082031 ci: build independent Electron targets together for E2E (#23378)
* ci: reuse parallel Electron targets for E2E builds and guard cache action setup

* test: recognize the top-level cache repository preload

* docs: record E2E build timings and exact output parity

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 01:17:27 -07:00
Neil 17690e6b9a style: settle oxfmt 0.70 drift and stop formatting vendored licences (#23377)
The oxfmt 0.65 -> 0.70 bump landed without a repo-wide reformat, so 36 files
already in the tree no longer matched what the new version emits. Anyone running
`pnpm format` picked all of them up alongside their own change.

Also excludes `resources/licenses/**`: `oxfmt --write .` was rewriting the
vendored PCRE2 licence, turning its `*` redistribution bullets into `-`. Third
party licence text has to be reproduced verbatim, so formatting must not touch it.
2026-09-27 01:14:53 -07:00
OrcaWinandm4air 47cebbf5d2 ci: use ARM unit runners, overlap web builds, and reuse verifier fixtures (#23376)
* test: reuse isolated mobile bundle fixtures for verifier checks

* ci: run PR unit shards on ARM and overlap independent web builds

* docs: record controlled CI overlap and runner measurements

* test: observe WebRTC packets with the host clock

* ci: isolate Windows installer CIM probe from native test load

* docs: record native probe scheduling validation

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 01:06:54 -07:00
Neil bc78acc43e fix(editor): detect all bundled Monaco language associations (#23371) 2026-09-27 00:32:55 -07:00
Jinjing f5f537ef14 Revert "Support mouse Back/Forward buttons in shortcuts (#23287)" (#23350)
This reverts commit a86fae0889.
2026-09-26 22:58:01 -07:00
OrcaWinandm4air 9f5a8a5b8a Reuse mobile recording compilation and refresh desktop CI timings (#23343)
* ci: reuse recording compilation, split families, and refresh shard timings

* Keep recording suite intact after hosted performance comparison

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 22:44:15 -07:00
github-actions[bot] cdff2a624e Update README downloads badge 2026-09-27 01:08:35 +00:00
Neil a86fae0889 Support mouse Back/Forward buttons in shortcuts (#23287)
* feat: support mouse Back and Forward shortcut bindings

* fix: ignore duplicate mouse shortcut presses until release
2026-09-26 17:53:58 -07:00
OrcaWinandm4air 4b6fe95943 fix(windows): preserve relocated terminals and native process scans (#22872)
* fix(windows): ship the process-table addon to the relocated daemon host

The Windows terminal daemon runs from a copy of the app under
%LOCALAPPDATA%\Orca\daemon-host\<version>. That copy took node-pty but not
@vscode/windows-process-tree, so the daemon's bare require of the addon found
nothing and every process-table read (foreground tracking, descendant sweeps)
fell back to a powershell.exe Get-CimInstance scan (#16905).

- Copy the addon's runtime files (package.json, lib/, the .node binary) into
  the host; the ~25MB of gyp intermediates beside them are filtered out.
- Treat a host missing those files as unmaterialized, so hosts built before
  this are rebuilt, and skip relocation if the install itself lacks them.
- Log the daemon's native/CIM capability at startup and warn once when the
  process table falls back to CIM.

Revives #19525 on current main.

* test(windows): locate update-survival loss before relaunch

* test(windows): preserve daemon tree before update-survival proof

* test(windows): distinguish Electron exit from launcher close timeout

* test(windows): verify process exit when inherited pipes delay close

* test(windows): trace installer process checks in isolated survival runs

* fix(windows): probe process-query capability before installer sweep

* fix(windows): match installer probe and process-check profile behavior

* fix(windows): use NSIS separators for the process-check include

* test(windows): dismiss session-search overlay in survival harness

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 14:31:07 -07:00
github-actions[bot] 3990076fa3 Update README downloads badge 2026-09-26 18:30:10 +00:00
OrcaWinandm4air d17a17684b Reduce redundant CI runs, pnpm uploads, and fixture startups (#23145)
* Reduce redundant CI runs, store uploads, and fixture processes

* Avoid repeating draft-independent mobile checks on readiness

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 01:17:51 -07:00
OrcaWinandm4air 9b30c7f60a ci: verify mobile disposal and balance unit-test costs (#23114)
* ci: verify mobile disposal and reduce unit scheduling costs

* docs(ci): clarify timing assignment validation

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 00:12:06 -07:00