Commit Graph
12345 Commits
Author SHA1 Message Date
Neil 8ff6296bc7 Speed up serializer checks and keep native caches stable (#24476)
* Reuse serializer oracle cells and isolate native cache policy

* Preserve native cache post-save paths and record hosted oracle gain

* Record native cache reuse and separate cancel-test startup budget
2026-10-01 21:43:56 -07:00
Brennan Benson 444f1952c7 ci: run every cross-version wire test, picked up by folder so new ones can't be skipped (#24499)
* ci(cross-version-wire): run the whole directory so no compatibility test is left out

Three cross-version tests ran in no CI job because the job named its files by hand.
Run the directory instead, ratchet that every file kept out of the unit shards
runs in some PR job, and re-run the job when the modules the newly running
tests guard change.

* test(cross-version): give the orchestration downgrade test its siblings' 120 s budget

* ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it

The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are
left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable
workflows those jobs call.

It also only proved that some step names each excluded file, not that the job runs when the file
changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path
trigger matched neither it, its harness nor its subject, so a PR touching only those ran it
nowhere. The check now asserts a change to each excluded file fires a gating job that names it,
and the shell trigger gains those three paths.

* ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver

A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to
orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the
job whose tests guard exactly those contracts. Also corrects the publish/read direction in the
turn-end comment.

* test(cross-version): state why the orchestration downgrade test needs 120 s

* test(ci): glob the unit tree once for the unit-exclusion coverage checks
2026-10-01 20:53:51 -07:00
Brennan Benson 9dea72f0ae fix(claude): an informational note is a warning row in its own words, never a raw frame row (#24471)
* fix(claude): an informational note is a warning row in its own words, never a raw frame row

Claude Code shows info, notice and suggestion notes as transcript chrome; only a warning earns a
row. The frame used to fall through to the provider fallback and print its opcode.

* fix(claude): name the real source of an informational warning in comments and tests
2026-10-01 20:52:29 -07:00
Jinjing 9e34bd06e7 Revert "fix(markdown): return focus to editor from find bar (#23175)" (#24500)
This reverts commit d9d5993cd5.
2026-10-01 19:55:27 -07:00
Jinwoo Hong 9f395208c8 fix(native-chat): move the chat-tab surface out of the session host so main passes lint (#24496)
The host file reached 302 lines after #24203 and #24311, over the 300-line limit,
so the static-analysis job failed on every commit to main. The five chat-tab
members now come from createStructuredAgentSessionTabSurface in the existing
structured-agent-session-host-tabs module; behaviour is unchanged.
2026-10-01 22:51:42 -04:00
Brennan Benson 453408746d fix(build): import the Electron remote capability list from its own module (#24494)
#24203 moved ELECTRON_REMOTE_RUNTIME_CLIENT_CAPABILITIES out of protocol-version.ts into
electron-remote-runtime-client-capabilities.ts. Three files that landed on main while it was in
review (SSH access links and managed orcad server work) still import it from protocol-version.ts,
and a test #24203 added predates main making the structured host's logger option required, so
main's node typecheck fails. Point the three imports at the new module and pass the logger.
2026-10-01 19:41:11 -07:00
Jinwoo Hong e3cb32791e refactor(runtime): read four agents' readiness from JSON rule files through one engine (#24348)
* test(runtime): add a readiness census pinning every tui-idle verdict

Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.

Refs STA-9098

* test(runtime): pin the census quiet probes to literal windows

A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.

Refs STA-9098

* test(runtime): say which census probe writes runtime state

Refs STA-9098

* test(runtime): observe the census through settled panes and caller-visible waits

- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
  of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
  coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
  work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
  directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.

* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix

Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.

* test(runtime): read the census baseline field without Reflect.get

The anti-slop lint rejects Reflect.get on parsed input.

* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files

Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.

The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.

Refs STA-9098

* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes

Refs STA-9098

* fix(runtime): refuse rule patterns that repeat an optional or alternating group

The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.

* refactor(runtime): give agent state rules and text anchors one when/answer shape

Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.

- Cursor's prompt is two anchors answering working and idle; the one-off
  workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
  matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.

* docs: point the readiness evidence docs at the agent state rule files

* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail

A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.
2026-10-01 22:16:06 -04:00
Brennan Benson 757736628f fix(native-chat): a paired server admits structured chat by client capability, not its own chat setting (#24203)
* fix(native-chat): a host admits structured sessions by client capability, not its own chat setting

A host's experimentalStructuredNativeChat decided whether any paired client could reach
agentSession.* at all, and whether session.tabs.* showed it structured tabs. That setting is the
host user's own launch preference: whether a new agent opens as a chat or a terminal is decided by
whoever launches it. Using it as admission control meant a client whose own preference was
"structured chat" was refused on a host whose preference was "terminal", and chats opened while
the setting was on were withheld from mobile once it was turned off.

The gate now asks one thing: did the client advertise agent-session.structured.v1 (in-process
callers negotiate nothing and are always admitted). Tab projection and restore follow the same
rule. With the setting no longer gating anything, the separate cleanup gate (close, cancel,
unsubscribe, release), which existed only so those kept working after the setting was switched
off, is identical to the main gate and is folded into it. The settings listener that republished
tabs when the setting changed is removed, since projection no longer depends on it.

The host setting still picks the default for launches that start on the host itself
(agent.launch from mobile, orchestration worker-start).

* fix(native-chat): negotiate client-chosen launch mode so released phones and old servers keep terminals

Hosts advertise agent-session.structured.client-launch-mode.v1: they admit
structured sessions by client capability alone. A remote client that does
not advertise it (phones released before agent.launch) asks createSupport
to pick the launch mode, so the host keeps answering that with its own
setting, exactly as before. Cleanup methods keep their own named gate so a
future admission condition cannot make close or cancel refusable.

* chore(native-chat): justify the two type assertions this change's lines touch

* fix(native-chat): chats that already exist keep showing whatever the chat setting says

The structured chat setting decides only what new agents open as. With it
off, this machine's structured chats used to be hidden while the host,
which no longer reads the setting, still reported them to the workspace
activation gate, so a workspace holding only a chat opened empty. The
local chat mirror and its startup restore now run whatever the setting
says, the continue-after-restart offer follows the chats that exist, and
the setting's copy says it applies to new agents.

* test(native-chat): pin that a host advertises the client-chosen launch mode

* fix(native-chat): mirror this machine's chats only where it holds them

Round 1 ran the local chat mirror for everyone so existing chats show
whatever the setting says. That gave every desktop a permanent
session-tabs listener, which turns on the runtime's phone replication
paths, plus two full session-tab censuses at startup, and made the
browser client mirror its remote host a second time.

The runtime now says whether it holds structured chats: its structured
host is built only when saved chats were restored at startup or a client
created one here, and it announces the moment one is built. The mirror,
the startup restore and the continue-after-restart offer run only when
the setting launches chats or the host holds some, and never in the
browser client. A chat a paired client creates here with the setting off
still appears at once. The chat behaviour settings show wherever chats
exist, and the setting's copy says it picks what new agents open as. The
toggle-off teardown this made dead is removed.

* test(native-chat): record install listeners without a cast

* fix(native-chat): mirror this machine's chats only once it holds one, not once its host is built

Session history, resume preparation, terminal resume commands and replay-safe phone launches all
build the structured host for users who never had a chat, which turned on the chat mirror and the
structured-only settings rows until the next restart. The signal is now derived from the host's
records (or a records file still owed its import) and pushed when the first chat is restored or
created. A throwing listener no longer fails the install that fired it.

* feat(native-chat): createSupport reports the saved selection a new chat on this host starts with

A chat on a paired server starts with the server's saved model and options, which the desktop could
not read, so its picker showed a guess. createSupport's answer, which the desktop already waits for
before a paired launch, now also carries that seed as a new optional field (older clients ignore it).
Create and createSupport read it through one resolver so they cannot drift.

* refactor(protocol): move the Electron remote client capability list into its own module

Merging main left protocol-version.ts one line over the max-lines limit on this branch. The list of
capabilities the desktop advertises to a paired host moves, unchanged, into
electron-remote-runtime-client-capabilities.ts, the module the next PR in the stack already uses
for it; importers point there.

* test(cross-version): stub the launch seed resolver createSupport now reads

* fix(native-chat): the desktop tells its own host it picks each launch mode, so retrying an existing chat works with the setting off

* docs(native-chat): name the real exit for the released-phone createSupport rule

* test(cross-version): a released client still gets the host-setting createSupport answer; a launch-mode client gets supported plus the seed
2026-10-01 18:55:32 -07:00
Brennan Benson 0d2300ca8e test(e2e): retry the crash probe's main-process reads through the transient-evaluate helper (#24479)
expect.poll does not retry a thrown read, so one spurious 'Resulting promise was
garbage collected' from Electron's main evaluate failed the crash-recovery test.
2026-10-01 18:06:42 -07:00
Neil 4e919f3b5a test: send Codex Ctrl+C to a live raw terminal fixture (#24480) 2026-10-01 17:55:59 -07:00
Neil 02790c53e8 Fix Codex status after Ctrl+C copy and side-chat navigation (#24339)
* fix: preserve Codex status on ambiguous Ctrl+C input

* fix: confirm Codex turn cancellations from host rollout records

* fix: keep ephemeral side hooks separate from the Codex main turn

* perf: watch active Codex rollouts and skip unrelated records

* fix: retain confirmed Codex cancellation across late relay events
2026-10-01 17:33:58 -07:00
Brennan Benson 6e7e964705 feat(orchestration): tell each agent its own orchestration address (#22636)
* feat(orchestration): report the caller's host-resolved orchestration address in orca status

orca status --json gains a caller block: the calling agent's address as the
host resolved it from the identity its environment carries. A structured
session is session:<id>; a terminal agent is its handle, with whether the host
still knows it. A session the host refuses reports that refusal instead.

The host answers through a new read-only orchestration.callerShow, so the
session claim runs through the same dispatch-entry resolver every verb uses.
An older host leaves caller unresolved. The help footer and the run/check
specs stop describing identity only in terminal terms.

* docs(orchestration): tell agents their address and give chat coordinators a non-waiting loop

The orchestration guide now states that a chat session's address is
session:<id> (never the provider's id), that orca status --json reports it,
and that no caller flag should name another agent. A consuming check no
longer tells every caller to name itself with --terminal. A chat coordinator
starts its wave, ends the turn, and on each turn Orca starts for new mail
runs a non-waiting check and ack; it never blocks in check --wait. The guide
also names ORCA_CLI_COMMAND as the executable in chat sessions.

* feat(native-chat): add Copy Orchestration Address to a structured chat's context menu

Copies session:<id>, the Orca-minted address other agents message the chat
by. The existing Copy Session ID still copies the provider's id and is left
as is; the new action is labelled so the two cannot be confused. Strings are
added to every locale catalog.

* feat(orchestration): tell every dispatched worker its own orchestration address

The worker preamble names the coordinator's address rather than a terminal
handle, and states the worker's own address. A structured worker is told it
is session:<id>, that its coordinator reaches it there or at its dispatch
mailbox, and that mail arriving while it is idle starts a new turn. Its
commands invoke the CLI through ORCA_CLI_COMMAND in its own shell's form, the
same rendering the pointer turn uses, because a bare orca in a login shell can
reach a different Orca.

* docs(orchestration): give the ORCA_CLI_COMMAND form for POSIX shells and PowerShell

A chat session's shell reads the variable as "$ORCA_CLI_COMMAND" in a POSIX
shell (Git Bash included) and as & $env:ORCA_CLI_COMMAND in PowerShell, the
same two forms the pointer turn and worker preamble render. The chat
coordinator loop now runs the check its pointer turn names.

* docs(orchestration): say that /clear gives a chat a new address and Orca moves its Runs

* fix(orchestration): keep CLI resolution in the shared skill stub and the orchestration kernel in budget

The guide-contract tests own two rules this PR broke: only the shared skill
stub may describe how to resolve the CLI, and the always-loaded orchestration
kernel stays within 202 lines. The ORCA_CLI_COMMAND text moves to the stub's
resolver block, which now covers chat sessions and login shells beside WSL and
gives the POSIX and PowerShell forms; every skill projection and the bundle
manifest are regenerated. The kernel keeps one line each for the caller's
address, the environment-resolved check caller and the chat coordinator's
non-waiting loop; the loop steps and the address details move to the
coordinator-loop and messaging references. The two kernel pins now assert the
new check contract and refuse the old --terminal <your_handle> shape.

* fix(orchestration): refuse a blocking check --wait from a native chat session

A chat runs turn by turn through a shell tool with its own timeout, so a
blocking wait is killed mid-wait and retried. The host now refuses it with
wait_requires_terminal and the turn-loop recovery, keyed on the session's
lease: a session a terminal view holds still runs in a PTY and may block.

* fix(orchestration): resolve orca status's caller with the verbs' ladder, host-side

callerShow now answers a terminal caller the way the coordinator verbs act:
the carried handle while it is live, else the handle its pane was reminted
as. The CLI always asks, so the host decides that a process has no identity
from the same envelope every verb sends; a pane key alone now resolves.

* fix(orchestration): show a structured worker as session:<id> wherever agents read mail

A structured worker was session:<id> in orca status and its preamble, but
structworker_<uuid> in check rows, banners, reply hints, its own check label
and a sub-worker's coordinator line. The minted handle is now only the
mailbox key: mailbox reads, the check label and preamble coordinator lines
spell the worker session:<id>, which the host binds back to that mailbox.
Send receipts still echo the stored row, whose sender key worker_done
settlement matches.

* fix(orchestration): teach a chat worker the turn loop and pin preamble parity at the contract

The worker preamble was byte-identical across modes except its address, so a
chat worker was taught a 600s blocking ask its shell tool kills before the
message ID for --resume prints, heartbeat exemptions for check --wait, and to
keep a shell open. Parity now pins the contract (sections, verbs, flags,
lifecycle ids); interaction discipline follows the mode: a chat asks with a
5s wait and ends its turn, owns sub-workers through the turn loop, and names
itself session:<id> in every command. The guide says a chat's address
survives /clear and that Orca refuses a chat's check --wait.

* test(orchestration): pass the db to preamble delivery and fence a terminal-view waiter

The coordinator line maps a structured coordinator's handle through the
orchestration db, so delivery takes it from its caller. The consumer-fencing
waiter test now waits as a terminal-view session, the only session kind that
may still block in check --wait.

* test(orchestration): read the Run id with the fixture's checked accessor

* feat(orchestration): copy a chat's conversation address, which /clear keeps

Copy Orchestration Address copied session:<live id>. A chat's address is its
conversation's, derived by the host from the session records, so the menu now
asks the host for it at copy time through orchestration.sessionAddress, the
same derivation a verb acting as that session binds to. A host that predates
the method has no /clear lineage, so there the live id is the address. The
guide's /clear text says the address survives and nothing moves.

* test(orchestration): pin that a cleared chat's successor copies its conversation's root address

* test(orchestration): give the mode-opacity fixture's record store the listing a lineage lookup reads

A structured worker's agent-visible address now resolves through its conversation's lineage,
which lists the session records; the fixture's partial store lacked that listing, so the
sub-worker start failed at dispatch input.

* refactor(orchestration): format a chat's copied and reported address from its root Orca session id

Carries the Orca session id rename into the self-address surfaces.
orchestration.sessionAddress, the copy action's fallback, callerShow and the address a
structured worker is shown now format `session:<id>` from the conversation's bare
root Orca session id with formatOrcaSessionAddress, and ids arriving as strings are
checked with isOrcaSessionId first. The CLI status line, the check caller label and
the dispatch preamble spell the prefix from the one exported constant.

* refactor(orchestration): resolve a session's reported address through the party resolver, and refuse every session's check --wait

- orchestration.sessionAddress, and the agent-visible spelling of a structured
  worker, resolve through the party resolver, so they format the lineage root
  the one id hook derives; sessionAddress.sessionId is classified as a target.
- With the terminal handoff gone every structured session runs turn by turn, so
  check --wait is refused for any session caller, a worker included, and the
  session caller no longer carries its lease's runtime kind.
- The coordinator loop no longer mentions a terminal view, and the messaging
  reference says a chat takes messages but is refused as a Dispatch assignee.

* fix(orchestration): cap a session caller's blocking wait below its shell tool instead of refusing it

A chat or structured worker runs each command under its provider's shell-tool timeout, so check
--wait was refused for every session caller and chats were taught a separate loop. The host now
caps check --wait and ask for a session caller below that timeout (Codex 10s one-shot exec
default, Claude Code Bash 120s) and answers the normal timed-out result, so the terminal
coordinator loop runs unchanged in a chat. A terminal caller's wait is untouched.

* refactor(orchestration): teach a chat worker the terminal worker's preamble, byte for byte but the address

One preamble for both modes: the chat variant (short ask, end your turn, this chat stays
available, ORCA_CLI_COMMAND invocation) is deleted. A structured worker's only difference is its
address, session:<id>; the byte-parity test between modes is restored with just that substituted.

* docs(orchestration): drop every chat-specific instruction; name the address once, generically

The guide, its references, the shared CLI-resolution stub and the help return to main's text,
with one kernel line saying `orca status --json` shows your address (the kernel stays at main's
length). The status caller block reports only the opaque address, the same shape for a chat and
a terminal agent. Guides regenerated.

* test(orchestration): pin that a chat and a terminal agent see the same preamble, pointer and guide

* test(orchestration): key the wait-cap fixture's records by plain session id strings

* chore(i18n): add the copy-address strings at the head of native-chat, clear of main's catalog edits

* test(orchestration): fail the capped-wait test on the settle, not on the test timeout

* test(orchestration): keep main's takeover assertions on a session coordinator's waiting check

With the session wait capped rather than refused, the test main extended runs as it is: the restack re-added the shorter pre-main version over it.

* fix(orchestration): show a /clear-ed chat its lineage root's address everywhere it reads its own

check labelled a session caller with session:<live id>, while orca status and the
preamble show the conversation's root. The CLI cannot read the lineage, so the label
now comes from the same host answer orca status prints (orchestration.callerShow),
asked only when there are messages to render, and falling back to the live id only
when the host cannot say. The host also spells a session address it shows an agent
with the lineage root: a dispatch preview filled in from the chat's own address, and
the provider-id refusal that names a session's address.

* test(orchestration): the parity test's gate facts resolve like the host's

* test(orchestration): the parity test's gate facts carry the submissions main's pointer lane reads

* fix(orchestration): wait a chat's check --wait and ask exactly as long as a terminal's

The host capped a session caller's blocking wait (Codex 6s, Claude 100s) so the
provider's shell tool would not kill it. Neither provider kills a long shell
call: default Codex's exec tool yields and keeps the command running, and Claude
Code moves a timed-out Bash call to the background. Terminal agents run the same
tools uncapped, so the cap only made a chat coordinator re-poll every few
seconds. A session caller's check --wait and ask now wait the budget asked for.

* fix(orchestration): show every agent one address, the mailbox address its mail is keyed by

A structured worker was told `session:<id>` in orca status and its preamble,
but its own send receipts, inbox, worker-list, dispatch previews and task rows
still showed the `structworker_` handle its mail is stored under; only some
reads were re-spelled. Instead of re-spelling reads, callerShow,
sessionAddress and the preamble now report the caller's stored mailbox
address (mailboxAddressOf): a terminal's handle, a structured worker's handle,
and a chat's `session:<lineage root>`. The read-side re-spelling layer
(withAgentVisibleAddresses and its check/banner/preamble call sites) is gone.

Dispatch previews spell the coordinator by its party's mailbox address, so a
`/clear`ed chat's dispatch-show still names its root.

* refactor(orchestration): label check output from what the CLI already knows

check asked the host for orchestration.callerShow after every non-empty check
by a session, only to fill a label used when a legacy row lacks to_handle,
which host rows never do. The label is again the caller's handle or its
injected mailbox address, with no second round trip after mail is consumed.

* refactor(native-chat): offer Copy Orchestration Address on chat tabs only

No mount passes both terminal-pane actions and an orchestration address: a
chat shown inside a terminal pane is that terminal's agent, copied by its
terminal ID. Drop the unreachable terminal-pane placement and its tests.

* fix(orchestration): have orca status report the handle the agent's own check reads

After a window reload a terminal agent keeps ORCA_TERMINAL_HANDLE=term_old while
its pane is reminted as term_new. callerShow reminted and advertised term_new,
but check, send and ask act as the carried handle and never remint, so mail
sent to the advertised address was never read by that agent. callerShow now
answers the carried handle with its liveness, and null for a pane key alone,
from which the mailbox verbs have no identity. Resolving terminal callers once
on the host for every verb is a separate follow-up.

* fix(orchestration): read the renamed coordinator line in the long-prompt repro, and trim round-one leftovers

The reliability repro's fake worker parsed "Your coordinator's terminal handle
is:", which the preamble now spells "Your coordinator's address is:", so it
silently skipped worker_done; it accepts both. Dispatch and its dry-run go back
to main's coordinator line (their `from` is already bound at the entry); only
dispatch-show, whose `from` is unbound, resolves it. Also drops a stale
status-caller comment, trims the wait test to its one uncapped-wait case, and
reverts comment-only churn in the worker opacity test.

* docs(orchestration): keep worker obligation 1 as main words it

The guide grows by the one caller.address line; the parity test bounds the
kernel at main's length plus that line instead of forcing a reword.

* test(native-chat): prove a structured chat tab offers Copy Orchestration Address

Renders the pane-commands hook as a structured chat tab and selects the item:
it asks orchestration.sessionAddress with the tab's target and session id.
Also corrects the menu item's comment to what it copies.

* test(orchestration): D5's tests expect the orca_session_id prefix and 'Orca session ID' wording

* fix(orchestration): name a session by its Orca session ID, and leave terminal agents as main has them

Terminal agents keep main's exact wording: a terminal worker's preamble is
byte-identical to main's, and orca status prints nothing new for them. A
session is named by its Orca session ID (orca_session_id:<id>, its /clear
root's): a structured worker's preamble says "Your Orca session ID is: …"
and its commands use that ID, and a session coordinator is "Your
coordinator's Orca session ID is: …". orca status shows a session caller's
`caller.orcaSessionId`; callerShow answers null for anyone else. The chat
tab menu item becomes "Copy Orca Session ID" with a tooltip saying what the
ID is, and its toasts match. No agent-read text calls this ID an address.
A structured worker's mail is still keyed by its minted handle.

* test(orchestration): check CLI help and status for "address" wording from a CLI test

The node project cannot compile src/cli, so the guard over CLI help, specs
and status text moves to src/cli; both halves share one pattern. Also brings
two comments and the long-prompt repro's coordinator-line regex to the Orca
session ID wording.

* fix(native-chat): keep the Orca session ID tooltip within the tooltip primitive's typography

Drops a restyle the design-system gate refuses on TooltipContent, keeps
"Agent" untranslated in the Japanese tooltip as that catalog does, and types
the test's tooltip mock without an assertion.

* fix(native-chat): the Orca session ID tooltip names the agent CLI's own session ID in the singular
2026-10-01 17:29:04 -07:00
Brennan Benson 8c1670f2c4 test(e2e): fake Codex answers the --no-daemon --help probe without a spawn (#24440)
* test(e2e): fake Codex answers the --no-daemon --help probe without a spawn

Since #23933 Orca runs `codex --help` before a path-named Codex launch. The
fakes logged it as an agent spawn and held the probe for its 5 s timeout.
Move the app-server refusal and --help answer into one shared
FAKE_CODEX_LAUNCH_PROBES_SOURCE used by every fake Codex.

* test(e2e): stop asserting the dispatch capability column a current worker no longer has

#23994 stopped minting the per-dispatch capability, so capability_hash is null.
The --help spawn failure used to stop this test before it got here.
2026-10-01 16:57:15 -07:00
Brennan Benson ccb63afc06 fix(native-chat): a Stop still reads as yours after Orca restarts, because the turn's end reads the Stop's event (#24311)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused

Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.

* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered

A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.

* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget

* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it

The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.

* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card

* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows

Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.

One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.

The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.

Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.

* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones

A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.

* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction

The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.

* fix(native-chat): stop creating the unused queue pause table

The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.

* fix(native-chat): a Stop's pause never hides the restart pause

A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.

Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.

* test(native-chat): pin the Stop's no-resend, lift and held-card rules

- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
  again" at one instant, before a queue ignoring the pause re-sends. They
  now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
  whether or not a person's turn lifts it; it now reads the Stop's pause
  before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
  queued before a rewind.

* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller

The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.

* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event

* test(native-chat): pin that Stop and Resume rows never reach apps or count as history

* test(native-chat): only a person's Stop event pauses the queue

* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop

Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.

* test(native-chat): a card held at a starting agent is checked before the Stop's timing

Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.

* test(native-chat): a released build keeps and folds a journal holding Stop events

Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.

* style(native-chat): format the Stop event changes

* test(native-chat): type the released build's exports through one checked helper

* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only

* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade

The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.

Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.

* fix(native-chat): a Stop that stops nothing new writes no Stop event

A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.

It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.

* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop

* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled

* fix(native-chat): any later Stop event ends a person's Stop pause

A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.

An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.

* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed

A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.

A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.

Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.

* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes

A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.

The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.

* fix(native-chat): a Stop still reads as yours after Orca restarts before the turn ends

Every stop that ends work now writes the Stop's event before it ends the child: a
person's close of the chat, an eviction (worktree teardown, orchestration stop, tab
cleanup) and the idle sweep's stop of a start that never landed. A stop that ends
nothing writes nothing, and quit writes none: its resume marker records why.

The turn-end write reads the latest Stop event where every turn row is built, so the
adapter's settle, the host's fallback and the relaunch's settle all agree: a turn a
person's Stop or close named, ending with no verdict of its own after that Stop, ends
as their cancellation. A relaunch's probe-bounded end is no earlier than a Stop that
found the turn running. When the provider refuses the interrupt and the turn runs on,
a refusal row answers the Stop, so a later crash still reads Failed; pressing Stop
again after a refusal is a new Stop.

* refactor(native-chat): a stop no longer carries its cause; the turn's end reads the Stop event

The cause of a stop was threaded in memory from each entry through the host's stop
step, the adapter router and each adapter's close onto the `ended` it settled with,
and Claude kept a per-turn copy of a Stop it sent. All of that is gone: adapters
settle a turn they cut as interrupted with no verdict, the host's fallback does the
same, and the one rule where a turn row is built (`turnEndAfterStop`) reads the
journal's latest Stop event to say whether it was a person's.

- `closeSession` / `disposeSession` take no cause; `ended` has no `stopCause`.
- Claude reads an error result after a person's Stop as their cancellation from the
  journal's Stop event (through the event sink), not from a per-turn slot, and a
  refused interrupt is the host's refusal row, not `withdrawTurnStop`.
- An owed wind-down keeps no cause: its retry's fallback reads the Stop event.
- The mutation context's Stop passes no cause: its step already wrote the event, and
  the delivery loop's child-end reason is read back from it.
- A Stop pressed before its turn showed applies to the turn that opens under it,
  unless a send a person made since was accepted.

* test(native-chat): a turn a later send opened is no Stop's that named no turn

* test(native-chat): the restart test's death proof carries its detail

* refactor(native-chat): a refused Stop leaves no record; a Stop only ever ends the turn it names

The stop-refused mark is gone: its tombstone kind, its fold, the clock-keyed match that tied it to
a Stop, and the exception that let a second press after a refusal write a new Stop. A Stop that
stops nothing writes nothing. A Codex refusal names a turn that is no longer its active one, and
the Stop names that turn, so the turn running instead never reads as the person's by its id alone.

* fix(native-chat): a Stop pressed before any turn showed stops only the turn opened next

A Stop that named no turn read as the person's cancellation for every later turn that opened
after it, until a send a person made was accepted. The queue's drain, orchestration mail and a
restart continuation send as the host, so a turn they opened long after, cut by a crash, read
"Interrupted" as if the person had stopped it. The Stop now applies only to the first turn
opened after it.

* fix(native-chat): an older Claude's error end after a Stop pressed before its echo reads Interrupted

Claude CLIs before 2.1.91 end an interrupted turn with an error result that names no reason. The
translator judged whether a person's Stop explained it by its own copy of the Stop rule, which
ignored a Stop that named no turn, so a Stop pressed before Claude echoed the send read "Failed".
The translator now writes such an end as interrupted with no verdict and no error row whenever a
person's Stop may name the turn, and the journal's one rule decides as it writes the end.

* fix(native-chat): a person's Stop and /clear each name why they end the agent

The host's mutation path ended the agent with one "recorded" ending for every caller, which read
back the reason of whatever Stop event the journal held last, however old. /clear writes no Stop
event, so its end took an unrelated earlier reason. Each caller now names its own: the chat's Stop
`user-stop`, whose event its own step wrote, and /clear `user-close`, the user replacing this chat.

* fix(native-chat): a host stop judges whether it ends work after the provider's rows land

A close, eviction or host stop decided whether it ended a running turn from the journal as it
stood, while the provider's own rows (the turn its echo opened) could still be in the session's
event sink. A close landing in that gap wrote no Stop event, so the turn it cut read as news. It
now reads after the sink drains, as a person's Stop does, through the same check; a drain that
fails or takes over a second reads working.

* fix(native-chat): a Claude Stop naming a turn that just ended still marks the follow-up it cuts

A phone names the turn it last saw. When that turn had ended and a follow-up was still unechoed,
Claude's Stop interrupted the follow-up and ended the child, but the Stop's event named the ended
turn, so the follow-up's turn the child's end cut read "Failed" under "Cancellation requested.".
A Stop that ends the provider's session ends whatever is in flight, so its event now names the
live turn or none, and a Stop that names none binds the turn opened next. Codex keeps naming only
the turn the Stop names.

The Claude Stop turn-end tests move to their own file, since the session-ending Stop suite is at
its line budget.

* fix(native-chat): the idle sweep reads working by the same rule as a stop's event

The sweep judged a chat resting while a send whose reply was lost was still unanswered, but the
stop's event writer counts that send as work. So the sweep evicted it and wrote an evict event,
which ends a person's Stop pause and let the cards behind it drain on their own. The sweep's owed
work now reads the main agent working the way every session list and the event writer do.

* test(native-chat): an aborted eviction's injected drain failure lands on the eviction's own drain

A host stop now drains the session's sink once to judge whether it ends work, so the tests that
fail the eviction's drain-published step skip that first drain.

* fix(native-chat): the idle sweep's rest writes no Stop event; it evicts a send that never echoes

The previous commit made the sweep count an unanswered send as owed work, which pins a chat whose
admitted send Codex never echoes forever, and the sweep exists to retire exactly that. That rule
returns. The sweep stops only an agent it judged resting, so its eviction now writes no Stop
event, whatever send it retires: a person's Stop pause holds through it.

* fix(native-chat): stopping a start that carries no send writes no Stop event

A host stop, eviction or close of a starting child wrote a Stop event whatever the start carried.
A start with a send already reads working, so the clause only mattered for a start with none,
which ends no turn and no send: its event only lifted a person's Stop pause and bumped the idle
clock, which is why the idle sweep had been changed to close the conversation in the same pass.
The clause goes and the sweep is #24072's again. The child's end still reads host-stop, as before.

* test(native-chat): a Stop's pause across a restart is tested with a restart that writes no event

The rig's restart closes the chat with an eviction, which now writes a Stop event when work runs
and so ends a person's Stop pause. "A Stop never hides a restart's pause" then passed with no Stop
pause left to hide anything. Those tests, and the pause-lift test whose dropped assertion returns,
restart as a process that dies with no close, which like a quit writes no Stop event, and assert
that both the Stop's and the restart's pauses are in force first.

* fix(native-chat): a host stop of a turn a person's Stop is still ending keeps that Stop's reason

An eviction or host stop that landed while a person's Stop or close was already ending the same
turn wrote a newer Stop event, and the turn's end reads only the latest, so the person's Stop of
that turn read as news. A host reason now writes nothing while a person's Stop still decides what
runs: the live turn it names or bound, or, with none, the turn a send opens next. The person's
own close still writes. The E2 tests now open and end the stopped send's own turn, as Codex does,
so the mail turn after it is not the turnless Stop's.

* fix(native-chat): an older Claude's error on a later turn keeps its error text after a Stop

The translator left an error result that names no reason to the journal's Stop rule whenever a
person's Stop named the turn or none, but the rule binds a Stop naming no turn only to the turn
opened next. So a real error on a later turn read "Failed" with its error text dropped. The
translator now asks the journal's rule itself (`personStopDecidesTurn`, the one core
`turnEndAfterStop` and a host stop's in-force check share), so the two cannot disagree.

* fix(native-chat): a Stop of a start that never landed binds no later turn, whatever sent it

A person's Stop pressed while the agent starts names no turn, and the send it stopped is
cancelled before it opens one. The Stop then bound the next turn anything opened (orchestration
mail, a restart continuation, the queue's drain, all of which send as the host), so a host
eviction of that turn wrote nothing and its crash or close read as the person's cancellation. A
Stop that named no turn now binds only a turn no send journaled after it opened: any send since,
of any origin and not refused, opens its own. The E2 test's mail send is accepted as Codex
accepts it, instead of opening the stopped send's own turn first.

* test(native-chat): a rewind's restated turnless Stop binds no turn opened after the rewind

A Codex rewind restates a person's Stop still in force after the turns it keeps, at a new
sequence, so by sequence alone it would bind the next turn opened after the rewind. A send
journaled after the restated row voids that binding (the previous commit), which this pins.

* fix(native-chat): a relaunch settles a person's stopped turn with no "stopped while in progress" row

After a restart, a turn a person's Stop ended reads "Interrupted after N" with the muted mark, but
the relaunch still added the error row saying the provider stopped mid-response, which a live Stop
never writes. The settle now skips that row when every turn it interrupts is the person's Stop's
by the journal's one rule; a crash nobody stopped keeps it.

* test(native-chat): the unexpected-exit settle's journal fake answers whether a person's Stop decides a turn

* fix(native-chat): a host stop whose sink drain fails reads the journal as it stands

A host stop drains the session's sink before judging whether it ends work, and a failed or slow
drain read as working. So an eviction of an agent at rest wrote a Stop event that ended nothing,
which lifts a person's Stop pause, and a close wrote a person's event naming no turn. The drain is
now best effort: the stop goes ahead either way and only its record is at stake, so a failed or
slow drain leaves the journal's read as it stands. A person's Stop keeps its own rule.

* fix(native-chat): a Stop that named no turn applies only to a turn a send it stopped opened

A person's Stop pressed before any turn showed names no turn. It bound the first turn opened
after it, then (5ead1f6bcc) any turn opened by no later send, so a turn the host started for
a card the Stop held, or for orchestration mail, read as the person's cancellation, and a host
eviction of it wrote no Stop event when its send had been abandoned by the close first.

The rule is now the concept itself: a Stop naming no turn applies to a turn opened by a send it
stopped, one already handed to the agent at the Stop's position. Nothing new is stored. The turn's
row names the send that opened it (Codex: the submission's key; Claude: the echo, which the journal
aliases to the submission), and a handed-over send's item sits at its handover, so the target set
is derived from the journal. A card the Stop held is handed over after it, so it is no target; a
Stop of a start whose send never opens a turn binds nothing; a rewind keeps no submissions, so a
restated Stop binds no turn opened after it. With no turn running, a host stop defers to the
person's Stop only while every unanswered send is one it stopped. Claude's translator, which asks
before its echo row lands, passes the send its echo acknowledged.

* fix(native-chat): a host stop whose sink drain runs long reads the agent working; a failed one reads the journal

A drain past its bound may still hold the turn's row, while the echo's acceptance has already
landed, so the journal as it stands read nothing running: a person's close of that turn wrote no
Stop event and the turn read as news. The two drain outcomes now differ: one that failed has
nothing more to deliver, so the journal's read holds (as before); one still running reads working.

* test(native-chat): a host stop with no turn running defers only while every unanswered send is the Stop's

The branch had no test. An eviction with only the stopped send unanswered writes nothing; one
with a send made after the Stop still unanswered writes its event.

* fix(native-chat): a slow sink drain reads working only while an accepted send's turn row is due

The previous commit read every drain past its bound as working, so a host eviction or stop of an
agent at rest during a sink backlog wrote a Stop event that ended nothing and lifted a person's
Stop pause. A slow drain now reads working only when the latest send the agent accepted has opened
no turn the journal holds, the race it was for; otherwise the journal's read holds.

* test(native-chat): the host-stop control keeps the stopped send unanswered beside the later one

With both unanswered, the host writes only because not every unanswered send is the Stop's; a rule
that deferred when any one was would pass the old control.

* fix(native-chat): a steer is no send owed a turn when a slow drain judges a host stop

A slow drain reads working when the latest accepted send has opened no turn yet. A Codex steer or
a Claude fold is accepted into the running turn and never opens one, so a chat at rest whose last
send was a steer still read working, and an eviction lifted a person's Stop pause. Sends delivered
into a running turn, whose item carries that turn's scope, are skipped.

* test(native-chat): type the Stop test envelope's fields narrowly

* test(native-chat): the retry of a close whose exit was unproven writes no second Stop event

The idle sweep finishes a stop left owed with that stop's own cause. It is the same stop, so its
event stands alone and the child's end keeps the cause, for a person's close and an eviction.
2026-10-01 16:56:13 -07:00
Brennan Benson 1553bc3b80 fix(terminal): reattach a background terminal to its own tab instead of opening a duplicate (#24458)
* fix(terminal): reattach a background terminal to its own tab instead of opening a duplicate

When a workspace with already-running terminals is opened and the renderer has
lost a tab's link to its terminal, the activation gate asks the host who owns
each unlinked terminal. The host's graph only carries mounted or provably live
panes, so an unmounted tab whose link was lost reads as "no surface", and the
gate opened a brand-new tab on the running terminal: the same terminal then
showed in two tabs once the original tab mounted and reattached.

The host now names the pane it last recorded for an orphaned terminal
(`recordedPaneKey`, optional). The renderer, which owns its tabs, rebinds the
terminal there when it still holds that pane free, and opens a tab only when
the pane is gone or holds another terminal.

* test(terminal): pin that an unowned PTY never rebinds to a vanished leaf or another worktree's tab

The rebind to the host-recorded pane relies on two existing guards in the exact-surface binder: the
recorded leaf must still be in the tab's layout, and the tab must belong to this worktree. Neither was
pinned on the unowned path. Both new cases mint a fresh tab and leave the recorded tab untouched.
2026-10-01 16:36:58 -07:00
Neil f69052e113 Reuse qualified Windows server builds and dependency verification records (#24448) 2026-10-01 16:19:40 -07:00
38c2d1dcb9 feat(ssh): update, roll back, recover and stop a managed orcad server (#16741 T6-5 follow-up) (#24463)
* feat(ssh): update, roll back, recover and stop a managed orcad server (#16741 T6-5 follow-up)

Builds managed-server maintenance on T6-2's deploy, rollback and recovery and
T6-4's journaled decommission. Each step reads a terminal census through the
server's tunnel; orcad answers it through a new capability-gated
orcad.terminalCensus RPC, and an older host or lost answer is unverifiable.
Update defers over live or uncounted terminals and status reports the last
deferral. A stop unlinks the server (deployment record, tunnel, SSH claim)
only after a proven exit. Inert until the T6-6 settings UI.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(rpc): load the orcad terminal census lazily so the dispatcher does not pull in the xterm window polyfill

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:53:24 -07:00
Jinwoo HongandClaude Opus 5.5 eefc49f7e3 feat(terminal): warn when a typed Codex joins Codex's shared server (STA-9051) (#24217)
* feat(terminal): warn when a typed Codex joins Codex's shared server

Orca adds --no-daemon to the Codex it launches and to a codex typed in
shells whose wrapper it controls, but a codex typed another way (fish,
cmd.exe, a path-named binary) still joins Codex's shared server, which
mixes up agent status across tabs.

When a local pane's Codex is on that server, show a banner at the top of
the pane with the command that turns auto-start off, a Copy button,
"Don't show again" (a new setting next to the Codex server setting) and
a per-pane dismiss. The banner takes layout space; the terminal refits
below it.

Main answers pty:isCodexOnSharedServer from the pane's outermost Codex
command line (flags and subcommands that keep Codex embedded rule it
out), the CODEX_HOME the pane launched with, and whether that home's
server is live: a socket connect on macOS/Linux, the server's pid record
plus creation time on Windows. The renderer asks only while the pane
already shows Codex, on a short bounded ladder.

Refs STA-9051

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: restore the Claude WSL trust-file fix (#23973) dropped by the banner commit

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(terminal): redesign the Codex shared-server banner and fix dialog

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(terminal): let the Codex shared-server fix run its commands

The fix dialog now runs each step with the shared server's own Codex on the
pane's CODEX_HOME, verifies the result (feature read back, server probed),
and falls back to a copyable command on failure. Stopping asks first.

Also: an apostrophe in a prompt no longer hides an opt-out flag, restored
panes fall back to the saved pty id, and the IPC guards have a table test.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(terminal): give each fix step its own card and label the command it runs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(terminal): simplify the Codex shared-server banner after review

- Read a subcommand only from Codex's first positional, so prompt words
  like "a", "update" or "review" no longer hide the banner.
- Probe the server fresh on every ask; drop the probe cache.
- Make the pty preload methods required and stub them on the web client,
  replacing the optional-method and paired-client checks.
- Render the banner from the existing Codex pane portal loop.
- Treat a non-zero or timed-out Codex command as failed; skip the
  read-back when the disable write failed.
- Reserve the banner's space with a CSS :has() selector instead of a
  data attribute.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(terminal): make the Codex shared-server fix persist on Orca's mirror home

- Step 1 now writes daemon_auto_start = false to the user's own Codex home
  when the pane runs on Orca's shared mirror home (as Windows panes do), then
  to the mirror home too; the mirror is rebuilt from the user's home on every
  launch, so a mirror-only write was lost.
- The server probe is three-state (live / absent / unknown); stop reports
  success only once the server is proven gone.
- The banner retires the one-time "runs Codex without its shared server"
  toast it contradicts.
- A command line with no Codex program never counts as joining the server.
- The fallback local PTY provider reports each pane's root pid.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(codex): keep Turn off in Orca's Codex home when ~/.codex has no config

A pane on Orca's mirror home wrote the setting to ~/.codex first. With no
~/.codex the spawn failed on its cwd and Codex rejects a missing CODEX_HOME;
and creating a config holding only this setting would make the next mirror
replace every setting made in Orca's Codex. The mirror skips a missing or
blank ~/.codex/config.toml, so write only the mirror home then.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(codex): promote [features].daemon_auto_start from Orca's Codex home

Promotion now carries one [features] key alongside the [tui] keys, so a
shared-server Turn off written in Orca's mirror home reaches
~/.codex/config.toml instead of being reverted by the next mirror. Table
keys share one <table>.<key> scan for read, removal and upsert. Orca's own
daemon socket override is never read as a user value, and a blank source
config is seeded from the runtime like a missing one.

* refactor(codex): run Turn off once, in the pane's own Codex home

Settings promotion now carries the setting to ~/.codex, so the separate
settings-home resolution and the two-home loop are gone.

* fix(terminal): offer Stop server only after sharing is turned off

Stopping while sharing is still on closes every sharing session, and the
next Codex starts a new shared server.

* fix(terminal): skip legacy mirror panes off Windows and quoted dotted keys

A retained shared-home pane on macOS/Linux points at a mirror that is no
longer promoted, so Turn off there would be reverted; name no home for it.
A quoted top-level key such as "tui.theme" is one key, not [tui].theme.

* fix(terminal): drop the Turn off note that promised the setting reaches Codex outside Orca

Orca's tabs are what this fix is for; carrying the setting to ~/.codex is best-effort.

* fix(terminal): keep the Turn off note that the setting also applies outside Orca

It holds for nearly everyone; the rare Windows upgrade gaps don't justify hiding it.

* fix(codex): promote Turn off to ~/.codex under an older Orca's baseline

A pane on Orca's Windows mirror home writes daemon_auto_start = false into
the mirror. A promotion baseline from an Orca that predates this key has no
entry for it, so the next mirror pass kept the write as a conflict and then
recorded it, and ~/.codex never got the setting.

Turn off on a mirror-home pane now runs the same mirror pass a terminal
launch runs before and after the write: the first records the key in the
baseline, the second promotes the write to ~/.codex. The passes are
synchronous, so they cannot interleave with a launch's pass. A failed pass is
logged and does not fail Turn off, since the write still fixes Orca's tabs.
Real-home panes are unchanged.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:50:43 -04:00
8b76683b40 feat(ssh): deploy and pair an empty managed orcad server over SSH (#16741 T6-5) (#24453)
* feat(ssh): deploy and pair an empty managed orcad server over SSH (#16741 T6-5)

Adds deploy + pair + status for a managed orcad environment on an empty SSH
host, the loopback tunnel it is reached through (rebuilt on reconnect and
after host resume), SSH provisioning of a new host, and SSH access for an
already paired server. The deployment link lives in the environment sidecar
so a downgraded build cannot strip it. Inert until the T6-6 settings UI.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(ssh): refuse a managed claim while saved state still references the host

Until the T6-8 census exists, a target is claimable only when no workspace
session, automation, worktree metadata or saved PTY lease (any status) points
at it. An unreadable store refuses as unverifiable. Refusals name what
blocked them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(ssh): keep electron out of the resume path; report a decommission journal in status

The caller now passes the profile path for managed-tunnel recovery after host
resume, so ssh-host-sleep-reconnect no longer reads electron's app. Status maps
T6-4's decommission transaction.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:25:33 -07:00
Jinwoo Hong b804c13934 test(runtime): add a readiness census that pins every tui-idle verdict (#24336)
* test(runtime): add a readiness census pinning every tui-idle verdict

Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.

Refs STA-9098

* test(runtime): pin the census quiet probes to literal windows

A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.

Refs STA-9098

* test(runtime): say which census probe writes runtime state

Refs STA-9098

* test(runtime): observe the census through settled panes and caller-visible waits

- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
  of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
  coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
  work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
  directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.

* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix

Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.

* test(runtime): read the census baseline field without Reflect.get

The anti-slop lint rejects Reflect.get on parsed input.
2026-10-01 17:21:48 -04:00
OrcaWinandm4air d3f8c5063b fix(ssh): orcad GC honors the activation journal; readiness requires proven daemon coverage (#16741 T6 follow-up) (#24451)
GC pins every slot an in-flight activation journal names and skips the pass entirely when a journal is unreadable or a fence is held without one. orcad's self-test now reports the coverage the daemon says its probe achieved, and remote readiness probes accept a slot only on pty-spawn coverage, or handshake on win32; builds without the field keep the identity gate.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 13:32:33 -07:00
OrcaWinandm4air 43d9b43d3f feat(ssh): remote orcad stop by request file and journaled decommission (#16741 T6-4) (#24449)
Clients stop an orcad that advertises health.stopRequests through its slot-local request file and keep SIGTERM for older builds. Decommission runs through the activation journal and fence: it refuses while the terminal census is live or uncounted, stops the instance with an instance-bound managed request, cancels a stop orcad never acted on, and deactivates the record only on proven exit. orcad gains --cancel-managed-stop and an exclusive per-transaction decision file so a cancel can never race a dispatched stop. POSIX-only and inert: no production caller.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 13:32:25 -07:00
Brennan Benson fd5804dd6a fix(native-chat): after a Stop whose exit can't be confirmed, the next message retries the stop instead of failing (#24333)
* feat(native-chat): a status for a message waiting on an exit Orca could not verify

Adds the `previousExitUnverifiable` failure fact, a status row only, never a
reason a message was not sent: Orca couldn't confirm the agent's previous
process ended, and the message will send once it has. Copy in every catalog.

* fix(native-chat): a message after a Stop whose exit was unproven retries the stop, then waits

When a Claude Stop could not prove its child gone, the child stayed in place
with a connection that refuses writes, so the next message failed with
"Orca couldn't hand this message to the agent" and every later one did too,
until the idle sweep retried the stop after 30 minutes of quiet.

The delivery loop now retries an owed stop before it starts or writes to a
child, once per new message. Still unproven, the message stays queued under one
warning row saying why, and the idle sweep's next tick retries the stop, child
or not, and hands the message over once the exit is proven.

* fix(native-chat): every operation that reaches the agent finishes an owed stop first

Option changes, card answers, background-task stops, goal changes and rewinds
went to a child a Stop could not prove gone, whose connection takes no input.
The retry now lives in one place, `finishOwedStructuredAgentSessionStop`:
`ensureStructuredAgentSessionAgent` calls it before reporting or starting a
child, and operations on the running child prepare through it. Still unproven,
it refuses with `previousExitUnverifiable` and the exit `unverifiable`; the
delivery loop turns that refusal into the visible wait, so a send still waits
under its one row instead of being refused.

* test(native-chat): type the option-change envelope in the unproven-stop test

* fix(native-chat): the stop that proves the exit hands over the message that waited on it

- The stop itself wakes delivery where its owed wind-down clears, so a message held on an
  unproven exit goes out whichever retry lands: an option change, a background-task stop, the
  sweep or the next message. Only the sweep woke it before, so an option change that proved the
  exit left the message queued with nothing to send it.
- The owed stop keeps where it was asked for, and the child's end is ordered there. A message
  accepted while retries ran is no longer rejected as "chat closed" (a tab close) or failed as a
  host stop when the retry that proves the exit lands after it.
- A wind-down owed by an earlier child never stops a different live child in front of it.

* docs(native-chat): say who retries a wind-down a Stop leaves owed

* fix(native-chat): a waiting message retries the unproven stop once, and its note stays true

- A waiting message holds against the newest pass that failed to prove the exit, kept on the
  owed-stop record (`failedAt`), not against the note's position. The note is written once per
  process, so after a second message every later journal commit re-ran a retry of up to 10 s.
- The note no longer promises this message will send: "Messages wait to be sent until Orca
  confirms it has ended." stays true after a Stop withdraws the message, a close, or a restart
  rejects it. Every catalog follows; es and zh lose the mismatched informal possessive, and the
  French reads naturally.
- Tests: commits after a second waiting message retry nothing; with follow-up queueing on (the
  default) a follow-up becomes a draft, and Steer retries once and waits under the same note.

* fix(native-chat): only a retry continues an owed stop; a new close is a new ask

A second tab close of a chat whose exit stayed unproven kept the first close's position, so a
message sent between the two closes was delivered once a retry proved the exit, even though the
second close closed it (its own rejection of what was queued had failed). Continuation is now
explicit: only the retry of an owed stop passes `retry`; every other stop stamps a new position.

Tests: a second close whose rejection fails still closes the message it closed; the default-
queueing test waits for a draft's commit to settle before asserting it retried nothing, and bounds
Steer's retries; the one-retry-per-message test has the budget to show each commit's retry.

* fix(native-chat): a held message always says why it waits

When another operation's retry of the unproven stop failed between a message's accept and its
delivery step, the step held the message (it had already waited through a failed retry) and
stopped before writing the note, so the message sat Working with no reason shown. The hold now
makes sure of the note too; it is written once per unproven process, so its own commit still
retries nothing.

* fix(native-chat): an option change waits only on an unproven exit, not on bookkeeping

When a stop proved the old process gone but a later step of its wind-down kept failing, an option
change retried that bookkeeping, and refused the change when it failed again. On main the change
was kept at rest. Operations that start no child (option change, card answer, background-task stop)
now hold back only while the exit itself is unproven, its child still on record; with the exit
proven, the bookkeeping retry is reported and the operation goes on as with nothing owed. Sends
still wait: a start needs the released lease that bookkeeping gives back.

* test(native-chat): type the unproven-stop test's envelope fields by the fingerprint's own shape

* test(native-chat): a message after a Codex Stop whose exit was unproven retries that stop, then goes to a fresh Codex
2026-10-01 13:32:02 -07:00
Brennan Benson 976dc00337 fix(native-chat): Stop's pause is worked out from the chat's history, so a steered message is never re-sent (#24072)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused

Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.

* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered

A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.

* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget

* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it

The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.

* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card

* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows

Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.

One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.

The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.

Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.

* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones

A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.

* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction

The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.

* fix(native-chat): stop creating the unused queue pause table

The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.

* fix(native-chat): a Stop's pause never hides the restart pause

A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.

Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.

* test(native-chat): pin the Stop's no-resend, lift and held-card rules

- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
  again" at one instant, before a queue ignoring the pause re-sends. They
  now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
  whether or not a person's turn lifts it; it now reads the Stop's pause
  before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
  queued before a rewind.

* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller

The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.

* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event

* test(native-chat): pin that Stop and Resume rows never reach apps or count as history

* test(native-chat): only a person's Stop event pauses the queue

* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop

Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.

* test(native-chat): a card held at a starting agent is checked before the Stop's timing

Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.

* test(native-chat): a released build keeps and folds a journal holding Stop events

Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.

* style(native-chat): format the Stop event changes

* test(native-chat): type the released build's exports through one checked helper

* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only

* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade

The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.

Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.

* fix(native-chat): a Stop that stops nothing new writes no Stop event

A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.

It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.

* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop

* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled

* fix(native-chat): any later Stop event ends a person's Stop pause

A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.

An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.

* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed

A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.

A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.

Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.

* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes

A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.

The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.

* refactor(native-chat): one reading of a Stop's turn for its event and its note

A Stop's event and its note each worked out the same two facts on their own:
which turn the Stop is about (the one it named, else the one running), and
whether a named turn is the one the journal shows running. The event decides
before the interrupt; the note and whether the session ends decide after the
provider's answer, so those decisions stay separate, but the facts they read
are now one helper each in structured-agent-session-turn-stop-notes.ts:
structuredAgentSessionStoppedTurnId and
structuredAgentSessionStopNamesTurnNotLive. The event's turn, the note's key,
the session-ending condition, the running-command check and the repeat check
all read them. No behavior change.

Tests: a Stop naming no turn records the running turn on its event, and
rewrites that turn's note as a Stop naming it does.

* refactor(native-chat): a failed-interrupt Stop reads its turn through the shared helper

The new branch that ends a Codex child after a failed interrupt asked
whether the Stop's turn still runs with `turnId ?? liveTurnId`, a third
copy of "the turn a Stop is about". It now reads
structuredAgentSessionStoppedTurnId, the value the note key already uses,
read at the same point before the cancel. No behavior change.

Test: a Codex Stop whose interrupt failed ends the child, holds the card
queued before it with the queue paused, and writes its Stop event before
the turn's end.
2026-10-01 13:21:58 -07:00
Jinwoo Hong 477e699922 fix(codex): install Codex's Interrupt hook so an Esc-cancelled turn settles (#24332)
* fix(codex): install Codex's Interrupt hook so an Esc-cancelled turn settles

Codex 0.150+ fires an Interrupt hook when the user presses Esc on an
approval prompt or mid-tool, and nothing else. Orca did not install it, so
the pane stayed blocked/working until the next prompt.

- Add Interrupt to the managed Codex events and label maps, written with
  Codex's 3s cap (a larger value triggers a startup clamp warning).
- Hash the timeout Codex hashes (Interrupt is clamped to [1,3], default 1)
  so self-computed trust matches Codex; pinned against a real 0.159.3 hash.
- Map a root Interrupt to the existing cancelled-turn record
  (markCodexLeadTurnInterrupted), keeping child work in the fold; a
  child-scoped Interrupt is ignored. Relayed rows take the same path.

* refactor(codex): let the hook builder own Codex's per-event timeout

The managed hook's timeout is now Codex's own normalization of the shared
budget, and every installer derives its trust entry from the hook it wrote,
so no installer repeats the Interrupt special case.

Claude-Session: codex-interrupt-hook review

* refactor(codex): route Interrupt through the Stop lead update with an outcome

Interrupt now writes the lead record through the same setCodexMainAgentTurnState
call as Stop, so markCodexLeadTurnInterrupted keeps its original signature.
Drops the child-scoped Interrupt guard: Codex never runs Interrupt hooks for
subagents and its input schema has no agent_id.

Claude-Session: codex-interrupt-hook review

* fix(codex): ignore an Interrupt from an earlier turn once the next turn has started

A replayed or late Interrupt carries the old turn's turn_id; matching it against the turn_id from the
running turn's UserPromptSubmit keeps it from cancelling the new turn. Missing ids still cancel.

* revert(codex): drop the Interrupt turn-id guard; delivery is already ordered

Hook events reach Orca in order: Codex waits for the Interrupt hook before the next turn, and the restart spool is one append-only file per pane, replayed in order before live events. The guard protected an unreachable case and could drop a real cancel if the ids ever differed.
2026-10-01 15:56:01 -04:00
OrcaWinandm4air b093d3ab20 feat(orcad): supervisable server: stop requests, managed stop receipts and a lifetime that keeps its lock on failed teardown (#16741 T6-3) (#24433)
orcad stops through slot-local and instance-bound request files, so a reused PID is never signalled. A managed stop is proven by its completion command and recorded as a receipt. Optional daemon retirement is best effort: an idle daemon retires, while a busy or unverifiable one stays up with its admission fence released. Runtime teardown runs in reverse order and keeps the instance lock and profile admission when any writer fails to stop. Browser discovery no longer delays readiness. Legacy worker recovery and watcher children are drained before the final flush. Headless terminal close no longer waits on a renderer tab that does not exist. No production deployment.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 12:45:34 -07:00
Jinwoo HongandClaude Opus 5.5 ae41eb414a fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup (#24284)
* fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup

A `codex` typed into a plain fish tab ran without --no-daemon because only
wrapped fish tabs (startup command / ready marker) got Orca's codex function.

Plain fish spawns now prepend an Orca data dir to XDG_DATA_DIRS and record the
exact prefix in ORCA_FISH_XDG_DATA_DIRS_PREFIX. Fish sources the dir's
fish/vendor_conf.d snippet, which first restores XDG_DATA_DIRS (unset again if it
was unset), erases the marker, drops its dir from fish's derived vendor/function/
completion paths, then defines the shared fish codex function at the first prompt
so the user's config.fish still wins. fish argv is unchanged; wrapped tabs keep
their existing -C path. A local fallback to another shell restores the user's
XDG_DATA_DIRS instead of deleting it.

Bumps the terminal daemon protocol to v39 so new tabs move to a daemon that
injects the env; v38 owners stay attachable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(fish): skip the XDG handoff for -N/--no-config and empty XDG_DATA_DIRS

fish never reads vendor_conf.d under -N/--no-config (also abbreviated or
clustered), so the snippet could not undo the prefix; and the restore cannot
tell an empty XDG_DATA_DIRS from an unset one. Both now launch untouched.
Run the real-fish handoff tests in the shell contracts job, where fish is
required, so they no longer skip in CI.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(fish): compare the unset-restore case against a fish without Orca

Ubuntu runners ship snapd's fish vendor snippet, which sets XDG_DATA_DIRS on
every fish start, so "unset" was never the right oracle there.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(fish): treat an empty XDG_DATA_DIRS like unset so the tab still gets the codex hook

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(pty): put back the user's own launch env on a shell fallback

The primary shell's launch config now records the pre-launch value of each
key it writes. A fallback shell restores those values (unsetting keys that
had none) instead of deleting the keys, which hands back an inherited
XDG_DATA_DIRS after a fish fallback and an inherited ZDOTDIR after a
zsh->bash fallback, with no per-shell special case.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(fish): drop the Node restore twin and simplify the vendor snippet

- Remove restoreFishXdgDataDirs; the generic fallback restore covers it.
- Snippet: read ":$XDG_DATA_DIRS:" directly and filter Orca's vendor dirs
  with one string match per variable.
- Require inheritedXdgDataDirs in both getShellLaunchConfig option shapes.
- Drop the test-only FISH_XDG_DATA_DIRS_HANDOFF_DAEMON_PROTOCOL_VERSION.
- Fix stale fish comments and trim redundant -N launch cases.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(fish): stop scrubbing fish's lookup paths after the handoff

Only XDG_DATA_DIRS is restored, by exact prefix; Orca's dir holds nothing but this snippet, so leaving it on fish's derived paths loads nothing else and drops the glob match.

* docs(fish): drop the comment for the removed vendor-dir cleanup

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:38:53 -04:00
Jinwoo Hong 07dad6739a refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run (#24443)
* refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run

A same-cap wave no longer consumes a 15-minute monitor dry-run and its sealed,
single-use, five-minute-fresh evidence. Each apply wave now samples fleet health
itself right before isolation, with the monitor's evaluator, thresholds, and
tolerances, for a window sized to the cell's host count (3/5/8 min), plus three
lookback rules: no cell container exit in 10 min, no minute over 500 director
503s in 10 min, and director concurrency p99 within the monitor bar over 4 min.

Removes the monitor-run inputs, the gate's consume/authorize steps, the
break-glass override, and the same-cap-only authorization shapes in
relay-monitor-evidence.mjs. The monitor workflow and the rehome enable path are
unchanged.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): bound the pre-drain sample overrun and keep the drain token fresh

Review follow-ups: alternating tolerated readings could hold the sample open
until its step timeout, so cap the overrun at three samples past the window;
record why a read failed; mint a fresh admin ID token for the drain after the
sample; raise the job timeout to 90 min so a long sample cannot cancel the
job past the failsafe.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* feat(relay): exempt the rolled cell and existing-only cells from the pre-drain crash rule

The exit rule counted every relay container exit fleet-wide, so a cell that
crashes every few hours (c25, 12 a week) blocked the very roll that fixes it,
and existing-only legacy cells (c5, 15 a week) blocked rolls they take no part
in. Exits are now grouped by instance, each instance is named by its own newest
runtime-metrics log line, and only exits on general or migration-only cells
other than the target count. An exit no configured cell can be named for trips
the rule; a failed lookup is a failed read. relay-observability.tf joins the
evidence-code set because the rule depends on its filter.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): cover re-asking for an unnamed exiting instance; note the boot-exit risk

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 15:36:17 -04:00
Jinwoo Hong 051ca4d34e feat(relay): declare US cells c32 and c33 at the 3,000-host shape (#24444)
* feat(relay): declare US cells c32 and c33 at the 3,000-host shape

Declares two us-central1 cells at the Asia shape (cap 3000, 6000 request
units, e2-standard-4) with the US default pool of 10, and generalises the
Asia topology and admission ladder to derive each wave's region from its
reviewed zone, leaving every Asia wave's behaviour unchanged.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* docs(relay): note the US canary tie-break and leave the fleet pool list to promotion

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): plan C32 and C33 as one topology wave

The live-image overlay refuses a declared non-target cell with no template,
so a lone C32 plan would fail on C33. Registration and promotion stay
one cell at a time.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 15:34:57 -04:00
Jinwoo Hong b82307937a fix(relay): admit drained hosts through their own lane and stagger their return (#24446)
* fix(relay): admit drained hosts through their own lane and stagger their return

A same-cap roll's drain sends every host on the isolated cell back to the
director at once. Those hosts reconnect through the sticky lane (one slot per
director), and each one's re-placement holds that slot for most of a second
behind the region-wide inventory lock, so ordinary reconnects time out behind
them and the drained hosts retry every 2 s: ~30k 503s per drain.

The reconnect verification read now also says whether the host's home cell is
isolated for a roll right now (same predicate re-placement uses). Those hosts
release the sticky slot after the read and take a separate drain-return lane
(1 per director, matching the store's per-director placement serialization).
When that lane is full the host gets a Retry-After that reserves the next
free service slot, paced by the measured re-placement time and capped at 300 s.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): keep drain returns inside the placement pool budget and their own slot

Review follow-ups for the drain-return lane:
- The lane now borrows placement permits (never placement's last, never ahead
  of a queued placement), so placement + sticky still bounds the database pool.
- A host's own early retry (row-busy redial, duplicate dial) gets the 2 s lane
  interval instead of a fresh slot behind the cohort, and a host that returns
  early to the same director keeps its reserved slot.
- Classification also excludes an open migration row whose lease counter
  lapsed, matching the re-placement rule.
- The load test now runs five directors behind random routing with a shared
  inventory lock.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 15:34:23 -04:00
OrcaWinandm4air 1a9ac0e955 feat(ssh): crash-safe orcad activation, rollback and recovery (#16741 T6-2) (#24423)
Journal every orcad activation and rollback under a host fence so an interrupted one recovers to exactly the slot the activation record names. D7: planOrcadUpdate and assessOrcadRollback refuse a restart whose incoming build cannot attach the live terminal daemon's protocol. POSIX-only and inert: no production caller.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 12:34:11 -07:00
Brennan Benson 8bf90ce2c2 test(e2e): give the sparse preset proof room to finish on CI (#24442)
The ~60-step screenshot proof takes 2-4 s per step on CI runners and hit the
120 s default on both failing nightlies, at different steps.
2026-10-01 12:18:45 -07:00
Brennan Benson 9912042812 test(e2e): select the onboarding Codex card by its exact name (#24439)
#22720 dropped the command subtitle from onboarding agent cards, so the
Codex card's accessible name is now "Codex" and /^Codex\s/ never matches.
2026-10-01 12:08:07 -07:00
Brennan Benson b3577b7c2a test(e2e): drag the manual-order worktree to a slot that changes the order (#24441)
Smart sort lists the new worktrees newest-first, so dropping the source before
the row that already follows it is a no-op and correctly keeps Smart sort.
2026-10-01 12:07:41 -07:00
Neil 197ea3a3b3 Free PR CI capacity by avoiding repeated setup and real-time test waits (#24355)
* Reduce repeated PR setup and transcript timing waits; add hosted comparisons

* Align parallelism contract with Node-only external rebuild toolchain

* Record hosted coverage and launch package, store, and cancellation comparisons

* Apply hosted Windows setup savings and remove measured test waits

* Keep measured PR package gains and remove completed comparison jobs

* Report measured test counts with precise units
2026-10-01 11:51:43 -07:00
OrcaWinandm4air 99db2bfae4 feat(runtime): SSH access links for paired servers in a downgrade-safe sidecar (#16741 T5-1+T5-2) (#24420)
* feat(runtime): SSH access links for paired servers in a downgrade-safe sidecar (#16741 T5-1+T5-2)

Paired runtime environments gain a durable two-phase SSH access link
(prepare link, verified link, prepare unlink, complete unlink, cancel),
plus the reconciliation record type and the runtime identity verification
helper that T6 needs. Nothing calls the link store yet; T6's managed tunnel
and runtime SSH access are its first writers.

Why a sidecar: v1.4.217 and v1.4.218 parse orca-environments.json with plain
z.object schemas, which strip unknown keys, and rewrite the whole file on
routine use (markEnvironmentUsed). Storing the link there, as #16741 did, would
let a downgraded build keep the tunnel endpoint but drop sshAccess, stranding
the server unlinkable and losing its pinned host-key fingerprint. #16741's own
answer (bumping the store version) makes those builds reject the whole file.

So orca-environments.json keeps exactly its shipped shape (version 1, persisted
fields only), and all T5 state lives in orca-environment-sidecar.json beside
it: the link and its tunnel endpoint, the pending operation, the reconciliation
record, the verified runtime id and a monotonic pairing-revision floor. Reads
overlay the sidecar; each entry is bound to the environment's createdAt,
pairing revision and preferred endpoint, so a re-pair, removal or edit by an
older build makes it stale (ignored, pruned on the next sidecar write). An
unreadable sidecar fails closed, like the main file.

Porting note (source: #16741 a68b6f3531):
- Taken: the link-store behavior and messages, the access-link schemas and
  refinements, the reconciliation record, identity verification, and the
  store/schema tests.
- Adapted: link state moved from orca-environments.json to the sidecar;
  existing mutators write persisted fields only; removal, re-pairing and
  runtime identity changes refuse while SSH access is linked or pending.
- Left for later slices: orcadDeployment and restoreManagedOrcadEnvironmentLink
  (T6), reconciliation store, integrity, catalog and UI (T5-3 to T5-5), the 24
  renderer terminal-input files (T7), runtime-identity and the managed-tunnel
  resolver (T6).

* test(runtime): read SSH access from the known view of a persisted environment

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 11:45:59 -07:00
github-actions[bot] 198fe72066 Update README downloads badge 2026-10-01 18:35:36 +00:00
Jinwoo Hong 9bcdb6ad86 fix(native-chat): report a failed Stop child end through the host logger (#24437)
#24334 reported it through onEventSinkError, which #24312 replaced with the
required diagnostics logger; main did not typecheck after both merged.
2026-10-01 14:35:01 -04:00
Jinwoo Hong 56c7642aae test(orcad): skip the live-terminal runtime hand-over across a protocol bump (#24429)
* test(orcad): skip the Bun-to-Node live-terminal hand-over across a protocol bump

The last Bun orcad's daemon reports protocol 38 forever, so asserting the
adopted daemon matches this checkout's PROTOCOL_VERSION failed every bump.
Ask the Bun slot's daemon for its protocol once, run the hand-over when it
matches, and skip with the two versions named when it does not: a daemon
at another protocol is never adopted across an update.

* test(orcad): clean up the Bun protocol probe even when its launch fails

The probe's cleanup ran only after a successful launch, so a launch that
timed out or threw left its orcad and daemon running. One finally now
stops the orcad, kills what it launched, and kills any daemon named by a
pid file in the probe's data root.
2026-10-01 14:28:05 -04:00
Brennan Benson beccdec74c fix(native-chat): a Codex Stop that Codex refuses or never answers ends the Codex process (#24334)
* fix(codex): a Stop whose interrupt Codex refused or never answered ends the session

A Codex Stop is the interrupt alone, so its background terminals keep running. When Codex refused
the interrupt, or never answered it, the turn kept running and the chat said "Codex didn't stop"
or "Cancellation was not confirmed." with no way to stop it short of closing the chat.

Now, after an interrupt that failed, the Stop re-reads the conversation once Codex's frames already
received have landed, and ends the child through the host's usual stop (proof of exit, then the
lease release) when it still runs what the Stop was sent for. It skips a conversation at rest, a
child whose exit already ended its turn, and one Codex has moved on to a different turn. The turn
then reads as the user's cancellation. If the child's exit is not proven, the failure is reported
and the Stop keeps its "didn't stop" or "not confirmed" row, which is still true.

A Stop Codex took keeps the child, unchanged.

* fix(codex): end the child only for an interrupt whose effect is unknown, on the turn the Stop meant

- Codex's invalid-request refusal (-32600: no active turn, another turn active, thread not
  loaded) states the named turn is not running, and can arrive before that turn's end frame.
  The Codex adapter now marks it `turnNotRunning`, and such a refusal never ends the child.
  Only an internal-error refusal (-32603), an unanswered interrupt, or a thrown cancel does.
- A Stop that named a turn, or an unnamed Stop that read one, ends the child only when the
  journal, after draining received frames, still shows that same turn. A session working on
  a follow-up whose turn has not opened is no longer ended by a Stop of the finished turn.
- When the child's exit was proven and only a later cleanup step failed, the Stop reads as
  requested; the failure is still reported.
- `AgentSessionCancelOutcome` moves beside the adapter's other Stop members (re-exported), which
  keeps the adapter file within its line budget.

* test(codex): pin that a failed interrupt decides on the drained journal

The frames Codex sent before the interrupt failed are held in the sink, so a read that skips the drain sees the stopped turn still running.

* test(codex): a refusal naming a turn Codex is not running carries turnNotRunning

* test(native-chat): fail the route release synchronously, as the adapter's acknowledgement is

* docs(codex): a -32600 Codex could not parse also reads as not running and keeps the child

* fix(native-chat): a named Stop whose child end is unproven says Codex didn't stop, not that the turn finished

The new branch has just read the turn running, so the named Stop's 'already finished' row was false.
2026-10-01 11:27:53 -07:00
OrcaWinandm4air 34a582bd39 feat(relay): capability-gated owner reset with a durable preparation journal (#16741 T3 R1) (#24418)
The relay gains an owner-reset surface that no current client calls. A
session-owner client will be able to ask its relay to prepare a shutdown
(`relay.reset`), have that preparation journaled durably beside the relay
endpoint, and later recover it (`relay.recoverPreparedReset`, or the read-only
`--read-reset-preparation` exec mode if the daemon is gone). Relay status
advertises `relay.ownerReset.v1` and, when a journal is configured,
`relay.durableResetPreparation.v1`.

Why ahead of its caller: relays and clients update independently, so the
host side has to be deployed before any client can rely on it. Old relays
answer method-not-found and advertise neither capability, which a client
reads as "unavailable", never as "reset".

- relay-grace-lifecycle: shutdown is split into prepareShutdown and
  finishShutdown so a reset can settle its response before the process exits.
  Idle and signal shutdown keep today's behavior, including the deferred
  retry when disposal fails, and still do not wait on admitted requests; only
  a reset initiator drains work around owned-process disposal.
- relay-work-drain-contract: both reset requests are admitted during a drain.
- relay.ts / relay-daemon.ts: register the reset, its journal under
  `<endpointDir>/owner-reset-preparations`, the status capabilities, and the
  reader argv (checked after the self-test and Windows breakaway launch).

Porting note (source: #16741 a68b6f3531, merge-base 277c289bd4):
- Based on #24414, which already adds the adapter's activeSessionOwner and
  assertOwnerPublicationSettled with code identical to #16741's.
- Dropped: network-tunnel fencing (T4), the PTY ownership-transfer fence and
  its refusal test case (T7), the Bun arm of the reader integration test,
  the Bun runtimeVersion status field, and #16741's lazy entry imports.
- Adapted: wire parsers use a record guard instead of casts; the idle path
  no longer gains an unbounded work drain.
- Added: relay-grace-lifecycle tests (idle exit, deferred retry, attached
  client, reset preparation, joining and retry) and a wire-contract test
  pinning the capability, method and flag names.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 11:25:37 -07:00
OrcaWinandm4air d53063d2b1 feat(ssh): track connection-manager drains, test probes and provider continuations (#16741 T2 P3+P8a) (#24407)
* feat(ssh): track connection-manager drains, test probes and provider continuations (#16741 T2 P3+P8a)

The SSH connection manager now registers each connection it allocates against
the connection's transport-closure notice, and tracks every connect,
disconnect, reconnect and teardown per target, so later slices can tell when a
target's local work and transports are really gone. Ordinary connect,
disconnect and quit behavior is unchanged apart from cleanup.

Manager (ssh-connection-manager.ts):
- registerConnection counts every connection a target allocated, retired pool
  entries included, until its transport-closure notice arrives.
- connect/disconnect/reconnect/disconnectConnection/disconnectAll run through
  trackTargetOperation; a failed teardown is remembered as unconfirmed.
- A failed startup is now disconnected, not just dropped from the pool.
- disconnectAndDrain(targetId, signal) and disconnectConnection(..., drain),
  bounded by a 10s timeout.
- disconnectAll(shouldDisconnect) only tears down targets the filter allows.

IPC:
- ssh-connect-attempt-registry: runSshTestConnectionProbe publishes a probe per
  target before it starts and clears testingTargets / credential flags only
  when the target's last probe settles. ssh:testConnection uses it, so quit
  still joins an in-flight probe within the shutdown budget.
- ssh-target-lifecycle-queue: runTargetLifecycle returns the operation's value.
- ssh-shutdown-drain: a mayDetach predicate (default: every target) is plumbed
  through detach, invalidation and disconnectAll.
- ssh-renderer-broadcast: targets owned by a runtime are hidden by owner too,
  not only by id prefix.
- ssh-target-registry: direct-authority resolver, installed by
  ssh-active-relay-sessions.
- Provider dispatch: unregister*IfCurrent for git and filesystem providers.

P8a, SSH provider continuations: ssh-provider-continuations tracks local
settlement of SSH filesystem writes/deletes, imports, detected-worktree
listings and worktree/folder removals per target. It records only local
settlement, never remote execution or exit.

Porting note (source: #16741 head a68b6f3531, merge-base 277c289bd4):
- Left for later slices: connectExclusive and assertTargetTransportsClosed
  (T8), hasTargetActivity (T3), assertProfileLifetimeAdmission (P8b, stubbed
  out of the manager and continuations), the reset admission checks in
  ssh-connection-handlers and the reset predicate in the shutdown drain (T3),
  the network-tunnel resolver (T4), and the managed-orcad provisioning test
  hunks in ssh-target-registry.test.ts (T6).
- Tests read the manager's bookkeeping through
  ssh-connection-manager-test-probes until T3/T8 expose readers. Two of the
  four deferred disconnect-drain cases are restored; the two that need
  connectExclusive, two closure cases that need it and three profile-admission
  cases stay deferred.
- The removal and detected-worktree call sites were rebased onto main's
  background removal and execution-host routing. filesystem-import-ssh moved
  its remote-existence helpers out to stay under 300 lines.

* test(ssh): name the filesystem continuation handler payload type

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 11:25:29 -07:00
Zun 35c8887a76 fix(i18n): correct Korean working and shell labels (#24341) 2026-10-01 14:22:05 -04:00
Brennan Benson c6cfcc034e refactor(native-chat): structured chat failures always reach the diagnostics log (#24312)
* refactor(native-chat): give the structured chat host one required logger

The structured chat runtime took an optional onError callback that the
desktop never passed, so a late dispatch settlement, an unanswered-dispatch
release, a journal event-sink write and a provider lifecycle delivery that
failed were dropped with no trace. Other host failures went to scattered
console.warn calls, which reach nothing in a packaged desktop build.

The runtime and host now take one required logger (warn/error with a scope
and fields). The production logger writes each entry as a failed span to
<userData>/logs/main.trace.ndjson, which the diagnostic bundle collects, and
to the console (stderr under a supervised headless host). The runtime and the
host wrap it so a logger that throws never fails what it reports, and the
install refuses without one. Sites that deliberately kept a recovery-capsule
error out of the log still log no error object.

* refactor(native-chat): hand the chat host's collaborators the logger, and give orcad its trace file

The delivery loop, idle sweep, queued-message drain, lease renewer, event
sink, conversation map and provider start/exit settlement each took an
internal error callback that the host mapped onto the logger. They now take
the logger itself and log under their own scope. The event sink keeps one
onFailed hook, which decides whether to stop the provider, not whether to
report. The dead-generation settlement returns its failure so each caller
logs it under its own scope.

orcad now installs the desktop's local trace sink under its own data root, so
a headless host's chat failures reach <data-root>/logs/main.trace.ndjson as
well as stderr.

Also passes the logger in the test fixtures the first commit missed, which
tc:node caught.

* fix(native-chat): keep repeated chat failures from flooding the trace file, and record their causes

- The production structured-chat logger writes a repeated failure (same level, scope, session,
  message and error text) once per 5 minutes, carrying how many repeats it swallowed; the
  tracked set is capped at 256.
- Trace entries now carry the error's code (and SQLite errcode) and up to three causes by name and
  message.
- A chat read whose conversation will not open is logged through the host's logger
  (open-for-read), and so are the runtime's chat-tab bookkeeping failures that already hold the
  host.
- orcad writes its own orcad.trace.ndjson, closes it after every quit handler, and flushes it on
  process exit; a trace file that cannot be opened leaves tracing off instead of stopping the app
  or orcad.
- Tests: the desktop wiring test proves the logger reaches the trace sink, and the privacy tests
  read every level the logger received.

* fix(native-chat): log a created chat's tab-publication and launch-prompt failures through the host's logger

* fix(native-chat): key a repeated chat failure on everything its entry writes

The repeat suppression keyed on the message and the error's text, so two refusals with the same
code but different causes, a plain error and a refusal of one code, or two object-valued errors
shared a key and the second was swallowed for five minutes. The key is now the entry's whole
written content (fields, code, errcode, refusal reason, cause chain, a stable rendering of a
non-error value) plus the error's name and message; a refusal's reason is also written.

* test(native-chat): pin that an error's name keeps two repeated failures apart

* test(native-chat): build the refusal in the repeat-key test as the wire does

* fix(native-chat): read an error's code and a refusal's reason by narrowing, not Reflect.get
2026-10-01 11:21:50 -07:00
OrcaWinandm4air 3fbdaba262 feat(orcad): migration manifest and dormant-state contracts (#16741 T6-7) (#24422)
Adds the shared contracts for converting a desktop catalog into a dormant
managed-server catalog: the versioned migration manifest, import receipts,
dormant worktree/workspace/automation/session/client-state validation,
scrollback snapshot descriptors, migration preflight categories, catalog
state and staged-catalog normalization. Nothing calls them yet.

- The manifest is a cross-version artifact: it carries version 1 and any
  other version is refused (orcad_migration_manifest_version_unsupported).
- Workspace references accept `folder:` keys alongside `worktree:` keys and
  bare worktree ids; each must belong to the manifest's own catalog.
- Every list and payload is bounded: the manifest at 768 KiB, scrollback at
  512 snapshots, each within the terminal store limit, 256 MiB in total and
  fixed-size chunks.
- Workspace sessions are validated with main's own parseWorkspaceSession, so
  the contract tracks the current session shape instead of a frozen copy.

Porting note (source: #16741 a68b6f3531):
- Adapted: production parsers no longer cast. Each structuredClone(...) as T
  became a type-guard predicate that runs the same field checks; host-id
  arrays narrow through isWorkspaceHostId, and manual repo order now requires
  a real workspace host id, as its type always claimed.
- Split to stay under 300 lines: manifest field helpers and the automation
  run field checks moved into their own modules.
- Excluded: terminal publications and live terminal bindings (T7/T8) and the
  source-cutover contract (T8).

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 11:09:19 -07:00
OrcaWinandm4air ece9e4d2e3 fix(runtime): fence runtime-environment subscriptions and status probes by identity (#16741 T5-3) (#24421)
Splits runtime-environment subscriptions and the status probe out of their
IPC modules and fixes the reconnect and teardown races the split exposed.

Subscriptions (runtime-environment-subscriptions):
- A subscription id is reserved while its socket opens, so a duplicate request
  for the same id is refused instead of racing the first.
- Every callback checks it still owns its id by token, so a late event or close
  from a superseded socket cannot reach a subscription that reused the id.
- A close that arrives during setup closes the new socket and fails the call
  instead of registering a subscription that is already dead.

Status (runtime-environment-status-probe, status-owner):
- A status request records the capability incarnation it started under; if the
  environment was re-paired or retired meanwhile, it answers
  runtime_environment_changed instead of publishing stale status, diagnostics or
  a runtime id.
- runtimeEnvironmentChangedFailure moves to the revision guard and
  shouldUseSharedControlEnvelope to support routing, so both have one home.

Status and transport routing still resolve environments with plain
resolveEnvironment; the managed-tunnel hook is T6.

Porting note (source: #16741 a68b6f3531):
- Left for T5-4 (reconciliation): the control/resource transport-generation
  scopes and preserveResourceStreams, retireRuntimeEnvironmentControlTransport,
  and their test cases.
- Left for T6/T8: the managed-tunnel resolution, the orcadDeployment
  disconnect guard and connectivity forget hook, and the managed, migration and
  ssh-access handler channels and registrations.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 11:09:12 -07:00
OrcaWinandm4air dd87ae578d feat(ssh): remote orcad primitives on the pinned Node runtime (#16741 T6-1) (#24419)
- orcad-remote-runtime-control: the one launch-and-poll-readiness loop that
  deployOrcad and rollbackOrcad now share, plus exec and recovery helpers.
- orcad-remote-record-file: bounded, marker-delimited reads and atomic writes
  for host records. A read that returns no verifiable answer rejects instead
  of reading as "no record".
- The activation record store reads through it (a lost read no longer becomes
  an empty record) and writeOrcadActivationRecord refuses to replace a record
  this client cannot read, such as a newer schema; deploy and rollback use it.
- orcad-active-readiness: prove a recorded-active or relaunched slot against
  the activation gate, reporting exited / unverifiable / rejected.
- orcad-remote-build-hash and orcad-remote-context (host, home, server target
  via orcad-deployment-target, activation record; POSIX-only).
- Readiness reads are capped at 256 KiB, and a finished invalid JSON line is
  malformed rather than pending forever.

Inert: nothing in the app calls managed orcad deploy yet. Ported from #16741
(a68b6f3531) onto main's pinned-Node deploy core; the branch's Bun runtime,
slot eligibility, target detection and bundle installation are superseded by
main's orcad-remote-runtime, orcad-deployment-target and orcad-remote-install.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 11:09:04 -07:00
OrcaWinandm4air 92cb71765e feat(ssh): add pty.resumeClient and split SSH PTY process listing (#16741 T2 P5+P6) (#24414)
* feat(ssh): add pty.resumeClient and split SSH PTY process listing (#16741 T2 P5+P6)

P6: a relay now answers pty.resumeClient, which admits only an exact resume of
the existing session owner (it never mints a fresh claim when that owner is
gone). resumeSshPtyConsumerSession calls it with cancellation and authority
checks; an old relay's method-not-found becomes a pty_consumer_resume_unsupported
refusal that leaves the channel usable for pty.openClient. Owner grant
publication now rolls back if the response-settlement hook cannot be armed, and
the adapter exposes read-only owner and publication-settled queries.

P5: SshPtyProvider.listProcesses moves to ssh-pty-process-list unchanged, and
the notification-routing tests split into a shared fixture plus recovery and
recovery-activation files.

Nothing calls pty.resumeClient yet (T6). Ported by hunk from #16741
(a68b6f3531) without ownership-transfer (T7) or Bun runtime hunks.

* refactor(ssh): type the consumer-session transport as the members it uses

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 10:39:38 -07:00
OrcaWinandm4air ff212dbbef feat(daemon): idle retirement, session census and recovery-only provider (#16741 T2 P4b) (#24409)
* feat(daemon): idle retirement, session census and recovery-only provider (#16741 T2 P4b)

- DaemonPtyRouter routes idle retirement through DaemonRouterRetirement: it
  fences new spawns, counts spawns already in flight, takes a census of every
  daemon generation and retires them only when all are idle. A lost reply
  keeps the fence; an unanswered census is unverifiable, never empty.
- Each adapter answers requestIdleRetirement through shutdownIfIdle and fences
  its own spawns while retiring.
- A recovery-only adapter/provider (createDaemonRecoveryProvider) reattaches
  and controls existing sessions but admits no new process, never respawns
  the daemon, and never prunes sessions of unknown worktrees.
- listLiveDaemonSessions and requestIdleDaemonRetirement report null /
  'unverifiable' when any generation cannot answer.
- ptySpawnHealth replies add optional coverage and Node runtime fields;
  checkDaemonHealthWithCoverage reads them and treats an absent coverage from
  an older Windows daemon as handshake-only.
- reconcileOnStartup moves to daemon-router-session-reconciliation unchanged.

No caller uses retirement, the census or the recovery provider yet (T6), so
they are inert. Ported by hunk from #16741 (a68b6f3531), without the PTY
ownership-transfer input fence (T7) and with a Node-only runtime report.

* test(daemon): narrow parsed frames and endpoint errors instead of asserting types

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 10:39:31 -07:00
OrcaWinandm4air d23ecef301 feat(session): retry failed renderer session writes and verify local folder PTYs (#16741 T2 P9) (#24406)
* feat(session): retry failed renderer session writes and verify local folder PTYs (#16741 T2 P9)

The renderer's session write subscriber now waits for the local session patch
to be accepted by main. While a write is in flight no second write starts; if
it rejects, the fields it carried go back into the pending set (newer edits to
the same fields are not overwritten) and the next store or gate wake retries
them instead of silently dropping them.

Boot PTY hydration verifies a folder workspace as local only when its id is
unique and it is pinned local, or its legacy scope resolves local with no
remote candidate repo. A folder under an SSH project group is never verified.

Ported by hunk from #16741 (a68b6f3531).

* test(memory): build typed folder-workspace fixtures instead of asserting types

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 10:39:24 -07:00