mirror of
https://github.com/stablyai/orca.git
synced 2026-09-30 16:02:56 +00:00
7afa4ee3dc69db4e7df6f8669c503261bf76c50e
1367
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
444e0b1cf9 |
fix(codex): recognise Codex's quoted spellings in config.toml, and repair Orca's duplicates (#22592) (#23958)
* fix(codex): recognise Codex's quoted project-trust spellings in config.toml (#22592) Codex's settings screen writes project trust as ["projects"."/p"] and "trust_level" = "trusted". Orca's matchers only knew the bare spelling, so a trust write appended a second [projects."/p"] table (or a second trust_level line) and every codex command then failed with "duplicate key". The config mirror kept both spellings in Orca-managed homes for the same reason. - Project table headers are now read through the existing TOML key-path parser, so bare, quoted, literal-quoted, mixed and spaced spellings are the same table for trust writes and the managed-home mirror/dedupe. - trust_level is found by decoded key, in both the trust writer and the mirror's trust reader, and an existing key is rewritten, never duplicated. - On the next trust write, a table older Orca appended (exactly [projects."<p>"] holding only trust_level = "trusted") that duplicates the user's table, or the bare line it inserted under a quoted "trust_level", is removed; the user's table wins and the atomic writer keeps config.toml.bak. Any other duplicate, or a repair that would still leave one, leaves the file untouched and logs once. * build(cli): list the new Codex trust modules in the CLI project * fix(codex): recognise Codex's quoted hooks.state spellings and repair Orca's copies (#22592) Codex writes hook trust as ["hooks"."state"."<key>"] (and the parent as ["hooks"."state"]). Orca's hook-trust writer, parent-table check and mirror only knew the bare spelling, so a hook-trust write appended a bare copy and the file failed to parse with "Cannot declare ... twice". - The hooks.state header, parent-table and mirror checks now use the TOML key-path parser, like project tables. - The duplicate repair now also removes Orca's own hooks.state tables (an exact [hooks.state."<k>"] with only enabled + trusted_hash, or an empty [hooks.state]) that repeat a table in another spelling, and runs on hook trust writes too, so a file with both project and hook duplicates is fully repaired. The Orca-shaped copy is removed whichever order the two tables are in, only when exactly one other table (the user's) remains; anything else is left untouched and logged once. * fix(codex): carry plain-Codex plugin and project hook trust into Orca's Codex homes (#22592) Codex keeps hook trust in $CODEX_HOME/config.toml under hooks.state, keyed by the hook's source. Plugin keys (`id@mkt:path`) and project keys (`<repo>/.codex/...`) are the same in every home, but the mirror dropped every hooks.state table from ~/.codex, so Codex inside Orca asked users to re-trust plugin and project hooks they had already trusted in plain Codex. - classifyHookTrustKey splits keys into home-scoped (the home's own hooks.json/config.toml, re-keyed by install as before) and shared. - The mirror now carries shared hook trust from ~/.codex in every spelling. A key the managed home already holds keeps the managed copy, a key repeated in ~/.codex is carried once, and the parent [hooks.state] table is never copied, so the result never declares a table twice. - mergeSystemCodexConfigIntoRuntime moves to codex-config-mirror-merge.ts to keep codex-config-mirror.ts under the line limit. - Tests cover plugin/project carry in each spelling, user-hook keys staying out, repeated launches, managed-copy precedence, Windows key spellings, parent tables, and user-hook trust re-keying (trusted_hash and enabled) from every ~/.codex spelling. * fix(codex): carry session_end and interrupt hook trust into Orca's Codex homes (#22592) The shared-trust classifier parsed hook keys with Orca's own trust-key parser, which only knows the ten events Orca installs hooks for. Keys for Codex's session_end and interrupt events did not parse, so their plugin and project trust was treated as home-scoped and left out of the managed home. The classifier now reads the source path from Codex's key shape `{source}:{event}:{group}:{handler}` for any event label. A key without that shape is still never carried. Tests cover both events for plugin and project keys in both spellings, user-layer keys for both events, and five unattributable key shapes. |
||
|
|
5b3366f78e |
test: unit tests can no longer write a developer's real agent or Orca settings (#23979)
* fix(agent-trust): write per-user trust under the home the launched agent reads The Cursor, Copilot, Qoder and Antigravity writers and the local Codex config list resolved ~ with os.homedir() at write time, so any test that reached them wrote into the developer's real ~/.codex, ~/.cursor, ~/.copilot or ~/.gemini. Each writer now takes the home, derived once from the launch env (HOME, or USERPROFILE on Windows, else this host's home) by launchedAgentHome, which the SSH relay already used. * test: give tests that wrote the real agent or Orca home a temp one The structured Codex adoption replay pre-trusted /repos/workspace-1 in the real ~/.codex/config.toml; it now runs with a temp HOME and userData. The Codex session-resume and WSL hook tests created Orca's managed Codex home under the live userData, and the Claude Agent Teams tests wrote their tmux shim into ~/.orca; each now runs against a temp userData or HOME. * test: fail any unit test that writes the real agent or Orca home A vitest setup file wraps the node:fs mutating calls and refuses a target under the account's real ~/.codex, ~/.claude(.json), ~/.orca, ~/.cursor, ~/.copilot, ~/.gemini, ~/.qoder or Orca userData, found through os.userInfo() so a test that swaps HOME cannot hide it. The refusal is recorded and rethrown after the test, since trust writers swallow errors. Reads are untouched. It stands down only while an opted-in real-agent suite's own switch is set. It also unsets what an Orca terminal exports toward the live app (userData, Codex and Claude homes, and the Codex launch preflight CLI, which a shell test would otherwise run), so a local run matches CI. * test: type the guarded fs call from its narrowed original |
||
|
|
d414033400 |
fix(packaging): stop shipping the relay bundles inside app.asar (#24027)
resources/relay is the only relay copy a packaged build resolves, but out/relay was also packed into app.asar — 14.2MB of unreachable duplicate. Kaspersky flagged app.asar as a compound object precisely because relay.js was inside it, so one script-heuristic verdict on relay.js gutted the whole install. Excluding it decouples app.asar from that verdict and drops the duplicate bytes. |
||
|
|
e594cb06af |
test(mobile): record RPC goldens without a pinned commit, and check recorded requests against the desktop's params rules (#23732)
* test(mobile): add rpc:diff to decode what a recording change moved
The RPC recording goldens are content-addressed JSON, so their raw git diff is
pool hashes. `pnpm --dir mobile rpc:diff [<base>]` decodes both sides and prints,
per golden, the checkpoint, field and JSON path that moved with both values,
grouped across checkpoints, plus added and removed goldens. `--summary <file>`
appends a Markdown report capped for GitHub's step-summary limit.
It reads any pooled format, so it can prove the next commit's format change
moves no recorded value. Checkpoints are matched by occurrence because an id can
repeat within one golden.
This commit adds files under the recorder directory, which moves the header
digest every golden pins; the next commit removes that header.
* test(mobile): record RPC goldens without a pinned commit or input digests
Every golden carried a pinned `baseline` commit plus digests of the recorder,
its mount adapter and its scenario, and the record script refused to run unless
the product tree matched the pin. So every behaviour change repinned to its own
branch commit and rewrote all ~790 files, the squash made that commit
unreachable, and main's pin job stayed red until a hand-made repin pull request
landed (22 of them in 12 days). The digests could only fail when an input moved
and the recording did not, which is exactly the change that carries no
information; every run already re-derives each golden from the current tree and
compares it.
Format 6 keeps the format version, operation, family, named deltas, the value
pool and the recording. Removed: the pin and fence, the three digest modules and
their test, the pin guard and its CI job, and the dead scenario `version` field
(the manifest reader now refuses `baseline` and `version` with a message).
- `pnpm --dir mobile rpc:record [<golden-id>...] [--prune]` records all or some
goldens; orphans are listed, and deleted only with `--prune`. Every derived
test title now starts with its golden id so an id selects it.
- `compareGolden` reports every difference in one failure (identity fields by
name, the checkpoint list, each checkpoint/field/path grouped), keeps the
final byte compare, and ends with the command to re-record that golden.
- `unhandled-recording.test.ts` now drives a detached rejection through
`runRecording` into a checkpoint and the cleanup checkpoint; no golden carries
one, and disconnecting the capture passed every suite before.
- Seam rules that existed only to keep a digest honest are gone; the
mutant-reachability, register-completeness and one-exposure rules stay.
- CI: `Mobile tests on main` runs the whole mobile suite on every merge that
touches mobile/, src/shared/, the root lockfile or the host RPC paths, since
`verify` never runs on main. A new `Mobile RPC Recording Replay` workflow
replays the recordings on pull requests that touch src/shared/ or the root
lockfile without touching mobile/. `verify` writes the `rpc:diff` report to
the job summary.
Proof: `rpc:diff` against the parent reports no recorded behaviour moved; each
golden only loses its ten header lines.
* test(mobile): check every recorded request against the host's params contract
The goldens script the host's replies, so a scenario could record a success
for a request the real host would refuse, and a desktop change that tightens a
params schema moved no golden at all.
`recorded-request-params.test.ts` parses every distinct request the corpus puts
on the wire with the host dispatcher's own `parseRpcRequestParams` and the
schema `rpc-params-catalog.generated.ts` binds to that method. It fails on a
method the host lacks, params it refuses, params sent to a method that takes
none (the dispatcher never reads them), and keys the schema silently strips
unless an inventory entry gives the reason; a stale entry fails too. Each rule
is also shown firing on a made-up request, since the corpus has no instance of
three of them. It imports the desktop dispatcher, so it sits beside the other
Node-side tests outside the RN test program, and the params-contract boundary
now exempts test files, which are never bundled.
It found twelve requests the host would refuse, all from invented fixture
values, not product code, fixed at their source:
- git.branchDiff sent `base-oid`/`head-oid`/`merge-base` where the host needs
full object ids (diff-review and source-control adapters, and the branch
compare replies in the manifest that feed them);
- an iOS push registration without `apnsEnvironment`, which a real iOS token
always carries (`push-token.ts`); the adapter now defaults to `production`;
- `settings.update` given Linear's `assigned` filter as a GitHub preset, which
the product type forbids; the scenario now picks `my-issues`;
- GitLab `projectRef` as a string where the host and the product type take
`{ host, path }` (7 methods, 5 adapters and the manifest).
46 goldens move, and a decoded comparison of every one of them shows no change
other than those substitutions; `rpc:diff` lists them.
* ci(mobile): detect a mobile change without a SIGPIPE-prone grep pipe
Under the runner's pipefail, grep -q exiting on its first match SIGPIPEs git
diff on a long file list, so a large pull request touching mobile/ read as
uncovered and replayed the recordings a second time.
* test(mobile): drop comments that still describe the golden header and digests
Eleven adapters justified an import rule by the header a golden no longer
carries, and that rule's test is gone. The census failure now names the
rpc:record and --prune commands.
* test(mobile): refuse a golden that keeps a key no recording writes
Decoding dropped unknown top-level keys, so an old header left behind by a
hand-resolved merge conflict passed every compare unseen.
* ci(mobile): summarize RPC recording changes after a failed test step too
* test(mobile): stream rpc:record output instead of capturing it
A captured run stayed silent for its whole duration and clipped its tail,
where the failure summary sits, past 8 MB.
* test(ci): let the Ruby-gate contract skip the always-run RPC summary step
|
||
|
|
fb52c0602a |
fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout (#23920)
* fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout
xterm paints nothing while DEC mode 2026 (synchronized output) is open and only
force-flushes after 1000ms. Codex wraps every draw in mode 2026, so any byte gap
or chunk split that loses the closing \x1b[?2026l freezes the pane for a full
second and then repaints in one burst.
Orca never emitted \x1b[?2026l anywhere, and three paths could destroy a TUI's:
the per-PTY pending cap drops buffered output wholesale (mode 2031 was already
salvaged there, 2026 was not), main sliced pending data at a blind 16KB offset
that can land inside an open frame or sever the 8-byte marker, and the renderer's
backlog warnings replace a queued tail that may hold the close.
- salvage the 2026 latch across dropped output, mirroring the existing 2031
salvage, and append the release on both delivery sites
- ground 2026 in RESET_AFTER_BYTE_GAP and the replay baseline, and in both
backlog warnings, so every drop path is self-healing
- make main's 16KB flush split frame-aware instead of a blind byte offset
- lift the synchronized-output scanner into shared/ so main and the renderer
use one implementation
Closing a frame early costs one premature repaint; leaving it open costs a
second of blank screen, so the asymmetry favours always closing.
Also adds the reproduction this needed: the pre-existing typing bench observes
the xterm BUFFER, which the parser fills while rendering is held, so it scored
these freezes as fast echoes.
* fix(terminal): stop the renderer's queue drain cutting inside an open DEC 2026 frame
takeQueuedChunk sliced a queued chunk at a blind byte offset to fit the 16KB
coalescing budget, which can strand a frame's closing \x1b[?2026l in the residual
until a later drain. Same defect as main's flush split, same fix: reuse the
frame-aware split helper.
Usually masked because the drain coalesces adjacent chunks and reassembles what
main split, but not when the budget boundary falls inside a frame.
* fix(relay): keep the SSH path's bounded slice outside an open DEC 2026 frame
pty-handler split pending output at a byte offset with a surrogate-pair guard but
no synchronized-output awareness, so a frame straddling the 16KB wire slice had
its closing \x1b[?2026l stranded in the remainder — the same defect just fixed on
the local path, on the path AGENTS.md requires us to consider.
Placed before the surrogate guard so that guard keeps the final say, and floored
at 2 so frame alignment can never walk a healthy slice into the guard's
decrement and then into the chunkChars <= 0 pause-and-retry path.
Also drops a dead `splitAt === 0` branch in takeQueuedChunk: both callers pass a
positive limit and the helper never returns 0 for one.
The two new split tests were each confirmed to fail without their fix.
* test(terminal): sweep the DEC 2026 split helper over escape-sequence shapes and every limit
Covers OSC 52, DCS, repeated open/close markers and limits 1..len+3, asserting the
result never exceeds the limit, never reaches 0, and stays byte-exact. Also pins
that a buffer beginning inside an open frame degrades to the blind offset rather
than doing something worse, and documents that callers do not thread latch state.
* fix(terminal): ground DEC 2026 on the daemon slice, the recovery replays, and the process boundary
Four more sites could strand the latch, found by sweeping every path that drops,
splits, or replays terminal bytes.
- daemon-stream-data-batcher: the 64KB bulk-write slice used a surrogate-only
clamp, and its remainder is HELD until 'drain' — "seconds for multi-MB
backlogs" per the file's own note. A frame straddling that boundary parked its
\x1b[?2026l behind the hold, blanking the pane past xterm's 1s timeout once per
frame for as long as the backlog lasted. This is the default daemon-backed pane
path, so it is the one users actually hit. The new
clampToSafeBulkWriteSplitIndex frame-aligns first and surrogate-clamps last,
and lives in daemon-stream-data-split alongside the policy it belongs to.
- replay-data-drain and remote-runtime-terminal-binary-snapshots wrote a bare
\x1b[2J\x1b[3J\x1b[H, which does not clear mode 2026 — so on the SSH/remote
reconnect path, the very event most likely to sever a frame, the whole replay
could paint nothing.
- ipc-pty-attach: trimIncompleteTerminalControlTail can cut a half-written
\x1b[?2026l while its opening marker survives in the replayed prefix.
- PROCESS_BOUNDARY_GROUND: the "process that armed these modes is gone" ground
omitted 2026, the last unexplained gap in that file. A disable, so it still
satisfies the recovery barrier's ownership scan (only ?25h may be an enable).
Recovery-path expectations updated where they pin the emitted bytes. Deliberately
NOT touched: apply-reattach-payload and ssh-snapshot-prepaint already ground via
buildSnapshotReplayPrologue.
Still unfixed, deferred with reason: terminal-output-frame-chunks.ts splits the
remote wire on accumulated UTF-8 byte width and needs a different shape than the
char-index helper; desktop clients reassemble in main's pending buffer, so the
exposure is mobile/web only.
* fix(terminal): emit the DEC 2026 release before the mode-2031 tail, and stop claiming the drop path writes it
Two corrections from adversarial review of the earlier commits.
1. Ordering bug I introduced. getDroppedMode2031RendererData ends with
`state.tail`, which extractPrivateModeScanTail deliberately retains as an
INCOMPLETE private-mode sequence so the next chunk can resolve it. Appending the
2026 release after it put an ESC behind a dangling CSI, aborting it and silently
losing whatever mode spanned the drop boundary. The release now goes first.
2. The drop-path release does not reach xterm in the dominant case, and the comment
now says so instead of implying otherwise. live-data-callback's droppedOutput
branch discards `data` and salvages only queries
(salvageRendererQueriesFromDiscardedRestoreData handles CPR/DA1/OSC colour;
\x1b[?2026l is not a query), so for hidden panes and visible panes outside
foreground-restore backpressure the synthesized release was dropped. The grounded
snapshot replay releases the latch instead.
I tried writing it through writePtyOutputToXterm there and reverted: it consumes
the pending hidden-output snapshot and broke
pty-connection-hidden-snapshot-resize-signals ("re-restores a skipped alt frame"),
so the release rides the restore rather than perturbing that state machine.
Residual gap, documented: a cap-dropped pane whose restore never arrives.
The salvage is still load-bearing on the fall-through path, so it stays.
* fix(terminal): release DEC 2026 on the reattach clears, floor the split, and correct the freeze framing
Remaining findings from adversarial review.
- apply-reattach-payload's three bare-clear branches (:63 daemon snapshot, :229
relay replay, :269 cold restore) had no release anywhere in their sequence: I
checked all seven POST_REPLAY_* profiles reachable via chooseReattachReplayReset
and none contains \x1b[?2026l. Only the buildMainModelSnapshotReplayWrites branch
was grounded, so covering the streamed replay path and not the main reattach path
was inconsistent. Verified no production code matches these clear strings — the
three test updates are mock equality, and each was confirmed to fail without the
source change.
- clampToSafeBulkWriteSplitIndex could return 0 (('\u{1F600}aaaa', 1) — alignment
returns 1, the surrogate clamp decrements to 0), which would leave a zero-length
slice that never shifts the batcher's queue entry and spin its drain loop.
Unreachable from today's only caller, but it is exported with an unstated
precondition. Floored at 1.
- Frame alignment could halve per-PTY flush throughput: main re-queues the
remainder with eligibleRound = round + 1, so the shortfall cannot be refilled in
the same round, and aligned size is floor(W/F)*F — 50% worst case in the 8-16KB
band, which is exactly the full-screen redraw burst that reaches the pending cap.
Alignment is now rejected below half the window, preferring throughput and
letting the reset profiles release the latch.
Framing corrected throughout: bufferRows records a row range and clears nothing, so
the pane freezes on its last painted frame — it does not go blank. The real trade is
"stale but coherent for <=1s" versus "immediate partial frame", and
RESET_AFTER_BYTE_GAP (written alone, with no repaint behind it in the same write) is
the one site that can newly flash a partial frame. Said so at the constant instead
of implying the release is free.
* fix(terminal): rename the shape-flagged symbols the anti-slop audit rejects
CI's anti-slop gate rejects "shape" in symbol names as structural rather than
domain language: `shapes` -> `outputSamples`, and
`writeCodexShapedEchoProbeScript`/`codexShapedEchoProbeScript` ->
`writeCodexEchoProbeScript`/`codexEchoProbeScript`.
|
||
|
|
45c63a66e9 |
test: delete the source-grep tests an earlier detector's regex missed (#23976)
A rebuilt detector found 195 source-grep candidates where the original found 111.
The gap was one over-specific regex: the first scanner required a literal `.ts`
path inside `readFileSync(...)`, so every test that built its path from variables
(`join(dirname, '..', 'foo.tsx')`) was invisible to it. Roughly 84 files of a
pattern an earlier wave reported as cleared had in fact survived.
Deleted whole, every case asserting on production source text:
- `app-startup-routing.test.ts` (27 cases) — exact import statements
(`"import('../components/UpdateCard').then"`), relative-path spelling, and
`indexOf` source ordering. A file move or a `lazy()` refactor breaks it.
- `pull-request-page-host-boundary.test.ts` (13) — `toContain` on whole argument
expressions concatenated across 20+ component files.
- `SmartWorkspaceNameField-source-boundaries.test.ts` (7) — placeholder copy, a
Tailwind class string, and `not.toContain` on an already-deleted symbol.
- `github-project-repo-list-load.test.ts` (9) — `indexOf` statement ordering
inside `loadTasks`.
- `github-enterprise-slug-routing-boundary.test.ts` (4) —
`toContain('host: githubProjectHost(parsed?.slug.host)')`.
- `web-viewport-shell.test.ts` (3) — a regex demanding exact CSS selector-list
ordering and whitespace.
- `agent-catalog-links.test.ts` (1) — restates two `homepageUrl` literals straight
out of `agent-catalog.ts` with nothing in between.
Trimmed, keeping only what nothing else can reach:
- `desktop-startup-ordering.test.ts` 549 -> 66 lines, retaining the three cases
named as `assertionRefs` by the `ssh-filesystem.stream-inactivity-lifecycle` and
`agent-browser.owner-boundary-cleanup` gates; 15 source-order greps went.
- `ResourceUsageStatusSegment.session-polling.test.ts` keeps its census that no
`setInterval` exists and `listSessions()` is called exactly once — an added poll
multiplies a global daemon scan and no behavioral test sees it. The
`indexOf('if (!open)')` ordering pair and four `not.toContain` lines went.
- `agent-skill-installed-command-callers.test.ts` 231 -> 86, keeping the
`readdirSync` census that discovers every `<AgentSkillSetupPanel` caller and
asserts set equality against the allowlist, so a new panel host cannot silently
show a default Update action.
Also in this wave, from the renderer lib/runtime sweep: 22 cases whose routing
signal the production path never reads — verified by mutation, stripping
`connectionId`, the WSL preference and the UNC path from four of them left all 29
tests passing — plus braille-spinner rows collapsed onto one regex range, copied
`WELL_KNOWN_LABELS` rows, and a whole `resolveAiVaultResumeStartupShell` describe
whose four darwin/linux fixtures all return before the login shell is read.
`config/reliability-gates.jsonc` drops the two `app-startup-routing.test.ts`
references; the manifest still validates for 140 gates.
|
||
|
|
ad2e1b5efa |
fix(terminal): restore the mouse format with mouse tracking, so phone swipes don't type into Codex (#23946)
* fix(terminal): restore the mouse encoding with mouse tracking in every snapshot Swiping to scroll Codex from the phone on a Windows host typed legacy `ESC [ M` mouse reports into the Codex composer (#23818). SerializeAddon re-arms mouse tracking (?1000h/?1002h/?1003h) but never the SGR encoding (?1006h/?1016h). Any snapshot taken from a desktop pane's xterm (the runtime seeds its headless model from it after a reattach, and serves it to remote viewers when no model exists) therefore restored "tracking on, legacy encoding", and the phone encoded wheel events as X10 bytes, which ConPTY hands to Codex as keystrokes. serializeWithAbsoluteCursor, the one wrapper every Orca snapshot producer uses, now appends the encoding xterm itself parsed, read from xterm's mouse state service. The daemon/runtime headless model reads tracking and encoding from xterm too, so its regex mirror of the DECSET stream is deleted (one source of truth; one less regex pass per PTY chunk). Mixed versions: no wire field changes. A new host's snapshot carries an extra DECSET that old desktop and phone clients already parse; an old host's snapshot restores exactly as before. With tracking off the encoding alone sends no reports, so the wheel still scrolls scrollback. * test(terminal): pin the mouse-encoding read against the renderer xterm build * fix(terminal): type the xterm mouse-state read behind named shapes |
||
|
|
26bb7c23f1 |
fix(terminal): run Orca's cmd.exe, path-named and setup-gated Codex launches without the shared server (#23933)
* fix(terminal): give plain shells and cmd.exe Codex launches --no-daemon Plain bash, zsh and fish tabs were never wrapped, so a typed codex skipped the shell function that adds --no-daemon. Wrap them (bash keeps its prompt and DEBUG trap untouched unless Orca asked for command markers), add --no-daemon host-side where no function can run (cmd.exe, path-named binaries), and move new tabs to a v38 terminal daemon so they get the new wrappers. * fix(terminal): keep plain bash a login shell and give the setup gate the codex function Plain bash and Git Bash tabs launch exactly as before again: the rcfile wrapper would have made every one a non-login shell. Plain tabs on the user's configured shell args stay unwrapped on both transports. The wait-for-setup gate's bash -lc now defines the codex function, so a sequenced Codex launch gets --no-daemon from the binary it actually runs. * fix(terminal): define the setup gate's codex function after setup finishes Setup can be what puts codex on PATH, so defining the function before the marker wait found no binary and skipped --no-daemon. * refactor(terminal): fold the SSH/WSL guard into the Codex launch planner and bound the gate test * fix(terminal): honour the pane's env deletions in the Codex opt-out check Also pin the setup-gate test's fake codex ahead of path_helper's PATH. * revert(terminal): launch plain zsh and fish tabs exactly as on main Drops the always-wrap for plain zsh and fish, the configured-args guard that only served it, and the v38 daemon bump: the daemon's launch configs and generated wrappers are byte-identical to main again. Keeps the host-side --no-daemon for cmd.exe and path-named launches and the setup gate's codex function. |
||
|
|
1aa943798a |
fix(codex): a Codex Stop is the interrupt alone, so it no longer says "Cancellation was not confirmed" (#23850)
* fix(native-chat): a Codex Stop that ended its turn no longer reads "not confirmed" Codex answers a turn's interrupt only as that turn ends, then sends the turn's own end. Orca also required its sweep of the turn's processes to prove them gone, and when that sweep could not read the process table it wrote "Cancellation was not confirmed." a few milliseconds before the turn's interrupted end arrived. A Stop that named its turn said "The provider had already finished this turn." in the same case. Codex's answer is now what confirms the Stop. The sweep still runs and still holds the turn's end until it finishes, but its result is logged, not shown. A Stop Codex never answers, which is a turn that never ends, still reads "Cancellation was not confirmed." * docs(codex): say why an interrupted turn's un-echoed send is withdrawn, for a steer and for a turn's own input * fix(codex): a Codex Stop is the interrupt alone A Codex Stop sent turn/interrupt and, alongside it, swept the turn's processes: it compared a snapshot of the app-server's descendants taken at each send and compaction against a fresh one and killed what was new, then waited for those processes to exit. The turn's end was held back until the sweep finished. Codex already owns this. It kills a turn's one-shot commands on interrupt and deliberately keeps unified-exec background terminals running across one, so the sweep killed work Codex means to keep, read the process table on every send, and delayed the interrupted end behind a kill-and-wait. A Stop is now turn/interrupt only. Codex's answer confirms it and the turn's end publishes the moment it arrives, through the same path as any other notification. A long-running command started in that turn keeps running until the thread ends, as it does in Codex itself. Removed: the turn-process module and its integration test, the snapshot taken at dispatch and compaction, the held turn/completed, the adapter dependencies that injected both, and the prompt-claim re-check that only covered the wait for the snapshot. * docs(crash-reporting): justify the self-kill ring size by the documented window-close burst |
||
|
|
9f4311598f |
fix(codex): trust the worktree Codex starts in, not a guessed repo root (#23937)
* fix(codex): trust the path Codex checks for bare-repo worktrees Codex keys a linked worktree's trust on the main checkout only when that checkout's .git leads back to the common git dir; otherwise (bare repo, --separate-git-dir) it keys on the worktree itself. Orca always wrote the main-checkout key, so Codex showed its trust prompt and worker-start failed at agent_readiness. Mirror trust.rs exactly, and pin it with a real-binary contract that runs in the existing Codex contract CI job. Fixes #23847 * fix(codex): trust the worktree path itself instead of mirroring trust.rs Codex looks up the cwd's own [projects] entry before any repo root (config_toml.rs get_active_project, loader decision_for_dir), so trusting the workspace realpath satisfies every git layout. Drops the resolve_root_git_project_for_trust mirror: simpler, cannot drift from Codex, and never widens trust past the folder Orca launched in. Cost is one config entry per worktree; entries older Orca wrote on main checkouts stay valid. The six real-git layout tests now assert the workspace key and, under the contract, that real Codex starts workspaceWrite for each. The contract probes the binary version once and fails at load when required but missing; its CI path filter now includes config-toml-trust. |
||
|
|
fed1eca486 |
test: stop restating internal tuning constants, keep the ones that are contracts (#23950)
Removes ~74 assertions of the form `expect(SOME_CONSTANT).toBe(<literal>)` where the literal is an internal tuning value — a timeout, retry count, debounce interval, cache TTL, circuit-breaker window, Tailwind class string. Those cannot fail for any reason a user would notice: they fail only when someone deliberately changes the number, and then the test is simply updated. They are copies of the declaration. The same pattern is NOT junk when the exact value is observable outside this process, so those were deliberately kept: - terminal byte contracts: `\r`, `\x03` ETX, Kitty escapes, `\x1b[?1;2c`; - wire and capability values: `agent.launch.v2`, protocol 3 / min-compatible 2, daemon per-feature boundary versions (a daemon survives app updates, so those pin what an old field daemon may be trusted with), relay header tokens; - security invariants: the `127.0.0.1` bind default, an empty iframe `sandbox`; - values external processes read: exit code 78 (EX_CONFIG) and exit code 3 (systemd `RestartPreventExitStatus`), `ORCA_AGENT_SESSION_SPAWN_TOKEN`, `npx skills …` commands users paste, on-disk journal schema versions, the `orca_<hash>` filename prefix the fish sweeper matches; - third-party names: expo-router's `unstable_settings` / `ErrorBoundary`, iOS Safari's 16px zoom threshold. Where a case asserted a relation rather than a literal — `A < B`, a sum of parts, a cap compared against a sibling budget — the relation stays and only the literal went. Test-only changes: no production file is touched and no test file is deleted. |
||
|
|
59b746ff3c |
feat(native-chat): one structured-chat journal database per host, owned by one process (#23613)
* feat(native-chat): one structured-chat journal database per host, owned by one process
Every structured chat on a state directory now lives in one SQLite file,
agent-session-journal.db, opened once by the process holding
agent-session-journal.owner: an empty SQLite file whose held BEGIN EXCLUSIVE is a
kernel byte-range lock, refused while another process holds it and released when
the holder dies.
- Stores own no connection: the per-chat handle, its close contract and the
close-retry registry are gone; closing a conversation drains its writes, and the
one connection closes last at teardown.
- The owner lock is taken at runtime start, before orca-runtime.json is written;
a process that does not own the chats is not published and refuses every
structured request with journalUnavailable and words that say what to do. It
retries the lock with backoff and runs the full install once it holds it.
- A journal that will not open fails the host install: every chat says "Unable
to load this chat." (journalCorrupt), and nothing is renamed, deleted or
rebuilt. A newer build's database is refused and left byte-identical.
- An append is one INSERT. The listing status is a column, written after the
rows it describes and keyed by (epoch, sequence).
- A per-chat journal from an earlier build is copied in verbatim (epoch UUID and
every sequence) on that chat's first open, and its directory is retired only
after the copy commits.
- auto_vacuum = INCREMENTAL, with freed pages handed back in bounded steps after
every delete.
* perf(native-chat): key journal rows by block so one chat's rows sit together
Each chat's live epoch owns a block of row ids, block * 2^32 + seq, so a chat's
rows share leaf pages with nobody else's, a replay is one range scan, and
replacing or rewinding a chat deletes one contiguous range. Measured on the
largest real chat (61 MB) beside 19 interleaved peers: 39 ms and 7.5 MB of WAL,
against 214 ms and 102 MB for a (session_id, epoch, seq) key.
- Ids are computed in Number arithmetic, never bitwise. A sequence is refused
outside [1, 2^32) and a block at 2^21, which keeps every id below 2^53.
- A replace, roll or import allocates a fresh block, moves the chat's pointer,
and deletes the old block in the same transaction, so no orphan block exists.
- The listing status write moves into its own writer beside the column.
* feat(native-chat): copy a chat's per-chat journal again when an older Orca wrote it after a downgrade
A per-chat journal.db that reappears after its chat was copied in is the newer
history: an older build, run after a downgrade, attached the chat and wrote it.
- journal_imports records the (epoch, tip) each chat was copied from, in the
copy's own transaction. A file already copied is never copied again, across
any number of restarts after a failed rename; a file that differs always is.
- Newest writer wins, per chat, with a row saying the chat was continued in an
older version of Orca. When both builds wrote past the recorded tip under one
epoch, the copy takes a fresh epoch, so readers reset instead of skipping rows.
- Each copied directory retires to its own .imported-<epoch8>-<ms> name, so a
second downgrade and re-upgrade never collides with the first.
* test(native-chat): fixture deps match the host journal database shape
Attach-flow and reconcile-attach fixtures stop passing a journal database those inputs do not take, and host and restore fixtures pass the one they now require instead of the removed journal root.
* test(native-chat): state why the runtime-state fixtures' existing casts are safe
* fix(native-chat): start up normally when this process cannot open the chat journal
A process refused the chat journal, because another Orca owns it or because its own journal will not open, failed startup restoration: the window booted in degraded no-save mode and a paired phone could not list any tabs. Startup restoration now treats the refusal structured requests are getting as having no structured host; terminals, tabs and saving go on, structured requests are still refused by the gate, and the install is retried on the next one. Any other install error fails startup as before.
* test(native-chat): name the owner-lock sweep test after the two sweeps it runs
* fix(native-chat): session history and terminal resume work while chats are refused
Session history (listing and preparing a resume) and a terminal typing a resume command only check whether a structured chat owns a provider session. In a process refused the chat journal they failed outright. They now take the refusal chats are getting as having no structured host, the same treatment startup restoration gets, through one shared helper; any other install failure still fails them. Chat requests keep the gate's refusal.
* test(native-chat): the first-work rename's fake journal saves the listing status
* fix(native-chat): open a chat whose per-chat journal file never got its schema
A crash between creating a chat's journal.db and creating its tables left an empty or schema-less file. Each chat used to open that file as an empty chat; the importer instead refused the open as "try again" forever. A file with no journal_sessions table is now read as never written, the same as one with no rows. A file that is not a database, or whose read fails, is still refused.
* fix(native-chat): let the event loop run between chats during startup restore
Opening a chat's journal is synchronous SQLite now that no per-chat directory
is created first, so the restore of every visible chat ran as one main-thread
task. Each chat now waits for a macrotask before it opens.
* fix(native-chat): import a per-chat journal in bounded batches
The one-time copy of an earlier build's per-chat journal ran as one
transaction, which blocked the main thread for 650 ms on the largest chat.
Rows now copy 512 at a time, each batch its own transaction, yielding to the
event loop between batches. The rows go into a block journal_import_blocks
reserves, which no reader follows and no other chat is allocated; the last
batch publishes the chat's pointer, repair marker and import marker together
and releases the reservation. A copy that stops midway leaves only that
block, which the next open clears and copies again. Two opens of one chat
import one after the other.
* fix(native-chat): refuse chats when the owner lock file cannot be opened
A lock file that is not a database, or cannot be opened, made the claim throw
before any refusal was recorded, so startup restoration failed on every
launch. The claim now sits in the same try as the database open and records
the same typed refusal.
* fix(native-chat): finish reclaiming pages a delete frees during a running pass
A reclaim pass ended as soon as the freelist stopped shrinking between steps,
so a second delete that freed more than one step's worth mid-pass ended it
early and left those pages on the freelist. A pass now ends only when a step
itself frees nothing, or the freelist is empty.
* test(native-chat): desktop session history is served while chats are refused
* fix(native-chat): a send to a chat holding a newer Orca's rows says to update
A chat opened read-only because a newer Orca wrote rows to it answered a send
with the generic write failure. It now refuses the way a database a newer Orca
wrote does, with the same reason and words.
* fix(native-chat): keep chat tabs while this process cannot list its chats
A process whose chats another Orca owns, or whose chat journal will not
open, has no structured host. Its session-tabs inventory still answered,
with no chat rows, and the renderer read that as "every chat was closed":
it removed the restored chat tabs and the next session save persisted
their placement away.
The inventory now says `agentSessionsUnverifiable` when the last tab
restore ran with chats on disk but no host to list them. The flag is set
and cleared at the per-client projection point beside the client-hosted
page hold, and the restore is memoised only once a host answered, so a
later lock takeover or journal open republishes the chats and clears it.
The renderer keeps agent-session tabs, and keeps cancellation tombstones,
against an inventory that does not affirm its chat set.
* fix(native-chat): say chats are open in another Orca, with this process's way past it
A process refused because another Orca owns the profile's chats sent the
generic `journalUnavailable` reason, so current desktop and phone surfaces
said "couldn't open this chat's history right now. Try again." — a step
that never helps while the other Orca runs.
The refusal now names its own reason, `journalOwnedElsewhere`, with the
refused process's kind (dev desktop, packaged, orcad) as a fact. Each kind
gets its own step: quit the other Orca, or give this one its own profile
(ORCA_DEV_USER_DATA_PATH) or data folder (ORCA_USER_DATA). The sentences
are added to the shared notice copy, the desktop catalogs in all six
locales, and the boot catalog.
A client that predates the reason reads it as none and keeps the code's
words; an unknown kind reads as the packaged app's step. The `message`
released clients print is unchanged. Which requests refuse does not change.
* fix(native-chat): restore chats on taking ownership, without a list to ask
A refused startup kept its hostless result, so after the owner quit this
process never installed a host, never reconciled restart leases, and kept
telling clients it could not list its chats until a desktop chat request.
Taking the lock now reruns startup restoration once and pushes the chats.
* test(native-chat): a navigation reply says chats are unverifiable while refused
* fix(native-chat): retry a refused owner lock at most every 5 seconds
The lock frees as its holder exits, but a refused process only learns that
on its next retry, and the 30 s cap left a second Orca refusing chats for up
to half a minute after the owner quit. One retry is an open and BEGIN
EXCLUSIVE on an empty file.
* fix(native-chat): install before deciding whether a takeover must republish chats
A list that landed on the refused startup after the lock was taken finished
after the takeover had already checked, so nothing republished. The takeover
now installs first, waits for any restore in flight, and restores only then;
the restore that clears "cannot tell" pushes the frames itself, so a list
that heals the inventory first reaches subscribers too.
* fix(native-chat): no takeover lands a host after the runtime stop
Quitting cancelled a refused claim's retry only at its end, so a retry firing
during the stop's awaits took the lock and installed a host the stop never
tore down, and the lock was then released under an open journal. The stop
now cancels the retry first, keeping the refusal, and repeats its teardown
while an install that began during it (a takeover already under way) is
pending, so no journal connection outlives the lock.
* fix(native-chat): show a thrown refusal in its own words, not its code
A refusal the host throws reaches the client as an RPC error whose message is
the bare code; its reason and facts ride only in the error's data, which no
client read. The chat pane's status line therefore printed
agent_session_journal_unreadable, a send took the bare "not sent" path, and
other writes said the outcome was unconfirmed.
One shared reader, agentSessionThrownRefusal, now reads the refusal from the
error data. A failed history read shows the refusal's read-history words, a
send keeps the refusal behind its Retry exactly as a returned refusal does, and
the other writes (desktop and phone) name the refusal instead of doubting the
outcome. The phone's read failure goes through the same reader.
* fix(native-chat): log a failed journal open once per distinct failure
Every chat request retries a journal open that failed, which is intended, but
each retry also logged the failure with its full stack: a junk database file
logged the same "file is not a database" error 189 times in a minute. The open
now logs a failure only when its code and message differ from the last one
logged, and forgets it once an open succeeds. The retry is unchanged.
The open moves to its own module beside the runtime, which had no room left.
* fix(native-chat): restore lists a chat from its per-chat file and copies it on first use
Startup restore opened every restored chat, and that open copied the chat's
per-chat file into the host database, so the first boot after an upgrade paid
the whole one-time copy before the chat list appeared.
A restore open now reads a chat that is still in its per-chat file straight
from that file, read-only, with the importer's own reader, and closes the file
before moving on. That read drives the listing, the status row and the
restart offer, as it did when every chat had its own file. The copy becomes
owed work on the chat's write queue: it runs before the chat's first write,
and a reader that reaches the chat awaits it. A chat the host already holds,
or that was copied before, still opens through the import and its reimport
rules, and so does a file whose read needs a repair written.
* fix(native-chat): no host stays registered after a stop an install spanned
Each teardown pass clears the registered host before it awaits an install in
flight, and that install registers its host when it finishes. The pass then
tore the host down but left it registered, so a request after the stop was
served by a host whose journal was closed. The stop now clears the slot once
its passes are done.
* fix(native-chat): checkpoint the journal with a full flush on macOS
synchronous = FULL fsyncs each commit, but macOS fsync leaves the drive cache
unflushed, so FULL alone does not survive a power loss there. With
checkpoint_fullfsync, each checkpoint uses F_FULLFSYNC; elsewhere it is a no-op.
The comment that said FULL alone was enough is corrected.
* fix(native-chat): delete a per-chat journal once its copy verifies
An imported chat's per-chat file was kept under an `.imported-*` name, which
doubled the disk its history takes. The copy now reads back from the host
database before it is published: its items, submissions, epoch and tip must
match the file's. Only then does one transaction publish the chat with its
import marker, and the file and its WAL files are deleted, the directory too
when nothing else is in it (a pre-SQLite transcript there is kept).
A copy that does not match is never published: the file stays, the chat is
refused as unreadable ("Unable to load this chat."), and the mismatch is logged
once. A file left behind by a failed delete or a crash matches the marker, so
the next open deletes it rather than copying it again; a file an older build
wrote after a downgrade still differs, and is still copied again.
* fix(native-chat): verify an imported chat a batch at a time
The check that a copied chat reads back as its per-chat file folded both whole,
each in one synchronous task: over half a second on the largest chat. Both
reads now go a batch at a time between turns of the event loop, like the copy
itself, and count rows as well, so a copy that lost a row with no item in it
is caught too.
* fix(native-chat): restore reads a chat's per-chat file a batch at a time
Restore folded a chat still in its per-chat file in one task, so the largest
chat's file held the main thread for about half a second at startup. The fold
now takes the file a batch of rows per turn of the event loop, into the same
fold a replay uses, and nothing reads it before it is done. The file is still
closed before restore moves on.
* fix(native-chat): end a per-chat copy on a turn of its own
A chat's first open ran the copy's last steps (the verified publish and the
per-chat file delete) and the replay of what was copied in one task. The copy
now yields before it returns, so the replay, which every open runs, is a task
of its own.
* fix(native-chat): commit a per-chat copy's batches without an fsync each
Each 512-row batch of a chat's first-use copy committed under synchronous =
FULL, so a large chat paid one fsync per batch, about a quarter of its first
open. The batches now commit under NORMAL, set and restored in the batch's own
task so no other chat's commit runs under it. The publish that makes the copy
visible still commits under FULL, and under WAL that sync makes every earlier
batch durable with it. A crash before it leaves only the unpublished block,
which the next open clears and copies again.
* fix(native-chat): roll back a chat journal transaction whose COMMIT fails
The shared connection's transaction rolled back only when its body threw. A
COMMIT that failed left the transaction open, so every later write, for any
chat, failed with "cannot start a transaction within a transaction", and reads
saw rows that never committed. Under the unsynced copy the failure also tried
to restore the sync level inside the open transaction, which SQLite refuses,
so the caller got that error instead of the COMMIT's.
One transaction helper now covers the body and the COMMIT, rolls back whatever
transaction survives, and rethrows the original error. Schema creation uses it
too. If that ROLLBACK fails as well, the connection is marked stranded: each
later use retries the ROLLBACK, and until one goes through every chat gets the
same "history unavailable, try again" refusal a journal that will not open
gives. The rollback that frees it also restores the FULL sync level.
* fix(native-chat): keep the chat journal connection until its close succeeds
Closing the journal dropped its connection handle before closing it. A close
that failed left the database reporting itself closed with the connection still
open, so the stop that retried the teardown found nothing to close and released
the owner lock over a live connection.
The handle is now dropped only once the close succeeds. A failed close keeps
the runtime pending and the lock held, and the next stop closes that same
connection before it releases the lock.
* fix(native-chat): publish the runtime only once it holds the chat journal lock
When this process could not open the owner lock file at all (a permission
error, or a file that is not a database), the runtime counted that as owning
the chats and wrote orca-runtime.json. That overwrote the real owner's entry,
so the CLI was sent to a process that cannot serve its chats.
A claim that throws is now refused like one another process holds: the runtime
starts but does not publish, the claim's existing retry keeps asking for the
lock, and discovery publishes once the retry takes it. Chats still get the
refusal for the failure itself, and startup restoration reruns on the takeover
the same way it does after another owner quits. A sole process whose lock file
never opens is not found by the CLI until it does.
* fix(native-chat): keep a chat's history when an older build started it over
The first copy deletes a chat's per-chat file, so an older build run after a
downgrade finds no file and starts the chat from nothing. On the re-upgrade that
fresh file was copied in as the newer history, replacing everything the shared
database held for the chat, and then deleted.
A file whose epoch is not the one last copied and that opens with
`session_created` is now kept: neither copied nor deleted, and the chat keeps
the history it has. A file that carried the copied epoch on is still copied
again, as before.
* test(native-chat): pin which chats startup restore copies
Restore copies a chat still in its per-chat file only when restore itself has
to write to it: settling what the last run left open, here a running tool call
or a send handed over and never answered. Every other restored chat stays in
its file until its first use.
* test(native-chat): pin the copy wait on a read that opens a chat restore opened
A read queued behind restore's open of the same chat reaches the conversation
through its own open rather than the listing. It must still wait for the
owed copy, or it reads the chat before its history is in the one database.
* fix(native-chat): record a set-aside per-chat file so no later open reads it
Setting aside a file an older build started over is decided once and kept in
the new `journal_set_aside` table (schema 2, additive), with the file's epoch
and tip as they were. Every later open of the chat skips the file without
opening it, across restarts and after the older build writes more to it:
anything written there grows from that build's own start, never from this
build's history.
The best-effort delete moves beside the per-chat file reader.
* fix(native-chat): set aside any per-chat file at an epoch this build never copied
A chat's per-chat file is deleted once its copy verifies, so a file that
reappears at another epoch was never this build's history, whatever its first
row says: an older build started the chat over, possibly rewinding it after
(`handle_forked`), or rolled the epoch of a file whose delete had failed.
Copying any of them would replace everything the chat holds, so each is set
aside. Only a file still at the copied epoch is copied again (it grew) or
deleted (it did not). The first-row check is gone.
* fix(native-chat): copy a reappearing per-chat file again only while this build has not written past the copy
A per-chat file that an older build carried on under the copied epoch was
copied again even when this build had also written to the chat since the
copy, or had rolled its epoch. The second copy replaced the chat's block,
so what was sent in this build after the copy was gone for good.
Now the file is copied again only when the chat still stands exactly as it
was copied: the same epoch and tip the import marker recorded. Otherwise it
is set aside like any other file that is not this build's history, left on
disk untouched and recorded so no later open reads it. A second copy
therefore never replaces rows this build wrote, keeps the file's own epoch,
and the fresh-epoch rewrite goes away. The row it adds now says the history
includes what the older version recorded, not that anything was replaced.
* test(native-chat): pin that a chat founded here keeps its history, and the v1 schema upgrade
A chat this build founded has a pointer and no import marker, so a per-chat
file an older build later starts for it is set aside. Nothing pinned that
half of the rule: letting such a chat be copied again replaced its history
and every test still passed. A second test pins that a database written at
schema version 1 upgrades in place, gaining the set-aside table and keeping
its import markers.
* test(native-chat): drop a lost copied row by patching the source, not wrapping it
* chore(mobile): restore the mobile lockfile to main's
* fix(native-chat): pass a classified journal refusal through a send or Stop unchanged
* fix(native-chat): refuse a read whose owed copy fails as a failed open does
* test(native-chat): measure only the replace's WAL in the block-key case
Opening the chats starts a free-page pass that waits one event-loop turn,
and the seed never yields one, so that pass was still pending when the
replace committed. It woke during the async stat and reclaimed the pages
the replace freed, adding ~500 KB of WAL whenever the stat lost the race
(Linux CI). Drain that pass before measuring and stub the replace's own.
* test(native-chat): the RPC fixture's status journal can save its listing status
The status feed now hands every projection to the journal, which decides whether it is worth saving.
* test(native-chat): state why the RPC fixture's status journal cast is safe
* fix(native-chat): refuse a per-chat copy whose rows differ from the file, not only its counts
* fix(bench): build the replay benchmark's baseline arm from the base tree and release its handles on failure
* fix(native-chat): retry a failed listing status save on the next read of a cached status
* refactor(native-chat): drop the chat journal owner lock; the process instance lock already guards the profile
The journal carried its own exclusive lock, with a retry loop, an in-process
takeover, lock-gated runtime discovery and a "chats are open in another Orca"
refusal. Every shipped process kind (packaged desktop, serve mode, orcad)
already refuses a second instance on one profile before the journal opens, so
the lock only ever mattered for dev desktops, which the next commit covers at
the process level instead.
The host now opens its one journal connection at install with no lock. What a
sole process whose journal will not open needs stays: the install refusal
recorded for the gate, the no-host startup path, and the unverifiable chat
inventory, now in structured-agent-session-host-refusal.ts. The unreleased
journalOwnedElsewhere reason, its processKind fact and their copy are removed.
* fix(startup): dev desktops take the single-instance lock, and a second one says why it quit
Dev skipped Electron's single-instance lock so parallel `pnpm dev` runs from
several worktrees would not quit silently, but two dev processes on the
default orca-dev profile then write the same stores at once. Dev now takes
the lock like packaged builds: a second launch on the same profile focuses
the first window and exits with code 3, printing one stderr line that names
the taken profile and how to run another copy (ORCA_DEV_USER_DATA_PATH).
Serve mode, the macOS diagnostic bypass and the E2E harness are unchanged:
an E2E launch still skips the lock unless it sets
ORCA_E2E_ENFORCE_SINGLE_INSTANCE_LOCK=1.
* refactor(native-chat): key journal rows by chat, epoch and sequence
Rows in the host's journal database are now addressed by the chat's own
identity, with `(session_id, epoch, seq)` as the primary key, the same
shape each per-chat file already used. The block-keyed layout goes with
everything built on it: the block column and its allocator, the 2^21
block ceiling, and the import's reserved block table.
A first-use copy writes its rows under the file's epoch, which the chat's
pointer does not name until the verified copy publishes it, so no reader
sees a half-copied chat. A try that stopped midway leaves only rows no
pointer names, and the next try deletes them before it copies again.
Replace, rollover and repair delete by (chat, epoch).
This build's history always wins: once a chat was copied or founded here,
any per-chat file that reappears is set aside, and the same-epoch copy
again after a downgrade is removed.
The bounded free-page reclaim after every delete is dropped;
`auto_vacuum = INCREMENTAL` stays at file creation, so a later periodic
reclaim can still be added. Session search keeps its own step.
The schema moves to version 3. Versions 1 and 2 were written only by
unreleased builds of this change and are refused as found, not migrated.
* fix(native-chat): open a chat journal a newer Orca wrote read-only instead of refusing it
After a downgrade, the host's journal database carries a newer user_version. It was refused
outright, so every chat's history disappeared. It now opens on a read-only connection, as the
per-chat journals did: each chat shows what this build can read, from the database or a per-chat
file never copied in, and every write is refused with "Chats were saved by a newer Orca. Update
Orca to keep using them." Nothing is written, copied, repaired or founded, and the file stays
byte-identical. A table the newer schema changed reads as the same read-only refusal, not damage.
* refactor(native-chat): leave the saved listing status to the change that reads it
Nothing in this change reads the per-chat listing status column: it was a stored copy of a fact
the status feed derives, written after every turn end and cleared on every epoch change. The
status_json / status_seq columns, their writer, the saved-status type, the status feed's save and
its retry on a cached projection all go, with their tests. The change that lists chats from a
saved status adds the column back beside its reader.
* fix(native-chat): a chat saved by a newer Orca says to update Orca, not to try again
When a newer Orca wrote the chat journal, this build opens it read-only. A send or a Stop was
refused with the reason `journalUnavailable`, so today's desktop and phone clients chose the
words for an open that can clear: "Orca couldn't open this chat's history right now. Try again."
Retrying never cleared it; only updating Orca does.
The refusal now names its own reason, `journalWrittenByNewerOrca`, whose words are "Chats were
saved by a newer Orca. Update Orca to keep using them." A read refused the same way names it
too. An older client does not know the reason, drops it, and falls back to the code's words
("Orca couldn't read this chat's saved history."), and released clients still print the message.
* fix(native-chat): a chat journal from an unreleased build reads as unusable, not as retryable
A chat journal database stamped with schema 1 or 2 was written only by unreleased development
builds of this change. Opening it threw a plain error, which every chat reported as "Orca couldn't
open this chat's history right now. Try again." Retrying never cleared it.
It now throws a named error that is classified as unusable, so every chat says "Unable to load
this chat." The one log line names the file, says an unreleased development build wrote it, and
says to move it aside. Nothing migrates or renames it.
* docs(native-chat): drop the second-Orca-owns-the-chats case from three comments
The chat-only owner lock is gone, so only a chat journal that will not open leaves a runtime
unable to list its chats.
* docs(native-chat): correct three chat-journal comments the redesign left behind
A per-chat file left without its WAL is set aside, not copied again; nothing runs an incremental
vacuum yet, so the auto_vacuum mode is kept for a later pass; and the idle sweep drops a chat's
in-memory fold, since a chat holds no journal connection.
* refactor(native-chat): stop exporting chat-journal names nothing imports
Each is used only inside its own module now; the teardown's export served a deleted test.
* test(native-chat): name the version-0 test for what it covers, and check every journal table
The test named 'migrates an older user_version forward' covers only a version-0 file that already
has its tables; versions 1 and 2 are refused. The table test now also checks journal_imports and
journal_set_aside.
* fix(startup): a second dev launch's exit line no longer claims it focused a window
The running dev instance may be a background launch or a server, which show no window. The line
now says only that this launch passed its request to that instance.
* fix(native-chat): a failed structured-chat install closes the journal connection it opened
The install opened the chat journal database and closed it only if the record store then failed
to open. A later failure, such as the model catalog wiring or the host constructor, left the
connection open, and the next install opened a second one in the same process. Every failure
after the open now closes it.
|
||
|
|
6194a7a1b6 |
test: drop private-internal and boundary-census tests with behavioral owners (#23941)
Fourth audit wave, cut short by a session restart, so this lands the verified subset rather than the full batch. Removes private-predicate cases whose behavior is already covered through the module's real entry point, and de-exports the seams they reached for. Also drops three whole files whose every case was a duplicate or a call-shape grep. The source-grep vein is close to exhausted. One auditor reviewed 15 remaining flagged files and deleted nothing: what is left is mostly legitimate architectural ratchets that no type checker and no behavioral test can reach — AST fences banning `as`/`any` in an RPC operation region, discovered-vs-listed set equality over subscription sites, count ceilings on unchecked reply readers, and assertions on generated WebView bundles (no CDN URL, no `</script` tokenizer escape, parses at the Chrome 74 floor). Those stay. |
||
|
|
21d4ae9448 |
feat(agents): pre-trust the folder wherever Orca starts an agent (#23744)
* feat(claude): pre-trust worktrees Orca creates Claude Code asks "Do you trust this folder?" on first launch in any folder it has not seen, which blocks unattended launches in worktrees Orca itself made. Orca now records where a worktree's content came from when it creates it, and before each Claude launch writes Claude's own folder-trust entry for that worktree's root (never the main checkout) when the new setting is on and the content is the user's repository. Forks, bare commits, folder workspaces and external checkouts keep Claude's prompt. The write takes Claude's lock, never creates or breaks the file, runs on the SSH host itself, and is revoked when the worktree is removed or the setting is turned off. Launches that already pass --dangerously-skip-permissions also skip the trust prompt for that one process only, via CLAUDE_CODE_SANDBOXED=1 on the command. * fix(claude): parse the relay trust request with a schema and ship its search keys * fix(claude): never write a WSL guest's trust into the Windows config A WSL worktree's Claude reads the guest's own config. Two paths still wrote its trust into the Windows host's ~/.claude.json instead: the Claude auth prep's fallback (runtime 'wsl' but the host config dir, when the WSL home cannot be resolved), which wrote a Linux-path key the removal revoke can never delete; and the Agent Teams leader, which passes no auth or distro and wrote a UNC key. Require the guest's own config dir, and treat any WSL worktree path as guest-only. * revert(claude): drop the skip-permissions trust shortcut Pre-trust stays limited to worktrees Orca creates from the user's own repository. The per-launch CLAUDE_CODE_SANDBOXED prefix skipped Claude's trust question in every folder for launches carrying the skip-permissions flag (Orca's default Claude args), including the user's own folders and fork PR worktrees, and the setting could not turn it off. Remove the prefix, the inherited-variable strip that existed only for it, and the Agent Teams leader-to-teammate propagation; restore the tests that pinned the prefixed launch string. * fix(claude): revoke SSH trust in the config file the grant used At spawn the relay resolves Claude's config from the launch env, which carries a CLAUDE_CONFIG_DIR set in Orca's Claude default env. claudeTrust.converge had only the relay's own process env, so removing the worktree or turning the setting off revoked in the default file and the grant outlived the worktree. Send the config-file keys with the request, as the local revoke already uses. * i18n(settings): translate the Claude worktree trust setting * fix(settings): say Claude trust applies when Orca starts Claude The description said Claude skips its trust prompt in any worktree Orca created. Trust is written only when Orca itself starts Claude there, so a `claude` typed by hand in a fresh worktree still asks. Say that, and bring the es/fr/ja/ko/zh translations in line with the new text. * fix(worktrees): treat a base on an Orca-added fork remote as fork content A worktree based on a named ref was always stamped as the repository's own content, so picking the fork remote Orca adds for a pull request (or a local branch tracking it) as the base made a fork's code eligible for Claude trust. At create time, read the repo's `remote.<name>.orca-created` markers and each branch's tracked remote in one `git config` call; a base on such a remote is stamped as a fork's content, and a read failure is not vouched for. Remotes the user added, such as `upstream`, stay first-party. * perf(claude): revoke worktree trust once per config file, not per worktree Turning "Trust worktrees Orca creates for Claude" off read and parsed the whole Claude config once per Orca worktree on the main process. Group the revocations by config file locally and by SSH connection, and make claudeTrust.converge take a batch of requests. * feat(settings): one agent-wide "trust the folder" setting in Settings > Agents Replace the Claude-only worktree trust toggle with a single setting, agentWorkspaceTrustEnabled (on by default; unreleased, so no migration). The row says what it does for every agent: agents Orca starts skip their "trust this folder?" prompt in that worktree or folder, turning it off stops new trust while existing trust stays, and while it is off unattended launches (orchestration workers, automations, the phone) stop at the agent's trust question until someone answers. Translations for es/fr/ja/ko/zh. Also restores the `awaitingUnnamed` chat catalog keys an earlier merge of main dropped from this branch. * feat(agent-trust): pre-trust the workspace for every preset agent at PTY spawn Every Orca-started agent PTY passes through one of the two spawn builders with its declared launchAgent, which survives setup-script wrapping. The builders now call one hook that, for a fresh launch (never a reattach or restored pane) with the setting on, applies the agent's trust preset to the worktree, folder workspace or main checkout it starts in. - One dispatcher, applyAgentWorkspaceTrust(preset, workspacePath, launch context), carries what a writer needs: the final spawn env, the Claude managed-account auth prep, the WSL distro and the SSH connection. - Claude joins the presets on both the claude and claude-agent-teams entries. Its writer stays grant-only in claude-folder-trust-file.ts: the file Claude reads (CLAUDE_CONFIG_DIR / custom-OAuth suffix / legacy .config.json / a WSL guest's own file), Claude's <file>.lock never broken and taken only when a write is due, atomic temp+rename keeping mode and symlinks, never creating the file, NFC + realpath keys. - SSH Claude launches forward the optional claudeFolderTrust spawn field so the relay grants with its own spawn env; old relays ignore it and Claude asks. Other presets keep the SFTP writer. A WSL launch never writes the Windows home: non-Claude presets skip it, Claude writes the guest's file or nothing. - Codex keeps the 20 s deadline its shared config lane needs; every other preset gets 1.5 s. A miss means the agent asks; trust bookkeeping never fails or blocks a launch. Removes the Claude-only machinery this replaces: the eligibility/host/ lifecycle/spawn modules, the persisted creation content-origin field and its classification, revoke-on-removal, the setting-off sweep, the claudeTrust.converge relay method and the Agent Teams leader special case (the leader pane now spawns through the hook with the claude preset). The agent config types move to tui-agent-config-types.ts so the config table stays under the line budget. * refactor(agent-trust): delete the pre-spawn trust writes the spawn hook replaces The spawn hook is now the only owner of agent folder trust, so remove every other writer: - the agentTrust:markTrusted IPC channel, its preload bridge and types, and all renderer callers (agent-trust-preflight and its callers in the background session, work-item direct launch, session continuation, worktree creation, folder workspace composer and session fork); - the main pre-spawn sites: the createdWithAgent preflight in worktree-remote.ts, markLocalWorktreeTrusted/markRemoteWorktreeTrusted and the runtime's markWorkspaceTrustedForAgent family with the markTrusted ports of the runtime create flows; - Codex's own launch-prep and resume-prep trust writes. Each of those launches reaches a spawn builder with launchAgent set, so the hook covers it. This also fixes a live gap: the worktree-remote.ts copy of the preset switch omitted Antigravity, so an agy agent started from a desktop worktree create still asked; the single dispatcher covers it. Trust is also written on the host the PTY actually spawns on, which removes the #11163 class of writing the wrong host's config. * test(agent-trust): type the spawn-builder trust fixtures and prove the spawn waits for trust The builder test passed untyped args (a string launchAgent) and cast its deps, which failed tc:node. It now builds both spawn states from a fully typed deps fixture and a typed restored pane, with no casts. Adds a case that holds the trust write pending and checks the builder does not finish until it settles, the ordering the deleted renderer and launch-prep tests used to cover. * fix(agent-trust): give SSH trust writes the 20 s deadline again The dispatcher gave every non-Codex preset a 1.5 s budget, including the SSH writers for Cursor, Copilot and Qoder, which make several round trips over the link. Before this PR those writes had 20 s (desktop) or no limit (runtime), so on a slow link an unattended SSH worker would now stop at the agent's trust question. SSH writes get the 20 s deadline back; local non-Codex writers keep the short budget, and Codex keeps 20 s. The relay's Claude grant keeps the short budget: it writes the relay host's own disk and does not cross the link. * fix(agent-trust): never pre-trust a home folder or a filesystem root A folder workspace can be the user's home folder or a disk root. Claude and Copilot let a trusted folder cover every folder under it, so pre-trusting one of those would silently trust everything on the machine for those agents. One check, isHomeOrFilesystemRoot, now refuses them for every preset: the dispatcher checks roots and this machine's homes (including the spawn env's HOME and a cached WSL guest home), the SSH writer checks the remote home it already resolves, and the relay checks its own home. The agent then asks, as it would without Orca. * refactor(agent-trust): drop the Codex launch plumbing that only carried trust The spawn hook replaced the trust writes in Codex launch prep and resume prep, which left the fields that fed them unread: CodexHomeLaunchContext.workspacePath and .launchAgent, the resume prep's workspacePath, and the structured Codex launch input's workspacePath (plus the extra target lookup that produced it). Remove them and their plumbing; unavailableManagedHomePath stays. Also removes test stubs of runtime trust methods this PR deleted, whose not-called assertions could no longer fail, and two comments that still described the old trust preflight. * chore(reliability-gates): point the trust gate at the spawn-time trust tests The agent-session trust gate still listed three test files this PR deleted (the renderer preflight, the Codex launch-prep deadline and the e2e trust completion suites), so check:reliability-gates, which runs in PR CI and in pnpm lint, failed on missing files. Its invariant also described the deleted IPC handler and pre-spawn writers. The gate now covers what replaced them: the spawn builders holding the spawn until trust settles, the fresh-launch and setting gates, the per-preset deadlines, and the home and root refusal. * fix(agent-trust): skip the relay Claude grant for a WSL shell On a Windows SSH host whose pane shell is wsl.exe, Claude runs inside the WSL guest and reads the guest's config. The relay still granted trust in the Windows host's own .claude.json, writing the Windows home for a WSL launch, which the local path never does. The relay now skips the grant there, so that Claude asks, as a local WSL launch does when Orca cannot reach the guest file. * perf(agent-trust): only agent launches wait on the trust hook Both spawn builders awaited the trust hook on every spawn, including plain shells, reattaches and agents without a preset. Awaiting even a resolved promise adds microtask ticks ahead of the pane-spawn reservation check, and this handler already keeps non-Codex spawns off an await because an extra tick reorders those reservation races. The hook now returns null when there is nothing to write, and the builders await only a real trust write. * test(agent-trust): keep the home and root cases off any real Claude config The home and root cases ran the real Claude writer with the test process's env, so a regression in the guard would have written trust for the home folder and / into whatever Claude config that env named. They now point CLAUDE_CONFIG_DIR at a folder that does not exist, and the writer never creates a config. * test(runtime): drop needless casts from the launch-host test The renamed launch-host test kept three `as never` casts on launch options that already match launchAgentTerminal's parameter type. The changed-lines casting gate reads the renamed file as new and failed on them. * test(agent-trust): type the Claude grant mock with the real writer's signature The mock took an unknown target, so installing the real writer as its implementation would not typecheck under strict function types. * fix(codex): drop the launch context the trust move left unread in local spawn env * fix(agent-trust): queue Claude grants per config file so a launch burst keeps them all Concurrent grants in one process retried Claude's file lock in lockstep, so each retry round admitted about one winner. Starting 12 Claude agents at once left 6 of them at the trust question with nothing logged. Grants for one config file now queue in-process; only Claude's own writes contend for the lock. The relay shares the writer, so bursts of SSH launches are covered too. * fix(agent-trust): never pre-trust a folder above a home either The guard refused only an exact home or a filesystem root. A folder workspace at /Users, /home or C:\Users was still pre-trusted, and Claude walks up parent folders for a non-git folder, so every non-git folder in the user's home became trusted. The guard now also refuses any folder that contains a home, on every host, and is renamed to say what it decides. * perf(agent-trust): skip the SSH round trips for Antigravity, which has no remote writer Every Antigravity launch over SSH now reaches the remote trust writer, which resolved the remote home and realpath'd the workspace over the link before writing nothing (the known remote gap). That delayed each launch by two SSH round trips, and up to the 20 s deadline on a stalled link. It now returns first. * fix(settings): keep the hidden folder trust row out of web-client settings search The paired web client hides the host-only "Trust the folder" row, but settings search still listed it, so searching "trust" opened the Agents pane with no matching row. Its search entry is now filtered the same way as Agent Awake. * fix(agent-trust): a failing breadth guard skips trust instead of failing the spawn * docs(qoder): New Tab now pre-trusts through the agent-wide spawn hook * fix(agent-trust): never pre-trust a home reached through a symlink Every trust writer stores the workspace's resolved path, but the breadth guard compared only the path as given. A folder workspace that is a symlink to the home folder (or a real home picked while HOME names a symlinked one, as on distros that link /home to /var/home) passed the guard, and Claude, Copilot and Cursor then trusted the home itself. The local dispatcher and the relay now compare given and resolved forms of both the workspace and each home. The SSH writer resolves the remote home alongside the workspace, in parallel, so it adds no round trip. Local non-Claude WSL launches still skip before any filesystem call. * fix(relay): a failing breadth guard skips Claude trust instead of failing the SSH spawn The relay ran its home/root guard and homedir() before its catch, so a throw there rejected the relay's terminal spawn. Same fix as the main dispatcher's: the whole grant, guard included, is best-effort. * fix(agent-trust): guard the path each writer stores, not the path Orca was asked to trust The breadth guard checked the launch's workspace while each writer stored a transformed path, so every new transformation opened a hole. Codex stores a linked worktree's main checkout: with a git repo rooted at the home, a Codex launch in one of its worktrees wrote trust for the whole home. One relay-safe host module now computes the stored path (Codex's main-checkout hop, then given and resolved forms of it and of each home), refuses a root, a home or a folder above one, and only then writes. Main uses it for local and WSL launches and the relay for Claude. An unknown home writes nothing, and the WSL home cache is keyed case-insensitively by distro. * fix(ssh): the relay writes every preset's trust on the SSH host itself Codex, Cursor, Copilot and Qoder trust over SSH was written from the desktop over SFTP: four or five round trips per launch, so it needed a 20 s deadline that outlasted the 8 s draft paste, the 10 s phone wait and the 15 s web-client create. It also skipped Claude's atomic rename, ignored CODEX_HOME, and stored the worktree where local Codex stores the main checkout. The unreleased `claudeFolderTrust` spawn field becomes `agentWorkspaceTrust`, sent for every preset. The relay derives the preset from the `launchAgent` it already receives and runs the same host writer main uses, on its own disk, within 1.5 s and with no extra round trip. Antigravity still returns early on the relay (its writer is unverified on SSH hosts), a WSL shell still skips, and any throw means the agent asks. Deleted: the SFTP preset writer, the remote Qoder writer, the SSH deadline clause and the desktop-side SSH root pre-check. * test(e2e): keep CLAUDE_CONFIG_DIR out of isolated Electron launches The spawn hook now writes Claude folder trust into the config CLAUDE_CONFIG_DIR names, so an e2e run started from a shell that sets it could add trust entries to the developer's real Claude config. Also drops a stale comment that still named Codex launch prep as the trust owner. * fix(agent-trust): guard Claude's resolve() form of the workspace too Claude's writer stores both resolve(path) and the realpath. The breadth guard compared only the given path and its realpath, so a workspace path that does not exist and climbs back with `..` (for example <home>/missing/..) passed the guard while Claude stored a key for the home itself. The guard now also compares resolve(path), so it sees every form a writer stores. * test(relay): pty.spawn writes agent trust before the agent's process starts Nothing exercised the relay handler's call into the trust writer, so removing that call, or no longer awaiting it, left every suite green while SSH launches silently stopped pre-trusting. The new case holds the trust call pending and checks the spawn waits for it, and that the call gets the request, the declared agent and the final spawn env. The reliability gate lists the new suite and records the resolve() form the breadth guard now compares. * fix(agent-trust): refuse a home only for agents that inherit trust from it The home and root refusal applied to every preset, so Codex, Cursor and Antigravity started asking in a home folder workspace, where they did not before. Only Claude, Copilot and Qoder let trust on a folder cover the folders below it; Codex matches its start folder or that folder's repo root, Antigravity the exact folder, and Cursor itself never inherits from a home, a folder above one or a shallow path. The refusal now reads a per-preset table in the host module, so the local and relay writers share the rule. * fix(agent-trust): trust Codex at the folder it starts in, as before Before this PR, Codex launch prep trusted the spawn's start folder. The spawn hook trusted only the workspace root and skipped terminals with no workspace, so Codex began asking in a floating terminal and in a subfolder of a non-git folder workspace: its lookup checks the start folder, then that folder's repo root, and a plain folder above it is neither. The hook now passes the resolved start folder for presets marked as keyed by it (Codex only), falling back to the workspace root. * fix(agent-trust): pre-trust a structured Codex chat's folder, as before Before this PR, creating a structured (native) Codex chat pre-wrote Codex trust for its folder through launch preparation. The PR removed that write and routed trust through the PTY spawn builders, which a structured chat never passes. Codex's app-server trusts the folder itself only when the chat's permissions can write it, so a read-only chat started running untrusted and ignored the project's .codex config. Creating the chat now calls the same dispatcher, behind the same setting, before launch prep. * fix(settings): plainer folder trust setting text |
||
|
|
2807332735 |
refactor(shared): bring constants.ts back under the max-lines limit (#23923)
* refactor(shared): move onboarding, notification and terminal platform defaults out of constants.ts * chore(lint): keep the shapedSidebar naming exemption on the file that now holds it * chore(i18n): regenerate the runtime catalog so it covers main's shipped keys |
||
|
|
9420d49bcb | fix(terminal): run Codex in Orca terminals without the shared background server (#23900) | ||
|
|
31012aeb09 |
test: remove assertion-free probes, copied inventories and export-shape checks (#23816)
Second audit wave, targeting three more junk patterns: - assertion-free cases that run code and assert nothing, so they pass no matter what the code does; - inventory literals re-typed from a production declaration, where the only way the assertion can fail is someone editing one of the two copies; - export key-set and export-shape loops (`typeof x === 'function'` over every export) that restate what TypeScript already enforces. Yield is much smaller than wave 1 on purpose: the assertion-free scanner has a high false-positive rate, because many flagged blocks assert through a shared helper or their oracle is "this must not throw". Those were kept. `mobileWebCheckArgs` in `config/scripts/run-mobile-web-app-checks.mjs` is de-exported — after the inventory comparison went away, nothing outside the module read it. |
||
|
|
6e1b7e7fa3 |
test: remove junk tests that assert source text instead of behavior (#23815)
Deletes 101 test files and trims 112 more, all matching documented junk patterns: exact source/import/string greps, copied inventories and export lists, duplicate invocations of a contract another test already owns, typeof-shape checks TypeScript already enforces, and self-comparisons. The largest group read a production `.ts` file and asserted on its text — for example a TaskPage test that required the source to contain `selectedRepos.find((r) => r.id === newIssueRepoId) ?? selectedRepos[0] ?? null`. Any behavior-preserving rename broke it; no behavior change ever did. Production-side follow-through: exports that only these tests imported are de-exported or deleted, stale comments pointing at removed censuses are dropped, and the reliability-gate registry, `cloud/package.json` test lists, and orphaned source-reading helpers are updated so nothing references a deleted file. Two files kept their real coverage and lost only the census scaffolding: `agent-status-producer-census.test.ts` now drives all five producers end to end instead of grepping the source tree, and `config-toml-trust-stale-writes` replaces an export-list parity check. |
||
|
|
e1362ada4c |
fix(terminal): stop inline-image decoders exhausting the renderer's wasm memory budget (#23499)
V8 reserves an 8 GiB guard region per wasm memory inside its 1 TiB sandbox, so an Electron renderer can hold only ~124 live wasm memories regardless of free RAM. @xterm/addon-image instantiated a SIXEL decoder per terminal at activation (and kept IIP decoders after the first image), so ~120+ terminals exhausted the budget: new panes raised 'WebAssembly.instantiate(): Out of memory' rejections, and the next Kitty/IIP image threw 'WebAssembly.Memory(): could not allocate memory' out of the parser, permanently wedging that terminal's write queue. The addon-image source patch now borrows SIXEL decoders from a shared pool only while a sequence is open (color registers stay on the terminal), drops IIP decoders after each image, and turns a failed decoder allocation into a dropped image instead of a parser throw. Bundles regenerated with regenerate-xterm-patches.mjs --write. Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
f8f656ca19 |
perf(ci): spend fewer concurrency slots per pull request (#23810)
A concurrency slot is charged per job, not per core, and the account's cap is the scarce resource: standard runner minutes are free and unlimited on a public repository. Two paths spent slots that bought nothing. The unit matrix ran eight fixed shards averaging 6.5 minutes each, 3384 job-slots a day and 68% of all slot demand, while the arm pool queued 10.5 minutes at p95 — the queue was the oversharding. Five shards run the same work in ~10.5 minutes each for three fewer slots per run. Bun profile persistence escalated to all six platforms on `config/`, `resources/` and `.github/` wholesale, which took 36.5% of the last 1100 commits through the full matrix where a platform-flavoured predicate takes 19%. A pull request now qualifies one platform unless the change is platform- flavoured, and the push to main re-qualifies all six, so an unescalated miss surfaces minutes after merge rather than at the next cron. Missing changed-file evidence and an unavailable dependency graph still fail closed to all six. |
||
|
|
25d9c57e3a |
fix(native-chat): keep a child turn settling on an error Codex will not retry (#23808)
#23801 removed the child-path reading of Codex's turn-ending `error` along with the import of the module #23682 deleted. That was the wrong half to remove: the primary journal path can rely on Codex's failed `turn/completed` arriving within ~32 ms, which is what #23682 established, but a child turn has no such guarantee, and without the error as its end the child's lifecycle row latches on `working` for the life of the session. Three tests assert exactly that and could not run, because the unresolved import had been skipping the unit matrix since #23682 merged. The reading is restored inline against `readCodexErrorWillRetry`, itself restored to `codex-structured-thread-facts.ts`, rather than by reviving the deleted module: its `thread-stopped-running` arm lost its only consumer when #23682 rewrote the primary path, so restoring the file would re-add dead code. Also drops `pr-workflow-parallelism.test.mjs`'s read of `.github/workflows/track-community-prs.yaml`, which #23796 deleted while leaving the assertion behind. Same failure class, and it fails the same shard. |
||
|
|
2ea3fb1d46 |
perf(ci): take advisory unit-selection evidence off the gate (#23776)
selection_evidence is continue-on-error on both the job and its comparison step, so it can never fail a PR -- it downloads the shard reports, compares selection against the full results and uploads a review artifact. But a caller's `needs: test` waits for every job in the called workflow, so living inside unit-tests.yml it held verify for ~36s after the last shard finished. It moves to its own reusable workflow called as a sibling, so it still runs on every PR and still uploads its artifact, but verify no longer waits for it. It is deliberately absent from verify's needs, and a contract test pins both that and its advisory status so it cannot drift back onto the critical path. Measured on a recent run: the shards finished, then selection_evidence ran 36s, then verify 3s. Only the last of those gates anything. |
||
|
|
c5fc0c6f26 |
fix(ci): keep a squash-merged RPC recording pin reachable through its pull request (#23720)
* fix(ci): keep a squash-merged RPC recording pin reachable through its pull request
Main's "RPC recording pin" check has been red since #22762: that branch pinned
the recording corpus to its own commit
|
||
|
|
3976ad4c59 | perf(test): remove obsolete structural snapshots (#23777) | ||
|
|
ec9f35e2ee |
perf(ci): plan the unit shards before the static-analysis gate instead of behind it (#23743)
A caller's `needs` gate the whole called workflow, so while the plan job lived in unit-tests.yml it could not start until static analysis and typecheck had both finished and passed -- and the shard matrix then waited on it. The two hops were serial when they did not need to be: planning reads the checkout, a git diff against HEAD^1, the import graph and the checked-in timing baseline in config/scripts/ci-shard-timings.json, and consumes nothing that static analysis, typecheck or the native-cache primer produce. Planning moves to its own reusable workflow so pr.yml can run it against code_paths alone, overlapping it with the gate. Measured across 99 runs, the shard matrix is created a median 93s earlier (p25 47s, p90 241s, never later). Planning stays a required predecessor of the shards, so an empty assignment cannot expand the matrix. The gate itself is deliberately left in place. It fires on 22% of runs, and the shard queue wait knees hard above ~9 concurrent ARM jobs -- 4s median below that against 218s at 15-19 -- so admitting 8 doomed shards per failed run would cost more in queue pressure than it returns in latency. Cost is one 37s ubuntu-latest job, which does not touch the ARM pool the shards contend for. A planning failure still fails the PR: the shards are skipped, and verify's check_job requires success whenever the classifier says tests should run, so it reports `test: expected success, got skipped`. |
||
|
|
ccdb324b63 |
Add CodeBuddy as a built-in coding agent (#23740)
* feat(agents): integrate CodeBuddy launch, status and session history * docs: record CodeBuddy lifecycle verification * fix(codebuddy): backfill scoped history and negotiate remote resume * test(cli): include CodeBuddy in known search agents |
||
|
|
d68eee3757 |
fix(runtime): retire an exited terminal before its stream end (#23492)
* fix(runtime): retire an exited terminal before its stream end An exit's durable retirement became asynchronous, so onPtyExit released the terminal stream before the retirement landed. A paired client answers a stream end by re-activating its pane; that activation still found the exited leaf, materialized it under the same session id, and registerPty dropped the pending retirement. The exited split pane came back as a fresh shell. The exit now stages the retirement into the in-memory session and publishes it synchronously, then notifies exit listeners, and only then makes it durable. A failed durable write is logged and left in memory for the next profile write instead of being rolled back, since the process is gone either way. This removes the pending-retirement latch and its post-await incarnation fence: there is no longer a window for them to guard. * test(runtime): a failed exit retirement still reaches disk Pins the no-rollback contract through a real Store and SQLite authority: when the retirement's own durable write fails, the in-memory retirement is carried by the next unrelated profile write, and by the app-quit flush when no other write happens. The delayed authority fixture can now fail its next write, and the acknowledged-retirement fixture reads the database a relaunch would load and models the quit flush. * test(runtime): a stream end observes the exit retirement already published The re-activation check alone passes with the listener ordering reverted, because activation awaits before its lookup. Record the session binding and publication count at the moment the exit listener fires so the ordering itself is pinned. * fix(runtime): an exit cleanup fault still ends the terminal stream * perf(runtime): exits retired together share one durable write * test(runtime): a refused staging write still retires the pane and ends the stream * refactor(runtime): describe exit retirement as staged, not durably accepted The retirement result is staged in memory before any write, and the removable-surface comment and the replacement-admission test name still described the old publish-after-durable rule. |
||
|
|
c5330d0d52 |
fix(native-chat): stop killing processes that only inherited a chat's spawn tag (#23460)
* fix(native-chat): stop signalling processes that only inherited a spawn token A spawn token is an environment variable, so every descendant of a provider child carries it. The Linux-only startup scan treated any carrier no lease claimed as a lost provider child and sent it SIGTERM, which also hit editors, tmux servers and nested Orca processes the agent had started. Remove that scan's killing consumer; the token scan stays for the reservation probe, and recorded owners are still stopped by identity during recovery. * fix(codex): remove the token-scan kill path from app-server teardown Every descendant inherits the spawn token, so killing each pid that carries it can reach processes the agent started that are not the provider. Production never injected this path; teardown always uses the process-group and descendant-snapshot proof. Drop it, its deps, and the now-unused spawn-token argument. |
||
|
|
2ca4ecbc61 |
feat(orchestration): let a structured chat run orchestration as itself (#22568)
* feat(orchestration): inject the Orca session id into structured children and let the CLI act as it Every structured session's child (native Claude, native Codex, and the terminal view) carries ORCA_AGENT_SESSION_ID and reaches the Orca CLI. The CLI sends the id in the orchestration envelope; when present it is the caller, and a caller flag naming anyone else is refused before any request. The id is stripped from inherited PTY env and from the SSH host-CLI passthrough, and crosses into WSL so the host can refuse the cross-host claim. * test(orchestration): pin session id injection for native Claude, native Codex, the terminal view, WSL, PTY inheritance and SSH * test(orchestration): pin one caller precedence rule across every CLI verb that names its caller Adds the per-verb table (flagless acts as the session; a conflicting --from or --terminal is refused before any request; the session's own spellings are accepted), the enumerated guess population with its positive control, the structured worker's own handle, the identity-less refusal for an older child, the unchanged terminal agent, and the envelope. dispatch-show's --from only fills preview text, so it passes through unfenced and a session's flagless preview names the address the real dispatch writes. * refactor(orchestration): keep the identity-less marker reader to the marker; the id is checked first * test(orchestration): pin that a host refusal of the session surfaces verbatim from the CLI * fix(orchestration): keep the identity-less marker beside the id for CLIs that predate it A CLI older than the id, reached through a global install when a shell rc resets PATH, would otherwise guess a sibling's terminal in a chat that no longer carries the marker. It refuses on the marker instead; a current CLI checks the id first, so the marker never makes a session with an id identity-less. * fix(orchestration): refuse a conflicting --from on gate-list and task-list scoped by --run A --run listing needs no caller, so both handlers skipped the resolver and a --from naming another actor was dropped silently under a session. The conflict check now runs on that branch too; terminal callers are unchanged. * fix(orchestration): name this app's CLI by absolute path for a structured session's login shells A provider can run each command in a login shell: Codex runs zsh -lc, and the profile rebuilds PATH, putting a global install (possibly an older Orca) ahead of the directory Orca prepended. ORCA_CLI_COMMAND, which an agent resolves the CLI from first, is now the absolute launcher in that directory (the native launcher on Windows), so no shell's startup files can swap it. The PATH prepend stays for shells that read no profile. Found by the live coordinator run of the next PR. * test(orchestration): pin a structured worker's CLI command as this app's absolute launcher * test(orchestration): run the zsh login-shell arm in the real-shell lane that installs zsh The ordinary Linux unit lane has no /bin/zsh, so the zsh arm failed there with ENOENT. It moves to a live-shell file registered in the shell-contracts lane; the bash arm keeps running in every lane. The lane guard's detector now also sees a zsh spawned through the ProcessSpec program field, which is how this test escaped it. * fix(orchestration): omit a structured child's CLI command when no launcher resolves, and pin its instance A bare `orca` fallback named GNOME's screen reader on packaged Linux, and an inherited value named another app's CLI. The builder now deletes any inherited value, sets the absolute launcher only when one resolved, and pins ORCA_USER_DATA_PATH so a current CLI dials the instance that minted the id. Renames the marker reader to hasStructuredSessionMarker and records why the terminal view carries the id without the marker. * fix(terminal): name this app's CLI launcher by absolute path in every local terminal ORCA_CLI_COMMAND meant three things by lane: an absolute launcher for a structured session, a bare name for WSL, and nothing for any other terminal, so a structured session's terminal view lost it. Local terminals now get the same absolute launcher the structured lane gets; WSL keeps its guest command name, and a terminal whose launcher does not resolve still gets none. * feat(cli): hand a command to the session's own CLI when another Orca CLI was invoked A login shell can reorder PATH behind a global install, and an agent or its helper script can run bare `orca`, so the binary that answered depended on the agent following instructions. Orca's packaged launchers and bare-orca shims now export ORCA_CLI_SELF (outermost wins). At the CLI entry, when it names a different launcher than ORCA_CLI_COMMAND, the command re-runs once through the named launcher with ORCA_CLI_REEXEC=1 and exits with its status; both variables are consumed so no child inherits them. Dev launchers export no self on purpose, WSL and SSH names never qualify, and a launcher that cannot start leaves the command to run here. The Windows launcher no longer rewrites ORCA_CLI_COMMAND; the legacy ask protocol normalizes its resume command itself. * refactor(orchestration): declare which flag names the caller on each spec and refuse at the CLI entry Each handler hand-classified its --from/--terminal as the caller or a target, and the refusal of a conflicting caller flag ran inside the caller resolver plus two standalone calls for --run listings, so a new verb that read its flag raw would pass a sibling's handle to a pre-session host. Specs now declare identityFlagRoles, the CLI entry refuses a conflicting caller flag once from the spec, the resolver only applies the id-wins rule, and a test fails any orchestration verb that accepts --from or --terminal without classifying it. * perf(cli): keep the session caller check off the actor codec's module graph The check runs at the CLI entry for every command, and the actor codec pulls zod through the session record. Compare the session's own spellings as plain strings instead. * refactor(cli): spell a session's address from the one prefix constant, off the codec's module graph The Orca session address prefix moves to a leaf module with no imports, re-exported by the address codec, so the CLI entry check derives `session:<id>` from that constant instead of re-typing it and still stays off the codec's zod graph. Prose and test names say caller or Orca session id, not actor. * refactor(orchestration): drop the session id's terminal-view spawn now that the handoff is gone The terminal handoff was removed, so no terminal is ever a structured session: - delete the terminal-view identity env and its WSL passthrough, and their tests; - strip the session caller keys from every terminal's env unconditionally; - the CLI's own-address spelling moves beside the injected id in src/shared, with a test pinning it to the address the host's party resolver gives that session. * fix(terminal): run the Codex launch preflight through the CLI the terminal names Packaged Linux names the userData shim in ORCA_CLI_COMMAND, while the preflight ran the bundled launcher behind it. The CLI saw a different launcher and handed the preflight off to the shim, booting Electron twice before every codex launch. * revert(terminal): keep terminals on main's ORCA_CLI_COMMAND and Codex preflight Only a structured session needs an absolute ORCA_CLI_COMMAND; local terminals go back to naming none (WSL keeps its guest command), and the Codex launch preflight goes back to the bundled launcher. The CLI handoff is scoped to sessions, so a terminal's preflight can no longer be handed off and start Electron twice. This reverts commit |
||
|
|
8c61a5df1f | fix(windows): require signed release binaries and identify CLI launcher (#23680) | ||
|
|
0f52bb8be5 |
perf(ci): use four ARM test workers and remove repeated compilation (#23685)
* ci: benchmark per-job Node compile caching on full unit shards * ci: measure unit shards with three and four workers * ci: benchmark localization extraction CLI patch * perf(build): reuse identical relay bundles across platforms * ci: compare Vitest 4 and 5 on complete ARM shards * perf(ci): upgrade localization extraction to skip irrelevant syntax walks * perf(ci): use all four ARM cores and remove benchmark workflows * ci: preserve failures while capturing unit source revision * fix(ci): preserve commented and escaped localization calls * ci: remove corrected localization benchmark harness |
||
|
|
a84bd1c4fd |
fix(claude): write only the hook events and statusLine the user's Claude accepts (#23614)
* refactor(claude): name the Claude version module after the hook events it gates Pure move of claude-session-end-hook-capability.ts and its tests; the next commit turns its one-event SessionEnd floor into a per-event version table. * fix(claude): write only the hook events the resolved Claude knows Claude 1.0.81 through 2.1.100 validate settings.json `hooks` against a closed event enum and discard the whole file on one unknown name, so Orca's install made Claude <= 2.1.77 silently ignore the user's env, permissions and hooks. Each managed event now carries the first Claude release that knows it (pinned to per-release enums read from the npm packages), and install, status and the SSH/WSL relay installer write only the events the resolved Claude accepts. An unresolved version gets the set every tabled Claude knows; a downgrade removes only Orca's own entry for an event the older Claude would reject. * refactor(claude): move the managed Claude hook events into their own module hook-settings.ts is at its line limit; the event list and its version gate move out whole so the next change has room. * fix(claude): an unresolved Claude version never removes Orca's hook entries A failed or timed-out version probe is no evidence of an old Claude, so it must not strip StopFailure, PermissionRequest and the other newer events a version-aware install wrote. With the version unknown, install adds only the set every tabled Claude knows and leaves every other entry exactly as it is; only a known version that lacks an event retires Orca's entry. * fix(claude): gate the core hook events on the Claude release that added them Claude validates hooks against a closed event list from 1.0.23, not 1.0.81. The table treated SessionStart, UserPromptSubmit, Stop, SubagentStop, PreToolUse and PostToolUse as known by every resolved version, so a Claude from 1.0.23 to 1.0.61 was still sent names it rejects, and it dropped the whole settings file. Pin each to its first release from the packed enums and keep the unresolved-version set as its own policy. * fix(claude): write Orca's statusLine only for a Claude that knows it Claude 1.0.49 through 1.0.66 also reject any unknown top-level settings key, and statusLine joined that schema only in 1.0.64. Orca wrote its statusLine for every Claude, so 1.0.49 to 1.0.63 still dropped the whole settings file even with the event gate. Gate statusLine on 1.0.64, pinned by the packed schemas; a known older Claude has Orca's own statusLine removed along with the opt-out marker, so an upgrade re-adds it. An unresolved version is now assumed to be 1.0.64, which knows the same core events and keeps the statusLine install it had before. * test(claude): a user statusLine opt-out survives a downgrade and upgrade Retiring Orca's statusLine for a Claude older than 1.0.64 forgets the install marker only when Orca's own statusLine was removed. Pin that, so a user who deleted Orca's statusLine is not opted back in by an upgrade. * test(claude): check the whole written settings file against each strict schema Claude 1.0.49 through 1.0.66 discard the whole settings file over any top-level key their schema lacks. The fixture recorded only whether each release knew statusLine, so a new top-level key Orca wrote would pass every test. Record each release's top-level keys instead (statusLine is derived from them), add the hook enums for every packed release in that window, and check that a real install and a downgrade write only keys and events each strict release accepts. |
||
|
|
95a16e3f67 |
fix(release): stop the release policy from deleting pipeline-cut releases (#23669)
* fix(release): stop the release policy from deleting pipeline-cut releases The policy judged a release by who created the release object. Cut Release reuses an existing draft, so a CI-built v1.4.216 whose draft a person had created was deleted (tag included) when its notes were edited, and Latest fell back to v1.4.214 because v1.4.215 was also published by a person. - Authorize a desktop release when its annotated tag was created by the release pipeline and points at its `release: vX` commit, not only by author. - Only delete on `published`; an edit never deletes a release or tag. - Pick Latest from the highest authorized stable using the same check. - Move the policy into config/scripts/release-policy.mjs with tests. * fix(release): load the policy module from the tagged commit Release events run the workflow file from the tag's commit, so checking out the default branch could pair an old workflow with a newer module. |
||
|
|
e8d8f2f0b3 |
fix(deps): take Electron 43.7.5 so detached webviews stop blanking browser tabs (#23586)
Electron 43.7.0 threw 'Invalid guestInstanceId' from <webview>'s disconnectedCallback for a loaded guest (electron/electron#53989), so a webview React removed and re-inserted kept a dead guest id: the tab went blank, reload did nothing, and the destroyed listener never fired. 43.7.4 (electron/electron#54097) returns early when the guest is gone. Raises the runtime floor test to 43.7.4 so a downgrade cannot re-ship it. Fixes STA-8757 |
||
|
|
21f8e0f9db |
ci: use faster gzip for temporary Linux test packages (#23609)
* ci: benchmark faster Linux package compression * ci: pass compression options through typed builder configuration * ci: retain original configuration for benchmark baseline * test: preserve release settings in CI compression configuration * ci: normalize generated changelog dates in package comparison * ci: remove completed Linux compression benchmark |
||
|
|
cf20423ff3 |
ci: skip unrelated installs and share xterm build dependencies (#23607)
* ci: pilot shared xterm installed dependencies * ci: bound xterm cache production to verified main entries * ci: benchmark xterm reuse on the production ARM runner * ci: avoid installing Orca dependencies for standalone xterm checks * ci: use Node-only setup in the production xterm job * ci: remove completed xterm benchmark workflow |
||
|
|
8b410b4893 |
feat: add first-class Qoder CLI support (#23581)
feat: add first-class Qoder CLI support Integrate Qoder launch, identity, canonical hook status, trust and resume. Verify with captured Qoder 1.1.64 transcripts and hidden Electron sidebar checks. Builds on and cross-reviews #7502, #8611, #9655, #12910, #13311 and #15291. Co-authored-by: dalveytech-vincent <vincent@dalveytech.com> Co-authored-by: Eridanus117 <45489268+Eridanus117@users.noreply.github.com> Co-authored-by: xingqingzzp-gif <xingqingzzp-gif@users.noreply.github.com> Co-authored-by: jyang2004 <jyang2004@users.noreply.github.com> Co-authored-by: yunqian <yunqian@alibaba-inc.com> Co-authored-by: huzhening.hzn <huzhening.hzn@alibaba-inc.com> |
||
|
|
400e4e7957 |
feat(agents): add Freebuff launch and sidebar status support (#23567)
Add Freebuff launch support and execution-host status reporting for the sidebar, including running, question, blocked, and settled states. Validate against captured CLI transcripts and real rendered sidebar evidence. Cross-referenced community implementations #17065, #20839, and the Freebuff portion of #18790. Preserve their agent/catalog/mobile/documentation coverage and add canonical status publication and regression tests. Co-authored-by: Harkaran Brar <18134082+harkaranbrar7@users.noreply.github.com> Co-authored-by: Prarambha369 <98906077+Prarambha369@users.noreply.github.com> Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com> |
||
|
|
d05080175b |
perf(ci): stop duplicating shared Linux download caches (#23578)
* perf(ci): pilot pnpm verification record caching on Linux * test(ci): review pnpm verification record in mobile cache audit * perf(ci): share Electron downloads and clean closed PR caches * test(ci): retain cache ordering checks for restore-only consumers * perf(ci): limit archive sharing rollout to primed Linux hosts |
||
|
|
e6fbbdf684 |
perf(ci): cache pnpm verification records on Linux (#23568)
* perf(ci): pilot pnpm verification record caching on Linux * test(ci): review pnpm verification record in mobile cache audit |
||
|
|
3dd7d29455 |
Revert "perf(ci): shard the anti-slop audit across processes instead of one JS runtime (#23543)" (#23575)
This reverts commit
|
||
|
|
ba858ee446 |
feat(mobile): the keyboard covers the page like a native screen, and the shell says its height (#23110)
The shell no longer shortens the WebView for the keyboard; it publishes the keyboard height like the safe-area insets, so native's keyboard lift, refit hold and dismiss key run on the page unchanged. Keyboard and inset arithmetic read the shell's OS through a host-os seam. One page-version floor (manifest pageVersion, shell floor 1) replaces per-feature accept negotiation; a page below the floor gets the existing update wall, a desktop with no bundle keeps native screens. iOS shell drops the form accessory bar and its own keyboard observers. Native session screens untouched. |
||
|
|
080c562898 |
perf(ci): diff against the merge commit's first parent so PR checkouts can be shallow (#23562)
Every changed-path gate asked git for `--merge-base "$BASE_SHA" "$HEAD_SHA"`, which needs the event payload's base SHA to be in the local graph. That is the only reason two jobs cloned all 8127 refs' history. On a pull_request checkout HEAD is already the merge commit, so its first parent is the base side and no merge base has to be computed. config/scripts/git-pull-request-diff-base.mjs resolved that for the two Node gates; the workflow's inline gates now use the same helper through a small CLI rather than open-coding it. code_paths gates all 22 jobs, so its checkout is charged to the start of every one of them: measured 20.7s to 1.6s, keeping blob:none because its sparse tree is ~7 files and leaves no blobs to refetch. Static analysis drops the filter instead, since populating all 30,226 files makes blob:none force a second promisor fetch: 23s to ~11s. Verified on a real merge ref. At depth 50 the old and new forms produce identical changed-file sets. At depth 2 the new form still works and the old one fails with `fatal: bad object`, which is the failure a stale base would have caused once the checkout stopped being complete. Also drops the dead resolveBase + merge-base prelude in the changed-code gate, whose result resolvePullRequestDiffBase already discarded on every PR. |
||
|
|
2077956254 |
fix(mobile): size a terminal's first subscribe from the document's reported cell box (#23080)
* fix(mobile): size a terminal's first subscribe from the document's reported cell box #22960 sent phone dims on a terminal's first subscribe by opening a throwaway empty terminal (init 80x24 ""), awaiting its ready and measuring, behind a per-document first-subscribe mark whose lifetime was tied to web-ready. That cost a second xterm/WebGL instance and ~150 ms per open, plus lifecycle state. The document now measures the cell box without a terminal (xterm 6's CharSizeService strategy, rounded as the renderer rounds it) for every text-size preset and reports it with its viewport in web-ready; a table, because the text scale only reaches the document after that notify. Each init's ready reports the box xterm actually laid out, which replaces the probe's entry. The controller answers fitDimensions/measureFitDimensions from that table and the view's layout with no message; without a table it asks the document as before. The session seeds an unmeasured viewport synchronously in subscribeToTerminal, so the first subscribe carries dims by construction. Deleted: the empty init, its awaitReady gate, deferFirstSubscribeUntilViewportMeasured and the subscribedDocuments mark. The fit pass is unchanged and still covers a document that reports no cell box. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): read the reported cell box through in-narrowing, not Reflect.get Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): correct the probe's cell-box guess from the box xterm lays out The web-ready probe is a guess: building the WebGL addon creates no context, so a context that fails on load lands on the DOM renderer, whose width is not snapped and depends on the column count. Before, a ready box that differed was only logged; the first subscribe had carried the wrong column count, the host echoed it, the fit pass saw the viewport equal to the host's dims, and the grid stayed slightly shrunk. The store also kept the WebGL width after a context loss. The document now reports the box xterm laid out whenever it changes (from onRender, which covers a renderer swap and a DPR change that onDimensionsChange does not fire for, and at ready). The store replaces the guess; when that changes the current text size's entry, the view calls onCellBoxChange with xterm's grid and the session re-fits, running the bounded fit pass if the dims moved (one resubscribe). Equal boxes do nothing. The RN layout box now survives a document reload; the document's own viewport only stands in until the view reports a layout (on the page, web-ready arrives first). The mismatch console.log is gone, and the probe's rounding names the xterm version it copies. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): correct each cell-box guess at most once, so a DOM renderer cannot loop On the DOM renderer the cell width is the rounded canvas width divided by the column count, so every re-init at new cols reported a new box. Each one counted as a correction, a floor over floats could flip the fit between two sizes, and each flip landed converged, which reset the resubscribe budget: an unbounded series of full-snapshot resubscribes. Only the first laid-out box for a guessed text size may be a correction; later reports still update the store, so fits stay truthful, but never resubscribe on their own. The fit's floor gains a 1e-6 epsilon so floating-point error at an exact boundary cannot flip a column or row. New document tests pin the render report after a renderer swap and the report at ready for a paused renderer. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): make xterm the only terminal cell measurer The document builds its real terminal before web-ready, at the app's text scale, and reports the box xterm laid out; the first init reuses that terminal. The page-side prediction, the per-scale guess table and the once-per-document correction are gone. The app remembers the box per text scale for its lifetime, so a later open at a known scale subscribes with phone dims at once. A box that changes at the same grid (renderer swap, pixel ratio) refits the open terminal in place. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): keep commands queued before the terminal WebView first loads A subscribe sized from the stored cell box can queue init before the native WebView reports its first load start, which cleared the queue and left the terminal blank. Only a reload now drops queued commands. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): re-init a document that lacks the subscription's init, and fit one frame width - Web-ready now says whether the document holds the terminal's latest init (a reload before the first ready drops a queued one); the session resubscribes any initialized terminal whose document lacks it. - One grid fit, shared by the app and the document, fed the unrounded frame width React Native laid out; it keeps exact fits whole at fractional pixel ratios. The document's viewport-width fits are gone. - The page builds every document at the scale the view mounted with, as the native WebView does. - A new document's first cell box is compared against the grid the subscribe fitted from the stored box. - The terminal built before ready stays hidden until its first init. - The cell-box census matches glyph-measurement techniques, not names; the store's unused clear() is gone. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): build the terminal before ready only for the view shown at mount A session mounts one terminal view per tab, and each built xterm and a WebGL context before ready: 20 tabs made 20 contexts at load, past the ~16 a page (or Android's shared WebView renderer) holds, and native logged 32 context losses. Only the view shown when it mounts builds early now; the rest build at their first init as before. Deferring the WebGL addon instead would change the reported box: the DOM renderer lays out 7.8x15 where WebGL lays out 7.667x15 at the same font. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): write a WebView document's start values into its page, not an injected script Android ran the pre-content injected script after the document's own in 1 of 22 documents on the emulator; that document started with no text scale or shown flag and built a terminal it should not have. The values now sit in the page ahead of the document script, one source object per start pair so a render never reloads the WebView. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): the pre-ready terminal measures and reports while hidden Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): measure only the laid-out frame, and refit on a new grid, not a new width - A measure needs both of the frame's dimensions from React Native; the document's viewport-height fallback is gone, and before the first layout the handle answers no fit without asking the document. - A frame width change that still fits the PTY's grid from the stored box is a no-op, so sub-pixel layout jitter no longer re-measures. The width ref is written in that effect rather than during render (react-doctor). - One "last grid" ref: the last reported grid, or the one a subscribe fitted from the stored box. - The page render rig measures through the frame it laid out, as the session does, and lets the replay's fit settle before its resize-refit witness. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): let only the current terminal document's ready flush A reload kept the WebView and its onMessage, so the old document's late web-ready flushed the queue into the reloading view and the new document got a second init. Each document now gets its own view (keyed on a generation the controller owns), every notify carries the generation of the view that received it, and a web-ready from a replaced document flushes nothing and stamps nothing. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): drop every notify from a replaced terminal document One rule at the receive boundary: a notify from any generation but the current one is dropped, whatever its type, not only web-ready. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): make fitDimensions a pure question; name each generation counter - fitDimensions no longer records the grid. A width change to a new grid asked it first, so the DOM renderer's report of that grid's box read as "same grid, new box" and refit again. Only the first-subscribe seed (seedFitDimensions) records the grid the document's first report is checked against. - viewGeneration counts the views, readyGeneration counts web-readies. - replaceDocument no longer resets the load flag; the load-start reset stays as the guard for a view that reloads itself. - The name-based lifecycle census is replaced by a behavioural test. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): typecheck the handle mocks, drop the unused cell-box get - The two handle mocks carry both fitDimensions and seedFitDimensions, and the fake-timer acts return nothing, so the three test files check under tsconfig.test.json again. - terminalCellBoxes.get had no product caller; the store's tests assert through fit. - The load-start comment says what the controller does now. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): hold the grid the document has, ignore a replaced view's load start, dispose a failed pre-ready terminal - The document reports a new grid even with an unchanged box, so an in-place reflow on WebGL is held before a later renderer swap at that grid; the swap then refits. The app's apply paths do not hold the grid themselves: the DOM renderer's box follows cols, and a grid held on apply would read its own box as a renderer change and loop. One writer (holdGrid) holds the seeded or reported grid. - A load start from a view a replacement unmounted is ignored, as its notifies already are. - A terminal whose open throws before ready is disposed, not only unreferenced. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): ignore every native event from a replaced terminal view One wrapper binds each WebView lifecycle event (load start, error, HTTP error, render process gone, content process terminated) to the view's generation, so a replaced view's late event cannot reset, replace or put an error over the current document. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): a DOM seed refits once on its first report, not on the refit's own Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): subscribe a terminal only after its document is ready The document still builds its terminal before ready and reports the cell box xterm laid out in web-ready; the app now subscribes after that ready and fits from that box, so nothing is sent to a document before it is ready. Everything that made a pre-ready subscribe safe goes: the app-lifetime box store, the seed fit, the per-document view generations and their event filtering, the init tracker and the hasInit resubscribe. The native view reloads in place again and web-ready keeps main's reload rule. Boxes are kept per view; the grid a document last reported still guards the in-place refit against the DOM renderer's cols-dependent box. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): hold one reported cell box and the grid the subscribe fitted The controller keeps only the box the current document last reported, not a per-text-size store: the document re-reports on a scale change. The subscribe after ready fits from that box and holds the grid it fitted, so the DOM renderer's first report at that grid (a new box) refits once in place and converges; refit and apply paths hold nothing. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): fit only a ready box at the app's scale; forget a reloaded document's box and grid A reload keeps the document's mount scale, so a ready after a text-size change reports a box at the old scale; that box no longer sizes the first subscribe, which then takes the no-box path. A readiness reset drops the old document's box and held grid, so the new document's first DOM report at the same grid does not refit. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): give the terminal document its frame at init, and fit text scale over it only A subscribe sized from the ready box sends no measure, so the document had no frame when the text size changed and reported the pre-refit row pitch. The app's init now carries the frame it laid out, in the fields a measure uses; the router takes it from either. The text-scale fit reads only that frame, with no viewport fallback. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * docs(mobile): say why a frameless text-scale change skips the resize Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): one cell box per terminal notify, not an array web-ready and cell-metrics carry `cellBox: {fontScale, cellWidth, cellHeight} | null`; the document's `laidOutCellBox` returns one or null and the parser validates one object. The text-scale match moves from web-ready into `handle.fitDimensions`, the one place a box is fitted. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): fit terminals in the app from the reported box; drop the measure round trip The app already holds the box the document reported, so the refit and the fit pass await the init's ready and call `handle.fitDimensions` instead of posting `measure` and waiting on `measure-result`. The document's measure, its retries, and the measure promise and timeout go. The document still resizes locally on a text-size change, so every grid the app sends (init, resize, reflow) carries the laid-out frame it was fitted to. `holdSubscribedGrid` replaces `subscribeFitDimensions`, so the only fits are `fitDimensionsFromCell` and `handle.fitDimensions`. The render rig reads its fit from the ready box. The recorder adapter mounts the new handle with the same recorded effects; the goldens it mounts move on their adapterSha256 header only. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): keep the terminal frame in one ref, and notify a new width imperatively The session held the frame in a height ref, a width ref, a width state and the refit's own width ref. It now holds one `terminalFrameRef` ({width, height} | null until the first layout; a hidden 0x0 layout keeps the last box). onLayout notifies a new width imperatively, as it does height, and the refit's notify skips a width whose fit is the grid the PTY has. `terminal-frame-width-refit.ts`, the width state and its effect go. The subscribe's layout gate reads "no frame yet" directly. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): subscribe a held-back terminal on the frame's first layout only `handleTerminalFrameLayout` ran on every onLayout; it now runs once, when the frame first has a size. Later layouts only notify a new width. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): size the first subscribe inline in subscribeToTerminal `sizeTerminalViewportFromCellBox` wrapped five lines in a 37-line module; the subscribe now fits the ready box against the frame, holds that grid and records the diagnostic itself. The helper's tests fold into the subscription tests, which move to the subscription's name. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): drop the unreachable font-size guard on the reported cell box xterm 6.1.0-beta.303 updates the render service's cell box in the same task that sets `options.fontSize`: CharSizeService.measure fires onCharSizeChange, and RenderService.handleCharSizeChanged runs the renderer's `_updateDimensions` (DomRenderer.ts:359, WebglRenderer.ts:229). `term.onRender` fires from RenderService._renderRows after the rows are drawn (RenderService.ts:213, CoreBrowserTerminal.ts:538), and the document writes its text scale and the font size in one task (text-scaling.ts applyTextScale, terminal-init.ts init). So no report can read a box between the font and the scale; the guard and its test go. A new test pins the real order: no report when the font is set, the new box at the new scale on the next render. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): one start seam, no source cache, the reported box as an object - `useState` already pins each view's WebView source at mount (a new test re-renders at another text scale and gets the same object), so the module-level `webViewSources` Map goes. - `initialTextScale` and `buildsTerminalBeforeReady` become one `start(): { textScale, shown }` seam. - `reportedCellBox` holds the last reported box and grid, not a string key. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): repin the RPC recordings to this branch and re-record The terminal refit now fits in the app from the reported box and reads one frame ref, so the recorder's terminal adapter mounts the new handle (`awaitReady` + `fitDimensions`) and options (`terminalFrameRef`), keeping its recorded effects. `baseline` is repinned to |
||
|
|
7a24d3d335 |
fix(native-chat): the conversation outlives its agent (#22835)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * fix(native-chat): name a chat write by its target, not the owner generation A write carried the fence of the last frame the pane read, and the host refused it unless that fence was still current. An idle release and the restart after it each move the fence, and the release publishes nothing, so a send after a release was refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a cold start was refused as stale. Every write already names what it acts on: a send its conversation, a cancel its turn, a prompt answer its item revision, a rewind its epoch; an option is last-writer-wins. So admission stops comparing the client's fence, and the rebase that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it. The writer-lease check stays, and so does the attach's compare-and-swap. Frames now stamp the fence read when each frame is sent instead of a copy each subscriber kept, which went stale on the same release. * fix(native-chat): every journal append reaches the chats that are open A journal write and its delivery to open readers were two calls, and some writers made only the first. A failed start whose lease could not be handed back, a provider revision with no frame behind it, and eviction's settlement were all journaled without reaching an open chat. A journal handle now reports every durable change, and the host's session map binds that report to the session's readers when the handle is set. Writers no longer publish what they append; the per-writer publish calls are deleted. * test(native-chat): an epoch replacement reaches the open chat * test(native-chat): each row reaches an open chat once, and a live handle enters only through the map * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. * perf(native-chat): a publish behind a delivered commit reads nothing Each commit now delivers itself, so the publish a provider frame still sends afterwards found every reader caught up but still read rows and rebuilt the timeline for each one. A caught-up reader now skips the read. * test(native-chat): state why the teardown test's fake journal is safe to cast * docs(native-chat): say mutation admission checks only the writer lease * docs(native-chat): drop the send rebase from comments that still described it * fix(native-chat): a message is accepted, then delivered A send to a chat with no running agent restarted the agent inside the send call, before the message was recorded, so the client waited for the whole start and a failed restart refused the message. Claude held prompts sent during startup, and those could settle as "unconfirmed". A send is now accepted inside the session's serialized queue: one ledger row and one submission row marked handoverRecorded, published, answered pending. A per-session delivery loop exists while a message is queued. It starts the agent through the same serialized attach a hold uses, waits outside the queue for a Claude child to prove its start, and hands the oldest queued message over as its own serialized step, writing dispatch{pending} before the adapter call. A start it needed and did not get writes one error-tone row and rejects every queued message with the same words; a start Stop cancelled writes none. Settlement follows from the rows. A queued message is provably unwritten, so a close, an eviction or an exit rejects it. A handed-over message stays in doubt. A queued row at or below the sequence a handle found when it opened was left by an earlier process and is rejected at open, with no latch. Stop withdraws queued messages with no writer lease and no fence. An attach failure keeps the conversation open, and the attach adopts its journal. Owed work counts the loop and queued rows. A compaction or rewind found prepared when a conversation opens was started under a child this process no longer has, so the open settles it rather than leaving it to refuse every send until a view attaches. The open cursor is scoped to its epoch, because sequences restart when an epoch is replaced. Deleted: restart-before-admission, recordFailedRestart, the fence rebase, Claude's startup gate, the attach's forget on failure and its own crash boundary. Clients without agent-session.accepted-send.v1 get their reply held until the handover; the desktop and paired desktop lists advertise it. * fix(native-chat): settle queued messages only for the child that ended A child that proved its start and then exited before its message was handed over left the message queued: the exit settlement returned early when nothing else was in flight. Delivery then started another child for it, and a child that died the same way started another, without end and without a row. A retried settlement for an earlier generation, run by the attach that delivery started, did the opposite: with that generation's turn unfinished it rejected the message queued for the child being attached. The settlement now takes the rejection for queued messages from its caller. The unexpected exit and the eviction pass one, and it applies even with no other work in flight; the retry for an earlier generation passes none. * fix(native-chat): an adoption that fails to import keeps the conversation open The attach now writes into the conversation's own open journal, but a failed transcript import still closed it as if it were the attach's provisional one. The conversation stayed indexed with a closed journal, so every later send answered "could not be recorded" and every attach failed again until the app restarted. The import now closes only a journal the attach opened for itself. * perf(native-chat): the recovering open reads the journal once Every conversation open now goes through the recovering open, including the read restore of every chat at startup, which used to replay its journal once. The recovering open replayed it twice: once to probe it and again inside the open. The probe is now handed to the open as its load. * fix(native-chat): an attach that fails after indexing its child leaves no child behind A failed attach now keeps the conversation open, but a failure after `onAttached` indexed the child (the rewind or compaction recovery, or the attach's own success record) left that entry claiming a child the failure path had already released. The next send found the phantom, skipped the start, and wrote at a fence the journal had moved past, so the message stayed queued for good. The entry now drops the released child and its event sink, and follows the record's fence, as a failure before indexing already did. * fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer The error strip for a message the host accepted and then did not deliver matched the entry before the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not send your message" with nothing to retry. It now reads the reconciled entry. A rejection the journal records before the send's own pending answer lands is final as well: that answer no longer puts the entry back to dispatching with no Retry. * fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down The preamble waits for its submission to be delivered while the worker's agent starts. When that wait ran out it threw operation_unknown, and the failed-start teardown then closed the session, which rejected the very preamble the host was about to deliver. It now reports a turn start nobody observed yet: the worker is start-unknown with its session kept, the host delivers the preamble when the agent starts, and the worker's report settles the dispatch as for any unobserved start. The receipt no longer suggests reading a screen a structured worker lacks. * fix(native-chat): a message rejected while its chat was closed reads as not sent A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked every later message behind a Retry and no reason, and the delivery probe, seeing the journal already answered, never ran. The reconcile now settles it as rejected like a dispatching one. * test(orchestration): name why the readiness settlement fakes are cast * fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent * docs(native-chat): drop the fence from the admission the send effects run behind * docs(native-chat): give the fence move on release the reason that still holds * docs(native-chat): stop citing a write fence check in launch and mailbox comments Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease. * refactor(native-chat): the provider child is its own record A conversation now outlives any number of provider children, so the child is one record on the conversation's entry instead of five loose fields beside its journal. It is written in one place: indexed only once an attach has fully succeeded, and ended through one function that an exit, a failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence. - A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence patch after it are gone. - Conversation writes read the record's fence, the way mutation admission already does; a child's own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of the conversation's fence, are gone. - The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer dropped when an attach replaced the whole entry. - Stop on a child still proving its start stops only the child: its lease goes back and the chat is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus the conversation's close. - The settlement retry uses the conversation's own journal, opened through the host's one open. * fix(native-chat): the delivery loop alone settles a message its start or child failed A queued message was settled by whichever path happened to end the child first: the loop, the unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that rejected every pending row. That gave two failure rows with different tones for one start, a loop that could hand over to a different child than the one it waited on, and a Claude start that died while starting reading unlike every other failed start. - The loop remembers the child it waited on. At handover, if that child is gone or replaced, it reads how it ended: a Stop continues; anything else writes one failure row and rejects every queued message with the same words, then stops. A child still starting whose start the adapter says did not land fails the same way. The exit, eviction and the settlement retry only settle the handed-over and legacy rows of the child that ended. - One failure row, always an error, keyed by the start. A start a view began that dies with nothing queued writes the same row through the same builder, so a second report revises it. - The open no longer rejects leftovers; the loop's first step does, and the open wakes it. - `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the failure before the exit is processed. - Quit closes every conversation the way closing a chat does: what is still queued is rejected as closed, with or without a child, and a start the loop already has in flight is waited for so the child it produces is stopped rather than left behind. * refactor(native-chat): a stopped child ends on the one reading of its stop The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that verdict to the child's ending, so the host never forms a second view of whether the root is gone. Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition, and a failed re-attach passes what its release saw. The end-of-child record can therefore also carry a stop whose root was not seen to go, which nothing ends on yet. * feat(native-chat): the host says it accepts a send before any agent has it The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same string capable clients already send. A client can then tell a host that answers a send at acceptance, and admits a Stop with no writer before a turn starts, from an older one that still restarts the agent inside the send. Additive: an older client ignores a capability it does not know. * refactor(native-chat): an attach never opens a journal of its own The attach adopts the conversation's open journal, which outlives it, so it no longer opens one for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag that told the two cases apart is gone. Tests that attach without a host open the conversation the way a host does. * fix(native-chat): a moved fence resends nothing on a host that accepts first The outbox treated any fence change as a new owner: it dropped the answer of a send in flight, queued that send to go out again under the same id, and unblocked a refused head. On an older host that is how a send the restart refused, unrecorded, gets another try. On a host that records every send before it starts an agent, a fence moves because that start ran, so the same rule resent into every failed start. With a fence stamped on every frame, that became a loop. The outbox now reacts to a fence change only when the host has not advertised that it accepts a send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed start reaches the client as a rejected message it keeps with its Retry. Against an older host, or before one has answered, the outbox behaves as it did. Desktop and paired web share this hook. * refactor(native-chat): a child's end says whether the user or the host stopped it The end-of-child record's cause now tells a user's Stop from the host stopping the child for a cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a user's Stop, as before, and fails the start it was waiting on after a host stop, with the one error row and every queued message rejected, in the stop's reason when it gave one. The reason stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet. * fix(native-chat): a chat whose only work is a queued message is not offered for resume A message accepted while the agent was starting counts as working in the chat, and quit rejects it as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a chat whose agent never had the message. The snapshot now reads only what was handed over. * fix(native-chat): the conversation outlives its agent Opening a chat no longer starts its agent. A conversation is reached through one host accessor that opens its journal at rest, and a send is what starts the agent, through the delivery loop. One idle sweep, every five minutes, stops an agent that has been quiet for thirty minutes and owes no work, then drops an open journal handle that is only a cache. Its record, tab, status row and readers stay. - hold and release are no-ops; hold still builds the host for shipped mobile builds. - The holders, the holds, the release clock and the exit respawn are deleted. - Options, the model list, the goal and the context meter answer at rest; a model pick at rest is recorded as intent for the next start. - Compact, rewind, clear and goal changes start the agent first. A send does too when a rewind is still in doubt after the conversation opens. - Orchestration routes mail and group addresses on ownership (the record plus the chat tab), not on whether the process runs. An open dispatch keeps its worker running. - The restart continuation is a send; Resume all holds each slot until the message is handed over or rejected. - A read error never replaces a loaded transcript, and shows the host's own words. * test(native-chat): type the queued-message fixtures in the resume-offer tests * fix(native-chat): a start that dies while a message waits on it is that message's failed start Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When that start died, its exit wrote the start's error row and left the message queued, so the delivery loop started a second agent into the same failure and wrote a second row. A child's end now records where the conversation's journal stood, and the loop settles a message accepted before a failed start ended with that start: one row, under its key, and no second start. A message sent after the failure still gets a fresh start. * docs(native-chat): say what an attach's open conversation and unconfirmed ids are now * test(native-chat): pin what a failed start settles, and what a resume offer names A view's child that dies while a sent message waits settles that message only when it died starting and no child has taken its place: a proven child's crash, or a second start since, gets the message delivered. The resume offer names the handed-over message, never a newer one still queued. * test(native-chat): the failed-start pins fail on what the message became, not on a timeout * fix(native-chat): a restart offer ends when the chat's agent starts again The offer used to end only when the chat's newest user message changed, because opening a chat started its agent and that start could not be told apart from real activity. Opening a chat starts nothing now, so the host reads the fact it already publishes: a chat's status row goes from not host-owned to host-owned exactly when its agent is started. At that edge the offer and any failure record for the chat are withdrawn, unless the start is a resume action's own (its continuation is the oldest undelivered message). A continuation and a message racing to be first are decided at acceptance: the continuation is refused, quietly and with nothing filed, when any other message was accepted since the restart. A failed continuation start leaves the offer retryable, and each resume action sends its own message id. Deleted: the newest-user-message comparison, its journal reader, the continuation filter, and the failure ledger's own "answered by the chat" check. The marker still carries its message id for one release, so the previous build can read it. * fix(runtime): end a transcript stream when its client unsubscribes Desktop: the IPC subscription controller was dropped as soon as the streaming handler returned, which for most streams is right after it binds. A later runtime:unsubscribe then found nothing to abort, so the host kept the subscriber and derived and sent every publish to a channel no one listened to. The controller now lives until the renderer unsubscribes, resubscribes the same id, or goes away. Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe with the stream's frame id, so the host ends that subscriber and leaves a sibling stream on the same socket running. The direct path now passes the frame id the relay path already passed. * test(orchestration): the preamble's host stub is typed, not cast The preamble send now takes only what it reads of the host, the send, the settlement wait and the record's fence, so its test builds that host with real types instead of `as never`. * fix(native-chat): one fact ends a restart offer: the chat moved on since the restart The offer is live while no other message has been accepted in the chat since the restart and its agent has not proved a start since. The offer list, the resume's reservation check and the continuation's acceptance check all read that one fact, so a message whose start then failed withdraws the offer too, and a stale click finds nothing to act on. The fact is read off the conversation's open handle, which the restart closed, so it is retired durably whenever it may have changed: a message accepted, a start proven. A close and reopen within the same run therefore cannot bring the offer back. A continuation rejected before it reached the agent does not count, so a retry after a failed start still runs. Deleted: the quit-time gate on withdrawal, which changed nothing because the withdrawal and the quit's own offer write share one queue; the per-action "withdrawn" flag and the separate acceptance check it paired with. * test(native-chat): an older build reads the restart offer this build records The offer lives in a file the previous release reads after a downgrade. Pin that against the pinned release's own capsule, and run the lane when the marker or the capsule changes. * fix(native-chat): read a restart offer against where the journal stood when it was taken "Since the restart" was read off the conversation's open handle, which the idle sweep closes: after a reopen, a message the user had already sent looked older than the handle and the withdrawn offer came back. The offer now records the journal position (epoch and sequence) at the moment it is taken, and a message accepted after that position, or a journal on another epoch, means the chat moved on. That is derived from the journal, so it holds across any number of closes and reopens. An older build's offer has no position; only a start withdraws it. Because the message half is now durable, the offer is no longer rewritten in the recovery file on every accepted message; a proven start still writes it, since only the host that saw the start knows of it. * test(native-chat): wait for the listing's retire write before reading the recovery file * fix(native-chat): keep the terminal-backed chat's read error over its local echoes Messages winning over a read error is right for the structured chat, whose read retries and whose messages came from the transcript. The terminal-backed view assembles its list from local echoes too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no error. Only the structured pane now keeps messages over an error. * fix(native-chat): a start retries the exit settlement a failed journal write left owed An agent exit whose journal settlement write failed releases the lease latched until a retry lands. Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried it before the next app launch, and every send was refused. The start the send needs now runs the retry first, where the attach would. * perf(native-chat): answer the owner check without opening the chat Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the answer comes from the session record alone. Reaching it through the accessor opened each resting chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it open for the idle window. It now checks the record and the adapter's support, as before this series, and opens nothing. * fix(native-chat): a read waiting on the session lock opens nothing once quit began The accessor checked for quit before queueing the open, so a read queued behind a session task ran its open after teardown had begun and indexed a journal no teardown step would close. The check now runs at the open itself. * fix(native-chat): read a failed resume's chat before calling it retryable Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the restart, read from its journal. The failure list read it only for a chat already open, so once the idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did nothing. The list now opens the failed chats first, as the offer list does. * test(native-chat): type the provider event sink the settlement test reaches for * fix(native-chat): say the structured read keeps trying only where it does The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an untranslated fallback whenever the read error had no text, and the empty state prefers any message. The view state now leaves the message out, so the structured pane shows that line and the terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an error frame, so it no longer makes the claim. * test(native-chat): await the send's settlement instead of polling for the start The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a loaded machine outran. They now await the host's own settlement of the message. * fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now dated by the resume action. Telling a rejected continuation from the user's own message read the operation ledger, whose rows expire after about a day; after that a failed resume stopped being retryable. The offer now records the continuation each action sends on its own capsule entry, bounded to the newest 16, so the ids end with the offer. The ledger read is deleted. * fix(orchestration): route no mail to a structured worker its orchestration released A structured worker is routed on ownership, and a resting worker's lease is released, so ownership held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one restarted its agent. Routing now also reads the orchestration's own resource row: once it is released, direct mail, group addressing and worker-show's addressable answer drop the worker, as they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored. * fix(native-chat): a failed retry names the user's prompt, not Orca's continuation A resume's continuation is written to the chat before its start, so after a failed attempt the chat's newest user message is that rejected continuation. A second failure then showed Orca's own restart text as the chat's prompt. A retry now keeps the prompt its first failure named. * fix(orchestration): read the released row optionally, as the authority does worker-show's observation called the row lookup directly, which a runtime double without it threw on and failed the structured tab-retirement release. * fix(native-chat): the status bar drops a restart offer the chat moved on from The renderer re-read the host's restart offer only when a failed chat showed activity, so after a message withdrew a pending offer the host answered no chats while the status bar kept counting one, and clicking it opened nothing. The same watch now covers pending offers: a status change in an offered chat asks the host again, once. * test(native-chat): a roster of idle or finished children does not keep an agent awake The sweep reads owed background work through the shared child-work liveness that upstream's release clock adopted; a child that went idle or finished is not work the agent still owes. * fix(orchestration): a task dispatched into a resting structured worker keeps it running The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's process incarnation now counts, derived from the existing rows. * docs(native-chat): comments stop describing the hold this PR removed Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it. Comment-only. * fix(native-chat): a restart offer keeps the start its own continuation made Whose start ended an offer was decided at read time, from whether the offer's continuation was still the queued message. Once the provider refused that continuation, the child it had started read as someone else's start, so the offer ended and its failure showed no Retry. The delivery loop now records which queued message a start is for on the in-memory child, and the child's end carries it; the offer counts a start as its own when that message is one of its continuations. * fix(native-chat): an agent gets a full idle window after its owed work ends The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can read done before the lead's wake-up turn writes anything, and stopping in that gap loses the wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full window afterwards, as the release clock it replaced did. * test(claude): the options-read fixture runs a live child The fixture marked its conversation running with a hasProviderChild field the session type does not have, so the read took the at-rest path and refused a session with no record. It now carries a child, which is what the read checks. * test(native-chat): host tests reach its collaborators through a typed seam The rest-test rig and three test files read the host's private members with Reflect.get and cast the result. The host now exposes one test-only accessor, collaboratorsForTests(), and the subscribers class a subscriberCountForTests() beside its existing retainedActivityCountForTests(), so the tests are checked against the real types and the casts are gone. * refactor(orchestration): one owner answers a structured worker's custody Routing, group addressing, worker-show and the idle sweep each composed their own reading of whether orchestration still holds a structured worker, so each new obligation or retirement state had to be added to every reader. structured-worker-custody now derives both answers from the worker-terminal list state coordinators see in worker-list: addressable is owned and not released, and owed work is an active custody or an unsettled task dispatched to the same incarnation. The owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests. * refactor(orchestration): owed work is an open dispatch on the worker's incarnation A supervised worker's own dispatch context stays open exactly while the worker is active, so the separate active-custody branch only repeated it. Owed work is now one fact, which also states the policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are written once at the top of the module. * fix(native-chat): a restart offer knows its continuations by a tag in their id The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running action's id in memory. Both could disagree with the journal: past the cap an old rejected continuation read as the chat moving on, and a crash during a retry restored the failure's older entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer (its teardown and chat), then the action's own part, so any continuation of this offer, queued or rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own continuation the start was for, read against the stored marker. * test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait The tui-idle probe reads through readTerminal, which now awaits the structured worker check before the PTY read, so the probe's snapshot request starts a microtask later. vi.waitFor missed it on its first check and polled again at 50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot then resolved after the wait had already timed out, so the test passed without judging it, and the rejection landed before any handler was attached. Vitest reported that as an unhandled error and failed the shard. Polling every 1 ms sees the request within a few ms, so the snapshot is judged while the wait is still pending. * fix(native-chat): a view never restarts a chat whose last start failed A Claude chat whose CLI exits during startup left one red row per start, and every time a view bound to it (the chat opening right after its create died, or the user switching back to it) the hold started the CLI again, so the same launch-failure row repeated. Only a send retries a failed start now, the same rule provider-exit recovery already applied; the rule lives in one predicate the hold, exit recovery and the delivery loop share. * fix(native-chat): the idle sweep reads owed work every tick Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it refreshed the clock at most once a window. Work that ended just before the next read left the agent to be stopped at that read, moments after the work ended, which is the gap the refresh was meant to cover. The sweep now reads owed work on every tick for a started agent, so the window always runs from the last tick that saw work owed. * fix(native-chat): a continuation handed to the agent stays sent The offer read its own continuation as not reaching the agent while its dispatch was pending, which also covered one already handed over and still unanswered. When the wait for that answer ended first, the failure it filed read as retryable, and a retry sent a second continuation to an agent that may have acted on the first. Only a continuation still queued, or rejected, is now read as unsent. * test(native-chat): start the child the loop waits on with an attach, not a second view A view no longer starts a child whose last start failed, so the R2 case that waits on a child started since the failure now gets that child from a client attach, the one non-send starter left. * fix(native-chat): settle a gone generation's turn wherever a conversation opens A send that opens a chat this process had not read yet (after a crash, from a phone or the CLI) went through the delivery open, which never settled what the dead generation left running; only the read restore and a successful acquire did. When the send's start then failed, the turn stayed running for every reader. The settlement now runs in the one journal open, at the crash boundary, for every opener except an acquisition, which settles from the evidence it read before its reserve; the read restore's separate step is gone. * test(native-chat): prove the next child's start settles the turn an earlier child left The R1 case lost its only settlement assertion when the latch it checked was deleted. It now seeds the running turn the earlier child left and asserts it ends at the exit's receipt, with the exit's row, before the message is handed to the new child. * test(native-chat): count a failed start's rows by row, not by text Comparing the set of texts passed when two different rows carried the same words, which is the duplicate the test exists to catch. * test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget * test(native-chat): the interrupted create's own retry continues again The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its own operation id, with a fresh start whose result nothing read. That fresh start passes with the released-reservation continuation deleted, so the case the fix exists for went untested. The retry and its assertion are main's again. * docs(native-chat): three comments that still had views starting agents A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction left alone would refuse every send, so no agent would ever start to finish it; and a current host raises the unattached read refusal only once quit began, with the attach window belonging to an older host. * test(native-chat): pin the open's and the send's start and row counts, however the view binds Opening a fresh chat whose starts fail makes one start and one row, with two views bound before or after the create's child died; one send makes one more of each. * test(native-chat): a reader's open settles the turn a failed exit settlement left running An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle sweep closed, and a read that opens the chat before the restart restore reaches it. * test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner A subscription reads the conversation before it returns, so under load the two views took longer than the create child's 300 ms start, which then exited before the test checked that it had not. The child now takes a second to fail. * fix(native-chat): settle a gone generation's turn at every open but an acquisition's The journal open skipped the settlement whenever the lease read reserved or live, to leave an acquisition's own open to the acquisition. But a lease a crashed process left in recovery also reads live, until the next acquire resolves it. A send that opened such a chat, from a phone or the CLI after a crash on a host that could not prove the old owner gone, skipped the settlement; when its start then failed, the dead turn stayed running for every reader. The acquisition now says it is the opener, and every other open settles, whatever the lease still claims. * test(native-chat): hold the create's start open until the views bind The "view binds while the create is still starting" case gave the create a 300 ms head start and asserted the views bound before it died. On a loaded runner the holds took longer, the create's exit landed first, and the case failed its own precondition. The create's initialize now waits on a gate the test releases once the views are bound. * test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the dead process's lease still reads live. The open settles the turn it left running anyway, and the restore that follows finds it settled. * refactor(native-chat): drop the composer's second error formatter After the merge with main, every chat write in the composer path reports its failure as a typed outcome worded by the refusal-notice table, so the send's catch sees only a local throw. The {code, message} formatter this branch added for it has no payload left to format, and its claim to be the one way a chat words a failure is no longer true. The composer send is main's again. * test(native-chat): pin the reason on a message rejected while its chat was closed The reopen test checked only that the message reads as not sent; it now also checks the Retry row carries the host's reason. * docs(native-chat): drop the removed dispatch hold from six comments A worker's session no longer takes a dispatch hold, and no release clock rests a chat by visibility; the agent-launch comments, the abandon test, the teardown test and the refusal census still said so. * test(native-chat): rest the owner-status chat through the idle sweep, not a hold The activation-gate test from #22808 put its chat at rest by holding and releasing it, and passed the release-clock grace. This branch deleted both, so the case threw before it reached its assertions. It now moves the host's clock past the idle window and lets the sweep stop the agent and close the conversation, then asserts the same owner answer and activation gate. * fix(native-chat): show the structured pane's retrying line when a read fails The read transport always hands the pane the host's words, so the error state's "Orca keeps trying to load it" line, which showed only when there were none, was never seen: the pane showed the host's text twice, as its subtitle and on the status line under it. The structured pane now always says its read keeps retrying, and the host's text stays on the status line. The terminal-backed chat is unchanged. * test(native-chat): wait for a send's background start before the refusal oracle removes its store An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals. * fix(native-chat): a start a message waited on gets one failure row, the delivery loop's When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice. The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
5e5f4f6603 |
fix(native-chat): keep a Codex ask's questions in the order it asked them (#23502)
* fix(native-chat): keep a Codex ask's questions in the order it asked them Codex journals every question of one ask in a single write, so the questions share a timestamp. The transcript list sorted rows by timestamp and broke ties by row id, and a question's id ends in the question id the model chose, so answered, cancelled and still-pending rows of one ask came out in the alphabetical order of those ids. The transcript projection now breaks timestamp ties by the order its source gave: the journal's order for structured sessions. The session assembler keeps its id tie-break, since the sources it merges share no order of their own. * refactor(native-chat): name the id tie-break comparator for what it does Two shared comparators differed only in whether they break timestamp ties by id, under near-identical names; spell the id tie-break in the name. * fix(native-chat): order structured chat rows by their journal position The desktop list and the host's conversation outline sorted structured rows by timestamp. The journal's contract is that the sequence orders the timeline and the timestamp is the provider's clock: a Codex ask writes all of its questions in one write (one sequence, one timestamp), and a row recovered after a crash carries an earlier clock at a later sequence. The host now records each item's place within the write that created it, keeps it across revisions like the sequence, and sends it as an optional field. Rows projected from the journal carry that position, and both the desktop list and the outline order journal rows by it. Rows outside the journal keep their rules: rank for the streaming and pending tail, outbox sends after every journal row, and terminal-backed chats keep time then id. * fix(native-chat): keep a refused send at its journal place, and keep list positions off worker reads A send the host journalled before the provider refused it is shown from the outbox, and it sorted after every journal row, so it dropped below whatever the agent wrote after it. It now takes the journal position of the submission the host recorded. Structured worker reads and their archives projected journal rows through the same projection, so they returned the list-only journal position on every message. The worker payload bound now drops it. * test(native-chat): a failed restart's row draws below the message it failed Since a message is accepted before its delivery starts the agent, a restart that fails is journalled after the message, and the refused message keeps that journal place in the chat. Both host paths now pin it through the chat's own projection: a start the host could not make, and a restarted child that exits before proving its start. Folds the journal reducer's batch item write onto fewer lines, which the merge of main pushed past the file's line limit. |
||
|
|
9179b93ebf |
ci: reduce repeated runner work and validate affected-test selection (#23540)
* ci: stage heavy checks and measure affected-test selection * fix(ci): exercise the real Git boundary in unit selection planning * Harden review cancellation and CI demand reporting |
||
|
|
fae0ae7a46 |
perf(ci): shard the anti-slop audit across processes instead of one JS runtime (#23543)
config/oxlint-anti-slop.json turns every native oxlint category off and runs its rules through jsPlugins, so oxlint's threaded Rust engine does no work and the pass is one JS runtime per process. Measured, it does not scale with --threads: 11.68s at 4 threads against 12.38s at 16. Parallelism has to come from more processes, so the audit now splits its file set across them. Locally, 11.72s single pass against 3.83s at 4 shards (3.1x) and 2.91s at 8. Sharding is sound because every anti-slop rule is a single-file analysis; the only mutable module state is a WeakMap keyed on each file's own Program node. Verified by running both shapes with all 18 rules enabled: 392,398 findings from one pass and from the shard union, identical as sorted multisets. |
||
|
|
67581dd090 |
perf(ci): stop the orcad smoke idling 15s and the static job fetching mobile packages it skips (#23541)
The shutdown race in the orcad terminal smoke never cleared its losing timer, so the process sat on a live 15s timer after PASS had already printed. Measured locally: 21.4s -> 6.52s, with the round trip and the shutdown assertion intact. Static analysis also asked for the mixed root+mobile pnpm store (537 MB, 8.6s to restore) on every run, while installing mobile dependencies only when the diff needs them. Most runs paid 216 MB for packages they never linked. |
||
|
|
45f3512a33 |
feat(agents): add first-class DeepSeek Harness (dsh) support (#22468)
* feat(agents): add first-class DeepSeek Harness (dsh) support Register DSH as a supervised Orca agent: catalog entry and detection for its dsh-tui profile, status/question hooks through DeepSeek's own Claude-Code hook bridge, composer-ready prompt delivery, session resume, headless Source Control AI, and title identity that no longer collides with Gemini's. * fix(dsh): reach Orca through DSH's credential scrub and stop reading its title as Gemini DSH runs command hooks through its own shell executor, which drops every env var whose name contains KEY, TOKEN, SECRET or PASSWORD — taking ORCA_PANE_KEY and ORCA_AGENT_LAUNCH_TOKEN with it, so every hook exited without posting. Mirror both onto scrub-safe aliases at spawn and restore them at the top of the DSH hook script. Its title collided too: DSH rests on the same glyph Gemini works on, so a resting DSH pane was relabelled Gemini CLI and reported working forever. Defer both the Gemini classifier and the title status detector on DSH's whale, in the base module both copies of that classifier read. * test(mobile): repin the session-route closure for the DSH agent icon * fix(dsh): address review — never splice user rows, cover remote panes, keep the diff off argv - findManagedDshPatchRegion paired an orphan start marker with a later block's end, so a truncated write made install/remove delete the user's own rows. Pair each end with the nearest preceding start; regression test fails without the fix. - The relay PTY env builder never applied the scrub-safe aliases, so remote DSH status silently never appeared even with the remote hook installed. - Source Control AI sent the whole diff on argv; send it over stdin with DSH's '-' marker. - dsh-tui/dst already chose the interactive profile, so a workspace folder named 'web' or 'plugin' no longer marks a live agent pane non-interactive. - Isolate USERPROFILE as well as HOME so a Windows run cannot edit the real home. - Drop the duplicate README badge and revert an incidental doc reformat. * refactor(dsh): share the managed-hooks reader and tighten the new modules Reuse before reimplementing: readManagedDshHookEvents was a near-verbatim copy of Muse's, with byte-identical private helpers. Both now call one readManagedHookEventsFromJson. Also: one readTextOrAbsent instead of two spellings of the same read (dropping an existsSync TOCTOU), one status() builder instead of four inline literals, rmSync(force) instead of exists-then-unlink, and a redundant empty-string guard before JSON.parse. The patch-file transforms lose their index juggling for a predicate plus a filter. * fix(dsh): refuse a flow-style patch file, keep its mode, and stop the relay inheriting a pane - applyManagedDshPatch matched only an exact `[]`, so `[] # keep empty` or a non-empty flow sequence got a block entry appended after it — invalid YAML that would leave DSH unable to load the user's own patch layer either. It now strips the token from an empty sequence (keeping a trailing comment) and returns null for a non-empty one; install reports that and changes nothing. - The patch rewrite dropped an owner-only file to the umask default (CWE-732); pass preserveMode. - The relay PTY env never dropped inherited pane identity the way the local and daemon builders do, so a spawn that specified none could inherit the relay's own and every agent's hook would report against that pane. * fix(dsh): keep the flow-style refusal in every status read, and scope the mode test to POSIX A refused patch file carries no managed region, so getStatus() fell through to a bare not_installed with detail null — the actionable 'rewrite it as a block sequence' message only ever reached the one-shot install() return. Export the predicate and check it first, behind one shared message constant. The owner-only mode assertion cannot hold on Windows, where chmod only toggles the read-only attribute and mode & 0o777 reads 0o666 for any writable file. * docs(readme): restore the DeepSeek Harness badge lost in the rebase * test(mobile): repin the session-route closure to the measured 4221 Measured, not derived: 4220 without the DSH icon entry, 4221 with it. Two of the three modules above main's 4218 pin are not this change's — they arrived with the mobile work after #22570 and were never repinned; the changelog records that split explicitly. * fix(dsh): settle tui-idle on the agent's own hook, so supervised workers see it ready Reported by a tester on the adhoc build: `terminal wait --for tui-idle` ran to its 90s timeout against an already-ready DSH composer, so a supervised worker never sees the agent as ready. Every existing tier reads the title, and DSH deliberately carries no title status: its rest prefix is Gemini's working glyph, so the detector reports none. A fresh first-party `done` is better evidence than any title anyway — it is the agent's own account of its own turn, and normalizeDshEvent drops subagent events, so it is the lead's. Scoped to DSH: for agents whose hooks report child turns, a mid-turn `done` is the #6011 class this file prevents. * test(daemon): record the DSH transcript's true-colour I2 divergences Adding the dsh-tui capture to __fixtures__ enrolled it in the serialize replay sweep, where it reports 10 I2 divergences and failed the unlisted-transcript default of 0. Every one is the same shape — visible-grid row=0, a 24-bit background the round trip does not restore to default — which is DSH's whale intro painting whole rows of true colour. Verified as an upstream limitation rather than a regression by replaying against the previous build (build-serialize-addon-at-ref.mjs --ref origin/main): I1 and I3 both hold. * fix(dsh): return the new tui-idle verdict from the first-party done lane Main refactored isTuiIdleSatisfied into evaluateTuiIdle, which returns a verdict rather than a boolean. The DSH lane still returned `true`; it is tier-1 positive evidence, so it returns READY_STRONG like the title/body lane above it. Re-verified the regression test still fails without the lane. * test(relay): pin the scrub-safe pane-identity aliases on the relay spawn path The relay builds a remote pane's env itself, so the alias mirroring there had no test: removing the call left every suite green while remote DSH status silently vanished. Both cases fail without it. * docs(dsh): point the hook service at the integration reference The reference doc had no inbound link from anywhere in the repo. |