Commit Graph
904 Commits
Author SHA1 Message Date
Jinwoo Hong 9f4311598f fix(codex): trust the worktree Codex starts in, not a guessed repo root (#23937)
* fix(codex): trust the path Codex checks for bare-repo worktrees

Codex keys a linked worktree's trust on the main checkout only when that
checkout's .git leads back to the common git dir; otherwise (bare repo,
--separate-git-dir) it keys on the worktree itself. Orca always wrote the
main-checkout key, so Codex showed its trust prompt and worker-start
failed at agent_readiness.

Mirror trust.rs exactly, and pin it with a real-binary contract that
runs in the existing Codex contract CI job.

Fixes #23847

* fix(codex): trust the worktree path itself instead of mirroring trust.rs

Codex looks up the cwd's own [projects] entry before any repo root
(config_toml.rs get_active_project, loader decision_for_dir), so trusting
the workspace realpath satisfies every git layout. Drops the
resolve_root_git_project_for_trust mirror: simpler, cannot drift from
Codex, and never widens trust past the folder Orca launched in. Cost is
one config entry per worktree; entries older Orca wrote on main
checkouts stay valid.

The six real-git layout tests now assert the workspace key and, under
the contract, that real Codex starts workspaceWrite for each. The
contract probes the binary version once and fails at load when required
but missing; its CI path filter now includes config-toml-trust.
2026-09-29 20:27:02 -04:00
Neil fed1eca486 test: stop restating internal tuning constants, keep the ones that are contracts (#23950)
Removes ~74 assertions of the form `expect(SOME_CONSTANT).toBe(<literal>)` where
the literal is an internal tuning value — a timeout, retry count, debounce
interval, cache TTL, circuit-breaker window, Tailwind class string. Those cannot
fail for any reason a user would notice: they fail only when someone deliberately
changes the number, and then the test is simply updated. They are copies of the
declaration.

The same pattern is NOT junk when the exact value is observable outside this
process, so those were deliberately kept:
- terminal byte contracts: `\r`, `\x03` ETX, Kitty escapes, `\x1b[?1;2c`;
- wire and capability values: `agent.launch.v2`, protocol 3 / min-compatible 2,
  daemon per-feature boundary versions (a daemon survives app updates, so those
  pin what an old field daemon may be trusted with), relay header tokens;
- security invariants: the `127.0.0.1` bind default, an empty iframe `sandbox`;
- values external processes read: exit code 78 (EX_CONFIG) and exit code 3
  (systemd `RestartPreventExitStatus`), `ORCA_AGENT_SESSION_SPAWN_TOKEN`,
  `npx skills …` commands users paste, on-disk journal schema versions,
  the `orca_<hash>` filename prefix the fish sweeper matches;
- third-party names: expo-router's `unstable_settings` / `ErrorBoundary`,
  iOS Safari's 16px zoom threshold.

Where a case asserted a relation rather than a literal — `A < B`, a sum of parts,
a cap compared against a sibling budget — the relation stays and only the literal
went.

Test-only changes: no production file is touched and no test file is deleted.
2026-09-29 17:06:09 -07:00
Brennan Benson 59b746ff3c feat(native-chat): one structured-chat journal database per host, owned by one process (#23613)
* feat(native-chat): one structured-chat journal database per host, owned by one process

Every structured chat on a state directory now lives in one SQLite file,
agent-session-journal.db, opened once by the process holding
agent-session-journal.owner: an empty SQLite file whose held BEGIN EXCLUSIVE is a
kernel byte-range lock, refused while another process holds it and released when
the holder dies.

- Stores own no connection: the per-chat handle, its close contract and the
  close-retry registry are gone; closing a conversation drains its writes, and the
  one connection closes last at teardown.
- The owner lock is taken at runtime start, before orca-runtime.json is written;
  a process that does not own the chats is not published and refuses every
  structured request with journalUnavailable and words that say what to do. It
  retries the lock with backoff and runs the full install once it holds it.
- A journal that will not open fails the host install: every chat says "Unable
  to load this chat." (journalCorrupt), and nothing is renamed, deleted or
  rebuilt. A newer build's database is refused and left byte-identical.
- An append is one INSERT. The listing status is a column, written after the
  rows it describes and keyed by (epoch, sequence).
- A per-chat journal from an earlier build is copied in verbatim (epoch UUID and
  every sequence) on that chat's first open, and its directory is retired only
  after the copy commits.
- auto_vacuum = INCREMENTAL, with freed pages handed back in bounded steps after
  every delete.

* perf(native-chat): key journal rows by block so one chat's rows sit together

Each chat's live epoch owns a block of row ids, block * 2^32 + seq, so a chat's
rows share leaf pages with nobody else's, a replay is one range scan, and
replacing or rewinding a chat deletes one contiguous range. Measured on the
largest real chat (61 MB) beside 19 interleaved peers: 39 ms and 7.5 MB of WAL,
against 214 ms and 102 MB for a (session_id, epoch, seq) key.

- Ids are computed in Number arithmetic, never bitwise. A sequence is refused
  outside [1, 2^32) and a block at 2^21, which keeps every id below 2^53.
- A replace, roll or import allocates a fresh block, moves the chat's pointer,
  and deletes the old block in the same transaction, so no orphan block exists.
- The listing status write moves into its own writer beside the column.

* feat(native-chat): copy a chat's per-chat journal again when an older Orca wrote it after a downgrade

A per-chat journal.db that reappears after its chat was copied in is the newer
history: an older build, run after a downgrade, attached the chat and wrote it.

- journal_imports records the (epoch, tip) each chat was copied from, in the
  copy's own transaction. A file already copied is never copied again, across
  any number of restarts after a failed rename; a file that differs always is.
- Newest writer wins, per chat, with a row saying the chat was continued in an
  older version of Orca. When both builds wrote past the recorded tip under one
  epoch, the copy takes a fresh epoch, so readers reset instead of skipping rows.
- Each copied directory retires to its own .imported-<epoch8>-<ms> name, so a
  second downgrade and re-upgrade never collides with the first.

* test(native-chat): fixture deps match the host journal database shape

Attach-flow and reconcile-attach fixtures stop passing a journal database those inputs do not take, and host and restore fixtures pass the one they now require instead of the removed journal root.

* test(native-chat): state why the runtime-state fixtures' existing casts are safe

* fix(native-chat): start up normally when this process cannot open the chat journal

A process refused the chat journal, because another Orca owns it or because its own journal will not open, failed startup restoration: the window booted in degraded no-save mode and a paired phone could not list any tabs. Startup restoration now treats the refusal structured requests are getting as having no structured host; terminals, tabs and saving go on, structured requests are still refused by the gate, and the install is retried on the next one. Any other install error fails startup as before.

* test(native-chat): name the owner-lock sweep test after the two sweeps it runs

* fix(native-chat): session history and terminal resume work while chats are refused

Session history (listing and preparing a resume) and a terminal typing a resume command only check whether a structured chat owns a provider session. In a process refused the chat journal they failed outright. They now take the refusal chats are getting as having no structured host, the same treatment startup restoration gets, through one shared helper; any other install failure still fails them. Chat requests keep the gate's refusal.

* test(native-chat): the first-work rename's fake journal saves the listing status

* fix(native-chat): open a chat whose per-chat journal file never got its schema

A crash between creating a chat's journal.db and creating its tables left an empty or schema-less file. Each chat used to open that file as an empty chat; the importer instead refused the open as "try again" forever. A file with no journal_sessions table is now read as never written, the same as one with no rows. A file that is not a database, or whose read fails, is still refused.

* fix(native-chat): let the event loop run between chats during startup restore

Opening a chat's journal is synchronous SQLite now that no per-chat directory
is created first, so the restore of every visible chat ran as one main-thread
task. Each chat now waits for a macrotask before it opens.

* fix(native-chat): import a per-chat journal in bounded batches

The one-time copy of an earlier build's per-chat journal ran as one
transaction, which blocked the main thread for 650 ms on the largest chat.
Rows now copy 512 at a time, each batch its own transaction, yielding to the
event loop between batches. The rows go into a block journal_import_blocks
reserves, which no reader follows and no other chat is allocated; the last
batch publishes the chat's pointer, repair marker and import marker together
and releases the reservation. A copy that stops midway leaves only that
block, which the next open clears and copies again. Two opens of one chat
import one after the other.

* fix(native-chat): refuse chats when the owner lock file cannot be opened

A lock file that is not a database, or cannot be opened, made the claim throw
before any refusal was recorded, so startup restoration failed on every
launch. The claim now sits in the same try as the database open and records
the same typed refusal.

* fix(native-chat): finish reclaiming pages a delete frees during a running pass

A reclaim pass ended as soon as the freelist stopped shrinking between steps,
so a second delete that freed more than one step's worth mid-pass ended it
early and left those pages on the freelist. A pass now ends only when a step
itself frees nothing, or the freelist is empty.

* test(native-chat): desktop session history is served while chats are refused

* fix(native-chat): a send to a chat holding a newer Orca's rows says to update

A chat opened read-only because a newer Orca wrote rows to it answered a send
with the generic write failure. It now refuses the way a database a newer Orca
wrote does, with the same reason and words.

* fix(native-chat): keep chat tabs while this process cannot list its chats

A process whose chats another Orca owns, or whose chat journal will not
open, has no structured host. Its session-tabs inventory still answered,
with no chat rows, and the renderer read that as "every chat was closed":
it removed the restored chat tabs and the next session save persisted
their placement away.

The inventory now says `agentSessionsUnverifiable` when the last tab
restore ran with chats on disk but no host to list them. The flag is set
and cleared at the per-client projection point beside the client-hosted
page hold, and the restore is memoised only once a host answered, so a
later lock takeover or journal open republishes the chats and clears it.
The renderer keeps agent-session tabs, and keeps cancellation tombstones,
against an inventory that does not affirm its chat set.

* fix(native-chat): say chats are open in another Orca, with this process's way past it

A process refused because another Orca owns the profile's chats sent the
generic `journalUnavailable` reason, so current desktop and phone surfaces
said "couldn't open this chat's history right now. Try again." — a step
that never helps while the other Orca runs.

The refusal now names its own reason, `journalOwnedElsewhere`, with the
refused process's kind (dev desktop, packaged, orcad) as a fact. Each kind
gets its own step: quit the other Orca, or give this one its own profile
(ORCA_DEV_USER_DATA_PATH) or data folder (ORCA_USER_DATA). The sentences
are added to the shared notice copy, the desktop catalogs in all six
locales, and the boot catalog.

A client that predates the reason reads it as none and keeps the code's
words; an unknown kind reads as the packaged app's step. The `message`
released clients print is unchanged. Which requests refuse does not change.

* fix(native-chat): restore chats on taking ownership, without a list to ask

A refused startup kept its hostless result, so after the owner quit this
process never installed a host, never reconciled restart leases, and kept
telling clients it could not list its chats until a desktop chat request.
Taking the lock now reruns startup restoration once and pushes the chats.

* test(native-chat): a navigation reply says chats are unverifiable while refused

* fix(native-chat): retry a refused owner lock at most every 5 seconds

The lock frees as its holder exits, but a refused process only learns that
on its next retry, and the 30 s cap left a second Orca refusing chats for up
to half a minute after the owner quit. One retry is an open and BEGIN
EXCLUSIVE on an empty file.

* fix(native-chat): install before deciding whether a takeover must republish chats

A list that landed on the refused startup after the lock was taken finished
after the takeover had already checked, so nothing republished. The takeover
now installs first, waits for any restore in flight, and restores only then;
the restore that clears "cannot tell" pushes the frames itself, so a list
that heals the inventory first reaches subscribers too.

* fix(native-chat): no takeover lands a host after the runtime stop

Quitting cancelled a refused claim's retry only at its end, so a retry firing
during the stop's awaits took the lock and installed a host the stop never
tore down, and the lock was then released under an open journal. The stop
now cancels the retry first, keeping the refusal, and repeats its teardown
while an install that began during it (a takeover already under way) is
pending, so no journal connection outlives the lock.

* fix(native-chat): show a thrown refusal in its own words, not its code

A refusal the host throws reaches the client as an RPC error whose message is
the bare code; its reason and facts ride only in the error's data, which no
client read. The chat pane's status line therefore printed
agent_session_journal_unreadable, a send took the bare "not sent" path, and
other writes said the outcome was unconfirmed.

One shared reader, agentSessionThrownRefusal, now reads the refusal from the
error data. A failed history read shows the refusal's read-history words, a
send keeps the refusal behind its Retry exactly as a returned refusal does, and
the other writes (desktop and phone) name the refusal instead of doubting the
outcome. The phone's read failure goes through the same reader.

* fix(native-chat): log a failed journal open once per distinct failure

Every chat request retries a journal open that failed, which is intended, but
each retry also logged the failure with its full stack: a junk database file
logged the same "file is not a database" error 189 times in a minute. The open
now logs a failure only when its code and message differ from the last one
logged, and forgets it once an open succeeds. The retry is unchanged.

The open moves to its own module beside the runtime, which had no room left.

* fix(native-chat): restore lists a chat from its per-chat file and copies it on first use

Startup restore opened every restored chat, and that open copied the chat's
per-chat file into the host database, so the first boot after an upgrade paid
the whole one-time copy before the chat list appeared.

A restore open now reads a chat that is still in its per-chat file straight
from that file, read-only, with the importer's own reader, and closes the file
before moving on. That read drives the listing, the status row and the
restart offer, as it did when every chat had its own file. The copy becomes
owed work on the chat's write queue: it runs before the chat's first write,
and a reader that reaches the chat awaits it. A chat the host already holds,
or that was copied before, still opens through the import and its reimport
rules, and so does a file whose read needs a repair written.

* fix(native-chat): no host stays registered after a stop an install spanned

Each teardown pass clears the registered host before it awaits an install in
flight, and that install registers its host when it finishes. The pass then
tore the host down but left it registered, so a request after the stop was
served by a host whose journal was closed. The stop now clears the slot once
its passes are done.

* fix(native-chat): checkpoint the journal with a full flush on macOS

synchronous = FULL fsyncs each commit, but macOS fsync leaves the drive cache
unflushed, so FULL alone does not survive a power loss there. With
checkpoint_fullfsync, each checkpoint uses F_FULLFSYNC; elsewhere it is a no-op.
The comment that said FULL alone was enough is corrected.

* fix(native-chat): delete a per-chat journal once its copy verifies

An imported chat's per-chat file was kept under an `.imported-*` name, which
doubled the disk its history takes. The copy now reads back from the host
database before it is published: its items, submissions, epoch and tip must
match the file's. Only then does one transaction publish the chat with its
import marker, and the file and its WAL files are deleted, the directory too
when nothing else is in it (a pre-SQLite transcript there is kept).

A copy that does not match is never published: the file stays, the chat is
refused as unreadable ("Unable to load this chat."), and the mismatch is logged
once. A file left behind by a failed delete or a crash matches the marker, so
the next open deletes it rather than copying it again; a file an older build
wrote after a downgrade still differs, and is still copied again.

* fix(native-chat): verify an imported chat a batch at a time

The check that a copied chat reads back as its per-chat file folded both whole,
each in one synchronous task: over half a second on the largest chat. Both
reads now go a batch at a time between turns of the event loop, like the copy
itself, and count rows as well, so a copy that lost a row with no item in it
is caught too.

* fix(native-chat): restore reads a chat's per-chat file a batch at a time

Restore folded a chat still in its per-chat file in one task, so the largest
chat's file held the main thread for about half a second at startup. The fold
now takes the file a batch of rows per turn of the event loop, into the same
fold a replay uses, and nothing reads it before it is done. The file is still
closed before restore moves on.

* fix(native-chat): end a per-chat copy on a turn of its own

A chat's first open ran the copy's last steps (the verified publish and the
per-chat file delete) and the replay of what was copied in one task. The copy
now yields before it returns, so the replay, which every open runs, is a task
of its own.

* fix(native-chat): commit a per-chat copy's batches without an fsync each

Each 512-row batch of a chat's first-use copy committed under synchronous =
FULL, so a large chat paid one fsync per batch, about a quarter of its first
open. The batches now commit under NORMAL, set and restored in the batch's own
task so no other chat's commit runs under it. The publish that makes the copy
visible still commits under FULL, and under WAL that sync makes every earlier
batch durable with it. A crash before it leaves only the unpublished block,
which the next open clears and copies again.

* fix(native-chat): roll back a chat journal transaction whose COMMIT fails

The shared connection's transaction rolled back only when its body threw. A
COMMIT that failed left the transaction open, so every later write, for any
chat, failed with "cannot start a transaction within a transaction", and reads
saw rows that never committed. Under the unsynced copy the failure also tried
to restore the sync level inside the open transaction, which SQLite refuses,
so the caller got that error instead of the COMMIT's.

One transaction helper now covers the body and the COMMIT, rolls back whatever
transaction survives, and rethrows the original error. Schema creation uses it
too. If that ROLLBACK fails as well, the connection is marked stranded: each
later use retries the ROLLBACK, and until one goes through every chat gets the
same "history unavailable, try again" refusal a journal that will not open
gives. The rollback that frees it also restores the FULL sync level.

* fix(native-chat): keep the chat journal connection until its close succeeds

Closing the journal dropped its connection handle before closing it. A close
that failed left the database reporting itself closed with the connection still
open, so the stop that retried the teardown found nothing to close and released
the owner lock over a live connection.

The handle is now dropped only once the close succeeds. A failed close keeps
the runtime pending and the lock held, and the next stop closes that same
connection before it releases the lock.

* fix(native-chat): publish the runtime only once it holds the chat journal lock

When this process could not open the owner lock file at all (a permission
error, or a file that is not a database), the runtime counted that as owning
the chats and wrote orca-runtime.json. That overwrote the real owner's entry,
so the CLI was sent to a process that cannot serve its chats.

A claim that throws is now refused like one another process holds: the runtime
starts but does not publish, the claim's existing retry keeps asking for the
lock, and discovery publishes once the retry takes it. Chats still get the
refusal for the failure itself, and startup restoration reruns on the takeover
the same way it does after another owner quits. A sole process whose lock file
never opens is not found by the CLI until it does.

* fix(native-chat): keep a chat's history when an older build started it over

The first copy deletes a chat's per-chat file, so an older build run after a
downgrade finds no file and starts the chat from nothing. On the re-upgrade that
fresh file was copied in as the newer history, replacing everything the shared
database held for the chat, and then deleted.

A file whose epoch is not the one last copied and that opens with
`session_created` is now kept: neither copied nor deleted, and the chat keeps
the history it has. A file that carried the copied epoch on is still copied
again, as before.

* test(native-chat): pin which chats startup restore copies

Restore copies a chat still in its per-chat file only when restore itself has
to write to it: settling what the last run left open, here a running tool call
or a send handed over and never answered. Every other restored chat stays in
its file until its first use.

* test(native-chat): pin the copy wait on a read that opens a chat restore opened

A read queued behind restore's open of the same chat reaches the conversation
through its own open rather than the listing. It must still wait for the
owed copy, or it reads the chat before its history is in the one database.

* fix(native-chat): record a set-aside per-chat file so no later open reads it

Setting aside a file an older build started over is decided once and kept in
the new `journal_set_aside` table (schema 2, additive), with the file's epoch
and tip as they were. Every later open of the chat skips the file without
opening it, across restarts and after the older build writes more to it:
anything written there grows from that build's own start, never from this
build's history.

The best-effort delete moves beside the per-chat file reader.

* fix(native-chat): set aside any per-chat file at an epoch this build never copied

A chat's per-chat file is deleted once its copy verifies, so a file that
reappears at another epoch was never this build's history, whatever its first
row says: an older build started the chat over, possibly rewinding it after
(`handle_forked`), or rolled the epoch of a file whose delete had failed.
Copying any of them would replace everything the chat holds, so each is set
aside. Only a file still at the copied epoch is copied again (it grew) or
deleted (it did not). The first-row check is gone.

* fix(native-chat): copy a reappearing per-chat file again only while this build has not written past the copy

A per-chat file that an older build carried on under the copied epoch was
copied again even when this build had also written to the chat since the
copy, or had rolled its epoch. The second copy replaced the chat's block,
so what was sent in this build after the copy was gone for good.

Now the file is copied again only when the chat still stands exactly as it
was copied: the same epoch and tip the import marker recorded. Otherwise it
is set aside like any other file that is not this build's history, left on
disk untouched and recorded so no later open reads it. A second copy
therefore never replaces rows this build wrote, keeps the file's own epoch,
and the fresh-epoch rewrite goes away. The row it adds now says the history
includes what the older version recorded, not that anything was replaced.

* test(native-chat): pin that a chat founded here keeps its history, and the v1 schema upgrade

A chat this build founded has a pointer and no import marker, so a per-chat
file an older build later starts for it is set aside. Nothing pinned that
half of the rule: letting such a chat be copied again replaced its history
and every test still passed. A second test pins that a database written at
schema version 1 upgrades in place, gaining the set-aside table and keeping
its import markers.

* test(native-chat): drop a lost copied row by patching the source, not wrapping it

* chore(mobile): restore the mobile lockfile to main's

* fix(native-chat): pass a classified journal refusal through a send or Stop unchanged

* fix(native-chat): refuse a read whose owed copy fails as a failed open does

* test(native-chat): measure only the replace's WAL in the block-key case

Opening the chats starts a free-page pass that waits one event-loop turn,
and the seed never yields one, so that pass was still pending when the
replace committed. It woke during the async stat and reclaimed the pages
the replace freed, adding ~500 KB of WAL whenever the stat lost the race
(Linux CI). Drain that pass before measuring and stub the replace's own.

* test(native-chat): the RPC fixture's status journal can save its listing status

The status feed now hands every projection to the journal, which decides whether it is worth saving.

* test(native-chat): state why the RPC fixture's status journal cast is safe

* fix(native-chat): refuse a per-chat copy whose rows differ from the file, not only its counts

* fix(bench): build the replay benchmark's baseline arm from the base tree and release its handles on failure

* fix(native-chat): retry a failed listing status save on the next read of a cached status

* refactor(native-chat): drop the chat journal owner lock; the process instance lock already guards the profile

The journal carried its own exclusive lock, with a retry loop, an in-process
takeover, lock-gated runtime discovery and a "chats are open in another Orca"
refusal. Every shipped process kind (packaged desktop, serve mode, orcad)
already refuses a second instance on one profile before the journal opens, so
the lock only ever mattered for dev desktops, which the next commit covers at
the process level instead.

The host now opens its one journal connection at install with no lock. What a
sole process whose journal will not open needs stays: the install refusal
recorded for the gate, the no-host startup path, and the unverifiable chat
inventory, now in structured-agent-session-host-refusal.ts. The unreleased
journalOwnedElsewhere reason, its processKind fact and their copy are removed.

* fix(startup): dev desktops take the single-instance lock, and a second one says why it quit

Dev skipped Electron's single-instance lock so parallel `pnpm dev` runs from
several worktrees would not quit silently, but two dev processes on the
default orca-dev profile then write the same stores at once. Dev now takes
the lock like packaged builds: a second launch on the same profile focuses
the first window and exits with code 3, printing one stderr line that names
the taken profile and how to run another copy (ORCA_DEV_USER_DATA_PATH).

Serve mode, the macOS diagnostic bypass and the E2E harness are unchanged:
an E2E launch still skips the lock unless it sets
ORCA_E2E_ENFORCE_SINGLE_INSTANCE_LOCK=1.

* refactor(native-chat): key journal rows by chat, epoch and sequence

Rows in the host's journal database are now addressed by the chat's own
identity, with `(session_id, epoch, seq)` as the primary key, the same
shape each per-chat file already used. The block-keyed layout goes with
everything built on it: the block column and its allocator, the 2^21
block ceiling, and the import's reserved block table.

A first-use copy writes its rows under the file's epoch, which the chat's
pointer does not name until the verified copy publishes it, so no reader
sees a half-copied chat. A try that stopped midway leaves only rows no
pointer names, and the next try deletes them before it copies again.
Replace, rollover and repair delete by (chat, epoch).

This build's history always wins: once a chat was copied or founded here,
any per-chat file that reappears is set aside, and the same-epoch copy
again after a downgrade is removed.

The bounded free-page reclaim after every delete is dropped;
`auto_vacuum = INCREMENTAL` stays at file creation, so a later periodic
reclaim can still be added. Session search keeps its own step.

The schema moves to version 3. Versions 1 and 2 were written only by
unreleased builds of this change and are refused as found, not migrated.

* fix(native-chat): open a chat journal a newer Orca wrote read-only instead of refusing it

After a downgrade, the host's journal database carries a newer user_version. It was refused
outright, so every chat's history disappeared. It now opens on a read-only connection, as the
per-chat journals did: each chat shows what this build can read, from the database or a per-chat
file never copied in, and every write is refused with "Chats were saved by a newer Orca. Update
Orca to keep using them." Nothing is written, copied, repaired or founded, and the file stays
byte-identical. A table the newer schema changed reads as the same read-only refusal, not damage.

* refactor(native-chat): leave the saved listing status to the change that reads it

Nothing in this change reads the per-chat listing status column: it was a stored copy of a fact
the status feed derives, written after every turn end and cleared on every epoch change. The
status_json / status_seq columns, their writer, the saved-status type, the status feed's save and
its retry on a cached projection all go, with their tests. The change that lists chats from a
saved status adds the column back beside its reader.

* fix(native-chat): a chat saved by a newer Orca says to update Orca, not to try again

When a newer Orca wrote the chat journal, this build opens it read-only. A send or a Stop was
refused with the reason `journalUnavailable`, so today's desktop and phone clients chose the
words for an open that can clear: "Orca couldn't open this chat's history right now. Try again."
Retrying never cleared it; only updating Orca does.

The refusal now names its own reason, `journalWrittenByNewerOrca`, whose words are "Chats were
saved by a newer Orca. Update Orca to keep using them." A read refused the same way names it
too. An older client does not know the reason, drops it, and falls back to the code's words
("Orca couldn't read this chat's saved history."), and released clients still print the message.

* fix(native-chat): a chat journal from an unreleased build reads as unusable, not as retryable

A chat journal database stamped with schema 1 or 2 was written only by unreleased development
builds of this change. Opening it threw a plain error, which every chat reported as "Orca couldn't
open this chat's history right now. Try again." Retrying never cleared it.

It now throws a named error that is classified as unusable, so every chat says "Unable to load
this chat." The one log line names the file, says an unreleased development build wrote it, and
says to move it aside. Nothing migrates or renames it.

* docs(native-chat): drop the second-Orca-owns-the-chats case from three comments

The chat-only owner lock is gone, so only a chat journal that will not open leaves a runtime
unable to list its chats.

* docs(native-chat): correct three chat-journal comments the redesign left behind

A per-chat file left without its WAL is set aside, not copied again; nothing runs an incremental
vacuum yet, so the auto_vacuum mode is kept for a later pass; and the idle sweep drops a chat's
in-memory fold, since a chat holds no journal connection.

* refactor(native-chat): stop exporting chat-journal names nothing imports

Each is used only inside its own module now; the teardown's export served a deleted test.

* test(native-chat): name the version-0 test for what it covers, and check every journal table

The test named 'migrates an older user_version forward' covers only a version-0 file that already
has its tables; versions 1 and 2 are refused. The table test now also checks journal_imports and
journal_set_aside.

* fix(startup): a second dev launch's exit line no longer claims it focused a window

The running dev instance may be a background launch or a server, which show no window. The line
now says only that this launch passed its request to that instance.

* fix(native-chat): a failed structured-chat install closes the journal connection it opened

The install opened the chat journal database and closed it only if the record store then failed
to open. A later failure, such as the model catalog wiring or the host constructor, left the
connection open, and the next install opened a second one in the same process. Every failure
after the open now closes it.
2026-09-29 16:42:29 -07:00
Neil 6194a7a1b6 test: drop private-internal and boundary-census tests with behavioral owners (#23941)
Fourth audit wave, cut short by a session restart, so this lands the verified
subset rather than the full batch.

Removes private-predicate cases whose behavior is already covered through the
module's real entry point, and de-exports the seams they reached for. Also drops
three whole files whose every case was a duplicate or a call-shape grep.

The source-grep vein is close to exhausted. One auditor reviewed 15 remaining
flagged files and deleted nothing: what is left is mostly legitimate
architectural ratchets that no type checker and no behavioral test can reach —
AST fences banning `as`/`any` in an RPC operation region, discovered-vs-listed
set equality over subscription sites, count ceilings on unchecked reply readers,
and assertions on generated WebView bundles (no CDN URL, no `</script`
tokenizer escape, parses at the Chrome 74 floor). Those stay.
2026-09-29 16:35:49 -07:00
Jinwoo Hong 9420d49bcb fix(terminal): run Codex in Orca terminals without the shared background server (#23900) 2026-09-29 10:22:27 -07:00
Neil 31012aeb09 test: remove assertion-free probes, copied inventories and export-shape checks (#23816)
Second audit wave, targeting three more junk patterns:

- assertion-free cases that run code and assert nothing, so they pass no
  matter what the code does;
- inventory literals re-typed from a production declaration, where the only
  way the assertion can fail is someone editing one of the two copies;
- export key-set and export-shape loops (`typeof x === 'function'` over every
  export) that restate what TypeScript already enforces.

Yield is much smaller than wave 1 on purpose: the assertion-free scanner has
a high false-positive rate, because many flagged blocks assert through a
shared helper or their oracle is "this must not throw". Those were kept.

`mobileWebCheckArgs` in `config/scripts/run-mobile-web-app-checks.mjs` is
de-exported — after the inventory comparison went away, nothing outside the
module read it.
2026-09-29 02:21:47 -07:00
Neil 6e1b7e7fa3 test: remove junk tests that assert source text instead of behavior (#23815)
Deletes 101 test files and trims 112 more, all matching documented junk
patterns: exact source/import/string greps, copied inventories and export
lists, duplicate invocations of a contract another test already owns,
typeof-shape checks TypeScript already enforces, and self-comparisons.

The largest group read a production `.ts` file and asserted on its text —
for example a TaskPage test that required the source to contain
`selectedRepos.find((r) => r.id === newIssueRepoId) ?? selectedRepos[0] ?? null`.
Any behavior-preserving rename broke it; no behavior change ever did.

Production-side follow-through: exports that only these tests imported are
de-exported or deleted, stale comments pointing at removed censuses are
dropped, and the reliability-gate registry, `cloud/package.json` test lists,
and orphaned source-reading helpers are updated so nothing references a
deleted file.

Two files kept their real coverage and lost only the census scaffolding:
`agent-status-producer-census.test.ts` now drives all five producers end to
end instead of grepping the source tree, and `config-toml-trust-stale-writes`
replaces an export-list parity check.
2026-09-29 01:21:53 -07:00
OrcaWinandm4air e1362ada4c fix(terminal): stop inline-image decoders exhausting the renderer's wasm memory budget (#23499)
V8 reserves an 8 GiB guard region per wasm memory inside its 1 TiB sandbox,
so an Electron renderer can hold only ~124 live wasm memories regardless of
free RAM. @xterm/addon-image instantiated a SIXEL decoder per terminal at
activation (and kept IIP decoders after the first image), so ~120+ terminals
exhausted the budget: new panes raised 'WebAssembly.instantiate(): Out of
memory' rejections, and the next Kitty/IIP image threw 'WebAssembly.Memory():
could not allocate memory' out of the parser, permanently wedging that
terminal's write queue.

The addon-image source patch now borrows SIXEL decoders from a shared pool
only while a sequence is open (color registers stay on the terminal), drops
IIP decoders after each image, and turns a failed decoder allocation into a
dropped image instead of a parser throw. Bundles regenerated with
regenerate-xterm-patches.mjs --write.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-29 01:20:37 -07:00
Neil f8f656ca19 perf(ci): spend fewer concurrency slots per pull request (#23810)
A concurrency slot is charged per job, not per core, and the account's cap is
the scarce resource: standard runner minutes are free and unlimited on a public
repository. Two paths spent slots that bought nothing.

The unit matrix ran eight fixed shards averaging 6.5 minutes each, 3384
job-slots a day and 68% of all slot demand, while the arm pool queued 10.5
minutes at p95 — the queue was the oversharding. Five shards run the same work
in ~10.5 minutes each for three fewer slots per run.

Bun profile persistence escalated to all six platforms on `config/`,
`resources/` and `.github/` wholesale, which took 36.5% of the last 1100
commits through the full matrix where a platform-flavoured predicate takes 19%.
A pull request now qualifies one platform unless the change is platform-
flavoured, and the push to main re-qualifies all six, so an unescalated miss
surfaces minutes after merge rather than at the next cron. Missing changed-file
evidence and an unavailable dependency graph still fail closed to all six.
2026-09-29 00:13:33 -07:00
Neil 25d9c57e3a fix(native-chat): keep a child turn settling on an error Codex will not retry (#23808)
#23801 removed the child-path reading of Codex's turn-ending `error` along with
the import of the module #23682 deleted. That was the wrong half to remove: the
primary journal path can rely on Codex's failed `turn/completed` arriving within
~32 ms, which is what #23682 established, but a child turn has no such
guarantee, and without the error as its end the child's lifecycle row latches on
`working` for the life of the session. Three tests assert exactly that and could
not run, because the unresolved import had been skipping the unit matrix since
#23682 merged.

The reading is restored inline against `readCodexErrorWillRetry`, itself restored
to `codex-structured-thread-facts.ts`, rather than by reviving the deleted
module: its `thread-stopped-running` arm lost its only consumer when #23682
rewrote the primary path, so restoring the file would re-add dead code.

Also drops `pr-workflow-parallelism.test.mjs`'s read of
`.github/workflows/track-community-prs.yaml`, which #23796 deleted while leaving
the assertion behind. Same failure class, and it fails the same shard.
2026-09-28 23:42:01 -07:00
Neil 2ea3fb1d46 perf(ci): take advisory unit-selection evidence off the gate (#23776)
selection_evidence is continue-on-error on both the job and its comparison step,
so it can never fail a PR -- it downloads the shard reports, compares selection
against the full results and uploads a review artifact. But a caller's
`needs: test` waits for every job in the called workflow, so living inside
unit-tests.yml it held verify for ~36s after the last shard finished.

It moves to its own reusable workflow called as a sibling, so it still runs on
every PR and still uploads its artifact, but verify no longer waits for it. It is
deliberately absent from verify's needs, and a contract test pins both that and
its advisory status so it cannot drift back onto the critical path.

Measured on a recent run: the shards finished, then selection_evidence ran 36s,
then verify 3s. Only the last of those gates anything.
2026-09-28 22:13:35 -07:00
Brennan Benson c5fc0c6f26 fix(ci): keep a squash-merged RPC recording pin reachable through its pull request (#23720)
* fix(ci): keep a squash-merged RPC recording pin reachable through its pull request

Main's "RPC recording pin" check has been red since #22762: that branch pinned
the recording corpus to its own commit 03995ae, and the squash-merge left that
commit out of main's history. Every behaviour-change squash did the same, and
each needed a hand-made repin PR to clear it (#23565, #23535, #23046 and more).

The guard now accepts a pin that is either in this history or in the head of the
pull request whose squash wrote it into the manifest. It finds that pull request
from the `(#n)` subject of the commit that added the pin and fetches
`refs/pull/<n>/head`, which GitHub keeps after the branch is deleted. The
reproduce step uses the same lookup, so it can still check the pinned tree out.

* fix(ci): give the recording pin lookup room to walk a blobless clone

In CI's blobless clone, `git log -S` fetches the manifest's blobs one commit at a
time, a few seconds each. Under the 30 s process default the walk was killed after
a handful of manifest commits, which main's history already exceeds (up to 7
manifest commits between a pin landing and the next pin change), and the guard
then failed with an empty "Could not find the commit that pinned" error. The
lookup and the pull request fetch now carry explicit budgets and say when they
timed out.

The not-an-ancestor instruction now names the pull request whose head was
checked, or says the commit that pinned it names none.

Adds the two merge-preview shapes the guard runs on: a branch opened after a
squash resolves main's pin through the squash's pull request, and a branch whose
rebase dropped its own pinned commit fails on its pull request instead of on main.

* fix(mobile): tell a missing recording pin apart from product drift

After a squash the pinned commit can live only in its pull request's head, so a
clone that never fetched it makes `git diff --quiet <baseline>` exit 128. The
recorder reported that as "Product sources or lockfile differ from the pinned
main baseline", which sends the developer to repin a tree that may match. It now
prints git's error and the command that fetches the pin.

* fix(ci): ask GitHub which pull request holds a squash-dropped recording pin

The recording pin guard found the pull request that keeps a squash-dropped
pin by walking main's first-parent history for the commit that wrote the pin
into the manifest and reading "(#n)" off its subject. A merger who edits the
squash title loses the number, and the push to main turns red anyway. That
already happened on main: of the 22 squashes that left a pin outside main's
history, #21674's title had no "(#n)".

The guard now asks GitHub for the pull requests associated with the pinned
commit (GET /repos/{owner}/{repo}/commits/{sha}/pulls) and, for each in turn,
fetches refs/pull/<n>/head and accepts only when git proves the pin is an
ancestor of that head. GitHub only nominates candidates, so a wrong answer can
fail the guard but never pass it. The endpoint named the right pull request
for all 22 historical cases, #21674 included, and names none for commits a
force-push orphaned.

This removes the pickaxe walk, its 600 s budget and its lazy blob fetches in
a blobless clone, the first-parent subtlety, and the subject regex. A revert
that restores an older pull-request-only pin now resolves too, because the
lookup is by the pin itself rather than by the commit that last wrote it.

CI passes the job token to both guard steps and grants the job
pull-requests: read. Local runs work without a token on this public repo and
send GITHUB_TOKEN or GH_TOKEN when set. A failed lookup throws with the HTTP
status, and names the rate limit when an unauthenticated call is refused.
2026-09-28 21:17:07 -07:00
Neil 3976ad4c59 perf(test): remove obsolete structural snapshots (#23777) 2026-09-28 20:55:05 -07:00
Neil ec9f35e2ee perf(ci): plan the unit shards before the static-analysis gate instead of behind it (#23743)
A caller's `needs` gate the whole called workflow, so while the plan job lived in
unit-tests.yml it could not start until static analysis and typecheck had both
finished and passed -- and the shard matrix then waited on it. The two hops were
serial when they did not need to be: planning reads the checkout, a git diff
against HEAD^1, the import graph and the checked-in timing baseline in
config/scripts/ci-shard-timings.json, and consumes nothing that static analysis,
typecheck or the native-cache primer produce.

Planning moves to its own reusable workflow so pr.yml can run it against
code_paths alone, overlapping it with the gate. Measured across 99 runs, the
shard matrix is created a median 93s earlier (p25 47s, p90 241s, never later).
Planning stays a required predecessor of the shards, so an empty assignment
cannot expand the matrix.

The gate itself is deliberately left in place. It fires on 22% of runs, and the
shard queue wait knees hard above ~9 concurrent ARM jobs -- 4s median below that
against 218s at 15-19 -- so admitting 8 doomed shards per failed run would cost
more in queue pressure than it returns in latency.

Cost is one 37s ubuntu-latest job, which does not touch the ARM pool the shards
contend for.

A planning failure still fails the PR: the shards are skipped, and verify's
check_job requires success whenever the classifier says tests should run, so it
reports `test: expected success, got skipped`.
2026-09-28 19:22:10 -07:00
Brennan Benson c5330d0d52 fix(native-chat): stop killing processes that only inherited a chat's spawn tag (#23460)
* fix(native-chat): stop signalling processes that only inherited a spawn token

A spawn token is an environment variable, so every descendant of a provider child
carries it. The Linux-only startup scan treated any carrier no lease claimed as a lost
provider child and sent it SIGTERM, which also hit editors, tmux servers and nested
Orca processes the agent had started. Remove that scan's killing consumer; the token
scan stays for the reservation probe, and recorded owners are still stopped by
identity during recovery.

* fix(codex): remove the token-scan kill path from app-server teardown

Every descendant inherits the spawn token, so killing each pid that carries it can
reach processes the agent started that are not the provider. Production never
injected this path; teardown always uses the process-group and descendant-snapshot
proof. Drop it, its deps, and the now-unused spawn-token argument.
2026-09-28 15:25:24 -07:00
Brennan Benson 2ca4ecbc61 feat(orchestration): let a structured chat run orchestration as itself (#22568)
* feat(orchestration): inject the Orca session id into structured children and let the CLI act as it

Every structured session's child (native Claude, native Codex, and the terminal
view) carries ORCA_AGENT_SESSION_ID and reaches the Orca CLI. The CLI sends the id
in the orchestration envelope; when present it is the caller, and a caller flag
naming anyone else is refused before any request. The id is stripped from
inherited PTY env and from the SSH host-CLI passthrough, and crosses into WSL so
the host can refuse the cross-host claim.

* test(orchestration): pin session id injection for native Claude, native Codex, the terminal view, WSL, PTY inheritance and SSH

* test(orchestration): pin one caller precedence rule across every CLI verb that names its caller

Adds the per-verb table (flagless acts as the session; a conflicting --from or
--terminal is refused before any request; the session's own spellings are
accepted), the enumerated guess population with its positive control, the
structured worker's own handle, the identity-less refusal for an older child,
the unchanged terminal agent, and the envelope. dispatch-show's --from only fills
preview text, so it passes through unfenced and a session's flagless preview
names the address the real dispatch writes.

* refactor(orchestration): keep the identity-less marker reader to the marker; the id is checked first

* test(orchestration): pin that a host refusal of the session surfaces verbatim from the CLI

* fix(orchestration): keep the identity-less marker beside the id for CLIs that predate it

A CLI older than the id, reached through a global install when a shell rc resets
PATH, would otherwise guess a sibling's terminal in a chat that no longer carries
the marker. It refuses on the marker instead; a current CLI checks the id first,
so the marker never makes a session with an id identity-less.

* fix(orchestration): refuse a conflicting --from on gate-list and task-list scoped by --run

A --run listing needs no caller, so both handlers skipped the resolver and a
--from naming another actor was dropped silently under a session. The conflict
check now runs on that branch too; terminal callers are unchanged.

* fix(orchestration): name this app's CLI by absolute path for a structured session's login shells

A provider can run each command in a login shell: Codex runs zsh -lc, and the
profile rebuilds PATH, putting a global install (possibly an older Orca) ahead of
the directory Orca prepended. ORCA_CLI_COMMAND, which an agent resolves the CLI
from first, is now the absolute launcher in that directory (the native launcher
on Windows), so no shell's startup files can swap it. The PATH prepend stays for
shells that read no profile. Found by the live coordinator run of the next PR.

* test(orchestration): pin a structured worker's CLI command as this app's absolute launcher

* test(orchestration): run the zsh login-shell arm in the real-shell lane that installs zsh

The ordinary Linux unit lane has no /bin/zsh, so the zsh arm failed there with
ENOENT. It moves to a live-shell file registered in the shell-contracts lane; the
bash arm keeps running in every lane. The lane guard's detector now also sees a
zsh spawned through the ProcessSpec program field, which is how this test
escaped it.

* fix(orchestration): omit a structured child's CLI command when no launcher resolves, and pin its instance

A bare `orca` fallback named GNOME's screen reader on packaged Linux, and an inherited value named
another app's CLI. The builder now deletes any inherited value, sets the absolute launcher only when
one resolved, and pins ORCA_USER_DATA_PATH so a current CLI dials the instance that minted the id.
Renames the marker reader to hasStructuredSessionMarker and records why the terminal view carries
the id without the marker.

* fix(terminal): name this app's CLI launcher by absolute path in every local terminal

ORCA_CLI_COMMAND meant three things by lane: an absolute launcher for a structured session, a bare
name for WSL, and nothing for any other terminal, so a structured session's terminal view lost it.
Local terminals now get the same absolute launcher the structured lane gets; WSL keeps its guest
command name, and a terminal whose launcher does not resolve still gets none.

* feat(cli): hand a command to the session's own CLI when another Orca CLI was invoked

A login shell can reorder PATH behind a global install, and an agent or its helper script can run
bare `orca`, so the binary that answered depended on the agent following instructions. Orca's
packaged launchers and bare-orca shims now export ORCA_CLI_SELF (outermost wins). At the CLI entry,
when it names a different launcher than ORCA_CLI_COMMAND, the command re-runs once through the named
launcher with ORCA_CLI_REEXEC=1 and exits with its status; both variables are consumed so no child
inherits them. Dev launchers export no self on purpose, WSL and SSH names never qualify, and a
launcher that cannot start leaves the command to run here. The Windows launcher no longer rewrites
ORCA_CLI_COMMAND; the legacy ask protocol normalizes its resume command itself.

* refactor(orchestration): declare which flag names the caller on each spec and refuse at the CLI entry

Each handler hand-classified its --from/--terminal as the caller or a target, and the refusal of a
conflicting caller flag ran inside the caller resolver plus two standalone calls for --run listings,
so a new verb that read its flag raw would pass a sibling's handle to a pre-session host. Specs now
declare identityFlagRoles, the CLI entry refuses a conflicting caller flag once from the spec, the
resolver only applies the id-wins rule, and a test fails any orchestration verb that accepts --from
or --terminal without classifying it.

* perf(cli): keep the session caller check off the actor codec's module graph

The check runs at the CLI entry for every command, and the actor codec pulls zod through the session
record. Compare the session's own spellings as plain strings instead.

* refactor(cli): spell a session's address from the one prefix constant, off the codec's module graph

The Orca session address prefix moves to a leaf module with no imports, re-exported by
the address codec, so the CLI entry check derives `session:<id>` from that constant
instead of re-typing it and still stays off the codec's zod graph. Prose and test names
say caller or Orca session id, not actor.

* refactor(orchestration): drop the session id's terminal-view spawn now that the handoff is gone

The terminal handoff was removed, so no terminal is ever a structured session:
- delete the terminal-view identity env and its WSL passthrough, and their tests;
- strip the session caller keys from every terminal's env unconditionally;
- the CLI's own-address spelling moves beside the injected id in src/shared, with
  a test pinning it to the address the host's party resolver gives that session.

* fix(terminal): run the Codex launch preflight through the CLI the terminal names

Packaged Linux names the userData shim in ORCA_CLI_COMMAND, while the preflight
ran the bundled launcher behind it. The CLI saw a different launcher and handed
the preflight off to the shim, booting Electron twice before every codex launch.

* revert(terminal): keep terminals on main's ORCA_CLI_COMMAND and Codex preflight

Only a structured session needs an absolute ORCA_CLI_COMMAND; local terminals go back to
naming none (WSL keeps its guest command), and the Codex launch preflight goes back to the
bundled launcher. The CLI handoff is scoped to sessions, so a terminal's preflight can no
longer be handed off and start Electron twice.

This reverts commit d2cefb6c03 and commit dd2853a5a9.

* fix(cli): hand off to the session's CLI only inside a structured session

The handoff ran whenever an Orca launcher's ORCA_CLI_SELF differed from an absolute
ORCA_CLI_COMMAND, so any process with both - a terminal, a script - ran another install's CLI
instead of the one invoked: a beta's --version lied, and an AppImage command from a terminal
that outlived its Orca failed. It now requires the injected session id, the identity it exists
to deliver. The launcher variables are still consumed in every process.

* fix(cli): name the packaged Windows command after the handoff decision

The launcher stopped writing orca/orca-ide over ORCA_CLI_COMMAND so the handoff could see a
session's absolute launcher, which also changed what every Windows terminal's CLI read. The CLI
entry now applies the launcher's rule itself once the handoff is decided, so terminals and the
legacy ask resume command see exactly what they saw before, and the resume-command reader
goes back to its original form.

* refactor(cli): decide the session handoff from the CLI's own entry, not a launcher export

Every packaged launcher, shim and dispatcher exported ORCA_CLI_SELF so the CLI could tell which
launcher ran it, and compared that with the session's ORCA_CLI_COMMAND. Two launchers of the same
app are different files, so a session that reached its own app through a global orca-ide on Linux
still handed off and started Electron twice, and the export rode artifacts every terminal uses.

A structured session now also names the JS entry its launcher runs (ORCA_SESSION_CLI_ENTRY), and
the CLI compares its own argv entry with it: any launcher of the same app stays, another install
hands off. The launcher scripts, Linux shim and dispatcher go back to main; the Windows launcher
keeps only leaving ORCA_CLI_COMMAND for the CLI to name after the handoff decision.

* refactor(cli): drop the session CLI handoff; the pinned instance and injected id already bind any current CLI

Every current Orca CLI dials the instance ORCA_USER_DATA_PATH names and sends the injected
session id in the orchestration envelope, so a bare `orca` that reaches another install's
current CLI already acts as the session. An older CLI has no handoff code and refuses on the
marker. The handoff only lined up versions between two current CLIs, and comparing two
separately derived paths kept misfiring (an AppImage's mount against its registered
extraction started the CLI twice on every call).

Removes the re-exec, ORCA_SESSION_CLI_ENTRY and ORCA_CLI_REEXEC, and the CLI-side Windows
command naming; the packaged Windows launcher rewrites ORCA_CLI_COMMAND again, as on main,
inside its own process only. resolveHostCliEntryPath goes back to the SSH passthrough.

* test(orchestration): say why the registered worker case pins the handle, now that every session's env is populated
2026-09-28 15:19:44 -07:00
Neil 8c61a5df1f fix(windows): require signed release binaries and identify CLI launcher (#23680) 2026-09-28 14:34:08 -07:00
Neil 0f52bb8be5 perf(ci): use four ARM test workers and remove repeated compilation (#23685)
* ci: benchmark per-job Node compile caching on full unit shards

* ci: measure unit shards with three and four workers

* ci: benchmark localization extraction CLI patch

* perf(build): reuse identical relay bundles across platforms

* ci: compare Vitest 4 and 5 on complete ARM shards

* perf(ci): upgrade localization extraction to skip irrelevant syntax walks

* perf(ci): use all four ARM cores and remove benchmark workflows

* ci: preserve failures while capturing unit source revision

* fix(ci): preserve commented and escaped localization calls

* ci: remove corrected localization benchmark harness
2026-09-28 14:05:30 -07:00
Jinjing 95a16e3f67 fix(release): stop the release policy from deleting pipeline-cut releases (#23669)
* fix(release): stop the release policy from deleting pipeline-cut releases

The policy judged a release by who created the release object. Cut Release
reuses an existing draft, so a CI-built v1.4.216 whose draft a person had
created was deleted (tag included) when its notes were edited, and Latest
fell back to v1.4.214 because v1.4.215 was also published by a person.

- Authorize a desktop release when its annotated tag was created by the
  release pipeline and points at its `release: vX` commit, not only by author.
- Only delete on `published`; an edit never deletes a release or tag.
- Pick Latest from the highest authorized stable using the same check.
- Move the policy into config/scripts/release-policy.mjs with tests.

* fix(release): load the policy module from the tagged commit

Release events run the workflow file from the tag's commit, so checking out
the default branch could pair an old workflow with a newer module.
2026-09-28 12:03:58 -07:00
Jinwoo Hong e8d8f2f0b3 fix(deps): take Electron 43.7.5 so detached webviews stop blanking browser tabs (#23586)
Electron 43.7.0 threw 'Invalid guestInstanceId' from <webview>'s
disconnectedCallback for a loaded guest (electron/electron#53989), so a
webview React removed and re-inserted kept a dead guest id: the tab went
blank, reload did nothing, and the destroyed listener never fired.
43.7.4 (electron/electron#54097) returns early when the guest is gone.

Raises the runtime floor test to 43.7.4 so a downgrade cannot re-ship it.

Fixes STA-8757
2026-09-28 13:58:48 -04:00
Neil 21f8e0f9db ci: use faster gzip for temporary Linux test packages (#23609)
* ci: benchmark faster Linux package compression

* ci: pass compression options through typed builder configuration

* ci: retain original configuration for benchmark baseline

* test: preserve release settings in CI compression configuration

* ci: normalize generated changelog dates in package comparison

* ci: remove completed Linux compression benchmark
2026-09-28 03:48:07 -07:00
Neil cf20423ff3 ci: skip unrelated installs and share xterm build dependencies (#23607)
* ci: pilot shared xterm installed dependencies

* ci: bound xterm cache production to verified main entries

* ci: benchmark xterm reuse on the production ARM runner

* ci: avoid installing Orca dependencies for standalone xterm checks

* ci: use Node-only setup in the production xterm job

* ci: remove completed xterm benchmark workflow
2026-09-28 03:46:22 -07:00
400e4e7957 feat(agents): add Freebuff launch and sidebar status support (#23567)
Add Freebuff launch support and execution-host status reporting for the sidebar, including running, question, blocked, and settled states. Validate against captured CLI transcripts and real rendered sidebar evidence.

Cross-referenced community implementations #17065, #20839, and the Freebuff portion of #18790. Preserve their agent/catalog/mobile/documentation coverage and add canonical status publication and regression tests.

Co-authored-by: Harkaran Brar <18134082+harkaranbrar7@users.noreply.github.com>
Co-authored-by: Prarambha369 <98906077+Prarambha369@users.noreply.github.com>
Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>
2026-09-28 02:32:41 -07:00
Neil d05080175b perf(ci): stop duplicating shared Linux download caches (#23578)
* perf(ci): pilot pnpm verification record caching on Linux

* test(ci): review pnpm verification record in mobile cache audit

* perf(ci): share Electron downloads and clean closed PR caches

* test(ci): retain cache ordering checks for restore-only consumers

* perf(ci): limit archive sharing rollout to primed Linux hosts
2026-09-28 01:58:46 -07:00
Neil e6fbbdf684 perf(ci): cache pnpm verification records on Linux (#23568)
* perf(ci): pilot pnpm verification record caching on Linux

* test(ci): review pnpm verification record in mobile cache audit
2026-09-28 01:50:30 -07:00
Neil 3dd7d29455 Revert "perf(ci): shard the anti-slop audit across processes instead of one JS runtime (#23543)" (#23575)
This reverts commit fae0ae7a46.
2026-09-28 01:47:07 -07:00
Jinwoo Hong ba858ee446 feat(mobile): the keyboard covers the page like a native screen, and the shell says its height (#23110)
The shell no longer shortens the WebView for the keyboard; it publishes the keyboard height like the safe-area insets, so native's keyboard lift, refit hold and dismiss key run on the page unchanged. Keyboard and inset arithmetic read the shell's OS through a host-os seam. One page-version floor (manifest pageVersion, shell floor 1) replaces per-feature accept negotiation; a page below the floor gets the existing update wall, a desktop with no bundle keeps native screens. iOS shell drops the form accessory bar and its own keyboard observers. Native session screens untouched.
2026-09-28 04:07:35 -04:00
Neil 080c562898 perf(ci): diff against the merge commit's first parent so PR checkouts can be shallow (#23562)
Every changed-path gate asked git for `--merge-base "$BASE_SHA" "$HEAD_SHA"`,
which needs the event payload's base SHA to be in the local graph. That is the
only reason two jobs cloned all 8127 refs' history. On a pull_request checkout
HEAD is already the merge commit, so its first parent is the base side and no
merge base has to be computed. config/scripts/git-pull-request-diff-base.mjs
resolved that for the two Node gates; the workflow's inline gates now use the
same helper through a small CLI rather than open-coding it.

code_paths gates all 22 jobs, so its checkout is charged to the start of every
one of them: measured 20.7s to 1.6s, keeping blob:none because its sparse tree
is ~7 files and leaves no blobs to refetch. Static analysis drops the filter
instead, since populating all 30,226 files makes blob:none force a second
promisor fetch: 23s to ~11s.

Verified on a real merge ref. At depth 50 the old and new forms produce
identical changed-file sets. At depth 2 the new form still works and the old one
fails with `fatal: bad object`, which is the failure a stale base would have
caused once the checkout stopped being complete.

Also drops the dead resolveBase + merge-base prelude in the changed-code gate,
whose result resolvePullRequestDiffBase already discarded on every PR.
2026-09-28 00:37:26 -07:00
Jinwoo Hong 2077956254 fix(mobile): size a terminal's first subscribe from the document's reported cell box (#23080)
* fix(mobile): size a terminal's first subscribe from the document's reported cell box

#22960 sent phone dims on a terminal's first subscribe by opening a throwaway
empty terminal (init 80x24 ""), awaiting its ready and measuring, behind a
per-document first-subscribe mark whose lifetime was tied to web-ready. That
cost a second xterm/WebGL instance and ~150 ms per open, plus lifecycle state.

The document now measures the cell box without a terminal (xterm 6's
CharSizeService strategy, rounded as the renderer rounds it) for every
text-size preset and reports it with its viewport in web-ready; a table,
because the text scale only reaches the document after that notify. Each
init's ready reports the box xterm actually laid out, which replaces the
probe's entry. The controller answers fitDimensions/measureFitDimensions from
that table and the view's layout with no message; without a table it asks the
document as before.

The session seeds an unmeasured viewport synchronously in subscribeToTerminal,
so the first subscribe carries dims by construction. Deleted: the empty init,
its awaitReady gate, deferFirstSubscribeUntilViewportMeasured and the
subscribedDocuments mark. The fit pass is unchanged and still covers a
document that reports no cell box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): read the reported cell box through in-narrowing, not Reflect.get

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct the probe's cell-box guess from the box xterm lays out

The web-ready probe is a guess: building the WebGL addon creates no context,
so a context that fails on load lands on the DOM renderer, whose width is not
snapped and depends on the column count. Before, a ready box that differed was
only logged; the first subscribe had carried the wrong column count, the host
echoed it, the fit pass saw the viewport equal to the host's dims, and the grid
stayed slightly shrunk. The store also kept the WebGL width after a context loss.

The document now reports the box xterm laid out whenever it changes (from
onRender, which covers a renderer swap and a DPR change that
onDimensionsChange does not fire for, and at ready). The store replaces the
guess; when that changes the current text size's entry, the view calls
onCellBoxChange with xterm's grid and the session re-fits, running the bounded
fit pass if the dims moved (one resubscribe). Equal boxes do nothing.

The RN layout box now survives a document reload; the document's own viewport
only stands in until the view reports a layout (on the page, web-ready arrives
first). The mismatch console.log is gone, and the probe's rounding names the
xterm version it copies.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct each cell-box guess at most once, so a DOM renderer cannot loop

On the DOM renderer the cell width is the rounded canvas width divided by the
column count, so every re-init at new cols reported a new box. Each one counted
as a correction, a floor over floats could flip the fit between two sizes, and
each flip landed converged, which reset the resubscribe budget: an unbounded
series of full-snapshot resubscribes.

Only the first laid-out box for a guessed text size may be a correction; later
reports still update the store, so fits stay truthful, but never resubscribe on
their own. The fit's floor gains a 1e-6 epsilon so floating-point error at an
exact boundary cannot flip a column or row. New document tests pin the render
report after a renderer swap and the report at ready for a paused renderer.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): make xterm the only terminal cell measurer

The document builds its real terminal before web-ready, at the app's
text scale, and reports the box xterm laid out; the first init reuses
that terminal. The page-side prediction, the per-scale guess table and
the once-per-document correction are gone. The app remembers the box
per text scale for its lifetime, so a later open at a known scale
subscribes with phone dims at once. A box that changes at the same grid
(renderer swap, pixel ratio) refits the open terminal in place.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep commands queued before the terminal WebView first loads

A subscribe sized from the stored cell box can queue init before the
native WebView reports its first load start, which cleared the queue
and left the terminal blank. Only a reload now drops queued commands.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): re-init a document that lacks the subscription's init, and fit one frame width

- Web-ready now says whether the document holds the terminal's latest init
  (a reload before the first ready drops a queued one); the session
  resubscribes any initialized terminal whose document lacks it.
- One grid fit, shared by the app and the document, fed the unrounded frame
  width React Native laid out; it keeps exact fits whole at fractional pixel
  ratios. The document's viewport-width fits are gone.
- The page builds every document at the scale the view mounted with, as the
  native WebView does.
- A new document's first cell box is compared against the grid the
  subscribe fitted from the stored box.
- The terminal built before ready stays hidden until its first init.
- The cell-box census matches glyph-measurement techniques, not names;
  the store's unused clear() is gone.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): build the terminal before ready only for the view shown at mount

A session mounts one terminal view per tab, and each built xterm and a
WebGL context before ready: 20 tabs made 20 contexts at load, past the
~16 a page (or Android's shared WebView renderer) holds, and native logged
32 context losses. Only the view shown when it mounts builds early now;
the rest build at their first init as before. Deferring the WebGL addon
instead would change the reported box: the DOM renderer lays out 7.8x15
where WebGL lays out 7.667x15 at the same font.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): write a WebView document's start values into its page, not an injected script

Android ran the pre-content injected script after the document's own in
1 of 22 documents on the emulator; that document started with no text
scale or shown flag and built a terminal it should not have. The values
now sit in the page ahead of the document script, one source object per
start pair so a render never reloads the WebView.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the pre-ready terminal measures and reports while hidden

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): measure only the laid-out frame, and refit on a new grid, not a new width

- A measure needs both of the frame's dimensions from React Native; the
  document's viewport-height fallback is gone, and before the first layout
  the handle answers no fit without asking the document.
- A frame width change that still fits the PTY's grid from the stored box is
  a no-op, so sub-pixel layout jitter no longer re-measures. The width ref is
  written in that effect rather than during render (react-doctor).
- One "last grid" ref: the last reported grid, or the one a subscribe fitted
  from the stored box.
- The page render rig measures through the frame it laid out, as the session
  does, and lets the replay's fit settle before its resize-refit witness.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let only the current terminal document's ready flush

A reload kept the WebView and its onMessage, so the old document's late
web-ready flushed the queue into the reloading view and the new document
got a second init. Each document now gets its own view (keyed on a
generation the controller owns), every notify carries the generation of
the view that received it, and a web-ready from a replaced document
flushes nothing and stamps nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop every notify from a replaced terminal document

One rule at the receive boundary: a notify from any generation but the
current one is dropped, whatever its type, not only web-ready.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make fitDimensions a pure question; name each generation counter

- fitDimensions no longer records the grid. A width change to a new grid
  asked it first, so the DOM renderer's report of that grid's box read as
  "same grid, new box" and refit again. Only the first-subscribe seed
  (seedFitDimensions) records the grid the document's first report is
  checked against.
- viewGeneration counts the views, readyGeneration counts web-readies.
- replaceDocument no longer resets the load flag; the load-start reset
  stays as the guard for a view that reloads itself.
- The name-based lifecycle census is replaced by a behavioural test.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): typecheck the handle mocks, drop the unused cell-box get

- The two handle mocks carry both fitDimensions and seedFitDimensions,
  and the fake-timer acts return nothing, so the three test files check
  under tsconfig.test.json again.
- terminalCellBoxes.get had no product caller; the store's tests assert
  through fit.
- The load-start comment says what the controller does now.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold the grid the document has, ignore a replaced view's load start, dispose a failed pre-ready terminal

- The document reports a new grid even with an unchanged box, so an
  in-place reflow on WebGL is held before a later renderer swap at that
  grid; the swap then refits. The app's apply paths do not hold the grid
  themselves: the DOM renderer's box follows cols, and a grid held on
  apply would read its own box as a renderer change and loop. One
  writer (holdGrid) holds the seeded or reported grid.
- A load start from a view a replacement unmounted is ignored, as its
  notifies already are.
- A terminal whose open throws before ready is disposed, not only
  unreferenced.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): ignore every native event from a replaced terminal view

One wrapper binds each WebView lifecycle event (load start, error, HTTP
error, render process gone, content process terminated) to the view's
generation, so a replaced view's late event cannot reset, replace or
put an error over the current document.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): a DOM seed refits once on its first report, not on the refit's own

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a terminal only after its document is ready

The document still builds its terminal before ready and reports the cell
box xterm laid out in web-ready; the app now subscribes after that ready
and fits from that box, so nothing is sent to a document before it is
ready. Everything that made a pre-ready subscribe safe goes: the
app-lifetime box store, the seed fit, the per-document view generations
and their event filtering, the init tracker and the hasInit resubscribe.
The native view reloads in place again and web-ready keeps main's reload
rule. Boxes are kept per view; the grid a document last reported still
guards the in-place refit against the DOM renderer's cols-dependent box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold one reported cell box and the grid the subscribe fitted

The controller keeps only the box the current document last reported,
not a per-text-size store: the document re-reports on a scale change.
The subscribe after ready fits from that box and holds the grid it
fitted, so the DOM renderer's first report at that grid (a new box)
refits once in place and converges; refit and apply paths hold nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): fit only a ready box at the app's scale; forget a reloaded document's box and grid

A reload keeps the document's mount scale, so a ready after a text-size
change reports a box at the old scale; that box no longer sizes the first
subscribe, which then takes the no-box path. A readiness reset drops the
old document's box and held grid, so the new document's first DOM report
at the same grid does not refit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give the terminal document its frame at init, and fit text scale over it only

A subscribe sized from the ready box sends no measure, so the document
had no frame when the text size changed and reported the pre-refit row
pitch. The app's init now carries the frame it laid out, in the fields a
measure uses; the router takes it from either. The text-scale fit reads
only that frame, with no viewport fallback.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say why a frameless text-scale change skips the resize

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one cell box per terminal notify, not an array

web-ready and cell-metrics carry `cellBox: {fontScale, cellWidth,
cellHeight} | null`; the document's `laidOutCellBox` returns one or
null and the parser validates one object. The text-scale match moves
from web-ready into `handle.fitDimensions`, the one place a box is fitted.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): fit terminals in the app from the reported box; drop the measure round trip

The app already holds the box the document reported, so the refit and
the fit pass await the init's ready and call `handle.fitDimensions`
instead of posting `measure` and waiting on `measure-result`. The
document's measure, its retries, and the measure promise and timeout go.
The document still resizes locally on a text-size change, so every grid
the app sends (init, resize, reflow) carries the laid-out frame it was
fitted to. `holdSubscribedGrid` replaces `subscribeFitDimensions`, so
the only fits are `fitDimensionsFromCell` and `handle.fitDimensions`.
The render rig reads its fit from the ready box. The recorder adapter
mounts the new handle with the same recorded effects; the goldens it
mounts move on their adapterSha256 header only.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): keep the terminal frame in one ref, and notify a new width imperatively

The session held the frame in a height ref, a width ref, a width state
and the refit's own width ref. It now holds one `terminalFrameRef`
({width, height} | null until the first layout; a hidden 0x0 layout
keeps the last box). onLayout notifies a new width imperatively, as it
does height, and the refit's notify skips a width whose fit is the grid
the PTY has. `terminal-frame-width-refit.ts`, the width state and its
effect go. The subscribe's layout gate reads "no frame yet" directly.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a held-back terminal on the frame's first layout only

`handleTerminalFrameLayout` ran on every onLayout; it now runs once, when
the frame first has a size. Later layouts only notify a new width.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): size the first subscribe inline in subscribeToTerminal

`sizeTerminalViewportFromCellBox` wrapped five lines in a 37-line
module; the subscribe now fits the ready box against the frame, holds
that grid and records the diagnostic itself. The helper's tests fold
into the subscription tests, which move to the subscription's name.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): drop the unreachable font-size guard on the reported cell box

xterm 6.1.0-beta.303 updates the render service's cell box in the same
task that sets `options.fontSize`: CharSizeService.measure fires
onCharSizeChange, and RenderService.handleCharSizeChanged runs the
renderer's `_updateDimensions` (DomRenderer.ts:359, WebglRenderer.ts:229).
`term.onRender` fires from RenderService._renderRows after the rows
are drawn (RenderService.ts:213, CoreBrowserTerminal.ts:538), and the
document writes its text scale and the font size in one task
(text-scaling.ts applyTextScale, terminal-init.ts init). So no report
can read a box between the font and the scale; the guard and its test
go. A new test pins the real order: no report when the font is set,
the new box at the new scale on the next render.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one start seam, no source cache, the reported box as an object

- `useState` already pins each view's WebView source at mount (a new
  test re-renders at another text scale and gets the same object), so
  the module-level `webViewSources` Map goes.
- `initialTextScale` and `buildsTerminalBeforeReady` become one
  `start(): { textScale, shown }` seam.
- `reportedCellBox` holds the last reported box and grid, not a string key.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to this branch and re-record

The terminal refit now fits in the app from the reported box and reads
one frame ref, so the recorder's terminal adapter mounts the new handle
(`awaitReady` + `fitDimensions`) and options (`terminalFrameRef`),
keeping its recorded effects. `baseline` is repinned to 21954dbd2f, the
last commit to touch a fenced path, and every golden is re-recorded.
Proof by class against HEAD: 787 header-only, 0 body moved, 0 added,
0 deleted. Header keys moved: `baseline` on all 787, and `adapterSha256`
on the 14 goldens `terminal-mount-adapters.ts` mounts (query-reply 3,
accessory-raw-send 4, takeover-report 4, viewport-refit 3). No recorded
traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold the reported cell box and its grid in one ref

The controller kept the box in `cellBoxRef`, the grid in a string
`lastGridRef` and wrote it through `terminal-held-grid.ts`. One
`heldRef` now holds `{ cellBox, grid }`, as the document's own
`reportedCellBox` does: web-ready writes the box, every cell-metrics
report writes both, `holdSubscribedGrid` writes the grid, and a
readiness reset clears it. Same write points, so the one-refit bound
holds; the DOM-loop and refit-once tests pass unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the hold-rule commit and re-record

H (f00bebba48) touched a fenced path after the last repin, so
`baseline` moves to it and every golden is re-recorded. Against the
corpus before this branch's refreshes (21954dbd2f): 787 header-only,
0 body moved, 0 added, 0 deleted; `baseline` on all 787 and
`adapterSha256` on the 14 goldens `terminal-mount-adapters.ts` mounts.
Against the previous refresh: `baseline` only. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the main merge and re-record

The merge (b3b1b0def2) is the last commit to touch a fenced path, so
`baseline` moves to it and every golden is re-recorded. Against
97b5bb2b9a: 787 header-only, 0 body moved, 0 added, 0 deleted;
`baseline` on all 787, and `adapterSha256` on the 14
session.diff-review-actions goldens whose adapter #22951 edited. Against
origin/main: 787 header-only, 0 body moved/added/deleted; `baseline` on
all 787 and `adapterSha256` on this branch's 14 terminal goldens.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): the terminal document holds the grid and decides each refit

The document already kept the last reported box and grid; the app kept
a mirror of both to decide the refit. Now the document decides: its
`cell-box` notify carries `{ cellBox, refit }`, sent only when the box
changes, with `refit` a box that changed at a kept grid. web-ready
records the pre-ready terminal's box at its 80x24 grid, and the first
init that reuses that terminal holds the init's grid, so the DOM
renderer's first report refits once, as the subscribe's hold did. A
re-init no longer clears the record, so a new renderer at the same grid
still refits. The app keeps one `cellBoxRef` and `holdSubscribedGrid`,
`heldRef` and the grid on the notify go. The one-refit, DOM-loop and
renderer-swap tests move to the document with the same scenarios.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one init options object, and a frame on every grid

`init` takes `{ cols, rows, data, preserveScroll, oscLinks, frame }`
instead of six positionals, and `init`, `resize` and `reflow` (handle
and messages) require `frame: TerminalFrame | null`. The refit's reflow
check reads `!dims` alone, and the controller's test file is named for
the `cell-box` notify it now covers.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one notifyTerminalFrame for the frame's layout

The frame's onLayout made four calls and held the classification
itself. It now calls `notifyTerminalFrame({ width, height })`, and the
session's terminal-webview hook keeps the one frame ref, notifies the
height, subscribes the document held back for the first layout, and
notifies a later width change. `handleTerminalFrameLayout` is named for
what it does: `subscribeIntendedActiveTerminal`. The layout tests move
to that hook.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-8 head and re-record

85d421963c is the last commit to touch a fenced path. Against
a676c1b65a: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): name the init option initialData, as the message does

The init option `data` becomes `initialData`, the message field's name,
so the controller passes it through unrenamed. The `preserveScroll` why
stays on the message type only, and the document test's title names the
three grids that carry the frame.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-9 head and re-record

486566c82b is the last commit to touch a fenced path. Against
3371c39715: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci: rerun checks against main with #23560 landed

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 03:26:49 -04:00
Brennan BensonandClaude 7a24d3d335 fix(native-chat): the conversation outlives its agent (#22835)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* fix(native-chat): the conversation outlives its agent

Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.

- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
  at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
  a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
  tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
  handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a restart offer ends when the chat's agent starts again

The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).

A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.

Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.

* fix(runtime): end a transcript stream when its client unsubscribes

Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.

Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart

The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.

The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.

Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.

* test(native-chat): an older build reads the restart offer this build records

The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.

* fix(native-chat): read a restart offer against where the journal stood when it was taken

"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.

The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.

* test(native-chat): wait for the listing's retire write before reading the recovery file

* fix(native-chat): keep the terminal-backed chat's read error over its local echoes

Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.

* fix(native-chat): a start retries the exit settlement a failed journal write left owed

An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.

* perf(native-chat): answer the owner check without opening the chat

Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.

* fix(native-chat): a read waiting on the session lock opens nothing once quit began

The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.

* fix(native-chat): read a failed resume's chat before calling it retryable

Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.

* test(native-chat): type the provider event sink the settlement test reaches for

* fix(native-chat): say the structured read keeps trying only where it does

The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.

* test(native-chat): await the send's settlement instead of polling for the start

The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.

* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations

The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.

Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.

* fix(orchestration): route no mail to a structured worker its orchestration released

A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.

* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation

A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.

* fix(orchestration): read the released row optionally, as the authority does

worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.

* fix(native-chat): the status bar drops a restart offer the chat moved on from

The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.

* test(native-chat): a roster of idle or finished children does not keep an agent awake

The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.

* fix(orchestration): a task dispatched into a resting structured worker keeps it running

The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.

* docs(native-chat): comments stop describing the hold this PR removed

Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.

* fix(native-chat): a restart offer keeps the start its own continuation made

Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.

* fix(native-chat): an agent gets a full idle window after its owed work ends

The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.

* test(claude): the options-read fixture runs a live child

The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.

* test(native-chat): host tests reach its collaborators through a typed seam

The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.

* refactor(orchestration): one owner answers a structured worker's custody

Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.

* refactor(orchestration): owed work is an open dispatch on the worker's incarnation

A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.

* fix(native-chat): a restart offer knows its continuations by a tag in their id

The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.

* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait

The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.

Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): the idle sweep reads owed work every tick

Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.

* fix(native-chat): a continuation handed to the agent stays sent

The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): the interrupted create's own retry continues again

The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.

* docs(native-chat): three comments that still had views starting agents

A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* test(native-chat): a reader's open settles the turn a failed exit settlement left running

An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.

* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner

A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn

On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* docs(native-chat): drop the removed dispatch hold from six comments

A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.

* test(native-chat): rest the owner-status chat through the idle sweep, not a hold

The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.

* fix(native-chat): show the structured pane's retrying line when a read fails

The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.

* test(native-chat): wait for a send's background start before the refusal oracle removes its store

An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals.

* fix(native-chat): a start a message waited on gets one failure row, the delivery loop's

When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice.

The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 23:46:14 -07:00
Neil 9179b93ebf ci: reduce repeated runner work and validate affected-test selection (#23540)
* ci: stage heavy checks and measure affected-test selection

* fix(ci): exercise the real Git boundary in unit selection planning

* Harden review cancellation and CI demand reporting
2026-09-27 23:25:20 -07:00
Neil fae0ae7a46 perf(ci): shard the anti-slop audit across processes instead of one JS runtime (#23543)
config/oxlint-anti-slop.json turns every native oxlint category off and runs its
rules through jsPlugins, so oxlint's threaded Rust engine does no work and the
pass is one JS runtime per process. Measured, it does not scale with --threads:
11.68s at 4 threads against 12.38s at 16. Parallelism has to come from more
processes, so the audit now splits its file set across them.

Locally, 11.72s single pass against 3.83s at 4 shards (3.1x) and 2.91s at 8.

Sharding is sound because every anti-slop rule is a single-file analysis; the
only mutable module state is a WeakMap keyed on each file's own Program node.
Verified by running both shapes with all 18 rules enabled: 392,398 findings from
one pass and from the shard union, identical as sorted multisets.
2026-09-27 23:15:19 -07:00
Neil 67581dd090 perf(ci): stop the orcad smoke idling 15s and the static job fetching mobile packages it skips (#23541)
The shutdown race in the orcad terminal smoke never cleared its losing timer, so
the process sat on a live 15s timer after PASS had already printed. Measured
locally: 21.4s -> 6.52s, with the round trip and the shutdown assertion intact.

Static analysis also asked for the mixed root+mobile pnpm store (537 MB, 8.6s to
restore) on every run, while installing mobile dependencies only when the diff
needs them. Most runs paid 216 MB for packages they never linked.
2026-09-27 22:51:59 -07:00
Brennan BensonandClaude 85067494a1 fix(native-chat): a request that failed reads as failed (#22944)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 22:23:49 -07:00
Jinwoo Hong 52dc32f9ea fix(mobile): the page pushes its terminal frame into the document from RN layout (#23079)
* fix(mobile): keep one mounted terminal frame under every session branch

The frame was keyed so react-native-web would attach its onLayout: a View
that gains onLayout after mount is never observed, and unkeyed the frame
reused the loading branch's View. One contentFrame View now wraps every
branch and carries the handler from the first mount, so no key is
needed. terminalFrame stays on the terminal branch, so nothing else is
clipped.

The parity pin moves: the key string leaves, one View and one style
reference arrive.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): push the page terminal's box into its document from RN layout

The page's document sized itself through a ResizeObserver on its host,
which duplicated RN layout and needed three rules in fit-scale: skip a
0-wide box, skip the last fitted box, and forget that box on any fit
request. The host View's onLayout now pushes into the mount, which is
the page's counterpart of the WebView's window resize.

react-native-web still lays a display:none screen out as 0x0 and its
return as the old box, so the mount treats neither as a change and keeps
answering the last real box while hidden, as a WebView keeps its size.
A fit asked for while hidden therefore lands at once, and fittedBox and
all three rules go. The zero-width wait in applyFitScale stays for a
host that mounts under a hidden screen before its first layout.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold a page terminal's fit while its host is hidden

A fit asked for under a covering screen read the last box, but where
xterm cannot measure cells on a display:none host (its DOM measure, with
no OffscreenCanvas) the retry loop ran out and committed scale 1, and
the show that followed was not a change, so 1 stuck.

The page's rect now says when the host is hidden, and the document holds
any fit asked for then as fitPending instead of committing. The mount
reports the same box coming back as 'shown', distinct from 'resized',
and the document runs a held fit on it and otherwise does nothing, so
pan and zoom still survive a plain hide and show. Native's window
resize reports 'resized' and is never hidden.

The mount also seeds its box from the host, so a first layout that
lands before the mount is not later mistaken for a resize. The hidden
text-scale case now asserts the resized column count, which stale
cells miss.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read the page terminal's grid once xterm has sized its cell

The init render check read the grid the moment .xterm-screen existed,
but xterm sizes its one-cell helper textarea only on a cursor move or
resize, after the replay drains. Under full-suite load the read won
that race and measured a 0-wide cell. In the failing runs the fit had
committed at 390/560 on the cell-width gate, so the page was right and
the read was early. It now waits for a sized cell: 8/8 under four-way
parallel load, where it was 4/8.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin that a held page terminal fit runs once

A held fit that lands on show must be spent: a second hide and show
re-running it would reset the pan and zoom the user set in between. The
case now hides and shows again and expects no new transform, which a
commit that stops clearing fitPending fails.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold the refit a narrow hidden text-scale change owes

A text-scale change whose new cells leave fewer than MIN_FIT_COLS in
the box skips the grid resize and returned before any fit. Base cleared
fittedBox there so the next box refit; with that gone, a hidden host
shown at the same box kept the old scale under the larger font. The
branch now asks for the fit while the host is hidden, which holds it
until show. A shown host, and every native one, returns as before.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): fit a text-scale change too large to resize the grid

A visible viewport too narrow for the new cells (280 px at 200%, 18
columns) skipped the resize and returned without a fit, so the larger
text overflowed the fit made for the smaller cells. The branch now fits
unconditionally; applyFitScale already holds the fit while hidden.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 00:06:35 -04:00
Jinjing 5a8237a7a5 Browser tab placement review (#23462)
* Implement source-following placement for browser tabs

- Add afterTabId and executionHostId parameters to track and anchor new tabs after their source
- Resolve source browser pages to unified tab wrappers for background link opens
- Implement anchor-based insertion that respects pinned tab boundaries
- Stage source-adjacent rows before host RPC for paired browser creation
- Ensure duplicated tabs and background links place directly after their source
- Add comprehensive test coverage for placement scenarios across single/split groups

* Preserve tab strip scroll anchor when tabs are added

Keep the viewed tab stable on screen when other tabs are inserted around it.
Records the active tab's on-screen position before insertions and restores it
after, so the user's focused tab doesn't jump unexpectedly. Matches the behavior
of VS Code and Chrome.

* Refine browser tab placement: fix host handling and sort-order gaps

- Avoid reordering unified tabs when the computed order matches the current order
- Do not substitute execution host for browser tab wrappers; let createUnifiedTab apply its active-workspace fallback instead
- Always apply sort values after tab insertion to prevent sortOrder gaps from preview replacement
- Fix import path for paired browser tab creation and clarify registration comments

* refactor(tests): use paired-browser-tab-creator registry for browser tab

- Replace direct web-runtime-session mock with paired-browser-tab-creator pattern
- Add test-rig file naming convention to localization audit skip list
2026-09-27 18:08:09 -07:00
Brennan Benson d6336be8db ci(e2e): run the SSH browser route e2e when its source changes (#23498)
A change to the SSH workspace browser route, its gate card, or the host-connection
phase it waits on selected no e2e specs, so the Docker SSH browser spec never ran
on the PR that changed it.
2026-09-27 17:51:16 -07:00
NeilandClaude 161bdf93c3 fix(bench): load benchmark modules under test through jiti (#23482)
`pnpm run bench:terminal-partial-escape-tail` and
`pnpm run bench:worktree-refresh-churn` both died at startup with
ERR_MODULE_NOT_FOUND. Each entrypoint static-imported a `src/` module with an
explicit `.ts` extension, but bare `node` type-stripping cannot resolve the
extensionless relative specifiers *inside* that module's graph
(`terminal-partial-escape-tail.ts` -> `./terminal-escape-introducer`,
`worktree-catalog-reconciliation.ts` -> `../../../../shared/structural-value-equality`).

Routes both through jiti, matching the four benchmarks that already load `src/`
TypeScript that way (`pty-source-ack-boundary`, `locale-collator-sort`,
`worktree-base-pending-marker`, `wsl-git-shell`). Adding the extension at each
import site was the alternative, but `config/tsconfig.node.json` does not set
`allowImportingTsExtensions`, so a `.ts` specifier in `src/shared` fails
typecheck with TS5097 -- and no file under `src/` uses that shape today.

Developer tooling only; no production code changed.

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 15:34:38 -07:00
Neil a4b60c2ea3 test(windows): stop the NSIS capability probe failing on a slow PowerShell cold start (#23381) 2026-09-27 13:24:48 -07:00
OrcaWinandm4air 27b823f934 ci: compile the E2E CLI once for all consumers (#23384)
* ci: share compiled CLI output across E2E consumers

* ci: preserve CLI setup and old-ref fallback for shared artifacts

* docs: record shared E2E CLI benchmark evidence

* docs: include final CLI reuse timing range

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 02:02:29 -07:00
OrcaWinandm4air c15f082031 ci: build independent Electron targets together for E2E (#23378)
* ci: reuse parallel Electron targets for E2E builds and guard cache action setup

* test: recognize the top-level cache repository preload

* docs: record E2E build timings and exact output parity

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 01:17:27 -07:00
OrcaWinandm4air 47cebbf5d2 ci: use ARM unit runners, overlap web builds, and reuse verifier fixtures (#23376)
* test: reuse isolated mobile bundle fixtures for verifier checks

* ci: run PR unit shards on ARM and overlap independent web builds

* docs: record controlled CI overlap and runner measurements

* test: observe WebRTC packets with the host clock

* ci: isolate Windows installer CIM probe from native test load

* docs: record native probe scheduling validation

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 01:06:54 -07:00
Neil bc78acc43e fix(editor): detect all bundled Monaco language associations (#23371) 2026-09-27 00:32:55 -07:00
OrcaWinandm4air 25c3ac400b ci: overlap shell setup, localization extraction, and mobile route preparation (#23368)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 00:28:02 -07:00
OrcaWinandm4air d8e2a694f6 ci: overlap package preparation and security scans; share localization parsing (#23364)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 23:53:19 -07:00
OrcaWinandm4air ccd1e87287 Overlap independent CI checks with native Actions background steps (#23351)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 23:09:40 -07:00
OrcaWinandm4air 9f5a8a5b8a Reuse mobile recording compilation and refresh desktop CI timings (#23343)
* ci: reuse recording compilation, split families, and refresh shard timings

* Keep recording suite intact after hosted performance comparison

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 22:44:15 -07:00
1c982ff2d9 fix(packaging): exclude root notes from app files (#23326)
Exclude root notes files from packaging while retaining nested runtime notes assets.

Co-authored-by: lurunzi <lurunzi@gmail.com>
Co-authored-by: Codex <noreply@openai.com>
2026-09-26 21:45:35 -07:00
OrcaWinandm4air b5dec85a4e ci: reuse mobile web route analysis and skip unrelated mobile tests (#23329)
* ci: share mobile route analysis and scope mobile test runs

* ci: cover mobile web runner process dependencies

* test: verify mobile web selectors through the new runner

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 21:37:40 -07:00
OrcaWinandm4air 1882458f44 ci: scope orcad smoke and parallelize Linux packages (#23314)
* ci: scope orcad smoke and parallelize Linux package formats

* ci: validate packaging when its copy dependency changes

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 20:44:15 -07:00