Commit Graph
2265 Commits
Author SHA1 Message Date
Brennan Benson 60bd1dfdea feat(native-chat): one shell-environment setting for every structured chat (#22387)
* feat(native-chat): one shell-environment setting for every structured chat

Structured Codex chats started from the login-shell environment, while
structured Claude chats started from Orca's own process environment, so a
variable exported in .zshrc reached one and not the other. Both now start
from the same base, chosen by a new setting:

- on (default): the whole login-shell environment, as a terminal gets
- off: Orca's environment plus PATH, locale, SSH_AUTH_SOCK, and the
  variable names the user lists

The setting is re-read each time a chat starts or resumes. It is shown
only when Chat UI, the Chat UI default view, and structured native chat
are all on. Terminal-backed chat is unchanged.

* fix(native-chat): normalize the shell-environment settings when a profile loads

A hand-edited settings file could store the variable list as something other
than an array, and the structured runtime called `.filter` on it per launch, so
a malformed value failed every structured chat create and resume, and the
settings pane render. Normalize both keys where the profile loads, the same way
the other array settings are, through one shared normalizer the runtime policy
also uses. Also pin that an uncommitted name draft survives an unrelated
settings re-render.

* fix(native-chat): keep the pinned account as the only source of a structured chat's Claude home

The session record owns which Claude home a structured chat uses, and the
acquisition pin (claudeConfigDirEnvPatch) is the only emitter of
CLAUDE_CONFIG_DIR, compared against what the child would otherwise inherit.
With the login-shell snapshot as the inherited base, a CLAUDE_CONFIG_DIR
exported only in a shell rc flipped that comparison and produced an explicit
pin to the CLI default home, which moves the CLI off its default Keychain item.

Drop the inherited CLAUDE_CONFIG_DIR in the Claude launch resolver before the
pin runs, as Codex already does for an inherited CODEX_HOME. A configured
per-agent overlay still passes through, since the record already honors it.

* fix(native-chat): drop Orca's own CLAUDE_CONFIG_DIR from a structured Claude child too

The process spawner merges Orca's process env under the launch env, so a
CLAUDE_CONFIG_DIR exported to Orca itself reached the child around the launch
resolver's drop and unseen by the account pin. One helper now strips it from
both inherited bases. Also declare the two shell-environment settings on the
runtime store contract and add the six new strings to every locale catalog.

* feat(native-chat): add shell variables one at a time with a removable list

* fix(native-chat): return focus to the name input after removing a shell variable

* fix(native-chat): use a neutral placeholder for the shell variable input

The empty input showed a grey HTTPS_PROXY as its placeholder, which reads as a
saved value, especially right after that exact entry is removed from the list.
Use "Variable name" instead, in every locale catalog.
2026-09-23 14:29:20 -07:00
Brennan Benson 845db9e5e2 fix(native-chat): underline only file links a click can act on (#22370)
* fix(native-chat): underline only file links a click can act on

A chat message could underline a bare file name such as `deck.md` that
resolved nowhere, and clicking it did nothing, so it read as a broken link.

- Inline code and quoted text become file links only when they name a path
  (contain a `/` or `\`), matching plain prose; a bare file name stays plain code.
- Every file link click now answers: it opens, or says the file was not found,
  that the host could not be checked, or that the path could not be resolved.
- Explicit links like [x](README.md:5) route as files, and linked text keeps
  `#`, `?` and `%XX` literally instead of re-parsing them as URL syntax.

* fix(native-chat): wrap the parsed file location so file URIs in chat text still open

Linkified prose, quoted text and inline code wrapped their display text, which the
literal wrapped-href route no longer URL-parses, so file:///... resolved as a relative
path under the worktree. Wrap pathText[:line[:col]] from the parsed link instead.
2026-09-23 13:29:24 -07:00
Brennan Benson 641a7f36d9 fix(native-chat): keep one live tool-run header from a call's start to the turn's end (#22432)
* fix(native-chat): keep one live tool-run header from a call's start to the turn's end

The collapsed tool run's header was two elements, one for "a call is running"
and one for "nothing is", chosen call by call. Every call start and end
remounted it, the count disappeared while a call ran and came back one
higher, and a call that finished inside a frame still bought the whole swap.
That is the 42→43 flicker in the report.

The header is now one element whose live state belongs to the turn, not to
any call: it stays live from the run's first call until the agent moves past
it (prose, a further run, or the turn's end), and settles in place. While
live the sentence speaks in the present tense and counts the call in flight
("Running 3 commands"), with the latest call's command beside it as a muted
preview; once settled it reads as before ("Ran 3 commands ✓"). The category
glyph is the run's in both states, and the completion mark only appears once
settled, so nothing pops between calls.

Which run is live is derived where the transcript is sliced into rows: the
last row that speaks or acts is the trailing one. A reasoning aside after it
leaves it live; an answer or a further run settles it.

Present-tense forms for the ten sentence categories are added to the shared
copy and the English catalog. The transcript-file lane, which renders with
the structured activity UI off, is unchanged.

* fix(native-chat): settle a run blocked on the reader, keep it live past an approval

- A run whose question is awaiting the reader's answer no longer pulses
  "Reading 1 file" while the agent is blocked; it falls back to its calls.
- An approval's receipt no longer moves past the run above it, so the call
  it just approved reads as running while it runs.
- The header button is the live region, so the count is announced too.
- Drop the unused live option and record from the shared English sentence;
  nothing renders it yet.

* fix(native-chat): stop the settled run's check from fading in on every mount

Windowing remounts settled rows as the reader scrolls, and a restored transcript
mounts them all at once, so the fade replayed where nothing had changed. Also
pin that the live header counts the next call on the same element.
2026-09-23 10:49:04 -07:00
Brennan Benson 563dd5487f feat(native-chat): show a Codex chat's goal above the composer, and set it from goal mode (#22377)
* feat(native-chat): show a Codex chat's goal above the composer and set it from goal mode

Structured Codex chat now treats the thread goal as session state: a banner above the
composer shows the current goal (pursuing / paused) with clear, pause/resume and expand;
/goal enters a goal mode whose send calls thread/goal/set; the objective is journaled as a
user message marked as sent as a goal. The banner is derived from the journaled goal rows,
which Codex's resume snapshot refreshes, so a reopened or adopted chat shows its goal.

Fixes STA-8159

* fix(native-chat): replace a recorded goal by clearing first, and recover a lost goal-change response

- A set while the journal records a goal (any status) clears it before setting,
  so the new goal starts with its own time and token counters instead of
  rewriting the old goal's objective in place.
- The threadGoal plan answers an unknown outcome from the goal the journal
  records and reruns otherwise, so one request timeout no longer refuses every
  later Clear/Pause/Resume as unknown for the mounted session.
- The goal-mode chip says "Exit goal mode"; "Clear goal" stays the banner's
  action on the provider goal.
- A typed bare /goal on Enter enters goal mode, the same as picking it.
- The renderer reads the goal off the tail of its ordered snapshot; the host's
  unordered map keeps the by-sequence reader.
- Drop the composer's duplicate in-flight guard; the goal controller already
  serializes changes.
- Pin that a counter-only revision reaches a subscriber's live page under its
  original sequence.

* fix(native-chat): keep a bare /goal inside goal mode as the entrance, and pin goal delivery and serialization

- A bare `/goal` submitted while already in goal mode re-enters the mode instead
  of setting a goal whose objective is the literal text "/goal".
- The counter-only revision pin now drives the host's own event sink bound to a
  real journal, so it goes red when the publish after a lifecycle transition is
  dropped; the previous fake sink never published.
- Pin that a set which threw after journaling its objective puts that objective
  back exactly once when the ledger reruns the same operation id.
- Cover the goal controller hook: absent without host support, the loaded window
  wins over the host's answer, a second change while one is unsettled answers
  false without a request, and a refused change frees the next one.

* fix(native-chat): resume a blocked or usage-limited goal, and keep goal-mode drafts honest

- The goal bar offers Resume on a blocked or usage-limited goal, which the
  provider resumes exactly as it resumes a paused one; a goal whose token budget
  is spent still offers only Clear. The rule lives beside the other goal facts
  in shared code so every reader answers it the same way.
- A `/goal <text>` typed inside goal mode sets the objective `<text>`, as it
  does outside goal mode, instead of a goal whose objective is the literal
  command.
- Setting a goal is a host round trip; a draft edited while it was in flight is
  no longer wiped when the goal lands, matching every other host command.
- Pin that a lost status-change response is read as applied only when the
  recorded goal is in that status, that a cleared row in the loaded window
  outranks the host's earlier answer, and that the PTY lane is untouched.

* fix(native-chat): keep the load-older anchor on the loaded window when a live revision lands below it

A live revision of a row keeps that row's original sequence. When the row is
older than the client's loaded window, the shared reducer merged it in and it
became the load-older anchor, so paging `before` it skipped every row between.
A goal's counter-only revisions during a long goal turn reach any client that
attached after the goal row left its window, so a reopened chat lost rows on
scroll-back.

The reducer now admits live rows only at or above the window's oldest row
while older rows remain on the host; the journal keeps the revision and the
page reader serves it once the window reaches the row. With nothing older on
the host the window is the whole journal, so a row below the head is admitted
as before.

Also drain accepted provider events before a goal set reads the journal to
decide whether it replaces a recorded goal.
2026-09-23 10:34:06 -07:00
Brennan Benson a375936c04 feat(agent-launch): let a caller reserve the pane its terminal launch creates (#22291)
* feat(agent-launch): let a caller reserve the pane its terminal launch creates

* fix(agent-launch): refuse a launch whose reserved pane is already live

* fix(agent-launch): refuse a live reserved pane before it is revealed

The live-pane refusal used to fire in the executor, after createTerminal
had already issued a handle, published the mobile snapshot and revealed
the tab. The reveal re-registered a fresh launch config over the running
agent's. agent.launch now passes requireFreshPane with a reserved pane,
and createTerminal throws AgentLaunchPaneAlreadyLiveError as soon as
spawn reports it attached to a live pane. That is before any handle,
snapshot or reveal. The spawn reattach itself is the one terminal.create
already uses, so the live PTY is never killed, and the stable-pane
create claim is still released in finally. The isReattach plumbing
added to the launch factory for the old check is gone.

A replay-safe launch refused this way on an existing workspace now
records a failed ledger row, the same way a name collision does.
Before, the row stayed claimed, so every retry got
agent_session_operation_unknown. agent.launchReplay passes the code
through. On create-worktree the workspace already exists when the
terminal is refused, so the row stays unknown. The code is added to
the runtime passthrough list so callers can branch on it.

The pane key is now in the replay fingerprint, deliberately. It is not
placement: group, anchor and focus still stay out of the request and
out of the ledger. It is identity. It is written into the pane's PTY
environment and names the tab the caller has placed. A retry that
reserved a different pane is therefore a different request. Replaying
the first answer would return a key the new reservation can never
find. This matches terminal.createAgentSession, which also fingerprints
its tab and leaf ids. The key is only folded in when present, so every
existing digest is unchanged, and a test pins that.

The wire schema now refuses a tab id the runtime would not adopt as
sent: one with surrounding whitespace, which the runtime trims, and one
longer than 512 characters, which the spawn reservation does not key
on. It reuses the tab-id schema that Placement uses.

* test(agent-launch): pin that a refused live pane issues no handle

The refusal test named handle issuance but only asserted the reveal, so a
throw moved to just before the reveal would still pass. Assert no terminal
is registered, with the attach test as the positive control.
2026-09-23 10:18:16 -07:00
Neil eb18eaf2b6 feat(usage): add Muse Code local usage provider (#22379)
* feat(usage): add Muse Code local usage provider

Scan Muse session logs (including subagent logs, which hold usage the parent
log does not) for model_completed token events and surface them as a fourth
local usage provider: shared scan worker, persisted per-file cache reused by
mtime/size, cross-log dedupe, Stats tab, and Usage Overview integration.
Muse logs carry no price, so the provider reports tokens only.

* fix(usage): name Muse in Stats & Usage copy; skip partial-cost warning when nothing is priced

* fix(usage): surface unreadable Muse sessions root; name Muse in remaining Stats & Usage copy

* fix(usage): count distinct same-content Muse records within one log
2026-09-22 22:44:10 -07:00
Brennan Benson 8757e40063 fix(native-chat): keep a structured agent's tool line between tool calls (#22349)
* fix(native-chat): keep a structured agent's tool line between tool calls

A structured session's status named a tool only while the call was still
running, so the sidebar's tool line went blank whenever the agent was
thinking or writing between calls. Terminal agents keep naming the finished
tool until the next one starts, and clear it after a failure. The structured
status projection now does the same: a running call wins, otherwise the
turn's newest root call if it completed.

* fix(native-chat): bound the structured tool line by the running turn, not the user row

A send made while a turn is running writes its user row into the journal
straight away, and the turn keeps going. Stopping the scan at that row
blanked the tool line while a tool was still running. The scan now runs to
the turn record and names a call only when that record is still running, so
a turn that already ended never lends its last tool to a pending follow-up.

This lookup was the running-only lookup's only production caller, so it
replaces that lookup instead of sitting beside it.

* fix(native-chat): keep naming a structured agent's failed tool until the next one

Clearing the tool line after a failed call brought the blank gap back for
much of a turn: Codex marks any nonzero exit as failed, so a search with no
match or a red test run is enough. The failure already shows on the tool's
own row in the transcript. The running turn's newest running call still
wins; otherwise its newest root call is named whatever it settled to.

* fix(native-chat): name a structured Codex edit on the tool line as the chat draws it

Once a Codex edit's changes exist, its apply_patch call becomes a diff row,
which the status lookup skipped, so the row named the command before the edit.
The chat's tool-call block for a journal row now comes from one shared builder,
and the status lookup reads the same definition: a diff is named as Diff with
its path, and counts as settled since it carries no lifecycle.

* docs(native-chat): describe the structured tool field as running-or-latest

The status summary's toolName/toolInput now name the running turn's
latest tool between calls, not only a running one. Update the wire type
and status bridge comments that still said "the running tool".
2026-09-22 22:42:31 -07:00
Jinwoo Hong 996f9cc306 feat(mobile-web-bundle): gzipped 384 KiB ranges over a capability-negotiated mobileWeb.bundle.range (OTA phase C follow-up) (#22381)
* feat(mobile-web): serve gzipped 384 KiB bundle ranges behind a capability

Adds mobileWeb.bundle.range with its own strict params and result, so
shipped chunk readers see no reply change. The host gzips each range at
level 6 and sends identity when gzip does not shrink it, sharing the chunk
method's read-slot budget and per-asset verification. status.get
advertises mobileWeb.bundle.range.v1 beside mobileWeb.bundle.v1, and the
method is allowlisted for paired phones.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): read the bundle range capability and range replies

Adds the range reply reader and operation, and picks range or chunk from
the status.get capabilities the connection already proved, so an older
desktop keeps being paged in chunks with no probe round trip.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* perf(mobile): keep four bundle chunk reads in flight across the whole manifest

The fetch ran one worker per asset and paged inside an asset sequentially, so
the largest script's 71 chunks were 71 serial round trips while the other
readers idled. One window of four chunk reads now covers every (asset, offset)
on the host's chunk grid, largest asset first. A read_limited refusal narrows
the window and retries the read; eof is still read from the reply.

Synthetic manifest (one 71-chunk asset, five small): 72 round trips -> 19.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore(mobile): add fflate 0.8.2 for gzip bundle ranges

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): decode gzip bundle ranges into a bounded buffer

Inflates each range into a buffer one byte past its window, so a gzip
bomb costs at most that allocation and an overlong body is visible. A
corrupt, truncated or unknown-encoding body refuses as range-undecodable;
a body of the wrong decoded length refuses as range-length-mismatch.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): pass the bundle read method from the session to the fetch

The download reads the capabilities of the gates the reducer decided
under and hands the fetch range or chunk. The fetch does not act on it
yet; the range read lands on the pipelined window.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the bundle-fetch family under pipelined reads

Baseline moves to a1ee317368, the pipelined fetch.
781 goldens change only their `baseline` header line. Six bodies move:
mobile-web-bundle-fetch-paged, mobile-web-bundle-build-changed, and the four
matrix-mobileweb.bundle-fetch-* goldens.

The two bundle-fetch scenarios now bind requests in pipelined order, largest
asset first, with every chunk sent before any reply: index.html@0 (#1),
index.html@16 (#2), assets/app.js@0 (#3).
- fetch-paged: the request set is identical, only reordered. The chunk
  sender names/ordinals and the scenarioSha256 moved; the replies and the
  fetched bytes did not.
- build-changed: the same reorder, plus one request that is new because
  pipelining puts it in flight before the refusal lands (index.html@16).
  The refusal and the checkpoint are unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): the bundle chunk comment no longer says a reply picks the next offset

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): read gzipped bundle ranges on the pipelined window

A host that advertised mobileWeb.bundle.range.v1 is paged in 384 KiB
ranges on the range grid, through the same four-read window and queue as
chunks; any other host keeps the chunk grid. Each range is decoded to its
exact window length before the fill checks and the asset hash.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus after the bundle comment fix

Baseline moves from a1ee317368 to e94bde327d,
the comment-only commit on a fenced path. All 787 goldens and the scenarios file
change only their `baseline` line. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the corpus at the range-read pin

Repins baseline to the range-read commit and re-records every golden.
Only the baseline and lockfileSha256 headers move: the lockfile gained
fflate, and the bundle-fetch adapter pages the chunk path, whose
requests and replies are unchanged, so no golden body moves.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): state the on-settle reason that holds for pipelined bundle reads

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus after the on-settle comment fix

Baseline moves from 252c592b52 to fe41226ef5,
the comment-only commit on a fenced path. Re-recorded: all 787 goldens and
the scenarios file change only their baseline line. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore(mobile): keep the lockfile's patch block in main's form

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus after the lockfile patch-block restore

Baseline moves from fe41226ef5 to 84d6fca6e7,
which restores main's patchedDependencies form in mobile/pnpm-lock.yaml.
Re-recorded: all 787 goldens change only baseline and lockfileSha256, and
the scenarios file only baseline. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): pool four workers over planned bundle chunks, report progress per chunk

Design-review fix round, sketch C: four workers take reads from one planned
chunk queue, largest asset first. They replace the central pump and the
read_limited narrowing. A host frees its read slot before it replies, so a lone
fetch capped at four cannot trip the limit. A refusal now fails the fetch, as
it did on base, and stops the other reads.

- Each asset's buffer is allocated when the plan is built. That removes the
  nullable buffer and its guard. The per-asset byte count is gone, and the
  hash is the oracle (S1, S2).
- The caller's signal is checked before each read and once after the pool
  drains, so an abort during the final window rejects with fetch-stopped
  (S3). The stopped check now covers only the caller's abort. The internal
  stop only makes late replies skip checks, hashing and progress (S4).
- Progress is reported per accepted chunk. completedAssets still counts on
  completion (S6).
- Renames: MAX_CONCURRENT_CHUNK_READS, and `reply` for the RPC reply (N1).
- The slot check states exact-slot acceptance once, then classifies the
  refusal (N3).
- Stale test titles and comments are renamed (N4).

The synthetic manifest still takes 19 round trips.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say why bundle reads settle at on-settle under pipelined chunks

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile-web): announce bundle ranges on the manifest reply

The manifest reply now names the range grid in an optional rangeBytes,
beside chunkBytes, and the status capability is gone. The range method
takes exactly the chunk params on that grid instead of a caller length.
Both methods share one verified read that returns the six-field header,
and the range handler checks the connection again before deflating.
Range schemas move into the bundle RPC contract; SHA256_PATTERN is shared
from the manifest contract. Shared refusals are tested once over both
methods.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): drop the range capability read and its session threading

The phone will read rangeBytes off the loose manifest reply instead, so
the read-method module goes and the session effects and hook return to
the pipeline branch's version. Range imports move to the bundle RPC
contract, the reply reader reuses the shared SHA256_PATTERN, and a new
test pins that node's level-6 gzip from the host encoder inflates with
fflate to the same bytes. The fetch and window-read modules still import
the deleted names until part 2.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the bundle-fetch family under per-chunk progress

Baseline moves to 34fda6f62e. 782 goldens and
the scenarios file change only their `baseline` line. Five bodies move:
mobile-web-bundle-fetch-paged and the four matrix-mobileweb.bundle-fetch-*
goldens. The only change is bundle-progress effects. One report now lands
after the first accepted chunk of index.html (0 assets, 16 bytes), and the
later progress ordinals shift by one. Requests, replies and fetched bytes
are identical. mobile-web-bundle-build-changed keeps its body, because its
one accepted chunk is the whole of assets/app.js.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): page bundle ranges through one window reader chosen by the manifest

The fetch keeps the pipeline's four-worker pool and builds one window
reader from the manifest reply: ranges on rangeBytes when the host names
it, chunks on chunkBytes otherwise. The reader returns the six-field
header and a lazy bytes() so the stop and misroute checks run before any
decode. A range that inflates to the wrong length now falls to the slot
checks, with the one-byte-over buffer as the memory bound, so
range-length-mismatch is gone. A rangeBytes this build cannot page reads
as absent. Fetch names say window, not chunk.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): drain the fake host after the fetch settles so the sibling-stop bounds can fail

The wave host stopped releasing replies once the fetch settled. Reads that
should have been stopped were never answered, so the read_limited bound (7)
and the chunk-failure bound (5) held even with no sibling stop at all. It now
drains until nothing waits. With the worker's stopped.abort() removed, both
bounds fail at 76 requests. assertChunkDescribesAsset's parameter is renamed
to `reply`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the window-reader commit

Baseline moves to a25a355546. Re-recorded:
every golden and the scenarios file move only baseline, and the five
bundle-fetch goldens also move lockfileSha256 to this branch's lockfile.
Every golden body is byte-identical to the pipeline branch's. The
recording adapter's scripted host names no rangeBytes, so the bundle
family still records the chunk path.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus after the sibling-stop test fix

Baseline moves from 34fda6f62e to 97b13ec7f0.
All 787 goldens and the scenarios file change only their `baseline` line. No
golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the second pipeline merge

Baseline moves from a25a355546 to 4d2cab31e5,
the merge of the pipeline's sibling-stop test fix. Re-recorded: every golden
and the scenarios file move only baseline. Against the pipeline branch, only
baseline and lockfileSha256 differ; every golden body is identical.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read the bomb inflation without depending on call order

With the fetch's sibling stop removed, a read left over from the previous
test inflated into the bomb test's record first, and indexOf(601) picked
it. The test now asserts some inflation stopped at 601 and none exceeded
its buffer, whatever else ran.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the bomb-test fix

Baseline moves from 4d2cab31e5 to 73fde15487,
the test-only commit on a fenced path. All 787 goldens and the scenarios
file change only their baseline line. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile-web): tighten the bundle window contract and pin the range sibling stop

The range bomb test reads only its own host's inflations, keyed by the
gzip bodies that host sent, and plans twenty reads so a missing sibling
stop is visible: with stopped.abort() removed it sends all twenty.
The range params are an alias of the chunk params, and the chunk data
bound is the exact base64 length of a full chunk. The phone's chunk and
range replies share one header shape. The window reader closes over the
client and bytes() takes the slot length the fetch computes. The host's
positional read is readMobileWebBundleAssetWindow.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the window-contract commit

Baseline moves from 73fde15487 to 7346e005a3.
All 787 goldens and the scenarios file change only their baseline line.
No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the main merge

Baseline moves from 7346e005a3 to 9c0fe1a546,
the merge of main at 98a6a5325c. Recorded with --record: all 787 goldens
and the scenarios file change only their baseline line. No golden body
moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile-web): drop a test cast and shape-named field maps

The bomb test's inflation log is typed by its hoisted factory's return
instead of an assertion, and the zod field maps shared by the bundle
window schemas are windowParamsFields and windowHeaderFields.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the lint fix

Baseline moves from 9c0fe1a546 to f4f0915e70.
Recorded with --record: all 787 goldens and the scenarios file change only
their baseline line. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): page the range fetch fixture on a small advertised grid

The fake host names a 4 KiB range grid and a 1 KiB chunk grid on its
manifest reply, which the phone pages as it would the real ones, so the
fixtures shrink to a few KiB with the same shapes and each asset is
hashed once. The file runs in about 360 ms instead of 3.5 s, which a
loaded CI runner pushed past the 5 s test timeout. The desktop range
suite still pins that the real host names MOBILE_WEB_BUNDLE_RANGE_BYTES.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the range fixture fix

Baseline moves from f4f0915e70 to d4d6aadea2.
Recorded with --record: all 787 goldens and the scenarios file change only
their baseline line. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-23 01:34:49 -04:00
Neil 51d3cafc4f feat(editor): add a setting to turn off preview tabs (#22398)
Single-clicking a file in the Explorer, or following a link in Markdown
source, opens it as a preview tab that the next preview open replaces.
There was no way to turn that off, so browsing files kept swapping one tab.

Adds `editorPreviewTabsEnabled` (General -> Navigation, on by default).
A caller's `preview` flag is now an intent that `resolveEditorPreviewIntent`
resolves against the setting, covering every open path - files, diffs,
history diffs, conflicts - in one place.

Preview-ness is derived rather than reconciled: readers treat a tab as a
preview only when the stored flag and the setting agree, so a flag left
over from a saved session, another window, or a host switch is inert while
previews are off. Nothing rewrites stored flags when the setting changes,
so no settings-landing path has to remember to clean up.

Fixes #22397
2026-09-22 22:33:06 -07:00
Neil 52a1e2875b feat(orchestration): accept Muse model and effort for supervised workers (#22383)
* feat(orchestration): accept Muse model and effort for supervised workers

`worker-start --agent muse` already launched, but `--model` was refused because
Muse had no session-option catalog. Add one that maps worker preferences to
`muse --model <id>` and `--reasoning-effort <level>`; it seeds no models, so
native-chat surfaces show no picker.

opencode stays without `--model`: the opencode 2 TUI (now shipped as
`opencode`) rejects the flag, so the refusal now tells callers to rely on the
agent's own config. Help, skill guide, and docs list valid `--agent` ids and
the agents that accept `--model`.

Refs #19823

* test(mobile): repin session route closure for the Muse option catalog
2026-09-22 22:20:35 -07:00
Neil 83dd047fd9 fix(explorer): find files by name in large local workspaces (#22369)
* fix(explorer): search local workspaces by file name across every file

The Explorer name filter only searched remote workspaces directly; local
workspaces still filtered the first 20,001 listed files, so files beyond
that cap never matched in large repos. Local name queries now rank the
whole workspace on the host, and fall back to an uncapped git listing
when ripgrep is not installed.

* fix(explorer): filter capped local listings on the host with the Explorer word rule

Replaces the Quick Open fuzzy top-32 routing, which dropped multi-word
matches and capped visible results. The Explorer keeps its instant
renderer-side filter; only when the local listing hits its cap does it
re-list on the host with the same word rule applied before the cap.

* fix(explorer): keep capped matches when host name filtering fails

- Fall back to the capped listing (and stop re-listing) if the host scan fails
- Keep primary matches when the ignored-file pass fails during a filtered scan
- Key host scans on normalized filter words; reset capped state per filter session
- Bound nameFilter size at the IPC boundary; drop the double readdir walk

* fix(explorer): match name filters without locale-dependent lowercasing

* fix(explorer): avoid render-time ref writes in the host name filter fallback
2026-09-22 19:39:52 -07:00
NeilandAdrien De oliveira ebed0964a2 feat(agents): add first-class Muse Code harness (#22216)
* feat(agents): add first-class Muse Code harness

Add Muse as a supervised Orca agent across desktop, mobile, session history, source control, local hooks, SSH, WSL, and native Windows. Preserve user settings, support Muse 1.3 hook environment allowlists, and recognize versioned foreground processes. Include question, waiting, completion, resume, and readiness coverage.

Co-authored-by: homesh-dev <300847526+homesh-dev@users.noreply.github.com>

Co-authored-by: jeffhuen <32542276+jeffhuen@users.noreply.github.com>

Co-authored-by: John Cusack <johncusackccm@gmail.com>

Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com>

* test(agents): cover Muse remote hook registration

* test(agents): cover Muse hook and source-control contracts

* test(agents): exclude Muse hook metadata from script mode check

* test(agents): keep Muse skill picker coverage stable

* test(ai-vault): include Muse in every-agent fixture

* test(mobile): repin Muse agent icon closure

* fix(muse): detect questions and approvals from structured Muse signals

Muse 1.3 fires no hook for request_user_input, so a pending question left
the pane "working". Its internal reminder subagents also post hooks with
their own session ids (even after Stop), which surfaced "tool failed" rows
and flipped finished panes back to working.

- Read pending questions from Muse's session log
  (user_input_prompt_requested/settled) via the existing transcript poll,
  now generalized from Codex subagents to Muse on main and relay.
- Drop child-session hooks (SubagentStart ids, or turn_id === session_id).
- Treat Notification permission_prompt as the approval wait; PermissionRequest
  also fires for auto-approved calls, so it only caches the approval card.
- Ignore Notification copy as the prompt; poll replays are not new prompts
  or turn boundaries.
- Allowlist USERPROFILE so Windows cmd AutoRun doesn't fail every hook.

* perf(muse): parse only question events from the session log

Most Muse session-log lines are large model/tool records. Filter raw lines
by the user_input_prompt_ marker before JSON.parse via an optional
readJsonlCursor line filter.

* fix(muse): unwrap batched log records and scope questions to the live turn

Review follow-ups: question events inside retained_frame batches were
skipped, and a question left open by a crash or interrupt stayed pending
for the pane's life. Share the history scanner's retained_frame unwrapper,
and only report a pending question whose run_id matches the hook turn_id.

* refactor(muse): drop type assertion in retained_frame unwrap

* fix(agent-hooks): satisfy exhaustive-switch lint in transcript poll policy

---------

Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com>
2026-09-22 19:13:11 -07:00
Jinwoo HongandDavid Bebawy 7c46a69049 feat(telemetry): report the macOS daemon's code identity on adoption and folder-denial events (#22171)
* feat(daemon): import the macOS process code-identity probe from PR #21826

Takes `daemon-mac-code-identity.ts` and its test verbatim from David Bebawy's
community PR #21826 (stablyai/orca). The probe asks Security.framework, via
`codesign --display --verbose=1 +<pid>`, where a live process's code lives on
disk — the question Node cannot answer, and the one that decides whether tccd
can still resolve a running daemon's code identity after an app update.

Imported unchanged here so the adaptation that follows is reviewable as a diff
against the author's original.

Co-authored-by: David Bebawy <david.ayad2@gmail.com>

* feat(telemetry): report the daemon pid's macOS code identity on the two adoption events

Community PR #21826 argues that macOS terminal daemons lose Documents/Desktop/
Downloads access after an update because the daemon's own executable is
unlinked — Squirrel parks the outgoing bundle under a ShipIt staging directory
and later deletes it — so tccd can no longer map the daemon pid to on-disk
code. Today's `spawner_path_class` and `tcc_attribution` read the binary that
forked the daemon, which an in-place update deletes and recreates, so neither
can see that state.

This adds the detector as a measurement only. `code_identity` rides on
`daemon_adopted` and `daemon_pty_cwd_denied`, the two events that already
describe an adopted daemon, so denied daemons can be cross-tabbed against
healthy ones. Nothing reads the verdict: no replacement, no notice, no UI.

The probe is David Bebawy's, narrowed from a path-carrying union to the closed
enum the wire allows, and memoised per pid so one codesign spawn answers for a
whole daemon generation. Off macOS, or with no pid, it reports `probe-failed`,
which keeps both schemas strict and non-optional.

Co-authored-by: David Bebawy <david.ayad2@gmail.com>

* fix(telemetry): read the daemon's code identity fresh on every adoption event

The probe memoised its verdict per pid and never expired it, so
`daemon_pty_cwd_denied` reported whatever the probe saw at adoption rather than
what was true at the denial. That breaks the measurement in both directions: a
transient codesign failure during startup pinned `probe-failed` for the rest of
the run, and the `parked` to `unresolvable` transition became invisible.
Squirrel leaves the parked bundle in place until the next update, which can be
days, so a daemon adopted as `parked` and denied as `unresolvable` is the exact
crossover this study exists to catch, and the cache hid it.

Now every ask runs its own codesign. Only concurrent asks about the same pid
share a probe, and that entry is cleared as soon as it settles, so nothing
survives to be reported later. Both events are rare enough that one spawn each
is not worth a cache.

* fix(telemetry): drop the dead existence check from the code-identity probe

The classifier stat'd the path codesign displayed and called a missing one
unresolvable. That path is unreachable: once the executable is unlinked,
`codesign --display` prints no `Executable=` line at all and exits 1 with
"No such file or directory", which the fallback below already classifies as
unresolvable. Verified directly on Darwin 25.5 against a signed binary deleted
out from under a running pid.

All the branch actually covered was the window between codesign reading the
path and this process stat'ing it, and it paid for that with a synchronous
stat on the main thread.

* fix(telemetry): never classify a timed-out codesign probe as a verdict

`runProcess` kills the child at the deadline and reports `timedOut`, but the
runner type dropped that field, so a codesign killed mid-display could still
have printed an `Executable=` line and been read as `resolved` or `parked`.
A half-written display proves nothing about where the daemon's code lives.

The runner result now carries `timedOut`, and a timed-out probe returns
`probe-failed` before the output is looked at.

* docs(telemetry): state what each code-identity verdict actually asserts

A reviewer read `resolved` as a claim that the executable sits inside the
installed app and asked for that to be validated. It is not that claim, and we
are not making it: proving containment needs the pid record's spawner path, and
deciding anything from where the code lives is #21826's proposed behaviour
rather than this measurement.

The enum doc now spells out all four verdicts in the terms the probe can
actually support, and says plainly why `resolved` stops at "exists and is not
parked". A matching note sits beside the parked-path pattern.

* docs(telemetry): stop asserting how long a parked bundle survives

The probe's rationale claimed Squirrel keeps the parked bundle "until the next
update". A reviewer claimed the opposite, that it is deleted at the end of the
same install. Neither holds up against this Mac's ShipIt log: the install moves
the outgoing bundle to a TMPDIR ShipIt directory and logs no removal of it at
all, and the one "Couldn't remove owned bundle" line names the incoming
download staging copy, not the parked one. Every parked bundle from the last
two days is nevertheless gone now.

So the rationale in the probe doc, the enum doc, and the reprobe test comment
now assert only what is established: the outgoing bundle is moved aside at
install and disappears later on a schedule we have not pinned down. That is
already enough to justify the design, since one pid's verdict can change
within an app run, which is exactly why every ask reads fresh.

* feat(telemetry): report readable TCC-gated spawns as the code-identity control

`daemon_pty_cwd_denied` gives code_identity's hit rate on denials, but a
readable spawn emitted nothing, so an `unresolvable` adoption with no denial
could not be told apart from a user who never opened a terminal in Documents,
Desktop, or Downloads. The false-positive rate that gates #21826's
auto-replacement was unmeasurable.

`daemon_pty_cwd_readable` now fires when a daemon reads a TCC-gated cwd, once
per daemon and folder class per app run, with the same origin properties as
the denial event. The read-out becomes a 2x2 of code_identity against
readable/denied on protected-folder spawns. Fire-and-forget on the spawn path
like the denial emit, and no app-side directory read.

* refactor(telemetry): one emitter and schema for both cwd verdicts, no dedupe state

The once-per-daemon dedupe on `daemon_pty_cwd_readable` was keyed before the
probe ran, so a daemon first seen readable while `parked` never reported again
once it turned `unresolvable` — the one cell that would count most against
#21826. It also counted per daemon while denials count per spawn, so the 2x2
mixed units.

Readable now reports every spawn, like denied, and both events share one
emitter (`trackDaemonPtyCwdVerdict`) and one schema. The TCC-folder gate lives
in the verdict branch. The origin fields are one shape spread into both
schemas. The codesign probe calls `runProcess` directly and tests mock it,
replacing a test-only runner parameter. The repeated "never cached" rationale
is now said once.

* fix(telemetry): rename the shared origin schema fields for the anti-slop gate

no-shape-in-symbol-names rejects daemonOriginShape; the fields are event props.

---------

Co-authored-by: David Bebawy <david.ayad2@gmail.com>
2026-09-22 22:10:51 -04:00
Brennan Benson 1b85be67d8 feat(native-chat): notify on every settled structured turn (#22105)
* feat(native-chat): notify on every settled structured turn

A structured chat that finished while you were elsewhere lit the sidebar
but never raised an OS notification, and a notification that did arrive
for one could not open the chat it came from.

Unread and delivery now come out of the single resolveAgentAttention
decision the terminal lane already uses: the structured dispatcher calls
applyAgentAttention instead of applyAgentAttentionUnread, so the same
policy that decides what to light also decides what to deliver, through
the same sound and blocked-permission tail.

Every settled turn notifies, as the CLI lane does. Success says
"finished"; failure and cancellation say "stopped" through the shipped
agentInterrupted flag rather than a second vocabulary. A turn whose
outcome the host never stated stays unknown and lights nothing.

The host now dedupes mobile fan-out by event identity (scope, session,
turn) beside the existing per-workspace burst cooldown, so a completion
two windows both saw reaches the phone once while each window still
decides its own banner. Clicking a structured notification reveals the
chat tab: its pane key's leaf is synthetic, so focusTerminal would hunt
a split-layout leaf that does not exist.

* fix(notifications): spend each mobile gate only when it actually notifies

Two review findings on the structured-chat notification lane, both real.

The mobile event gate consumed its reservation before the per-workspace
burst cooldown ran. Two chats in one workspace share that cooldown key,
so the second chat's completion could burn its event key and then lose
the cooldown to the first chat — never announced, yet permanently marked
as announced, so a later window dispatching it could no longer reach the
phone. The gate now peeks first and records the event at dispatch, which
also keeps a known duplicate from burning the cooldown slot.

A notification id is minted from the status row's stateStartedAt, and the
row re-projects that field as the turn settles: the working episode's
start moves into stateHistory and the settled start takes its place. A
banner raised in the window before that re-projection therefore carried
an id acknowledgement never rebuilt, leaving it on screen for good.
Acknowledgement now collects ids for the row's left episodes too — the
same episodes the unread check beside it already scanned, so the two
halves finally read the same turns. Lane-neutral: the terminal lane
mints its ids the same way and had the same gap.

* fix(notifications): drop the mobile event gate and reveal chats in folder workspaces

The per-event mobile dedupe defended against one completion being
dispatched by several Orca windows. Only one renderer mounts the
structured attention bridge, the completion feed is live-only with no
replay, and any in-process duplicate lands inside the existing 5s
per-workspace burst cooldown, which already collapses mobile and
desktop alike. The gate never acted on a real sequence, so the wire
field, the shared ledger and its tests go; mobile delivery is back to
main's behavior.

A folder workspace id ("folder:<id>") has no "repoId::" prefix, so the
click binding was skipped and clicking a chat notification there did
nothing. The chat route selects its workspace itself through
ui:focusEditorTab, so it now binds without a repoId; the terminal
route is unchanged.

* fix(notifications): retire the banner ids actually dispatched, not ids rebuilt from a moved row

A banner's id is minted from the status row's stateStartedAt at dispatch, and that field moves
afterwards: a completion can outrun the settled re-projection, and a settled structured row is
re-stamped with no history entry by any later journal row (a cancel appends a status note after
the turn settles). Rebuilding ids from the row's episodes at acknowledgement missed the second
case and fanned out up to 21 mobile dismissals per pane for ids never raised.

The shared delivery tail now records each dispatched id per subject; acknowledgement retires
those plus the current-row rebuild it always had. The acknowledgement collector is back to
main's single-field form.

* refactor(notifications): retire announced notifications by subject in main

Main now records, per pane, the ids it actually announced (a desktop banner
shown or a phone alert sent) and an acknowledgement passes the acknowledged
pane keys so main retires all of them. This replaces the renderer-side record
of dispatched ids: main is where the announcement happens, so it records only
real announcements, including phone alerts whose desktop banner focus
suppressed. The id rebuilt from the current row stays as the fallback after
a restart empties the in-memory record.
2026-09-22 18:35:18 -07:00
Jinwoo Hong dfff3915c4 fix(browser): scope back/forward/reload/zoom/grab shortcuts to the originating split (#22340)
* fix(browser): scope back/forward/reload/zoom/grab shortcuts to the originating split

With two browser panes visible in a split, Back, Forward, Reload, Hard
Reload, page zoom and Focus Address Bar fired in every visible pane. Main
forwarded these guest chords without the page id, and each split's active
pane subscribed. The renderer-side listeners for the same chords were also
window-wide per pane, so a key pressed in the toolbar (or in a terminal in
another split) reached every active browser pane.

Guest-forwarded chords now carry the originating browserPageId; preload
admits only well-formed payloads and each pane ignores ids that aren't its
own. Toolbar-path listeners use the same focused-split scope Find already
uses. The streamed remote pane's history chord moves onto that scoped hook.

Cmd/Ctrl+C grab (STA-3319) gets the same scope and no longer arms while a
text selection exists outside the browser pane, so copying from the native
chat transcript works again.

* refactor(browser): simplify split shortcut scoping per review

Drop the preload payload admission (main and preload ship together), fold
the three inline scope checks into browserChromeShortcutOwnsEvent, and
replace the outside-overlay selection check with a plain live-selection
rule so Cmd+C copies from surfaces that do not move split focus.

* refactor(browser): share one zoom command type and tidy shortcut comments

BrowserPageZoomEventDetail and BrowserPageZoomCommand were the same shape;
keep one in shared/browser-page-zoom.ts and route guest and local zoom
through a single handler.

* refactor(browser): narrow the zoom event with instanceof instead of a cast

* test(e2e): pin split-scoped browser shortcuts

Two browser splits (and a terminal beside a browser) now prove that Back,
Forward, Reload, Hard Reload, page zoom, Focus Address Bar, and the element
grab chord act only on the split that sent them, from both the guest page and
the browser toolbar. A native chat selection proves Cmd/Ctrl+C copies instead
of arming grab. Split fixtures move to a shared helper so both specs reuse them.
2026-09-22 20:14:19 -04:00
Brennan Benson 9ece273056 fix(native-chat): journal rows name the agent that produced them (#22299)
* feat(native-chat): carry producer linkage on every journal row

A journal is the durable record of one agent SESSION, and a session that runs
subagents journals their rows into the same timeline with nothing on the row
saying which agent wrote it. Add that: a per-row linkage bundle naming the
producing agent, its parent, the provider's raw parent reference as provenance,
the kind of work, and which run of the agent produced the row.

The bundle rides the row BASE, not the body: two nested prompt shapes are
strict, so an unknown key on a body makes the whole row parse as malformed. It
is deliberately not a schema-version bump either — an unknown `v` makes a row
unreadable and latches the host read-only, while an unknown key is ignored, so
an older host reads a stamped row and behaves exactly as it does today.

One reader predicate interprets absence, by presence and not by truthiness: an
id that failed to resolve is still an id, and a truthy test would read it as
root and put the child's content back on the parent. The parent-facing status
scans — thinking, the running tool call, the latest assistant line and the
quoted prompt — now skip rows a subagent produced. The transcript is left
unscoped on purpose: it shows every agent's output.

No producer stamps anything yet; this is the carrier and the reader.

* fix(claude): attribute a subagent's journal rows to the subagent

The Claude translator already parsed `parent_tool_use_id` on every envelope and
threw it away. It now resolves that reference to the producing agent's canonical
task id — never to the reference itself, which names the tool CALL and is
re-minted on every resume, so a row stamped with it would split one child into
two the moment it resumed. The raw reference is kept beside it as provenance.

Resolution is its own module rather than more roster: the roster maintains the
spawn-group row a user reads, while this answers, for one frame's parent
reference, whether the rows it produces are the session's own agent's, some
child's, or nobody's yet. It reads the alias table directly to tell "an
announcement named this spawn call" from "this id is simply unknown", which
comparing the canonical id against the raw one cannot do when the two match.

Where the identity is not final the row waits rather than guessing. A top-level
spawn whose `task_started` has not landed is the one case that can still
resolve, so its rows are held — bounded at 64, oldest written first — and
released when the announcement arrives or when nothing can name the producer any
more. Nothing is dropped and nothing is written as the parent's.

A release that announces no tasks at all is decided immediately instead of held:
nothing stable is ever reachable for its children, and an id that rotates is
worse than no id because it is silently wrong rather than visibly absent. Those
rows read as the session's own, exactly as they do today. Holding them instead
would strand every row that REVISES an earlier one — a tool result would leave
its tool row reading "running" for the rest of the turn.

The attempt counter moves in exactly one place, the existing reactivation branch
where a new spawn alias reopens an entry, and is gated on that observed alias
change rather than on the counter, so a late duplicate cannot advance a settled
run. The first run carries no attempt at all.

The group row keeps its own module's write path, now named there, because it is
the one row written from a child's frame that is deliberately the parent's.

* test(native-chat): pin producer linkage end to end, and fix the harnesses first

Four harnesses in this area silently discarded the append options they were
handed, so every assertion about attribution would have passed against
`undefined`. Two are fixed here — the journal double behind the real deferred
sink, and the Claude subagent translator's sink — and each records the options
beside its existing call log rather than on it, so the assertions about call
order stay about call order.

The read side is pinned first, because linkage correct in the store and never
read by the projection is the way this ships looking finished and fixing
nothing. The three defects are asserted through the live-turn and projection
readers: a parent no longer reads as thinking because its child is reasoning, no
longer shows its child's running tool, and no longer quotes its child's prose or
prompt. The opposite direction is pinned too — the transcript still renders the
child's output, and a parent's own line is never suppressed.

Also covered: a resumed child keeps one identity while its spawn call id
rotates; a child row arriving before its announcement is held and then written
linked rather than dropped; a held row is written under the raw reference when
no announcement ever comes; the buffer's bound writes the oldest row rather than
losing it; an unresolvable id reads as a child rather than as the parent; a row
with no linkage reads as root; the schema version is unchanged; and a strict
prompt shape still parses, with a positive control proving that strictness is
real and is why linkage rides the row rather than a body.

* test(native-chat): give the journal double's cast its SAFETY rationale

Editing inside the object literal re-attributes the pre-existing assertion to
changed lines, and the changed-code gate requires a line-specific rationale.

* fix(claude): name the agent that spawned a nested subagent

A grandchild's rows carried an agent id but no parent, and under this
journal's semantics an absent parent is not silence — it is the claim that
the session's own agent spawned the row's producer. For a task spawned from
inside another subagent's sidechain that claim was simply false.

The frames from such a task carry exactly one handle: the nested call's tool
id. That id was journaled once already, as a tool-use block on the row of the
child that made the call, so the child is recoverable from it — but only if
something remembers which row carried it. The registry that already tracks
which tool calls reached the top-level transcript now records the sidechain
ones too, against the reference naming their owner, and the resolver follows
that reference to name a row's parent.

The reference is recorded, not an identity, and it is resolved through the
same path the owner's own rows resolve through, so a parent id always matches
the agent id the parent's rows carry however either was settled. A row now
persists only once BOTH its producer and its parent are final; a grandchild
whose child has not been announced yet waits in the same buffer, and leaves
it through the same three doors.

* test(claude): pin the streamed-text lane's attribution instead of only capturing it

The checkpoint harness was fixed to record the append options, and then nothing
asserted them: all five of its tests passed unchanged against an implementation
that resolves no producer at all, so the lane's attribution was covered by a
capture and no claim.

These assert it: a block streamed inside a child carries that child's linkage,
the session's own carries no keys at all, a checkpoint is held rather than
written while the producing agent is provisional, the flush before settlement
writes a held block under the raw reference rather than losing it or filing it
as the parent's, and a block's producer is resolved once and kept — every
checkpoint rewrites the same row, so a producer that moved would file one
agent's prose under two identities.

* refactor(journal): name the linkage row fields for their role, not their shape

The anti-slop gate rejects "Shape" in a symbol name. These are the linkage
fields a render item carries, and the name now matches the sibling helper
that builds them.

* fix(claude): hold the session's first subagent instead of filing it as the parent

A child's first frames can arrive before the `task_started` that names it, and
the resolver treated "this release has announced no task" as a settled fact
about the CLI. Before its own first announcement every session looks exactly
like that, so the FIRST subagent's pre-announcement rows were written straight
out as the session's own — the whole defect, for the first child of every
session, persisted with no backfill to repair it.

A spawn call the session forwarded at top level is positive evidence that an
announcement is still expected, so it now outranks the release check. A release
that genuinely announces nothing is unchanged: its rows reach the release
verdict at settle and still read as root, just written a little later.

Also stops a malformed owner chain that loops back from naming an agent its own
parent; the depth guard bounded that walk but could not make its answer mean
anything, and absence is the truthful claim.

* fix(claude): read a tool result as its caller's row, not a child's

Every top-level tool call is a forwarded tool id, not just a spawn, so a result
frame naming its own call as parent resolved as a child awaiting an
announcement that is never coming. The row was parked until the turn settled
and the tool sat `running` in the meantime.

A frame delivering the result of the very call it names is the caller consuming
its own output; only a spawn call ever gets a sidechain. Pins added for that and
for the first-subagent hold, and the never-announced case now asserts the row
was WRITTEN as root rather than that it carries no agent id, which an absent row
also satisfied.

* fix(journal): refuse an empty producer id, and drop a bad one without losing the row

The reader that scopes a parent's surfaces tests PRESENCE, so `agentId: ''` is
present: a row carrying it reads as a subagent's and disappears from its own
author's surfaces for good. Neither validator caught it — the wire schema
accepted any string, and the persisted-row guard type-checked nothing in the
bundle at all, against that file's own stated policy.

The wire schema now requires a non-empty id. The persisted side sanitises
instead: a bad linkage field is DROPPED and the row is kept. Rejecting there
would turn a tightened validator into a whole-store kill switch, and degrading
a row to the session's own agent is what every row said before linkage existed.

* fix(native-chat): answer the turn activity line for the session's own agent

`selectStructuredAgentTurnActivity` is a "what is this agent doing right now"
reader and was not scoped by producer. It builds a label set from every
tool-call row in the turn — a subagent's included — and both readers below it
use that set to suppress a line that repeats it. So a CHILD's tool label could
blank the PARENT's activity line: child data deciding the parent's surface.

Live, not latent: the provider-activity branch is populated in this lane, and
it consults the label set without ever consulting `providerFrame`, which is
what the status fallback loop relies on. Scoped once at the top, so both
readers share one interpretation point. Renderer and mobile share this
function, so both are covered.

* fix(journal): stop a lifecycle batch stamping one producer onto N mutations

A lifecycle-batch row carries N mutations but stamped linkage at ROW level, so
a future mixed-producer batch would silently attribute every mutation to
whoever opened it. Both callers are single-producer today, so this was latent.

The write path no longer accepts linkage for a batch, which removes the failure
mode by construction rather than guarding it. The reducer still READS linkage
off a batch row — a row may arrive from a host that writes one — and a genuinely
mixed batch would have to stamp per mutation, which nothing needs yet. Chosen
over adding a per-mutation field because that would persist a new key forever
with no writer and no reader.

* docs(journal): say why each unlinked write site is unlinked, and drop two false claims

Completes the write-site audit the PR claims. Prompt rows carry no linkage and
CANNOT: a prompt arrives through the SDK's permission callback, whose options
carry a request id and the tool awaiting approval and no parent reference of
any kind — unattributable at that site, not deliberately root. Turn rows are
deliberately root and now say so.

Two comments justified decisions by mechanisms this store does not have. The
linkage docblock cited compaction dropping a start row and a pagination
boundary; there is no compaction, and pagination is complete-or-reset. Per-row
repetition is still right, for the reason that is actually true: every reader
scans back from the tail and stops at the turn. A test carried the same false
framing. `claudeFrameParentRef` claimed to read the field by the same rule as
`isRootClaudeFrame`; it is deliberately stricter on the empty string.

* refactor(claude): write a child's rows through, then correct the attribution

Four misattribution paths shared one cause: the lane committed to an
attribution verdict at write time and could never revise it. That followed from
"the journal has no backfill", which is false — re-appending an `itemId` bumps
its revision, the reducer rebuilds linkage from the newest row, and it pins
`sequence`/`observedAt` so a correction does not move the bubble. This lane
already relied on that twice.

So the order inverts. A row whose producer is still provisional is written
immediately, stamped with the spawn call's own id, and re-attributed in place
when the announcement names it. Bookkeeping no longer gates a user's view of
what an agent said.

The hold buffer is deleted rather than left as a pass-through. Corrections are
bounded and die four ways: the announcement, turn settle, teardown, or the
bound. Passing the bound gives up on that producer WHOLESALE — correcting some
of a child's rows and not the rest splits one child across two ids, which is
worse than correcting none. A correction that would change nothing is dropped
rather than burning a revision.

The streamed lane loses its producer latch, which pinned the first verdict
permanently and is why an announcement one frame later could never reach the
row. Every checkpoint rewrites the same identity, so there is one row per block
and re-resolving can only revise it; the latch was guarding against a split
that cannot happen on this path. A block that stops streaming before its
announcement is re-attributed explicitly, since nothing else revisits it, and
the announcement is now observed BEFORE the forced flush that would otherwise
stamp it a line too early.

Also narrows the no-announcements-at-all escape so it no longer swallows a
forwarded spawn call. That escape now applies only to a sidechain id no spawn
call ever forwarded, where there is genuinely no handle to stamp.

* fix(journal): move two test doubles onto the signatures they pin

Both failed typecheck while passing at runtime, which is what a test double
gets to do: vitest never typechecks them.

The sink's lifecycle-batch double still read producer linkage off the batch
input after that input stopped carrying any, so it had no property in common
with the linkage type. The fence is now all it records, which is what the
narrowed contract actually forwards — and what the test beside it already
asserts.

The row-schema helper returned the whole six-arm `JournalRow` union while every
caller reads `body`. It now narrows to the item arm it always builds, so the
assertions read it directly rather than through a cast.

* fix(claude): resolve a tool result to its real caller, not to the session root

A nested tool row could end stuck `running` with its result content dropped.

Cause was in the result-frame attribution, not in the correction ledger. A frame
delivering the result of the call it names as parent is the CALLER consuming its
own output — but the code read "the caller" as "the session's own agent", which
is only true when the caller is the root. A call a child made is owned by that
child. Collapsing it to root both misattributed the row and made the result's
write resolve through a different reference than the call's, so the correction
owed to that row was left holding the body it had BEFORE the result landed, and
re-attribution then reverted the row.

The caller is now resolved through the registry that already records which agent
journaled a tool call, so both writes to one row resolve through the same
reference and the newest body wins.

A settled write also supersedes any correction owed to its row. One `itemId` is
legitimately written under two references — `claudeToolIdentity` is keyed on the
tool id alone — and a settled write already carries a final verdict, so an
outstanding correction could only restamp it from a reference that write did not
use. Dropped rather than re-bodied for that reason.

Adds the ledger's first unit tests, including the invariant this defect broke: a
correction changes a row's attribution and never its content.

* fix(claude): keep a correction owed when the sink refuses it

A correction went out through the plain append, which discards the queue's
admission. Under backpressure the write was refused and `retry` had already
dropped the entry, so the obligation died with nothing re-deriving it — the
failure class this work exists to refuse. It is self-feeding too: a correction
costs a commit on the same serialized writer that carries live rows, so the
burst that generates many corrections is what builds the backlog that drops
them.

It now uses the admission-returning path the sink already exposes, keeps the
entry outstanding on a refusal, and lets `abandon` try once more. A refusal
there ends it: the row keeps the spawn call's own id, which is usable, and an
obligation with no exit is worse than one that settles for less. `settle` also
reports what actually happened instead of always claiming it wrote, so publish
no longer fires for a write nobody accepted.

Also records why the live-turn scans may read the turn record before checking
the producer. A turn is the session's unit of work and no producer of a
turn-bearing body stamps linkage: Claude's turn rows carry none, Codex has no
linkage concept, the compact row passes only a fence, and the stale-turn sweep
goes through the lifecycle-batch path, which cannot carry linkage by type. The
ordering is safe by construction rather than by accident, and the comment says
so, so a future producer knows what it would break.

* fix(claude): never read a row naming a parent as the session's own

A non-null `parent_tool_use_id` names a child, always. The resolver still had
one branch that read such rows as the session's own agent's — a release that
had announced no task, where the comment claimed "there is no handle to stamp".
There is one: the reference itself. The branch was buying a false attribution to
avoid an id nothing joins on, which is the trade already reversed once for
forwarded spawn calls, and every reader of this field is a presence test.

So the branch goes, and with it the `root` arm of the verdict and the resolver's
whole dependency on whether the release announces tasks. Two states remain:
linked now, or linked now and owed a correction. The type deleted a stale test
double on sight, which is the argument for removing the arm rather than the
branch alone.

This also closes the severe half of the tool-origin eviction exposure. A spawn
id evicted from the bounded top-level set used to flip the release check on and
stamp a child's rows as the parent's; with nothing returning root that cannot
happen. What remains is a missed correction, which splits one child across two
ids — the same end state as passing the correction bound, benign in kind and
disclosed.

`isForwardedParentTool` stays where it gates PENDINGNESS. It now decides only
whether a correction is owed, never whether a row is a child's, so a stale
answer costs precision rather than correctness.
2026-09-22 13:57:46 -07:00
Brennan Benson 4c696a1e2a fix(agent-status): a structured session with live child work reads as working (#22295)
* fix(agent-status): a structured session with live child work reads as working

An idle native-chat session whose subagent was still running showed a green
check in the sidebar, the collapsed worktree pill, and worktree ps, while a
terminal Claude session in the same situation showed working. The two lanes
folded child work into the parent's status with different code: the hook
listener did, the structured lane did not.

Both lanes now share one child-work liveness vocabulary and one lead-status
fold. Live agent work makes a settled lead working; shells and monitors alone
make it monitoring. The structured lane derives liveness from the background
task list already on the wire, in both its readers, so the sidebar, the CLI,
the dashboard and mobile agree. The Claude task-kind table is one shared file
covering both the hook inventory and SDK stream names, and the renderer bridge
reuses the shared child-work projection instead of carrying its own copy.

* fix(agent-status): a blocked or out-of-contact subagent still holds its session working

Child-work liveness retired an agent-kind child on any state but working/monitoring,
while the shell beside it stayed live on everything except done/idle. A subagent
waiting on a permission prompt, or one whose host lost contact, therefore counted
for less than a backgrounded sleep and let the session read done. Both kinds now
share the settlement rule `resolveAgentChildWorkFreshness` already reads rows by:
only an explicit done/idle retires child work.

Also keep empty task labels out of the shared background-task projection candidate,
so a host that publishes `name: ''` cannot beat the child-row fallbacks.

* test(agent-status): pin the widened hook-inventory agent names, and correct two stale claims

The hook inventory now classifies through the shared kind table, which also maps the
SDK stream's `local_agent` / `local_subagent`. Nothing pinned that widening, so add
cases for all four agent names — including `teammate`, whose pane state stays `done`
under the #8825 idle-squat rule.

Two comments the fold made false:
- the teardown marker rule's comment claimed it could not disagree with what the UI
  calls working; it is deliberately lead-only, so now it says that and why;
- the agent-status store reference still described the structured row's `state` as the
  deleted `structuredAgentSessionStatusState`, and omitted the `workingMode` the ingest
  now writes.

* fix(agent-status): the state clock restarts when monitoring becomes a real turn

`stateStartedAt` carried forward whenever the prior `state` matched, which was sound
while `state` meant "a turn is running". Now that it folds in child work, an idle lead
watching a `sleep 3600` publishes `working`/`monitoring`; the user's prompt 45 minutes
later keeps `state: 'working'`, so the row inherited the watch loop's clock and read
"Working for 45m" the instant the turn began. Monitoring is its own displayed label
(`worktree-card-compact-agent-row.tsx:40`), so the continuity key is now the whole
published work identity — state AND workingMode — in both writers.

Also record two facts the code stated wrongly: the structured lane's `interrupted: false`
is inert (a projected session status has no interrupted member) rather than a decision,
and the child-work liveness rule's escape hatch is the roster's session lifetime, not a
settled state.

* fix(agent-status): a workflow is watch work, and child work dates itself

Two defects the fold introduced.

`isAgentChildWorkKind` counted `workflow` as agent work, so a structured session
whose only live task was a backgrounded `local_workflow` published a full working
spinner while the children projection — which admits `kind === 'agent'` only —
rendered nothing to expand, and the same workflow in a terminal pane showed the
monitoring badge instead. The repo already decides this: `isClaudeSubagentTask`
excludes workflows by name, and MATERIALIZED_TASK_KINDS leaves "the backgrounded
shell command and the workflow" to the non-agent owner. The predicate is now
`kind === 'agent'`, and the three sites that restated the same test route through
it, so a new kind is decided in one place instead of three that merely agree.

`evidenceObservedAt` dated every row by `summary.updatedAt`, the journal's last
activity. The journal cannot date child work: its clock stopped when the lead's
turn did, so a genuinely live roster aged past the 30-minute staleness window and
mobile's dot decayed a running session to idle. The fold now reports whether child
work alone holds the row open, and only then does the host's observation clock
stand in — keeping "a restart's republish is not new evidence" for lead turns.

* fix(agent-status): the sidebar dates child work the same way the host does

`fromChildWork` reached the host ingest but not the renderer bridge, so after ~30
minutes of live child work with no journal activity the sidebar's row aged into
staleness while `worktree ps` and mobile stayed fresh — two writers for one session
answering differently, which is the defect this PR exists to remove. For a remote
host the client's own receipt time is also the more honest clock, since the journal
stamp is the host's and is never comparable against this machine's now.
2026-09-22 13:48:07 -07:00
Brennan Benson 60c43695e5 feat(agent-launch): report the pane a terminal launch created (#22108)
* feat(agent-launch): report the pane a terminal launch created

A `term_*` handle is a main-side mapping the renderer cannot resolve
(terminal-handle-links.ts:309), so a client that draws its own tabs had no
way to name the tab it had just asked `agent.launch` to build. The runtime
already mints that pane, bakes it into the PTY's environment and hands it
to its own reveal; the surface factory then dropped it on the floor.

Carry it through as `paneKey` on the terminal outcome. Identity, not
placement: where the pane goes — which group, what order, whether it takes
focus — stays with whichever client is drawing, and nothing here rides the
wire for it. One field rather than a tabId/leafId pair, because the key
already holds both and two copies of one fact can disagree.

Absent when this launch minted no pane: a reused terminal was already
running, and a worktree-create startup terminal is built by the create,
which reports only a handle. Naming the wrong surface is worse than naming
none.

Optional on the wire and optional on the read side. Mobile parses the
receipt with a loose object and is deliberately mode-blind, so it ignores
the field; the persisted-row guard checks it when present and accepts a row
written before it existed, because a read rule stricter than the write side
turns one odd row into a refused replay.

* fix(agent-launch): retain startup terminal pane identity
2026-09-22 09:35:11 -07:00
OrcaWinandm4air ba742a86bb fix(linux): release orphaned processes when their owner exits (#22247)
* fix(linux): release orphaned processes when their owner exits

* fix(linux): handle inhibitor errors until streams close

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-22 05:10:03 -07:00
90d363afc9 fix(renderer): contain Monaco initialization failures (#21555)
* fix(renderer): add defensive error handling for Monaco editor crashes

Analyzed 34 crash reports for v1.4.205 released 2026-09-17. Identified and
added defensive fixes for React error boundary crashes in Monaco editor setup.

- Error: ReferenceError: thũs is not defined
- Location: Monaco editor initialization (editor.api2 bundle)
- Platforms: Linux, Windows, macOS
- Root cause: Undefined variable in Monaco setup or language registration
- Status: Added try-catch to prevent cascade crash

- Pattern: Cascading process deaths (network service + GPU service)
- Platforms: Primarily Windows
- Root cause: Infrastructure/concurrent process failure (not code defect)
- Status: Documented, requires Electron/Chrome infrastructure review

- Pattern: Renderer memory grows to 851MB on low-RAM Windows systems
- Root cause: Memory exhaustion on systems with <2GB free RAM
- Status: Existing memory monitoring detected; needs leak investigation

- Status: Requires minidump analysis with source maps

1. Added try-catch to Monaco editor mount callback (use-monaco-editor-mount.ts)
   - Catches errors during editor initialization
   - Logs file path and error for better diagnostics
   - Prevents crash cascade to React error boundary

2. Added try-catch to Monaco language registration (monaco-setup.ts)
   - Catches errors during Vue/Svelte/Astro/Nim language registration
   - Logs failures without crashing Monaco setup
   - Allows app to continue even if optional features fail

- Analyzed 34 crash reports across 3 categories
- Examined crash dumps, diagnostics, and memory profiles
- Reviewed Monaco setup and editor component code
- Checked git history for recent changes

- Crash breadcrumbs (memory, process state, user actions)
- Process metrics (heap, private memory, system memory)
- Component stacks (React error boundaries)
- Exit codes and system signals

- Error silently continues instead of crashing: Users get degraded experience
  instead of app crash, can still use editor in most cases
- May hide underlying issues: Errors are logged for crash reports, but won't
  be surfaced as prominently

- Type checking: pnpm tc:renderer (passed)
- Changes preserve existing error reporting through crash breadcrumbs
- Defensive coding only adds try-catch, no behavior change for success path

* fix(renderer): keep Monaco mount failures inside error boundary

* fix(renderer): isolate Monaco setup failures

* fix(renderer): contain Monaco mount failures at the editor surface

The try/catch around the onMount body did the opposite of containment: React
already routed that throw to the page boundary, so swallowing it left a
half-wired editor and hid the crash from the reporting pipeline. It also never
saw the reported failure, which is raised inside @monaco-editor/react's own
create effect before onMount runs.

Revert the hook to main and wrap the editor element in
RecoverableRenderErrorBoundary instead, so either throw degrades the file pane
only, still files a crash report, and retries by remounting on the existing
pane+path key.

* refactor(renderer): drive Monaco setup steps from one guarded table

Ten near-identical guarded calls, each repeating its own function name as a
label, become one [label, step] table run by a single loop. Same behaviour: an
optional registration that throws is logged and the rest still run.

loader.config and the editor model registry stay unguarded — they are
load-bearing, so catching there would only move the failure later.

* fix(renderer): breadcrumb swallowed Monaco setup-step failures

A guarded registration that throws was console-only, so a lost language or
behaviour guard never reached crash reports. Record a breadcrumb so the
containment stays visible in the field.

Claude-Session: ab8ff806-4870-4ea8-bbf5-bbd123b1166e

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
2026-09-22 02:57:18 -04:00
Jinwoo Hong e47ef8cc28 feat(mobile): the shell tells a page which optional capabilities it has (OTA phase D, C8.1) (#22141)
* fix(mobile): publish page-route pairs the strict host schema accepts (OTA phase D, C8.1)

`routeViewOf` handed the manifest's own route entries to the host as
`pageRouteGrants`. The phone reads a manifest route loosely, so an entry
arrives carrying whatever field the desktop that wrote it knew about, and
`BridgePageRouteGrantsSchema` is `.strict()`: one unread key refuses the
pairs, `createBridgeHost` refuses the route with them, and the page gets no
`init` at all rather than losing one field.

Fixed before any route carries an optional grant (ruling 37.4), so the
manifest field the next commits add costs an installed shell nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore: drop the closure and bundle probe scripts from the tree

Scratch measurements for C8.1 (which route closures reach the HTML preview,
and what the preview render rig costs to bundle with a client provider). They
belong outside the repository and were swept in by the previous commit's
`git add -A`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): a manifest route may declare optional grants (OTA phase D, C8.1)

Design B of design-ota-c8-1.md, ruling 37. `MobileWebBundleRouteSchema` grows
`optionalGrants` under the required lane's own grammar, with the 16-name
ceiling applied over the union of the two lists rather than to each. Serving a
route still reads `grants` alone, so a capability a screen cannot work without
stays required and takes the route native; a session's granted list is
`[...grants, ...optionalGrants]` narrowed to what this shell implements, from
one helper that both `grantsForRoute` and the `pageRouteGrants` publish read.

The ruling's compatibility rationale is corrected in place. `z.looseObject`
passes unknown members through rather than dropping them (measured, zod
4.4.3), so a shell older than the field still receives the key; what it lacks
is a policy that reads one. What makes the lane safe against such a shell is
therefore the previous commit's publish fix, not the reader.

BRIDGE_PROTOCOL_VERSION stays 1. No new notify, verb or frame field.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): name the shell's cancelled-navigation behaviour as a grant (OTA phase D, C8.1)

`externalNavigation` joins `MOBILE_WEB_SHELL_GRANTS` beside `screencastBinary`
and `haptics`, declared in `cancelled-navigation-target.ts` because that is
the module holding the rule which acts on it. A third token that is neither a
verb nor a notify: the page posts nothing to make a cancelled top-frame
navigation happen, so this list is the only thing that can tell a page whether
a tap inside the sealed HTML-preview frame escapes at all.

A constant and not a platform read (ruling 37.1): both engines dispatch the
event, `ios/MobileWebShellView.swift:481` and Android's
`MobileWebShellView.kt:382`, so an app build carries the behaviour on both or
on neither.

The policy census grows the half that was only pinned by the verb table: the
implemented set is that table plus exactly three non-verb tokens, each read
off the module that declares it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): the bundle builder carries a route's optional grants (OTA phase D, C8.1)

`resolveMobileWebPageRoutes` maps each declaration member by member, so a
field the declaration grows reaches a phone only once the map names it: until
now `optionalGrants` would have been dropped in silence and every route would
have declared nothing optional. Omitted when the route declares none, because
absent and empty are the same answer to a shell.

The declaration suite grows the rule rather than a row: the map carries the
lane through and writes no key without one, and the lane is held to the
manifest's own grammar and to the ceiling over the union of the two lists.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): the HTML preview hides its links on a shell that cannot open one (OTA phase D, C8.1)

The session route declares `externalNavigation` on the optional lane, and the
preview asks for it before it renders an artifact's links as links. Ruling
37.2's three readings are what "hide" means here, and removing `href` is what
delivers all three at once: `a:any-link` stops matching, so the UA stylesheet
stops underlining, the element leaves the tab order, and there is no dead
anchor a tap does nothing on. The text the author wrote stays where it was,
the artifact paints, and the Preview/Source toggle is untouched.

Done with the browser's own parser rather than over the source text: an `href`
inside a comment or a `<template>` is text to a browser, and a pass that
rewrote either would be editing the artifact instead of its links. The frame
also loses `allow-top-navigation-by-user-activation` on that path, so a link
the pass somehow missed is refused by the browsing context as well.

One route, measured rather than assumed: the design said two, and the file
preview route's closure does not reach the HTML preview at all - it renders
`MobileFilePreviewScreen`. The new closure census derives that list from the
hook's callers.

The render rig grows the case on both engines and the readings it needs, and
`mobile-web-app-preview-frame-readings.mjs` is split out of it at the
readings/arms boundary, because the two were over the 600-line cap together.
Two engine findings are recorded in the rig: an `<a>` with no `href` still
answers `tabIndex` 0 on both, so focusability is asked by focusing; and WebKit
computes `cursor: auto` for a real link, so that reading is pinned where it
discriminates and its blindness pinned where it does not.

The hop-coverage census now reads the effective set, because that is what the
running rule compares. Inert today: the session route is the only declarer and
an opener into every other route.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the preview's hidden link path where the unit suite can reach it (OTA phase D, C8.1)

The mobile suite runs in a `node` environment whose resolver has no `.web`
precedence, so `MobileHtmlPreview.web.tsx`'s import of the grant hook lands on
the native sibling, which answers yes unconditionally. That is why the
existing component suite still measured the granted frame without knowing a
grant exists, and it means the hidden path had no coverage in the sharded
`test` job, where the render rig is skipped for want of the bundler's
dependencies.

So the wiring gets its own file with the module replaced: that the component
asks, and that both the frame's sandbox and the document it is handed follow
the one answer. happy-dom rather than the suite default, because the inerting
pass parses with the browser's own `DOMParser`.

`String(node.type)` rather than a literal comparison: `node.type` is
`ElementType`, which overlaps a real intrinsic tag and not the host strings
these mocks render, so `=== 'Pressable'` is a no-overlap error under
`tsconfig.test.json` and the tests-typecheck ratchet reds on it.

Also replaces a `Reflect.get` the anti-slop gate refuses with an `in` check.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the session page closure at 4,362 for C8.1's three modules

Measured on both sides with `mobileWebAppRouteClosure(SESSION_ROUTE)` at base
`841d06a969` with all five postinstall generators run first, and the two
`local` lists diffed rather than the total inferred: 4,359 -> 4,362 modules,
1,017 -> 1,020 local.

All three are local source modules and none is vendored: the page's read of
`init.grants.native`, the pass that turns an artifact's links back into text
without the grant, and the module declaring the token beside the rule that
acts on it - reached both by that hook and by `page-route-policy.ts`. The
`bridge-caps.ts` it imports was already in this closure, and the hook's native
sibling is replaced rather than joined.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): allowlist the preview's grant sibling among the .web.* overrides

`mobile-web-app-web-overrides.test.mjs` pins the allowlist against the `.web.*`
files on disk, so a new web sibling reds it until the file says why the page
needs one. Red before: `expected [ …(36) ] to deeply equal [ …(37) ]`, naming
`src/components/use-html-preview-link-grant.web.ts`.

The preview's own entry is corrected with it: its reason said
`allow-top-navigation-by-user-activation` is granted, and that token is now
conditional on the shell answering that it can open such a navigation.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the hidden-link render case waits on the frame's own reading (round 1)

CI's chromium arm timed out at the full 240 s on this case alone while the
WebKit sibling passed in 1.5 s and it passed 26/26 locally. The cause is the
third arm: it tapped the granted link and waited through
`expectNavigation: 'main-frame'`, and `waitForRecordedNavigation` has no bound
but the case's own timeout. Under CI load the click missed its 2 s
actionability window, no navigation was ever recorded, and the arm sat in that
wait until vitest gave up - `recorded []`, with the frame attached only at
38.9 s. Three arms sharing one budget is what made this the case to find it.

The arm is dropped rather than its wait lengthened or retried. Every verdict
left is a reading the frame itself publishes: the anchors its document holds,
the style the engine computed for one, whether focus lands on it, and now
whether the tap this arm made landed at all - `actError` is asserted null, so
a click that never reached its target is no longer the same three zeros as a
tap that did nothing.

Nothing is lost. The tap's outcome on a granted shell is the next case, on
these same counters from this same rig and with a budget of its own, which is
the presence precondition this file already uses elsewhere for the same
reason. The WebKit sibling's discriminating reads are untouched.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): the inert-link pass changes nothing an engine renders but the links (round 2)

Round 2's ruling: the hidden-link path may change nothing about the artifact's
rendering except that links are not links. A parse and a reserialise is not
free of that by default, and all four findings reproduced on Chromium 147 and
WebKit 26.4.

A same-document fragment link is kept. It starts no navigation at all, so it
goes on working inside the sealed frame whatever the shell can do, and taking
it away would be degradation over a capability it never needed - an artifact's
own table of contents is the case. Its `target` still goes, because a fragment
aimed at another frame is a navigation rather than a scroll, and `href=""` is
not a fragment: it resolves to the frame's own URL.

Links inside `template.content` are reached, recursively. `<template
shadowrootmode>` is a declarative shadow root the frame's parser attaches and
renders, and `querySelectorAll` does not walk into template content, so those
links arrived live inside a sandbox that refuses their navigation - the dead
anchor ruling 37.2 forbids. Measured: `parseFromString` attaches no such root
on either engine or in happy-dom, so the pass can reach them.

The leading newline of a `pre`, `listing` or `textarea` is written back. A
parser drops one after the start tag and the serialiser is specified to put it
back; measured, neither engine's does, so a round trip lost a blank line from
every such block.

The doctype is carried whole, and the reason is corrected from the one the
finding gave. It cannot move this frame between layout modes: a `srcdoc`
document takes its mode from its embedder, and measured, a quirks doctype, the
bare name and no doctype at all all read `CSS1Compat` inside the frame. What
rewriting it does is change the document the author wrote for no reason, with
`document.doctype` observable beside a Source tab showing the original. The
render case pins `compatMode` as the blind reading it is and reads the frame's
own doctype identifiers as the one that discriminates.

Option B was not available: the frame has no `allow-scripts` and inherits
`script-src 'self'`, so nothing runs inside it and there is no injection to
carry the work.

Also drops a vacuous half of the affordance test. `renderSource()` is called
with no argument, so the markup a Source view shows is the caller's own
closure and asserting it equals the fixture passed whatever the component did.
What the component decides is whether the rewritten frame stays mounted
underneath, and that is what is read now.

`mobile-web-app-preview-arm-driver.mjs` is split out of the render rig at the
boundary the readings module already names - the rig holds what each case
claims, the driver how an arm is driven, the readings what it reports - since
the three were over the 600-line cap together. No max-lines disable or bump.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): a fragment link is a frame navigation in this preview, so it is inerted too (round 3)

pullfrog is right, reproduced on both engines before believing it. Round 2
kept `#`-prefixed hrefs on the theory that they are same-document scrolls. In
this frame they are not: the document's URL is `about:srcdoc` while its base
URL is inherited from the embedder, so `#section` resolves against the shell's
own URL and the destination differs from the document's by more than a
fragment - which makes activating it a frame navigation, and the shipped
`frame-src 'none'` refuses it.

Measured under the shipped policy, one tap, with something to scroll:

  Chromium 147   scrollY 0, frame becomes chrome-error://chromewebdata/,
                 artifact gone, embedder reports frame-src <origin>/preview
  WebKit 26.4    scrollY 0, frame stays about:srcdoc and intact, same report

So the destruction is Chromium-only but the absence of a scroll is not: there
was no working affordance to carve out for, and the carve-out left a live link
that destroys the preview - worse than the inert text it was meant to avoid.
Both sandbox values behave the same, so this is the base URL and the policy
rather than the sandbox.

The same tap does the same thing on the granted path, where this pass does not
run, so an artifact's internal links have never worked in the preview. That is
not this change's to fix; it is recorded in
`followup-html-preview-fragment-links.md`, and the render case reads the
granted arm's violation as its presence precondition so the behaviour is
pinned rather than merely known.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 02:11:40 -04:00
Jinwoo Hong 7650abe224 fix(macos): tell the user when Orca's terminal service can't read their folder, and walk them through the fix (#21923)
* fix(macos): tell the user when Orca's terminal service can't read their folder

On macOS, a terminal daemon that survived an app update can be refused access to
a workspace under Documents, Desktop, or Downloads while the Orca app itself can
still read it. Terminals opened there die with "Operation not permitted" and
nothing on screen explains why. The daemon has reported `cwdReadableByDaemon` on
every create since #18043 and main has emitted `daemon_pty_cwd_denied` on proven
divergence since then; the field data says 1,438 users hit it in 21 days. What
was missing was the notice.

The verdict itself moves off `access()`. A grant-less probe on an affected
machine showed a TCC mode where `access(R_OK|X_OK)` passes on `~/Documents` and
`opendir` still fails, so the check now does what a shell listing its cwd does:
`opendirSync`, one `readSync`, `closeSync`. Only EPERM/EACCES reads as denial —
a missing path, a non-directory, or an unexpected error still reads as readable,
so a non-permission failure can never masquerade as one. The same probe is what
the app side compares with, through one oracle shared by the telemetry emitter
and the notice, so the spawn path reads the directory once.

Proven divergence now also records evidence in main: one entry, keyed by the
daemon's pid, start time and launch nonce, carrying an opaque digest of that
identity and the folder class. No path leaves main. The existing focus-time
`macTccAttribution` poll carries it to the renderer, which raises a second toast
latched per daemon scope: dismissed stays dismissed, and a restart mints a new
identity so the poll returns null and the toast clears with no post-restart
probe. If the replacement daemon is denied too, about 31% of cases, the next
spawn re-records under the new scope and the notice returns, now with the
re-allow sentence doing the work.

No new IPC channel, no daemon protocol field, no polling change, and nothing new
on the spawn path beyond one `opendir`. `daemon_folder_access_notice` counts
shown, dismissed and open_manage_sessions against `daemon_pty_cwd_denied` as the
denominator; `shown` is emitted from main the first time a scope leaves the IPC
handler, so the renderer carries no telemetry plumbing for it.

* fix(macos): clear folder-access evidence only when the same folder class reads back

A readable spawn in ~/code said nothing about a Documents denial but was
hiding the notice; retire the evidence only when the daemon reads a folder
of the class it was denied on.

* fix(macos): say what a terminal-service restart actually does

The Manage Sessions restart confirmation still described the product as it was
before agents resumed themselves: it promised panes showing "Process exited"
that the user reopens by hand, and mentioned legacy-protocol sessions nobody
outside the daemon code can act on. Open terminals and agents come back on
their own now, so the old copy made a routine remedy sound like data loss.

It also called the thing a "daemon". The same restart is about to be offered
from a user-facing fix dialog, so both surfaces now say "terminal service", and
the confirm button is just "Restart".

The new body adds the one fact the old one never stated: terminals on remote
hosts are not affected. Translations of the two changed strings are dropped so
the five non-English locales fall back to English rather than keep showing copy
that is now wrong.

* feat(macos): give the denied-folder notice a fix the user can follow

The folder-access toast told the user their terminal service could not read
Documents and then handed them a paragraph: restart from Manage Sessions, and
if that does not work, re-allow Orca in System Settings. Both halves were
guesses. Roughly a third of restarts do not fix it, and the user had no way to
know which case they were in before spending every open terminal on finding
out.

Main can now answer that. `daemon-folder-access-probe.ts` forks a short-lived
child of the app binary the same way the daemon itself is forked, runs one
opendir/readdir/closedir against the denied path, and prints a single JSON
line. macOS attributes a TCC grant to the process that forked the child, so a
child of the app running now answers exactly the question the running daemon
cannot: would a replacement daemon get in? The child goes through the shared
child-process wrapper, never a shell, with a 3s deadline, a 1KB output cap and
an environment scrubbed to PATH/HOME/TMPDIR. Every failure — timeout, bad
output, spawn error — reads as `unknown`, never as a verdict.

That answer rides out as `restartWillHelp` on the evidence the existing
focus-time poll already carries, and the toast becomes a title and two buttons:
Fix… and Not now. Fix opens a dialog with the two real steps. When the grant is
already in place, step one is shown as done and Restart is live. When it is
not, step one is open and Restart is disabled until it completes — which it
does by itself, because the poll re-probes while the answer is still no, and
returning from System Settings is the moment that lands. An unanswered probe
never accuses the user of a missing grant; it leaves both steps open.

Restart calls the management API directly rather than stacking the Manage
Sessions confirmation on top, since the dialog already states the consequence.
Success replaces the steps with a done line and takes the toast down; failure
says so inline and leaves the button usable.

System Settings opens through the existing developer-permissions pane opener,
which takes an id rather than a URL, with Files and Folders added to it. The
event's action enum now also counts fix_opened, settings_opened,
restart_clicked and — emitted from main when a replacement daemon's first spawn
lands in the folder class the previous one was denied on — whether the restart
actually worked.

* fix(macos): let the folder-access notice return after a poll that read no daemon

A daemon identity reads as null during any reconnect blip, and the poll reports that as
"no mismatch". The notice dismissed itself and then never showed again for that daemon,
because the once-per-daemon latch still held its scope. Only "Not now" should latch.

* fix(macos): say what the folder-access notice costs the user

One line read like a stray warning. The toast now says who is blocked and what fails,
and still leaves the steps to the fix dialog.

* fix(macos): give the folder-access toast one action and the X, like every other toast

"Fix" is the only button; the X dismisses. Sonner fires onDismiss for programmatic
dismissals too, so the post-restart takedown now goes through the store and the hook,
and only a user's X is counted as dismissed.

* fix(macos): keep the fix dialog's steps a checklist and put the one action in the footer

Buttons inside each step made the list look like a form, and a footer Close duplicated
the X. The footer now carries the active step's action, with a ghost Cancel; a probe
that could not answer says so under step 1 instead of showing a check.

* fix(macos): let the checklist show the fix landed instead of saying so

A hedged sentence addressed to the user read like chat. On success both steps check
off and the footer offers Done; the unanswered-probe helper is a status, not advice.

* chore(i18n): drop the fix dialog's unused close key

* Revert "chore(i18n): drop the fix dialog's unused close key"

This reverts commit 365915df48.

* chore(i18n): drop the fix dialog's unused close key

* fix(macos): tell step 1 what to do when the folder toggle is already on

Users who need step 1 usually find Orca already allowed in System Settings; the grant
is recorded but not honoured for the daemon. Re-toggling re-records it.

* fix(macos): drop the unverified toggle instruction from step 1

Nothing has been confirmed to fix a grant that is already on, so the step says only
what the probe knows.

* refactor(macos): share the tccutil reset and bundle-id read behind one module

Clearing a macOS TCC row is about to have a second caller: the daemon
folder-access fix (STA-7948) needs the exact `tccutil reset` the computer-use
helper already issues. Extract both it and the PlistBuddy bundle-id read into
src/main/macos-tcc-reset.ts so the two remedies cannot drift apart.

The extracted calls go through runProcessSync rather than a fresh
node:child_process import: the spawn chokepoint's ratchet holds the direct
importer count at a pin, and a new module with its own spawnSync would raise it.
Behaviour is unchanged except that both calls now carry a 10s bound, and the
computer-use test asserts the same argv against the chokepoint's options.

* feat(macos): offer a permission reset when restarting the terminal service cannot help

About a third of the users who see the folder-access notice are still denied by
a freshly forked daemon even though Orca itself is allowed under Files and
Folders, so the restart the dialog offers cannot fix anything for them. That
state previously had one action: open System Settings, where the toggle they
would look for is already on.

The denied state now offers "Reset permission". Main clears Orca's TCC row for
that folder class with tccutil, then reads the folder from the app itself so
macOS raises its prompt against Orca rather than the daemon, then forces a
fresh-daemon re-probe that bypasses the poll's reuse interval. The dialog
re-renders from that verdict: allowed turns step one green and offers Restart,
still denied says so, and a refused reset points back at System Settings.

Nobody has confirmed this remedy on an affected machine, which is why main emits
the re-probe's verdict as reset_outcome_allowed/still_denied/unknown. Those
three, plus reset_clicked, are the evidence that decides whether the feature
stays.

* fix(macos): say what the permission reset does, and keep System Settings as the fallback

Step 1 was labelled like a Settings task while the button did something else, with two routes
in the footer for one step. The denied state now names the step for what the reset does,
explains it under the step, and shows System Settings only after a reset fails or leaves
things blocked.

* fix(macos): count a folder-access restart only against evidence that survived

The stored denial is the prior denial, so a second copy of it outlived the
one event that retires it: a daemon that read its own folder back cleared the
entry but left the copy, and the next daemon's first denial was then reported
as a restart that had never happened.

Track the outcome on the entry itself, drop the spawn-path probe (ten denied
terminals forked ten probe children the focus-time poll re-runs anyway), and
stop emitting `shown` from a getter the reset path calls for data. The
renderer's toast latch is what decides a scope is shown, so it emits it.

Both accessors now read one identity-matched entry.

* refactor(macos): name the folder-access verdict instead of encoding it as a tri-state

`restartWillHelp: boolean | null` re-encoded a verdict the probe already
returns as a named union, so every reader had to remember that `false` meant
"Orca itself must be re-allowed" and `null` meant "no answer".

`freshDaemonAccess: 'allowed' | 'denied' | 'unknown'` says it, end to end
through main, the IPC payload, the preload mirror and the dialog. The reset's
outcome event becomes a lookup. No user-visible string changes.

* refactor(macos): give the folder-access notice one latch instead of three

Two refs in the hook and a field in the store tracked the same fact, and the
dialog reached the hook through a store field plus an effect just to take its
own toast down before sonner echoed the dismissal back.

The store now holds the visible scope and the scopes the user closed, and
exposes the three things that happen to a notice: it is shown, someone else
retires it, or the user dismisses it. The dialog calls retire directly and the
effect is gone. `settingsIsFallback` loses an argument that was always true at
its only call site, so it becomes the local it always was.

* refactor(macos): stop blocking main on the tccutil reset

Two spawnSync calls with a ten-second timeout sat inside an async IPC handler,
so clearing a TCC row held main's event loop for as long as either binary took.

Both now run through runProcess. The computer-use caller that shared them was
already async, so it awaits them.

* test(macos): run the folder-access probe script against real paths

Every other test mocks the spawn away, so the minified child script — the one
piece that duplicates enumerateDirectoryOnce's errno mapping — had no oracle.
It now runs against a temp directory, an absent path, a file, and a directory
whose mode withholds it, which is skipped for root and on Windows.

* refactor(macos): read the folder-access entry through one identity match

All four callers that ask "is this evidence still this daemon's?" now go
through the same private accessor, so the rule the canonical path depends on
lives in one place.

* fix(macos): keep folder evidence through a failed health read, and make a forced re-probe always probe

A rejected attribution-health read nulled the folder evidence on the same poll, which the
renderer read as "cleared". A forced refresh after a reset returned early on an older
settled verdict. The dialog also closes when a reset finds the evidence gone, and stops
showing the unverified helper once the restart is done.

* fix(macos): name the folder in the access-notice scope

One daemon denied two protected folders kept one scope, so the toast, the
fix dialog, and the tccutil reset could each be about a different folder.

* refactor(macos): derive the folder-access dialog from the latest verdict

The store held an `open` flag and a mismatch frozen at the moment the toast
was raised, so the dialog could open on a stale verdict and its remedy state
could survive a close. It now keeps the latest verdict and the scope the user
opened, and the dialog is shown only while the two agree.

* fix(macos): offer the permission reset only where there is a row to reset

A workspace symlinked out of Documents or on an external volume can be denied
too, and the dialog offered a reset that main refuses. One shared list of the
TCC-backed folder classes now decides both.

* fix(macos): give the permission prompt's read a deadline

An unanswered macOS sheet blocks the app's folder read for as long as the user
ignores it, and the fix dialog is modal and busy until that read returns. The
wait now ends after a minute and reports an unknown outcome rather than
probing under the sheet.

* fix(macos): count the folder-access notice once per scope

A reconnect blip reports no daemon, which takes the toast down and lets the
same scope raise it again. Both raises counted as separate notices, inflating
the denominator behind the affected-user rate. The two latches are now one
map from scope to phase, and the count follows first insertion.

* fix(macos): drop the restart warning once the restart is done

Step two ticked green while its helper still warned that open terminals and
agents would restart, which had already happened.

* fix(macos): keep the folder-access toast up when the fix dialog opens

Sonner deletes a toast after its action button runs unless the handler
prevents the event, and it does so without calling onDismiss. Clicking Fix
therefore took the notice off screen while the scope stayed latched as
visible, so cancelling the dialog left no way back to it.

* refactor(preload): reuse the shared daemon cwd class instead of copying it

The five folder classes were hand-mirrored in preload behind a comment saying
preload cannot depend on main-only modules. The enum lives in src/shared,
which preload already imports from elsewhere, so the copy could drift.

* refactor(macos): close the fix dialog when its evidence disappears

A null verdict left the opened scope set, so the same scope coming back
remounted a checklist nobody had opened. Clearing it on a null verdict also
makes the dialog's scope key redundant, so it goes.

* fix(macos): let each fix-dialog button report its own work

The footer swaps the reset for a restart as soon as a poll says the grant
landed, which can happen while the reset is still running. Both buttons read
their label off the dialog being busy at all, so the restart button appeared
spinning as "Restarting…" for a restart nobody had started.

* fix(macos): clear the reset failure once the permission is granted

"Couldn't reset the permission" stayed on screen after the user granted it in
System Settings and the probe read allowed, contradicting the ticked step
above it. Its sibling line was already gated on the same verdict.

* fix(macos): end the folder-access remedy with the evidence it is about

Two ways out were missing. A reset that cleared the evidence closed the dialog
but left the toast on screen, because only the poll retired it; the store now
retires the notice whenever a verdict comes back null, so both callers get it
and the hook's own branch goes. And the opened scope survived a verdict for a
different scope, so the original one returning later reopened the dialog with
nobody having asked for it.

* fix(macos): keep the folder prompt off main's spawn path

The app-side readability check moved from accessSync to opendir when the
notice was added. TCC lets accessSync through but gates opendir, so on a
machine that has never granted Orca the folder, spawning a terminal there
raised the macOS sheet and froze main until the user answered it. The read is
async now and the spawn no longer waits for it. The blocking variant keeps a
name that says so, and the reset module's own copy of the read is gone.

* fix(macos): only say a folder is still blocked when something re-read it

Two paths reached "Still blocked after the reset." with no verdict behind it:
an unanswered prompt, where the reset returns the verdict stored before it
ran, and a re-probe that could not answer. The reset now returns the same
access it reports to telemetry, and the line waits for a real denial.

* refactor(macos): let the folder-access refresh decide when to skip itself

The poll handler re-implemented the refresh's own two guards, a null entry
and a settled allowed verdict, so each had to be kept in step by hand.

* fix(macos): stop the daemon blocking on its own folder read

The daemon reads the requested cwd before forking a shell to report whether
it can list it. That read is the one macOS gates, so on a folder the daemon
is refused it could hold the daemon's event loop behind a prompt. It is
awaited now, which leaves the blocking enumerator with no callers.
2026-09-22 01:09:26 -04:00
eb92222e7f feat: support Antigravity as supervised worker (#21705)
* feat: add supervised Antigravity worker support

* fix: address Antigravity worker review findings

* fix: stabilize Antigravity readiness detection

* fix: allow Antigravity resume footer after readiness

* fix(antigravity): make agy reach worker_done as a supervised worker

Three defects each blocked `orchestration worker-start --agent antigravity
--worktree new-child` at the agent_readiness stage.

1. Readiness never fired. The composer check required the trimmed line to be
   exactly one character, but agy 1.2.7 launches in accept-edits mode and paints
   it into the caret row (`> Accept-edits mode: ...`). Widened narrowly to a bare
   `>` or `> <name> mode:`; matching any `> <text>` would make every menu dialog
   read as ready, since they all prefix their highlighted row the same way.

2. No trust artifact for agy. Added markAntigravityWorkspaceTrusted, writing
   ~/.gemini/antigravity-cli/settings.json under `trustedWorkspaces` — verified
   empirically against agy 1.2.7, and distinct from the Gemini CLI's
   trustedFolders.json, which agy does not consult. Trust is exact-path and not
   inherited by subdirectories, so each child worktree needs its own entry.

3. The orchestration path skipped the preset. Orca has two trust dispatch
   chains: the renderer's preflightAgentTrust and the main-process
   markLocalWorktreeTrusted. worker-start only takes the second, which matched
   cursor/copilot/codex and fell through for antigravity, so the trust write
   never happened while renderer-side tests passed.

Verified live end to end: the dispatch settles `succeeded` with worker_done
carrying the right task and dispatch ids, and the worktree is appended to agy's
settings with sibling keys untouched.

Known gap: remote-agent-trust-presets.ts has no antigravity branch. The SSH
artifact path is unverified, so agy over SSH still stalls at agent_readiness.
Recorded in a comment there rather than guessed at.

* fix(antigravity): wire trust preset through preload safely

* fix: preserve Antigravity readiness across transcript tails

---------

Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: LielinaH <lielinah@gmail.com>
2026-09-21 20:22:08 -07:00
0677271709 fix(orchestration): reap leaked worker terminals via process-incarnation fallback — stops an unbounded PTY/process leak on Remote Server (OOM / cgroup PID exhaustion) (#18790)
* fix(orchestration): remint live handle from process incarnation on worker release

When a durable terminal handle goes stale (rendererGraphEpoch fence),
inspectWorkerTerminal re-mints a live handle via
resolveTerminalHandleByProcessIncarnation + matchesProcessIncarnation so
release/stop/read act on the still-running PTY instead of reporting
missing and leaking the agent process tree.

- keep main shared host-scope re-exports; add matchesProcessIncarnation
- wire observation.terminalHandle through control/stop/release
- rebuild release-completion on main structured paths
- on missing/unattached + provably exited: settleDead fence first, then
  same-incarnation settleWorker fall back (archive may block settleDead
  mid-request); settle before recovery defer

* fix(orchestration): derive SSH host scope from the reminted handle; reuse fresh-request recovery guidance for structured workers

Addresses two open CodeRabbit review comments on PR #18790.

inspectWorkerTerminal read the dispatch authority with the stale durable
terminalHandle, so after a remint the lookup resolved nowhere and
currentHostScope was always undefined — an SSH worker with no liveness
verdict and no persisted host_scope got classified from terminal.connected
instead of unverifiable. It now reads the same effectiveHandle every other
observation in the function uses.

stopStructuredWorkerForRelease told the caller to repeat the release with
the same --retry-request, which only replays the stale release_unknown
receipt and made a structured-worker close failure permanently unretryable.
It now sources releaseUnknownRecovery from worker-release-completion so the
fresh-request-ID guidance lives in one place.

Pre-commit lint-staged (oxlint + oxfmt) run manually: clean.

* test(orchestration): exercise incarnation recovery through runtime paths

* test(orchestration): pin the incarnation read scenario to the reminted terminal

The read scenario only asserted that the call resolved, so it documented
nothing about which handle the read reached. Assert that the handle
readTerminal received resolves to the registered pane and incarnation, so
the scenario proves the read went through the reminted terminal instead of
passing on the incarnation fence's throw.

* refactor(orchestration): drop redundant incarnation prefix check; require liveTerminalHandle

* feat: add freebuff as a first-class TUI agent (#42)

<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every
commit. -->

| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 0 | 0 | 0 | 0 |
| Prod | 28 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$​37 | 0 |
$\color{#1a7f37}{\Huge{\mathbf{+}}}$​37 |

<!-- /orca-pr-loc -->

## ELI5

Add Freebuff (`freebuff`) as a recognized first-class TUI coding agent
in Orca alongside Codebuff and other supported agents.

## What Changed

- Registered `freebuff` across shared TUI agent definitions,
configuration catalogs, display names, and telemetry schemas.
- Added agent icons, favicons, status mappings, and mobile asset
references for Freebuff.
- Added localization strings across supported language packs (`en`,
`es`, `fr`, `ja`, `ko`, `zh`) and updated locale translation policy.
- Documented Freebuff CLI in README agent table (`npm i -g freebuff`).

## Why

Freebuff is a CLI coding agent twin of Codebuff (`npm i -g freebuff`).
Adding it to the catalog enables users to launch worktrees, run
automated sessions, and pick Freebuff directly within Orca.

## Linked Issue

N/A

## Visual Proof

`N/A` - Catalog registration and metadata definition for CLI agent
launch; UI rendering uses existing TUI agent picker and status
components.

## Testing

- Verified TypeScript contracts, schemas, and catalog configurations.
- Tested CLI detection / agent picker integration locally on Linux
(`worktree create --agent freebuff`).

## AI Disclosure

Assisted by AI coding tooling.

## Checklist

- [x] This PR is small and focused
- [x] I explained what changed and why (including ELI5)
- [x] Before/after screenshots or videos attached for UI changes, or
`N/A` with reason
- [x] Self-reviewed for correctness, security, and performance
- [x] Cross-platform, SSH/remote, and path/shortcut impact considered
(or N/A)

---------

Co-authored-by: Lesley Murfin <lesley@revivebusiness.ca>

* test(orchestration): erase method overloads in worker reap fixtures

* test: document worker fixture type boundaries

* test: simplify worker fixture typing

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: svc-orca[bot] <313947298+svc-orca[bot]@users.noreply.github.com>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-21 17:23:33 -07:00
88f2f01061 fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY (#19430)
* fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY

Root cause: daemon-launched-child.ts forks the detached terminal daemon with
detached: true, which escapes the POSIX process group (setsid) but never the
systemd cgroup. Every PTY the daemon owns is itself an undetached direct
child of the daemon (native-pty-spawn.ts). Under a combined systemd unit
(Type=simple, KillMode=mixed, per docs/reference/headless-linux-server.md),
a systemctl restart/stop SIGKILLs every process still in the cgroup at the
stop timeout -- the daemon and every live terminal -- even though the
codebase already has a fully-built adoption/reattachment path for a
surviving daemon (orcad-entry.ts's refreshRestoredOrchestrationAuthority +
reconcileLegacyWorkerTerminals, gated on daemonOwnsFreshPersistentPtys()).
That path never fires today because the daemon never survives long enough.

Fix: when systemd is actually supervising the process and the OS user has a
reachable systemd --user manager (isDurableDaemonScopeSupported(), Linux
only), launch the daemon via systemd-run --user --scope so it lands in a
cgroup that is a sibling of the service unit's cgroup, not a descendant of
it. A systemctl restart of the combined unit then never reaches it. Any
failure of the scoped launch (no reachable bus, D-Bus policy rejection,
etc.) falls back transparently to the existing plain fork() launch, so
every platform/environment without this capability is unaffected.

The daemon self-detects its own resulting cgroup scope via /proc/self/cgroup
(detectOwnCgroupScopeUnit()) rather than trusting the launcher's intent, and
publishes it as cgroupUnit in its pid record and orcad's health/readiness
payload (health.terminalDaemon.cgroupUnit), so a running deployment can be
observed to confirm the fix actually engaged.

No new session registry is added: the existing daemon pid-record + adoption
protocol (publishDaemonPidFile, daemon-pid-record-quarantine.ts's
dead-record reclaim, refreshRestoredOrchestrationAuthority) already
implements durable, crash-safe reattachment for a surviving daemon -- it
was simply never exercised against a full unit restart before now.

Proven via a systemd-in-Docker recovery test: a live PTY session's shell
process, its daemon, and the daemon's cgroup scope were all confirmed
unchanged across a real systemctl restart of a Type=simple/KillMode=mixed
unit, while the main process pid changed (confirming the unit actually
restarted) and the new process's health payload recognized the surviving
daemon as adopted and live. A fresh write into the same PTY post-restart
reached the same running shell. Ordinary terminal create/work/release and
the #18789/#18790 worker-release reap-fix regression tests are unaffected.

Fixes stablyai/orca#19408

* fix(daemon): probe the real per-UID XDG_RUNTIME_DIR before trusting the process's own env

isDurableDaemonScopeSupported()/buildDurableDaemonScopeCommand() trusted the current
process's own XDG_RUNTIME_DIR env var first, falling back to /run/user/<uid> only when
that var was unset entirely. On mtl-02, orca-serve@factory.service's RuntimeDirectory=
hardening directive makes systemd export XDG_RUNTIME_DIR=/run/orca_serve/factory into the
unit's process -- a private scratch dir that shares the env var's name but has nothing to
do with the user session bus. /proc/<pid>/environ on that host confirmed exactly that path
plus DBUS_SESSION_BUS_ADDRESS=disabled:, while the real bus was reachable the whole time at
/run/user/985 (confirmed via systemctl --user is-system-running with that dir exported by
hand). The probe treated the hardened override as authoritative, found no bus socket there,
and reported unsupported on every launch -- so the cgroup-escape fix from #19408/#19430
never actually engaged on real hardware, even though tonight's factory deployment picked it
up.

Fix: resolveUserRuntimeDir() now always tries the conventional /run/user/<uid> path first
(computed independently via getuid(), never trusted from env), checking for a genuinely
connectable bus socket via statSync(...).isSocket() rather than a bare existsSync. It falls
back to the process's own XDG_RUNTIME_DIR only when that canonical path has no reachable
bus -- covering hosts that legitimately have no /run/user/<uid> at all but do have a
working bus wherever their own environment points. buildDurableDaemonScopeCommand() now
explicitly sets XDG_RUNTIME_DIR to whichever path this resolution picked, rather than
inheriting the spread env's (possibly hardened-wrong) value.

Both isDurableDaemonScopeSupported() and buildDurableDaemonScopeCommand() gained an
injectable canonicalRuntimeDir parameter (defaulting to the real computed path) so tests
can exercise the hardened-override scenario deterministically with a real, connectable
AF_UNIX socket fixture instead of the live host's actual runtime directory.

Docker's stock jrei/systemd-ubuntu test container never had this hardening directive, so
this gap was structurally invisible to the container-based verification in #19430 -- only
caught against real mtl-02 hardware.

* fix(daemon): report the daemon's own pid over the ready handshake, not systemd-run's

The launcher used to infer the daemon's identity pid from the immediate
spawned child (`child.pid`). On the durable-scope path that child is
`systemd-run --user --scope`, not the daemon, so the launcher was asserting
an identity it had no authority over.

`DaemonReadyIdentity` now carries a required `pid` populated from
`process.pid` inside the daemon itself, and `daemon-launched-child.ts` takes
`launchedIdentity.pid` from that self-report. Both sides of the
`holdDaemonAdoptionLease` pid comparison therefore originate inside the
daemon process, which is the idiom this branch already uses for cgroup
membership (`detectOwnCgroupScopeUnit` reads `/proc/self/cgroup` rather than
trusting what the launcher intended).

Note on the reported consequence: `systemd-run --scope` registers its *own*
pid on the transient scope unit and then `execvpe()`s the target command --
same pid, no intermediate process -- so adoption did not in fact fail on
systemd >= 206 (verified against systemd 255.4-1ubuntu8.17 and current main,
`src/run/run.c` `start_transient_scope()`). The fix stands on its own merits:
it removes a silent dependency on that exec-vs-fork implementation detail,
which a `systemd-run` shim earlier in PATH or any future systemd change would
have broken with no diagnostic.

`terminateLaunchedDaemonChild` was audited and deliberately left on
`child.pid`: for the same execve-preserves-pid reason that pid is either
still systemd-run mid-scope-setup (killing it correctly aborts the launch) or
already the daemon, so it targets the right process either way.

Regression coverage: `daemon-launched-child-identity.test.ts` pins the
identity source, and `daemon-ready-identity.test.ts` gains pid-validation
cases. Ready-message fixtures across the `daemon-init-*` suites were updated
for the now-mandatory field.

Addresses:
https://github.com/stablyai/orca/pull/19430#discussion_r3953722704
https://github.com/stablyai/orca/pull/19430#discussion_r3954346518

* test(daemon): assert cgroupUnit in the pid-file parse contract

`parseDaemonPidFile` returns `cgroupUnit` on every branch as of the
durable-scope commit on this branch, but five exhaustive `toEqual`
assertions in daemon-health.test.ts still described the pre-scope shape, so
they failed on the branch independently of any later change.

Adds the field to those expectations. Deliberately not relaxed to
`toMatchObject`: asserting the full parsed shape is what makes these tests
catch a field silently dropped from the pid-file contract.

* refactor(daemon): resolve the canonical user runtime dir at one point

The per-UID path cannot change for a live process, so compute it once into a module
const instead of threading the same default call through three signatures, and drop
the try/catch around a getuid() that cannot throw once it exists. Trims the module
prose to the non-obvious facts and corrects the pid-file record comment: an unscoped
daemon writes null; only records no daemon wrote are absent.

* test(daemon): clean up the cgroup-scope fixtures and assert a verdict

The cgroup fixture tracked only the file it wrote, leaking one temp dir per case.
Drains both fixture lists with splice so the pop-may-be-undefined guards go away,
and replaces a not-throw/typeof-boolean pair with the verdict it was circling:
no resolvable runtime dir means unsupported.

* refactor(daemon): share the detached child options across both launch paths

cwd, detached and stdio were repeated in the fork and systemd-run branches, which
left the two comments explaining them hovering over the env block instead. Names
them once so each branch carries only its own delta.

* refactor(daemon): validate the ready pid like every other field

typeof-first narrows the value, so the two 'as number' casts the isSafeInteger check
needed disappear and the pid guard reads like the startedAtMs guard below it.

* fix(daemon): don't retry the launch unscoped after losing the endpoint race

A scoped attempt that lost the endpoint to another daemon was retried unscoped: a
second doomed fork, a misleading 'cgroup-scope launch failed' warning, and the same
DaemonEndpointUnavailableError the caller was already going to adopt on. Rethrows it
instead, since no launch mode can win a race that is already lost.

Also drops a private alias for DaemonChildSpawnOptions and the two 'as number' casts
on child.pid in the startup-failure cleanup.

* fix(daemon): unlink the pid record by the pid the daemon published

The record holds the daemon's self-reported pid, so match on that rather than on the
immediate child's, which is the systemd-run wrapper's until it execs.

* fix(daemon): route the scope launch through the child-process chokepoint

The two files this PR added imported `node:child_process` directly, which
`child-process-import-boundary.test.ts` fails on deterministically: the
offender count went 155 -> 157 against a pin of exactly 155. Raising the pin
or listing the files is what that test explicitly forbids, and the allowlist's
own note says a split "moved the import, it did not add one" -- so the fix is
to get both new files off the module and put the count back at 155.

- `daemon-cgroup-scope.ts`: the `systemd-run --version` probe now uses
  `runProcessSync` instead of `execFileSync`, so it gets the shared spawn
  decisions. Kept synchronous deliberately: `launchDaemonChild` attaches the
  readiness listener in the same tick it is called, and an await before the
  spawn moves the child past that tick. A non-zero exit is data rather than a
  throw here, so the verdict now checks `code === 0 && !timedOut`.
- `daemon-launched-child-spawn.ts`: the scoped launch uses `spawnProcess`, and
  the long-standing unscoped launch keeps `fork` semantics through a new
  `forkProcess`.
- `src/shared/child-process/fork-process.ts`: the fork arm of the chokepoint.
  `spawnProcess` cannot express a Node child with an IPC channel started from
  a module path under an overridden `execPath`, and the existing launch tests
  are written against `fork`'s contract, so a spawn rewrite would have changed
  module resolution, `execPath` and `execArgv` at once. It passes
  `windowsHide: true` -- the flag every other call site in that directory
  sets, reachable via an assertion because `ForkOptions` omits it -- which
  keeps `windows-console-visibility.test.ts` at its pin of 65 too.

Both ratchets pass with both pins and both allowlists untouched.

Docs: `orcad-operations.md` and `headless-linux-server.md` still described the
limitation this PR removes as permanent. Both now describe the durable-scope
survival path and its preconditions (systemd as PID 1, a reachable user bus /
`loginctl enable-linger`, `systemd-run` on PATH), and scope the old text to
the unscoped-fallback case, pointing at `health.terminalDaemon.cgroupUnit` as
the way to tell the two apart on a running host.

* fix(daemon): seal the cgroup capability probe from the host and correct KillMode=mixed docs

The capability probe consulted the host's own /run/systemd/system marker and
spawned the real systemd-run binary, so the hermetic unit tests could only pass
on a systemd host (and fail closed otherwise, even with faked bus sockets).

- Thread systemdBootPath and runVersionProbe as test seams through
  isDurableDaemonScopeSupported, defaulting to the real boot marker and
  systemd-run --version probe in production.
- Narrow the injected probe to the ProcessResult slice it consumes.
- Cover: no-systemd-boot, non-zero probe exit, and probe-timeout cases.
- Correct KillMode=mixed semantics in the docs: the cgroup-wide SIGKILL fires
  the instant the main process exits, not after TimeoutStopSec; document the
  Docker-container caveat and add KillMode=mixed to the multi-service template.

* fix(daemon): satisfy assertion checks in scoped launch

* fix(daemon): satisfy anti-slop and console guards

* test(serve): update shutdown docs assertions for daemon scope

* fix(daemon): migrate adopted legacy scopes

* docs: qualify restart safety by daemon scope

* docs(daemon): qualify Upgrade restart prose with durable scope caveat

Align the Upgrade section in docs/reference/headless-linux-server.md with
the earlier preservation section and docs/reference/orcad-operations.md:
a service restart terminates live processes only when running under the
unscoped fallback, and stops should be treated as destructive unless
health.terminalDaemon.cgroupUnit names an orca-daemon-*.scope.

Update the shutdown workflow test assertion in
config/scripts/headless-serve-shutdown-workflow.test.mjs to match.

* fix(daemon): harden legacy scope migration

---------

Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-21 17:23:30 -07:00
Brennan Benson 15472cd4c6 feat(native-chat): keep restart recovery available in status bar (#21397)
* feat(native-chat): keep restart recovery available in status bar

* fix(native-chat): source the restart offer from the host and retire it on recovery

Closing the reconnect dialog spent the durable recovery offer, so looking around
before deciding lost the recovery for good. The offer now survives a close, and
the status bar carries it — but a durable offer needs a way to die, and it only
had a reconnect, an explicit dismiss, and a 24h expiry.

The claim's launch-scoped lifecycle moves into its own collaborator, which splits
what the host ADVERTISES from the evidence it holds. A resume-capable hold that
hands a marked chat its provider child back is the recovery the offer existed to
perform, so it stops being advertised and stops being written back at quit, while
the marker stays valid evidence — a user who reopened a chat can still ask the
agent to carry on. Teardown re-derives the snoozed offer rather than round-tripping
raw markers, and this teardown's own witness now outranks the stale claim for the
same chat instead of being overwritten by it, which was silently persisting an old
turn id and making the next launch refuse the chat that was actually mid-turn.

On the renderer the candidate list gets its own producer against
agentSession.restartResumable, so the status entry and the dialog read one
host-owned answer instead of the dialog pushing its local state at a sibling. The
entry re-reads the host before reopening, so a reopened list can never name a chat
the host would now refuse; dialog open becomes the external one-shot request
rather than a flag mirrored into render state, which is what let a reopen replay
the launch answer and re-offer chats already reconnected. Dismiss all is quiet
rather than destructive, saves the preference like every other exit, and reports a
write the host never confirmed instead of trapping the dialog open.

* fix(native-chat): keep a durable offer a launch never read, and settle the one a continuation spent

Teardown replaced the recovery capsule with whatever this launch still owed,
and a launch that never read the offer owes nothing — so a quit after a
failed first read, a disabled flag, or a window that never mounted deleted a
recovery the user was never shown. The write-back now distinguishes "claimed
and still owed" from "never claimed": the first is re-derived as before, the
second carries forward verbatim, because nothing revealed those sessions and
the predicate would refuse every one for want of a journal nobody opened.

Reconnect and continue spent the same claims Reconnect does but never shrank
the offer, leaving the status bar counting chats the host had already handed
back and sending the user to an entry that re-reads, finds nothing and does
nothing.

* fix(native-chat): stop a teardown answering for an offer it could not read

Two ways the write-back deleted a durable recovery offer nobody had seen.

A take that FAILED left the claim holding an empty list and reporting that
this launch had answered for the offer. The markers were still on disk,
unread and unknowable, and teardown then overwrote them with its own empty
list. It now writes nothing at all unless it has a witness of its own.

`owed()` read "has the capsule been touched" where it meant "did anything
here LOOK at the offer" — and its own write-back read counted. Teardown is
retried when a phase fails, so the second attempt re-derived carried markers
against a session map eviction had already emptied, refused every one, and
wiped what the first attempt had just carried forward. The flag is now set
only by the paths that actually read or act on the offer.

The mock guard for the carry could not fail: it indexed the session it
claimed nothing had revealed, so re-deriving passed and the verbatim carry
was never the reason it went green. It now runs against no indexed session,
which is what an unread offer looks like.

Also drops the `Not now` row from the preference table, where it was paired
with a dismiss method it no longer calls, and asserts the same thing where
the snooze is already covered. Splits the marker predicate's journal reader
out of the resume host, which was at its line ceiling.

* fix(native-chat): clear the corrupt recovery capsule the take refused

A capsule whose contents no longer parse made take() throw before it ever
reached the clear, so the bad file survived every launch. Nothing else
rewrites it now that a teardown owing nothing readable declines to write, and
the freshness filter runs after the parse, so the 24h window could not release
it either: one corrupt file refused recovery forever.

Clear it inside the same transaction that failed to read it, then rethrow, so
the poison dies on the next launch while callers still see why the take failed.
A clear that fails is swallowed rather than allowed to mask the parse error.
Refusing to expose partial candidates is unchanged, and a read that fails for
any other reason still writes nothing.

* feat(native-chat): make resume the one restart action, and make it actually resume

The restart prompt offered two actions: "Reconnect all", which reattached
and sent nothing — exactly what opening the chat already does — and
"Reconnect and continue", which reattached and asked the agent to carry
on. The vacuous one is gone, the "Not now" button and the info popover
with it, and the feature is now called resume throughout.

"Don't ask again (resume automatically)" now runs the action the button
runs: the launch calls agentSession.restartContinue instead of
agentSession.restartResume, so the preference means what it says. Several
comments asserted the opposite as a structural guarantee and are
corrected. agentSession.restartResume stays: no in-app caller is left,
but it is a published wire method a non-desktop or older client can call.

* fix(native-chat): label the resume button with the number of chats selected

The button read "Resume all" whenever every chat happened to be ticked,
which described the selection rather than the action. It always acted on
the selected chats only. Now it always names that count, with a singular
variant so one chat does not read "1 chats".

* refactor(native-chat): drop the reconnect vocabulary the resume action left behind

Resuming became one action — reattach and ask the agent to carry on — so the
notification helpers no longer need to be told which action they are reporting.
Every caller passed `continue`; the `reconnect` branch, its helper and its
catalog keys are gone.

The dialog and the launch path had grown two copies of the same call: same RPC,
same response shape, same announce-and-settle. That now lives once in the store
module that owns the offer, which also takes the dismiss call, leaving the modal
presentational. The two copies had drifted — only the dialog's caught a
malformed payload — and the unified one keeps the defensive reading.

No behaviour change. `agentSession.restartResume` stays: it is a published wire
method even though nothing in the app calls it.

* refactor(native-chat): derive the resume selection instead of intersecting it

The modal's selection was intersected back against the host's candidate list
before every action, as a guard against naming a chat the host never offered.
That guard could never fire: the selection was already derived from that same
list, so the intersection was the identity. The array of chosen ids is now the
derived value and the lookup set falls out of it, which makes the property
structural rather than checked. The helper had no other caller and is gone,
along with its three tests.

Three tests mocked the resume response in the shape the old API returned. Two
never reached that branch at all; the third only passed because the unreadable
shape happened to exercise the malformed-payload path. All three now use the
real shape, and the malformed-payload behaviour — report an unconfirmed
delivery, leave the offer standing — gets a test that says so.

Also: the candidate reader took two trailing optional parameters, so one caller
passed a placeholder `false` to reach the second; they are an options object
now. `isFolderWorkspaceId` had no caller outside its own module and is no
longer exported. `RestartActionOutcome` only ever describes a continuation row,
so it is named for that. `dismissAll` set a busy flag that nothing could
render, since it closes the dialog first. Several comments repeated an argument
already made in the module they point at.

Settings: the automatic-resume description is one sentence again.

No behaviour change.

* fix(native-chat): make restart recovery explicitly durable

* fix(native-chat): preserve dismissal fence across new interruptions
2026-09-21 16:38:05 -07:00
Brennan Benson cd59678394 refactor(agent-launch): assemble host startup-plan inputs in one resolver (#22082)
* refactor(agent-launch): assemble host startup-plan inputs in one resolver

buildAgentStartupPlan was already one shared implementation, but every host
re-derived its argument object by hand from the same four settings
(agentCmdOverrides, agentDefaultArgs, agentDefaultEnv, terminalWindowsShell),
and the copies had drifted.

resolveAgentStartupPlanInputs owns that assembly. What genuinely varies per
launch stays a parameter: the host (platform, isRemote), a requested shell, the
per-launch agentArgs override, and the picked session options.

Fixes a live divergence on the agent.launch path: orca-runtime-create-agent-session
passed sessionOptions without sessionOptionsOverrideAgentArgs, so a configured
`--model` in agentDefaultArgs reached argv alongside the picked model and won on
argv order, while the same launch through worktree.create honored the pick.
The plan also reported no applied sessionOptions, so the chat surface could not
name the model the user chose.

Migrates the four host sites; the eleven renderer sites are unmigrated and still
assemble their own inputs.

* fix(agent-launch): preserve picked options in draft launches

* test(agent-launch): assert draft option precedence
2026-09-21 16:15:44 -07:00
297cfe0cf3 fix(usage): price GPT-6 Astra, and declare when the Codex cost total omits a model (#22073)
* feat(usage): price GPT-6 Astra token usage

gpt-6-astra was missing from MODEL_PRICING, so normalizeModelForPricing
returned null and estimateCostUsd dropped every event on that model from
the total. Stats & Usage showed ~$0 for hundreds of millions of tokens
with no unpriced indicator, since hasInferredPricing only covers a
missing model name, not a missing table entry.

Rates are the published ones: $10 input / $1 cached input / $50 output
per 1M, with the >272K long-context tier at 2x input and cache and 1.5x
output, which the existing tier fields already express.

Fixes #22005

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(usage): say when the Codex cost total omits an unpriced model

A daily row whose model has no `MODEL_PRICING` entry gets a null cost, and
`buildSummary` simply skips it. As long as one other row is priced,
`hasAnyBillableCost` is true, so the Codex card prints a confident dollar
figure that silently leaves those tokens out. That is how GPT-6 Astra usage
read as near-$0 before the entry landed, and it is how the next unpriced
model will read too.

`hasInferredPricing` does not cover this: it only fires when a rollout has no
model name at all, and its label ("inferred pricing") describes a guess, not
an omission.

So the summary now carries `hasUnpricedModels`, set when a row has a model
name and no price, and the estimated-cost card appends
"• excludes unpriced models" — the same bullet-suffix idiom the breakdown
rows already use for "• inferred pricing". The number stays; it stops
claiming to be the whole bill.

* fix(usage): caveat the Overview total too, and only when a remainder exists

Review of #22073 found the Codex caveat stopped at the Codex tab. The
Overview tab prints a combined total across providers and already has a
"- some model prices are unavailable" line, but `hasPartialCost` only
noticed a provider whose whole cost was null. A Codex range with one
unpriced model among priced ones kept a real number, so the line stayed
hidden and the Overview repeated the same confident, incomplete figure.
`UsageProviderOverview` now carries `hasPartialCost` — set from
`hasUnpricedModels` for Codex, false for the providers that cannot yet
report it — and the reduction ORs it in. No new string.

Second, the Codex card could read "n/a • excludes unpriced models" when
nothing at all was priced. "Excludes" promises a remainder, and there was
none. The suffix now also requires a non-null total; that case is still
declared, on the Overview, through the null-cost path.

`unpricedCostLabel` becomes `costCardLabel`, since it holds the plain
label whenever there is nothing to qualify.

---------

Co-authored-by: Alfred212121 <58665898+Alfred212121@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 17:37:50 -04:00
Brennan Benson d60043787b feat(agent-launch): carry the launch inputs the host cannot derive (#22037)
* feat(agent-launch): carry the launch inputs the host cannot derive

Desktop's launch call sites cannot move onto `agent.launch` while the wire
drops inputs they depend on. This adds the three the host genuinely cannot
work out for itself, and deliberately adds nothing the host can.

- `agentArgs` — the host read only `settings.agentDefaultArgs`, so a saved
  launch recipe's arguments had no way across. Tri-state is preserved: `null`
  is "no arguments", absent is "use the settings default".
- `cwd` — `TerminalCreateOptions.cwd` already reached the spawn, but nothing
  on the wire filled it. It also decides the route: only a terminal can start
  somewhere other than its workspace, so the host now feeds it to
  `requiresTuiLaunchCommand` and downgrades with `tui_launch_command` rather
  than running a structured session in the wrong directory.
- `launchSource` — telemetry, and the only member of the `agent_started`
  triple the host cannot derive; `agent_kind` and `request_kind` are computed
  host-side. Typed `z.string()`, not the closed enum: params are validated by
  the HOST, so a closed arm set would let an older host refuse a newer
  client's launch over a label. Attribution must not gate a user action.

Not added, because the host already derives them: `launchPlatform`
(`getAgentLaunchPlatformForWorkspace`, from the same connectionId/path/
projectRuntime the renderer uses) and `startupCommandDelivery` (a pure
function of the agent inside `buildAgentStartupPlan`).

Fingerprint: `agentArgs` and `cwd` are in — they change what the call does, so
a retry carrying different ones must conflict rather than replay.
`launchSource` is out — two buttons producing the same launch are one
operation, and folding it in would refuse an honest re-attributed retry. A
caller sending none of the new fields digests exactly as before, because the
canonicalizer drops undefined keys, so launches admitted by an older build
still replay across the upgrade.

Arguments reaching a structured route are ignored by an existing deliberate
decision (the Agent SDK and app-server version their option sets separately
from the interactive CLI), so the host reports it in `warning` instead of
overriding the user's preference on the strength of a field that is not
evidence about the surface.

* fix(agent-launch): forward create-target launch inputs
2026-09-21 13:58:16 -07:00
Brennan Benson 2739246058 feat(native-chat): light the unread indicators when a structured chat finishes (#21924)
* feat(native-chat): light the unread indicators when a structured chat finishes

A structured native chat had no attention producer. The PTY lane reaches the
unread markers through use-notification-dispatch, whose liveness reads PTY
state and whose admission requires terminal panes, so a structured session —
which runs on the execution host with no renderer PTY — could finish a turn
with nothing lighting anywhere. A backgrounded chat was the worst case: with
no mounted pane there was no reader to notice at all.

The host derives the completion, because only the host can. The journal keeps
committing whether or not a renderer holds a reader, so the new feed observes
each commit at StructuredAgentSessionClientDelivery.publishJournal and emits on
every running -> settled transition. That edge runs after the subscriber loop
and independent of it, which is exactly why a chat nobody is watching can still
complete. It is a separate capability-gated stream rather than a field on the
status summary: the summary carries no turn identity and no outcome, and is
re-broadcast on every status change, so folding a completion into it would make
every status consumer a completion consumer.

ONLY `success` LIGHTS ANYTHING. Outcome is A0's provider verdict and is never
inferred: a turn the host merely watched stop carries no outcome and produces
no event, because absent means UNKNOWN. `completed` alone proves nothing — a
provider reports its own API error as a finished turn — so the host emits
nothing for it and the renderer filters again on the way in.

RECOVERY IS LIVE-ONLY. Nothing is retained, queued or replayed on either side.
A subscriber learns what settles while it is subscribed and nothing else; on
reconnect it re-opens an empty stream and whatever landed during the gap is
gone. A retained completion would be a durable "unread is owed" obligation with
nothing to retire it, and a reconnect would then light the dot for work the
user already read. Tests on both sides pin this so a later refactor cannot
quietly turn it into catch-up.

The dot itself reuses the neutral policy in attention/agent-attention-policy
and #21274's structured surface adapter, so suppression, acknowledgement and
addressing keep exactly one implementation and the surface key is never omitted
to evade a check. No second suppression rule is introduced. OS delivery is
deliberately not wired: this calls applyAgentAttentionUnread, not
applyAgentAttention.

Also narrows the completion feed's journal dependency to the newest-turn reader
it actually uses, and adds journal.newestTurn() beside the existing
activeTurnId() on the one shared by-sequence scan rather than a second scan.

* test(cross-version): register the turn-completion subscribe on the wire manifest

The cross-version gate asserts the structured surface's method list by name and
count, so an additive method has to be declared there deliberately. Adding the
entry makes the suite call it in both skew directions and stubs the host side,
which is the statement the gate exists to force.

* fix(native-chat): rebaseline completion feed after rewinds
2026-09-21 13:05:24 -07:00
Brennan Benson 91e6e1f355 fix(native-chat): collapse a finished turn to its answer (#22029)
* fix(native-chat): collapse a finished turn to its answer

A finished turn's "Worked for N" row hid the turn's tool runs and nothing
else. Every sentence the agent said on the way to its answer stayed in the
transcript, so the resting state of a long chat was the narration, not the
reply — one 16m 56s review turn left 21 assistant messages and roughly
seven screens of scrolling behind a control that reads as if it had put
the work away.

The fold's unit is now the turn. A settled turn draws its prompt, its
duration, and its answer; the narration and activity that produced it sit
behind the caret. The answer is the turn's last assistant row that renders
prose — derived, because the journal carries no marker saying which message
is the reply.

Collapsed stays derived rather than stored: nothing closes the disclosure
when a turn ends, it arrives closed because the turn gained a duration. A
running turn therefore folds nothing and the reader watches the work as it
happens, which is what already happened and is now stated rather than
inherited.

Rows that outlive the turn that started them stay outside the fold — a
spawn roster and a background task are often the only record of how that
work ended. So does the reader's own message, question receipts, and the
turn's diff rollup. A turn that produced no prose folds whole, its status
row standing as the anchor.

Two presentation changes come with it, both about the opened view:

- A settled run's header was a call count followed by a monospace list of
  tool names and arguments. It is now one sentence in the transcript's own
  type — "Read 7 files, ran 17 commands, and searched 4 times" — built on
  the tool-category vocabulary that already picks the row's glyph, so the
  words and the icon cannot claim different things. A run of one command
  keeps that command as its header.
- A tool call now owns its result instead of standing beside a separate
  `Result` row, so an opened run lists the work rather than twice as many
  rows half of which say `Result`. Output is one more click. Pairing is
  positional — a result answers the most recent unanswered call — because
  result blocks carry no call identifier to match on.

Command previews also lose the `/bin/zsh -lc "…"` wrapper they all opened
with. The unwrap happens inside `summarizeToolInput`, before truncation,
because the clip at 80 characters removes the closing quote that proves the
wrapper; one site fixes the header, the rows, and the running label.

Measured on a real session journal at 1200x900: the turn above goes from
6,300px across 51 rows to 452px across 2, the whole session from 8,151px
to 2,138px, and the same turn opened from 18,540px to 11,131px.

The fold derivation lives in `src/shared` so the mobile transcript can read
the same rule; wiring mobile's list to it is not part of this change.

* fix(native-chat): preserve FIFO tool result pairing

* fix(native-chat): keep tools collapsed when opening turn

* test(native-chat): clarify independent tool disclosures
2026-09-21 12:13:30 -07:00
Brennan Benson c49b8cd534 Revert "fix(native-chat): stop a subagent's output speaking for the agent tha…" (#22058)
This reverts commit 33ba1ff3df.
2026-09-21 12:11:32 -07:00
Brennan Benson 33ba1ff3df fix(native-chat): stop a subagent's output speaking for the agent that spawned it (#21398)
* docs(attr-parent-label): record the attribution defect and its constraints

* docs(attr-parent-label): add reference findings and the feasibility fact

* docs(attr-parent-label): verify at source and decide the attribution mechanism

Re-baselined against origin/main (one unrelated commit; no drift in any cited
file). Confirmed the two unverified items at source, found a third reader with
the same defect and a fourth append path a naive fix would miss, and recorded
the producer-attribution decision with its field shape, migration behaviour,
wire category, tests and implementation order.

* fix(native-chat): stop a subagent's output speaking for the agent that spawned it

One journal is the durable record of one agent session, but a session that runs
subagents journals their rows into it too, with nothing on the row saying which
agent wrote it. Every "what is this agent doing right now" reader is a backward
scan bounded by markers only the root agent writes, so the window is guaranteed
to hold foreign rows and, while a subagent runs, the newest row in it is the
child's. The sidebar therefore showed a child's prose and a child's running tool
on the parent's row.

Attribute at the producer instead of guessing at the reader. The Claude
translator already parses `parent_tool_use_id` on every envelope and threw it
away; it now stamps `producedBySubagent` on every row that envelope produces,
including the streamed-text path, which persists from a callback with no
envelope in scope and takes the flag from the block identity registry that
already scopes itself on that id. The three status readers skip non-root rows
through one shared predicate. The transcript is deliberately left unscoped: it
shows every agent's output.

No schema version bump, no upcaster, no backfill. An unknown `v` makes a row
unreadable and latches the host read-only, while an unknown key is ignored, so
an older host reads a stamped row and behaves exactly as it does today. Rows
written before the flag read as root, which reproduces today's behaviour for
that history exactly.

* docs(attr-parent-label): add the PR body for the producer-attribution change

* style(native-chat): apply formatter to the merge resolution

* fix(native-chat): preserve producer attribution in resolved appends

* chore: keep attribution review artifacts under docs

* chore: remove review artifacts
2026-09-21 11:31:43 -07:00
Jinwoo Hong f07bf8544c feat(session-search): sort search results by newest, and break relevance ties by recency (#21863)
* feat(session-search): sort search results by newest, and break relevance ties by recency

Results were ordered by match score alone with the session id as the
tiebreak, so equally good matches came out in an arbitrary order and
nothing ever favoured recent work. The Sort menu now offers Most relevant
and Newest while the box has text; the engine already knew both orders and
the all-computers merge already honoured the newest one, so only the panel
had to ask. Under Most relevant, equal scores now go to the newer session.
The choice persists with the other view options, separately from the
list's own Last updated / Created sort.

* fix(session-search): label results by the order they are in

The header subtitle and the results group said "best matches" whichever
sort was chosen; under Newest they now say so. The panel's scope state and
its two context effects move to use-ai-vault-panel-scope.ts, which keeps
the panel under the line cap and gives that behaviour a name.

* feat(session-search): move search sort onto a results bar above the hits

Search mode gets a bar in the group header's place: the hit count on the
left, a ghost menu button on the right that names the current order and
opens the two-item radio group. The filter menu's Sort section keeps one
meaning again (Last updated / Created), the header subtitle stops
reporting sort, and search rows run flat with no group header.

* style(session-search): drop the icons from the results-bar sort menu and match its text size

* feat(session-search): one sort bar above the list in both modes

Filters stay behind the header filter icon; sort moves onto the bar
directly above the session list, in browse mode as well as search.
The bar is mode-agnostic: it takes a label, the selected value, a typed
option list, and a callback, and the panel configures it twice.

- rename AiVaultSearchResultsBar to AiVaultSessionListBar and generalize it
- add ai-vault-sort-options for the two option lists and their aria labels
- drop the Sort section from the filter menu and stop counting sort in the badge
- header subtitle now reads "Indexed history" in both modes

* feat(session-search): count sessions plainly and offer Show more when the scan fills its depth

* fix(session-search): step history depth 250 at a time and keep Show more visible while the rescan runs

* style(session-search): let the sort menu hug its two options

* fix(session-search): show more reads the depth its rows came from

The row inferred "a deeper rescan is running" from the selected depth minus one
page, which at the default depth is zero, so every foreground scan with at least
one session painted a disabled "Loading more sessions…" footer the scan had room
for.

The scan now publishes the depth it ran at beside its sessions, and the row
compares the two: it survives the rescan because that depth trails the selected
one until the deeper scan lands. Drops the stepping arithmetic and
nextAiVaultSessionLimit, and moves the row out of the menu file it was sharing.

* refactor(session-search): an untitled group is what hides a header

Search mode said "no group headers" twice, in two files, both keyed off the same
flag: an empty label in the filters hook and a hideGroupHeaders prop on the list.
The label is now the only fact. A null label means the group has no header of
its own, the list renders its rows flat, and the prop is gone.

The shared group type keeps its string label so the mobile sections that map it
are untouched; the nullable label is the renderer list's own type.

* refactor(session-search): plain labels, and a browse bar that can report zero

Three small simplifications around the list bar:

- The browse bar is guarded on the loaded history rather than the filtered rows,
  so "0 of 250 sessions" can actually appear when filters hide everything and
  the sort control stays reachable. Search keeps its own guard.
- The two count labels were components whose whole body was a ternary over
  translate; they are functions returning a string, and the bar's label prop is
  a string.
- The persistence guards stop being exported with no caller outside the file,
  and the search-sort guard reads the AI_VAULT_SEARCH_SORTS list instead of
  respelling the union.
2026-09-21 14:11:40 -04:00
Brennan Benson 663d670878 feat(agent-launch): deliver a launch prompt to a terminal agent (#21891)
`agent.launch` could hand its initial text to a structured session but not to
a terminal. The contract already anticipated the terminal half — the
`handed-to-terminal` arm has been declared in agent-launch-intent.ts since the
receipt was written and had zero producers — and the executor's own docstring
recorded the assumption behind the gap: that a terminal's paste belongs to the
pane owner. That assumption is what this overturns. The host owns the PTY, so
it can write into one whether or not any window is open on it, which is why
mobile and the CLI got an agent and no prompt.

A terminal takes its prompt one of two ways, and which one is not a
preference. `argv` exists so multi-line and special-character text reaches a
CLI as one argument rather than keystrokes, and it has no readiness race
because the text is in the process's arguments at exec time. So an agent whose
CLI accepts a prompt argument gets it on the launch command, and only a
`stdin-after-start` agent — plus any reused terminal, whose process started
before the launch existed — is written to as a bracketed paste.

That fork is asked once. `agentPromptRidesLaunchCommand` is derived from the
same injection table `buildAgentStartupPlan` branches on, and
tui-agent-prompt-transport.test.ts pins the two against each other for all 37
agents, so adding an agent or changing its mode fails loudly instead of
silently dropping that agent's prompt.

Reused rather than rebuilt: `sendTerminalAgentPrompt` is the runtime's one
agent-prompt writer (bracketed paste, per-PTY serialization, lifecycle
generation pinning, per-agent submit timing, and local/WSL/SSH routing), gated
by `waitForTerminal('tui-idle')` — the same pair orchestration's worker
dispatch already delivers a preamble through. The agent-first create path
needed no new mechanism at all: `startupPrompt` already flows to
`buildWorktreeStartupForAgent`, and the launch had simply been stripping it as
a reserved field without re-supplying its own.

Receipts stay consequences of the act they name. `handed-to-terminal` is
reported only from a launch command that carried the text or a PTY write that
returned; everything unproven under-claims as `not-delivered`. No fourth arm.
The one inversion is a stalled submission, which the verifier raises after the
write: that is reported as delivered, because a resend would paste the whole
prompt a second time into an agent already working on it.

A prompt the launch command cannot carry is refused at the terminal-create
resolver rather than dropped, since that path returns options and has no PTY
to fall back to.

`delivery: 'draft'` remains `not-delivered` for both surfaces. The host could
paste a terminal draft without submitting it, but it cannot observe that the
composer accepted it, so a receipt claiming delivery would be a guess.

No call site is migrated, nothing is added to the wire, and placement and tab
creation are untouched.
2026-09-21 09:19:34 -07:00
Neil 4feaaf5c5c feat(terminal): configure interactive Unix shell arguments (#21904)
* feat(terminal): configure interactive Unix shell args

* fix(settings): clarify Unix shell argument defaults

* fix(settings): improve Unix shell argument guidance

* fix(settings): simplify shell argument guidance

* fix(settings): clarify empty shell args

* fix(settings): explain empty shell args

* feat(settings): make shell argument modes explicit

* fix(settings): keep no args inside custom mode

* fix(terminal): apply configured shell args on the renderer spawn path

The renderer's pty:spawn handler builds options in ipc/spawn-options, not
the runtime controller, so the configured profile never reached a terminal
pane. The local launch plan also dropped the args whenever shellOverride
was set -- which the spawn path always fills from terminalDefaultShell.

Both spawn paths now share one resolver.

* chore(i18n): allowlist the new terminal shell argument strings

Matches how the sibling Terminal shell settings strings are already handled.
2026-09-21 01:35:58 -07:00
Neil ac4d6b407a fix(devin): skip workspace trust for Orca launches (#21925)
* fix(devin): skip workspace trust in yolo launches

* fix(devin): migrate existing workspace trust defaults

* fix(devin): persist migrated launch arguments

* fix(i18n): include required source control stop label
2026-09-21 00:43:27 -07:00
Neil 98299d879b fix(terminal): persist a parked remote pane's scrollback across a hard restart (#21295) (#21367)
* fix(terminal): route a parked pane's scrollback patch to the remote host's partition

A park capture changes only terminalLayoutsByTabId, so its debounced session
patch carries no tabsByWorktree. splitWorkspaceSessionByHost built its
tab->worktree index from the patch alone, resolved nothing, and routed every
layout to the 'local' partition, where main's pruneLocalTerminalScrollbackBuffers
strips scrollback it cannot attribute to a remote worktree. The remote host's
runtime:<id> partition never received the capture, so anything parked since the
last clean checkpoint was lost on a crash, SIGKILL, or a forced kill during an
app update (#21295).

Route tab-keyed patch fields with the renderer's live tab catalogs as a fallback
when the payload names no tab rows. Payload rows still win, so full-payload
writes are byte-identical. Once routed to runtime:<id>, main merges the
partition's own prior tabsByWorktree and the prune preserves.

Proven by tests/e2e/paired-remote-terminal-parked-scrollback-restart.spec.ts: a
hard kill (no checkpoint) then relaunch, asserting the capture is in the remote
host's partition on disk. Mutation: reverting the routing fix turns that
assertion red and fails the 3 catalog-dependent unit routing tests.

(cherry picked from commit 58a344c1d4)

* test(terminal): read both scrollback homes in the restart spec, and ratchet the resolver to the cap's home list

The restart spec read only buffersByLeafId, but the ordinary park now writes
localOnlyScrollbackByTabId, so its own proof reported a false zero and both tests failed for the
wrong reason. Both readers now go through resolveLeafScrollbackBuffers: the on-disk reader calls
it directly (it is a pure function), and the store reader — which runs inside page.evaluate —
reaches it through a new window.__terminalParkingDebug.resolveLeafScrollback(tabId) handle.

resolveTabScrollbackBuffers is typed off TERMINAL_SCROLLBACK_SESSION_HOMES and its unit test
enumerates that constant, so adding a third home fails to compile and fails a test until the
resolver reads it — the 'no consumer reads a home directly' invariant becomes enforceable.

The clean-quit control no longer asserts tokenAfterReveal (measured true, true, false on identical
product code; the live host can serve the reveal from its own tail). It keeps the five
deterministic fields and logs the reveal; the hard-kill test still asserts it, because there the
host is forced unavailable and the reveal must come from the client copy.

(cherry picked from commit 0cd3db1489)

* docs(persistence): pin why the local-only scrollback home stays outside full normalization

The two scrollback homes look symmetric (TERMINAL_SCROLLBACK_SESSION_HOMES), so the missing key
reads as an oversight. It is load-bearing: adding it would route the field through the fail-closed
strip and reintroduce the loss this branch fixes. The renderer prunes it with attribution before
the patch is sent, so the cap still holds without main as a second line.

(cherry picked from commit 1092e35b0a)
2026-09-20 22:57:33 -07:00
Jinwoo Hong 5d13a70ea3 fix(mobile): keep an in-page hop local only when the session's grants cover it (OTA phase C, C2.9) (#21723)
* feat(mobile): carry what each page route declared in init (OTA phase C, C2.9)

The page decides an in-page hop from `init.pageRoutes`, which says which patterns
this shell would render and nothing about what each one costs. So a push kept
local on the strength of the pattern alone runs the target under the opener's
grants — which is how the tasks page is reached from the wide-layout sidebar
without `native.clipboard.write`, and why its copy actions refuse silently.

`init` now also carries `pageRouteGrants`, the manifest's own route/grant pairs,
from the manifest the shell already holds. Optional in both directions: an older
shell omits it and an older page ignores it, and a page that receives none keeps
today's rule. No new frame kind, no cap change, no protocol bump.

The grammar is the manifest's, imported rather than restated
(`MobileWebBundleGrantNameSchema`, now exported for this), so a grant name the
bundle could not have declared cannot reach the page through this field either.
The host validates the pairs before it builds the frame and refuses the session
when they fail, for the reason it already refuses a malformed route: an `init`
the page would reject whole is worse than no session at all.

Two files were at their line ceiling and are split rather than bumped. The pairs
schema moves to `bridge-page-route-grants.ts`, which is read by both the envelope
and the host, so it belonged in one place anyway. In the session reducer the
three sites that each spelled out "patterns, their grants, this route's grants"
become one `routeViewOf`; that is a net reduction and removes the fourth spelling
before it is written.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep a hop local only when the session's grants cover it

The rule the page was using is "the shell would render this pattern", and that is
not the question. Grants are resolved once, from the route the shell opened, so a
push kept local runs the target under the opener's list. On a wide layout the
sidebar renders beside every `/h` route and pushes `/h/<id>/tasks` through this
seam, so from the worktree list, agent history or the files pages the tasks page
ran without `native.clipboard.write` and its copy actions refused with nothing on
screen to say why.

`servedHere` now means served here *and* covered: the target's declared grants
must be a subset of this session's. An uncovered page route is handed to the
shell exactly like a non-page route, and the shell opens it as its own session
with its own grants — which is the mechanism that already exists, rather than a
new one.

Three answers, not two, because an absent field is not an empty one. A shell that
sent no pairs keeps the old behaviour: `null` is "nobody told me", and an older
shell has to keep working. A target the shell lists but names no entry for is
*not* covered — the page cannot justify that hop, so it hands it over rather than
guessing in the direction that loses grants.

This is C3.1's explorer ⊇ preview finding without its pairwise pin: that hop is
covered by this rule and stays local, and the rule scales to the sidebar, which
reaches every route and which no pairwise list can keep up with.

Red first on the two cases only the new rule answers; the other four are the
regression guards and passed before and after. Two whole-session assertions
gained `pageRouteGrants: null`, which is what the reader now returns.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): prove the sidebar hop in a browser, under the session's own grants

The unit tests pin the decision; only a browser shows the control exists, is
reachable at the viewport where the sidebar renders, and that the document does
not move when the hop is handed over.

Four cases on the shared harness, which now forwards `pageRouteGrants` (omitted
when a caller names none, because an absent field is not an empty one and the
page reads the difference).

- Wide, session without `native.clipboard.write`: tapping Tasks posts exactly one
  `navigate` notify, the document stays on the worktree list, and **no new chunk
  is fetched** — which is what says the page did not quietly render tasks under
  the wrong grants.
- Wide, same tap with the grant added: no notify, the document moves to `/tasks`.
  Without this the first case would pass on a page that simply never navigates.
- Wide, shell sending no pairs at all: the old behaviour, local. An older shell
  must not start handing every hop over on a field nobody sent.
- Narrow: asserts the absence rather than a tap. `app/h/_layout.tsx` renders the
  sidebar only on a wide layout, and only that header branch labels its Accounts
  and Tasks controls; the narrow header's are unlabelled pressables. So the hop
  does not exist at that viewport, and `getByLabel('Tasks')` finding nothing is
  the honest assertion. That unlabelled narrow header is a real accessibility gap
  and is not this lane's to fix.

Registered in `pr.yml`'s `mobile_web_app` job beside the other render checks.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): census the in-page hops a session's grants cannot cover

The rule landed in the commit before this one decides each hop; this says which
hops those are, so a route's grants growing — or a new push between two page
routes — shows up here rather than as a verb that silently refuses on a device.

Openers are every page route, not the one that happens to push. On a wide layout
`app/h/_layout.tsx` renders the worktree-list sidebar beside every `/h` route and
its header pushes tasks, which is exactly why a pairwise pin is the wrong shape:
the sidebar reaches everything, so the census has to be the cross product of what
the manifest declares against what the source actually builds.

Targets come from the hrefs the app builds, read out of `mobile/src` and
`mobile/app` and reduced to route patterns, so a hop nobody writes is not pinned
and a hop someone adds is. A presence case asserts the sidebar's tasks push is
among them, because a census that stopped finding hops would go quietly green.

Two hops are pinned as handed off today, both into tasks, which is the only route
declaring more than `navigate` and `storage`. A third case asserts the other half
of the rule on the manifest: a target asking for no more than its opener stays in
the document.

Checked that it discriminates rather than assuming: widening the worktree list's
grants to cover tasks fails the pin, and restoring them passes it.

**No pin was deleted.** The brief expected C3.1's pairwise explorer/preview pin to
be replaced here, but C3.1 is not on this base — `MOBILE_WEB_PAGE_ROUTES` has
three routes and no `files` entry, so there is nothing to remove. When C3.1 lands,
its pin is this census's to subsume.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop an unused import from the hop census

`statSync` was imported and never used; `oxlint` fails it. My error: I committed
the census on a green test run without waiting for lint, the same order mistake I
made earlier in this lane. Fixed forward rather than amended, because the lane
forbids rewriting a commit that exists.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): fold C3.1's pairwise grant pin into the hop census

C3.1 landed while this branch was open, and it brought the case this lane
generalises: the explorer pushes to its own preview, that push stays in the
document, so the preview runs under the explorer's grants. Its pin asserted that
one pair by name.

The census now covers it as a consequence rather than a rule. With the files
routes in the manifest the cross product finds six more hops the session cannot
cover — the sidebar into files from the worktree list and from agent history, and
both files routes into tasks — and it does **not** find explorer → preview,
because the preview declares no more than the explorer. That absence is the
pairwise pin, derived.

So the pairwise block is deleted, with its import. The rest of that file stays:
its external-link seam checks and its clipboard-absence control are about what
the files closure contains, which this census says nothing about.

Checked the extended census still discriminates: granting the explorer
`native.clipboard.write` fails the pin, restoring it passes.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): prove the sidebar hop from a files route, not only the worktree list

The defect is not "the worktree list pushes tasks". On a wide layout the sidebar
renders beside every `/h` route, so the same hop exists from the files explorer,
whose session carries `externalLink` but not `native.clipboard.write`. One opener
proving the rule would have left the general case to inference, which is the
inference C3.1's pairwise pin already made once.

Opened on `/h/<id>/files/<wt>` with the files route's own grants, the sidebar's
Tasks control posts exactly one `navigate` notify, the document stays on the
files route, and no new chunk is fetched.

The harness helper now takes the route and the text to wait for, so a case can
open on something other than the worktree list without a second copy of it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make the render helper wait on the text its caller named

The `awaitText` parameter I added in the commit before this one was never wired
into the wait, so it was dead and `oxlint` failed it. The case still passed,
because the files route renders the host name in its sidebar and that is what the
helper was still waiting on — which is exactly the kind of accident a dead
parameter hides.

Third time in this lane I have committed on a green test run before lint
finished. Fixed forward, not amended.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): carry route grants through the download path

`onManifestRead`'s download branch set `pageRoutes` and `routeGrants` from the
new manifest and dropped `pageRouteGrants`; nothing downstream recomputes it, so
every first install and every OTA update reached `ready` with the default or the
previous generation's pairs. The page then read each target as listed-with-no-
entry and handed off every in-page hop.

`routeViewOf` moves to `page-route-policy.ts`, beside the two functions it calls,
to keep the reducer under its line cap without a bump; its stale neighbouring
comment, which described a filter that moved into it, goes.

Red first: the cold-cache and generation-change cases failed, the cached-hit case
already passed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): derive census targets from navigation call sites

The reachability filter was inert. Harvesting every `/h/${…}` template caught the
five screens that declare their own mount pathname, two `pathname ===`
comparisons and the route template types, so every declared route was reachable
through its own mount: the pinned table was the all-pairs one, eight hops with
the filter and eight without.

Targets now come from the arguments of `router`/`navigation` `push`, `replace`
and `navigate`, and of `navigateFromHostList`; mounts, comparisons and types are
excluded by construction because they are not navigation arguments. Two real
hops are not written as a literal, so a local binding or a call is followed one
step to the function that returns the pathname: the files explorer is pushed as
`{ pathname: descriptor.pathname }` and the preview as
`push(createMobileFilePreviewHref(...))`. A call site whose target cannot be read
is returned rather than dropped.

Derived patterns go from 11 to 10; the pinned table stays at eight because all
five page routes are genuinely pushed to. What changes is that the filter now
discriminates: deleting the header's two tasks pushes reds the presence case and
drops the four `-> tasks` rows from the pin, where the old derivation stayed
green on the same deletion because `app/h/[hostId]/tasks.tsx` still declared the
pathname. A push added at a real call site appears in the set.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): restore the preview-declares-something guard

The pairwise pin this case replaced asserted the preview declares at least one
grant before asserting the explorer covers them all; without it two empty lists
satisfy the subset check and a route that lost its grants passes.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): describe the route list under the handoff rule

Two passages described the world before this PR: the explorer's note said the
census pins its pair with the preview, and a closing paragraph left the sidebar's
tasks hop open for a later PR. This is that PR. Covering the preview now buys the
in-document hop rather than making it correct, an uncovered target is handed to
the shell and reopened under its own grants, and the census reads the explorer to
preview relation off this list rather than pinning it by name.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): mirror the manifest's tasks grants in both fixtures

CodeRabbit on #21723: both fixtures declared the tasks route as `navigate`,
`storage`, `native.clipboard.write` while the manifest also declares
`externalLink`, so no covered-session case ever required it.

Both now mirror the manifest's four, and the covered sessions hold them. That
alone does not make an `externalLink`-blind rule fail, since those sessions hold
every grant either way, so the unit suite gains the case that does: a session
holding the clipboard but not `externalLink` must still hand the hop off.
Mutating the rule to treat `externalLink` as always held reds that one case and
no other.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): make a stalled hop name its own cause

Both waits for the hop to land read as a bare 30 s timeout when it does not. The
CI failure that sent this file back was a `TypeError` inside React Navigation
that blanked the document, and it was invisible here because the error
assertions run after a wait that never returns.

The wait now throws with the page's own account: the pathname it stayed on, the
collected page and console errors, the `navigate` notifies posted, the first 300
characters of the body, and every `.js` response since the click with its status.
The response listener records every script answer rather than only the 200s, so a
chunk the navigation waits on can be seen failing; the 200-only list the
no-new-chunk assertions read is unchanged, as is everything the five cases
assert. Kept in this file because no other render file waits on the pathname
moving.

Proved by mutating the rule to hand every hop off: the covered case fails naming
the pathname it stayed on, an empty error list, the notify it posted and no
scripts since the click — which is the handoff signature, distinct from the
crash signature CI saw.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): aim the narrow hop at the control C2.10 named

The narrow case asserted the absence of a labelled Tasks control, which
was true only because the narrow toolbar carried no accessibility props.
C2.10 gave it the wide sibling's role and label, so the assertion was
red on the merge and, worse, the rule this file is about went unproven
on the branch the phone actually presses.

It taps that control now: at 390 px there is exactly one, and the tap
posts exactly one navigate notify for the tasks route while the document
stays on the worktree list and fetches no new chunk. Red first against
the merged header (count 1, expected 0); with the session given
native.clipboard.write the hop goes local and the case reds, which is
what says the assertions discriminate.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): drop the handoff predicate's contradicted one-liner

The pre-C2.9 summary said the answer is whether this document renders
the target, which is exactly the claim the block comment below it
replaced: the predicate now also requires the target's grants to be
covered. Two doc comments on one declaration, the first of them wrong.

Comment only; the 35 handoff cases are unchanged and green.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): assert which field a route refusal blames

The host builds `pageRouteGrants: <issue>` so a refusal says which of the
two checked inputs failed, and nothing read it: the case counted
refusals, so a host that reported the route's own verdict for a malformed
pair would have stayed green while sending whoever reads the refusal to a
pathname that was never the problem.

The case pins the prefix, a non-empty issue behind it, and that the
diagnostic and the callback carry the same string. The control is an
opener that fails the other way: a malformed route reports its own issue
and does not take this prefix, without which the pin would hold on any
reason at all.

Red first with the field branch dropped from the reason: the prefix
assertion fails and the control stays green.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): stop exporting the route filter the reducer stopped calling

`implementedPageRouteEntries` and `implementedPageRoutes` were the
reducer's two ways in before it moved to `routeViewOf`. The entries form
had no caller anywhere afterwards and the patterns form had only this
test, so the module's public surface advertised two functions no product
code reaches. Both are module-local now; the surface is
`matchesRoutePattern`, `pageRendersRoute`, `grantsForRoute`,
`routeViewOf` and the grant list.

The test reads the same list through `routeViewOf(...).pageRoutes`, which
is the reducer's own view of it, so no assertion changed and no export is
kept for a test.

Red first: with both un-exported and the test untouched, seven cases fail
with `implementedPageRoutes is not a function`; routed through the view
all nineteen pass. Still discriminating, as a control: with the grant
filter dropped from the entries helper, four of them fail.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): route the merged haptics cases through the policy view

PR E's two haptics cases arrived with the merge calling
`implementedPageRoutes`, which this branch had already made module-local,
so the merged file was red with `implementedPageRoutes is not defined`
on both of them. They read the same list through `pageRoutesOf`, the view
the rest of the file already uses, so neither assertion changes.

PR E's paragraph named that function for the filter it describes; the
filter now sits in the entries helper the view is built on, so the
sentence says that instead of naming a function the reader cannot see.

Red: the two cases above on the merge. Green: all 21, PR E's two included.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): mirror the haptics token in every handoff fixture

PR E put `haptics` on all five manifest routes, and these fixtures still
carried the pre-E grant lists: tasks with four grants where the manifest
now declares five. A fixture that is short the same token on both sides
of the subset check agrees with the rule by accident, and would have gone
on agreeing after the token stopped being universal.

The pairs mirror the manifest now, and each session carries what its
opener route would actually be granted, since the host narrows a route's
declared grants to what the shell implements and the shell implements the
token.

Red first, with the token added to the pairs alone: the two covered-hop
cases flip to handed-off, `stays in this document when the session
already covers the target` and `keeps the hop in the document when the
session covers tasks`. Green once the sessions carry it, 35 and 5.

The hop census needed nothing: it reads `MOBILE_WEB_PAGE_ROUTES` itself.
Measured there, all 5 routes declare the token and it is the missing
grant in 0 of the 8 uncovered pairs, so it cannot decide a hop and the
rule still reads only `pageRouteGrants`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(config): count C2.9's two bridge modules in the session route closure

#21908 recorded this pin at 4,324 for the haptics notify module. C2.9
adds two more that the same closure reaches: the page-route-grants schema
and the manifest contract whose grant grammar it imports rather than
restates, both pulled in by `bridge-envelope.ts`, which the page reads to
parse `init`.

Named in the docstring beside #21908's sentence rather than folded into
its number, because the three modules arrived from two PRs and a single
count with one reason invites the next author to assume the rest.

Red first against 4,324: expected 4,326. Measured on this head, not
inferred -- a control worktree at pristine main gives 4,324, so the two
are this branch's.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 01:29:02 -04:00
Neil 253f0e3946 Fix Antigravity source-control model discovery and retired defaults (#21606)
* fix(antigravity): discover current source-control models and use CLI defaults

* fix(antigravity): gate configured models on remote runtime support

* fix(runtime): forward default TUI agent for remote git generation

* test(runtime): cover inherited agent forwarding
2026-09-20 22:12:02 -07:00
Pablo Werlangandorca-agent 646fa3645f fix(opencode): attribute shared-server sessions to their panes (#21577)
* docs: allow-list opencode tool-readout follow-up note

* fix(opencode): attribute shared-server sessions to their panes

The v2 shared server stamps every hook post with its own frozen pane,
so all panes' status lands on the starter pane (#21359).

- shared: session->pane registry plus ingest-time envelope rewrite;
  bound sessions resolve to their real pane, tab and live launch token
  before disposition, unbound sessions keep the stamped identity.
- main: binder poll (SQLite session store, PTY-registry pane snapshots,
  argv-aware client sweep) with directory-containment plus
  client-lifetime correlation; 60s loop plus debounced SessionStart kick,
  wired into the hook server lifecycle.

* fix(opencode): newest-wins pane dedupe, macOS private/tmp normalization

Live verification against the dev instance found two binder gaps: remint
rows for one pane counted as an ambiguous tie, and /tmp vs /private/tmp
spellings never met on macOS.

* fix(opencode): review fixes — newest-wins worktree, drop dead constant

- applyBinderOwnerships now overwrites per-pane worktree, matching the
  round's newest-wins pane dedupe; a remint's live row wins over a stale
  row (pinned by test).
- remove the unused OPENCODE_CLIENT_PRE_CREATE_WINDOW_MS export and the
  nowMs residue from clientCouldCreate.
- give the per-pane launch-token cache its own named cap constant.

* fix(opencode): address thread review — cursor, native table, tokens, lifecycle

- composite (time_created, id) store cursor advanced past handled rows
  only, so same-millisecond pagination and full unbound maps no longer
  drop sessions silently.
- Windows sweep reads the native process table instead of forking
  powershell.exe; quote-aware argv parsing on both platforms.
- directory keys via normalizeRuntimePathForComparison (Windows
  case-fold, POSIX backslash literals) plus narrow macOS /tmp|/var|/etc
  aliases and lexical dot-segment resolution.
- bound sessions always take the stored pane token (never the frozen
  stamp); token tracking runs after resolution.
- binder generation guard discards post-stop rounds; first round runs
  immediately at loop start.
- unbind/move use exact pane-key match; pane launch-token cache gets its
  own cap constant.
- move the tool-readout note out of this PR for its own branch.

* fix(opencode): second review round — executable field, worktree scope, round lifecycle

- POSIX sweep reads comm= alongside args= and classifies on the
  kernel executable name, so unquoted install paths with spaces no
  longer split argv[0] and reject the client; Windows rows carry the
  native table name. Degrades to argv[0] when comm is unavailable.
- bound sessions take only the binding's worktree (never the stamped
  pane's), so a worktree-less binding cannot file a row under the
  wrong worktree.
- the binder generation is captured before the round body and the
  running flag clears only for the current generation, so an obsolete
  post-stop round cannot admit an overlapping round.

---------

Co-authored-by: orca-agent <orca-agent@local>
2026-09-20 20:06:55 -07:00
Seongho BaeandCursor ea5152f1c2 fix(orchestration): line-settle delay for antigravity multiline paste (#21665)
* fix(orchestration): retry Enter after cursor-agent worker-start paste

Worker-start dispatches through bracketed paste in the main process; cursor-agent
can leave long prompts as "Pasted text +N lines" and swallow the first Enter.
Apply the same submitRetryDelayMs path Codex uses in the renderer, but only for
agents without the Claude/Codex render gate so hook turn-start reservation stays intact.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(orchestration): line-settle delay for antigravity multiline paste

Antigravity 1.2.x expands long bracketed paste slowly ("↑ N more lines") while
Orca only waited for byte ingest (~500 ms on macOS). Add submitLineSettleMsPerLine
and retry Enter for antigravity; wire agent-aware submit scheduling through the
main-process prompt writer and plain terminal.send suffix path.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(orchestration): antigravity line-settle only; drop unverified retry

Address PR review: revert accidental pnpm-lock.yaml churn, remove cursor and
antigravity submitRetryDelayMs until live-verified, keep submitLineSettleMsPerLine
for agy multiline paste, and move the regression test out of the 900+ line runtime
submission suite.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-20 18:01:41 -07:00
Neil 438744ca77 fix(opencode): preserve global config discovery (#21854) 2026-09-20 17:58:37 -07:00
OrcaWinandm4air 4085e1cf60 fix(memory): release stale session registries (#21734)
* fix(memory): bound session and lifecycle registries

* fix(memory): bound transient filesystem registries

* fix(memory): cap path and locale caches

* fix(memory): bound runtime recovery registries

* fix(memory): bound host mirror gap verdicts

* fix(memory): bound shell startup env cache

* fix(memory): bound gitlab host context cache

* fix(memory): release removed ssh generations

* fix(memory): expire cloud refresh replay guards

* fix(memory): release retired plugin generations

* fix(memory): bound plugin log key retention

* fix(memory): bound automation authority generations

* fix(memory): bound native chat enrichment cache

* fix(memory): bound web session tracking generations

* fix(memory): bound codex credential absence paths

* fix(memory): bound WSL canonical path cache

* fix(memory): bound sparse checkout cache

* fix(memory): bound shared directory cache

* fix(memory): bound advertised URL scan snapshots

* fix(memory): bound automation manager cache

* fix(memory): bound web session reorder intents

* fix(memory): bound web session focus intents

* fix(memory): bound web session handoffs

* fix(memory): bound automation dispatch tokens

* fix(memory): bound host mirror waiters

* fix(memory): bound retained session activity

* fix(memory): bound retained session activity

* fix(memory): bound web session close intents

* fix(memory): bound cloud session cache

* fix(memory): bound WSL home cache

* fix(memory): bound SSH capability cache

* fix(memory): bound trust grant cooldowns

* fix(memory): bound WSL auth drain state

* fix(memory): bound Linear workspace credential cache

* fix(memory): bound local Git capability cache

* fix(memory): bound WSL Git environment cache

* fix(memory): bound WSL Git environment cache

* fix(memory): bound WSL preflight cache

* fix(memory): keep hot cache entries warm

* fix(memory): preserve generation fences across eviction

* fix(memory): close remaining eviction fences

* fix(memory): align evicted upstream generations

* fix(memory): trim successful capability probes

* fix(auth): retain expired refresh replay evidence

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-20 14:41:50 -07:00
Jinwoo Hong d5dc7b9cf8 feat(mobile): budget the terminal snapshot on serialized bytes and hold live output instead of ending the stream (OTA phase C, C7.3) (#21785)
* fix(mobile): budget the mobile terminal snapshot on the bytes it serializes to (OTA phase C, C7.3, ruling 1)

The desktop trims a mobile snapshot to 512 KiB of raw terminal text. A client
reading it through the page bridge measures the serialized event against a
640 KiB frame cap, and an ANSI snapshot is mostly ESC bytes, each of which JSON
spends six on. Measured here on a colour-dense 80-column screen: the raw budget
hands back 465,766 bytes that serialize to 669,268 — 102.1% of the cap — so
`deliver` answers `cancel(id, 'overflow')` and the terminal is dead before its
first live byte, with no recovery that does not reproduce it.

`terminal.subscribe` gains an optional `snapshotByteBudget`. A subscriber that
sends one is trimmed against the JSON its payload will really cost: the escaped
text, plus the metadata it cannot bound from its own side — a path, the OSC-link
list, the pending escape tail. A subscriber that sends none, which is every
socket client and every older page, keeps the raw byte rule exactly.

No negotiation, and none is needed: the field is additive and optional, so an
older desktop ignores it and trims as it always did. The page then still has a
snapshot over its cap, the shell still ends the stream with `overflow` (C0.3
stands), and the terminal renders its stream-error state rather than a blank
pane. The page derives the number from the cap less the event envelope rather
than writing it down, so a cap that moves takes the budget with it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): hold and coalesce terminal output instead of ending the stream on the window (OTA phase C, C7.3, ruling 2)

The shell's backpressure window ends a stream when the page falls 4 MiB behind.
That is right for a stream whose reader can survive a gap and wrong for a
terminal, whose reader cannot see the hole a dropped chunk leaves — and the
window does not wait for a page to go wrong. Measured by the design: the host
produces 70.3 MiB/s of JSON and real xterm applies 2.2 MiB/s, so an ordinary
`cat` crosses the window in 62 ms. Replayed here through the real ledger against
a page draining at that rate, a 5 MB transcript ends the stream after 85 of 107
chunks plain and after 40 of 107 under `grep --color`.

Keyed by method on the shell, since the page cannot pick its own window,
`terminal.subscribe` now holds what it cannot send, merges consecutive output in
escaped bytes under the frame cap, and delivers as the page acks. Nothing is
dropped: merging concatenates, and the only exit that loses bytes is ending the
stream, which the page is told about. Both transcripts now arrive whole and in
order, in 104 and 81 frames, with the largest frame at 622,551 bytes against the
655,360-byte cap.

It ends only on the two things that are not slowness: a page that has acked
nothing for 20 s, an order of magnitude above the 1.9 s a full window takes to
drain, and a backlog past 32 MiB, which at that drain is about 15 s of catching
up. Both reach the page as `overflow`, because the shell is the installed app
and its page comes from the desktop, so a reason the page's reader has never
heard of is a frame it drops rather than an end it acts on. Which one fired, the
coalesced-frame count and the peak pending bytes go to the diagnostic log, which
is the device proof's only oracle for any of this.

Every other stream keeps the byte window exactly, and an event over the frame cap
still ends any stream, terminal or not (C0.3). The landed window cases now name a
stream the window still governs, so the two rules are never read off each other.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): narrow the event arm the backlog replay reads

A binary event carries no `payload`, so the tests-typecheck ratchet refused the
reach into it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): narrow the snapshot serializer to the buffer source it reads

The changed-code casting gate refused the test's stub runtime, and it was right
to: a service-wide type for a function that calls one method is what made the
stub need an assertion. The parameter now says what it needs, and the fixture
path is no longer one a machine-path grep reads as a leaked local checkout.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): measure the snapshot budget by building the payload, not by summing fields (OTA phase C, C7.3, ruling 14)

Round one summed the escaped text and four metadata fields. The payload a
bridged client assembles carries nine more — `kind`, `cols`, `rows`,
`requestId`, `displayMode`, `reason`, `seq` and both truncation flags — plus the
`type` and `streamId` it adds, the `serialized` key and the object's own braces.
So a snapshot this host accepted at exactly the budget, with
`truncatedByByteBudget` false because nothing had trimmed it, published over the
cap and the stream ended with `overflow` before a byte was painted.

Measured here on a screen sized to land exactly on round one's budget: the
published payload is 655,446 bytes against a 655,273-byte budget, 173 over, and
the frame it makes is over the 640 KiB cap by the same amount.

The metadata is now built by one function that `sendSnapshotFrames` and the
budget both call, and the budget stringifies the payload that function produces.
Nothing is summed and nothing is estimated, so a field added to the frame is paid
for by the budget the moment it is sent. Where a value is not yet known — the
truncation flags, and `seq` or `requestId` at a site that has not fixed them —
it is measured at the widest `JSON.stringify` can write it, which is a bound
rather than a guess, and forcing `seq` to a number also opens the three fields it
gates so those are counted too.

The budget therefore travels with the publication fields, because the payload
cannot be built without them.

On the page, the event envelope is now derived in one place in the protocol
module and read by both the snapshot budget and the shell's own merge budget, so
the two cannot drift; the page pins the number it sends and the host's cases name
that pin, since the two programs cannot import from each other.

The case that re-implemented the host's measure is gone: it could not have seen
this, because it was the same arithmetic twice.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): arm a held terminal's silence clock only while something is pending (OTA phase C, C7.3)

The invariant is "armed implies waiting on the page", and round one broke it in
the one direction that kills: an ack re-armed the clock and the drain that
followed emptied the queue without clearing it. A terminal that had delivered
every byte and gone quiet — which is what a terminal does between commands —
would die on `overflow` twenty seconds later.

The clock is now synchronised after every change to the queue, so it is armed
exactly while something is held. A rule that only ever arms is a rule that only
ever ends more streams.

Red-first: with round one's arming, an idle stream whose queue has drained still
reports its clock armed, and firing it ends a healthy terminal.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the held-stream cases the rulings name (OTA phase C, C7.3)

Six cases nothing covered. Two subscriptions on one shell keep separate backlogs,
so a busy terminal cannot end a quiet one. A stream the page unsubscribed mid-
backlog posts nothing after, and neither does one that has already ended, however
much was still held. A payload that is not output breaks a merge run and keeps
its place, because a resize is state the reader applies in order. And the budget
boundary is checked on the side that enforces it: a payload at exactly the number
the page asks the desktop for is delivered inside the cap, and one the cap cannot
hold ends the stream under C0.3.

The replay no longer acks unconditionally in its catch-up loop. That was the page
behaving better than a page can — it acks on reading frames — and it is what hid
the silence clock left armed over an empty queue. The held-stream cases close the
window on its frame count rather than on four megabytes of string work.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): give the event-envelope derivation its own module (OTA phase C, C7.3)

`bridge-envelope.ts` is at its line cap and is the protocol's schemas; what a
frame costs around its payload is a derivation over them, and two budgets read
it — the snapshot the page asks the desktop for, and the output the shell merges.
One module, so they cannot drift and neither file is pushed over its limit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test: narrow the budget fixtures instead of asserting them

The changed-code casting gate refused six `as NonNullable<...>` in the new
budget cases, and it was right to: a fixture that serialized nothing is a broken
case rather than a null to assert away.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor: give the snapshot payload shape its own module (OTA phase C, C7.3)

`terminal-snapshot-publication.ts` crossed the root config's 300-line cap, which
mobile's own lint does not apply and CI does. The frame's shape and what it costs
a client reading it as one payload is a description the budget and the sender
both need, so it is the part that leaves.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix: empty a snapshot the budget cannot fit instead of posting it over (OTA phase C, C7.3, ruling 15)

Both trimming loops published the zero-row candidate whatever it measured, and
zero scrollback is not a small screen: a wide colour-dense viewport still carries
its 24 live rows. A capped subscriber could get one frame over its cap, end the
stream on `overflow` and paint nothing — worse than a blank terminal, because a
blank one repaints on the next byte of output and a stream that never opened does
not reopen.

Ruling 15: a budgeted subscriber gets that frame with its text emptied and
`truncatedByByteBudget` true, never over and never refused. The raw rule keeps its
fallback, so an older page and every socket client are served exactly what they
were before. Below the metadata the frame must carry there is nothing left to give
up, and that boundary is pinned rather than claimed away.

The renderer loop is the same walk reached by a different caller and had no test
at all; its runtime parameter is narrowed to the two methods it reads so a case
can stub it without a cast.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): report what a held terminal stream did instead of calling it an outlived view

The backlog report had no branch in the reporter, so it fell through to the "a
view outlived its host" warn and every field it exists to carry was discarded.
The key made it worse: keyed by kind alone, one backlog per host was ever logged,
and a shell holds one stream per open terminal.

That report is the only oracle the coalescing rule has. Nothing crosses to the
page saying how much was held or how many frames its bytes arrived inside, and
both ways a held stream dies reach the page as `overflow`, because a reason its
reader has never heard of is a frame it drops. In production the two rules were
indistinguishable. They are now a line each, per stream.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test: give the renderer fixture the source its serializer returns

`serializeRendererTerminalBuffer` answers `renderer`, and vitest does not
typecheck, so the stub's `headless` passed every run and failed the node
typecheck instead.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix: budget the frame the publication actually sends (OTA phase C, C7.3)

The budget and the publication were written out twice, five lines apart, and had
drifted at every site: a budget for `{kind:'scrollback'}` approved a frame sent as
`kind:'resized'` with a `reason` beside it, and the live module budgeted
`pending-output-overflow` while sending `renderer-mount-ready`. It held only
because the padded `requestId` and `seq` are absent from those frames and more
than covered the difference. Each site now builds one object and hands it to
both.

`displayMode` cannot travel that way and was a third under-measure nobody had
named: the subscribe flow re-reads it from the runtime after the snapshot is
serialized and before the frame is sent, so no caller can tell the budget which
mode the publication will carry. It joins `seq`, `requestId` and the truncation
flags as a field taken at its widest. The mode list resolves the constant to
`never` if the runtime gains a mode it does not carry, so a new one is weighed
here rather than found on a phone.

Red-first needed a second attempt: the first fixture had trimming slack, so three
extra bytes fit and the probe could not see the defect it was written for. The
case now budgets a fixed screen at exactly its `auto` measure, where the margin
is the whole of the test.

One figure for the overshoot everywhere, with its basis: 169 bytes over the
655,360-byte cap on a frame carrying an 8-character request id, 247 with a
24-character one. Three places said 169 and one said 173.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): delete two backlog guards no input can reach

Both survived mutation because neither is reachable, and neither became
reachable when I tried to write a case for it.

`next` narrowed the merge ceiling to one frame, but its only caller,
`drainTerminalBacklog`, has already narrowed it: the parameter is what one
payload may occupy, not what the window holds, so the second narrowing could
never change the answer. The parameter now says so and the class no longer needs
the frame size at all. The bound still lives in the caller and is still covered:
removing it there reds a delivery case.

The merge run also compared stream ids, but a backlog belongs to one subscription
and every `data` payload on it carries that subscription's single stream id, so
the comparison could not fail. The run still stops at anything that is not
output, which is reachable and pinned.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): record the invariant the deleted stream-id guard rested on

The merge run compares no stream ids because it cannot need to: a backlog belongs
to one subscription and every `data` payload reaching it carries that
subscription's single stream id. Written down where the run is, because the thing
that would break it is a change made somewhere else — multiplexing two streams
onto one record would merge their output into one payload under the first id.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-20 10:19:41 -04:00
Jinwoo Hong f5d2d6e757 feat(mobile): carry browser screencast frames over the bridge as base64 (OTA phase C, C6.1) (#21758)
* feat(mobile): carry screencast frames over the bridge as base64 (OTA phase C, C6.1)

`bridge-screencast-binary.ts` landed in C0 as the page's half of the binary
lane and named C6 as the owner of the encoder that satisfies it. This is that
encoder, plus the host honouring `wantsBinary`: a subscribe that asked for
binary gets an `onBinaryFrame` on the native stream, and each frame crosses as
the envelope's `event.binary` on the same `seq` ledger as the stream's JSON
events, because the page acks by that count.

The base64 encoder is grouped rather than per byte or per `fromCharCode`
window. Its docstring carries the measurement, including the part that
contradicts the design note this came from: on V8 the per-byte form is the
fastest of the three, not the quadratic one, and the chunked form it was meant
to beat is the slowest. The grouped one is here because its cost does not
depend on how an engine ropes `+=`, and Hermes is what the shell runs.

No new opcode, no `v` bump, no negotiation added: `wantsBinary` is already in
the contract and is the negotiation. Over-cap behaviour is unchanged in this
commit — a binary event over the frame cap still ends the stream, which is what
C6.2 changes.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): drop an over-cap screencast frame instead of ending the stream (OTA phase C, C6.1)

Measured at the pane's own request parameters, a screencast frame exceeds the
640 KiB envelope on a phone layout whenever the page will not compress: JPEG's
worst case is 0.545 bytes per pixel at quality 72, so mobile view mode at
780x1424 is 811,289 bytes, 124% of the cap. Ending the stream there blacks out
a browser tab for the life of the pane over one frame.

So the two kinds of event part at the cap. A JSON event that will not fit still
ends the stream with `overflow`, because its reader cannot see the hole it
would leave; a screencast frame is dropped and the stream lives, because the
next frame is one throttle interval away and the pane is still showing the last
one. Both are asserted side by side so neither turns into the other.

A drop leaves no other trace: the diagnostic beside it prints once per host, so
a stream shedding a frame a second and one that shed a single frame read the
same. The host therefore counts them per stream for the diagnostic and keeps a
session total, and the shell's dev facts carry that total — the surface that
already shows build state, with the line moved into its own module so what it
says is pinned rather than inferred from a template. The 12-character build
prefix it has always shown is unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): name the binary screencast lane as a grant (OTA phase C, C6.1)

Ruling 5's negotiation, and the check it asked for first: no reader of a grant
is a closed enum, so there is no blocker and nothing an older page has to
tolerate. `BridgeGrantsSchema.native` and the shell's manifest reader are both
open string arrays, and the shell reader's own docstring already states the
degradation — a grant name a build does not know leaves that one route native
rather than refusing the bundle.

What does constrain the name is the host contract's `GRANT_NAME_PATTERN`: a
grant is one camelCase token or a `native.<domain>.<action>` verb with at least
two dot segments. So `browser.screencast` and `native.screencast` are both
refused, and the lane is `screencastBinary`. `screencast` alone would be wrong:
the page can already subscribe to `browser.screencast` and receive its JSON
events, and only the binary frames need the encoder.

Added to the shell's implemented set, which is the same list `init.grants.native
` offers, so a route declaring it is served by a shell that has the encoder and
left native by one that does not. No route declares it here; C7's session route
does.

The contract-side case is a characterisation pin, not a red-first one: the
pattern already admitted this name, and the test records that the two tempting
spellings are the ones it refuses.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): check the dropped-frame total through the bridge hook (OTA phase C, C6.1)

The hook gained a required `onBinaryFramesDropped` two commits ago and this
test kept calling it without one, so the tests-typecheck ratchet went red on
that commit — caught here rather than in CI because an exit code was read off a
pipeline's last stage instead of the script.

Fixed by wiring the callback into the probe rather than by a cast, and with the
case that makes the wiring evidence instead of types: a dropped frame raises
the total the screen receives, and the stream stays subscribed while it does.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): keep the dropped-frame counter with the ledger it belongs to (OTA phase C, C6.1)

Declared between a getter and a method, which is not where this class keeps
state: the subscription map is at the top and the counter is the same kind of
thing. Move only.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): serve the binary screencast lane only to a route granted it (OTA phase C, C6.1)

Reported as a gap after C6.1's third commit and ruled on: the host honoured
`wantsBinary` from any page, so a route that never declared `screencastBinary`
could still make the shell encode base64 on its behalf. That is the hole
per-route grants exist to close — the same class as a route granted only
`navigate` and `storage` reaching the clipboard.

The rule now reads the session's resolved list, which is what its route
declared narrowed to what this shell implements, and is the same set
`init.grants.native` is built from. So the host offers the lane in `init`
exactly when it will serve it.

Ungranted is not a refusal. The subscription proceeds and its JSON events cross
as before, which is the silence every other grant gives at the call site; a
page that reads its own grants never reaches that state. Both branches are
pinned beside each other, and `grantsForRoute` is pinned dropping a grant this
shell does not implement — granted-but-unimplemented and never-granted arrive
at the host as the same absence, so its rule reads one case.

The grant name moves into the module that holds the rule reading it, so the two
cannot drift. `bridge-host.ts` is at 298 of its 300-line cap after this.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): move the page's stream-frame rules out of the host (OTA phase C, C6.1)

`bridge-host.ts` reached 298 of its 300-line cap, so the next main merge that
touched it would have crossed under CI pressure on someone else's PR. Split
deliberately instead, at the boundary the growth came from.

`bridge-host.ts` is the host's lifecycle and its dispatch. Opening a stream is
the only frame kind whose handling is more than one line of delegation — four
refusals and, since C6.1, the binary-lane decision — so it moves whole, and
`cancel` and `ack` move with it so all three stream frames are decided in one
place. The host's `cancel` arm still chooses between a stream and a request
where it always did: a page's `cancel` names one or the other, and splitting
that choice would leave half an arm in each module.

Counted without blank lines or comments, as the rule counts them:
bridge-host.ts 298 -> 270, and the new module is 59.

A pure move. No test changed and none was added, which is what makes the
existing suites the proof: 45 files and 745 tests green on the same assertions
as before.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): report a page that asked for screencast frames it was not granted (OTA phase C, C6.1)

An ungranted `wantsBinary` is not a refusal on the wire, so nothing crosses
back: the subscription proceeds and its JSON events cross as they always have.
That left a page which did ask getting JSON for the life of the document with
no side able to say why. `notify-refused` has covered the equivalent notify
case since C0; this is the same shape for the one frame kind that lacked it.

The rule now answers a verdict rather than a boolean, because `not-asked` and
`ungranted` are the same answer for different reasons and only one is worth
reporting. So the decision and the report read one rule, and a page that never
asked stays silent — pinned, along with a granted route staying silent, so the
line cannot start firing on either.

The wire is unchanged and pinned unchanged: the case beside this one still
asserts one JSON event delivered and zero error frames.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): reset the dropped-frame total with the host that counts it (OTA phase C, C6.1)

Round 1 on #21758, three findings.

The real one: the count is per host and the screen's copy was not. A rebuilt
host starts its own total at zero, so the screen kept the retired host's number
until the new one dropped a frame and then read *lower* — a falling count looks
like frames coming back, which is worse than starting over. The hook now
announces a fresh count as it builds a host. That also reports zero on the
first build, where the screen is already at zero and React bails out of the
render; the two hook cases pin that leading zero rather than leave it to be
rediscovered.

Two docstrings that described nothing: `BUILD_ID_PREFIX_LENGTH`'s stayed behind
when the constant moved to the dev-facts module and had drifted above
`failureMessage`, and `page-route-policy.test.ts` kept the docstring of the
test it replaced above the one that replaced it. Both deleted; the first's text
lives on the new module.

Red-first for the reset, checked against its final expectations rather than its
first: with the one line reverted both hook cases fail on the missing zero, and
both pass with it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the dropped-frame total out of a production build's render path (OTA phase C, C6.1)

CodeRabbit's Major on #21758. The total went into React state on every dropped
frame in every build, and outside a development build the line that reads it
renders null — so an over-cap page re-rendered the whole shell screen up to ten
times a second for a fact nobody can see. Measured, not argued: five drops,
five extra renders.

Fixed at the seam rather than with a ternary at the call site. The dev-facts
module owns the line, so it now owns the number behind it and the rule that the
number is only state where something renders it. The screen holds no flag and
no counter; it asks for both and passes the reporter on. The reporter is stable,
so the bridge host is never rebuilt for it.

`isDevelopmentBuild` becomes a call rather than a module constant. A build flag
never changes at runtime so this costs nothing, and as a constant the branch was
unreachable to anything that did not set the global before the module loaded —
which is why the production case could not be written at the screen at all.

Also fixed, found while writing that case: the screen test's
`usePageHostSnapshot` double returned a fresh object on every render, so the
host effect's identity changed each time and the bridge host was torn down and
rebuilt on every render of the screen, settling every pending request with it.
The real hook holds the snapshot in `useState` and is stable. One object for the
file now. This was masking the fold under test — the count reset to zero on
every render — and every other case in that file was measuring a rebuild storm.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* perf(mobile): price a screencast frame before encoding it (OTA phase C, C6.1)

Round 2 on #21758, two lows.

The encode is a base64 pass over the whole image and the window decides whether
the frame can be posted at all, so deciding after encoding made a page that had
stopped acking pay for every frame the shell then threw away — the reviewer's
case is ten 300 KB frames against a closed window, 3 MB encoded and nothing
sent. The size is knowable without encoding: base64 is ASCII, so JSON escapes
none of it and the frame is its header serialized plus exactly the image's
encoded length. `encodeBridgeScreencastFrame` is now built from that header
rather than beside it, so the shape measured and the shape sent cannot drift,
and the window arithmetic is one rule read before the encode and again on the
frame that was.

Exact, not conservative, so the drop diagnostic still reports the whole frame
and the committed byte pin is untouched.

Red-first with the real encoder wrapped in a counter: window full, ten frames,
ten encodes before and zero after, with the drop count still ten. An over-cap
frame likewise goes from one encode to none. A third case holds the other
direction — two carryable frames still encode twice — so the fix cannot pass by
encoding nothing.

Second low: the dev-facts block sat outside the only `beforeEach` and left
`routeGrants` and `client` mutated, inert only because it runs last. The shared
setup moves to file level where the mutable dependencies actually live, resets
both, and a case at the end of the file pins it — deleting the reset fails
there and nowhere else, since nothing else runs after a case that mutates them.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-20 04:19:00 -04:00
NeilandXiro The Dev ee354a35d7 feat(agents): add OpenCode 2 beta support (#21418)
* feat(agents): add OpenCode 2 beta support

Co-authored-by: Xiro The Dev <lethanhtrung.trungle@gmail.com>

* fix(opencode2): support current plugin lifecycle and session storage

* fix(opencode2): preserve lifecycle ordering and full session capture

* test(opencode2): cover setup event bridge

* test(opencode2): cover setup event bridge

* test(browser): satisfy anti-slop naming check

* test(opencode2): cover live form lifecycle

* fix(relay): preserve OMP config directory selection

* test(opencode2): avoid assertions in bridge fixture

* fix(rebase): retain OMP resume and fresh launch behavior

* test: align upstream OMP resume expectations

* test(opencode2): verify rejected form closes waiting state

---------

Co-authored-by: Xiro The Dev <lethanhtrung.trungle@gmail.com>
2026-09-19 17:49:03 -07:00
a445abadd4 fix(browser): bound CDP output for stalled clients (#20949)
* fix(browser): bound CDP output for stalled clients

* fix(browser): log CDP outbound overflow before terminating the client

The outbound queue terminated the automation client silently on overflow, so
the client saw a socket close indistinguishable from a crash. Surface the cap
that tripped and the backlog held when it did.

The queue dropped its backlog before invoking onOverflow, so the counters were
already zero at the callback. Snapshot them first and pass them through.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Neil <neil@stably.ai>
2026-09-19 17:24:33 -07:00
Neil e4c7632db2 perf(terminal): skip kitty scans for plain PTY output (#21643)
* perf(terminal): skip kitty scans for plain output

* fix(terminal): keep the kitty scan fast path total for absent chunks

The new escape-byte fast path dereferences the chunk before the string
concatenation that used to coerce a nullish value, so an unchecked
caller now throws instead of no-opping. Normalize once at the top.

Also type the AgentTerminalPreview connect mock against the real preload
signature, which turns the stale bare-string replay fixture that tripped
this into a compile error.
2026-09-19 16:26:32 -07:00
3e7da29767 feat(editor): add opt-in collapsed unchanged regions for file diffs (#11955)
* feat(editor): add opt-in collapsed unchanged regions for file diffs

The combined "View All Changes" diff already collapses unchanged lines into
expandable bands (DiffSectionBody sets Monaco's hideUnchangedRegions), but a
single-file diff opened from Source Control renders the whole file. Reviewing
one changed line in a long file means scrolling past everything else.

Adds a General > Editor setting, default off, that applies the same Monaco
option to the single-file diff viewer. Off keeps today's full-file rendering.

The option is always emitted rather than omitted when off: Monaco retains the
last applied value across an options update, so dropping the key would strand
an open diff in collapsed mode after the setting is turned back off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(settings): register collapse unchanged search entry

* fix(editor): keep diff viewer under line limit

* fix(editor): satisfy diff viewer line budget

---------

Co-authored-by: Dan Cieslak <dcieslak19973@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Neil <neil@stably.ai>
2026-09-19 15:53:28 -07:00