Commit Graph
15 Commits
Author SHA1 Message Date
Brennan BensonandClaude d443320af2 refactor(native-chat): remove the unused terminal handoff (#22783)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-25 10:17:36 -07:00
Brennan Benson 25d7c21fcb feat(native-chat): show context window usage in the composer (#22301)
* refactor(native-chat): move the composer's stop action into its own hook

The composer is at its line budget; lifting the stop action out makes room
for the context usage ring without changing what Stop does.

* feat(native-chat): record the Claude CLI's context window facts on the structured journal

A structured Claude session now keeps what the CLI says about its context
window on journal rows, so every client reads the same answer and a restart
replays it:

- Each main-thread assistant response keeps its API usage. A subagent's
  response measures its own window, so it carries none.
- The turn a result settles records the session's window: the largest
  contextWindow across the result's per-model usage, since side calls to a
  smaller model report their own smaller window.
- After a result and after a compaction boundary the host asks the CLI for
  its /context breakdown (5s bound) and records the answer on the current or
  last turn. An answer is dropped when the main conversation moved, a send was
  accepted, a newer request was issued, or the session was released while it
  was in flight; a failure or an older CLI leaves the row unchanged.
- A compaction boundary or conversation reset records that the used count is
  unknown until the next response or report, so the pre-compaction size is
  never shown as current.

Every part carries its own host clock, and the reader takes the newest, since
a revised turn row keeps its place in the transcript. The persisted validator
admits every value the writer can write, including a zero auto-compact
threshold: a row replay rejects truncates the journal from that row.

* feat(native-chat): show context window usage in the composer

A structured Claude chat shows a ring beside send once the journal can state
the session's context usage. Hovering shows used/window with a bar and, when
the CLI has reported its breakdown, one row per CLI category as a share of the
window, largest first. Between reports the ring shows the newest response's
usage against the newest window the CLI reported, marked as estimated. Before
the CLI has reported any window, and after a compaction or reset until the next
response, there is no ring. A terminal-backed chat shows none.

* fix(native-chat): measure the context ring against the main thread's model window

The result's per-model usage is cumulative across the session and includes
subagents and side calls, so the largest window was often not the one the
main conversation runs in: after switching from a 1M model to a 200k one, or
when a subagent ran on a larger-window model, the ring read against the
wrong window after every turn. Pick the entry named by the model that served
the newest main-thread response, and among its [1m]/non-[1m] entries the one
the result moved; fall back to the largest only when nothing names it.

* fix(native-chat): keep the context ring moving through tool-only responses

The live estimate lived on assistant message rows, and a response with only
tool calls or thinking writes no message row, so the ring froze through long
tool loops and stayed hidden after a mid-turn auto-compaction until the next
text reply. Record every main-thread response's usage on its turn row
instead, once per response, so the selector sees each one.

* fix(native-chat): read the ring's window from the model the turn's init names

The CLI keys per-model usage by the main loop's model string, [1m] included,
and every turn's system/init frame carries that exact string, while a
response drops the suffix. Match the init's key first, so a session that
switched between the 1M and 200k variants of one model reads the right
window; fall back to the newest response's model, then the largest entry.

* fix(native-chat): correct the composer's control-order note for the context ring

* fix(native-chat): keep a running turn's context facts when the host settles it

A turn row now carries the live context estimate while it runs. When the host
settles a running row itself (a crashed or stale generation, a close the
translator never saw), it rebuilt the record field by field and dropped those
facts, so after a crash the ring fell back to an older turn's size, or to a
pre-compaction size the dropped reset had superseded.

* refactor(native-chat): revise Claude turn rows from the journal so the ring survives a restart

The context ring's facts were written to turn rows through an in-memory list
of recent turns. A new translator is built on every acquisition, so after a
restart or reattach that list was empty and every fact for a turn that was
not open was dropped: a /compact as the first action after a restart never
cleared the ring and never showed the fresh breakdown.

Every Claude turn-row write is now a revision of the row as the bound journal
holds it when the write runs. The sink gains a resolved revision that reads
the target row and its body at execution; the queue runs one operation at a
time, so the read-modify-write cannot interleave, and revisions are never
coalesced. Lifecycle writes own the lifecycle fields and context writes own
contextUsage; each keeps every other field. Only the open turn is kept in
memory. A fact with no open turn lands on the newest turn row, and a report
lands on the turn it was requested for.

The persisted facts are simplified to a window, which now names the model it
was measured for, and a single used part (report, estimate or unknown) that
each write replaces. The ring reads the newest turn row carrying each part,
and hides an estimate whose model the window was not measured for instead of
dividing by another model's window. A turn opening, and a reset, count as
activity, so a late report can never land behind a newer turn.

Host settlement of a stale running turn now drops only the fields its verdict
owns, so context facts and any field a newer build wrote survive it.

* fix(native-chat): keep a turn row whose context facts this build cannot read

Context facts are validated deeply, so one malformed or future-shaped fact
made the whole turn row malformed, and replay truncates the journal from that
row on. Replay now drops unreadable facts from a turn row, in item rows and in
settlement batches, and keeps the row, the same way it already drops producer
linkage it cannot trust. The ring shows nothing for that turn instead of the
session losing its history.

* test(native-chat): pin that a child exit mid-turn keeps the ring's last size

A lifecycle-only revision, the end a turn gets when its child exits without a
result, must keep the context facts the row already carries.

* perf(native-chat): revise a named Claude turn row by key instead of scanning the journal

Every Claude turn-row write walked every reduced journal item to find its
row, even when it already knew the row's identity, so a long session paid
O(items) per write on the main process. The journal now answers a keyed read,
and a context report names its turn by row identity rather than turn id, so
only a write made while no turn is open still scans.

* fix(native-chat): tell a 1M window from a 200k one of the same model

Responses drop the [1m] suffix, so after a switch between the 1M and 200k
windows of one model the running turn was measured against the previous
turn's window until its result arrived. An estimate now records the turn's
init model, which keys the window exactly, and the reader requires the full
model id to match.

* fix(native-chat): show no ring for a context kind a newer host writes

A paired client reads turn rows from the host unvalidated, so a used-count
kind this build does not know fell through to the estimate branch and threw
reading its missing usage. Only the kinds this build can measure now produce
a ring.

* fix(native-chat): keep the context ring through plan-mode turns on another model

Plan mode can run a turn on a model the turn's init does not name (opusplan
upgrades to Opus's 1M window). The estimate then carried only the response's
id, which drops [1m], and the exact comparison against the window hid the ring
for every plan-mode turn. The estimate now records the response's id beside
the init's exact key, and the reader matches the base model only when no exact
key was recorded.

* refactor(native-chat): pair the context ring's window by model change, not by model id

The ring divided the newest response's size by the newest window only when
their model ids matched, which meant comparing ids from the init frame, the
response, per-model usage keys and canonical ids. Those disagree in plan mode
and across 1M and 200k windows of one model.

The writer now knows when the model may have changed: after a model or
permission-mode write that changes the value, when a restore cannot put the
stored model back, and when a main-thread response comes from a different
model than the one the window serves (an approved plan). It then marks the
size unknown, holds estimates, and asks the CLI for its context report, which
states the new model's window. Any new window, from a report or a turn
result, releases the hold. The reader compares nothing: a report, or the
newest estimate over the newest window.

Turn rows no longer store window.model, window.canonicalModel,
estimate.model or estimate.responseModel.

* fix(native-chat): keep a late context report's window when only its count went stale

* test(native-chat): pin that each turn's init lets its result restate the context window

* fix(native-chat): open the context card on click and tap

* fix(native-chat): wait a beat before a mouse hover opens the context card

* fix(native-chat): publish each context write in the operation that makes it

A context report answers after the turn's last frame, so a revision that waited
for the next frame's publish reached live clients only on the next turn. Context
writes now queue their revision and its publication as one operation.

* fix(native-chat): write million-token counts with a capital M

A lowercase m read as minutes on the context card.

* fix(native-chat): keep the context card open while the pointer crosses into it

The card closed the moment a mouse left the ring, so the pointer could not
cross the gap into the card. Leaving now waits a beat, and entering the card
cancels the close.

* fix(native-chat): let Escape close the context card without stopping the agent

The card keeps focus in the composer, so the Escape that closed it also
reached the composer and interrupted the running turn. The composer now
skips an Escape an open layer already handled.

* fix(native-chat): show the context ring when the chat has not loaded the turn it belongs to

The ring read context facts only from the rows the chat had loaded, so a
reopened chat whose recent page started after the last turn row, or a live
turn longer than the retained window, showed no ring until the next turn.

The host now derives the newest context facts from its whole journal with
the same selector the chat uses, and returns them on agentSession.options
for sessions that write them. The chat prefers each fact its loaded rows
carry and takes the host answer for a fact they lack. When a live batch
revises a turn row older than the loaded window, the chat asks for options
again so that answer stays current.

* fix(native-chat): bound context refresh reads and refresh when the turn row is trimmed

Each turn-row revision the loaded window missed started its own options
read. Those reads share the session's host queue with sends and interrupts,
and each asks the CLI for its settings, so a burst could pile reads in front
of a user action and discard every answer before it landed. The chat now
keeps one options read in flight and at most one behind it.

A live turn longer than the retained window also lost its turn row to the
trim without asking for a fresh host answer, so the ring fell back to the
answer read at turn start until the next response. Trimming a turn row now
asks again, like a dropped revision does.

* fix(native-chat): show the context ring from the first response, sized from the session's model

A new session has no measured window until its first result, so the ring
stayed hidden for the whole first turn. The host now keeps the window the
applied model's name implies (1M for a [1m] name, unknown for default, 200k
otherwise) and writes it beside an estimate when the journal holds no window,
or after a model write, until the result or the CLI's report replaces it.

* fix(native-chat): imply a context window only from a [1m] model name

A bare model name does not fix the window: first-party runs today's opus,
sonnet and fable models natively at 1M while a gateway or cloud provider runs
them at 200k, and opusplan and haiku run another model in plan mode. Sizing
their first response at 200k read the ring about five times too full, so only
a [1m] name implies a window now; any other name waits for the result.

* fix(native-chat): size the first response from a report taken before any turn

A model picked in a chat with no turn yet asks the CLI for its context
report, but with no turn row the report's write lands nowhere. Recording it
still marked the journal as holding a window, so the first response wrote
none and the ring stayed hidden until the turn's result.

The report's window now serves as the fallback a response writes while the
journal holds no window, and recording a report no longer assumes its write
landed.

* test(native-chat): move the fake Claude connection out of the structured integration suite

The context-report delivery case pushed the suite past the 800-line limit,
failing repo-wide lint. The fake child now lives in its own fixture.
2026-09-24 00:12:08 -07:00
Brennan Benson 563dd5487f feat(native-chat): show a Codex chat's goal above the composer, and set it from goal mode (#22377)
* feat(native-chat): show a Codex chat's goal above the composer and set it from goal mode

Structured Codex chat now treats the thread goal as session state: a banner above the
composer shows the current goal (pursuing / paused) with clear, pause/resume and expand;
/goal enters a goal mode whose send calls thread/goal/set; the objective is journaled as a
user message marked as sent as a goal. The banner is derived from the journaled goal rows,
which Codex's resume snapshot refreshes, so a reopened or adopted chat shows its goal.

Fixes STA-8159

* fix(native-chat): replace a recorded goal by clearing first, and recover a lost goal-change response

- A set while the journal records a goal (any status) clears it before setting,
  so the new goal starts with its own time and token counters instead of
  rewriting the old goal's objective in place.
- The threadGoal plan answers an unknown outcome from the goal the journal
  records and reruns otherwise, so one request timeout no longer refuses every
  later Clear/Pause/Resume as unknown for the mounted session.
- The goal-mode chip says "Exit goal mode"; "Clear goal" stays the banner's
  action on the provider goal.
- A typed bare /goal on Enter enters goal mode, the same as picking it.
- The renderer reads the goal off the tail of its ordered snapshot; the host's
  unordered map keeps the by-sequence reader.
- Drop the composer's duplicate in-flight guard; the goal controller already
  serializes changes.
- Pin that a counter-only revision reaches a subscriber's live page under its
  original sequence.

* fix(native-chat): keep a bare /goal inside goal mode as the entrance, and pin goal delivery and serialization

- A bare `/goal` submitted while already in goal mode re-enters the mode instead
  of setting a goal whose objective is the literal text "/goal".
- The counter-only revision pin now drives the host's own event sink bound to a
  real journal, so it goes red when the publish after a lifecycle transition is
  dropped; the previous fake sink never published.
- Pin that a set which threw after journaling its objective puts that objective
  back exactly once when the ledger reruns the same operation id.
- Cover the goal controller hook: absent without host support, the loaded window
  wins over the host's answer, a second change while one is unsettled answers
  false without a request, and a refused change frees the next one.

* fix(native-chat): resume a blocked or usage-limited goal, and keep goal-mode drafts honest

- The goal bar offers Resume on a blocked or usage-limited goal, which the
  provider resumes exactly as it resumes a paused one; a goal whose token budget
  is spent still offers only Clear. The rule lives beside the other goal facts
  in shared code so every reader answers it the same way.
- A `/goal <text>` typed inside goal mode sets the objective `<text>`, as it
  does outside goal mode, instead of a goal whose objective is the literal
  command.
- Setting a goal is a host round trip; a draft edited while it was in flight is
  no longer wiped when the goal lands, matching every other host command.
- Pin that a lost status-change response is read as applied only when the
  recorded goal is in that status, that a cleared row in the loaded window
  outranks the host's earlier answer, and that the PTY lane is untouched.

* fix(native-chat): keep the load-older anchor on the loaded window when a live revision lands below it

A live revision of a row keeps that row's original sequence. When the row is
older than the client's loaded window, the shared reducer merged it in and it
became the load-older anchor, so paging `before` it skipped every row between.
A goal's counter-only revisions during a long goal turn reach any client that
attached after the goal row left its window, so a reopened chat lost rows on
scroll-back.

The reducer now admits live rows only at or above the window's oldest row
while older rows remain on the host; the journal keeps the revision and the
page reader serves it once the window reaches the row. With nothing older on
the host the window is the whole journal, so a row below the head is admitted
as before.

Also drain accepted provider events before a goal set reads the journal to
decide whether it replaces a recorded goal.
2026-09-23 10:34:06 -07:00
Brennan Benson a4c11f1889 fix(native-chat): stop a bounded tail read from moving the chat cursor past unapplied rows (#20581)
* fix(native-chat): stop a bounded tail read from moving the chat cursor past unapplied rows

A structured chat pane could latch "Working for N" forever after the agent had
finished, showing the send arrow rather than Stop, while the sidebar and
`worktree ps` correctly read idle.

The client replica has one position (`state.cursor`) and one body. Two
operations keep those consistent: replace (both from one host snapshot) and
append (rows contiguous with the cursor). The `tail-page` branch was a third
thing: it took the cursor from the journal head, the items from a bounded page
(200 items, byte-capped), then merged retained client submissions over the
page's. Under continuous journal writes the client is always slightly behind,
so the branch ran on every window focus and on every pane re-activation. When
more than a page of rows had landed since a send, that send's user item fell
off the page, its submission was not carried, the retained `pending` survived,
and the cursor jumped past the dispatch-acceptance row. Nothing re-sends it: a
batch carries only touched items and that submission is never touched again.

Delete the third operation rather than guard it. A live subscription is now the
only thing that moves the cursor, and `subscribe({ cursor })` already replays
exactly the missed rows.

- remove the window `focus` listener and the owner/transport `refresh` contract
- skip warm hydration: a retained owner subscribes at its applied cursor
- cold hydration keeps its history read, applied as the existing `snapshot`
  (replace) event rather than `tail-page`
- delete the `tail-page` action and its reducer branch
- delete `resumeCursor` and `shouldAdvanceStructuredResumeCursor`; two cursors
  with two advancement rules were how position and body drifted apart

`older-page`/`loadOlder`, the unattached-refusal grace, generation guards and
the coalescer are unchanged. No host, wire or schema change.

Also fixes a second cost of the same branch: focus during a busy turn discarded
paged-in older items, shrinking the transcript to one bounded page mid-turn.

* fix(native-chat): preserve unavailable mixed-version session fences
2026-09-14 10:28:16 -07:00
Brennan BensonandMerge Sim 2626e2eca4 Make the structured turn lifecycle row durable so completed durations survive (#19695)
* Make the structured turn lifecycle row durable so completed durations survive

A structured-chat turn used to end by tombstoning its running lifecycle item,
which threw away the only durable record of when the turn ended. Completed
"Worked for" labels therefore depended on the renderer having observed the
turn finish, and vanished on reopen.

The lifecycle item is now revised in place, never tombstoned:
- running, with startedAt, at the provider's turn start
- completed or interrupted, with completedAt, at the provider's terminal frame,
  a user stop, or a child exit the host observed
- unverifiable, with no end, when a cold acquire finds a running row from a
  generation whose exit nobody observed

Both timestamps are the execution host's clock at receipt, captured before the
deferred sink, so the completed value is identical on every client and needs
no client clock. Codex history restore uses the provider's own second-granular
endpoints for turns that predate this change. Desktop and mobile read settled
durations off the journal through one shared selector, and anchor the live
counter on the host start with the client's local receipt so a skewed client
clock never leaks into the label. Locally observed durations remain the
fallback for hosts that still tombstone.

Timestamps live inside the existing turnLifecycle field, which old clients
strip, and every working-state consumer keys on state === 'running', so no
capability negotiation is needed.

* native-chat: avoid stale working status on settled turns

* test: align settled turn status expectations

* Name settled lifecycle rows by their terminal state

An interrupted or unverifiable turn must not read as completed for any
consumer that renders status text raw. One shared helper builds the text for
both providers from the lifecycle state.

* test: deduplicate turn lifecycle suites

Each behavior keeps one test; duplicated harnesses and restated cases go.

* Key lifecycle rows to their user item and record the provider's measured duration

A lifecycle row now names the user item that opened the turn by its provider
key, so clients attribute timing explicitly and fall back to journal order
only for rows from older hosts. A provider-initiated turn with no prompt can
no longer claim the previous prompt's duration.

When the provider measures the turn itself (Codex turn.durationMs, Claude
result.duration_ms) the terminal row records it and clients prefer it over the
host interval, so a turn shows the same number live and after a history
restore. Host receipt times remain the live-counter anchor and the fallback.

* Record a turn as a first-class journal item

The turn record is now its own item kind rather than a status row carrying a
lifecycle field: no text to misuse, and the fold matches the durable turn
record other systems keep. Rows that carry it are stamped journal schema v3;
every other row stays v2, so an older host keeps reading them and latches
read-only at the first v3 row instead of truncating the epoch.

Clients that predate the item would paint an unknown kind as a text bubble,
so the host publishes the legacy status form to any client that does not
advertise agent-session.turn-item.v1, through the same per-client seam
background tasks use. The downgrade is transitional and goes once no
supported release lacks the capability. The shared projection now renders
unknown item kinds as nothing, so later kinds need no gate. One shared reader
handles both forms for old journals and old hosts.

* Preserve observed turn end across settlement retries

* Retain turn attribution for loaded chat history

* Preserve Codex exit receipt across close retries

* Register completed turn duration reliability gate

* Keep earlier turns through a Codex rewind and count a mid-turn attach from the real start

Findings from an independent adversarial review of the typed turn record:

- A Codex rewind adopted the provider's item list as the new epoch, and the
  provider never returns the host's own turn rows, so every duration before
  the rewind point vanished. The host's turn rows are now spliced back beside
  the item each followed, and recovery no longer expects the provider to
  prove rows it never owned.
- The epoch row was stamped with the current schema version, so an older host
  latched read-only at row 1 of every new session, defeating the mixed
  version design. It carries no body and stays at v2; a stored-row test now
  reads SQLite directly, because the reader upcasts every row on read.
- A send Codex folds into a running turn shares the opening prompt's provider
  key, and the alias map credited the duration to the later prompt. The
  earliest submission naming a key now wins.
- The live counter anchored on first sight, so a client attaching mid-turn
  counted from zero. Published frames now carry the host's clock, the reducer
  keeps the last sample with its local receipt time, and both clients anchor
  on how long the host says the turn has run.

* Correct turn duration gate assertion reference

* Respect authoritative unknown native chat duration

* Preserve unverifiable timing across older host upgrade

* Record final completed turn duration reliability evidence

* Fix the CI failures the merge left behind

- A merged import list named the same module twice, which the native code
  quality plugin fails on.
- A running turn is now reported by the host with no duration, so the settled
  map carries an explicit null for it; the hook test still expected the entry
  to be absent.
- main gave the older-page action a cursor with a head-trim guard, so the
  retention test's epoch-only action no longer typechecks; it now passes an
  unbounded sequence, which is what the old shape meant.
- The roster comparator moved into the extracted module, leaving its import
  unused in the reducer.

* Split two files back under the line cap after the merge

Merging main put both one effective line over 300, and the cap forbids a
disable or a shave. The wire module's refusal vocabulary moves to its own file
and is re-exported, so its consumers are untouched; the host's four thin
mutation delegates move next to the functions they call.

* Advertise the turn-item capability on every client transport

Local IPC and mobile advertised it; the remote and web transports did not, so a
desktop paired to a remote host, the CLI, and web silently ran on the legacy
carrier forever and the canonical row was never exercised there. The renderer
that paints it is the same build on every transport.

* Update the web auth-frame expectation for the new capability

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 14:32:50 -07:00
Neil f2d5711b2d fix(native-chat): keep an older page from punching a hole in the transcript (#19845) 2026-09-10 03:10:04 -07:00
Brennan BensonandMerge Sim 2bf298d1dc feat(native-chat): the background-tasks strip says what is running (#19311)
* feat(native-chat): name, group, and state the background-tasks strip

The strip above the composer described five different kinds of background
work as "Monitoring background tasks", with identical flat-dot rows. Now:

- Wire: additive optional `name`, `state`, `startedAt` on
  AgentSessionBackgroundTask, plus `settledTasks` on the state object so
  terminal siblings of a live fan-out stay visible without changing what
  old clients render (they keep exactly the live `tasks` list).
- Reducer equality learns the new fields, so a publish whose only change
  is a task's state is no longer judged equal and dropped.
- Header counts by kind and lists states within a kind; past three kind
  segments (or on a narrow strip, measured by its own border-box against
  the live root font size) it falls back to an honest total, never a
  partial enumeration, and the strip stays expandable whenever the header
  is lossy.
- Rows group by kind (Agents / Shell / Monitors / Workflows / Tasks),
  stable-sorted first-seen-then-id, each with a kind icon, its own state
  dot, a resolved name (description -> name -> kind label), and elapsed.
- Claude producer: task frames now carry name (agent_type/subagent_type),
  a run state mapped from patch status, and first-seen startedAt. Terminal
  statuses settle a task (completed->done, failed->blocked,
  killed/stopped->idle) instead of deleting it; settled tasks render only
  beside still-live work and flush when the last live task ends, so the
  strip exits exactly when it does today. An unreadable patch leaves a
  task open, never settled.
- Turn gating moves off the strip: the tracker no longer zeroes its
  roster during a foreground turn, and the client renders the strip
  whenever it has contents while the idle-only flag now gates just the
  animated monitoring indicator and conversation commands.

* feat(sidebar): indent native-chat subagents under their session row

buildSubagentChildRows() has always rendered indented children from
parentEntry.subagents, and the structured-session status bridge has
always published an AgentStatusEntry for native chat — it just never
populated subagents. Connect them:

- Wire: additive optional `backgroundTasks` on AgentSessionStatusSummary
  (live tasks only), projected by the host status feed from the
  provider's backgroundTaskState hook and republished on task edges via
  the background-task channel, with the shared task equality suppressing
  no-op re-projections.
- Bridge: maps agent-kind tasks onto the sidebar's own
  AgentSubagentState (working/waiting/blocked, terminal -> idle) — kinds
  stay distinct, so a backgrounded shell never lands in a subagent
  count — and extends its pre-write equality with the existing
  agentSubagentsEqual.
- parentIsFresh for a bridge entry means "the host feed reported a
  change inside the sidebar's ordinary evidence window": every publish
  restamps evidenceObservedAt, and a dead feed stops restamping, so
  children decay to idle on lost contact instead of pinning 'working'.

* fix(native-chat): settle tasks the aggregate roster evicted first; carry usage

Real-agent QA showed settledTasks never rendered. A frame capture from the
SDK (probe against claude 2.1.261) explains it: when a backgrounded child
finishes, the producer emits `background_tasks_changed` FIRST — with the
task already absent — and only then `task_updated`/`task_notification`
with the outcome, in the same tick. The tracker's settle path looked the
task up in the live roster the aggregate had just evicted, so retention
lost the race 100% of the time.

Fix: aggregate eviction of a live backgrounded task now parks its details
in a bounded recently-removed map (new claude-settled-background-tasks.ts,
which also owns the settled roster), and the trailing terminal edge
consumes it. A removal whose outcome frame never arrives still vanishes —
nothing is guessed into a finished state. A second terminal edge for the
same task re-derives the settled state and can add final usage. The
captured sequence is replayed verbatim as a tracker test, including the
kill-at-exit tail proving the strip still exits with the last live task.

The same capture disproved the PR's earlier claim that Claude task frames
carry no usage: task_progress and task_notification both carry
usage.total_tokens. Additive optional `totalTokens` on the wire task,
covered by the shared equality; the tracker takes usage (never the
transient "Running <tool>" description) from task_progress, and rows
render the mock's "18.1k · 2m" meta — settled rows keep final usage with
no still-growing clock.

* chore(i18n): sync runtime-required catalog for backgroundTasks.runningList

* fix(native-chat): preserve background task lifecycle and bound update work

* fix(native-chat): transfer resumed background tasks to one live owner

* fix(native-chat): bring structured session host under the line cap and restore subscribe fixture

* fix(native-chat): complete journal stubs and stop notifying on feed teardown

The status feed's projection cache calls journal.cursor(); the rename test's
stubs are cast through unknown, so the missing method only surfaced at runtime.

Teardown runs only once nothing is activated, so there is no mounted reader to
notify - clearing confirmed sessions is what prevents a stale live on reactivation.

* feat(native-chat): lead each strip header count with its kind icon

The header carried one aggregate state dot, so a fan-out of agents and a
monitor looked alike. Each count segment now leads with its own kind glyph;
a collapsed total spans kinds and takes none.

Monitor is the heartbeat AgentStateDot already draws for monitoring, so the
strip and the agent sidebar speak one vocabulary.

* feat(native-chat): give the strip's monitor heartbeat the sidebar amber

The glyph matched AgentStateDot but the colour did not, so a monitor in the
strip did not read as the monitor in the agent sidebar. One shared tone helper
now serves the header segment and the expanded row, so they cannot diverge.

Monitoring is a state the app already colours; the other four kinds are plain
markers and stay neutral. A running turn still dims the whole set.

* fix(native-chat): draw the strip header separator in a visible tone

The separator used `text-border`, a divider-line token that is 7% white in
dark mode - an order of magnitude fainter than the counts on either side, so
the dot between them read as absent. main.css already records that token as
too faint for a visible mark.

* fix(native-chat): give the worktree-ps journal stub a cursor

The status feed's projection cache calls journal.cursor(); this stub is cast
through unknown, so the missing method only surfaced at runtime. Its journal
never changes, so a real one would hold the cursor steady.

* refactor(native-chat): split the sidebar subagent rows out of this PR

The strip stands alone: the sidebar mapping, its observation plumbing and the
AgentStatusEntry.subagents wiring move to a stacked follow-up. No wire field
here is sidebar-only - the strip's rows read name, state, elapsed and tokens.

* perf(native-chat): keep task usage out of the session status summary

A `task_progress` frame ticks a background task's `totalTokens`, which
failed the status feed's equality check and re-broadcast a full summary to
every `agentSession.subscribeStatus` subscriber — paired-web and SSH/relay
clients included — for a number no session list renders. The projection now
drops usage; tokens keep flowing on the background-task channel the strip
reads.

* fix(native-chat): correct token unit rounding and drop the unused dot state

`formatBackgroundTaskTokens` rounded before choosing the unit, so 999_950
rendered as "1000k" instead of "1m"; pick the unit from the rounded value.

`backgroundTasksDotState` has no caller on this branch or the stacked
sidebar PR, and its multi-kind branch would report 'monitoring' over an
attention state. Delete it rather than leave it to be wired up.

* fix(i18n): drop the orphaned backgroundTasks.runningList key

The strip rewrite removed its only call site, and an unreferenced key gets
promoted into the eagerly parsed boot catalog. Delete it from en.json and
regenerate en-runtime-required.json.

* fix(native-chat): show the reason on every attention row

The row guarded the reason line on 'waiting', so an 'unverifiable' child
("no contact") and a 'blocked' one ("failed") rendered bare while the
collapsed header named exactly those reasons. `backgroundTaskStateReason`
already returns null for the non-attention states, so the guard was only
lossy — the SSH boundary requires the unverifiable verdict stay legible.

Also keys the header segments off their kind discriminant instead of the
translated display text.

* fix(native-chat): make the strip header agree with its own count

The headline counts live AND settled rows, but the state breakdown omitted
'done', so one working agent beside four settled ones read "5 agents — 1
working": the count said five, the breakdown accounted for one. Done now
appears in the muted detail (never as an emphasised segment) so the two
agree.

The single-command header also drew an elapsed clock on a settled task,
which the row already refuses as a lie about finished work.

* perf(native-chat): memoize the background-task roster grouping

The 1 Hz elapsed tick re-rendered the strip, and the render body regrouped,
re-sorted and re-translated every task each time only `now` had changed.
The header still derives from `now` on purpose.

* test(native-chat): cover settled rows and the mid-turn mounted strip

Neither headline behaviour had component coverage: every strip render passed
`settledTasks={[]}`, and the `showBackgroundTasks` seam was never set true,
so the strip staying mounted through a running turn was exercised nowhere.

Adds a settled-beside-live row test (final usage kept, no clock, no stop) and
a mid-turn mount test (strip present, turn owns the voice). The background-task
tests share one session-element helper so the file stays under its line cap.

* refactor(claude): keep MAX_TASK_ID_LENGTH module-private

Nothing outside claude-background-task-frames.ts references it; the export
was residue from this PR's split.

* test(native-chat): give the mid-turn strip test a real turn

main now gates the composer's stop button on a provider-minted turnId rather
than the send-time working signal, so a test claiming a running turn has to
supply one. The controller mock hardcoded turnId null.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-09 23:47:47 -07:00
Neil 0b60b0dcb1 perf(native-chat): bound retained items on the structured session path (#19841)
`mergeSubmissions` caps submissions at 256, but `mergeItems` had no
equivalent bound, so `state.items` grew for the whole life of a long
structured session while the live path caps itself to its read window.

Head-trim `items` to a retained-item limit when a live batch merges, and
set `hasOlder` so anything trimmed is still reachable by paging. Paging
older raises the limit to what the page produced, so a live batch slides
the widened window instead of collapsing it back to the cap -- the same
shape as the live path's growing `limitRef`.
2026-09-09 22:12:52 -07:00
Brennan BensonandMerge Sim 5868fdc9e3 feat(native-chat): report Codex background tasks in the chat strip (#19346)
* feat(native-chat): report Codex background tasks in the chat strip

The background-tasks strip works for Claude only; a structured Codex
session shows nothing in it. Feed it from the Codex app-server stream.

The strip stands for work that OUTLIVED a turn, which is what the
monitoring header, Claude's foreground suppression, and the conversation
command gate all already assume. Codex has no `is_backgrounded` flag, so
that fact is derived from the turn boundary: a `subAgentActivity` child or
a primary-thread `commandExecution` becomes visible once the turn it
belongs to completes and it is still unsettled.

`turn/completed` only reveals a task here, never settles one — measured on
`codex app-server` 0.153.4, a spawn_agent child reported `completed` 95.8s
after its parent turn ended. Only a child's own activity kind settles it.

Codex exposes no honest stop: `turn/interrupt` on a child ends its turn
without emitting a terminal activity item and leaves its shell running. So
the state carries a new optional `supportsStopAll: false`, the strip hides
a control that could not act, and the blocked-command message asks the user
to wait rather than to press a button that does not exist.

* refactor(codex): move session teardown out of the structured adapter

Merging main crossed the 300-line cap on
`codex-structured-session-adapter.ts`: the rewind backend (#19235) and this
branch's close-time strip clear both landed in it. The four close paths move
verbatim into `codex-structured-session-teardown.ts`, where they funnel
through one `settled` helper instead of repeating the notification-retry and
background-task cleanup at each call site. No ratchet bump.

Also normalize a background task's description once at receipt rather than on
every projection; the roster is re-projected on each observed frame.

* fix(codex): drop the shell row the journal already settles

A `commandExecution` still `inProgress` when its turn ends was reported as a
`command` task. But `settleCodexJournalTurn` writes exactly those items to the
journal as `state: 'failed'` on `turn/completed` and forgets them, so the strip
row would have claimed a shell was still running at the same instant Orca
recorded that it was not — two surfaces contradicting each other about the same
process.

A subagent is the opposite case and stays: the roster pointedly does not sweep
at a turn boundary, because children measurably outlive it. That leaves the
producer making exactly one claim — these spawn_agent children are still live
after their turn — which the durable roster row corroborates.

* fix(native-chat): track Codex background execution lifetimes

* fix(native-chat): keep running tool groups from claiming completion

* Fix runtime catalog and capability expectation

* fix(codex): keep a child's name on the command row that outlives it

A child agent's commands stay hidden behind its agent row while the child
works. Once the child's turn settles with a command still running, that
command surfaces as its own row labelled from the raw command string, so
'long_probe' became "/bin/zsh -lc 'ping -c 300 127.0.0.1 > /dev/null'"
at the moment that row was the only remaining signal for the work.

Qualify a child's command row with the child's label. Resolved on read,
so a label registered after the command still lands, and bounded by the
existing description cap so admission accounting stays valid. Primary-
thread commands are left unqualified: they have no child to name.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-09 00:13:20 -07:00
Brennan BensonandMerge Sim bf4e270504 fix(native-chat): list the slash commands and skills a structured Claude session actually loaded (#19127)
* fix(native-chat): list the slash commands and skills a structured Claude session actually loaded

The chat composer's `/` menu was built from a curated five-command catalog plus a
host disk scan of skill roots. Neither is what the running session can do: the
session reports its own `/` surface, which carries this repo's `.claude/commands`,
the skills that only reach it through plugin roots, and a hide-list of commands
that mean nothing outside a terminal UI. On one local session the menu offered 6
commands and 17 skills where the session reported 62 commands and 33 skills.

Read that surface per session and let it drive the picker:

- A per-session catalog seeded from the frame that proves the session and kept
  current by every later report, exposed over a new `agentSession.commands` read.
- The report is the authority on WHICH skills exist; the disk scan stays the
  source of scope and description for the names both know about, so a skill the
  session never loaded is no longer offered and one it loaded from a root the
  scan cannot see now is.
- A host that predates the read answers `method_not_found` and the composer keeps
  its curated catalog, so mixed versions and the PTY lane are unchanged.

* test: register agentSession.commands on the three surface ratchets

The structured method count, the mobile allowlist, and the cross-version call
table each enumerate the agentSession surface on purpose, so an additive method
has to be declared in all three rather than counted around.

* fix: preserve session catalog authority and publish live updates

* fix(native-chat): publish authoritative command catalogs on session updates

* fix: seed Claude slash catalog before the first prompt

* test: verify unclassified catalogs survive session publication

* test: complete structured rename journal fixtures

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 20:33:14 -07:00
Brennan BensonandMerge Sim f7d5216016 Show provider activity in chat turn tails (#19055)
* feat(chat): show turn-scoped activity tail

* fix(chat): keep turn activity broad

* feat(chat): surface provider activity in turn tail

* fix(chat): keep reasoning headline as activity and widen redaction

A Codex reasoning summary streams as a bold headline followed by body text.
Folding the whole summary into the tail leaked literal ** markers and body
prose; only the first non-empty line is activity copy, and an unterminated
bold header mid-stream is unwrapped too.

Redaction used a hyphen for GitHub token prefixes (they use an underscore),
and missed fine-grained GitHub tokens, AWS access key ids, JWTs, URL
userinfo passwords, and bare token= values.

* fix(chat): wait for a complete reasoning headline

A bold headline still streaming has no closing marker yet; holding the
previous activity copy until it lands avoids flashing a half word.

* refactor(chat): drop bespoke secret redaction from activity copy

Reference agent hosts render provider-derived status text unredacted;
this table was the only one of its kind and its GitHub pattern matched
no real token. Bounding and the reasoning-headline extraction stay.

* Bound provider headline updates and clear activity on reconnect

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 16:42:18 -07:00
Brennan BensonandMerge Sim f8780a2c86 feat(native-chat): stop monitored tasks individually (#18807)
* feat(native-chat): stop monitored tasks individually

* test: expect Claude task stop capability

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-05 15:35:39 -07:00
Brennan BensonandMerge Sim e89deb63c9 Show Claude background task status in Native Chat (#18757)
* feat(native-chat): show Claude background task status

* fix(native-chat): carry background task fence forward

* fix(claude): bound background task stop requests

* Show running Claude background task details

* Harden Claude background task status updates

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-05 00:21:42 -07:00
Brennan BensonandMerge Sim 98e77ef1a7 feat(mobile): structured native Codex chat (#18074)
* feat(mobile): finalize structured native Codex chat

* fix(mobile): close structured chat lifecycle gaps

* wip(mobile): fence stale structured inventory and bound operation-id retention

Fence local structured-session inventory and subscription responses with a
sync generation so a toggle-off clear, reconnect restore, or retry cannot
apply a mirror from a superseded instance. Bound mobile ambiguous
operation-ID retention at 128 with unmount cleanup.

Staged on the reconcile branch only: the sync module is now 312 lines and
needs a real split before this can reach the PR head.

* fix(ci): split the structured session-tabs sync and give static analysis mobile types

The local structured session-tabs sync module outgrew the 300-line cap once it
took on generation fencing, so split it along its real seams instead of raising
the cap: the generation/cursor fence, snapshot projection, snapshot apply,
inventory refresh, and the subscription loop. The original path stays as a
barrel so no importer moves.

Repoint the host-session-mirror settle census at the apply module, which owns
two receipts now — the snapshot it mirrors in, and the toggle-off teardown that
retracts what it published. The teardown receipt is named rather than anonymous
so the pin says which direction it settles.

The changed-code quality gate lints mobile files and resolves their types from
mobile/node_modules, but mobile is a separate pnpm project that the root install
never populates, so every mobile type degraded to an `error` type and the gate
reported phantom findings. Install mobile dependencies in static analysis when
the diff touches mobile, gated on a new classifier output.

* fix(mobile): let a slow capability handshake still reach connected

The mobile capability update is an advisory whose result is discarded, yet an
unanswered one was fatal while an explicit rejection was tolerated. A 5s timeout
on the direct client force-closed the socket, and on the relay path it failed
`confirmResume` before `connected` was ever published, so a consistently slow
link redialled forever. Both paths now share one helper that settles every
ambiguous outcome (timeout, mid-flight drop) like a rejection and rejects only
when the frame never reached the wire — the one case nothing else recovers from,
since the socket's own desync force-close is gated on already being connected.
The generation guard still keeps a replaced session from connecting.

Retained structured-session operation ids were capped at 128 with oldest-first
eviction, but every retained id belongs to a send whose outcome is unknown, so
eviction turned a user's retry into a second message on the host. Bound the map
by expiry against the id's own embedded timestamp instead, mirroring the host's
operation ledger, so no id is released while the host would still honour it.

Also give the mobile CI install the root install's lockfile drift guard (mobile's
lockfile carries patchedDependencies a silent rewrite would drop), gate
mobile_dependencies on should_run, and key the pnpm store cache on both lockfiles.

* refactor(mobile): extract the relay pending-request registry

The merge composed two independently-sized changes — this branch's capability
handshake settle and main's dial-stage tracking — pushing the relay session file
to 304 lines against a 300 cap. Neither side broke it alone.

Move the in-flight request registry (id generation, tracking, settlement, and
reject-all with its delivery-ambiguity marking) into RelayPendingRequests,
matching the existing collaborator pattern alongside RelayDialStageTracker and
RpcSessionLivenessWatchdog. No behavior change.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 15:19:26 -07:00
Brennan Benson fd9125ea8c feat(native-chat): Codex structured native chat restructure (#16729)
* feat(native-chat): port structured Codex sessions from restructure-recovery

Rebuilds the desktop structured native-chat implementation from
brennanb2025/native-chat-restructure-recovery (tip 4e31c08db3) on top of
current main as a single commit, scoped to the local Codex path.

Ported:
- Structured agent-session core: durable record store + single-writer lease,
  canonical journal, agent-session wire host/attach/eviction/subscribers,
  `agentSession.*` RPC surface (registered via ALL_RPC_METHODS; host-side
  mobile allowlist included for wire compat), pty write gate, transcript
  additions, and the Codex app-server adapter/launch resolution.
- Renderer: NativeChatStructuredSession view/composer stack, structured
  launch path with the single-flight guard, local structured session tabs
  sync, activation gate + structured inventory (read-only
  `agentSession.handoffStatus` probe), agent-session tabs in the tab strip,
  AI-vault structured session activation, and the settings pane with the
  parent Experimental Chat UI toggle plus the nested "Use updated structured
  native chat" toggle. New sessions require both flags, agent codex, no
  prompt, and a local non-WSL, non-Windows-host execution host
  (structured-native-chat-availability).
- Fixes 72c013cea6 (verified Codex launch recovery), 8ddbaf5e3d (defer
  native terminal view switching affordances), and 4e31c08db3 (release the
  launch gate after a visibility retry) with their regression tests,
  including the third-launch-after-retry guard case.
- Cross-version agent-session wire test + CI lane, packaging entries
  (proper-lockfile, agent-tooling asar excludes), and the wire-compat doc
  section.

Deliberately not ported: mobile/ changes, the Claude structured runtime
(only the claude-transcript-branch-proof and claude-structured-owner-identity
leaf modules remain, backing the kept TUI-recovery arms), the terminal↔chat
adoption/handoff flow (`agentSession.adoptTerminal`/`requestHandoff`, the
handoff request engine, TUI adoption machinery, orca-runtime adoption
methods), renderer switching affordances and their dead leftovers, the
hook/subagent-status refactor cluster, and unrelated branch changes. The
crash-during-acquisition recovery path (restart handoff adjudication,
restore/reverse re-acquire, lease schema handoff keys) is kept because every
plain direct launch depends on it; a trimmed handoff coordinator exposes
only status/restore/close.

Branch edits that targeted files main has since split (ipc/pty.ts,
worktrees.ts, rpc/methods/terminal.ts, useIpcEvents, pty-connection,
store/slices/terminals.ts, runtime-types, web preload) were re-applied to
the split modules, preserving main's newer logic (Windows CIM fallback,
browser tab close rework, cold-restore resume flow, dispatcher threading).

Known seam: the mobile clipboard image-provenance CONSUMER gate ships
(agentSession.send refuses unproven mobile image refs with
agent_session_image_untrusted) but the producer hunk in
rpc/methods/clipboard.ts stays with the unported mobile cluster, so mobile
image sends into structured chat fail closed until that side ports.

* fix(native-chat): trust only authenticated local image uploads

* fix(build): preserve Windows process-tree patch application

* test(windows): include process creation time in addon fixture

* fix(build): run windows-process-tree node-gyp from the physical package dir

gyp expands the node-addon-api dependency by probing node, whose cwd
resolves to the package's physical directory in the store, so the emitted
target is a store-relative ../../../../node-addon-api@... hop. gyp then
resolves that hop against the rebuild cwd; from the node_modules
symlink/junction it escapes the store and configure fails with
"node_addon_api.gyp not found" (run 32999886072).

Rebuild from realpath(package dir) so both bases agree, matching how the
package manager itself runs native install scripts. The regression test
replays gyp's expansion+resolution against the planned cwd and fails
without the fix.

* fix(native-chat): keep chat tabs visible through terminal closes and empty-worktree launches

Two proven blockers in the native Codex tab contract:

closeTerminalTab pre-empted the canonical unified close. With one terminal
left it deactivated the worktree on a terminal/editor/browser-only check,
blanking a workspace that still held a renderable agent-session tab; with
two or more it pre-picked a successor from terminal entities only,
re-stamping the group active before closeUnifiedTab's MRU/neighbor repair
could land on the chat tab. Successor choice now defers to the unified
contract whenever the terminal has a unified row, and deactivation is
gated on the unified renderable count (matching leaveWorktreeIfEmpty),
with the legacy pre-pick kept only for terminals without a unified row.

A structured session created on an empty worktree was published into the
host's headless group while preserveLocalLayout froze the local layout,
leaving the tab in store but permanently off screen. A preserveLocalLayout
owner now always takes client-owned placement — repairing a rendered
leaf whose group record is missing, or materializing a rendered group on a
truly empty worktree — and applies the client-derived layout repair while
still rejecting host-authored layout.

Regression tests drive the real store through closeTerminalTab (git
worktree and folder workspace) and the real snapshot applier for the
empty-worktree adoption states; all fail without the fixes.

* fix(native-chat): close stale turns and retry rejected sends

* fix(native-chat): retire hosted rows on structured tab activation

* fix(native-chat): preserve rpc defaults across main merge

* chore: format remote wire compatibility guide

* test(native-chat): cover retry after unconfirmed send

* fix(native-chat): reload outbox on session switch

* docs(settings): disclose structured chat platform limits

* fix(native-chat): await Codex launch-home preparation

* fix(codex): align child-process allowlist with async trust bridge

* test(identity): update inventory for tab surface refactor

* fix(windows): preserve process-tree CRLF patch sources

* fix(native-chat): anchor an unmatched chat echo where it was sent (#16117)

* fix(native-chat): anchor an unmatched chat echo where it was sent

The reported symptom was old user messages replaying below every new turn, so the
conversation read as scrambled. The cause was not that the echo failed to match a
transcript row. Claude consumes a mid-turn send through a `queued_command`
attachment and writes no `type:"user"` record for it, so some echoes can never
match, and no amount of matching will change that. The cause was WHERE an
unmatched echo rendered: buildMobileNativeChatTransientData appended every pending
item after the entire transcript, so it re-read below each turn that landed
afterwards.

Render each echo directly after the transcript row it was sent against, using the
baseline the send already captures. An unmatched echo is then at worst a duplicate
in the right position rather than a scrambled one, and it stays visible. Echoes
sharing an anchor keep send order; a send with no baseline, or one whose anchor
folding dropped, still falls back to the tail.

Deliberately NOT fixed by deleting the echo. Inferring from send ordering that an
echo can never match, then removing it, loses the user's own text for a message
the agent did receive, and it cannot fire in the common case anyway - measured
drain groups are 1,017 of size 1 against 55 larger. It also escalates an existing
gap: the count pass has no baseline-tail guard, unlike the glue pass, while
`messages` is a 40-row window that head-trims, resets on reconnect and grows at
the front on loadEarlier, so a false landing there would license deleting a
DIFFERENT outstanding message.

That count-pass gap is real and left for a separate change; anchoring makes its
worst case a duplicate in place rather than a scrambled conversation.

* fix(native-chat): preserve folded echo anchors

* fix(native-chat): preserve forward-folded echo anchors

* fix(native-chat): keep leading folded echoes in place

* fix(workspace-cleanup): show git status for every row (#16690)

* fix(native-chat): refuse structured chat on every Windows execution path

canUseStructuredNativeChat only refused win32 when a project runtime
resolved, so folder-workspace keys (and other keys with no project
runtime) failed open into structured chat on Windows. Fail closed on
win32 unconditionally after the host check, matching the settings copy:
local macOS/Linux only; Windows/WSL/SSH stay on terminal chat.

* fix(native-chat): restore runtime refusals behind the win32 gate

506d375de3 replaced the project-runtime checks with a bare platform test,
so a WSL or repair-required runtime resolution would no longer refuse
structured chat off-win32. Keep the unconditional win32 refusal and
re-run the runtime resolution after it, so the gate does not depend on
the resolver's own platform guard. Tests inject WSL and repair-required
resolutions on darwin/linux and fail against the regressed gate.

* fix structured session journal durability

* fix structured tab active pointer after restart

* fix(native-chat): await optional lease renewal callbacks

* refactor(skills): extract install error messages

* fix(agent-session): harden recovery ownership

* fix(native-chat): retain panes across tab activation

* fix(native-chat): address round-one review findings

* test(native-chat): align integration coverage after main merge

* fix(native-chat): harden round-two reliability

* fix(native-chat): harden round-three reliability

* fix(native-chat): close round-four recovery gaps

* fix(native-chat): separate bounded journal key forms

* fix(native-chat): reset outbox error in render on session switch

The switch effect adjusted error state after the sessionId prop changed,
tripping react-doctor's no-adjust-state-on-prop-change on the changed-code
gate and flashing the old session's banner for a frame. Reset it with the
render-time previous-value guard instead.

* fix(native-chat): invalidate stale outbox settlements

* test(native-chat): restore settled-error session-switch regression

a6e2379bd1 replaced this test with the in-flight settlement race test,
leaving the render-time error reset unpinned: deleting the reset block
still passed the whole native-chat suite. Keep both scenarios pinned;
they are distinct (settled error clears on switch vs stale settlement
invalidated in the commit-to-passive window).

* test(wire): make release checkouts race safe

* test(wire): pin cross-process checkout single-flight and importer specifier contract

* test(wire): harden release checkout lifecycle

* fix(build): drop CR-byte residue from windows-process-tree patch

The two trailing CR bytes on the patch's deletion lines are a proven
no-op: pnpm hashes patches CRLF-normalized (both forms hash to the
lockfile's 946ffb2b) and materializes this package without applying the
patch in either form, so the load-bearing build edits come solely from
applyWindowsProcessTreeBuildFixes() (#16947), which handles both source
EOL forms. Restore byte-identity with main and repin the contract test
to the post-#16947 reality: LF-only patch bytes plus lockfile hash sync.

* fix(native-chat): skip empty startup recovery
2026-08-28 16:45:58 -07:00