Commit Graph
10646 Commits
Author SHA1 Message Date
Merge Sim 5acb404bf2 Preserve unverifiable timing across older host upgrade 2026-09-09 23:17:50 -07:00
Merge Sim dd196ac75a Respect authoritative unknown native chat duration 2026-09-09 23:11:19 -07:00
Merge Sim 1677a5dba4 Merge origin/main into durable native chat turns 2026-09-09 23:04:46 -07:00
Merge Sim 98cfde633b Correct turn duration gate assertion reference 2026-09-09 22:13:28 -07:00
Neil 0b60b0dcb1 perf(native-chat): bound retained items on the structured session path (#19841)
`mergeSubmissions` caps submissions at 256, but `mergeItems` had no
equivalent bound, so `state.items` grew for the whole life of a long
structured session while the live path caps itself to its read window.

Head-trim `items` to a retained-item limit when a live batch merges, and
set `hasOlder` so anything trimmed is still reachable by paging. Paging
older raises the limit to what the page produced, so a live batch slides
the widened window instead of collapsing it back to the cap -- the same
shape as the live path's growing `limitRef`.
2026-09-09 22:12:52 -07:00
Merge Sim e2b3abe9e4 Keep earlier turns through a Codex rewind and count a mid-turn attach from the real start
Findings from an independent adversarial review of the typed turn record:

- A Codex rewind adopted the provider's item list as the new epoch, and the
  provider never returns the host's own turn rows, so every duration before
  the rewind point vanished. The host's turn rows are now spliced back beside
  the item each followed, and recovery no longer expects the provider to
  prove rows it never owned.
- The epoch row was stamped with the current schema version, so an older host
  latched read-only at row 1 of every new session, defeating the mixed
  version design. It carries no body and stays at v2; a stored-row test now
  reads SQLite directly, because the reader upcasts every row on read.
- A send Codex folds into a running turn shares the opening prompt's provider
  key, and the alias map credited the duration to the later prompt. The
  earliest submission naming a key now wins.
- The live counter anchored on first sight, so a client attaching mid-turn
  counted from zero. Published frames now carry the host's clock, the reducer
  keeps the last sample with its local receipt time, and both clients anchor
  on how long the host says the turn has run.
2026-09-09 22:09:25 -07:00
Neil a067cccd38 perf(terminal): resume OSC terminator search past the carried frame (#19839)
A single unterminated OSC 9999 marker split across many PTY chunks
re-scanned the whole accumulation for a terminator on every chunk, so
work grew with the square of the frame length.

Carry how much of `pending` already failed the search and resume one
character before it, which is enough for an `ESC \\` straddling the
chunk boundary.
2026-09-09 22:06:31 -07:00
Brennan BensonandMerge Sim 4b4acf26a4 fix(mobile): enable patch-free iOS text selection in native chat (#19769)
* fix(mobile): make every native-chat text node selectable

Long-press selection worked on some chat text and not others. Markdown
paragraphs — the default block for agent prose — were the one block type
left out when headings, quotes, code, lists and table cells gained
`selectable`, and tool result output, diff rows, the unloadable-image
placeholder, permission/question bodies and the send-error banner never
had it at all.

Selection is now set on every content Text in the chat surface, on the
outermost block Text so nested inline spans inherit it. Labels inside a
Pressable (option rows, tool-line headers, buttons) are deliberately left
alone: selection there would swallow the tap they exist for.

Extracting MobileNativeChatEmptyState keeps the view under its max-lines
cap and matches desktop, where NativeChatEmptyState is already its own
component.

Tests render each surface and assert selection on the block that carries
the prose; both files were ablated against the unfixed source (4/10 and
3/5 red) so they pin the defect rather than the current behavior.

* fix(mobile): support native text range selection on iOS

* fix(mobile): remove persistent assistant message controls

* fix(mobile): scope patch-free text selection to chat

Use the stock react-native-uitextview dependency behind an iOS adapter and opt assistant Markdown into range selection only in native chat. Preserve the existing React Native Text behavior elsewhere and remove the persistent assistant controls.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-09 21:50:31 -07:00
Merge Sim 4ae26deb11 Register completed turn duration reliability gate 2026-09-09 21:47:39 -07:00
Merge Sim 45036bb52c Preserve Codex exit receipt across close retries 2026-09-09 21:46:06 -07:00
Brennan BensonandMerge Sim 2f828e4462 fix(native-chat): show Claude working from the send, not the provider echo (#19822)
* fix(native-chat): show Claude working from the send, not the provider echo

A structured session read as working only once a turnLifecycle row existed.
Codex writes that row ~150ms after the send; Claude cannot write it until the
SDK echoes the user message back, measured at a 3.4s median and 18s at p90, so
the chat and every session list read idle for the whole wait.

The journalled submission is the host's own evidence a turn is owed, so the
shared projection reads it too. `unknown` still counts -- the ack budget
elapsing answers delivery, not whether work is owed -- while a recovered
`unknown` does not, which needed the existing row flag carried onto the
projected submission.

Claude's activity line now stays the generic fallback. Its only turn-wide frame
carries a bare token, and its task_* prose describes a spawned task rather than
this turn; compaction is kept because it explains an otherwise silent wait.

* Fix structured chat pending-work lifecycle and mobile cancellation

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-09 21:38:35 -07:00
Merge Sim c6caab5fb7 Retain turn attribution for loaded chat history 2026-09-09 21:29:30 -07:00
Merge Sim 1274bdac40 Preserve observed turn end across settlement retries 2026-09-09 21:24:24 -07:00
Merge Sim 88f407d0f1 Merge remote-tracking branch 'origin/main' into brennanb2025/native-chat-turn-lifecycle-durable 2026-09-09 21:08:11 -07:00
Merge Sim 1d90b77d6e Record a turn as a first-class journal item
The turn record is now its own item kind rather than a status row carrying a
lifecycle field: no text to misuse, and the fold matches the durable turn
record other systems keep. Rows that carry it are stamped journal schema v3;
every other row stays v2, so an older host keeps reading them and latches
read-only at the first v3 row instead of truncating the epoch.

Clients that predate the item would paint an unknown kind as a text bubble,
so the host publishes the legacy status form to any client that does not
advertise agent-session.turn-item.v1, through the same per-client seam
background tasks use. The downgrade is transitional and goes once no
supported release lacks the capability. The shared projection now renders
unknown item kinds as nothing, so later kinds need no gate. One shared reader
handles both forms for old journals and old hosts.
2026-09-09 19:44:18 -07:00
Jinwoo Hong aac38d698f fix(push): isolate deployment and validate candidates before activation (#19771)
* fix(push): isolate deployment and validate candidates before activation

* test(push): classify dedicated rollout outside shared SQL lock census

* test(push): verify independent deployment identity and lock
2026-09-09 16:22:17 -04:00
Merge Sim a133af6b7b Key lifecycle rows to their user item and record the provider's measured duration
A lifecycle row now names the user item that opened the turn by its provider
key, so clients attribute timing explicitly and fall back to journal order
only for rows from older hosts. A provider-initiated turn with no prompt can
no longer claim the previous prompt's duration.

When the provider measures the turn itself (Codex turn.durationMs, Claude
result.duration_ms) the terminal row records it and clients prefer it over the
host interval, so a turn shows the same number live and after a history
restore. Host receipt times remain the live-counter anchor and the fallback.
2026-09-09 12:30:46 -07:00
Merge Sim a898e0a7d3 test: deduplicate turn lifecycle suites
Each behavior keeps one test; duplicated harnesses and restated cases go.
2026-09-09 12:14:56 -07:00
Jinwoo Hong 7dd183d82d fix(native-chat): keep worktree active during chat creation (#19753)
* fix(native-chat): activate worktree before chat session

* test: update native chat activation census
2026-09-09 13:47:59 -04:00
Jinjing 73d0521410 Replace the sidebar create dropdown with two direct action buttons (#19653) 2026-09-09 10:42:18 -07:00
github-actions[bot] b94bcdc632 Update README downloads badge 2026-09-09 12:37:29 +00:00
ed9d76178d perf: check tunnel queue capacity before copying frame payloads (#19497)
* perf: check tunnel queue capacity before copying frame payloads

* fix(browser-tunnel): derive writer admission size from the frame encoder

Shares one encoded-length helper so the pre-encode capacity check cannot drift
from what encoding allocates, and covers the exact byte-cap boundary.

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-09 03:59:50 -07:00
Neil 946b1fc078 perf(terminal): own retained tail rows so they stop pinning their chunks (#19528)
Every row in the 2000-row retained tail is sliced from the PTY chunk it
arrived in, so a spinner workload — CR-redraw frames, exactly what Claude
Code and Codex emit — makes a ~5 KB tail pin one whole chunk per row.
Measured on 400 x 64 Ki-char chunks: a 5,090-char tail retained 37.5 MB.
At 16 Ki x 1600 it was 46.9 MB.

Own the two values the plain tail path actually retains, newlyCompletedLines
and the partial line, at the point they enter the tail: 37.5 MB -> 0.18 MB
and 46.9 MB -> 0.09 MB, with per-chunk cost unchanged within noise. The
transcript keeps the same row objects, so it is covered transitively.

The multiline redraw builder never slices normalizedChunk — it writes rows
character by character — so only its partial line, which is re-sliced from
its own row on every frame, needs owning. Owning its rows as well measured
no retention gain and cost CPU on fullscreen TUI floods.

Also own the wait-blocked keywordCarry: 31 characters that pinned a full
lowercased copy of the chunk, per PTY.

Follow-up to #19396, which introduced ownRetainedString for the much smaller
pending-control fragments.
2026-09-09 03:50:45 -07:00
Neil 373a670aba perf(terminal): skip ordinary output in preview control scans (#19400)
* perf(terminal): skip ordinary output in preview control scans

* test(terminal): cover scan resumption after parsed controls
2026-09-09 03:49:14 -07:00
Neil 1e2ed20b35 perf(terminal): release oversized backing strings behind pending controls (#19396)
* perf(terminal): release oversized backing strings behind pending controls

* test(terminal): record reproducible pending-storage gate evidence

* perf(terminal): own retained control fragments with a fast copy primitive

The pending-control ownership landed with a charCodeAt block copier (10 us
at 4 Ki, 170 us at 64 Ki), so it needed a "copy only when discarded output
dominates the tail" gate to stay affordable. That gate was the whole cost
problem: on adversarial streams it fires every chunk and pays the slow copy
(+45% on 16 Ki ANSI chunks, +81..111% on 194 Ki status chunks), and it also
skipped ownership on fragments too small to be sliced strings anyway.

ownRetainedString replaces it with a Buffer utf16le round trip (0.57 us at
4 Ki, 21.9 us at 64 Ki) and returns anything below V8's SlicedString
kMinLength unchanged. Buffer is absent in the renderer and on mobile, so the
copier is resolved once behind a lone-surrogate round-trip self-check and
falls back to the block copier. With a ~1 us copy the gate is unnecessary:
ownership is now unconditional at all three retention sites and the
adversarial cases land within noise of the un-owned parsers.

The three forced-GC threshold fixtures are replaced by one forced-GC test
for the primitive plus deterministic spy assertions that each site routes
its retained value through ownRetainedString. All fidelity and differential
coverage is kept.

* fix(terminal): escape the NUL in the round-trip probe

A raw NUL byte in the source made git treat the file as binary, so its
diffs and blame were unreadable. Escapes are equivalent at runtime.
2026-09-09 03:40:03 -07:00
Neil 042cc5266c perf(terminal): skip plain text between partial escape sequences (#19393)
* perf(terminal): skip plain text between partial escape sequences

* perf(terminal): take the ESC at hand before searching for one

The unconditional ground-state indexOf regressed dense back-to-back
SGR/CSI streams, where the code unit at the cursor is already the ESC and
the search pays call plus SIMD setup to find it in place. Check the
current unit first and fall back to the native search otherwise.

0.9 MiB dense SGR/CSI medians: 2.04 ms before this PR, 2.56 ms with the
unconditional search, 1.92 ms with the hybrid. Sparse colored logs and
plain text keep the full search win (0.36 / 0.016 ms vs 1.54 / 1.41 ms
baseline). Differential over the full VT alphabet with lone and split
surrogates matched 388,416 cases against both prior implementations with
zero mismatches; the ground-scan work budget now records 16 inspected
code units and 2 native searches.
2026-09-09 03:31:17 -07:00
91f00b32ef perf: stop merging OS-opened files at the pending queue cap (#19508)
* perf: stop merging OS-opened files at the pending queue cap

* fix(os-open): report markdown opens dropped at the pending queue cap

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-09 03:29:51 -07:00
Neil 8c18e42be2 test(ci): replace fixed teardown waits with bounded polls (#19720)
The two flakiest tests on main both guessed at a duration instead of
waiting for the condition.

- windows-pty-job.win32.test.ts assumed job teardown finished in 1.5s;
  under load on a Windows runner it does not. Poll isAlive up to 30s
  instead -- the assertion is unchanged, so a real leak still fails.
- structured-agent-session-claude-options-round-trip.test.ts relied on
  vi.waitFor's 1s default for a two-hop handoff; give it 10s.

Both are test-only and strictly widen an existing wait.
2026-09-09 03:28:29 -07:00
069bdae283 perf: materialize only the requested recent plugin audit lines (#19498)
* perf: materialize only the requested recent plugin audit lines

* test(plugins): prove the recent audit window matches the full-split selection

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-09 03:23:55 -07:00
f1f2d61d0e perf: resolve current-workspace document addresses before catalog scans (#19495)
* perf: resolve current-workspace document addresses before catalog scans

* test(browser): lock current-workspace precedence in doc address resolution

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-09 03:23:47 -07:00
bf1e1b9004 perf: probe requested pane keys instead of enumerating records (#19494)
* perf: probe requested pane keys instead of enumerating records

* perf(agent-status): drop the requested-key array from pane removal

Probing the pane keys still beat enumerating the record, but materializing
the requested set allocated on every call including the common no-match
path, where a dozen records are swept per retirement. Copy lazily on first
match instead, and cover the set-disagreement and prototype-key cases.

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-09 03:23:38 -07:00
Neil 23f38ddf7f perf(workspaces): skip unchanged heartbeat subscriber work (#19392)
* perf(workspaces): reuse unchanged heartbeat status projections

* perf(workspaces): skip unchanged title-sync session collections

* test(workspaces): cover constant-size key swaps in status projection reuse

Also point the title-input gate at a tracked source file; the previous
motivatingLink referenced an untracked .agents skill path.
2026-09-09 03:19:55 -07:00
Neil db66ab4ac8 perf(runtime): skip hibernation inventories without completed agents (#19391)
* perf(runtime): skip hibernation inventories without completed agents

* perf(runtime): skip hibernation status scan when no runtime owners

* fix(runtime): require host evidence for workspaces resolved mid-inventory

`runtimeLivenessRequiredWorktreeIds` was sampled before the runtime inventory
await, while the plan is built from the state after it. A workspace that gained
tabs or resolved its runtime owner during that window was therefore absent from
the required set, so the planner did not demand fresh host evidence for it and
fell back to client PTYs — client bookkeeping answering for the execution host.

Union the post-await targets into the required set inside `snapshotFromState`.
Union rather than replace: the set only ever grows, so the planner can only skip
more workspaces, never authorize a hibernation it would previously have refused.
An absent inventory stays a skip; nothing reads it as an exited PTY.

Extract the coordinator test fixtures so the regression lives in its own file
without pushing the coordinator suite past the 800-line test cap.

* fix(types): annotate hibernation fixture mock exports for declaration emit

TS2883: the inferred `Mock<Procedure>` types of the fixture's exported
`vi.fn()` bindings reference `Procedure` from a transitive `@vitest/spy`
path that cannot be named.

* chore: keep local-file-sink-memory test formatting as on main

The merge commit's pre-commit hook reformatted a file this branch does not own.
2026-09-09 03:17:20 -07:00
OrcaWinandm4air c875941d7c perf: stop queued metadata work after watcher cancellation (#19446)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-09 03:17:15 -07:00
Merge Sim 6b4b547a1b Name settled lifecycle rows by their terminal state
An interrupted or unverifiable turn must not read as completed for any
consumer that renders status text raw. One shared helper builds the text for
both providers from the lifecycle state.
2026-09-09 01:00:15 -07:00
Merge Sim 5060f54584 Merge origin/main into native-chat-turn-lifecycle-durable
Routes both Codex turn notifications through the background-task roster before
the lifecycle boundaries, and lets the provider-exit test expect the open
turn's interrupted revision instead of a tombstone.
2026-09-09 01:00:00 -07:00
Jinwoo Hong 65631e449a refactor(ai-vault): split the session scanner into a transcript reader and consumers (#19666)
* refactor(ai-vault): split the session scanner into a transcript reader and consumers

The scanner's only output was the Session History summary; a second reader
of the same transcripts (a search index) had nowhere to plug in without
hooking the parse itself. Extract a reader that owns each file read, keeps
the resumable cursor and publishes every decoded message to registered
consumers. The session list stays a fold inside the parser and is the only
consumer here. Parsers take an optional message sink instead of a scope.

Also: read Cursor chats/<md5>/<uuid>/meta.json for cwd, title and
timestamps (Cursor transcripts carry only role and message); share the
lazily spawned worker-thread host between the OpenCode SQLite reader and
the port-scan probe; probe OpenCode's schema before querying; keep the
newest-N discovery set with a bounded insert instead of sort+slice.

Session list output is byte-identical to main across all 18 providers
cold and append-resumed; the one Cursor session gains cwd/timestamps
from meta.json.

* fix(ai-vault): serialize per-path parses and report unpublished reads

Overlapping parses of one transcript share the cached resume point's
message channel, so the second beginRead dropped the first read's
consumers and the first finishRead handed them the wrong outcome. Two
callers really do overlap: a forced refresh restarts a scan while the
aborted scan's parse is still in flight, and the title reader parses
outside any scan. Restore the per-path lane around the whole
lookup-read-store sequence.

OpenCode's SQLite sessions are decoded on a worker thread the channel
cannot reach, so their reads published no messages while reporting a
complete span. Finish those reads as incomplete instead, so a consumer
never records a cursor for a stream it did not receive.

* fix(ai-vault): degrade a refused cursor chats read instead of dropping sessions

A refused WSL read of Cursor's chats tree rethrew, and the per-file catch
in discovery then recorded an issue and skipped the transcript. Before
the meta.json join Cursor had no content dependency, so a stalled distro
could not hide a Cursor session at all. Degrade to no metadata for the
scan and report the chats root once. The parse cache stays honest without
the throw: discovery stats no meta.json on a refused scan, so the entry's
recorded size omits it and the next healthy scan re-reads the transcript.

The per-scan index scope covered discovery only, so every Cursor finalize
re-read the chats root to validate the module cache. Move the scope to
scanAiVaultSessions, which spans discovery and parse.

Also drop the unused signal parameters the sink threading added to the
Devin and Hermes content parsers, by giving each file parser a private
record parser instead.

* fix(ai-vault): do not cache a cursor parse whose meta.json read was refused

Discovery stats meta.json into the candidate's cache key, so when only
the meta.json read is refused the un-enriched session was stored under a
key that looks unchanged and reuseCachedSession never re-ran the enrich
hook. The session stayed without cwd until Cursor rewrote the file. The
enrich hook now reports 'refused', the resumable state exposes
isCacheable, and the parse cache drops the entry instead of storing it,
so the next healthy scan re-parses. The index-read branch is unaffected:
it never stats meta.json, so its key is honest already.

* fix(ai-vault): separate the transcript's size from its cache key

sizeBytes folds a content dependency's size in, so it is a cache key
rather than a file length. The reader compared a transcript byte offset
against it and reported it as a whole-file read offset, which for Cline
handed consumers an offset past the end of the file it read. Carry the
dependency's own size on FileWithMtime and subtract it in the reader.

A refused sibling stat rethrew, so discovery recorded an issue and
skipped the transcript, the same drop removed for the readdir and read
paths. Degrade to no dependency, note the tree once, and mark the key
untrustworthy.

An untrustworthy key no longer costs the resume cursor: the entry is
stored under an mtime no stat can produce, so unchanged is false while
the resume point survives and the next scan resumes instead of re-reading
the whole transcript.

* test(ai-vault): pin the untrustworthy-key mechanism, not just its effect

Both refusal tests asserted that a later healthy scan re-enriches, which
a plain store would also satisfy once the resume cursor was preserved.
Assert the cache entry directly: its mtime is the unmatchable sentinel
and its resume point survives. The sentinel is exported so the tests name
the contract instead of repeating -1.

* refactor(ai-vault): track a session's sidecar file apart from its transcript

Folding Cursor's meta.json stat into the transcript's mtime/size made one
key mean two things, and every round of review found another consequence:
a byte offset could not be compared against it, a refused sibling read
took the transcript down with it, and an un-enriched parse cached under
it looked current forever. Main already had the answer for a file the
transcript key cannot see: Codex titles are refreshed at reuse time over
the cached session, not folded into the key.

Discovery now records the sibling as its own observation, unknown when it
could not be read. A cache hit needs both the transcript key and the
sidecar to match. When only the sidecar moved, Cursor re-merges it over
the stored un-enriched fold result and never re-reads the transcript;
Cline, which reads its sibling as part of the parse, re-parses.

Merging over the fold result rather than the accumulator makes enrichment
pure, so a meta.json rewritten with a new cwd replaces the old one
instead of losing to it. That was unreachable while the merge used ??= on
a session it had already enriched.

Cline and the remote scanner move to the same field, so the fold is gone
from both discovery paths.

* fix(ai-vault): tell an absent sidecar from an unreadable one

Three places collapsed the two. sidecarUnchanged returned true for any
observed 'none' without reading the entry, so a sidecar that was deleted,
or one that was unreadable last scan, both read as cache hits. Native
discovery mapped every non-WSL stat failure to 'none', so an EACCES on
meta.json left a session enriched from a file nobody can see, with no
scan issue. Remote discovery could not tell a missing sibling from a
failed stat, because statRemoteSessionFile returns null for both.

'none' is now a claim: absent-now is a hit only when it was absent before
or the agent never had a sidecar, and only ENOENT/ENOTDIR reads as
absent. statRemoteSessionFile grows an opt-in rethrow so its caller can
distinguish the two failures it already reports.

Also rewrites three comments in the cursor chat-meta reader that still
described the deleted fold.
2026-09-09 03:36:41 -04:00
Merge Sim b256cb5a69 test: align settled turn status expectations 2026-09-09 00:24:23 -07:00
Brennan BensonandMerge Sim acd501486d Unify tab surface selection across workspace activation (#19635)
* Unify tab surface selection across workspace activation

* Cover the folder activation entry point and name its selection contract

Rewrite the folder-workspace selection tests to drive setActiveFolderWorkspace,
the entry point this PR rewrote; they previously went through setActiveWorktree
and exercised the git-worktree projection instead, so none of them failed
against pre-PR code. Add the layout-only ownership case.

Hoist the remembered-file condition out of a three-deep nested ternary and pin
the remembered agent-session/simulator cases that make it load-bearing, and
replace the Parameters<typeof ...> indirection with a named
ActiveSurfaceSourceState.

* Pin the folder-path openFiles fallback

The folder path now reaches the shared openFiles fallback: with no groups, no
layout and the remembered browser tab gone, an open file selects the editor
surface instead of falling through to terminal. That is parity with the
long-shipped git path, and nothing covered it.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-09 00:17:27 -07:00
Merge Sim 9a1af944cf native-chat: avoid stale working status on settled turns 2026-09-09 00:13:48 -07:00
Brennan BensonandMerge Sim 5868fdc9e3 feat(native-chat): report Codex background tasks in the chat strip (#19346)
* feat(native-chat): report Codex background tasks in the chat strip

The background-tasks strip works for Claude only; a structured Codex
session shows nothing in it. Feed it from the Codex app-server stream.

The strip stands for work that OUTLIVED a turn, which is what the
monitoring header, Claude's foreground suppression, and the conversation
command gate all already assume. Codex has no `is_backgrounded` flag, so
that fact is derived from the turn boundary: a `subAgentActivity` child or
a primary-thread `commandExecution` becomes visible once the turn it
belongs to completes and it is still unsettled.

`turn/completed` only reveals a task here, never settles one — measured on
`codex app-server` 0.153.4, a spawn_agent child reported `completed` 95.8s
after its parent turn ended. Only a child's own activity kind settles it.

Codex exposes no honest stop: `turn/interrupt` on a child ends its turn
without emitting a terminal activity item and leaves its shell running. So
the state carries a new optional `supportsStopAll: false`, the strip hides
a control that could not act, and the blocked-command message asks the user
to wait rather than to press a button that does not exist.

* refactor(codex): move session teardown out of the structured adapter

Merging main crossed the 300-line cap on
`codex-structured-session-adapter.ts`: the rewind backend (#19235) and this
branch's close-time strip clear both landed in it. The four close paths move
verbatim into `codex-structured-session-teardown.ts`, where they funnel
through one `settled` helper instead of repeating the notification-retry and
background-task cleanup at each call site. No ratchet bump.

Also normalize a background task's description once at receipt rather than on
every projection; the roster is re-projected on each observed frame.

* fix(codex): drop the shell row the journal already settles

A `commandExecution` still `inProgress` when its turn ends was reported as a
`command` task. But `settleCodexJournalTurn` writes exactly those items to the
journal as `state: 'failed'` on `turn/completed` and forgets them, so the strip
row would have claimed a shell was still running at the same instant Orca
recorded that it was not — two surfaces contradicting each other about the same
process.

A subagent is the opposite case and stays: the roster pointedly does not sweep
at a turn boundary, because children measurably outlive it. That leaves the
producer making exactly one claim — these spawn_agent children are still live
after their turn — which the durable roster row corroborates.

* fix(native-chat): track Codex background execution lifetimes

* fix(native-chat): keep running tool groups from claiming completion

* Fix runtime catalog and capability expectation

* fix(codex): keep a child's name on the command row that outlives it

A child agent's commands stay hidden behind its agent row while the child
works. Once the child's turn settles with a command still running, that
command surfaces as its own row labelled from the raw command string, so
'long_probe' became "/bin/zsh -lc 'ping -c 300 127.0.0.1 > /dev/null'"
at the moment that row was the only remaining signal for the work.

Qualify a child's command row with the child's label. Resolved on read,
so a label registered after the command still lands, and bounded by the
existing description cap so admission accounting stays valid. Primary-
thread commands are left unqualified: they have no child to name.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-09 00:13:20 -07:00
Merge Sim ff55fccd80 Make the structured turn lifecycle row durable so completed durations survive
A structured-chat turn used to end by tombstoning its running lifecycle item,
which threw away the only durable record of when the turn ended. Completed
"Worked for" labels therefore depended on the renderer having observed the
turn finish, and vanished on reopen.

The lifecycle item is now revised in place, never tombstoned:
- running, with startedAt, at the provider's turn start
- completed or interrupted, with completedAt, at the provider's terminal frame,
  a user stop, or a child exit the host observed
- unverifiable, with no end, when a cold acquire finds a running row from a
  generation whose exit nobody observed

Both timestamps are the execution host's clock at receipt, captured before the
deferred sink, so the completed value is identical on every client and needs
no client clock. Codex history restore uses the provider's own second-granular
endpoints for turns that predate this change. Desktop and mobile read settled
durations off the journal through one shared selector, and anchor the live
counter on the host start with the client's local receipt so a skewed client
clock never leaks into the label. Locally observed durations remain the
fallback for hosts that still tombstone.

Timestamps live inside the existing turnLifecycle field, which old clients
strip, and every working-state consumer keys on state === 'running', so no
capability negotiation is needed.
2026-09-09 00:06:41 -07:00
Jinwoo Hong 0fe132ea29 fix(orchestration): file mail from terminals in no Run under an unbound Run (#19696)
* fix(orchestration): file mail from terminals in no Run under an unbound Run

#19542 deleted the fallback that filed such mail under the legacy Run, because a
live row there makes the schema-skew probe read the database as pre-Runs and
replay adoption on the next open. That refusal also broke the first command in
the guide: `orca orchestration send --to <handle>` between two plain terminals,
which worked in v1.4.198.

Restore delivery by filing under `run_unbound`, a Run the probe never matches,
created on first use so `run list` shows it only to a user who has such mail.

Claude-Session: 1fec75fd-224b-46ab-95fe-d88e0f3d9ff9

* fix(orchestration): create the unbound Run only for a null Run id

Claude-Session: 1fec75fd-224b-46ab-95fe-d88e0f3d9ff9
2026-09-09 03:06:38 -04:00
Jinwoo Hong 2ee4053ae6 fix(native-chat): merge duplicate native-chat-types import (#19698)
#19230 added a second import of the same module, and the focused
code-quality gate (import/no-duplicates, --deny-warnings) fails every PR
opened on main since it merged.

Claude-Session: 1fec75fd-224b-46ab-95fe-d88e0f3d9ff9
2026-09-09 03:05:10 -04:00
Brennan BensonandMerge Sim 7197593e31 fix(lint): preserve deliberate collator benchmark baselines (#19686)
Co-authored-by: Merge Sim <sim@local>
2026-09-08 23:53:17 -07:00
Brennan BensonandMerge Sim d15a6df224 Render task checklists with update diffs and a composer progress panel (#19230)
* Render native chat task lists with incremental checklist updates

* Keep native chat task-list review plan out of repository root

* Render live Codex plan notifications through task checklists

* fix(native-chat): keep checklist test fixture within shared boundary

* Keep agent task progress in one composer panel

* Restore inline task checklists and historical update diffs

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-08 23:50:32 -07:00
Brennan BensonandMerge Sim addd9f3da7 fix(orchestration): accept a missing Run id from older federation coordinators (#19689)
* fix(orchestration): accept v1.4.198 coordinators on federationAttachStart

#19542 made runId required on the attach RPC and backfilled existing
attachments with '', on the premise that federation was unreleased. It
shipped in v1.4.198, so a v1.4.198 coordinator got 'Missing Run ID' from
an upgraded worker host, and every pre-upgrade attachment lost its mailbox
because home_run_id='' matches no Run.

- runId is optional on the wire; an absent id mints a per-attachment stub
  Run (run_federated_<dispatch>) through the same INSERT OR IGNORE path.
- migrate-v40 backfills existing attachments with the stub and inserts the
  stub Runs, so in-flight workers keep reporting back.
- create SQL gives home_run_id a DEFAULT '' so a rolled-back v1.4.198 host
  can still insert into a v1.4.199-created table.

Stub Runs never carry run_legacy_local, and #19542's attachment-mailbox
exclusion in the skew probe is untouched, so no adoption replay path is
reintroduced (probe test added).

* fix(orchestration): repair empty federated home Run ids on every open

A host on v1.4.199 that rolls back to v1.4.198, attaches workers (rows
land with home_run_id='' via the DEFAULT), then upgrades again never
re-runs the v40 backfill because user_version is already 40, so those
attachments stay without a Run and their control mail is refused.

Lift the two idempotent set-based statements out of migrate-v40 into
backfillFederatedStubHomeRuns and run it from both the v40 migration and
the OrchestrationDb constructor after migrate(), matching the existing
on-open rememberCurrentRunCoordinatorHandles repair.

Tests: reopen a v40 file DB holding a v1.4.198-shaped '' row through the
constructor and assert the stub Run and mailbox; pin the wire schema
accepting a v1.4.198 request with no runId ('' -> undefined, whitespace
passes the schema and is refused at the DB layer).

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-08 23:41:46 -07:00
Brennan BensonandMerge Sim 750e6ffada test(orchestration): pin the Run-required contract for unbound direct mail (#19684)
* test(orchestration): pin the Run-required contract for unbound direct mail

* test(orchestration): pin absent recovery keys and settle the push window for unbound mail

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-08 23:02:05 -07:00
Brennan BensonandMerge Sim 3a801d213d fix(orchestration): worker-start settles readiness on observed turn start, not write acceptance (#19423)
* fix(orchestration): worker-start settles readiness on observed turn start, not write acceptance

A dispatched PTY worker whose agent wedged at startup (six codex workers on
2026-09-07) was reported 'ok: true, state: ready, stage: input_accepted': the
preamble write was acknowledged with observationTimeoutMs: 0 and nothing ever
verified a turn began. The corpse and the healthy worker produced identical
receipts.

worker-start now runs the existing second-stage prompt observer
(observeTerminalAgentPrompt) after acceptance, inside the 30s window the
client RPC grace already budgets for (orchestration-worker-start-prompt-budget):

- turn observed (or provider ack for structured sessions) -> ready
- permission prompt -> ready; positive liveness, surfaced in the receipt
- provider without a turn-start signal -> ready; observation: unsupported
- observation supported and nothing started -> worker state start_unknown,
  response state outcome_unknown with nextCommands. Honest 'unverifiable',
  never a death claim: the capability and terminal are kept, and
  worker-report settlement already reconnects a start_unknown worker that
  recovers and reports.

Also fixes the effect-verb lie that misdirected the first diagnosis of this
incident: agent-first worktree creation labeled its own brand-new agent
terminal 'reused_agent_terminal' (a role test picking a lifecycle verb) on
both the local and federation paths. It now says 'created'; readers keep
accepting the retired verb for rows persisted before the rename.

* fix: preserve worker authority through start observation

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-08 22:53:48 -07:00
Brennan BensonandMerge Sim fe9d538b6f fix(native-chat): keep collapsible chevrons beside their header text (#19656)
Co-authored-by: Merge Sim <sim@local>
2026-09-08 22:45:30 -07:00