- A merged import list named the same module twice, which the native code
quality plugin fails on.
- A running turn is now reported by the host with no duration, so the settled
map carries an explicit null for it; the hook test still expected the entry
to be absent.
- main gave the older-page action a cursor with a head-trim guard, so the
retention test's epoch-only action no longer typechecks; it now passes an
unbounded sequence, which is what the old shape meant.
- The roster comparator moved into the extracted module, leaving its import
unused in the reducer.
Main renamed the background-tasks view binding and grew its roster equality to
compare settled tasks; the extracted comparator adopts that logic so the
reducer keeps both it and the host-clock sample.
* feat(native-chat): name, group, and state the background-tasks strip
The strip above the composer described five different kinds of background
work as "Monitoring background tasks", with identical flat-dot rows. Now:
- Wire: additive optional `name`, `state`, `startedAt` on
AgentSessionBackgroundTask, plus `settledTasks` on the state object so
terminal siblings of a live fan-out stay visible without changing what
old clients render (they keep exactly the live `tasks` list).
- Reducer equality learns the new fields, so a publish whose only change
is a task's state is no longer judged equal and dropped.
- Header counts by kind and lists states within a kind; past three kind
segments (or on a narrow strip, measured by its own border-box against
the live root font size) it falls back to an honest total, never a
partial enumeration, and the strip stays expandable whenever the header
is lossy.
- Rows group by kind (Agents / Shell / Monitors / Workflows / Tasks),
stable-sorted first-seen-then-id, each with a kind icon, its own state
dot, a resolved name (description -> name -> kind label), and elapsed.
- Claude producer: task frames now carry name (agent_type/subagent_type),
a run state mapped from patch status, and first-seen startedAt. Terminal
statuses settle a task (completed->done, failed->blocked,
killed/stopped->idle) instead of deleting it; settled tasks render only
beside still-live work and flush when the last live task ends, so the
strip exits exactly when it does today. An unreadable patch leaves a
task open, never settled.
- Turn gating moves off the strip: the tracker no longer zeroes its
roster during a foreground turn, and the client renders the strip
whenever it has contents while the idle-only flag now gates just the
animated monitoring indicator and conversation commands.
* feat(sidebar): indent native-chat subagents under their session row
buildSubagentChildRows() has always rendered indented children from
parentEntry.subagents, and the structured-session status bridge has
always published an AgentStatusEntry for native chat — it just never
populated subagents. Connect them:
- Wire: additive optional `backgroundTasks` on AgentSessionStatusSummary
(live tasks only), projected by the host status feed from the
provider's backgroundTaskState hook and republished on task edges via
the background-task channel, with the shared task equality suppressing
no-op re-projections.
- Bridge: maps agent-kind tasks onto the sidebar's own
AgentSubagentState (working/waiting/blocked, terminal -> idle) — kinds
stay distinct, so a backgrounded shell never lands in a subagent
count — and extends its pre-write equality with the existing
agentSubagentsEqual.
- parentIsFresh for a bridge entry means "the host feed reported a
change inside the sidebar's ordinary evidence window": every publish
restamps evidenceObservedAt, and a dead feed stops restamping, so
children decay to idle on lost contact instead of pinning 'working'.
* fix(native-chat): settle tasks the aggregate roster evicted first; carry usage
Real-agent QA showed settledTasks never rendered. A frame capture from the
SDK (probe against claude 2.1.261) explains it: when a backgrounded child
finishes, the producer emits `background_tasks_changed` FIRST — with the
task already absent — and only then `task_updated`/`task_notification`
with the outcome, in the same tick. The tracker's settle path looked the
task up in the live roster the aggregate had just evicted, so retention
lost the race 100% of the time.
Fix: aggregate eviction of a live backgrounded task now parks its details
in a bounded recently-removed map (new claude-settled-background-tasks.ts,
which also owns the settled roster), and the trailing terminal edge
consumes it. A removal whose outcome frame never arrives still vanishes —
nothing is guessed into a finished state. A second terminal edge for the
same task re-derives the settled state and can add final usage. The
captured sequence is replayed verbatim as a tracker test, including the
kill-at-exit tail proving the strip still exits with the last live task.
The same capture disproved the PR's earlier claim that Claude task frames
carry no usage: task_progress and task_notification both carry
usage.total_tokens. Additive optional `totalTokens` on the wire task,
covered by the shared equality; the tracker takes usage (never the
transient "Running <tool>" description) from task_progress, and rows
render the mock's "18.1k · 2m" meta — settled rows keep final usage with
no still-growing clock.
* chore(i18n): sync runtime-required catalog for backgroundTasks.runningList
* fix(native-chat): preserve background task lifecycle and bound update work
* fix(native-chat): transfer resumed background tasks to one live owner
* fix(native-chat): bring structured session host under the line cap and restore subscribe fixture
* fix(native-chat): complete journal stubs and stop notifying on feed teardown
The status feed's projection cache calls journal.cursor(); the rename test's
stubs are cast through unknown, so the missing method only surfaced at runtime.
Teardown runs only once nothing is activated, so there is no mounted reader to
notify - clearing confirmed sessions is what prevents a stale live on reactivation.
* feat(native-chat): lead each strip header count with its kind icon
The header carried one aggregate state dot, so a fan-out of agents and a
monitor looked alike. Each count segment now leads with its own kind glyph;
a collapsed total spans kinds and takes none.
Monitor is the heartbeat AgentStateDot already draws for monitoring, so the
strip and the agent sidebar speak one vocabulary.
* feat(native-chat): give the strip's monitor heartbeat the sidebar amber
The glyph matched AgentStateDot but the colour did not, so a monitor in the
strip did not read as the monitor in the agent sidebar. One shared tone helper
now serves the header segment and the expanded row, so they cannot diverge.
Monitoring is a state the app already colours; the other four kinds are plain
markers and stay neutral. A running turn still dims the whole set.
* fix(native-chat): draw the strip header separator in a visible tone
The separator used `text-border`, a divider-line token that is 7% white in
dark mode - an order of magnitude fainter than the counts on either side, so
the dot between them read as absent. main.css already records that token as
too faint for a visible mark.
* fix(native-chat): give the worktree-ps journal stub a cursor
The status feed's projection cache calls journal.cursor(); this stub is cast
through unknown, so the missing method only surfaced at runtime. Its journal
never changes, so a real one would hold the cursor steady.
* refactor(native-chat): split the sidebar subagent rows out of this PR
The strip stands alone: the sidebar mapping, its observation plumbing and the
AgentStatusEntry.subagents wiring move to a stacked follow-up. No wire field
here is sidebar-only - the strip's rows read name, state, elapsed and tokens.
* perf(native-chat): keep task usage out of the session status summary
A `task_progress` frame ticks a background task's `totalTokens`, which
failed the status feed's equality check and re-broadcast a full summary to
every `agentSession.subscribeStatus` subscriber — paired-web and SSH/relay
clients included — for a number no session list renders. The projection now
drops usage; tokens keep flowing on the background-task channel the strip
reads.
* fix(native-chat): correct token unit rounding and drop the unused dot state
`formatBackgroundTaskTokens` rounded before choosing the unit, so 999_950
rendered as "1000k" instead of "1m"; pick the unit from the rounded value.
`backgroundTasksDotState` has no caller on this branch or the stacked
sidebar PR, and its multi-kind branch would report 'monitoring' over an
attention state. Delete it rather than leave it to be wired up.
* fix(i18n): drop the orphaned backgroundTasks.runningList key
The strip rewrite removed its only call site, and an unreferenced key gets
promoted into the eagerly parsed boot catalog. Delete it from en.json and
regenerate en-runtime-required.json.
* fix(native-chat): show the reason on every attention row
The row guarded the reason line on 'waiting', so an 'unverifiable' child
("no contact") and a 'blocked' one ("failed") rendered bare while the
collapsed header named exactly those reasons. `backgroundTaskStateReason`
already returns null for the non-attention states, so the guard was only
lossy — the SSH boundary requires the unverifiable verdict stay legible.
Also keys the header segments off their kind discriminant instead of the
translated display text.
* fix(native-chat): make the strip header agree with its own count
The headline counts live AND settled rows, but the state breakdown omitted
'done', so one working agent beside four settled ones read "5 agents — 1
working": the count said five, the breakdown accounted for one. Done now
appears in the muted detail (never as an emphasised segment) so the two
agree.
The single-command header also drew an elapsed clock on a settled task,
which the row already refuses as a lie about finished work.
* perf(native-chat): memoize the background-task roster grouping
The 1 Hz elapsed tick re-rendered the strip, and the render body regrouped,
re-sorted and re-translated every task each time only `now` had changed.
The header still derives from `now` on purpose.
* test(native-chat): cover settled rows and the mid-turn mounted strip
Neither headline behaviour had component coverage: every strip render passed
`settledTasks={[]}`, and the `showBackgroundTasks` seam was never set true,
so the strip staying mounted through a running turn was exercised nowhere.
Adds a settled-beside-live row test (final usage kept, no clock, no stop) and
a mid-turn mount test (strip present, turn owns the voice). The background-task
tests share one session-element helper so the file stays under its line cap.
* refactor(claude): keep MAX_TASK_ID_LENGTH module-private
Nothing outside claude-background-task-frames.ts references it; the export
was residue from this PR's split.
* test(native-chat): give the mid-turn strip test a real turn
main now gates the composer's stop button on a provider-minted turnId rather
than the send-time working signal, so a test claiming a running turn has to
supply one. The controller mock hardcoded turnId null.
---------
Co-authored-by: Merge Sim <sim@local>
* fix(native-chat): one `/` picker for every agent, anywhere in the prompt
The composer only opened its picker when `/` was the first character of the
draft, so a skill named mid-sentence ("validate it with /electron") offered
nothing. Codex was worse: its `/` menu listed commands only, and skills lived
on a separate `$` trigger, so the prompt box behaved differently per agent.
`/` is now the whole composer grammar. It opens one grouped commands+skills
menu for every agent with a known grammar, both at the start of the draft and
mid-prompt after whitespace. The `$` trigger is gone.
Per-agent invocation is preserved where it belongs — in what a pick writes.
Each row carries its own token, so choosing a skill in Codex inserts
`$electron` while Claude inserts `/electron`, and the text that reaches the
agent stays the text that agent actually invokes. Only a draft-leading command
is dispatchable; picking one mid-sentence completes the token instead of
sending the command on its own and discarding the draft.
Name collisions now key on whether both kinds share a sigil, so a Codex
`/review` command and a `$review` skill stay separate rows.
* test(native-chat): model dismissal inputs on the live `/` grammar
The trigger-key swap cases still used `$:4` keys. editReplacesTriggerToken is
sigil-agnostic so they passed, but they modelled an input the composer can no
longer produce.
* test(native-chat): pin inline picker dispatch and discovery reuse
---------
Co-authored-by: Merge Sim <sim@local>
The partial-tail state machine and the preview normalizer's per-sequence
parser each carried their own copy of the byte-after-ESC table, so the
DCS/SOS/PM/APC set was stated twice and could drift.
Move it to `terminal-escape-introducer.ts` and have both read it. The
normalizer now classifies from `charCodeAt` instead of `value[i]`, which
drops the one-char string it minted on every escape it parsed -- the same
allocation the ground scan already avoids.
No behaviour change: `parseAnsiControlSequence` matches its previous
implementation on every 2- and 3-byte sequence after ESC and on 20k
random control-dense streams (differential check, not committed).
`mergeSubmissions` caps submissions at 256, but `mergeItems` had no
equivalent bound, so `state.items` grew for the whole life of a long
structured session while the live path caps itself to its read window.
Head-trim `items` to a retained-item limit when a live batch merges, and
set `hasOlder` so anything trimmed is still reachable by paging. Paging
older raises the limit to what the page produced, so a live batch slides
the widened window instead of collapsing it back to the cap -- the same
shape as the live path's growing `limitRef`.
Findings from an independent adversarial review of the typed turn record:
- A Codex rewind adopted the provider's item list as the new epoch, and the
provider never returns the host's own turn rows, so every duration before
the rewind point vanished. The host's turn rows are now spliced back beside
the item each followed, and recovery no longer expects the provider to
prove rows it never owned.
- The epoch row was stamped with the current schema version, so an older host
latched read-only at row 1 of every new session, defeating the mixed
version design. It carries no body and stays at v2; a stored-row test now
reads SQLite directly, because the reader upcasts every row on read.
- A send Codex folds into a running turn shares the opening prompt's provider
key, and the alias map credited the duration to the later prompt. The
earliest submission naming a key now wins.
- The live counter anchored on first sight, so a client attaching mid-turn
counted from zero. Published frames now carry the host's clock, the reducer
keeps the last sample with its local receipt time, and both clients anchor
on how long the host says the turn has run.
A single unterminated OSC 9999 marker split across many PTY chunks
re-scanned the whole accumulation for a terminator on every chunk, so
work grew with the square of the frame length.
Carry how much of `pending` already failed the search and resume one
character before it, which is enough for an `ESC \\` straddling the
chunk boundary.
* fix(mobile): make every native-chat text node selectable
Long-press selection worked on some chat text and not others. Markdown
paragraphs — the default block for agent prose — were the one block type
left out when headings, quotes, code, lists and table cells gained
`selectable`, and tool result output, diff rows, the unloadable-image
placeholder, permission/question bodies and the send-error banner never
had it at all.
Selection is now set on every content Text in the chat surface, on the
outermost block Text so nested inline spans inherit it. Labels inside a
Pressable (option rows, tool-line headers, buttons) are deliberately left
alone: selection there would swallow the tap they exist for.
Extracting MobileNativeChatEmptyState keeps the view under its max-lines
cap and matches desktop, where NativeChatEmptyState is already its own
component.
Tests render each surface and assert selection on the block that carries
the prose; both files were ablated against the unfixed source (4/10 and
3/5 red) so they pin the defect rather than the current behavior.
* fix(mobile): support native text range selection on iOS
* fix(mobile): remove persistent assistant message controls
* fix(mobile): scope patch-free text selection to chat
Use the stock react-native-uitextview dependency behind an iOS adapter and opt assistant Markdown into range selection only in native chat. Preserve the existing React Native Text behavior elsewhere and remove the persistent assistant controls.
---------
Co-authored-by: Merge Sim <sim@local>
* fix(native-chat): show Claude working from the send, not the provider echo
A structured session read as working only once a turnLifecycle row existed.
Codex writes that row ~150ms after the send; Claude cannot write it until the
SDK echoes the user message back, measured at a 3.4s median and 18s at p90, so
the chat and every session list read idle for the whole wait.
The journalled submission is the host's own evidence a turn is owed, so the
shared projection reads it too. `unknown` still counts -- the ack budget
elapsing answers delivery, not whether work is owed -- while a recovered
`unknown` does not, which needed the existing row flag carried onto the
projected submission.
Claude's activity line now stays the generic fallback. Its only turn-wide frame
carries a bare token, and its task_* prose describes a spawned task rather than
this turn; compaction is kept because it explains an otherwise silent wait.
* Fix structured chat pending-work lifecycle and mobile cancellation
---------
Co-authored-by: Merge Sim <sim@local>
The turn record is now its own item kind rather than a status row carrying a
lifecycle field: no text to misuse, and the fold matches the durable turn
record other systems keep. Rows that carry it are stamped journal schema v3;
every other row stays v2, so an older host keeps reading them and latches
read-only at the first v3 row instead of truncating the epoch.
Clients that predate the item would paint an unknown kind as a text bubble,
so the host publishes the legacy status form to any client that does not
advertise agent-session.turn-item.v1, through the same per-client seam
background tasks use. The downgrade is transitional and goes once no
supported release lacks the capability. The shared projection now renders
unknown item kinds as nothing, so later kinds need no gate. One shared reader
handles both forms for old journals and old hosts.
A lifecycle row now names the user item that opened the turn by its provider
key, so clients attribute timing explicitly and fall back to journal order
only for rows from older hosts. A provider-initiated turn with no prompt can
no longer claim the previous prompt's duration.
When the provider measures the turn itself (Codex turn.durationMs, Claude
result.duration_ms) the terminal row records it and clients prefer it over the
host interval, so a turn shows the same number live and after a history
restore. Host receipt times remain the live-counter anchor and the fallback.
* perf: check tunnel queue capacity before copying frame payloads
* fix(browser-tunnel): derive writer admission size from the frame encoder
Shares one encoded-length helper so the pre-encode capacity check cannot drift
from what encoding allocates, and covers the exact byte-cap boundary.
---------
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Every row in the 2000-row retained tail is sliced from the PTY chunk it
arrived in, so a spinner workload — CR-redraw frames, exactly what Claude
Code and Codex emit — makes a ~5 KB tail pin one whole chunk per row.
Measured on 400 x 64 Ki-char chunks: a 5,090-char tail retained 37.5 MB.
At 16 Ki x 1600 it was 46.9 MB.
Own the two values the plain tail path actually retains, newlyCompletedLines
and the partial line, at the point they enter the tail: 37.5 MB -> 0.18 MB
and 46.9 MB -> 0.09 MB, with per-chunk cost unchanged within noise. The
transcript keeps the same row objects, so it is covered transitively.
The multiline redraw builder never slices normalizedChunk — it writes rows
character by character — so only its partial line, which is re-sliced from
its own row on every frame, needs owning. Owning its rows as well measured
no retention gain and cost CPU on fullscreen TUI floods.
Also own the wait-blocked keywordCarry: 31 characters that pinned a full
lowercased copy of the chunk, per PTY.
Follow-up to #19396, which introduced ownRetainedString for the much smaller
pending-control fragments.
* perf(terminal): release oversized backing strings behind pending controls
* test(terminal): record reproducible pending-storage gate evidence
* perf(terminal): own retained control fragments with a fast copy primitive
The pending-control ownership landed with a charCodeAt block copier (10 us
at 4 Ki, 170 us at 64 Ki), so it needed a "copy only when discarded output
dominates the tail" gate to stay affordable. That gate was the whole cost
problem: on adversarial streams it fires every chunk and pays the slow copy
(+45% on 16 Ki ANSI chunks, +81..111% on 194 Ki status chunks), and it also
skipped ownership on fragments too small to be sliced strings anyway.
ownRetainedString replaces it with a Buffer utf16le round trip (0.57 us at
4 Ki, 21.9 us at 64 Ki) and returns anything below V8's SlicedString
kMinLength unchanged. Buffer is absent in the renderer and on mobile, so the
copier is resolved once behind a lone-surrogate round-trip self-check and
falls back to the block copier. With a ~1 us copy the gate is unnecessary:
ownership is now unconditional at all three retention sites and the
adversarial cases land within noise of the un-owned parsers.
The three forced-GC threshold fixtures are replaced by one forced-GC test
for the primitive plus deterministic spy assertions that each site routes
its retained value through ownRetainedString. All fidelity and differential
coverage is kept.
* fix(terminal): escape the NUL in the round-trip probe
A raw NUL byte in the source made git treat the file as binary, so its
diffs and blame were unreadable. Escapes are equivalent at runtime.
* perf(terminal): skip plain text between partial escape sequences
* perf(terminal): take the ESC at hand before searching for one
The unconditional ground-state indexOf regressed dense back-to-back
SGR/CSI streams, where the code unit at the cursor is already the ESC and
the search pays call plus SIMD setup to find it in place. Check the
current unit first and fall back to the native search otherwise.
0.9 MiB dense SGR/CSI medians: 2.04 ms before this PR, 2.56 ms with the
unconditional search, 1.92 ms with the hybrid. Sparse colored logs and
plain text keep the full search win (0.36 / 0.016 ms vs 1.54 / 1.41 ms
baseline). Differential over the full VT alphabet with lone and split
surrogates matched 388,416 cases against both prior implementations with
zero mismatches; the ground-scan work budget now records 16 inspected
code units and 2 native searches.
The two flakiest tests on main both guessed at a duration instead of
waiting for the condition.
- windows-pty-job.win32.test.ts assumed job teardown finished in 1.5s;
under load on a Windows runner it does not. Poll isAlive up to 30s
instead -- the assertion is unchanged, so a real leak still fails.
- structured-agent-session-claude-options-round-trip.test.ts relied on
vi.waitFor's 1s default for a two-hop handoff; give it 10s.
Both are test-only and strictly widen an existing wait.
* perf: probe requested pane keys instead of enumerating records
* perf(agent-status): drop the requested-key array from pane removal
Probing the pane keys still beat enumerating the record, but materializing
the requested set allocated on every call including the common no-match
path, where a dozen records are swept per retirement. Copy lazily on first
match instead, and cover the set-disagreement and prototype-key cases.
---------
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
* perf(workspaces): reuse unchanged heartbeat status projections
* perf(workspaces): skip unchanged title-sync session collections
* test(workspaces): cover constant-size key swaps in status projection reuse
Also point the title-input gate at a tracked source file; the previous
motivatingLink referenced an untracked .agents skill path.
* perf(runtime): skip hibernation inventories without completed agents
* perf(runtime): skip hibernation status scan when no runtime owners
* fix(runtime): require host evidence for workspaces resolved mid-inventory
`runtimeLivenessRequiredWorktreeIds` was sampled before the runtime inventory
await, while the plan is built from the state after it. A workspace that gained
tabs or resolved its runtime owner during that window was therefore absent from
the required set, so the planner did not demand fresh host evidence for it and
fell back to client PTYs — client bookkeeping answering for the execution host.
Union the post-await targets into the required set inside `snapshotFromState`.
Union rather than replace: the set only ever grows, so the planner can only skip
more workspaces, never authorize a hibernation it would previously have refused.
An absent inventory stays a skip; nothing reads it as an exited PTY.
Extract the coordinator test fixtures so the regression lives in its own file
without pushing the coordinator suite past the 800-line test cap.
* fix(types): annotate hibernation fixture mock exports for declaration emit
TS2883: the inferred `Mock<Procedure>` types of the fixture's exported
`vi.fn()` bindings reference `Procedure` from a transitive `@vitest/spy`
path that cannot be named.
* chore: keep local-file-sink-memory test formatting as on main
The merge commit's pre-commit hook reformatted a file this branch does not own.
An interrupted or unverifiable turn must not read as completed for any
consumer that renders status text raw. One shared helper builds the text for
both providers from the lifecycle state.
Routes both Codex turn notifications through the background-task roster before
the lifecycle boundaries, and lets the provider-exit test expect the open
turn's interrupted revision instead of a tombstone.
* refactor(ai-vault): split the session scanner into a transcript reader and consumers
The scanner's only output was the Session History summary; a second reader
of the same transcripts (a search index) had nowhere to plug in without
hooking the parse itself. Extract a reader that owns each file read, keeps
the resumable cursor and publishes every decoded message to registered
consumers. The session list stays a fold inside the parser and is the only
consumer here. Parsers take an optional message sink instead of a scope.
Also: read Cursor chats/<md5>/<uuid>/meta.json for cwd, title and
timestamps (Cursor transcripts carry only role and message); share the
lazily spawned worker-thread host between the OpenCode SQLite reader and
the port-scan probe; probe OpenCode's schema before querying; keep the
newest-N discovery set with a bounded insert instead of sort+slice.
Session list output is byte-identical to main across all 18 providers
cold and append-resumed; the one Cursor session gains cwd/timestamps
from meta.json.
* fix(ai-vault): serialize per-path parses and report unpublished reads
Overlapping parses of one transcript share the cached resume point's
message channel, so the second beginRead dropped the first read's
consumers and the first finishRead handed them the wrong outcome. Two
callers really do overlap: a forced refresh restarts a scan while the
aborted scan's parse is still in flight, and the title reader parses
outside any scan. Restore the per-path lane around the whole
lookup-read-store sequence.
OpenCode's SQLite sessions are decoded on a worker thread the channel
cannot reach, so their reads published no messages while reporting a
complete span. Finish those reads as incomplete instead, so a consumer
never records a cursor for a stream it did not receive.
* fix(ai-vault): degrade a refused cursor chats read instead of dropping sessions
A refused WSL read of Cursor's chats tree rethrew, and the per-file catch
in discovery then recorded an issue and skipped the transcript. Before
the meta.json join Cursor had no content dependency, so a stalled distro
could not hide a Cursor session at all. Degrade to no metadata for the
scan and report the chats root once. The parse cache stays honest without
the throw: discovery stats no meta.json on a refused scan, so the entry's
recorded size omits it and the next healthy scan re-reads the transcript.
The per-scan index scope covered discovery only, so every Cursor finalize
re-read the chats root to validate the module cache. Move the scope to
scanAiVaultSessions, which spans discovery and parse.
Also drop the unused signal parameters the sink threading added to the
Devin and Hermes content parsers, by giving each file parser a private
record parser instead.
* fix(ai-vault): do not cache a cursor parse whose meta.json read was refused
Discovery stats meta.json into the candidate's cache key, so when only
the meta.json read is refused the un-enriched session was stored under a
key that looks unchanged and reuseCachedSession never re-ran the enrich
hook. The session stayed without cwd until Cursor rewrote the file. The
enrich hook now reports 'refused', the resumable state exposes
isCacheable, and the parse cache drops the entry instead of storing it,
so the next healthy scan re-parses. The index-read branch is unaffected:
it never stats meta.json, so its key is honest already.
* fix(ai-vault): separate the transcript's size from its cache key
sizeBytes folds a content dependency's size in, so it is a cache key
rather than a file length. The reader compared a transcript byte offset
against it and reported it as a whole-file read offset, which for Cline
handed consumers an offset past the end of the file it read. Carry the
dependency's own size on FileWithMtime and subtract it in the reader.
A refused sibling stat rethrew, so discovery recorded an issue and
skipped the transcript, the same drop removed for the readdir and read
paths. Degrade to no dependency, note the tree once, and mark the key
untrustworthy.
An untrustworthy key no longer costs the resume cursor: the entry is
stored under an mtime no stat can produce, so unchanged is false while
the resume point survives and the next scan resumes instead of re-reading
the whole transcript.
* test(ai-vault): pin the untrustworthy-key mechanism, not just its effect
Both refusal tests asserted that a later healthy scan re-enriches, which
a plain store would also satisfy once the resume cursor was preserved.
Assert the cache entry directly: its mtime is the unmatchable sentinel
and its resume point survives. The sentinel is exported so the tests name
the contract instead of repeating -1.
* refactor(ai-vault): track a session's sidecar file apart from its transcript
Folding Cursor's meta.json stat into the transcript's mtime/size made one
key mean two things, and every round of review found another consequence:
a byte offset could not be compared against it, a refused sibling read
took the transcript down with it, and an un-enriched parse cached under
it looked current forever. Main already had the answer for a file the
transcript key cannot see: Codex titles are refreshed at reuse time over
the cached session, not folded into the key.
Discovery now records the sibling as its own observation, unknown when it
could not be read. A cache hit needs both the transcript key and the
sidecar to match. When only the sidecar moved, Cursor re-merges it over
the stored un-enriched fold result and never re-reads the transcript;
Cline, which reads its sibling as part of the parse, re-parses.
Merging over the fold result rather than the accumulator makes enrichment
pure, so a meta.json rewritten with a new cwd replaces the old one
instead of losing to it. That was unreachable while the merge used ??= on
a session it had already enriched.
Cline and the remote scanner move to the same field, so the fold is gone
from both discovery paths.
* fix(ai-vault): tell an absent sidecar from an unreadable one
Three places collapsed the two. sidecarUnchanged returned true for any
observed 'none' without reading the entry, so a sidecar that was deleted,
or one that was unreadable last scan, both read as cache hits. Native
discovery mapped every non-WSL stat failure to 'none', so an EACCES on
meta.json left a session enriched from a file nobody can see, with no
scan issue. Remote discovery could not tell a missing sibling from a
failed stat, because statRemoteSessionFile returns null for both.
'none' is now a claim: absent-now is a hit only when it was absent before
or the agent never had a sidecar, and only ENOENT/ENOTDIR reads as
absent. statRemoteSessionFile grows an opt-in rethrow so its caller can
distinguish the two failures it already reports.
Also rewrites three comments in the cursor chat-meta reader that still
described the deleted fold.
* Unify tab surface selection across workspace activation
* Cover the folder activation entry point and name its selection contract
Rewrite the folder-workspace selection tests to drive setActiveFolderWorkspace,
the entry point this PR rewrote; they previously went through setActiveWorktree
and exercised the git-worktree projection instead, so none of them failed
against pre-PR code. Add the layout-only ownership case.
Hoist the remembered-file condition out of a three-deep nested ternary and pin
the remembered agent-session/simulator cases that make it load-bearing, and
replace the Parameters<typeof ...> indirection with a named
ActiveSurfaceSourceState.
* Pin the folder-path openFiles fallback
The folder path now reaches the shared openFiles fallback: with no groups, no
layout and the remembered browser tab gone, an open file selects the editor
surface instead of falling through to terminal. That is parity with the
long-shipped git path, and nothing covered it.
---------
Co-authored-by: Merge Sim <sim@local>
* feat(native-chat): report Codex background tasks in the chat strip
The background-tasks strip works for Claude only; a structured Codex
session shows nothing in it. Feed it from the Codex app-server stream.
The strip stands for work that OUTLIVED a turn, which is what the
monitoring header, Claude's foreground suppression, and the conversation
command gate all already assume. Codex has no `is_backgrounded` flag, so
that fact is derived from the turn boundary: a `subAgentActivity` child or
a primary-thread `commandExecution` becomes visible once the turn it
belongs to completes and it is still unsettled.
`turn/completed` only reveals a task here, never settles one — measured on
`codex app-server` 0.153.4, a spawn_agent child reported `completed` 95.8s
after its parent turn ended. Only a child's own activity kind settles it.
Codex exposes no honest stop: `turn/interrupt` on a child ends its turn
without emitting a terminal activity item and leaves its shell running. So
the state carries a new optional `supportsStopAll: false`, the strip hides
a control that could not act, and the blocked-command message asks the user
to wait rather than to press a button that does not exist.
* refactor(codex): move session teardown out of the structured adapter
Merging main crossed the 300-line cap on
`codex-structured-session-adapter.ts`: the rewind backend (#19235) and this
branch's close-time strip clear both landed in it. The four close paths move
verbatim into `codex-structured-session-teardown.ts`, where they funnel
through one `settled` helper instead of repeating the notification-retry and
background-task cleanup at each call site. No ratchet bump.
Also normalize a background task's description once at receipt rather than on
every projection; the roster is re-projected on each observed frame.
* fix(codex): drop the shell row the journal already settles
A `commandExecution` still `inProgress` when its turn ends was reported as a
`command` task. But `settleCodexJournalTurn` writes exactly those items to the
journal as `state: 'failed'` on `turn/completed` and forgets them, so the strip
row would have claimed a shell was still running at the same instant Orca
recorded that it was not — two surfaces contradicting each other about the same
process.
A subagent is the opposite case and stays: the roster pointedly does not sweep
at a turn boundary, because children measurably outlive it. That leaves the
producer making exactly one claim — these spawn_agent children are still live
after their turn — which the durable roster row corroborates.
* fix(native-chat): track Codex background execution lifetimes
* fix(native-chat): keep running tool groups from claiming completion
* Fix runtime catalog and capability expectation
* fix(codex): keep a child's name on the command row that outlives it
A child agent's commands stay hidden behind its agent row while the child
works. Once the child's turn settles with a command still running, that
command surfaces as its own row labelled from the raw command string, so
'long_probe' became "/bin/zsh -lc 'ping -c 300 127.0.0.1 > /dev/null'"
at the moment that row was the only remaining signal for the work.
Qualify a child's command row with the child's label. Resolved on read,
so a label registered after the command still lands, and bounded by the
existing description cap so admission accounting stays valid. Primary-
thread commands are left unqualified: they have no child to name.
---------
Co-authored-by: Merge Sim <sim@local>