Commit Graph
10236 Commits
Author SHA1 Message Date
Merge Sim aa9acc160f test(native-chat): retire the subagent-visibility guards now the roster renders
Two tests from the sibling item-coverage PR asserted that subagent items stay
on the generic gray row, explicitly gated on "until a real renderer exists".
This branch is that renderer, so both guards fire on merge — the handoff they
were written to mark rather than a regression.

They now pin the other side of it: subAgentActivity is suppressed because the
spawn-group roster renders it, and collabAgentToolCall deliberately stays
visible, since nothing guarantees a session reports subagent work as
subAgentActivity at all.

Git merged both files without conflict; only running the suite surfaced this.
2026-09-05 15:39:42 -07:00
Merge Sim cfdcf87fa7 Merge remote-tracking branch 'origin/main' into brennanb2025/codex-subagent-worklog
# Conflicts:
#	src/main/native-chat/agent-session-wire/provider-frame-disposition.ts
2026-09-05 15:34:17 -07:00
Brennan BensonandMerge Sim 471a5f4aa7 feat(native-chat): model Codex MCP and web-search items instead of leaking opcodes (#18763)
* feat(native-chat): model Codex MCP and web-search items instead of leaking opcodes

Codex's app-server sends 19 thread-item types; the structured translator handled
six. The rest fell through to a generic gray `codex · item:<type>` row, even
though the disposition table's own comment says it exists so a new item type
cannot leak like that — the table had one entry.

Give `mcpToolCall` and `webSearch` real tool-call bodies, and chrome `sleep`,
which carries only a duration and renders as nothing in Codex's own TUI.

`subAgentActivity` and `collabAgentToolCall` deliberately keep their generic
rows. They arrive in real sessions today and are currently the only visible
sign a subagent is running; hiding them before the subagent UI lands would
render minutes of work as an idle turn. Tests pin that they stay visible.

MCP tool names pass through verbatim when they contain `:`, `.`, `/` or `__`,
so `mcp__server__tool` survives instead of being title-cased into nonsense.

* fix(native-chat): keep Codex MCP tool identity and web-search results on the row

Four fixes to the Codex MCP / web-search item bodies:

- Drop the title-casing display name. `get_forecast` became `Get Forecast`,
  which no longer matches the raw snake_case identifiers that the diff
  renderer, question parsers, and tool-input previews dispatch on, and does not
  match how the Claude lane or the sibling `shell`/`apply_patch`/`web_search`
  bodies name a tool. The row name is now `server/tool` verbatim, the bare
  `tool` when no server is given, and `mcp` when the item names no tool at all.
  Server-qualifying also stops an MCP tool that happens to be called
  `apply_patch` from hijacking the diff renderer.

- Pass the MCP call's own `arguments` as the tool input instead of wrapping it
  in `{server, tool, arguments}`. Row-label derivation only reads top-level
  keys, so the wrapper degraded every MCP row to a truncated raw JSON blob.
  A non-object `arguments` stays addressable under a key rather than being
  dropped; an absent one becomes null, which labels as empty rather than `{}`.

- Carry a web search's `results` as the call output, bounded like every other
  inline payload and omitted when there are none. They were being dropped
  entirely, which showed less than the generic fallback row it replaced.

- No streaming branches were added for these two item types: the Codex delta
  stream is a closed set of six methods that neither can reach, so such
  branches would be unreachable.

* fix(native-chat): label Codex web searches and argument-less MCP calls

A row label is derived from top-level `input` keys only, so a webSearch
whose detail lives inside `action` — an opened page, an in-page find, or
a bare `other` — fell through to the raw JSON of the whole input, as did
the empty `query` Codex leaves on a completed search. Hoist the action's
`url`, `pattern` and `type` beside the query, keep the full `action`
object so the expanded detail loses nothing, and emit no input at all for
the start frame.

An MCP tool that takes no arguments sends `arguments: {}`, which passed
straight through and labelled the row a literal `{}`; treat it as absent
so the row reads as a bare `server/tool`.

Split the durable-identity half of the item translator into
`codex-thread-item-identity.ts`, re-exported so every existing import is
unchanged, to keep both files under the max-lines cap.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-05 15:33:04 -07:00
Neil 239e3c7e0b test: select seeded workspace and confirm sidebar reveal (#18921) 2026-09-05 15:29:46 -07:00
Neil 9faa27c5f4 test: align desktop platform oracles with native behavior (#18915) 2026-09-05 15:12:19 -07:00
Neil d7767fb196 perf(worktree): remove redundant creation and terminal startup work (#18793)
* perf(worktree): remove redundant creation and terminal startup work

* test(worktree): cover optimized creation call signatures

Preserve explicit branch adoption, WSL callback routing and sparse cleanup expectations.

* perf: preserve user Git checkout worker settings

* perf(git): skip malformed remote base probes

* perf(cli): avoid loading other agent hooks for Codex preflight

* fix(build): retain Codex preflight entry for packaged CLI

* test(ssh): wait for replacement PTY before lease recovery input

* test(ssh): verify recovered shell execution and lease ownership

* test(electron): reap isolated macOS crash reporters on teardown

* test: allow either observed self-exit snapshot ordering

* test: capture frozen-host input recovery evidence
2026-09-05 15:08:22 -07:00
Neil 7bec98466b test: canonicalize setup fixture paths before worktree lookup (#18912) 2026-09-05 15:04:52 -07:00
Neil 51eed5a1bc feat(cli): report SSH host platforms (#18896)
* feat(cli): report SSH host platforms

* feat(cli): include SSH connection status

* fix(cli): preserve unknown SSH connection state
2026-09-05 14:35:45 -07:00
Neil 2afc8b55ef test: pin worker visibility fixture command and handle (#18897) 2026-09-05 14:34:35 -07:00
Neil dce5ebd83d test: isolate native crash restoration and refresh stale fixtures (#18883)
* test: isolate native crash restoration and seed current integration facts

* test: await scoped GitLab preflight before URL transition checks
2026-09-05 14:18:22 -07:00
Jinjing cd70048092 Fix favicon retention across same-origin navigations (#18879)
* fix: retain favicons across same-origin navigations

Move favicon clearing from did-start-loading to did-start-navigation and
only clear when origin changes. Chromium re-announces favicons only when
the icon URL list changes, so clearing on every load orphans same-origin
navigations. Extract favicon URL validation into a shared module.

* fix: drop favicon on cross-origin redirects

When a same-origin navigation redirects to a different origin, the favicon should be cleared to prevent stale icons from displaying the wrong site's identity.
2026-09-05 14:10:43 -07:00
Merge Sim 3c0c9eedc4 Merge remote-tracking branch 'origin/main' into brennanb2025/codex-subagent-worklog
# Conflicts:
#	src/renderer/src/components/native-chat/NativeChatToolRun.tsx
2026-09-05 14:08:47 -07:00
Neil 5cec2c2dfc test: preserve Docker context in isolated VM recipes (#18884) 2026-09-05 14:07:08 -07:00
Brennan BensonandMerge Sim ddc5b75ac7 feat(native-chat): label Codex tool rows by what the command actually did (#18760)
* feat(native-chat): label Codex tool rows by what the command actually did

Codex's app-server `commandExecution` item carries `commandActions`, which
already classifies each command as a read, a search, or a directory listing
with the target path, name, or query extracted. Orca ignored the field, so
every shell call rendered as an undifferentiated row of raw argv.

Read it and name the row by its class, keeping the raw command and cwd for the
expanded view. Unclassified commands are untouched: absent, null, or malformed
`commandActions` produces byte-identical output to before.

Rank the search term above the command in the shared label keys so a classified
search row reads by what it looked for rather than the shell text that ran it.
No first-party tool input carries both keys today, so this only reaches the new
rows; an MCP tool supplying both would prefer its search term.

Note `commandActions` is the app-server spelling. `parsedCmd` is the rollout-file
shape and never arrives on this lane; a test pins that it stays ignored.

* feat(native-chat): give tool rows a category glyph beside their word

A row named only by a word makes the reader parse text to tell a read from
a search. Pair the word with an icon: icon for category, word for action,
argument for target.

Name the full eight-category vocabulary in `src/shared/native-chat-tool-icon.ts`
now — read/search/listFiles/unknown/fileChange/webSearch/mcpToolCall/
subAgentActivity — even though only the classified shell categories reach a row
today, so the MCP and web-search rows landing separately inherit these names
rather than coining their own. Glyph ids are the lucide spelling shared by
`lucide-react` and `lucide-react-native`, so mobile can resolve one name to its
own component when it adopts this; mobile rows stay text-only for now.

The glyph is decorative and `aria-hidden`: the word is the accessible name, and
never renders without it. One glyph per category, fixed across running,
completed, and failed — a row that swapped icons on completion would read as
changing identity — so the run header's active row also takes its category glyph
instead of the generic wrench it fell back to once these rows stopped being
called `shell`. A word outside the vocabulary gets the terminal glyph rather
than a blank slot, so rows stay left-aligned.

Also stand `.` in for a `listFiles` action whose `path` is null, which is what a
bare `ls` sends. The row named the action and then showed the raw argv as its
target; now it names the directory it listed.

* fix(native-chat): hold the tool run header's glyph fixed and size its slot to 16/14

The header swapped its leading glyph on settle: the active tool's icon while
running, a check once done. That is the identity swap a fixed per-category glyph
exists to prevent — the row appeared to become a different thing when it
finished. Name the header by the run's latest tool in both states and move the
completion check to the trailing edge, where the rest of the state signal already
lives.

Size both header slots to the mock's 16px slot with a 14px glyph, matching the
tool rows beneath them and the subagent summary row landing separately. They were
24/16, so the icon columns sat 8px apart and broke the left alignment the icon
treatment depends on.

The fixity test walks running, completed, and failed and pins the leading glyph
of every row by lucide's own class name, so a swap shows up as a different name
rather than a still-present icon.

* fix(codex): stop a classified shell row from asserting facts the command doesn't support

Three claims the `commandActions` row model was making on its own:

- `listFiles` with a null path was given `path: '.'`. Codex sends null for a
  recursive walk and for the repo root, and the invented path flows into
  `createToolInputDisplay().filePath`, which mobile turns into a tappable
  "open file" link onto a directory — an affordance that can only fail. The row
  now keeps the raw command, which is what the label logic already falls back to.
- A command whose actions classify as two different things (`cat a.txt && ls src`)
  was named after the first one, silently dropping the rest. Recognized actions
  must now agree on one class; a repeat of one class keeps the class and only a
  target every entry names.
- `read` lifted `name` into the journal payload, where no label ever reads it —
  `path` always wins — so it was bounded weight carrying nothing.

* fix(native-chat): give an unmodelled tool row a generic glyph, not a terminal

The row-word vocabulary named seven words, and everything else fell through to
the terminal glyph — which reads as "a shell ran here" for rows where nothing
says one did. Codex's own `apply_patch` row, `Grep`/`Glob`/`Task`/`WebFetch`/
`TodoWrite`, and every `mcp__*` tool all rendered a terminal, leaving the
declared `mcpToolCall` and `subAgentActivity` categories unreachable.

- Split the vocabulary: `unknown` stays the shell command Codex could not
  classify and keeps the terminal, while a new `other` carries the generic
  wrench that unmodelled words now fall back to.
- Read the edit family from `EDIT_TOOL_NAMES` and the command tools from
  `isCommandToolName` rather than restating either. Command tools resolve first:
  `isEditToolName` counts `shell`/`exec` as possible patch carriers, and a shell
  row is not an edit.
- Result rows get no category glyph. Their word is `translate(…, 'Result')`, so
  keying a category off it resolved a different glyph per locale; an empty slot
  keeps the rows aligned.
- The header and the row now resolve through `NativeChatToolIcon`, so one `Grep`
  run can no longer show a wrench in the header and a terminal on its line. The
  glyph map and the unused `category` prop go with the duplication.

* fix(native-chat): give the projected Diff row the file-change glyph

Every Codex fileChange item projects to a tool call named `Diff`, which the
edit set does not name — it names the tools that carry the edit in their own
input. So a run whose body renders an edited-file card was headed by the
generic wrench.

* fix(codex): stop a classified shell row offering a folder as a file to open

A listFiles action's path is a directory, and a search action's path is the
root it scanned. Lifted under `path`, both became the row's file target, which
mobile renders as a tappable open-file link that can only fail — the same dead
link the removed `{ path: '.' }` stand-in would have produced. They lift to
`directory` instead, which still labels the row but is never a file target.

* fix(mobile): keep the terminal glyph on a classified Codex shell row

Mobile's run header picks between a terminal and a generic glyph by tool
name. Now that the host publishes `read`/`search`/`list` for the same
commands it used to publish as `shell`, that name check answers false and
a command that really ran heads its run with a wrench.

Ask the shared category vocabulary instead. Mobile keeps its two icons —
porting the full glyph set is a separate lane.

* fix(native-chat): say what the run header's glyph actually guarantees

The comment claimed the header names the same tool in both states, so its
glyph cannot change on settle. It can: the live header names the running
call while the settled one names the run's last tool call, and with
out-of-order completion those differ. The glyph is fixed for whichever
tool the header names — say that, and drop the never-taken running branch
from the settled header's call.

Also pin the other half of the file-target rule: `read` keeps `path`, so
its row stays tappable, where `list`/`search` lift a folder to
`directory` and offer no target at all.

* fix(native-chat): give a rollout-transcript shell row the terminal glyph

`exec` and `local_shell` are what the Codex rollout transcript names a
shell call — `native-chat-edit-normalize` already treats those three
words as the command tools — but the activity set the glyph vocabulary
reuses carries neither, so both rows headed a real command with the
generic-tool wrench.

Named in the vocabulary rather than in that activity set, because that
set also picks the running row's copy and this is only about the glyph.

* fix(mobile): pick the run-header glyph from the call's input, not its word

Codex now names a classified shell row `read` / `search` / `list`, which
lowercase to Claude's own `Read` / `Grep` / `Glob`. Mobile has only a terminal
and a wrench, so keying that choice on the row word gave Claude's filesystem
tools a terminal for a shell that never ran.

The input separates them: Codex keeps the raw command on a classified row,
while Claude's `Read` carries only a file path. `isShellActivityToolCall`
replaces `isShellActivityToolRow` and asks the command tool names first, then
the call's input.

* fix(native-chat): give the projected diff fixture its required digest

* fix(native-chat): head a settled run with a glyph the whole run shares

The settled run header drew the glyph of the run's last tool call while the
text beside it summarizes the run's first three, so a ten-call run ending in a
`read` showed an eye above "shell npm test · shell git status · …" — a category
the summary never described.

Resolve the header's glyph from every call in the run instead: the shared
category's glyph when all agree, the generic tool glyph when the run spans
categories, and no glyph when there are no tool calls. The running header still
names the active call, whose glyph is true of it.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-05 14:03:45 -07:00
Neil 1924c8f5b1 feat(perf): lint repeated sort setup and schedule regression contracts (#18822)
* feat(perf): audit comparator setup and schedule performance contracts

* test(sqlite): close readers after expected busy failures

* ci(perf): trigger contract workflow on the contract files themselves

Without these paths a contract rename lands green on PR CI and only
breaks the next nightly, where nobody owns the failure. Also run the
OS-independent source audit once instead of on all three runners.
2026-09-05 13:56:06 -07:00
Neil d8c4c83063 test(e2e): fence SSH recovery and release exited Electron pipes (#18880) 2026-09-05 13:53:27 -07:00
Jinwoo Hong 61ebffa86e fix(runtime): bound terminal-wait blocked-prompt rules to the live screen bottom (#18817)
The Codex prompt rules in terminal-wait-detection scanned the whole retained
tail (up to 256 KiB) with lastIndexOf, so any quoted prompt phrase in
scrollback registered as a live prompt. A Codex agent working on Orca prints
rg hits from this very file; one such line ~300 lines above an idle input box
made `orca terminal send` refuse with agent_prompt_blocked and `terminal wait
--for tui-idle` report codex-interactive-prompt. `clear` did not help because
the detector reads the retained tail, not the visible screen.

A prompt that owns the terminal is at the screen bottom, so every blocked
rule now runs over the last 12 non-blank lines (real Codex dialogs are 4-8
lines), the way the cursor approval rule already was. The returned index is
offset back into full-tail coordinates so ready-header comparisons keep
working. The sentinel fast path is unchanged.

Line-window primitives move to terminal-wait-tail-window.ts to keep the
detector under the max-lines cap.
2026-09-05 16:49:59 -04:00
Neil 79350f4551 test(e2e): follow current sidebar project and activity actions (#18878)
* test(e2e): follow current sidebar project and activity actions

* test(e2e): reopen activity after revealing a workspace
2026-09-05 13:39:56 -07:00
Jinwoo Hong a3c1d32995 fix(relay-ops): per-region cell latency bar and attributable preflight failures (#18877)
The incident monitor froze three healthy 15-minute production gates on
2026-09-05 because asia-east2 cells are judged against a bar calibrated
for us-central1. A cell's /ready fetches the auth JWKS and runs SELECT 1
against Cloud SQL, both in us-central1, so from the US GitHub runner the
asia-east2 round trip measures p50 0.88 s / max 2.7 s against 0.08-0.5 s
for us-central1 cells.

Give cell.<id>.latency_ms a per-region threshold (us-central1 2000,
asia-east2 4000) carried on IncidentCellExpectation from the tfvars
region. Director and auth latency rules keep the flat 2000 bar, and hard
faults are still caught by the .health/.ready equal-1 checks and the
probe's 8 s fetch timeout.

Also name the signal and its observed/threshold in the live preflight
failure message, keeping the source/code tokens other tooling matches on.
2026-09-05 16:21:50 -04:00
Neil 9c3d957ae3 perf(automations): reuse collation setup for list name sorting (#18823) 2026-09-05 13:19:25 -07:00
Jinjing 062db77118 reduce: lower tab minimum width from 88px to 72px (#18871) 2026-09-05 12:17:40 -07:00
Merge Sim c9af4fa64b fix(codex): key the subagent label collision ordinal on what the row draws
`codexSubagentLabel` tested the trailing segment trimmed but returned it
untrimmed, and `claimLabel` keys its collision ordinal on that string. Two
children at `/root/read` and `/root/ read ` therefore both drew as `read`
with no ordinal — the one thing the ordinal exists to prevent. Return the
trimmed segment so labels that render identically collide.

Also correct the legacy-clause note on the twin recognizer. It claimed shipped
journals hold the old `— N working` sentence; the feature is unreleased, so the
only journals holding one are dev worktrees of this branch. The branch still
earns its place — those rows replay too, and it adds no false-positive surface
the bare shape does not already carry — but the stated reason was wrong.
2026-09-05 05:43:27 -07:00
Merge Sim c25c630511 fix(native-chat): stop the roster's durable twin from claiming live subagents
The spawn-group row is written once and revised in place, but the row itself
is durable and replayed on every reconnect. Its plain-text twin — the only
thing a client that cannot draw the block ever sees — froze a live count into
that row: `Kicked off 4 subagents — 2 working`. The desktop renderer never
shows it, and reconciles the block's `working` to `unverifiable` outside the
live turn. A text-only reader does neither. When the writing process dies
mid-flight the turn-end sweep never runs, so the sentence keeps asserting two
running children forever, with nothing left that could re-check them. That is
the collapse `docs/reference/ssh-execution-boundary.md` forbids: loss of
contact reported as a live state.

Fix it at the source rather than per client: the durable sentence now states
only what survives its process — that the group was spawned, plus whatever
outcome had latched. `Kicked off` vs `Ran` stays, because it reports whether an
outcome was recorded at write time; saying `Ran` while children were in flight
would assert they exited, the same error inverted. The adverse count stays so a
failing fan-out still reads as failing. Reconciliation stays in the renderer,
where the block still needs it.

The twin recognizer keeps matching the legacy `— N working` shape: journals
already hold those sentences and their rows replay forever, so dropping the
branch would print every one of them twice, once as the block and once as prose
the reader meant to drop.

Also align the two functions that read `agentPath`. The root check compared the
raw string while the label normalized separators, so `/root/` was both the turn
itself and a child of it — a phantom row labelled `root` inflating the group by
one. Compare normalized segments instead, keeping `/morpheus` a child. And a
trailing segment with nothing visible in it survives the empty-segment filter
and would draw a nameless row, so it now reads as no label and falls back to the
placeholder.
2026-09-05 05:27:59 -07:00
Merge Sim 47fa9bd720 docs(codex): justify the subagent wire notes from the live probe alone
The roster and disposition comments explained themselves in terms of a
provider-internal path taxonomy rather than anything this repo can observe.
Restate them from the evidence Orca actually has: the live app-server probe
saw `agentsStates` arrive empty, so nothing reads it; and `collabAgentToolCall`
stays substantive because nothing guarantees a session reports subagent work as
`subAgentActivity` at all — one that only emits the collab tool call gets no
roster row, and suppressing that too would leave its fan-out blank.

Same behaviour, same tests; comments and one test name only.
2026-09-05 05:08:02 -07:00
Merge Sim 42f21468c3 Merge remote-tracking branch 'origin/main' into brennanb2025/codex-subagent-worklog
# Conflicts:
#	src/renderer/src/i18n/locales/en.json
2026-09-05 05:02:37 -07:00
Merge Sim cea5af049d docs(codex): restore the roster's evictionated trigger to its KNOWN LIMITATION
The previous rewrite dropped both triggers the old comment named and kept only
the restart one, but eviction is the reachable half: `groupFor` caps `groups` at
MAX_CODEX_SUBAGENT_GROUPS and drops the oldest-INSERTED entry (it returns an
existing group without re-inserting, so this is not LRU), which can evict a
still-live group in-process. The row identity is keyed on the group id alone, so
the next activity item rebuilds that row from one child — the same N-to-1
rewrite, with no restart, and with the sweep skipped so the children never latch
`unverifiable`. Also softens "every real turn id is freshly minted" to the
provider assumption it is: turn ids are read verbatim off provider frames and
nothing in this repo mints or asserts them.
2026-09-05 04:52:42 -07:00
Merge Sim b77d615f14 test(native-chat): pin the roster twin recognizer against prose
Both readers use it to decide the twin is already printing, so a false positive
eats a message's real prose and a false negative prints the roster twice.
2026-09-05 04:38:32 -07:00
Merge Sim 07770427b8 fix(native-chat): add the subagent roster's localization keys and narrow its twin filters
The roster row called 16 `components.native-chat.subagents.*` keys that were
never added to the catalog, failing the localization gate. Synced en.json; the
English strings are the component's own inline fallbacks, so nothing renders
differently.

Also tightens the twin/block handoff on both readers. The renderer dropped
every text block once a roster was present, which is safe only because Codex
writes a roster as its own message — the block is provider-agnostic, so a lane
folding prose in beside one would have lost it on desktop while mobile kept it.
And both readers decided "the twin is already printing" by recomputing the
sentence and comparing bytes, which a roster from a newer build never matches:
its unknown state normalizes to `unverifiable` here, so the CLI printed the
roster twice with two different verdicts. Both now recognize a twin by shape.
2026-09-05 04:35:31 -07:00
Merge Sim 410f684755 fix(cli): stop worker read printing the subagent roster sentence twice
The producer ALWAYS writes a roster block beside a plain-text twin carrying the
same sentence, for clients that cannot draw the block. The renderer honours that
contract from one side — it draws the block and drops the twin. The CLI honoured
neither side: it printed the twin as prose AND rendered the block as
`[subagents] <same sentence>`, so a real roster message read

    [system] Ran 2 subagents (1 failed)
    [subagents] Ran 2 subagents (1 failed)

Take the mirror of the renderer's rule, which is the cleaner half for a text
client: the twin IS the sentence, so print it and drop the block it stands in
for. A block that arrives WITHOUT its twin — a shape the wire admits and no
producer writes — still stands in for itself, because dropping it
unconditionally would lose the roster entirely. Either way the sentence prints
exactly once, off the same `subagentGroupFallbackText` helper both sides use.

Unreachable through `readWorkerTranscript` today, whose provider rollout decoder
never emits a `subagent-group` block — but the formatter is the CLI's contract
for any transcript source, and the shape is already producible.

The test pinned a TWIN-LESS group, a body `codexSubagentGroupBody` never writes:
it asserted the exact double-print this fixes was correct output, and would have
blessed either behaviour. Rebuild the fixture as the producer's real two-block
row, with the sentence taken from the shared helper rather than hardcoded so it
cannot drift, and assert the sentence appears exactly once. The twin-less shape
keeps a test of its own, labelled as the wire-only fallback it is.

Also record why `settleTurn` keys on the RAW `turnId` while `groupFor` remaps
off-primary activity onto the primary's active turn. The asymmetry is
load-bearing, not an oversight: were `settleTurn` to remap, a child thread
ending its own turn would sweep the parent group and settle every still-working
sibling to `unverifiable`. The lookup missing is the intended no-op.
2026-09-05 04:05:35 -07:00
Merge Sim 6196fc4594 fix(native-chat): make "counts as renderable" and "actually draws" agree for a spawn group
`MessageRow` counts any `subagent-group` block as renderable, but
`NativeChatSubagentRun` renders null for a childless roster. A group with
`agents: []` therefore mounted a row that drew nothing — an empty div that still
costs the transcript one `gap-5` slot. The Codex producer never writes one (every
`write()` call site operates on a group that already holds an entry), but the
block schema admits `agents: []` with no `.min(1)`, and the wire is where such a
shape would arrive.

Narrow `subagentGroupBlocks` — whose only production caller IS that renderable
check — to the groups that will draw, behind a named `isRenderableSubagentGroup`
that `NativeChatToolRun` now shares in place of its own copy of the predicate, so
the two guards cannot drift apart again. A childless group carrying its
plain-text twin now prints the twin, which is what the twin is for; a bare one
skips the row entirely.

Also correct four comments that had stopped describing the code:
  - the roster header called `agentsStates` "always empty", contradicting the
    probe note in `codex-subagent-activity.ts` — it is empty on the MultiAgentV2
    path that emits these items, and the V1 path does populate it;
  - `tokensByThread` was documented "retained UNCONDITIONALLY" while
    `handleTokenUsage` LRU-caps it 65 lines below;
  - the sweep is not "the LAST event a group ever gets": neither `settleTurn` nor
    `settleSession` removes the group, so a later `thread/tokenUsage/updated`
    naming a swept child still writes it. The retry condition is right; only its
    stated reason was wrong;
  - the `subAgentActivity` classification is not reached "for every event — and
    every one of them arrives twice". `handleSubagentItem` intercepts those items
    before `items.handle`, so the live path never consults the catalog;
    `restoreThread` replays them straight through, and is the real consumer.

Comment-only apart from the childless-group guard.
2026-09-05 03:55:34 -07:00
Neil af82126058 fix(native-chat): give the Claude exit barrier a handle on unpublished exits (#18826)
A first-hand Claude exit is not published where it is observed. `handleExit`
re-enters the close ladder and persists the transcript cursor before it emits
`ended`, and only that emission reaches the runtime's recovery chain. So the
runtime's `waitForRecovery` — whose whole job is to drain an in-flight recovery
before teardown stops children — returns immediately for an exit that is still
climbing the ladder, and nothing outside the adapter can tell an observed exit
from a published one.

The integration test for fenced host reconciliation had no handle on that
barrier, so it bounded-polled the lease for 100ms instead. Measured under 16x
local concurrency, publication alone takes 77-204ms: 19/24 runs failed.

Retain the ladder-then-settle tail on the exit record and expose
`drainObservedExits`, fold it into `waitForRecovery`, and export the barrier so
a caller that needs the settled lease can await it. Codex publishes inside its
own exit callback and needs nothing. The test now awaits the barrier: 0/24
under the same load, and it fails on an idle machine without the drain.
2026-09-05 03:50:19 -07:00
Merge Sim 9326b401c6 test(native-chat): cover the subagent roster at the message-list level
Every defect this feature has shipped so far lived in the assembly between
rows, and the row-level suites kept passing through all of them. Loop 4's
regression — a settled roster swallowed by the completed-turn disclosure —
was found by reading the code, not by a test, and an independent visual-proof
run observed the same symptom in the real UI and routed around it rather than
reporting it. `NativeChatToolRun` rendered alone is handed `expandOverride`
and `activeTurnIsWorking` by the test author, so it agrees with whatever the
caller was assumed to pass.

Drive the real component instead. The roster is its own `role: 'system'`
journal row carrying the producer's two blocks (structured + plain-text twin),
so what reaches the DOM depends on `foldToolMessages`, the turn-key mapping
and the disclosure state `NativeChatMessageList` owns — none of which a row
test exercises.

Three cases, on one assembled transcript that holds tool calls AND a roster:
  - a settled turn with activity collapsed, the resting state of the whole
    transcript, still shows the row (fails with loop 4's reorder reverted);
  - tool activity stays behind that disclosure and appears only on expand,
    and expanding draws no second roster (fails with the guard removed);
  - a working turn reads as a live spawn.

The first also pins that the plain-text twin is dropped rather than printed
beside the row it stands in for.

Timestamps are explicit and ascending: the list re-sorts by (timestamp, id),
so rows sharing a millisecond tie-break alphabetically and the user turn can
sort last, stranding the roster outside its own turn and reconciling live
children to `unverifiable`.

No production code changed.
2026-09-05 03:29:43 -07:00
Neil 265871c53d fix(native-chat): stop a settling handoff throwing an unhandled rejection at teardown (#18824)
* fix: stop a handoff flow from outliving the host that owns its session

A structured handoff runs on the session's serialized chain and nothing in
production awaited it. When the client-side deadline for the switch expired
first, teardown dropped the session map out from under a live flow, and the
flow's own failure notification then threw `agent_session_ownership_unknown`
out of a status publish — an unhandled rejection, plus journal rows written
into a directory that was already being removed.

Three fixes, each with a regression test that fails without it:

- The status publish is a notification, not a mutation: it now reads the fence
  without requiring an attached session, so an evicted or torn-down session
  makes it a no-op instead of a throw.
- `track` used `.finally`, which forwards a rejection onto a promise nobody
  awaits. `drain` settles flows through `allSettled`, so the bookkeeping chain
  is now settle-only and cannot resurface one.
- Host teardown drains in-flight handoffs before dropping the session map.
  `drain` existed for exactly this and was never wired up.

The integration test's `vi.waitFor` is dropped rather than widened: the request
enqueues the flow on the session's serialized chain before it returns, so the
status read is already ordered behind it. The poll only added a wall-clock
deadline that a loaded runner missed.

* fix: bound the handoff drain so a wedged flow cannot hold the quit open
2026-09-05 03:15:17 -07:00
Neil a823f97d63 perf(worktree): skip remote probes with no possible result (#18821) 2026-09-05 03:01:12 -07:00
Merge Sim 1333d95280 fix(native-chat): stop the subagent roster vanishing from every settled turn
`NativeChatToolRun` bailed out for a completed turn whose activity disclosure
is collapsed before it reached the branch that draws a roster-only run. That
guard exists to push TOOL activity behind the turn-status disclosure, and it
fires on exactly the shape a spawn group has: a roster message carries no tool
blocks, so `selectActiveToolCall` returns null and `isSettled` is true, while
the list passes `expandOverride={expandedTurnIds.has(turnKey)}` — false until
the reader opens that turn — and `activeTurnIsWorking={false}`.

That is the default state of every finished turn in the transcript, so the one
compact row this feature exists to leave behind ("Ran 3 subagents") disappeared
the moment its turn ended. Worse, `MessageRow` counts a spawn group as
renderable specifically so the row survives, then rendered a wrapper around a
component that returned null — the empty ghost bubble its own guard is written
to prevent.

Order the roster branch before the disclosure guard. A roster has no tool
activity to hide, and the guard's reasoning ("a failed child command looked
like the whole response was still running") does not reach it. Runs that do
carry tool blocks still fall through to the guard unchanged, and in practice a
roster never shares a message with them: it is its own `role: 'system'` journal
row and `isToolOnlyMessage` is false for it, so `foldToolMessages` never merges
tool blocks into it.

Also drop childless groups when building the rows, so `subagentRows.length`
stays an honest test of "something will draw" — the roster-only branch returns
a margin-bearing wrapper on the strength of it, and a group with no children
renders null.

Both tests fail with their fix reverted; the existing NativeChatToolRun suite
still passes, so the completed-turn disclosure behaviour is unchanged.
2026-09-05 02:59:57 -07:00
Jinwoo Hong 9f2a9a248e fix(cloud): validate protocol-0 same-cap cell plans without rehome trust lines (#18818) 2026-09-05 05:42:37 -04:00
Merge Sim fd2d2b7b0d fix(native-chat): retry a refused roster publish, and stop two wrong readings
Four defects from a third review pass over the Codex subagent roster.

`write()` set `lastSerialized` before the append and rolled it back only when
the APPEND was refused. A refused PUBLISH left it set, so an identical replay
short-circuited and the revision was never published again. The repo's own
pattern is the opposite: `codex-structured-item-streams.ts` advances
`checkpointLengths` only once the append AND the publish are both accepted.
Roll back on either half.

That alone did not cover the sweep, which is the LAST event a group ever gets:
its `changed` guard skips the write on a retry because every child has already
latched, stranding the settled roster's final revision. Write when the previous
attempt was refused part-way, too.

`formatWorkerTranscriptMessage` read `block.agents` as its exhaustive fallback.
The journal schema deliberately admits block types this build does not know and
`client.call` casts the RPC result instead of validating it, so a newer remote
host's block reached that line and threw `agents is not iterable`, taking down
the whole `worker read`. It printed a harmless `[image omitted]` before. Match
`subagent-group` explicitly and degrade the unknown case.

The elapsed clock measured to `now` whenever no child carried a terminal
timestamp. That is exactly the roster restored from the journal after the host
died: the reconciler latches `unverifiable` without a `settledAt`, so a child
that ran four seconds reported the time since the crash as its run length, on a
row that is not even counting. Show no duration when none is known.

Also restores package.json to origin/main: the merge had deleted one of main's
two duplicate `bench:terminal-partial-escape-tail` keys. Behaviour-preserving
(JSON is last-wins and the deleted line was the dead one), but unrelated to this
PR and better left to its own change. No gate rejects duplicate JSON keys.

The new refusal tests also cover the append-side rollback, which had none.
2026-09-05 02:39:57 -07:00
Soshiro Narita 1a76a11e39 feat(i18n): localize onboarding flow to Japanese (#18787) 2026-09-05 02:37:06 -07:00
Neil 0bbf86bb7a fix(renderer): stop a Node-only process-table module blanking the app at boot (#18814)
Every Electron E2E spec that boots the app has been failing on
`workspaceSessionReady did not become true`, and the app itself has been
launching to a blank white window: the renderer threw
`ReferenceError: process is not defined` while evaluating a shared chunk,
so React never mounted and no startup step ever ran.

`agent-completion-poll-interval.ts` (renderer) imported one constant,
`PROCESS_TABLE_SNAPSHOT_MAX_STALENESS_MS`, out of
`shared/process-table-snapshot-reader.ts` — a `node:child_process` /
`node:fs/promises` module whose dependency evaluates `process.platform` at
module scope to pick `ps` columns. The renderer runs sandboxed with
contextIsolation, where `process` is undefined, so that module-scope read
threw and took the whole chunk with it. Introduced by #18742; #18780 added a
second module-scope read next to the first.

The constant now lives in `shared/process-table-snapshot.ts`, the
environment-neutral half of the pair, and the reader re-exports it so host
callers are unchanged. The two `ps` column sets read the platform behind a
`typeof process` guard, which defuses the same landmine for any future
renderer import of that module — only hosts ever run the argv.

The regression test walks the renderer import graph (lazy routes included)
from all three entries and refuses any module that reaches a `node:` builtin.
It fails on the pre-fix import with the full 10-hop chain from `main.tsx`.
2026-09-05 02:07:37 -07:00
Merge Sim 3f880d493c fix(native-chat): stop the subagent roster announcing a new duration every second
The roster row is an `aria-live="polite"` region and it contains the elapsed
clock, which reticks once a second for as long as the fan-out runs. A screen
reader therefore reads out a fresh duration every second, burying the state
changes the live region exists to report — the headline, the verdict, and the
`+1 failed` alert.

No other live region in the transcript does this. `NativeChatToolRun`'s live
button holds only the active tool label, and in `NativeChatWorkingStatus` the
variant that shows a duration is precisely the one with no `aria-live`.

Hide the clock from the accessibility tree only while it is moving. Once the
group settles the duration is fixed, so it stays readable and costs no
announcements.
2026-09-05 01:54:32 -07:00
Jinwoo Hong 12e05203a4 fix(cloud): let the same-cap roll isolate Asia cells (#18811)
The same-cap wave validator approves 19 cells (c7-c26 plus the Asia cells
c27-c29), but the canary script it drives hard-rejected anything outside the
16 US capacity cells, so the first Asia same-cap canary failed closed at
isolate. Give the canary an explicit --approved-cells switch that selects the
same-cap allowlist, and pass it from the four same-cap job invocations. With no
switch the behaviour is unchanged, so the US-only capacity workflow keeps its
scope.
2026-09-05 01:51:06 -07:00
Neil 06ca54ae7c fix(test): stop detached git maintenance racing the divergence fixture teardown (#18810)
`worktree-base-divergence-real-git.test.ts` builds cap-sized histories (100 and
101 commits). Every `git commit` detaches `git maintenance run --auto`, whose
commit-graph task arms at 100 new commits, so the fixture reliably spawns a
background `git commit-graph write --split` that keeps creating
`.git/objects/info/commit-graphs` entries after the synchronous exec returns.
The `afterEach` recursive remove is then deleting `.git/objects` underneath a
live writer and dies with ENOTEMPTY — which is how "counts drift in both
directions" failed on main.

Reuse the existing `GIT_FETCH_SKIP_AUTO_MAINTENANCE_CONFIG_ARGS` (it already
covers modern maintenance and legacy auto-gc, so it holds at the Git 2.25
baseline) in the fixture's git helper. Traced spawns of
`git commit-graph write` over a full run of this file: 4 before, 0 after.

The production path under test only runs `rev-list` and `merge-base`, neither of
which triggers auto-maintenance, so there is nothing to fix outside the fixture.
2026-09-05 01:49:48 -07:00
Brennan BensonandMerge Sim cec26336e5 fix(native-chat): say when a structured launch fell back to a terminal (#18762)
* fix(native-chat): say when a structured launch fell back to a terminal

A definitive refusal already opened a terminal instead of the requested
structured chat, but said nothing — indistinguishable from the bug where the
wrong surface opens. Notify at message severity, since nothing failed.

Also stop putting the raw error in the failure toast's description: it carried
errnos and absolute paths straight into the UI. The detail moves to a warn log
and the toast gets catalog copy, matching how the coded refusals already read.

* fix(native-chat): avoid overstating terminal fallback

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-05 01:47:43 -07:00
Merge Sim c96dd59904 fix(native-chat): correct the Codex subagent roster's build, journal write, and failure reporting
* Restore the exhaustive block handling that adding `subagent-group` to
  `NativeChatBlock` broke. `formatWorkerTranscriptMessage` and `boundBlock`
  both fell through to `image-ref` field access, so `tsc -p` failed for the
  CLI and node projects and `build:cli` could not emit. Both now guard on
  `image-ref` explicitly and give the roster block its own branch.

* Stop the roster's publish from evicting its own append. The sink queue
  coalesces by `coalescingKey` alone with no op-kind check, so passing the
  append's key to `tryPublish` spliced the queued append out and the row
  never reached the journal — permanently, since `lastSerialized` was
  already set. `tryPublish()` now takes no argument, matching every other
  call site. The regression test's fake sink honours the key, which the
  previous fake did not.

* Keep `collabAgentToolCall` substantive. Only the MultiAgentV2 path emits
  `subAgentActivity`, so a V1 turn has no roster row; suppressing its collab
  tool calls too would have left a V1 fan-out showing nothing at all.

* Surface a settled failure while siblings still work. The summary now
  reports the worst adverse outcome independently of the group verdict, so
  the row shows `3 working +1 failed` with a failed-coloured dot instead of
  a neutral pulsing dot. The plain-text twin names it too.

* Treat `/morpheus` as a child. Only `/root` is the turn itself; the old
  segment-count test silently dropped a valid single-segment agent.

* Refresh token-usage recency on update so an active thread is not evicted
  as the oldest entry, and scope the `agentsStates` comment to the V2 path.
2026-09-05 01:38:41 -07:00
Neil 9927edd631 fix(ssh): let the host say whether it armed the ready marker (#18802)
#18796 made every SSH Codex background launch wait for the shell-ready marker,
but the client cannot see the remote shell. On a host that never publishes one --
fish, sh, Windows, or a relay predating #18796 -- no marker arrives and delivery
falls back at 1.5s where it used to write at 50ms.

The relay already computes whether it armed the marker; publish that as an
optional `shellReadyArmed` on the spawn reply and let the client skip a wait it
now knows is pointless. Absent stays UNKNOWN and keeps the client's own guess, so
an older host behaves exactly as before; false is only ever an answer a host gave.
It rides every reply, false included, or absent would stop meaning "old host".

A host that did not arm the marker did not arm bracketed paste either, so the
released path still submits raw.
2026-09-05 01:34:38 -07:00
Merge Sim 05e4fa037d Merge remote-tracking branch 'origin/main' into brennanb2025/codex-subagent-worklog
# Conflicts:
#	src/renderer/src/components/native-chat/NativeChatMessageList.tsx
#	src/renderer/src/components/native-chat/NativeChatMessageRow.tsx
2026-09-05 01:27:22 -07:00
Neil e95d247be1 perf(terminal): cheap-tier process inspection for anchored local agent panes (#18780)
* perf(terminal): cheap-tier process inspection for anchored local agent panes

Every idle local pane's completion cadence ran a full whole-host `ps` (with
`tty=` and `command=`, 0.34-0.50s on a 1,900-process Mac, 1.15s on Linux)
purely to build `foregroundProcessEvidence` that the renderer then discards
for local ids. Add a cheap tier (same job-control columns, no tty/command,
0.03s) gated so that it introduces no user-facing trade-off:

- Only a pane whose last FULL capture proved a recognized agent may take the
  cheap tier. Panes with no anchor always take the full capture, so start
  discovery keeps today's exact behaviour.
- The cheap tick compares a per-pane fingerprint (root shell pid+start, tpgid,
  every descendant's pid+start+pgid+job-control state). Any change, a changed
  node-pty foreground name, an unreadable capture, or an incarnation mismatch
  escalates to the full capture. A recognized agent's exit is always a pid
  vanishing, which the fingerprint always sees.
- A cheap answer OMITS evidence rather than fabricating a tty-less fence.
  Remote/restore consumers never send `steadyState`, so they keep the full
  capture unchanged.
- `steadyState` is a new optional request field; an old daemon ignores it and
  answers with the full capture.

Measured (8 idle panes, 60s, idle cadence, forks counted by column set):
30 full -> 1 full + 29 cheap.

* fix(terminal): route the cheap ps capture through runProcess

The cheap-tier reader imported node:child_process directly, which the
child-process import-boundary and windowsHide ratchet tests reject (CI shards
1/8 and 3/8). Use Orca's single spawn entry point instead; it pins windowsHide
and encodes argv. Map its result onto the capture-error vocabulary:
outputTruncated -> capture_truncated, timedOut -> capture_timeout, non-zero
exit -> ps_exit_<code>. Tests mock at the runProcess seam.

* fix(perf): refuse a pane fingerprint when any descendant start marker is missing

`buildPaneProcessFingerprint` rejected only a missing root start marker; a missing descendant
marker was stamped as `?`. Two captures that both failed to read the same descendant therefore
compared equal, which removes the pid-reuse protection the fingerprint exists to provide: a
recycled pid could make a vanished agent look unchanged, and the cheap tier would keep serving
its name instead of escalating.

Reachable on Linux, where `readLinuxProcStartTime` legitimately returns null when a process
exits between the `ps` capture and the `/proc/<pid>/stat` read.

Every subtree member now needs a start marker or the fingerprint is refused, which sends the
caller to the full capture — the same conservative default every other uncertain path takes.

Reported by CodeRabbit on #18780. The two new tests fail against the previous code with
`expected '4242@2400#4300:|4300@?:4300:+' to be null`.
2026-09-05 00:57:02 -07:00
Neil ccf3e27800 perf(relay): bound the symlink directory probes a remote readDir fans out (#18752)
`readRelayDir` issued one `stat` per symlinked entry and awaited them all in a single
`Promise.all`. A pnpm `node_modules` is hundreds-to-thousands of package symlinks in one
directory, so expanding it over SSH put that many stats in flight at once, saturating libuv's
four-thread pool and delaying every other relay filesystem operation — including the
interactive reads `fs-list-files-scan-coordinator` exists to protect.

The probes now run through `forEachWithConcurrency` at 8, the cap every other bounded probe in
this codebase already uses (`GIT_COMMON_SNAPSHOT_CONCURRENCY`,
`PRUNABLE_EXISTENCE_PROBE_CONCURRENCY`, `SPARSE_CHECKOUT_DETECTION_CONCURRENCY`).

Results and ordering are unchanged: every symlink still resolves to its target's kind, and
`sortDirEntries` still runs afterwards. The new test builds a 60-symlink directory and asserts
the same 60 stats happen with exactly 8 in flight at peak — the probes overlap, and never past the cap.
2026-09-05 00:56:50 -07:00
Brennan BensonandMerge Sim cc07249e78 fix(agent-session): refuse a pre-commit structured create with an envelope (#18697)
* fix(agent-session): refuse a pre-commit structured create with an envelope

The create route refused by throwing, which reaches a client as a generic
transport error indistinguishable from a lost answer — so desktop parked the
launch as visibility-unknown with no chat and no terminal. Convert the whole
pre-commit span, everything before `attach`, into a refusal envelope carrying a
code, and name the definitive-refusal allowlist the fallback decision needs.

* fix(agent-session): gate legacy fallback on definitive refusals

* fix(mobile): preserve unknown structured create outcomes

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-05 00:50:37 -07:00
Merge Sim daa105e83c feat(native-chat): give the subagent summary row its bot glyph
The row led with a glyph that swapped on state — a check once every child
completed, a group icon otherwise — so a group appeared to change identity
the moment it settled. Per the approved mock, the glyph names the category
and never moves: state is carried by the status dot and the tone of the
words beside it.

Use lucide `bot`, the same glyph the individual `subAgentActivity` rows take
in the eight-category vocabulary, so the summary reads as their parent. Slot
and glyph are the mock's 16px/14px, muted by default, and the svg is
`aria-hidden` — the headline is what a screen reader announces, so the icon
never stands alone.
2026-09-05 00:36:47 -07:00