mirror of
https://github.com/stablyai/orca.git
synced 2026-09-30 00:03:15 +00:00
47fa9bd720fbdc03f5adebd404e86c52baca72ab
10213
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
47fa9bd720 |
docs(codex): justify the subagent wire notes from the live probe alone
The roster and disposition comments explained themselves in terms of a provider-internal path taxonomy rather than anything this repo can observe. Restate them from the evidence Orca actually has: the live app-server probe saw `agentsStates` arrive empty, so nothing reads it; and `collabAgentToolCall` stays substantive because nothing guarantees a session reports subagent work as `subAgentActivity` at all — one that only emits the collab tool call gets no roster row, and suppressing that too would leave its fan-out blank. Same behaviour, same tests; comments and one test name only. |
||
|
|
42f21468c3 |
Merge remote-tracking branch 'origin/main' into brennanb2025/codex-subagent-worklog
# Conflicts: # src/renderer/src/i18n/locales/en.json |
||
|
|
cea5af049d |
docs(codex): restore the roster's evictionated trigger to its KNOWN LIMITATION
The previous rewrite dropped both triggers the old comment named and kept only the restart one, but eviction is the reachable half: `groupFor` caps `groups` at MAX_CODEX_SUBAGENT_GROUPS and drops the oldest-INSERTED entry (it returns an existing group without re-inserting, so this is not LRU), which can evict a still-live group in-process. The row identity is keyed on the group id alone, so the next activity item rebuilds that row from one child — the same N-to-1 rewrite, with no restart, and with the sweep skipped so the children never latch `unverifiable`. Also softens "every real turn id is freshly minted" to the provider assumption it is: turn ids are read verbatim off provider frames and nothing in this repo mints or asserts them. |
||
|
|
b77d615f14 |
test(native-chat): pin the roster twin recognizer against prose
Both readers use it to decide the twin is already printing, so a false positive eats a message's real prose and a false negative prints the roster twice. |
||
|
|
07770427b8 |
fix(native-chat): add the subagent roster's localization keys and narrow its twin filters
The roster row called 16 `components.native-chat.subagents.*` keys that were never added to the catalog, failing the localization gate. Synced en.json; the English strings are the component's own inline fallbacks, so nothing renders differently. Also tightens the twin/block handoff on both readers. The renderer dropped every text block once a roster was present, which is safe only because Codex writes a roster as its own message — the block is provider-agnostic, so a lane folding prose in beside one would have lost it on desktop while mobile kept it. And both readers decided "the twin is already printing" by recomputing the sentence and comparing bytes, which a roster from a newer build never matches: its unknown state normalizes to `unverifiable` here, so the CLI printed the roster twice with two different verdicts. Both now recognize a twin by shape. |
||
|
|
410f684755 |
fix(cli): stop worker read printing the subagent roster sentence twice
The producer ALWAYS writes a roster block beside a plain-text twin carrying the
same sentence, for clients that cannot draw the block. The renderer honours that
contract from one side — it draws the block and drops the twin. The CLI honoured
neither side: it printed the twin as prose AND rendered the block as
`[subagents] <same sentence>`, so a real roster message read
[system] Ran 2 subagents (1 failed)
[subagents] Ran 2 subagents (1 failed)
Take the mirror of the renderer's rule, which is the cleaner half for a text
client: the twin IS the sentence, so print it and drop the block it stands in
for. A block that arrives WITHOUT its twin — a shape the wire admits and no
producer writes — still stands in for itself, because dropping it
unconditionally would lose the roster entirely. Either way the sentence prints
exactly once, off the same `subagentGroupFallbackText` helper both sides use.
Unreachable through `readWorkerTranscript` today, whose provider rollout decoder
never emits a `subagent-group` block — but the formatter is the CLI's contract
for any transcript source, and the shape is already producible.
The test pinned a TWIN-LESS group, a body `codexSubagentGroupBody` never writes:
it asserted the exact double-print this fixes was correct output, and would have
blessed either behaviour. Rebuild the fixture as the producer's real two-block
row, with the sentence taken from the shared helper rather than hardcoded so it
cannot drift, and assert the sentence appears exactly once. The twin-less shape
keeps a test of its own, labelled as the wire-only fallback it is.
Also record why `settleTurn` keys on the RAW `turnId` while `groupFor` remaps
off-primary activity onto the primary's active turn. The asymmetry is
load-bearing, not an oversight: were `settleTurn` to remap, a child thread
ending its own turn would sweep the parent group and settle every still-working
sibling to `unverifiable`. The lookup missing is the intended no-op.
|
||
|
|
6196fc4594 |
fix(native-chat): make "counts as renderable" and "actually draws" agree for a spawn group
`MessageRow` counts any `subagent-group` block as renderable, but
`NativeChatSubagentRun` renders null for a childless roster. A group with
`agents: []` therefore mounted a row that drew nothing — an empty div that still
costs the transcript one `gap-5` slot. The Codex producer never writes one (every
`write()` call site operates on a group that already holds an entry), but the
block schema admits `agents: []` with no `.min(1)`, and the wire is where such a
shape would arrive.
Narrow `subagentGroupBlocks` — whose only production caller IS that renderable
check — to the groups that will draw, behind a named `isRenderableSubagentGroup`
that `NativeChatToolRun` now shares in place of its own copy of the predicate, so
the two guards cannot drift apart again. A childless group carrying its
plain-text twin now prints the twin, which is what the twin is for; a bare one
skips the row entirely.
Also correct four comments that had stopped describing the code:
- the roster header called `agentsStates` "always empty", contradicting the
probe note in `codex-subagent-activity.ts` — it is empty on the MultiAgentV2
path that emits these items, and the V1 path does populate it;
- `tokensByThread` was documented "retained UNCONDITIONALLY" while
`handleTokenUsage` LRU-caps it 65 lines below;
- the sweep is not "the LAST event a group ever gets": neither `settleTurn` nor
`settleSession` removes the group, so a later `thread/tokenUsage/updated`
naming a swept child still writes it. The retry condition is right; only its
stated reason was wrong;
- the `subAgentActivity` classification is not reached "for every event — and
every one of them arrives twice". `handleSubagentItem` intercepts those items
before `items.handle`, so the live path never consults the catalog;
`restoreThread` replays them straight through, and is the real consumer.
Comment-only apart from the childless-group guard.
|
||
|
|
af82126058 |
fix(native-chat): give the Claude exit barrier a handle on unpublished exits (#18826)
A first-hand Claude exit is not published where it is observed. `handleExit` re-enters the close ladder and persists the transcript cursor before it emits `ended`, and only that emission reaches the runtime's recovery chain. So the runtime's `waitForRecovery` — whose whole job is to drain an in-flight recovery before teardown stops children — returns immediately for an exit that is still climbing the ladder, and nothing outside the adapter can tell an observed exit from a published one. The integration test for fenced host reconciliation had no handle on that barrier, so it bounded-polled the lease for 100ms instead. Measured under 16x local concurrency, publication alone takes 77-204ms: 19/24 runs failed. Retain the ladder-then-settle tail on the exit record and expose `drainObservedExits`, fold it into `waitForRecovery`, and export the barrier so a caller that needs the settled lease can await it. Codex publishes inside its own exit callback and needs nothing. The test now awaits the barrier: 0/24 under the same load, and it fails on an idle machine without the drain. |
||
|
|
9326b401c6 |
test(native-chat): cover the subagent roster at the message-list level
Every defect this feature has shipped so far lived in the assembly between
rows, and the row-level suites kept passing through all of them. Loop 4's
regression — a settled roster swallowed by the completed-turn disclosure —
was found by reading the code, not by a test, and an independent visual-proof
run observed the same symptom in the real UI and routed around it rather than
reporting it. `NativeChatToolRun` rendered alone is handed `expandOverride`
and `activeTurnIsWorking` by the test author, so it agrees with whatever the
caller was assumed to pass.
Drive the real component instead. The roster is its own `role: 'system'`
journal row carrying the producer's two blocks (structured + plain-text twin),
so what reaches the DOM depends on `foldToolMessages`, the turn-key mapping
and the disclosure state `NativeChatMessageList` owns — none of which a row
test exercises.
Three cases, on one assembled transcript that holds tool calls AND a roster:
- a settled turn with activity collapsed, the resting state of the whole
transcript, still shows the row (fails with loop 4's reorder reverted);
- tool activity stays behind that disclosure and appears only on expand,
and expanding draws no second roster (fails with the guard removed);
- a working turn reads as a live spawn.
The first also pins that the plain-text twin is dropped rather than printed
beside the row it stands in for.
Timestamps are explicit and ascending: the list re-sorts by (timestamp, id),
so rows sharing a millisecond tie-break alphabetically and the user turn can
sort last, stranding the roster outside its own turn and reconciling live
children to `unverifiable`.
No production code changed.
|
||
|
|
265871c53d |
fix(native-chat): stop a settling handoff throwing an unhandled rejection at teardown (#18824)
* fix: stop a handoff flow from outliving the host that owns its session A structured handoff runs on the session's serialized chain and nothing in production awaited it. When the client-side deadline for the switch expired first, teardown dropped the session map out from under a live flow, and the flow's own failure notification then threw `agent_session_ownership_unknown` out of a status publish — an unhandled rejection, plus journal rows written into a directory that was already being removed. Three fixes, each with a regression test that fails without it: - The status publish is a notification, not a mutation: it now reads the fence without requiring an attached session, so an evicted or torn-down session makes it a no-op instead of a throw. - `track` used `.finally`, which forwards a rejection onto a promise nobody awaits. `drain` settles flows through `allSettled`, so the bookkeeping chain is now settle-only and cannot resurface one. - Host teardown drains in-flight handoffs before dropping the session map. `drain` existed for exactly this and was never wired up. The integration test's `vi.waitFor` is dropped rather than widened: the request enqueues the flow on the session's serialized chain before it returns, so the status read is already ordered behind it. The poll only added a wall-clock deadline that a loaded runner missed. * fix: bound the handoff drain so a wedged flow cannot hold the quit open |
||
|
|
a823f97d63 | perf(worktree): skip remote probes with no possible result (#18821) | ||
|
|
1333d95280 |
fix(native-chat): stop the subagent roster vanishing from every settled turn
`NativeChatToolRun` bailed out for a completed turn whose activity disclosure
is collapsed before it reached the branch that draws a roster-only run. That
guard exists to push TOOL activity behind the turn-status disclosure, and it
fires on exactly the shape a spawn group has: a roster message carries no tool
blocks, so `selectActiveToolCall` returns null and `isSettled` is true, while
the list passes `expandOverride={expandedTurnIds.has(turnKey)}` — false until
the reader opens that turn — and `activeTurnIsWorking={false}`.
That is the default state of every finished turn in the transcript, so the one
compact row this feature exists to leave behind ("Ran 3 subagents") disappeared
the moment its turn ended. Worse, `MessageRow` counts a spawn group as
renderable specifically so the row survives, then rendered a wrapper around a
component that returned null — the empty ghost bubble its own guard is written
to prevent.
Order the roster branch before the disclosure guard. A roster has no tool
activity to hide, and the guard's reasoning ("a failed child command looked
like the whole response was still running") does not reach it. Runs that do
carry tool blocks still fall through to the guard unchanged, and in practice a
roster never shares a message with them: it is its own `role: 'system'` journal
row and `isToolOnlyMessage` is false for it, so `foldToolMessages` never merges
tool blocks into it.
Also drop childless groups when building the rows, so `subagentRows.length`
stays an honest test of "something will draw" — the roster-only branch returns
a margin-bearing wrapper on the strength of it, and a group with no children
renders null.
Both tests fail with their fix reverted; the existing NativeChatToolRun suite
still passes, so the completed-turn disclosure behaviour is unchanged.
|
||
|
|
9f2a9a248e | fix(cloud): validate protocol-0 same-cap cell plans without rehome trust lines (#18818) | ||
|
|
fd2d2b7b0d |
fix(native-chat): retry a refused roster publish, and stop two wrong readings
Four defects from a third review pass over the Codex subagent roster. `write()` set `lastSerialized` before the append and rolled it back only when the APPEND was refused. A refused PUBLISH left it set, so an identical replay short-circuited and the revision was never published again. The repo's own pattern is the opposite: `codex-structured-item-streams.ts` advances `checkpointLengths` only once the append AND the publish are both accepted. Roll back on either half. That alone did not cover the sweep, which is the LAST event a group ever gets: its `changed` guard skips the write on a retry because every child has already latched, stranding the settled roster's final revision. Write when the previous attempt was refused part-way, too. `formatWorkerTranscriptMessage` read `block.agents` as its exhaustive fallback. The journal schema deliberately admits block types this build does not know and `client.call` casts the RPC result instead of validating it, so a newer remote host's block reached that line and threw `agents is not iterable`, taking down the whole `worker read`. It printed a harmless `[image omitted]` before. Match `subagent-group` explicitly and degrade the unknown case. The elapsed clock measured to `now` whenever no child carried a terminal timestamp. That is exactly the roster restored from the journal after the host died: the reconciler latches `unverifiable` without a `settledAt`, so a child that ran four seconds reported the time since the crash as its run length, on a row that is not even counting. Show no duration when none is known. Also restores package.json to origin/main: the merge had deleted one of main's two duplicate `bench:terminal-partial-escape-tail` keys. Behaviour-preserving (JSON is last-wins and the deleted line was the dead one), but unrelated to this PR and better left to its own change. No gate rejects duplicate JSON keys. The new refusal tests also cover the append-side rollback, which had none. |
||
|
|
1a76a11e39 | feat(i18n): localize onboarding flow to Japanese (#18787) | ||
|
|
0bbf86bb7a |
fix(renderer): stop a Node-only process-table module blanking the app at boot (#18814)
Every Electron E2E spec that boots the app has been failing on `workspaceSessionReady did not become true`, and the app itself has been launching to a blank white window: the renderer threw `ReferenceError: process is not defined` while evaluating a shared chunk, so React never mounted and no startup step ever ran. `agent-completion-poll-interval.ts` (renderer) imported one constant, `PROCESS_TABLE_SNAPSHOT_MAX_STALENESS_MS`, out of `shared/process-table-snapshot-reader.ts` — a `node:child_process` / `node:fs/promises` module whose dependency evaluates `process.platform` at module scope to pick `ps` columns. The renderer runs sandboxed with contextIsolation, where `process` is undefined, so that module-scope read threw and took the whole chunk with it. Introduced by #18742; #18780 added a second module-scope read next to the first. The constant now lives in `shared/process-table-snapshot.ts`, the environment-neutral half of the pair, and the reader re-exports it so host callers are unchanged. The two `ps` column sets read the platform behind a `typeof process` guard, which defuses the same landmine for any future renderer import of that module — only hosts ever run the argv. The regression test walks the renderer import graph (lazy routes included) from all three entries and refuses any module that reaches a `node:` builtin. It fails on the pre-fix import with the full 10-hop chain from `main.tsx`. |
||
|
|
3f880d493c |
fix(native-chat): stop the subagent roster announcing a new duration every second
The roster row is an `aria-live="polite"` region and it contains the elapsed clock, which reticks once a second for as long as the fan-out runs. A screen reader therefore reads out a fresh duration every second, burying the state changes the live region exists to report — the headline, the verdict, and the `+1 failed` alert. No other live region in the transcript does this. `NativeChatToolRun`'s live button holds only the active tool label, and in `NativeChatWorkingStatus` the variant that shows a duration is precisely the one with no `aria-live`. Hide the clock from the accessibility tree only while it is moving. Once the group settles the duration is fixed, so it stays readable and costs no announcements. |
||
|
|
12e05203a4 |
fix(cloud): let the same-cap roll isolate Asia cells (#18811)
The same-cap wave validator approves 19 cells (c7-c26 plus the Asia cells c27-c29), but the canary script it drives hard-rejected anything outside the 16 US capacity cells, so the first Asia same-cap canary failed closed at isolate. Give the canary an explicit --approved-cells switch that selects the same-cap allowlist, and pass it from the four same-cap job invocations. With no switch the behaviour is unchanged, so the US-only capacity workflow keeps its scope. |
||
|
|
06ca54ae7c |
fix(test): stop detached git maintenance racing the divergence fixture teardown (#18810)
`worktree-base-divergence-real-git.test.ts` builds cap-sized histories (100 and 101 commits). Every `git commit` detaches `git maintenance run --auto`, whose commit-graph task arms at 100 new commits, so the fixture reliably spawns a background `git commit-graph write --split` that keeps creating `.git/objects/info/commit-graphs` entries after the synchronous exec returns. The `afterEach` recursive remove is then deleting `.git/objects` underneath a live writer and dies with ENOTEMPTY — which is how "counts drift in both directions" failed on main. Reuse the existing `GIT_FETCH_SKIP_AUTO_MAINTENANCE_CONFIG_ARGS` (it already covers modern maintenance and legacy auto-gc, so it holds at the Git 2.25 baseline) in the fixture's git helper. Traced spawns of `git commit-graph write` over a full run of this file: 4 before, 0 after. The production path under test only runs `rev-list` and `merge-base`, neither of which triggers auto-maintenance, so there is nothing to fix outside the fixture. |
||
|
|
cec26336e5 |
fix(native-chat): say when a structured launch fell back to a terminal (#18762)
* fix(native-chat): say when a structured launch fell back to a terminal A definitive refusal already opened a terminal instead of the requested structured chat, but said nothing — indistinguishable from the bug where the wrong surface opens. Notify at message severity, since nothing failed. Also stop putting the raw error in the failure toast's description: it carried errnos and absolute paths straight into the UI. The detail moves to a warn log and the toast gets catalog copy, matching how the coded refusals already read. * fix(native-chat): avoid overstating terminal fallback --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
c96dd59904 |
fix(native-chat): correct the Codex subagent roster's build, journal write, and failure reporting
* Restore the exhaustive block handling that adding `subagent-group` to `NativeChatBlock` broke. `formatWorkerTranscriptMessage` and `boundBlock` both fell through to `image-ref` field access, so `tsc -p` failed for the CLI and node projects and `build:cli` could not emit. Both now guard on `image-ref` explicitly and give the roster block its own branch. * Stop the roster's publish from evicting its own append. The sink queue coalesces by `coalescingKey` alone with no op-kind check, so passing the append's key to `tryPublish` spliced the queued append out and the row never reached the journal — permanently, since `lastSerialized` was already set. `tryPublish()` now takes no argument, matching every other call site. The regression test's fake sink honours the key, which the previous fake did not. * Keep `collabAgentToolCall` substantive. Only the MultiAgentV2 path emits `subAgentActivity`, so a V1 turn has no roster row; suppressing its collab tool calls too would have left a V1 fan-out showing nothing at all. * Surface a settled failure while siblings still work. The summary now reports the worst adverse outcome independently of the group verdict, so the row shows `3 working +1 failed` with a failed-coloured dot instead of a neutral pulsing dot. The plain-text twin names it too. * Treat `/morpheus` as a child. Only `/root` is the turn itself; the old segment-count test silently dropped a valid single-segment agent. * Refresh token-usage recency on update so an active thread is not evicted as the oldest entry, and scope the `agentsStates` comment to the V2 path. |
||
|
|
9927edd631 |
fix(ssh): let the host say whether it armed the ready marker (#18802)
#18796 made every SSH Codex background launch wait for the shell-ready marker, but the client cannot see the remote shell. On a host that never publishes one -- fish, sh, Windows, or a relay predating #18796 -- no marker arrives and delivery falls back at 1.5s where it used to write at 50ms. The relay already computes whether it armed the marker; publish that as an optional `shellReadyArmed` on the spawn reply and let the client skip a wait it now knows is pointless. Absent stays UNKNOWN and keeps the client's own guess, so an older host behaves exactly as before; false is only ever an answer a host gave. It rides every reply, false included, or absent would stop meaning "old host". A host that did not arm the marker did not arm bracketed paste either, so the released path still submits raw. |
||
|
|
05e4fa037d |
Merge remote-tracking branch 'origin/main' into brennanb2025/codex-subagent-worklog
# Conflicts: # src/renderer/src/components/native-chat/NativeChatMessageList.tsx # src/renderer/src/components/native-chat/NativeChatMessageRow.tsx |
||
|
|
e95d247be1 |
perf(terminal): cheap-tier process inspection for anchored local agent panes (#18780)
* perf(terminal): cheap-tier process inspection for anchored local agent panes Every idle local pane's completion cadence ran a full whole-host `ps` (with `tty=` and `command=`, 0.34-0.50s on a 1,900-process Mac, 1.15s on Linux) purely to build `foregroundProcessEvidence` that the renderer then discards for local ids. Add a cheap tier (same job-control columns, no tty/command, 0.03s) gated so that it introduces no user-facing trade-off: - Only a pane whose last FULL capture proved a recognized agent may take the cheap tier. Panes with no anchor always take the full capture, so start discovery keeps today's exact behaviour. - The cheap tick compares a per-pane fingerprint (root shell pid+start, tpgid, every descendant's pid+start+pgid+job-control state). Any change, a changed node-pty foreground name, an unreadable capture, or an incarnation mismatch escalates to the full capture. A recognized agent's exit is always a pid vanishing, which the fingerprint always sees. - A cheap answer OMITS evidence rather than fabricating a tty-less fence. Remote/restore consumers never send `steadyState`, so they keep the full capture unchanged. - `steadyState` is a new optional request field; an old daemon ignores it and answers with the full capture. Measured (8 idle panes, 60s, idle cadence, forks counted by column set): 30 full -> 1 full + 29 cheap. * fix(terminal): route the cheap ps capture through runProcess The cheap-tier reader imported node:child_process directly, which the child-process import-boundary and windowsHide ratchet tests reject (CI shards 1/8 and 3/8). Use Orca's single spawn entry point instead; it pins windowsHide and encodes argv. Map its result onto the capture-error vocabulary: outputTruncated -> capture_truncated, timedOut -> capture_timeout, non-zero exit -> ps_exit_<code>. Tests mock at the runProcess seam. * fix(perf): refuse a pane fingerprint when any descendant start marker is missing `buildPaneProcessFingerprint` rejected only a missing root start marker; a missing descendant marker was stamped as `?`. Two captures that both failed to read the same descendant therefore compared equal, which removes the pid-reuse protection the fingerprint exists to provide: a recycled pid could make a vanished agent look unchanged, and the cheap tier would keep serving its name instead of escalating. Reachable on Linux, where `readLinuxProcStartTime` legitimately returns null when a process exits between the `ps` capture and the `/proc/<pid>/stat` read. Every subtree member now needs a start marker or the fingerprint is refused, which sends the caller to the full capture — the same conservative default every other uncertain path takes. Reported by CodeRabbit on #18780. The two new tests fail against the previous code with `expected '4242@2400#4300:|4300@?:4300:+' to be null`. |
||
|
|
ccf3e27800 |
perf(relay): bound the symlink directory probes a remote readDir fans out (#18752)
`readRelayDir` issued one `stat` per symlinked entry and awaited them all in a single `Promise.all`. A pnpm `node_modules` is hundreds-to-thousands of package symlinks in one directory, so expanding it over SSH put that many stats in flight at once, saturating libuv's four-thread pool and delaying every other relay filesystem operation — including the interactive reads `fs-list-files-scan-coordinator` exists to protect. The probes now run through `forEachWithConcurrency` at 8, the cap every other bounded probe in this codebase already uses (`GIT_COMMON_SNAPSHOT_CONCURRENCY`, `PRUNABLE_EXISTENCE_PROBE_CONCURRENCY`, `SPARSE_CHECKOUT_DETECTION_CONCURRENCY`). Results and ordering are unchanged: every symlink still resolves to its target's kind, and `sortDirEntries` still runs afterwards. The new test builds a 60-symlink directory and asserts the same 60 stats happen with exactly 8 in flight at peak — the probes overlap, and never past the cap. |
||
|
|
cc07249e78 |
fix(agent-session): refuse a pre-commit structured create with an envelope (#18697)
* fix(agent-session): refuse a pre-commit structured create with an envelope The create route refused by throwing, which reaches a client as a generic transport error indistinguishable from a lost answer — so desktop parked the launch as visibility-unknown with no chat and no terminal. Convert the whole pre-commit span, everything before `attach`, into a refusal envelope carrying a code, and name the definitive-refusal allowlist the fallback decision needs. * fix(agent-session): gate legacy fallback on definitive refusals * fix(mobile): preserve unknown structured create outcomes --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
daa105e83c |
feat(native-chat): give the subagent summary row its bot glyph
The row led with a glyph that swapped on state — a check once every child completed, a group icon otherwise — so a group appeared to change identity the moment it settled. Per the approved mock, the glyph names the category and never moves: state is carried by the status dot and the tone of the words beside it. Use lucide `bot`, the same glyph the individual `subAgentActivity` rows take in the eight-category vocabulary, so the summary reads as their parent. Slot and glyph are the mock's 16px/14px, muted by default, and the svg is `aria-hidden` — the headline is what a screen reader announces, so the icon never stands alone. |
||
|
|
4c5077d57a |
perf(persistence): skip rewriting unchanged terminal scrollback snapshots (#18764)
* perf(terminal): tighten the partial-escape-tail benchmark and equivalence test
* perf(terminal): spell the ESC gate the same way as the sibling ingest gates
* test(terminal): differential-fuzz the ESC-free partial-escape-tail gate against the unguarded fold
* test(terminal): make the escape-tail fuzz exhaustive at symbol depth, and cap the fold expectation
Two review findings on the differential fuzz, both about the test faithfully modelling the
function it guards.
The odometer generated strings by symbol depth but the caller filtered on `chunk.length`, which
is the UTF-16 code-unit count. An astral symbol is two code units, so every depth-4 string
containing one was silently skipped and the corpus was not exhaustive at depth 4 the way the
test name claimed. The generator now yields `{ depth, text }` and the caller filters on depth.
That restores the missing strings and takes the pinned corpus from 516,566 to 593,468 - exactly
the count CodeRabbit derived for the intended corpus.
The pairing assertion in the sibling suite compared the capped `advancePartialEscapeTail`
against an uncapped `extractPartialEscapeTail(pending + chunk)`. It passed only because no
pairing in that corpus crosses MAX_PARTIAL_ESCAPE_TAIL_LENGTH; it would have stopped modelling
the function the moment one did. The cap now lives in the expectation, matching the fuzz
oracle.
Re-verified the fuzz still fails on a wrong guard: mutating the gate to a bracket check fails
all four tests with a `gate diverged` assertion on a lone ESC chunk.
Reported by CodeRabbit and pullfrog on #18748.
|
||
|
|
6dc7e9fb6a | fix(sidebar): stop ⌘⇧↓ worktree navigation jumping to the first row (#18804) | ||
|
|
e89deb63c9 |
Show Claude background task status in Native Chat (#18757)
* feat(native-chat): show Claude background task status * fix(native-chat): carry background task fence forward * fix(claude): bound background task stop requests * Show running Claude background task details * Harden Claude background task status updates --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
320f4c8b87 |
perf(renderer): index four projections that rescanned their inputs per keystroke (#18747)
Four renderer projections scanned a collection inside a loop over another collection, each on a
path that reruns per keystroke or per store write. All four now build the index once per stable
input, which is what the surrounding code already does for its other lookups.
`workspace-kanban-search.ts` called `searchWorktrees`, the convenience wrapper that builds the
palette document index inline. The board's filter hook memoized the whole call on the query, so
every character re-normalized and re-segmented every indexed field of every worktree — and did
it again on every agent-status tick while a query was active, since those churn board
identities. `buildWorkspaceBoardPaletteDocuments` splits out, memoized on
`[worktrees, repoMap]`; only the match reruns per keystroke. This is the shape
`worktree-jump-palette-document-index.ts` already provides for Cmd-J.
`useTabGroupItemProjections` resolved each editor tab against `state.openFiles` — the global
list across every worktree — and each `tabOrder` entry against the group's tabs, with a `.find`
per element. Both are now `Map` lookups, alongside the `terminalTabById` index the same file
already built. The memo key `groupTabs` gets a new identity on any unified-tab write, so this
ran on title, label and colour changes.
`buildSourceControlTree` rebuilt every ancestor path with `segments.slice(0, i + 1).join('/')`
per segment, making tree construction O(files x depth^2) in characters copied — on the path the
Source Control file filter rebuilds per keystroke. The path now accumulates.
`worktree-header-section-boundaries.ts` ran a full `findIndex` over the render rows for every
header row, plus an `indexOf` over the bucket ordering, in two `useMemo`s keyed on `renderRows`
— so it recomputed on every sidebar row-model change, not just during a drag. One indexing pass
each, first-match-wins to match `findIndex`/`indexOf`. The successor index is keyed per bucket
because a repo or group id can appear in more than one bucket ordering; a flat id-keyed map would
pick whichever bucket was iterated first. `worktree-header-section-boundaries.test.ts` pins that
and the first-wins duplicate-header case.
Measured by `pnpm bench:renderer-quadratic-scans`. Three scenarios time the production export
against a reproduction of the pre-change function; the tab-group scenario is modelled on both
sides because the projection lives inside a React hook. Each asserts before/after agree first.
| projection | drives | scale | before | after | |
| --- | --- | --- | --- | --- | --- |
| workspace board filter (per keystroke burst) | production | 300 worktrees x 12 keystrokes | 15.4 ms | 4.4 ms | 3.5x |
| tab-group projections (per unified-tab write) | modelled | 60 tabs x 120 open files | 1.23 ms | 0.34 ms | 3.6x |
| source-control tree build (per filter keystroke) | production | 5000 changed files | 6.2 ms | 3.9 ms | 1.6x |
| sidebar header boundaries (per row-model rebuild) | production | 80 repos x 600 rows | 2.0 ms | 0.8 ms | 2.6x |
The tree and sidebar wins are smaller than the scans they remove because the rest of each
function (tree finalize/sort, per-row size estimation) is linear and now dominates.
4,537 existing sidebar, tab-group and right-sidebar tests pass unmodified.
|
||
|
|
c36dd5f4c7 |
perf(terminal): skip the partial-escape-tail walk on ESC-free PTY chunks (#18748)
`advancePartialEscapeTail` runs once per PTY chunk, on the main thread, for every terminal — visible, hidden or parked — inside `HeadlessEmulator`'s write path. It unconditionally concatenated the pending tail with the whole chunk and then walked the result one code unit at a time through a VT500 state machine. `extractPartialEscapeTail` only leaves `ground` on an ESC byte, so with no pending tail and no ESC in the chunk the answer is always ''. Taking that case up front skips both the full-chunk concat and the walk. `String.prototype.includes` is a native scan, so the gate costs essentially nothing on the chunks it does not short-circuit. This is the same gate its two neighbours on the very same ingest path already apply — `TerminalOscCwdTitleScanner.scan` and `TerminalMouseModeMirror.scan`, both carrying a comment citing their measured share of a 2.2x ingest regression. This call was simply missed. Measured by `pnpm bench:terminal-partial-escape-tail` over 640 x 16 KB chunks (10 MB), median of 7 rounds: | stream shape | before | after | | | --- | --- | --- | --- | | ESC-free (build logs, `cat`, piped output) | 40.3 ms | 0.19 ms | 213x | | SGR-coloured output (gate does not apply) | 30.6 ms | 32.1 ms | 1.0x | The benchmark proves equivalence over a 226-case corpus before timing, and a new unit test pairs every pending-tail state the scanner can be left in against every chunk shape, asserting the gate is indistinguishable from the unconditional fold. |
||
|
|
12a3e64177 |
perf(native-chat): stop rebuilding every transcript row on every stream frame (#18744)
The structured-session read owner emits per SDK event with no coalescing, so a working turn re-renders `NativeChatMessageList` tens of times a second. `MessageRow` was a plain function component, so every one of those frames re-ran `nativeChatProseToMarkdown`, an image-block scan, and a provider-frame lookup for every message in the window — up to the 300-message pagination limit — even though only the streaming tail had changed. `MessageRow` moves to its own module and is memoized, and its four derivations fold into the `useMemo` that already keyed on `message.blocks`. All of its props are primitives or already-stable references (`onScrollMessageToTop` is a `useCallback`, `onLinkClick` comes through `useShallow`), so the memo holds for every settled row. The move also takes `NativeChatMessageList` back under the max-lines limit rather than bumping it. `NativeChatToolRun` gets the same treatment: it was unmemoized and made four independent passes over its block list per render. `useNativeChatTurnStatus` allocated a fresh slice of the whole current turn and ran a nested `.some()` over it on every render; both now memoize on `[messages, latestUserIndex]`. Measured by the new `NativeChatMessageList.stream-render.perf.test.tsx`, a 120-message transcript with a 20-frame streaming turn: | markdown rebuilds per stream frame | before | after | | --- | --- | --- | | 120-message transcript | 120 | 1 | The test counts real per-row work rather than a render counter, and it fails on the pre-change component (`expected 120 to be less than 12`), so it is a genuine guard. Pure memoization on unchanged props: no visual, ordering, scroll-anchoring, or lifecycle change. All 997 native-chat tests pass unmodified. |
||
|
|
e1599c94b8 |
perf(terminals): let idle panes share one process-table capture instead of forking their own (#18742)
* perf(terminals): let idle panes share one process-table capture instead of forking their own Every visible local pane runs an agent-completion cadence that resolves through `getStrictProcessTableSnapshot`, and the inspection queue already collapses every shared-observation task enqueued in the same tick onto a single whole-host `ps`. Independent ±10% jitter per pane defeated that: the jitter was re-rolled on each reschedule, so panes drifted permanently apart, each landing in its own tick and each missing the snapshot's 500 ms TTL. Four idle panes cost four captures where one would have served all of them. Idle panes now aim at a deadline grid anchored at the epoch. The pull-forward is clamped to the snapshot TTL, so no interval is ever longer than its tier and none is more than 500 ms shorter: a pane off the grid walks onto it over at most `tier / TTL` steps, costs at most one extra inspection in total, and no inspection is ever delayed. Scoped deliberately. A pane with a foreground agent, or one still inside the 10 s post-activity hot window, keeps its exact interval and its own phase, so the bounded hot cadence is unchanged. The error-backoff path keeps its jitter, where spreading retries across panes is the point. Measured by `pnpm bench:agent-inspection-cadence` — whole-host `ps` captures over 60 s at the 2 s idle tier, median of 21 rounds: | visible panes | before | after | reduction | | --- | --- | --- | --- | | 1 | 29 | 29 | 0% | | 2 | 42 | 30 | 29% | | 4 | 62 | 31 | 50% | | 8 | 82 | 32 | 61% | `process-table-snapshot-reader.ts` measures the `command=` column at 1.15 s of work for 1,948 processes, so these are captures a quiet app was paying for continuously. All 4,202 existing terminal-pane tests pass unchanged, including the no-evidence cadence suite that pins the relaxed and hot intervals. * test(terminals): report n/a instead of dividing by a zero baseline in the cadence benchmark A window shorter than one cadence tier leaves the baseline capture count at zero, and the reduction line then divided by it and printed a meaningless percentage. Reported by CodeRabbit on #18742. |
||
|
|
7856e5a677 |
fix(sidebar): confirm filter reset before revealing active workspace (#18708)
* fix(sidebar): confirm filter reset before revealing active workspace * Polish workspace reveal confirmation and focus primary action * Fix reveal confirmation CI: commit-phase ref and sidebar test provider |
||
|
|
821c8b7df0 |
Bump mobile app.json to 0.0.48 (#18801)
Co-authored-by: Merge Sim <sim@local> |
||
|
|
36a826ff48 |
fix(ssh): compile node-pty from the host's own Node headers instead of nodejs.org (STA-6674) (#18774)
* fix(ssh): compile node-pty from the host's own Node headers instead of nodejs.org
STA-6674: a Linux SSH host that cannot reach nodejs.org never came up. node-pty
ships no Linux prebuild, so npm hands it to node-gyp, and node-gyp's default is
to download node-v<ver>-headers.tar.gz before configuring. The host refused
that connection (ECONNREFUSED) and the relay deploy failed inside npm install,
which the UI showed only as "Disconnected".
Every official Node build and every version manager that unpacks one already
has those exact headers at <prefix>/include/node. Export node-gyp's nodedir to
that prefix, on every command that can compile node-pty (npm install, npm
rebuild, the cloexec patch's rebuild), when the shipped node_version.h matches
the running Node. Both npm_config_nodedir (node-gyp 10, Node 20) and
npm_package_config_node_gyp_nodedir (node-gyp >= 11.4) are set so every Node
the relay runs on reads it. A version mismatch leaves it unset, which is the
existing behaviour.
When a host is both header-less and offline, name that in the deploy error
instead of forty lines of gyp http output, with the two remedies.
Reproduced and verified with a Docker sshd whose nodejs.org resolves to
127.0.0.1, on node:24.12.0 (the user's version), node:20 and node:26:
ssh-relay-offline-node-headers.docker.test.ts.
* fix(ssh): fail loudly when node-gyp ignores the exported Node headers dir
The headers export relies on npm forwarding npm_config_nodedir /
npm_package_config_node_gyp_nodedir into lifecycle scripts. If a future npm
drops that, node-gyp would silently fall back to downloading, and an offline
host would fail with the same "install an official Node" diagnosis -- wrong,
since the host did ship headers.
The prefix now echoes ORCA-NODE-HEADERS:<dir|none> into the command's output
before the compile, and the download-failure diagnosis reads it back: an
exported dir plus a download attempt is reported as an Orca defect naming
the dir, not as a host problem. Nothing else changes when it works.
* fix(ssh): address review on the relay node-headers export
- Unset any inherited npm_config_nodedir / npm_package_config_node_gyp_nodedir
before the conditional export, so a stale header dir from the remote profile
cannot bypass the version check and build a wrong-ABI binding (CodeRabbit).
- Require `gyp ERR! configure error` and a real network errno in the
headers-download matcher; node-gyp's fetch client logs retried attempts it
recovers from, and a FetchError can be a non-2xx mirror answer (pullfrog).
- Say "no local headers matching its own version", since the probe also
rejects a version mismatch, not only absent headers (CodeRabbit).
- Log the same diagnosis from the non-fatal `npm rebuild` fallback (CodeRabbit).
- Docker test waits for the SSH banner on the mapped port before connecting
instead of trusting `docker run -d` (CodeRabbit).
* fix(ssh): read the node-headers marker from the host output, not the quoted command
execCommand rejects with `Command "<command>" failed (exit N): <output>`, and
<command> quotes the whole prefix, marker echo included. The first-match
parser hit that copy and returned `${ORCA_NODE_HEADERS_DIR:-none}"; ...` as a
"dir", so every real no-headers failure was misreported as an Orca defect
(measured by an independent Docker exercise of
|
||
|
|
58553bfe1c | fix(recovery): fail a renderer recovery reload that never loads, instead of leaving a dead window (#18466) | ||
|
|
172aa1ac35 |
feat(native-chat): render agent file edits as inline diff cards (#18765)
* feat(native-chat): render agent file edits as inline diff cards An agent's file edit rendered as a flat list of every removed line followed by every added line, with no interleaving, no file header, and no line numbers. A Codex edit on the transcript lane rendered no diff at all: the patch arrives wrapped in the source string of its `exec` tool, which matched none of the shapes the old parser looked for. Adds one diff model shared by every edit shape the supported agents produce: - `native-chat-edit-lcs` interleaves a snippet pair, falling back to a linear prefix/suffix diff above the quadratic guard. - `native-chat-unified-patch` keeps the `@@` ranges as per-row line numbers instead of parsing them into display text and discarding them. - `native-chat-begin-patch` recovers the `*** Begin Patch` envelope from the JavaScript string literal Codex sends it in, so that lane renders a diff. - `native-chat-edit-normalize` folds all of it into one model, including the two Codex shapes that do not look like diffs: add and delete arrive as raw file content, and a rename is appended to the body as prose. Claude reports an edit as a snippet pair, which cannot locate the change in the file, so its result's resolved hunks are now carried on the tool-result block and preferred when present. The field is optional, so an older client reading a newer journal simply drops it. Where no resolved ranges exist the gutter stays blank rather than showing a snippet-relative number, which would read as a file position. The card renders the verb from the observed change kind rather than the tool name, pairs an edit's call and result into a single row, and takes its row and gutter grounds from new tokens derived from the git status palette, replacing the hardcoded Tailwind tints the old view used. Desktop only; mobile chat keeps its existing renderer and parser untouched. * fix(native-chat): stop the diff card from asserting an edit it cannot prove Every defect here shares one failure mode: the card stated something the input did not support, and stated it confidently. Parsing: - A hunk no longer ends on `--- `, `+++ ` or `\ No newline`. The first two are what a removed `-- comment` (SQL/Lua/Haskell) looks like once the marker is prepended, so they truncated the whole diff; the no-newline marker is emitted mid-hunk, between the removed old last line and the added new one. Real headers are recognised through `isFileHeaderPair`, lifted out of `native-chat-diff` so the rule has one home. - A `*** Begin Patch` envelope with no `*** End Patch` is declined. With no closing marker `indexOf` returned -1 and the slice swallowed the rest of the command line, so `… +y" && echo ok` rendered as file content the agent never wrote. - One splitter serves every shape, so a CRLF patch no longer keeps a `\r` on each row, in the phantom-row guard, or in the clipboard. It also tests for the trailing newline on the clipped body: on the un-clipped string that test deleted a real line whenever the slice fired. - Truncation is carried from each slice site to the card, so content past the character cap can no longer render as a complete unchanged file with no "Diff truncated" footer. Attribution: - A failed or still-running edit renders no card. It kept the generic tool view, whose result block carries the provider's own error — the card had been drawing "Edited file +1 −1" from the input while hiding the red error body, which is worse than what preceded this feature. - The result-as-patch fallback is scoped to `Diff`, the one tool whose call carries only a path. Any command tool's output could previously be read as a patch, so `git diff` through `exec` was reclassified as an edit of a file named "file" and its command line disappeared with the result. - A whole-content write claims a creation only on evidence — the editor tool's own `create` command, or the provider reporting one. Overwriting a large existing file had always read as "Added file". - `MultiEdit` reads its `edits[]`, and `NotebookEdit` leaves the set: it carries only the new cell source. Both previously fell through to the old renderer, so one turn could show two diff presentations at once. - Snippet-relative numbers are dropped at the model layer rather than hidden by a zero-width gutter, which the flex min-width floor re-exposed on top of the marker and the first characters of the row. The run memoizes its edit model, so a collapsed group no longer re-diffs on every streaming token, and the card's copy button says what it copies. * fix(native-chat): keep every edited file, and mark where the diff breaks A run of hunks was concatenated into one flat row list, so the gutter jumped from one region of the file to a distant one with nothing between them and the reader saw two unrelated spans as one continuous block. Rows now carry an explicit break: it holds no text and no position, counts toward neither side of the change, is trimmed from the end where it would mark nothing, and is left out of the copied text. The patch envelope lost files, and lost them silently: - An update chunk may carry no hunk header at all. The parser required one, returned nothing, and the caller dropped that file from a multi-file envelope with nothing to say it had gone. A header-less body now opens as a hunk of unknown position, and whether the rows are locatable is read off the rows themselves rather than off the header. - The envelope's own control lines rendered as content rows in the card. - A delete names its file and carries no body, which rendered as a card with an empty expandable row list. The header states the change and offers no disclosure behind it. - The header patterns are anchored and `.` excludes a carriage return, so a CRLF envelope matched no header at all and produced no card whatsoever. The envelope is split on both newline forms once, up front, rather than each pattern having to tolerate the extra character. A tool call's argument payload arrives as a string holding JSON. It was passed along undecoded, which is the only reason this code carried a hand-rolled string-literal unescaper. It is decoded once at the transcript decoder now — defensively, since the transcript is untrusted, so anything that is not a JSON object is left exactly as it arrived — and the unescaper is gone. Recovering the envelope no longer guesses at argument names either: it looks at the values, including the words of an argument vector, which is where the envelope actually sits once the payload is decoded. * fix(native-chat): only read a patch where a patch was actually run Recovering the patch envelope from any value of a tool's payload meant a write's own content was searched for one. A file documenting the patch format rendered a card for the file its example names, while the file actually written never appeared at all — the call and its result were consumed by that card, so nothing was left to correct it. Two changes: the envelope is recovered only for the tools that run one, never for a file edit whose payload is content; and only patch- or command-bearing arguments are searched, still including the words of an argument vector, which is where the envelope sits when a command tool applies it. The call payload is decoded back where it is needed rather than at the transcript decoder. Decoding it there changed the shape every reader of a tool's input sees, including the surface that recognises a question payload from any tool by shape alone: a tool whose arguments happened to carry that shape raised a question card pinned over the composer. That decode now happens inside the envelope recovery, the one consumer that needs the structure. A card also states an edit as made, so it now takes evidence that it landed — the provider reporting the call complete, or a result that is not an error. A turn that stopped before its call was answered reported an edit that may never have applied. This replaces the working-turn heuristic in the view, so the rule lives in one place. Two files still went missing. A multi-file patch has no per-file split, so it rendered as one card under the first file's name, with the later files' rows and their gutter numbers beneath it — a card asserting a false file position. Patch text is now split on its file boundaries, one card per file, each named by its own header, with a rename and a `/dev/null` side read from the same headers. And an envelope section that names a file but carries no body was dropped rather than reported, which is the same silent loss the delete case was fixed for. * fix(native-chat): type the patch-section scan and its test helper call The section under construction was only ever assigned inside the helper that opens one, which control-flow analysis does not see, so the variable stayed narrowed to its initial null and reading a field off it did not compile. The helper now only builds and records a section; the loop owns the assignment, which also fixes a real leak in the fall-through row: it opened a section it never made current, so the next row opened another one. The multi-file case also passed a possibly-undefined slice to a helper that takes an array or null. * fix(native-chat): stop the patch lane naming files it cannot name Splitting a patch into its files only ever looked for a boundary outside a hunk, and nothing reopened that state once the first hunk began, so every file after the first was swallowed as the first one's body. A `--- `/`+++ ` pair inside a hunk is now a boundary too, but only when a hunk header follows it immediately: a removed `-- x` over an added `++ y` is never followed by a column-0 header, which is what keeps the guard against reading content as structure intact. One producer cannot be recovered by any parser: it joins several files' patches and keeps a count where the path goes, so nothing in what reaches here names a file. That shape is refused rather than rendered under a name no file has. Recovering the per-file paths belongs to the producer and is filed separately. A clipped body carries its own marker in its text, and the bound that clips it is six times smaller than this module's, so it fires first. Read as content, the marker became a numbered line of the file and the rows before it were reported complete. It is recognised at the end of the text, removed, and reported as the truncation it is — the footer says so and the copied text no longer carries it. Also: a move appended to the body as prose is now read as a rename on every lane that carries the body as text, not just the one that also carries the destination as a field, where it had been rendering as a numbered line of the file it moved. The call's own path no longer wins over a rename's destination, which is only ever in the header, and only sections that name a file count toward deciding whether the call names the one file at hand. A command that merely quotes an envelope — writing documentation about the format — must now also invoke the tool that applies one. And two compared directories are no longer called a rename: only a header that states both sides as such is evidence of a move. * fix(native-chat): anchor the move marker to its own line The marker a producer appends to say where a file moved was matched anywhere on the body's last line, so a row whose own content mentions a move was cut in half at that point and the file it named claimed as the destination of a rename that never happened. It is now anchored to the start of the final line, on both lanes that carry the body as text. The command that applies a patch envelope has a second spelling the runner accepts and runs; requiring the first one refused a patch that really landed. Both are accepted, still matched against whole argument words rather than the payload at large. A clipped diff also said so only under its own rows, where a collapsed card — or one clipped down to no rows at all — showed nothing. It sits beside the change counts now, which are visible either way. * refactor(native-chat): tidy what the diff-card work left behind The copy text is joined from every row of the diff, which a collapsed card renders none of, and it was rebuilt on every render to seed a prop. It is memoized on the rows, matching how the run memoizes its edit model. The two scanners that read patch text kept the same file-section alternation verbatim, so they could drift apart while both looking correct; there is one definition now, beside the header-pair rule that already lives there. Also: the row that marks a break between regions is built in one place, so it is no longer exported; the move destination in the envelope reader was a function-wide binding written and read within one iteration, which read as if a move carried between sections; and a test comment named the wrong mechanism for keeping a card collapsed. Adds the missing pin on what the copy affordance actually copies. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
149732df6d | perf(persistence): stop the session write re-scanning and rebuilding unchanged state (#18739) | ||
|
|
b0c67eaf88 |
feat(mobile): port the restructured native-chat turn status and live tool progress (#18761)
* feat(mobile): port the restructured native-chat turn status and live tool progress Mobile chat had a single static "Agent is working" row and no live tool activity, while the desktop restructure (#17597, #18705) replaced that with a per-turn status row and a running-tool label. This brings mobile to parity and puts the derivation in one place instead of two. Shared (new, pure, RN-safe — desktop uses them as i18n fallbacks, mobile directly, matching the native-chat-empty-state pattern): - `native-chat-turn-status.ts`: duration formatting, label selection, the turn-timing state machine, and the active/settled split. - `native-chat-tool-activity.ts`: command-tool classification, the running-tool label descriptor, and running-call selection. Desktop now consumes both; `NativeChatWorkingStatus`, `NativeChatToolRun` and `use-native-chat-turn-status` keep their existing behavior and strings. Mobile gains the "Thinking" / "Working for 12s" / "Worked for 3m 4s" row with a caret that discloses the turn's tool activity, the pulsing "Running npm test" row with terminal-vs-wrench glyphs, and desktop's rule that a completed turn's tool run hides behind the turn caret. The bridge lane is untouched and keeps its three-dot indicator. Headings, quotes, code, lists and table cells are now selectable. Files at their max-lines cap were split rather than bumped: the tool-run subtree, the prompt card, the session-lane wiring, and the turn-disclosure state each move to their own module. * perf(mobile): stop the turn-status rows from re-rendering the whole transcript A streaming turn re-renders the chat list many times a second. The disclosure wiring handed every row a fresh status object and a fresh toggle closure on each of those renders, so `MobileNativeChatMessage`'s memo never held and every visible row re-rendered per tick — including settled turns that had not changed. Memoize the status selection on the timing map, and keep one stable toggle handler per turn (pruned when a turn leaves the transcript) attached only to the settled rows that can actually disclose anything. Now only the live turn's row changes identity while the agent works. * fix(mobile): keep the turn clock running when the optimistic echo is replaced An accepted send renders as `pending-N` until the transcript echo lands under its real message id. That flips the active turn key mid-turn, and the timing reducer treated the new key as a new turn — so a turn that had reached "Working for 8s" visibly restarted at "Working for 0s". The reducer now carries the start over when the previous key names a turn that has since left the transcript, which is exactly the echo-replacement case. A genuinely new turn (the previous key still in the transcript) and a turn that had already settled both keep their own clock; both are pinned by tests. Desktop does not pass the new key and is unaffected. * fix(mobile): keep the Tools toggle working on settled turns Hiding a settled turn's tool run behind the turn caret (desktop parity) also made the composer's global Tools control a no-op on every completed turn: the run it wanted to expand was not rendered at all. Let that toggle override the hiding, so it still reveals every run at once the way it did before. * fix(mobile): re-key the turn timing instead of only carrying its start The previous fix carried the start forward only while the turn was still working. When the transcript echo landed after the turn had already settled, the new key inherited nothing, the settled timing was pruned with the old key, and the turn's "Worked for N" row disappeared entirely. Move the timing onto the new key instead, which covers both orderings: an in-flight turn keeps counting from its original start (and later settles against it), and an already-settled turn keeps its duration. Both orderings are pinned. * test(mobile): pin the structured turn-status wiring at the view level Emulator QA could not reach the structured lane (mobile's Create Tab -> Codex falls back to a terminal tab when agentSession.createSupport says unsupported), so the view's own lane wiring had no coverage — the one seam between the shared turn-timing reducer and the rendered rows. Assert what the view hands each row: the live user turn gets a status object and the three-dot indicator is gone on the structured lane; the bridge lane keeps the indicator and gets no status; a finished turn settles to a numeric duration with a toggle; and an assistant row never carries a status row of its own. * fix(mobile): isolate structured chat turn state * fix(mobile): let the capability RPC actually store what a phone advertises `runtime.clientCapabilities.update` records the advertised set by assigning `authenticatedSocket.clientCapabilities`, but the socket handed to the dispatcher defined that property with a getter only. In strict mode the assignment throws `TypeError: Cannot set property clientCapabilities ... which has only a getter`, so the RPC answered `runtime_error` and the set was never stored. The consequence is not subtle: `supportsStructuredAgentSessions` requires the capability, so `projectSessionTabAgentStatus` removed every `agent-session` tab from a phone that had advertised it correctly. A paired phone saw ZERO tabs on a worktree whose only tab was a structured Codex chat — structured native chat was unreachable on mobile over this transport, not just missing its new turn UI. Give the socket a setter that writes through to the channel, which already owns the set for the connection's lifetime, so later requests on the same socket see it. Found while trying to capture emulator screenshots of the turn-status port: two full QA runs reported the new UI "missing" because the phone could only ever get a bridge/PTY tab. * fix(mobile): carry the turn key instead of caching a handler in a ref Builds on the scope-isolation fix: that kept (and extended) a ref that is written during render — once to memoize a per-turn handler, once to prune dead turns, once to reset on a scope change. React Doctor's "Ref mutated during render" is what CI's `check:react-doctor:changed` was failing on (x2), and on mobile it is a real hazard rather than a style note: react-freeze discards renders, and a discarded render would leave the cache mutated. Pass the settled turn's key down the row instead and let it call one stable handler with it. That preserves both properties the cache was bought for — per scope isolation, and identity stability so a streaming transcript does not defeat the row's memo — with no ref writes and no pruning to get wrong. The scope-keyed expanded set and the 128-turn cap are untouched; their tests move to the new contract and one now pins handler identity across a re-render. Note for future changes here: `check:code-quality:changed` does NOT cover this. CI additionally runs the standalone react-doctor CLI, which has rules the oxlint plugin config does not enable. * fix: ship native chat status translations * test(native-chat): pin the shared copy against the English catalog The shared constants are desktop's i18n fallback and mobile's actually-rendered string. If one changes without the other, desktop keeps rendering en.json while mobile renders the constant — and nothing fails, because a fallback is only used when the key is missing. That silent divergence is the exact thing the shared module exists to prevent, and it is now reachable precisely because these strings are runtime-required rather than statically extracted. Assert every key in both shared copy objects matches en.json byte for byte, plus the interpolation placeholders the catalog interpolates on. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
3f84c358f0 |
fix: recover from an alternate shell install, and close the relay duplicate-echo gap (#18796)
* fix: recover from an alternate shell install, and close the relay echo gap #18768: a startup profile that `exec`s a second install of the same shell keeps the pid but loses the wrapper's ready marker, and the recovery probe rejected the replacement's different canonical path -- costing plain Codex the full 15s barrier. The probe now also accepts an install that the pane's own PATH resolves, so a binary merely named bash/zsh outside it stays rejected. #18767: the SSH relay left plain Codex on early startup delivery, displaying the launch twice under a slow profile. The shell is the host's to know, so the relay now folds it into the same rule the daemon uses, and the SSH background client waits for the marker on any Codex launch. Bracketed paste is now gated on an observed marker rather than on the intent to wait, so a fallback release on a host shell that never publishes one submits raw. The non-daemon local provider needs no change: it hands Codex to the wrapper's own prompt hook and never writes it into the PTY. * review: correct an overclaiming comment, announce a silent skip, drop a shim Readiness review of #18796. The delivery comment claimed waiting "costs nothing", which is only true on a host that arms the marker -- fish, sh, Windows and hosts predating #18767 release on the fallback instead. Say so. The alternate-install recovery tests skip on usrmerge hosts, where /usr/bin/bash resolves back to /bin/bash; announce that rather than reading as coverage that does not exist. Import the line-editor predicate from shared directly instead of through a re-export left on daemon/shell-ready. |
||
|
|
b33d1972bc |
docs(relay): correct why the ConPTY teardown asset diverges from the desktop patch (#18636)
The divergence pinned by #18601 is real and worth keeping, but its stated reason was wrong. It claimed the desktop patch carries the early conin placement "and therefore the +2 File / +1 Process regression, measured against its exact installed tree" -- i.e. that the shipped desktop app leaks because of its own leak fix. It does not, for any terminal a user opens. node-pty defaults `_useConptyDll` to false. Every desktop site that opens a pane sets it true (`local-pty-utils.ts` twice, `native-pty-spawn.ts`), as does the `windows-conpty-warmup.ts` warm-up, so they take the `else` branch, where upstream already destroys the input socket. The relay passes no such option (`src/relay/pty-handler.ts`) and takes the `!useConptyDll` branch -- the one both this asset and the desktop patch edit. The desktop is not entirely off that branch, though: the hidden rate-limit probes in `src/main/rate-limits/claude-pty.ts` and `codex-pty-rate-limit-probe.ts` omit the option, recur, and tear down through `kill()`, so the hunk is live there -- just never for a visible pane. Whether the early placement costs the same +2 File / +1 Process across a probe's lifecycle is unmeasured; the numbers in this comment were taken on relay-style spawn/kill cycles, and the comment now says so. What is settled is the replaced claim: not every Windows user, and not every terminal. Those two probes were missed three enumerations running because they use `await import('node-pty')`, which no static-import grep finds. The comment now tells the next reader to grep for `node-pty` instead. The measurement that produced the wrong claim was taken by a standalone harness that passed no `useConptyDll` and so defaulted into the branch it was not trying to measure -- the same standalone-is-not-the-real-host trap #18601's own body warns about, one level down. Also refreshes the self-exit paragraph, which #18635 made stale. That leak is now fixed for the desktop, and the note records why the fix cannot reach a Windows relay. The fix is mostly native (`src/win/conpty.cc`) and this asset only rewrites `lib/*.js`, and all three delivery paths stop short of Windows: pnpm patches do not cross the SSH boundary; `MATRIX_SLOTS` in `build-orcad-prebuilds.mjs` has no win32 entry; and the one relay asset that does patch native source and rebuild on the host (`node-pty-1.1.0-master-cloexec-patch.cjs`) returns `skipped:unsupported-platform` for anything but linux/darwin. #18635's flat self-exit relay numbers were measured against a locally rebuilt binary, so they describe the relay code path on a patched tree, not the tree a relay host installs -- the note says so explicitly rather than leaving the next reader to conflate them. Assertion and hashes unchanged: the relay must still release conin after the console-list fork, and a patch sync must still not copy the early placement onto the relay's branch, where it does cost +2 File and +1 Process per terminal. Only the justification changes, plus the test name, which said "like the desktop patch" where it meant "unlike the desktop patch placement". |
||
|
|
41b520259e |
ci(package): retry apt fetches and docker builds behind the Ubuntu mirror (#18797)
The package job builds three Docker images whose apt-get update/install hit archive.ubuntu.com with no retry, timeout, or mirror fallback. When the mirror is mid-sync every build dies in one of three ways: - per-package fetch stalls (~64 s each, `Ign:` lines) until the runner's 10-minute docker build timeout fires: https://github.com/stablyai/orca/actions/runs/33935104546/job/101221447425 - `apt-get update` exit 100 with `Hash Sum mismatch` on noble-updates/restricted/Packages.gz: https://github.com/stablyai/orca/actions/runs/33935104546/job/101226099525 - `apt-get update` exit 100 with `File has unexpected size ... Mirror sync in progress?`: https://github.com/stablyai/orca/actions/runs/33935244026/job/101231083497 Each Dockerfile now retries `apt-get update` up to five times with Acquire::Retries and a 30 s HTTP timeout, clearing /var/lib/apt/lists between attempts so a half-synced index is never reused, and passes the same acquire options to `apt-get install`. Each runner script retries the whole `docker build` once when the first attempt fails or times out. |
||
|
|
ba4bbacd6b | fix(relay-ops): align the cloud-data freshness bar with Cloud Monitoring publish lag (#18798) | ||
|
|
436ef827dd |
fix(browser): present Electron's own user agent so Cloudflare Turnstile clears (#18749)
Orca rewrote every browser session's UA to look like plain Chrome by stripping the Electron and app tokens. That rewrite is what Cloudflare rejects: a Chrome UA that ships no client hints reads as a spoof and Turnstile returns 600010, while the same binary on the same IP clears every challenge with its stock UA. PR #885 added the rewrite to fix 600010 and was treating a symptom it created; issue #11518 later found the same rewrite is what broke Google sign-in. - Keep the stock Electron UA on every partition. The webRequest handler now only owns the host-scoped Google auth Firefox switch, which stays unchanged. - Delete the anti-detection script. Measured on Electron 43: plugins are already a real PluginArray, window.chrome exists, and navigator.webdriver is false even with the debugger attached, so three of its four premises were wrong, and the overrides it installed (instance-level webdriver, non-native Permissions.query, stubbed chrome.csi/loadTimes) are themselves published bot signatures. - Stop attaching a CDP debugger to every browsing guest. Only the auth-UA detach listener remains, because a detach clears Chromium's standing UA override. - Stop sending Runtime.enable into cross-origin iframes when the agent bridge auto-attaches. The challenge widget is one, nothing reads iframe Runtime events, and the Runtime domain's serialization side effect is the documented Cloudflare CDP tell. - Add a real-Electron test proving the wire identity: stock UA to ordinary hosts, Firefox with no client hints to accounts.google.com. Verified in the dev build: dash.cloudflare.com/login no longer shows "There was a problem with verification" and scrapingcourse.com's managed challenge clears, both failing deterministically before. Fixes #13822 |
||
|
|
040c3e5b32 |
fix(browser): match loading surfaces to the Orca theme (#18738)
* fix(browser): theme unavailable guest surfaces without recoloring pages * test(browser): keep generated loading evidence out of the PR diff * test(browser): freeze recovery clock during artificial attach gate * test(browser): await painted content after network recovery |
||
|
|
974acc901c | fix(relay-ops): retry freshness-only preflight failures on the first same-cap wave too (#18778) | ||
|
|
cb7f7dd11a |
fix(native-chat): tell old mobile builds why a structured chat is missing (#18756)
* fix(native-chat): tell old mobile builds why a structured chat is missing A structured native chat started on desktop was simply absent on a paired phone running any shipped App Store build. The host strips every `agent-session` tab from a client that does not advertise `agent-session.structured.v1`, and no released mobile build advertises it — so the chat had no representation at all and no way to explain itself. Keep the row and retitle it instead of deleting it. The shipped client does not filter unknown tab types and renders whatever title the host sends, so an old build now shows the chat's slot with a title naming the fix. Nothing is removed, so the tab order, groups and layout it belonged to are left intact. The prompt is keyed on the capability for that specific agent, not on the combined policy boolean: a capable phone whose desktop simply has the experiment off would otherwise be told to take an update that cannot help it. Claude rows are prompted too — mobile cannot render them yet and a later build can, so the message is true for that client as well. Restore is no longer gated on the caller's capability. It stayed gated on the host setting, which is what decides whether there is anything to reach at all, but gating on capability left an old client with nothing to project after a desktop restart: neither the chat nor the prompt. Tab titles are capped at 128px on one line in every shipped build, so the string is sized for ~15 characters rather than a sentence. Prompted rows are visible rows, so the host now permits all five session-tab mutations on them, close included. That is intended: a mobile close runs the same teardown as the desktop's own Close button. * fix(native-chat): keep fallback tabs safe and truthful --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
5a1acfec17 |
fix(native-chat): make document paths and links clickable in chat (#18712)
* fix(chat): make assistant file paths clickable * fix(chat): tighten native file link handling * fix(chat): link prose-joined relative paths separately * fix(chat): preserve complete Unicode file links * fix(chat): preserve links before sentence punctuation * test(chat): align structured session link props * fix(native-chat): harden generated file links * fix(native-chat): reject reference-number false positives * test(native-chat): align structured session parity * perf(native-chat): bound file-link detection --------- Co-authored-by: Merge Sim <sim@local> |