mirror of
https://github.com/stablyai/orca.git
synced 2026-10-09 00:02:39 +00:00
e1362ada4c33e7ef14709dff56564ea5dbcf5d98
12094
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e1362ada4c |
fix(terminal): stop inline-image decoders exhausting the renderer's wasm memory budget (#23499)
V8 reserves an 8 GiB guard region per wasm memory inside its 1 TiB sandbox, so an Electron renderer can hold only ~124 live wasm memories regardless of free RAM. @xterm/addon-image instantiated a SIXEL decoder per terminal at activation (and kept IIP decoders after the first image), so ~120+ terminals exhausted the budget: new panes raised 'WebAssembly.instantiate(): Out of memory' rejections, and the next Kitty/IIP image threw 'WebAssembly.Memory(): could not allocate memory' out of the parser, permanently wedging that terminal's write queue. The addon-image source patch now borrows SIXEL decoders from a shared pool only while a sequence is open (color registers stay on the terminal), drops IIP decoders after each image, and turns a failed decoder allocation into a dropped image instead of a parser throw. Bundles regenerated with regenerate-xterm-patches.mjs --write. Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
36c473ea1a |
fix(runtime): wait out Codex 0.157's startup screen, and stop at Codex's startup dialogs, before typing a worker brief (#23745)
* fix(runtime): wait for Codex's live chat before typing a worker brief Codex 0.157 draws a provisional startup screen (header reads model: loading) and discards typed input while it starts its shared daemon behind it; a fresh Codex home makes that window seconds long, so worker-start pasted briefs that were truncated or never submitted. Codex 0.158 dropped the header labels Orca matched, so worker-start stopped seeing Codex as ready at all. Readiness now requires Codex's live chat on both layouts: the provisional header vetoes a text match unless the live status row is already painted, and 0.158's greeting layout counts once that status row appears. Codex 0.158's model announcement dialog is reported as a blocked prompt instead of receiving the brief. * fix(runtime): recognise Codex's provisional screen from the text copy and the screen probe Live worker-start on a fresh Codex 0.157 home still typed during the daemon start: Codex leaves its alternate screen for that window, so the live screen showed no header and the screen-based veto never fired. The text copy keeps the provisional header until the live chat paints its status row, so the veto now reads it there. The tui-idle visible-screen probe used the bare text rule on the rendered screen; it now goes through the same body rule. * test(daemon): register the new Codex captures' known serializer divergences The serialize round-trip replay picks up every fixture under runtime/__fixtures__, and the three new Codex 0.157/0.158 captures showed 48/9/8 "new-fail" checkpoints against an expected 0, turning CI red. They are the existing live-pen colour leak on restored cells, the same class as the other Codex and DSH entries; this branch changes no serializer code. * fix(runtime): keep Qoder off the Codex screen probe change; drop an unbacked row filter - The tui-idle visible-screen probe now classified Qoder panes with isQoderComposerReady, which skips the working veto evaluateTuiIdle applies first. Qoder paints its composer mid-turn, so an adopted Qoder pane whose hooks said "working" settled the wait immediately. Only Codex and unknown panes take the body rule there; every other agent keeps its old verdict. - The live status-row check skipped rows containing "waiting for startup", a string Codex 0.157/0.158's TUI never prints. The line-folded text copy keeps a whole screen on one line, so the filter could only ever veto the real status row. It is now a bounded includes() with no split. - Lowercase the wait text once in isKnownReadyPromptBody. - Restore the per-frame "screen never takes a settled header away" check, guarded on the provisional veto, instead of checking the final frame only. * fix(runtime): stop reporting Codex 0.158's model announcement once it is answered The announcement's choices stay in the text copy after the user answers it, and the existing dismissal check needed the model:/directory: labels that 0.158's header lacks, so tui-idle waits and the agent-status query kept reporting codex-model-migration-prompt over a live chat. Codex repaints its whole screen, header included, when a startup dialog closes, so the header after the dialog now marks it answered. Also corrects the live-chat marker comments: the middle dot also comes from the daemon session's agents hint row and the warnings notice, not only the status row. * docs(runtime): note that Codex startup dialogs also draw the live-footer dot * fix(runtime): recognise Codex 0.157/0.158 startup dialogs by the rows they really print Codex 0.157 and 0.158 no longer print `Press enter to continue/confirm` on their startup dialogs; they print key rows instead (`enter continue · esc skip`, `enter confirm · esc skip`, `enter/esc continue · ctrl+c quit`). The update, hooks-review and model-migration matchers still required the old wording, so none of these dialogs was reported as blocked. On 0.157 the dialog's `·` also satisfied the live-footer check, so a tui-idle wait read the update dialog as ready and worker-start would type the brief into it, where Enter picks "Update now" (npm install -g, Codex exits). On 0.158 the wait timed out instead of reporting blocked. The matchers now accept the old wording or the new row, tolerating the spaces the line-folded text copy drops around `·`. Each one matches from the dialog's first `·` (for the update dialog that is its title row, `Update available · 0.157.0 → …`), so the dialog is blocked from the same character that would otherwise make the provisional header read as live. The retired-model notice without choices has a catalog-supplied heading (`GPT-5.4 is no longer available`), so it is matched by its own key row. No new blocked-reason value. The startup-dialog matchers move to startup-dialog-blocked-signals.ts to keep terminal-wait-detection.ts under the line limit. Backed by six real captures (update available, hooks review, retired model without choices, each on 0.157.0 and 0.158.0), replayed frame by frame and through a tui-idle wait; the serializer round-trip replay registers their existing live-pen colour divergences. * fix(runtime): keep reporting Codex's retired-model notice after a relaunch in the same pane The retired-model notice is matched by its key row alone, and the matcher took the first `enter/esc continue ·` in the live window while every other startup-dialog matcher takes the last. Quitting Codex from the notice and relaunching it in the same pane leaves the old copy ahead of the new launch's header, so the header read as having dismissed the new notice: 0.157 then read ready and a worker brief would be typed into the dialog. Take the last key row, and replay each captured dialog quit-and-relaunched to pin all six. * fix(runtime): match Codex startup dialogs by the rows the text copy keeps intact Codex 0.157+ paints each startup dialog over its startup screen by cell diff, so Orca's line-folded text copy can drop letters and spaces from a heading: #23765's 0.157.1 capture reads `Updat available`. The update matcher needed `update available`, so on that capture tier 1 read the dialog's own `·` as the live chat's footer and a tui-idle wait settled ready on the update dialog, whose Enter picks "Update now". Match each dialog from its first `·` by rows Codex prints as fixed literals: `available · <version>` and `enter continue · esc skip` (update), `enter confirm ·` (hooks review), `enter/esc continue|confirm ·` (model notices, which also covers 0.158's new-model announcement, so its choice-text matcher goes). Legacy `Press enter to …` wording still matches. Add the new rows to the blocked-signal prefilter, and replay #23765's 0.157.1 update capture in the dialog suite. * fix(runtime): don't name Codex's mid-session pickers a hooks review Codex's rate-limit reset popup (and its other pickers) end their key row with `enter confirm · esc back`, which the hooks-review row matched now that it no longer needs the heading. Exclude `esc back` instead of requiring `esc skip`, so a half-painted hooks-review row still blocks. |
||
|
|
bb91e39eb9 |
test(native-chat): main's Codex child tests expect a turn to end on its own turn/completed (#23783)
* fix(native-chat): a Codex child's turn ends on its own turn/completed #22553 ended a child thread's turn on an `error` Codex will not retry, reading the verdict module #23682 deleted when it made turn/completed the only end of a Codex turn. The two landed minutes apart with no textual conflict, so main no longer typechecks. Codex runs every thread's events, spawned children included, through the same per-thread handler: a turn-ending error is recorded as the turn's last error and the turn then completes as failed. So a child's failed turn/completed is its end, as on the primary thread, and a closed thread stays the one child ending with no completion. * refactor(native-chat): a closed Codex child thread is its own frame With a child's turn now ending only on its own turn/completed, the turn-ended frame carried a fixed turnId (null) and state (unverifiable), and endTurn took parameters nothing passed. Name the one remaining case: a thread-closed frame, and closeThread ending the running turn as unverifiable. Drop a test step that no longer exercised anything. |
||
|
|
b6c2de9f92 |
fix(codex): index a new account home before bridging history into it (#22971)
* fix(codex): index a new account home before bridging history into it Codex indexes every rollout present when it first creates its state DB, and the TUI gives up after 30s, so large histories broke a new account's launch. Fixes #20669 * test(codex): keep the account migration test from starting the real Codex binary Selecting an account starts the history bridge, which spawned codex app-server on the fixture homes and raced the test's temp-dir cleanup. * fix(codex): index a new account's most recent bridged history first Indexing a large history takes minutes, and Codex's /resume hides unindexed threads once a directory has any indexed one, so recent conversations stayed missing until the heal reached them in directory-listing order. * test(codex): cover the history bridge's quit, no-history and failed-link guards Drop the stop check before creating the state DB: it ran in the same tick as the check at bridge start, so it could never observe a quit. * test(codex): make the account heal tests fail closed instead of reaching a real codex Fake invocations now point at a nonexistent binary and a throwaway CODEX_HOME, and cover per-home failure memory and a session that fails during quit. --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
bd133058a9 |
refactor(sidebar): one subagent row for CLI and structured children, shared with the chat strip (#22565)
* test(sidebar): pin today's subagent rows and chat strip rows from legacy shapes Captured on the unmodified renderer: the compact and full sidebar child rows built from a legacy subagents snapshot, and the expanded chat strip built from a legacy background-task roster. A host that sends only these shapes must keep rendering exactly these rows. * refactor(sidebar): one subagent row for CLI and structured children, shared with the chat strip A child row's dot, name and detail are now decided once, by a shared row model (src/shared/agent-child-row-model.ts), and rendered by one piece (AgentChildRowContent) that both the sidebar's child rows and the chat strip use. The model reads the host's child views when a session publishes them and today's legacy shapes otherwise (the subagents snapshot for the sidebar, the task roster for the strip), so old hosts and CLI panes render exactly as before. From views it keeps what the legacy path lost: a finished subagent whose shell still runs reads monitoring through the shared child fold, a settled child reads by its outcome, the tool it runs shows as the CLI row shows it, and each row keeps its own clock. In the strip, work a child owns renders nested beneath it. * i18n: add the child row's Ended and ended-ago strings to every catalog * test(sidebar): a child reads the same in the sidebar and the chat strip A row-model table for every display state, and a parity table that renders the same child views through the compact sidebar row and the chat strip and asserts both show the same dot, name and detail. Covers a finished subagent whose shell still runs (monitoring on both, no stale tool text), sibling rows with their own clocks, a settled row timed from when it ended, and owned shells nested under their owner in the strip. The strip's view input is named childViews so it cannot be mistaken for React children. * fix(sidebar): keep the child display derivation loadable in the renderer The row model imported deriveAgentChildDisplayState from the view module, whose owner resolution reaches status subjects and, through them, agent-hook-relay's node:crypto. The renderer cannot evaluate that chunk, so the app booted blank (the renderer node-builtin boundary test fails on the previous commit). The display derivation (agentChildWorkOwnedLiveness, deriveAgentChildDisplayState, AgentChildDisplayState) now lives in agent-status-child-work-display.ts, which imports only the fold, liveness and the one-pass grouping both modules share (grouped-by.ts); the view module keeps the host-side projection. * fix(sidebar): the strip reads its parent's verdict, and the full row says Ended Review fixes: - The strip's view path took no freshness input, so once views are wired the same lost child would read unverifiable in the sidebar and working in the strip. agentChildRowContextForParent builds the one context a parent row gives its children; the sidebar uses it, and buildBackgroundTaskGroupsFromViews / the strip's new optional childRowContext prop accept it (absent: claims stand as reported, as before). - The full sidebar row showed no word for a child that ended with no outcome; its message line now reads the row model's (the message, or Ended). An unlabeled child falls back to its display state there too. - The summary order now includes failed, so a failed row is never dropped from the counts. - Parity now covers the full row for every state, a live parked child, an unlabeled child, and a lost or stale parent on both surfaces. * refactor(agent-status): the subagent snapshot, its normalization and equality get their own module agent-status-types.ts crossed the file-size limit once the row gained child views beside the main agent fact. The subagent snapshot shape, its admission normalization and the array equality move into agent-status-subagent-snapshot.ts (re-exported, so importers are unchanged); the three field caps it shares with the row move to the field normalization. * refactor(sidebar): one legacy builder family, one reader clock, frozen settled rows - The chat strip's task-roster rows are built in the shared row model beside the other two builders, with one placeholder set and the one detail rule. The module header states when each legacy builder is deleted. - A child's "No update" duration reads the parent's reader clock (receipt time for a mirrored parent); a view child's own stamp is moved onto that clock. The full sidebar row reads the same value as the compact row. - The failed->blocked / interrupted->idle collapse has one owner, shared by the sidebar row state and the strip header. - A settled strip row shows its run frozen at settledAt instead of a growing "ended N ago", so a strip of only finished work never wakes the 1 Hz tick. * test(sidebar): the sidebar row state and the strip header share one lifecycle word * test(sidebar): a child running a shell in its turn shows its Bash line with the shell nested beneath it While a child's turn runs its shell, the host records the shell both as the child's operation and as a live command the child owns. The strip shows the child working its Bash line with the shell row nested beneath it, the header counts only the agent, and the sidebar shows the child alone with the same text. * test(sidebar): re-pin the legacy chat strip golden to main's scoped goal-dock selector Main's #23725 rewrote the strip's goal-dock variants from `group-has-[...]/tasks:` to an ancestor-scoped `[[data-native-chat-background-tasks]:has(+[data-native-chat-thread-goal])_&]:` selector. The merged head renders byte-identically to main for this fixture; only the golden was captured before that change. |
||
|
|
357a2fed08 |
fix(native-chat): group chat rows by the turn that produced them (#23671)
* fix(native-chat): keep a turn's bar on the prompt that opened it A message sent while a structured turn runs appears in the transcript at once, so "the newest user message" is not the running turn's owner. The live "Working for" bar moved to the mid-turn message and counted from the earlier prompt's start, and a send queued behind the running turn counted its wait twice: once in the previous turn and again from its own send. Derive both from the host's turn records in one ordered pass: - The running turn's bar belongs to the user message its lifecycle row names (resolved exactly as settled timing resolves it). A message sent mid-turn gets no bar until its own turn opens; a send folded into the running turn never gets one. Surfaces fall back to the latest user message only when the host names no opener. - A turn counts from its send, but never before the previous turn in the journal ended (its recorded end, else its row's last host revision), capped at the turn's own start. The same origin feeds the live counter and the settled duration. Desktop and mobile share the derivation; no wire, host, or storage change. * feat(native-chat): derive each transcript row's owning turn from the journal Rows between a turn record and the next belong to that record's turn, so a message the provider folds into a running turn no longer captures the rows produced after it. A turn whose opener the host cannot name in the loaded window anchors to its own record instead of a bystander prompt, and shared nativeChatRowTurnKeys keeps positional preceding-user grouping for anything the host does not attribute (older hosts stay pixel-identical). * fix(native-chat): fold and time transcript rows by their owning turn on desktop A settled turn's bar now folds every row the turn produced, including rows after a mid-turn send; the steered bubble stays visible and carries no bar. A provider-opened turn renders its bar above its first row instead of borrowing the newest prompt, and row liveness follows the owning turn rather than the newest user message, so a running turn's rows stay live while a send waits. * fix(mobile): group phone transcript rows and bars by their owning turn Same shared derivation as desktop: the opener's bar owns every row of its turn across a mid-turn send, a steered bubble never grows a bar, a provider-opened turn's bar sits above its first row, and a running turn's tool rows stay live while a newer message waits behind it. * fix(mobile): declare the turn ownership map on the chat controller contract * fix(native-chat): one diff rollup per turn, and no wake-turn clock on a later prompt A turn's rows are no longer contiguous once rows are grouped by owning turn: a prompt sent before the running turn's last row lands among its rows. The diff rollup was drawn at every run boundary, so such a turn showed its rollup twice and the later prompt's rollup appeared under its bubble. It now renders once, under the turn's last row. On the phone, a turn keyed to its own record is not a user message, so when it ended its clock was treated as a replaced optimistic echo and handed to the newest prompt - a message sent during a wake turn got a bar with the wake turn's duration. Host-attributed turn keys now count as live turn keys. * fix(native-chat): keep a Codex turn's bar on its send until the echo lands Codex reports turn/started before it runs hooks and prewarm and before it echoes the send, so for that gap the turn names a provider key no alias resolves yet. Anchoring it to its own record left the running turn with no bar at all; treat the send still in flight ahead of the record as its opener. * fix(codex): restore each turn's record ahead of that turn's items Rows are grouped by the nearest turn record before them in journal order, which holds on the live path because a turn's record is written when the turn opens. Full-history restore (the fallback for Codex app-servers that reject excludeTurns) wrote each turn's items first and its record after, so every restored turn's rows were credited to the previous turn and turn 1's answer folded away. The restore now appends the record before the turn's items. The record itself is unchanged: same identity, state, outcome, opener key, endpoints and duration, one append each. The restore-order test now expects the record first, because that order is what keeps grouping correct; its old order was incidental, not a contract any reader relied on. Readers that key turns by id or opener key, or that scan for the newest record (all restored records are settled), read the same result in either order. Journals already imported in the old order stay as written; no migration. |
||
|
|
b7e3bbf3d1 |
fix(browser): restore hover after leaving a mobile viewport preset (#22846)
Turning touch emulation off sent maxTouchPoints: 0, which Chromium rejects (it only accepts 1-16, even when disabling). Touch emulation stayed on, so pages kept reporting no hover and a coarse pointer, and the desktop user agent was never restored. Omit maxTouchPoints when disabling so Chromium restores the original value. |
||
|
|
b8a7dafafc | fix(codex): guard the daemon socket in the home a resumed Codex pane launches in (#23724) | ||
|
|
f4829f0363 |
fix(renderer): stop the modal toast rule from freezing large diffs (#23721)
* fix(renderer): stop the modal toast rule from freezing large diffs The rule that lifts toasts above a dialog or sheet backdrop matched body:has(<overlay> anywhere). Chromium re-evaluates a descendant :has() on body for DOM changes anywhere in the page, so with large Monaco diffs open every editor mutation re-scanned the document and the renderer stopped responding. Dialog and sheet overlays portal straight into body, so matching them as direct children keeps the same behaviour while only changes to body's own children can affect the rule. * fix(renderer): raise modal toasts through a body variable, not a descendant :has() The child-combinator form still ran a whole-page walk after unrelated DOM changes (about 3x cheaper than the original, but still proportional to page size). Setting a custom property from a :has() on body alone takes the rule off that path; only opening or closing a backdrop restyles. * test(renderer): fail the toaster layering guard when the child combinator is dropped The guard only rejected a :has() inside the toaster's selector. Dropping the `>` from the body :has() (or removing the lift rule) still passed, and the unscoped form re-runs on every data-slot element change anywhere in the page, the same freeze the PR fixes. Assert the exact child-only rule shape. |
||
|
|
33f661b745 |
fix(native-chat): an empty workspace opens the default agent as a chat when chat is the default view (#23693)
* fix(native-chat): an empty workspace opens the default agent as a chat when chat is the default view With "open agent tabs in chat" on, clicking a workspace with no tabs still seeded a bare shell. The seeding path now opens the user's default agent through the shared launch funnel when that launch routes to a chat, and keeps the shell otherwise. A Blank Terminal pick at create time is now passed to activation as agent: null so it still gets a shell. * fix(native-chat): only a deliberate workspace open starts the default agent chat Round-1 review of the empty-workspace default chat: activation treated any call with no agent and no caller surface as user navigation, so CLI/phone creates, fallbacks after a failed agent launch, fork fallbacks, and the move to a neighbour after a delete all opened an agent chat nobody picked. - Opt in instead: activation options gain navigationIntent: 'user-open', set only by user navigation (sidebar row, keyboard cycling, Cmd+J, the workspace digit shortcut, back/forward, open-parent, jump-to-workspace, and the open-attached-workspace / space-manager actions). Every other activation keeps today's shell, so the Blank Terminal pass-through in the create flow is no longer needed and is reverted. - The Electron gated reseed now waits (bounded, 5 s) for the workspace host's agent list when it has not loaded, holding a per-workspace claim that the passive seed and other reseeds respect, then re-checks the active workspace, host, and emptiness. Host resolution reuses the detection target key, so an unresolved owner stays unknown. - Tests cover the gated path for worktrees and folders, non-opted activations, history navigation, the detection wait and its failure, and the passive seed deferring to the claim. * fix(native-chat): a reopen during the agent-list wait seeds for the latest open Round-2 review of the empty-workspace default chat: while a user open waited for the host's agent list, a second activation of the same workspace was dropped. Opening A from the sidebar, moving to B, then pressing Back to A during the wait left A active with no chat and no shell: the wait reseeded with the first open's host id, which no longer matched, and the passive seed had already marked A handled. - The pending wait now stores the latest activation's intent; a later activation (user open or plain) replaces it, and the wait seeds with it. - The passive seed no longer marks a workspace handled while a wait owns it, so leaving mid-wait and returning seeds a shell instead of nothing. - The wait and the open share one check, so a Blank Terminal default (or chat not being the default view) no longer waits up to 5 s before its shell. |
||
|
|
f8f656ca19 |
perf(ci): spend fewer concurrency slots per pull request (#23810)
A concurrency slot is charged per job, not per core, and the account's cap is the scarce resource: standard runner minutes are free and unlimited on a public repository. Two paths spent slots that bought nothing. The unit matrix ran eight fixed shards averaging 6.5 minutes each, 3384 job-slots a day and 68% of all slot demand, while the arm pool queued 10.5 minutes at p95 — the queue was the oversharding. Five shards run the same work in ~10.5 minutes each for three fewer slots per run. Bun profile persistence escalated to all six platforms on `config/`, `resources/` and `.github/` wholesale, which took 36.5% of the last 1100 commits through the full matrix where a platform-flavoured predicate takes 19%. A pull request now qualifies one platform unless the change is platform- flavoured, and the push to main re-qualifies all six, so an unescalated miss surfaces minutes after merge rather than at the next cron. Missing changed-file evidence and an unavailable dependency graph still fail closed to all six. |
||
|
|
5560e534ff | Update README downloads badge | ||
|
|
25d9c57e3a |
fix(native-chat): keep a child turn settling on an error Codex will not retry (#23808)
#23801 removed the child-path reading of Codex's turn-ending `error` along with the import of the module #23682 deleted. That was the wrong half to remove: the primary journal path can rely on Codex's failed `turn/completed` arriving within ~32 ms, which is what #23682 established, but a child turn has no such guarantee, and without the error as its end the child's lifecycle row latches on `working` for the life of the session. Three tests assert exactly that and could not run, because the unresolved import had been skipping the unit matrix since #23682 merged. The reading is restored inline against `readCodexErrorWillRetry`, itself restored to `codex-structured-thread-facts.ts`, rather than by reviving the deleted module: its `thread-stopped-running` arm lost its only consumer when #23682 rewrote the primary path, so restoring the file would re-add dead code. Also drops `pr-workflow-parallelism.test.mjs`'s read of `.github/workflows/track-community-prs.yaml`, which #23796 deleted while leaving the assertion behind. Same failure class, and it fails the same shard. |
||
|
|
2d358df8af |
fix(i18n): prune the wait-for-setup help copy #23799 removed (#23802)
#23799 removed both `waitForSetupBeforeAgentHelp` call sites but left the entry in all six locale catalogs. The runtime-required generator treats an entry with no literal-default call site as one only the catalog can serve, so the two dead entries had to ship in the boot bundle for the catalog check to pass, and `verify:localization-catalogs` failed on every pull request until they did. Deleting them is the resolution rather than regenerating `en-runtime-required.json`: no call site can reach either string, so shipping them would add dead weight to the boot bundle to satisfy a check about what the bundle must contain. The sibling `waitForSetupBeforeAgent` heading keys stay; #23799 removed only the help paragraphs. |
||
|
|
643f93f009 |
fix(native-chat): drop the Codex error path #23682 removed (#23801)
#23682 established that only `turn/completed` ends a Codex turn and deleted `codex-structured-journal-provider-verdicts.ts`, but the child-turn reader kept importing `readCodexProviderVerdict` and branching on its `turn-failed` verdict. The module is gone, so the import resolves to nothing: typecheck and static analysis fail on every pull request, which skips the whole unit matrix behind them, and the esbuild pass behind Bun profile scope detection throws, so every run falls back to all six platforms. A closed thread is now the only child-turn end without `turn/completed`, which is what #23682 intended: Codex follows a turn-ending `error` with a failed `turn/completed` for the same turn 0-32 ms later, and that completion carries the duration and receipt time the error does not. |
||
|
|
ceffecf12a | fix(setup): remove wait-for-setup helper text (#23799) | ||
|
|
f6324f242a |
chore(ci): stop auto-filing community PRs onto the project board (#23796)
Removes the Track Community PRs workflow. Community pull requests will no longer be added to project stablyai/13 automatically. |
||
|
|
2ea3fb1d46 |
perf(ci): take advisory unit-selection evidence off the gate (#23776)
selection_evidence is continue-on-error on both the job and its comparison step, so it can never fail a PR -- it downloads the shard reports, compares selection against the full results and uploads a review artifact. But a caller's `needs: test` waits for every job in the called workflow, so living inside unit-tests.yml it held verify for ~36s after the last shard finished. It moves to its own reusable workflow called as a sibling, so it still runs on every PR and still uploads its artifact, but verify no longer waits for it. It is deliberately absent from verify's needs, and a contract test pins both that and its advisory status so it cannot drift back onto the critical path. Measured on a recent run: the shards finished, then selection_evidence ran 36s, then verify 3s. Only the last of those gates anything. |
||
|
|
7445feaca8 |
fix(terminal): only diagnose disk exhaustion from capacity errors
Merged after required checks passed. |
||
|
|
18bd9cb1f6 | test(status-bar): cover re-probing a roomier level after the tightest level's fit moves (#23793) | ||
|
|
c8af48d8a4 |
fix(runtime): settle a quiet Codex composer as ready on every version (#23765)
* fix(runtime): settle a quiet Codex composer as tui-idle on every version Codex 0.158 dropped `model:`/`directory:` from its startup header, which both Codex readiness rules require, so `worker-start --agent codex` timed out; an idle Codex pane after a turn also had no readiness signal once the header left the screen. Generalize the Muse tier-1b lane into a quiet-ready-screen lane: a Codex (or agent-unknown) pane whose live screen shows the empty composer placeholder, no `to interrupt)` status row, no header `loading`, and no dialog wording in its live window, settles once the stream has been quiet for the tui-idle quiescence window. Additive only: the tier-1 rules and the Muse rule are unchanged. Fixtures: codex 0.150.1-0.158.0 captures at 120x40, including chunk-timed turns. * refactor(runtime): anchor the Codex quiet lane to the empty composer line Move the Codex screen rules into codex-terminal-readiness.ts and the quiet-screen body beside isKnownReadyPromptBody. The composer rule now matches only the `› Ask Codex to do anything` line and drops its dialog markers: every Codex dialog replaces the composer, and an answer ending "Would you like to…?" above a live composer must not hold the lane forever. The quiet lane checks quiescence before reading the screen. Trim the redundant startup and untimed turn fixtures. * fix(runtime): read Codex's busy row above the composer and scope the lane to codex panes * fix(runtime): read only Codex's live status row above the composer |
||
|
|
c30c8f9d77 |
fix(native-chat): a message Claude folds into its running turn no longer splits the turn (#23621)
* fix(native-chat): treat a folded mid-turn Claude replay as a receipt, not a turn boundary A message sent while a Claude turn runs is folded into the running turn by the CLI and replayed mid-turn with the client uuid. The replay-driven opener treated that replay as a new turn: the running turn was marked interrupted and the 'Worked for' bar split. The dispatch waiter now captures the open turn at write time (sentDuringTurnId, volatile); a replay that adopts the client uuid while that exact turn is still open settles delivery and opens no boundary. Plural result user_message_uuids settle each waiter under its own uuid; every other relation (miss, replaced turn, provider-resumed root, idle write, fresh replay uuid) keeps the opener path. Replay/turn resolution split out of the dispatch module to hold the line budget. * fix(native-chat): state the fold receipt's uuid-adoption limit as unmeasured The receipt admits only a replay that adopts the client uuid. The comment and test named a fresh replay uuid as "the CLI starting the queued send's own turn", but the measured miss case adopts the client uuid too, so adoption does not tell a fold from a later turn. Say what the rule actually is: the measured fold shape qualifies, and an unmeasured fresh-uuid replay keeps the opener path it always had. No behavior change. * fix(native-chat): decide Claude fold receipts from the provider's own request cycle Measured (p3 captures, CLI 2.1.280): the CLI folds any send that arrives while its request cycle runs — including one written before the first replay — and it announces every new cycle (sequential turn, queued turn, background wake, /compact) with a root system/init; each result names the sends its cycle ran in user_message_uuids. So the fold decision now reads provider state: an adopted replay while the open turn's cycle is still live (no root init since it opened) is a delivery receipt. The send-time bookkeeping (sentDuringTurnId) is removed; it missed the measured early-steer fold and guessed at what the provider states outright. A lost result is covered by the init staleness mark, and CLIs below the per-turn-init version floor (or with no reported version) keep today's opener path via an explicit gate on the init frame's claude_code_version. * fix(native-chat): read Claude fold membership from the cycle's own work A finished background task revises its row before the wake cycle's init, so the wake turn opened ahead of that init and was marked stale by it: a send folded into the wake still split the bar and interrupted the wake turn (captured p3-background-wake order). A send replayed after the current cycle's first root work (send echo or model output) is the fold; the cycle's first send is its opener. Init and every settle start a cycle with no work. * fix(native-chat): adopt the CLI version from any init frame, not only the startup proof Live sessions prove the session from a SessionStart hook frame that arrives BEFORE system/init, so startup facts read a frame with no claude_code_version and the fold receipt's version gate never passed: the real app still split the bar and marked the running turn interrupted while every fixture test — whose harness proves startup with the init frame itself — stayed green. The version is now adopted from whichever init frame carries it, at the same site that already adopts the per-turn model report, and the startup proof never clobbers a version a real init already reported. Pinned twice: a fixture run whose startup proof is the SessionStart hook frame, and a real-CLI fold test (skipped without a signed-in CLI) that fails on the unfixed branch at the interrupted assertion and passes with the fix. * test(native-chat): run the real-CLI fold test in the config dir it probed The connection strips an inherited CLAUDE_CONFIG_DIR and inherited auth vars, and with no launch env the pin compared against that same inherited value and emitted nothing, so the fold test always ran against ~/.claude whatever CLAUDE_CONFIG_DIR the availability probe checked. It also relied on the user's own SessionStart hooks for the live proof order and on their permission rules for the Bash steps; an isolated home hung at startup. The test now launches with the probed home and env auth, and pins a SessionStart hook and the sleep permission itself. * fix(native-chat): drop the CLI-version gate and prove the fold against full adapter captures The fold rule stands on frame-derived facts alone: an adopted client uuid (the capability check, read off the replay itself) while the open turn's cycle has done root work. The claude_code_version floor guarded only an unobserved triple fault — an adopting CLI without per-turn init AND a lost result — and its cost was a silent latch that already fired once live; all cliVersion plumbing is removed. Mid-turn auto-compaction was measured (forced via CLAUDE_CODE_AUTO_COMPACT_WINDOW): it emits status/compact_boundary and a synthetic continuation but NO root init, so the init cycle reset stays and the capture is pinned. The fake connection now defaults to the live startup shape (SessionStart hook proof; one init when the first command starts a cycle; capabilities on the initialize result), with init-at-startup an explicit unmeasured opt-in. Full frame streams recorded through Orca's own adapter against the real CLI are committed and replayed verbatim, asserting the provider's membership fact (each result's user_message_uuids) rather than design internals. * test(native-chat): drop the removed CLI-version gate from fold test comments Also reattaches the fake harness's initProof doc to initProof (it had landed on contextUsage, replacing that field's own doc) and renames a plural-result test whose title described a retired waiter it never creates. * test(native-chat): name the fold test's settings for their role * chore: take main's lockfile back after the merge * test(claude): route the real-CLI probe through the spawn chokepoint; give the slow-init test a live startup report The shared real-CLI availability probe moved out of a test file, so the child_process and CLI-runtime-pairing ratchets now scan it. It spawns through runProcessSync with the CLI paired to its own node, as the structured launch does. The provider-started test's CLI default now reaches startup the live way: get_settings reports it, since system/init only arrives with the first command. |
||
|
|
50a8ef18e4 |
fix(native-chat): keep a turn's bar on the prompt that opened it (#23573)
A message sent while a structured turn runs appears in the transcript at once, so "the newest user message" is not the running turn's owner. The live "Working for" bar moved to the mid-turn message and counted from the earlier prompt's start, and a send queued behind the running turn counted its wait twice: once in the previous turn and again from its own send. Derive both from the host's turn records in one ordered pass: - The running turn's bar belongs to the user message its lifecycle row names (resolved exactly as settled timing resolves it). A message sent mid-turn gets no bar until its own turn opens; a send folded into the running turn never gets one. Surfaces fall back to the latest user message only when the host names no opener. - A turn counts from its send, but never before the previous turn in the journal ended (its recorded end, else its row's last host revision), capped at the turn's own start. The same origin feeds the live counter and the settled duration. Desktop and mobile share the derivation; no wire, host, or storage change. |
||
|
|
c5fc0c6f26 |
fix(ci): keep a squash-merged RPC recording pin reachable through its pull request (#23720)
* fix(ci): keep a squash-merged RPC recording pin reachable through its pull request
Main's "RPC recording pin" check has been red since #22762: that branch pinned
the recording corpus to its own commit
|
||
|
|
e899809ff8 |
fix(mobile): unsubscribe session tabs by request on the direct connection (#22943)
* fix(mobile): unsubscribe session tabs by request on the direct connection * fix(mobile): hold a direct session tabs unsubscribe until the first snapshot The desktop registers a tab-list stream only as it emits the first snapshot. A direct unsubscribe sent before that found nothing, and with per-request addressing no later worktree-wide sweep collects the late stream, so it kept its desktop listener until the socket closed. Hold the unsubscribe until the snapshot arrives, as the relay connection does. * test(mobile): cover a held session tabs unsubscribe whose subscribe fails * fix(mobile): keep the session tabs hold within the registry line budget after merging main Move the pre-snapshot hold into the session tabs stream module, note that only older hosts need it, and give the unsubscribe test the real registration version now that a worktree-wide unsubscribe spares later streams. |
||
|
|
2ae7c00e84 |
fix(opencode): keep OpenCode 2 panes Working across plugin reloads (#23700)
* fix(opencode): keep OpenCode 2 panes Working across plugin reloads OpenCode 2 disposes and re-sets-up every plugin whenever its plugins dir changes, while sessions keep running. The status plugin published a final Idle on dispose, so a pane read Done mid-turn. Orca also rewrote the plugin file on every PTY spawn, so opening any terminal triggered that reload. Dispose now releases the factory's bookkeeping without publishing a verdict; the next lifecycle event settles the pane, and Orca's ended-process reconciliation still retires panes whose agent exited. The plugin file is written only when its bytes differ. * fix(opencode): skip rewriting an unchanged plugin in the SSH relay install too The relay's canonical-config install still unlinked and rewrote the status plugin on every OpenCode launch over SSH, which restarts every plugin in a remote OpenCode 2 server. Share one install-currency check (lstat + the existing byte comparison) between the local and relay writers, and pin write-if-changed with mtime so the tests also fail on filesystems that reuse a freed inode. * fix(opencode): keep the final Idle when OpenCode 1 tears its instance down OpenCode 1 disposes a plugin only when it tears the instance down, and that teardown cancels every running session, so the Idle published on dispose is true there; the cancelled run's own idle may never reach the plugin. Only OpenCode 2 disposes on a hot reload while turns keep running. The generated module serves both hosts, so the OpenCode 2 setup() entry point now tells the shared factory that sessions outlive disposal; the server() path keeps the previous disposal behaviour, including the hand-off to a surviving factory. * fix(opencode): compare a symlinked plugin by its target before rewriting OpenCode 2 loads plugins through file-level symlinks and reads the revision from the target's mtime, so a user whose Orca plugin file is a symlink (per-file dotfile managers) failed the regular-file check and got a write through the link, and a reload, on every spawn. The config-dir and relay installs now skip the write when the resolved target already has Orca's bytes; when stale they behave as before. Only the per-source overlay keeps the regular-file check, since a link there mirrors a user entry. Installers also skip the write inside a guarded block rather than returning early, so later install steps still run. * test(opencode): skip the plugin symlink tests on Windows like their neighbours Creating a file symlink on Windows needs Developer Mode or admin rights. * test(opencode): stub fetch without a type assertion in the dispose host test |
||
|
|
3976ad4c59 | perf(test): remove obsolete structural snapshots (#23777) | ||
|
|
b776e9ac99 |
fix(browser): give a tab's identity one owner so viewport presets stop dropping client hints (#23718)
* fix(browser): give a tab's identity one owner so viewport presets stop dropping client hints A desktop viewport preset installed a CDP user-agent override with no userAgentMetadata. Chromium then drops navigator.userAgentData and every sec-ch-ua header for that tab: a Chrome UA with no client hints. Identity was decided separately by the session request hook, the Google sign-in switch and the viewport code, and nothing decided per tab who it should claim to be. resolveBrowserTabIdentity now derives it from the process identity mode, whether the URL is a Google auth host, and whether a mobile preset is requested. applyTabIdentity is the one writer: it keeps the WebContents UA on the process or Firefox identity and clears the CDP override whenever that layer already presents the identity. Viewport emulation only records the requested preset; its metrics and touch steps log failures independently, so a rejected step can no longer skip the identity restore. * test(browser): read the presented identity instead of casting the guest stub * test(browser): drop a comment that described desktop presets writing a UA * fix(browser): keep same-document navigations and unapplied presets off the tab identity A same-document navigation (pushState/replaceState) now never rewrites the WebContents user agent. Chromium reloads a still-loading document when its user agent changes, so an OAuth callback that strips its code with replaceState after a redirect off Google sign-in was requested twice, replaying the one-time code. Measured on Electron 43.7.5: the callback URL hits the server twice with the write, once without. The session request hook now derives the mobile identity from the CDP override the tab actually holds instead of the requested preset. A preset whose write never landed (debugger attach refused while DevTools is open, a failed write, a detach) no longer puts the iPhone user agent and mobile client hints on the wire while the document reports desktop. * fix(browser): restore identity after a failed navigation without reloading the error page did-fail-load fires while the failed URL's error page is still loading, and WebContents.setUserAgent() at that moment makes Chromium reload it. After a redirect onto or off the Google sign-in host (identity moved over CDP only), the restore rewrote the WebContents UA there and replayed the failed request. The restore now goes over CDP; the next navigation rewrites the WebContents UA. * refactor(browser): let only a navigation start write the WebContents user agent Two review rounds each found a caller that asked the identity writer to rewrite the WebContents UA at a moment Chromium reloads or cancels the page (a same-document navigation, a failed load). A boolean at every call site left that decision to the callers. The writer now has two entry points: presentTabIdentityAtNavigationStart, the only one that may write the WebContents UA and only for a cross-document navigation, and retargetTabIdentity, which goes over CDP only and serves redirects, failed loads and preset changes. A table test pins the rule for every entry point. * test(browser): reject touch emulation regardless of payload in the identity-restore test The mock rejected only maxTouchPoints 0, so the test would stop exercising a failed touch step once the touch payload is fixed. |
||
|
|
b4c19f12c4 |
fix(claude): run structured Claude under the POSIX provider supervisor (#23476)
* fix(codex): the provider supervisor outlives its provider group when stopped A signalled supervisor forwards the signal to the provider group, escalates to SIGKILL after the grace, and exits only once the group is gone, so recovery's proof that the recorded pid is dead also proves the provider is. It refuses to spawn when its parent is already not the owner named in its spec, and watches that owner rather than whichever parent it first saw. The grace is a spec field. Recovery's SIGTERM stage now outlasts the supervisor's own stop, since a SIGKILL that lands first cannot be handled and leaves the group running. * fix(codex): a closed owner pipe no longer ends the supervisor before its provider group When Orca dies, the supervisor's stdout pipe has no reader. Provider output in the window before the parent-death watch fired raised an unhandled EPIPE that exited the supervisor with the provider group still running. * fix(codex): bound the supervisor grace so recovery's SIGTERM stage always covers it Recovery sized its SIGTERM stage from the default grace, so a launch with a longer grace would be SIGKILLed mid-stop and orphan its group with no test noticing. The spec now refuses any grace above one exported maximum, and recovery derives its SIGTERM stage from that maximum. * fix(codex): every supervisor stop asks the provider with SIGTERM first Owner death, stdin end after the grace, and a signal to the supervisor now all take one path: SIGTERM the provider group, SIGKILL it after the grace, and exit only once it is gone. The signal handlers are registered before the provider is spawned, so a stop that lands in the spawn window still reaps it. The longest stop grows to two graces plus the reap wait, and both recovery's SIGTERM stage and the connection's graceful close now wait that long before forcing, since forcing the supervisor sooner can orphan its group. * fix(claude): run the structured Claude child under the POSIX provider supervisor A close now stops Claude with a SIGTERM through the supervisor instead of letting stdin end finish the turn, and Orca's death stops it through the supervisor. * test(claude): pin the supervised stop against a real Claude CLI, opt-in * test(claude): a requested stop reads interrupted through the frames the supervised SIGTERM makes Claude emit * test(claude): show what the real CLI did when it never ran the tool * test(claude): Orca's death now reaps Claude's own tool through its SIGTERM * refactor(claude): take supervision from the spawn spec so the close ladder cannot disagree with the spawn createProviderSpawnSpec now reports whether it wrapped the provider in the supervisor, and the Claude spawner reads that instead of repeating the platform check. The close ladder's SIGTERM follows the process actually spawned. * fix(native-chat): derive quit's chat-eviction bound from the longest supervised provider close Quit's child-eviction phase was a hand-picked 8 s. It is now the sink drain plus the longest supervised close over Claude and Codex plus a named 1 s margin, so a provider close that grows widens it instead of silently outrunning it. A close's tree-kill fallback stays outside the bound: once main exits, the supervisor stops its provider group on owner death, which a new test now proves for a clean owner exit, and next launch's recovery settles the lease. |
||
|
|
2dc2693953 |
fix(native-chat): a turn a proven crash cut short reads interrupted (#23456)
* fix(native-chat): a turn a proven crash cut short reads interrupted, ending when it was last seen working * fix(native-chat): end a probe-proven turn at the last row the journal wrote live A revised item keeps its first sighting's timestamp, so a long command or a streamed reply read as ending when it started. The reducer now tracks the latest live row the same way it tracks all activity. * test(native-chat): give the unexpected-exit fake journal its live-activity read * refactor(native-chat): read the journal's live bound only for a probe-proven death * fix(native-chat): mark what a journal open settles for a gone host as crash reconciliation A crashed host's working roster is retired when the journal reopens. That row was written live, so a probe-proven turn ended at the relaunch and counted the downtime. * fix(native-chat): bound a probe-proven death with the last time Orca proved the owner alive A crash mid-tool left the turn ending at the tool call's start, because Claude writes nothing while a Bash call runs. The death evidence now records the lease's last renewal before the death as lastProvenAliveAt, and the turn ends at the later of that and the last live row, capped at the probe. Parking a lease in recovery no longer stamps lastRenewedAt, since nothing proved the owner alive then; a child that outlived Orca would otherwise have its turn count the downtime. * test(native-chat): a failed acquisition parked in recovery keeps its last proof of life, and the timing read goes through the display selector * refactor(native-chat): move the submission dispatch folds out of the journal reducer Main grew the reducer to its line limit, so the live-activity bound tipped it over. The dispatch row and echo-acceptance folds move unchanged into their own module. * test(native-chat): a send or reader that opens a crashed chat settles a proven death interrupted Main's open-time settle test still asserted the old rule, where only a watched exit proved a death. * docs(native-chat): say which proofs of death record a last proof of life * fix(native-chat): a proof of life bounds only the owner that wrote the turn A start after a crash that spawned a child and then failed without proving it gone parks that child for recovery; when recovery finds it gone, the proof of death carries its proof of life, which is after the crash. The older turn then ended there and counted the downtime. The journal now derives the fence of its newest live writer, and the lease that holds the proof names the owner it released by the fence it moved to. The last renewal counts only when that move was one step past the writer of the turn. * test(native-chat): a crash with a send in doubt still ends at the last proof of life The reopen settles that send at the new fence, so the owner check must read the fence of live rows only. * fix(native-chat): a proof of death judges only the turn its own owner wrote The settle read the record's latest proof of death for whatever turn a gone generation left running. After a crash, a start that reserved a new fence cleared the relaunch's proof, and if it then failed, its own child's death (a watched exit at the failure, or a probe finding the child it left for recovery gone) ended the older turn an hour after the crash. Every proof of death now records ownerFence, the fence of the owner or reservation it is about; a fence names exactly one owner. The journal derives the fence each item was created at, and a running turn is interrupted only by a proof naming its own owner; otherwise it is unverifiable. Evidence older builds wrote keeps their rule. This replaces the derived one-step fence check. * test(native-chat): every writer of a watched exit names the owner it released A watched exit that names no owner reads the older rule, so the settle alone cannot tell a dropped field; the writers are pinned directly, including past a recovery floor. * test(native-chat): give the fake journal's cast its safety rationale * fix(native-chat): a proof of death written after a chat opened revises the turn it left unverifiable On desktop the chat on screen at relaunch opens before the startup reconcile has probed its owner, so the open settles the cut-off turn unverifiable. When the reconcile then records the proof, it re-runs the same settle for every open conversation, which revises that owner's unverifiable turns to interrupted with the proof's end. Any later open re-runs it too, so a failed write converges. Only upward, only for a proof that names the turn's own owner. * test(native-chat): a chat read before the reconcile reads unverifiable, then interrupted Covers the reconcile revising an open chat to the last renewal (27 s) and a subscriber being sent both states, a start after the crash whose running turn the queued revision leaves alone, a failed revision write converging at the next open, a proof about another owner or from an older build never revising, and a second settle writing nothing. * fix(native-chat): revise an open chat's turn wherever a proof of death is written The store tells its listeners, once committed, of each record a transaction gave a new proof of death, so every writer (the startup reconcile, a recovery that stopped a child which outlived Orca, a failed start, a watched exit) triggers the same serialized resettle for a chat already open. The reconcile's own callback is gone. Quit stops listening first, and a queued resettle is drained with the starts. * test(native-chat): a failed exit settlement is retried in place once the exit is recorded Recording a watched exit now queues the same settle an open runs, so the turn converges without waiting for the chat to be reopened. The reopen and read-after-restart cases now refuse that retry too, so they still pin the open's own settle. * test(native-chat): tests that pin a send settling a failed exit refuse the in-place retry too The exit's release now queues the same settle, so two tests named for the send's settle refuse that retry as well; the comments that said nothing retries it now say what does. * fix(native-chat): name the explanation row by the death it explains, so a retried settle adds no second row * fix(native-chat): end a crashed turn at its owner's provider output, never at a later send A send accepted into a crashed chat before the proof of death wrote a submission row live, and the journal-wide last-live-row bound counted it, so the revised turn ended at the send and counted Orca's downtime. The bound is now the last row the owner's provider child wrote, per writer fence: submissions, dispatch rows and crash reconciliation are Orca's or the user's, and a newer owner's work says nothing about the dead one. * fix(native-chat): end a crashed turn at its last proof of life, never at a timeline row The end of a probe-proven death was the later of the last renewal and the last live timeline row. A send accepted into a crashed chat before the proof writes a row live, so the revised turn ended at the send and counted Orca's downtime. Rows cannot tell the agent's output from Orca's or the user's, so the end is now the last renewal alone, never after the probe, and never before the turn began. The journal's live-activity bound and the reopen's recovered marker, which existed only for it, are gone. * test(native-chat): drop a stale reference to the removed live-activity bound |
||
|
|
b283a09688 |
fix(native-chat): stop flashing "still starting" on every chat launch (#23666)
* fix(native-chat): stop flashing "still starting" on every chat launch Every structured chat passes through a short startup phase, and the pane showed "<agent> is still starting…" for all of it, so a normal launch flashed the notice for a fraction of a second. The notice now goes through a keyed delayed status: it appears only once startup outlasts a grace period, stays up for a minimum time once shown, and resets per session. * fix(native-chat): reset startup notice for each provider child |
||
|
|
94b5a7d256 |
fix(native-chat): only Codex's turn completion ends a Codex turn (#23682)
* fix(native-chat): the sink queue keeps a settlement's first batch, as the journal does The journal applies a lifecycle batch's settlement id once and skips any later batch with the same id. The deferred sink queue coalesced the same key the other way: a second batch replaced the first while it was still queued. So which record survived depended on whether the first had drained yet. A lifecycle batch now keeps the queued operation with its key, and a later one is accepted and dropped, which is what the journal does once the first is written. * fix(native-chat): only turn/completed ends a Codex turn Codex follows every turn-ending `error` (willRetry=false) with a failed `turn/completed` for the same turn, 0-32 ms later. That was captured from the real app-server on 0.141.0 and 0.158.0 across eight failure scenarios, and it is how Codex builds a failed turn: it records the error as the turn's last error, records any pending input, and then derives `failed` from that error when it completes the turn. The translator ended the turn twice: once on the error, and again on the completion, with a guard to make the first end final. Ending on the error threw away what only the completion carries: Codex's duration, and the completion's receipt time. It also forgot the turn before Codex recorded the turn's pending input. Now the error is only the row the user reads, inside the still-open turn, and `turn/completed` is the turn's only live end. A process exit between the two is the existing exit sweep's observed end, recorded as interrupted. A failed completion is stored as completed with outcome failure, live and on restore alike. Only `interrupted` maps to the interrupted state. The first-end-final guard is gone. Codex sends one completion per turn, the only redelivery Orca has is the retry of a refused frame (which changes nothing), and the settlement id already keeps the first record in the queue and the journal. * refactor(codex): delete the unreachable oversized-notification settlement The translator settled a streamed item when the transport rejected its notification as oversized. Nothing can produce that frame. The Codex stdio reader frames with `maxLineBytes: Number.POSITIVE_INFINITY` (codex-app-server-record-reader.ts), which it has done since the app-server records were uncapped. With an infinite limit the framer never reports `line-too-long`: no line, pending suffix or paused queue can exceed it. So the dispatcher never emits `frame:oversized-notification`, and the arm that settles it never runs. The arm, its helper module and its test go. In place of the test, the connection test now proves the reason: a notification past the old 16 MiB wire limit arrives whole, and no oversized frame is reported. * test(codex): replace the captured ids in the turn-endings fixture with synthetic ones The replay reads ids only to group frames, so the real thread, turn and response ids from the capture account carry nothing the test needs. The fixture moves beside the Codex tests that read it. * test(codex): use a neutral made-up status as the unknown-status example 'cancelled' read as a stop being recorded as a completion. * test(codex): a restored turn with a status Orca cannot place ends with no verdict Codex's history carries the same status field as the live completion, so the restore path is pinned to the same mapping: completed, and no outcome. |
||
|
|
a5ce8251e3 |
Agent launches carry the surface that started them (#23697)
* feat(agent-launch): every launch carries the surface that started it The host now attributes every agent it builds to the surface that asked for it, resolving a missing or unrecognized surface to 'unknown' in one place instead of silently skipping it. The CLI names itself on worktree.create and orchestration workers name themselves host-side. * fix(agent-launch): attribute the agent a startup-draft create launches The host builds a third kind of agent launch: a worktree.create with a startupDraft and no startupAgent, where the host picks the agent itself. It carried no launch record at all and ignored the caller's launchSource. Route it through the same resolver as the other two builders, and derive the startupAgent terminal record only from the resolver so no prebuilt record can stand in for it. * fix(agent-launch): attribute the agent a host-built agent session launches terminal.createAgentSession builds a fresh agent's launch on the host, like the other startup builders, but spawned it with no launch record, so those launches were never counted. Record them through the same resolver; the request names no surface, so they count as unknown. * test(agent-launch): require an attribution decision for every host-built agent startup |
||
|
|
85ad292930 | fix(status-bar): re-measure collapsing levels when only the collapsed width moves (#23773) | ||
|
|
a134d1259e |
feat(mobile): tell users when a newer app binary is installable (#23755)
* feat(mobile): tell users when a newer app binary is installable With OTA page updates, store releases get rare and users stop looking. The shell now asks the channel that installed it. Android sideload reads GitHub's mobile-android-v* tag refs and proves the release has an APK. iOS reads the App Store lookup. A home card above Desktops, dismissible per version, and Settings rows surface the result. The check runs on the desktop updater's cadence: cold start, foreground once 24 h have passed, and a 1 h retry after a failure. The releases atom feed was not used because it lists only the 10 newest releases, which are all desktop builds, so it never carries a mobile tag. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): parse update replies with zod schemas The anti-slop gate refuses Reflect.get on dynamic input. The GitHub refs, the release, the App Store lookup and the stored update record are now parsed into named schemas before they are read. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * docs(mobile): say why the Android update source reads tag refs Record why the Android source reads tag refs. The releases atom feed and /releases?per_page=100 are both newest-first windows that desktop releases fill. Either would silently report "current" once a run of desktop builds pushes the newest mobile release out. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): load update state once and apply review rulings Every check and dismissal now awaits one shared store load. This replaces the merge that guessed whether a check had landed during the load. A manual check that fails while the store loads therefore keeps its 1 h retry instead of re-checking at once. - checking is derived from the in-flight check. - start() uses a per-start flag, so a StrictMode double start applies one load. - A check that finishes after stop() writes nothing. - A corrupt stored update record loses only itself. - Tag refs are parsed with a single schema. - The runtime wiring is folded into one file, and the card moves to home/. - The recorded App Store fixture is oxfmt-formatted, with the same parsed value. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): keep the update timer armed across a stop and restart A check that stayed in flight across stop and restart returned 'failed' without rescheduling. The restart skipped arming because a check was in flight, which left a live checker with no timer until the next foreground. The stop counter is removed. schedule() already arms nothing while no start is active, and saving a real result after a stop is harmless. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * chore: retrigger CI after the RPC recording repin (#23757) landed on main Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): trim the update checker and Settings rows - The load sets prefs and the due time only. start() re-arms the schedule after it. - The Settings result hides through one effect keyed on the result. - onUpdate receives the URL. - The version row is bound once. - The retry and timeout constants are no longer exported. - The unused AppUpdateChecker type is deleted. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): run the update check when its timer fires The armed timer is the due time. Re-checking the wall clock when the timer fired meant a clock stepped back skipped the check and re-armed nothing. The due-time guard now applies only on foreground. The binary version still comes from expoConfig.version. SDK 55 removed Constants.nativeAppVersion, so the no-expo-updates invariant is now named in the comment. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): recover from a future check time and use Apple's page URL If the device clock was ahead when a check ran and was corrected later, the stored check time is in the future. Cold starts then armed a timer for the whole skew, and foreground never came due. The stored state now reads as never checked in that case. The iOS link is the lookup's trackViewUrl instead of a URL built from trackId. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * style(mobile): fit the future-check-time comment in the print width Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
ec9f35e2ee |
perf(ci): plan the unit shards before the static-analysis gate instead of behind it (#23743)
A caller's `needs` gate the whole called workflow, so while the plan job lived in unit-tests.yml it could not start until static analysis and typecheck had both finished and passed -- and the shard matrix then waited on it. The two hops were serial when they did not need to be: planning reads the checkout, a git diff against HEAD^1, the import graph and the checked-in timing baseline in config/scripts/ci-shard-timings.json, and consumes nothing that static analysis, typecheck or the native-cache primer produce. Planning moves to its own reusable workflow so pr.yml can run it against code_paths alone, overlapping it with the gate. Measured across 99 runs, the shard matrix is created a median 93s earlier (p25 47s, p90 241s, never later). Planning stays a required predecessor of the shards, so an empty assignment cannot expand the matrix. The gate itself is deliberately left in place. It fires on 22% of runs, and the shard queue wait knees hard above ~9 concurrent ARM jobs -- 4s median below that against 218s at 15-19 -- so admitting 8 doomed shards per failed run would cost more in queue pressure than it returns in latency. Cost is one 37s ubuntu-latest job, which does not touch the ARM pool the shards contend for. A planning failure still fails the PR: the shards are skipped, and verify's check_job requires success whenever the classifier says tests should run, so it reports `test: expected success, got skipped`. |
||
|
|
28942eed5a |
test(mobile): repin the RPC recording corpus to main after #22762 (#23757)
#22762 squash-merged as |
||
|
|
ccdb324b63 |
Add CodeBuddy as a built-in coding agent (#23740)
* feat(agents): integrate CodeBuddy launch, status and session history * docs: record CodeBuddy lifecycle verification * fix(codebuddy): backfill scoped history and negotiate remote resume * test(cli): include CodeBuddy in known search agents |
||
|
|
da48d98040 |
fix: bound combined diff editors and scope chat style invalidation (#23725)
* fix(editor): bound offscreen combined diff rendering by height * perf: scope native chat relational styles to their ancestor |
||
|
|
d606be3ade |
test: wait for remote terminal grid convergence after reveal (#23738)
* test: wait for revealed remote PTY grid convergence * test: retain reveal diagnostics on geometry failure |
||
|
|
bd5dca4406 | fix(editor): preserve combined diff scroll on line focus (#23735) | ||
|
|
8bb78f0ddd | test(browser): create probe body before starting frame requests (#23730) | ||
|
|
64dbe87de6 | test(terminal): wait for decoy panes before host parking (#23729) | ||
|
|
4b622e1b13 |
fix(runtime): reopen the quiet-foreground tui-idle lane for agents with no other rest signal (#23598)
* fix(runtime): reopen the quiet-foreground tui-idle lane for agents with no other rest signal A tui-idle wait could never settle on a pane running amp, goose, crush, kimi, qwen-code, rovo, auggie and other agents whose titles Orca cannot classify: the quiet-foreground lane was closed for every launched agent, and it was the only lane those agents could reach, so worker start failed at agent_readiness after 60s. Model each agent's rest signal, derived from the tables that already encode it (synthetic ready titles, the title classifier, the DSH hook and Muse ready screen lanes). The lane stays closed where a stronger signal will arrive and reopens for agents with none. On a reopened lane, silence counts only after the TUI has painted: an agent that has painted nothing is still booting. Linear: STA-7440 * fix(runtime): count only the command's own output as an agent's paint on the tui-idle foreground lane The after-paint lane accepted any output, and the shell's prompt and echoed launch command always land before the agent starts, so a silently booting agent could still settle and lose its first prompt. The runtime now reads the shell integration's command-start marker and requires visible output after it; panes whose shell emits no marker keep the any-output rule. Also skip the foreground-process read while the pane cannot settle, and register the new title-classifier call site in the pane agent identity inventory. * fix(runtime): classify Freebuff's rest signal and skip the backward marker scan on chunks without one Main added the Freebuff agent after this branch point, so the full per-agent rest-signal table no longer matched on the merge ref. Freebuff derives `none`: its screen reports a first-party `done`, which tui-idle trusts only for DSH, so the quiet-foreground lane is its only one. The command-start scan ran a backward search over every PTY chunk; a forward check first cuts that to the cost of a plain substring test on chunks with no marker. * fix(runtime): classify Qoder's rest signal after merging main Main added Qoder with its own readiness branch returning a boolean quiet-foreground flag; map it to the lane type and classify Qoder by its ready screen so the full rest-signal table and lane-agreement check stay exhaustive. Say what `none` actually means: no stronger lane tui-idle trusts, not no hooks at all. * refactor(runtime): track command paint with the shared OSC 133 scanner The command-paint tracker had its own split-unsafe 133;C parser. Reuse the chunk-boundary-safe scanner, which now reports where in the chunk the marker ended, so a marker split across reads is still found. Correct the unmarked-launch list: bash and zsh mark typed launches after the echo. * fix(runtime): drop command-paint state on an output gap or a new process A dropped chunk can cut a command-start marker in half, and the scanner's carry then completes it on unrelated output after the gap, leaving the pane waiting for a paint that already happened. Reset it with the other cross-chunk carries. * fix(terminal): keep the command-start offset out of renderer lifecycle callbacks |
||
|
|
2aed2cf64a |
fix(terminal): reveal splits while the source pane binds (#23692)
* test(e2e): observe passive terminal restoration before activation * fix(terminal): reveal persisted splits during source binding publication * test(terminal): handle nullable persisted layout roots |
||
|
|
d0db35c18c |
fix(terminal): preserve typing while a remote pane reattaches (#23701)
* fix(terminal): retain typing while a parked remote pane reattaches * test: persist restored remote terminal screenshots * fix(remote): buffer recovery reconnect input * fix(remote): retain input across restored pane attach * fix(remote): flush attach input after subscription * fix(remote): flush reattach input after attach readiness * fix(remote): stop buffering after reattach readiness * test(remote): trace parked reattach input lifecycle * test(remote): forward paired client lifecycle diagnostics * fix(remote): preserve restored typing before connect starts * chore(i18n): refresh runtime required catalog * fix(i18n): ship compact agent runtime label * fix(i18n): merge required label into existing sidebar catalog |
||
|
|
d68eee3757 |
fix(runtime): retire an exited terminal before its stream end (#23492)
* fix(runtime): retire an exited terminal before its stream end An exit's durable retirement became asynchronous, so onPtyExit released the terminal stream before the retirement landed. A paired client answers a stream end by re-activating its pane; that activation still found the exited leaf, materialized it under the same session id, and registerPty dropped the pending retirement. The exited split pane came back as a fresh shell. The exit now stages the retirement into the in-memory session and publishes it synchronously, then notifies exit listeners, and only then makes it durable. A failed durable write is logged and left in memory for the next profile write instead of being rolled back, since the process is gone either way. This removes the pending-retirement latch and its post-await incarnation fence: there is no longer a window for them to guard. * test(runtime): a failed exit retirement still reaches disk Pins the no-rollback contract through a real Store and SQLite authority: when the retirement's own durable write fails, the in-memory retirement is carried by the next unrelated profile write, and by the app-quit flush when no other write happens. The delayed authority fixture can now fail its next write, and the acknowledged-retirement fixture reads the database a relaunch would load and models the quit flush. * test(runtime): a stream end observes the exit retirement already published The re-activation check alone passes with the listener ordering reverted, because activation awaits before its lookup. Record the session binding and publication count at the moment the exit listener fires so the ordering itself is pinned. * fix(runtime): an exit cleanup fault still ends the terminal stream * perf(runtime): exits retired together share one durable write * test(runtime): a refused staging write still retires the pane and ends the stream * refactor(runtime): describe exit retirement as staged, not durably accepted The retirement result is staged in memory before any write, and the removable-surface comment and the replacement-admission test name still described the old publish-after-durable rule. |
||
|
|
2af897d7ea |
fix(codex): the provider supervisor outlives its provider group when stopped (#23466)
* fix(codex): the provider supervisor outlives its provider group when stopped A signalled supervisor forwards the signal to the provider group, escalates to SIGKILL after the grace, and exits only once the group is gone, so recovery's proof that the recorded pid is dead also proves the provider is. It refuses to spawn when its parent is already not the owner named in its spec, and watches that owner rather than whichever parent it first saw. The grace is a spec field. Recovery's SIGTERM stage now outlasts the supervisor's own stop, since a SIGKILL that lands first cannot be handled and leaves the group running. * fix(codex): a closed owner pipe no longer ends the supervisor before its provider group When Orca dies, the supervisor's stdout pipe has no reader. Provider output in the window before the parent-death watch fired raised an unhandled EPIPE that exited the supervisor with the provider group still running. * fix(codex): bound the supervisor grace so recovery's SIGTERM stage always covers it Recovery sized its SIGTERM stage from the default grace, so a launch with a longer grace would be SIGKILLed mid-stop and orphan its group with no test noticing. The spec now refuses any grace above one exported maximum, and recovery derives its SIGTERM stage from that maximum. * fix(codex): every supervisor stop asks the provider with SIGTERM first Owner death, stdin end after the grace, and a signal to the supervisor now all take one path: SIGTERM the provider group, SIGKILL it after the grace, and exit only once it is gone. The signal handlers are registered before the provider is spawned, so a stop that lands in the spawn window still reaps it. The longest stop grows to two graces plus the reap wait, and both recovery's SIGTERM stage and the connection's graceful close now wait that long before forcing, since forcing the supervisor sooner can orphan its group. * fix(codex): give the provider 1 s after stdin end and 3 s after SIGTERM to flush before SIGKILL The supervisor's stop was stdin end, 1.25 s, SIGTERM, 1.25 s, SIGKILL. Codex now gets 3 s after SIGTERM to flush its state. The two graces are separate constants, the longest stop they derive becomes 5.5 s, and a test keeps it inside quit's 8 s child-eviction bound. * test(codex): count eviction's pre-stop drain in the quit budget test Eviction drains the sink for up to 1 s before it stops the child, inside the same 8 s bound. |
||
|
|
64569a8183 |
fix(ssh): don't overwrite remote agent config after a failed read (#22644)
* fix(ssh): don't overwrite remote agent config after a failed read A flaky read was treated as an empty file, wiping the user's config. Fixes #22638 * test(runtime): model remote missing-config reads as relay ENOENT errors The runtime harness stubbed isENOENT as code-only, and the remote Codex startup specs rejected with a generic error that only passed while any read failure seeded an empty config. Use the real isENOENT and the message-only shape the relay actually delivers. --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
9b11b1d594 |
fix(agent-hooks): compare status rows structurally instead of serializing both (#23585)
Status-row change detection stringified two full IPC payloads on every status write, including an up-to-8 KB lastAssistantMessage re-posted on every OpenCode streamed part. Compare the same published field set with the existing structural-equality helper, with a same-reference fast path. Linear: STA-7432 |