mirror of
https://github.com/stablyai/orca.git
synced 2026-10-09 08:02:35 +00:00
dff65d55a3eae91c64d791d2cb1265f8f9f2903e
12941
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
dff65d55a3 |
feat(claude): prepare account profiles and shared history (Step 1 of 4) (#24300)
* feat(claude): add dormant profile setup and history sharing * fix(claude): make profile setup one gated, typed, fail-safe entry Review round 1 of the dormant profile setup found that the pieces could be called without their safety checks, that one failed write or an unreadable bookkeeping file could silently stop sharing for good, and that Windows prompt history could bring back history the user cleared. - One entry, provisionClaudeAccountProfile: the profile gate (namespace, no linked components, outside ~/.claude and ~/.config/claude, and an ownership marker beside the home naming the account and target) runs first and refuses before creating anything; then history sharing, config provisioning, and the hook install after the settings merge. Results come back per surface with closed warning codes instead of message text. - The sharing ledger is keyed by surface name, records a value only after its write succeeded, and an unreadable ledger starts empty and is rewritten instead of blocking every surface. - The profile state file goes through the same locked writer as folder trust (Claude's <file>.lock plus the in-process queue), generalized as updateClaudeGlobalConfig. Onboarding and trust are still applied when the personal state file is unreadable. - WSL descriptors build guest POSIX paths; the state-file path style follows the injected platform. - Orca's managed statusLine has one owner in a profile: the settings merge never shares it, a user's own statusLine is shared over it, and the profile installer follows the default home's slot so a default opt-out reaches every profile. remove() takes the same destination; the remote installer cannot accept one. - Prompt history compares file identity (bigint dev+ino) on every platform, never drains the shared file into itself, drains retained copies in generation order, never reuses a stale cursor, and on Windows keeps a replaced default's old copy aside instead of replaying it. Directory merges keep going past a failed entry. * fix(claude): share the user's own hooks and keep merged history whole A user's own Claude hooks in ~/.claude (notifications, formatters) did not run under a managed account, because the whole hooks key stayed private. They are now shared like any other settings key: Orca's own hook entries and its managed statusLine are stripped from both the personal value and the profile's current value before the per-key ledger comparison, so they never travel through the merge and never make the key look user-owned. Orca entries already in the profile are kept on write, and the profile hook installer adds them on top as before. Prompt history: merged bytes that lack a final newline are terminated, so Claude's next record no longer fuses onto the last merged line. When a CLI rewrote the profile's history file (old records plus new), only the lines past the part it shares with the default history are added, instead of the whole file again. * fix(claude): close review round 2 gaps in profile setup Hooks and statusLine sharing: - When ~/.claude holds only Orca's hook entries, the user's shared hooks now read as an empty value instead of a missing key. Removing the user's last own hook in ~/.claude therefore reaches profiles that never edited it, and deleting the only shared hook inside a profile stays deleted. - A custom statusLine Orca shared, and the profile never edited, goes away when the default home drops it. When a shared custom line replaced Orca's line in a profile, the profile's statusline marker is dropped so Orca's line comes back once the default returns to it; a profile that opted out stays opted out. No other key gains deletion. - install/remove/getStatus with a profile directory refuse when it is the default home, or its settings.json resolves to the default one, instead of editing System Default's hooks and opt-out state. - The profile statusline rule reads the default settings under the userHome passed to the setup entry, not os.homedir(). Profile state and ownership: - A malformed `projects` value skips only folder trust (new warning code trust-refused); onboarding and shared keys still apply. - The ownership marker stores only host-local facts (account, runtime, distro). The execution host id is the caller's view of the host, so it stays in the in-memory descriptor and is not compared. Prompt history interruption paths: - With no cursor yet, a retained copy starts past the bytes it shares with the default history, so an interrupted share no longer replays the whole history. - A retained name for the shared file itself is removed with its cursor instead of lingering until a later scrub makes it look new. - The Windows link record is read three-state: unreadable stops the share instead of reading as "no link". If the record cannot be written after linking, the fresh link is undone. - An unreadable retained copy is reported and no longer blocks linking. * build(cli): list the new Claude hook modules in the CLI project hook-service.ts and hook-settings.ts are compiled into the packaged CLI project, which lists every file explicitly. The statusline policy and profile destination modules they now import were missing, so the CLI typecheck failed with TS6307. The CLI still loads hook-service through the existing managed-agent-hook-controls build entry, which bundles both modules; neither imports electron. * fix(claude): close review round 3 regressions in profile setup - A profile whose hooks hold only Orca's entries and that sharing never recorded is no longer treated as a user edit, so the user's first own hook in ~/.claude reaches it (for example when the profile was set up before ~/.claude had any hooks). - A retained prompt-history file is removed as a second name for the shared file only when the default history does not itself link to it; otherwise it holds the only copy and is kept. - Default-home checks compare file identity: the profile hook destination check uses device and inode, and the profile/default separation check resolves on-disk case, so a case-only alias of ~/.claude is refused on case-insensitive filesystems. - A test pins that an unreadable leftover session tree no longer blocks linking. * fix(claude): let shared keys leave a profile when ~/.claude drops them QA found that removing a setting from ~/.claude never reached a managed account: deleting the whole `hooks` block left the user's hook running there. Only statusLine followed the default away. Every shared key now follows the same rule through the existing per-key ledger: when a key disappears from ~/.claude/settings.json (or mcpServers/theme from the personal state file), it is removed from the profile if the profile still holds exactly what Orca last shared. A value changed inside the account is kept. Keys Orca never shared, including denylisted ones, are never touched. Deleting the whole hooks block removes the user's shared hooks and keeps Orca's own entries. A missing source counts as empty; an unreadable source removes nothing. * fix(claude): share personal rules, themes, workflows and keybindings into account profiles A managed account launches Claude with its own config folder, so user-level rules/, custom themes/ (which a shared `custom:<slug>` theme points at), personal workflows/ and keybindings.json silently stopped applying. Link the three directories like skills and commands, and copy keybindings.json with the same edit-preserving ledger as CLAUDE.md. routines/ stays unshared: routines belong to the claude.ai account and the folder holds per-run state. * fix(claude): import the personal CLAUDE.md into account profiles instead of copying it Claude also loads ~/.claude/CLAUDE.md as a parent folder's memory for any project under home, so a copied account CLAUDE.md made every such session read the user's instructions twice (checked live with Claude 2.1.288). An @~/.claude/CLAUDE.md import resolves to the same real file, which Claude loads once from home, from projects under home and from folders outside it. * refactor(claude): simplify account profile setup toward the prior art - Windows keeps each account's history private; drop the hardlink, link record and conflict-copy machinery that only Windows reached. - Share hooks and statusLine as ordinary settings keys: Orca writes the same entries into every folder, so the installer finds them present. Drops the Orca-entry carve-out, the per-profile statusline follow logic and its marker. - Unreadable ledger is just an empty ledger. - Share from the user's own CLAUDE_CONFIG_DIR when they set one (marked so Orca's injected value is never mistaken for it), and refuse a profile at or around it. - Pin the one canonical profile path spelling in a test. * fix(claude-accounts): dedupe merged prompt history, drop drained copies, link setup folders by path - Prompt-history drain appends only lines the shared file lacks, so a purge never re-adds lines. - A set-aside history copy whose saved offset reaches its end is deleted on the next run. - Setup folders link to the default home's own entry, not its resolved target. - The profile gate and folder creation run once, in provisionClaudeAccountProfile. - installHooks receives only configDir; drop a duplicate test key that fails CI. * fix(claude-accounts): record installed hooks as Orca-shared; skip symlink tests on Windows After Orca installs its hooks into an account, record the account's hooks in the settings ledger so a later run can still bring the user's own hooks in. Tests that create real symlinks now skip on Windows. --------- Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> |
||
|
|
55c86ceb91 |
refactor(terminal): report breaches of the two layout invariants; add an unrun load repair (#25673)
* refactor(terminal): report breaches of the two layout invariants; add an unrun load repair
The binding write now checks, after the fast lane, that a terminal is bound to at most one
leaf and a leaf id is in at most one tab, and records a breach on the persistence.pty-binding
span as binding.owner_conflict. It never refuses or changes the write (D16). The load repair
for saved duplicates is added and unit-tested but not called on load; the B2 core turns on
refusal and repair together.
* refactor(terminal): keep the binding write unchanged if the owner check throws; plain-language comments
* fix(lint): merge duplicate removal import in delete-worktree failure toast
Main (#25668) introduced a duplicate import that fails the focused code-quality lint.
* refactor(terminal): keep only the report-only owner check; one definition of same terminal
Defer the unrun load repair to the change that runs it. Repeating relay ids without both
incarnations are no longer one terminal, a leaf held on another host is reported apart, and the
refusal check takes host partitions directly.
* fix(lint): bring structured-agent-session-host back under max-lines
(cherry picked from commit
|
||
|
|
d3e1494674 | test: bound memory used by the runtime Electron audit (#26049) | ||
|
|
7c9c3d431a | fix(lint): bring structured-agent-session-host back under max-lines (#26050) | ||
|
|
70c6876635 |
fix(rovo): detect and launch Rovo Dev through acli (#25981)
* fix(rovo): detect and launch Rovo Dev through acli Rovo Dev ships as the `acli rovodev` subcommand, with no `rovo` binary on PATH, so Orca never listed it as installed. Detect `acli`, launch `acli rovodev run` (prompt still typed after start, since a positional instruction is one-shot), and build resume as the launch command plus `--restore <id>` so a full-command settings override composes too. Fixes STA-9645 * test(skills): detect Rovo Dev through acli in the skills CLI fixture |
||
|
|
923ef7cff1 |
fix(native-chat): refuse a Command that can't run instead of silently running the stock CLI (#25667)
* fix(native-chat): honor custom Claude and Codex launch commands * fix(native-chat): sync launch failure localization catalogs * feat(native-chat): run Settings → Agents → Command as the native chat program Native chat ignored the Command setting, so a user who pointed it at a custom Claude or Codex build still got the stock CLI. The host now reads the Command per acquisition and per model-catalog probe, resolves it on the execution host as a program (an absolute path, ~ path, or a name on PATH), and spawns it with the normal structured arguments. A set Command that is not a runnable program refuses the start with a new agentCommandNotRunnable failure that names the setting; it never falls back to the stock CLI. * fix(native-chat): find a configured Claude Command on the launch PATH; clearer copy A Claude Command given as a bare name was looked up on Orca's own PATH, not the PATH the Claude child launches with (login shell plus Settings → Agents environment), so a name only that environment provides was refused. Codex already resolved against its launch environment. Claude now reads its launch env first and resolves a configured name against its PATH and home; the stock lookup with no Command is unchanged. The failure copy now states the rule (a program path or name, no arguments or variables) and names the Settings control (Reset). The catalog fingerprint doc notes that the program is not keyed, so a chat on an older program keeps refreshing the shared entry until it ends. * fix(native-chat): one PATH key for the Claude lookup and the child on Windows On Windows, an inherited `Path` and a Settings → Agents env `PATH` both reached the Claude child. The configured-program lookup read the first inserted twin (inherited), while Node's child_process keeps the lexicographically first (`PATH`, the user's), so a bare-name Command found only on the Settings PATH was refused. The Claude launch env now drops inherited case-twins of the overlay's variables (the rule the structured shell env already applies), so the lookup and the child read one PATH. pathEnvOf follows Node's win32 key rule, and Codex's lookup uses it too, so both agents read PATH one way. es/fr/ja copy now says the Command must be, as the English does. * fix(native-chat): refuse a Command that can't run instead of running the stock CLI A saved Settings → Agents → Command that did not resolve to a program (a command line such as `codex --profile work`, `FOO=1 claude` or `npx …`, a missing or relative path, a non-executable file) silently ran the stock Claude/Codex in native chat and its model-catalog probe. The user's choice was ignored with no sign of it. A set Command that does not resolve now refuses the start with a new agentCommandNotRunnable failure naming the setting, beside the existing Retry; the catalog probe refuses the same way and never lists the stock CLI. A blank Command keeps the stock lookup. On Windows a resolved file without a spawnable extension (.exe/.com/.cmd/.bat) refuses too, instead of failing to spawn with no word about the setting. * chore(native-chat): correct comments that say native chat ignores the Command After #25721 native chat runs a runnable Command, and this PR refuses one it can't run; four comments still said the Command applies to terminal launches only. The pre-spawn test fixture also drops the saved value from its error message, matching the resolver. |
||
|
|
5b1bb78d14 |
[STA-6940] Prepare file drops for element owners (PR 2 of 6) (#25748)
* feat(file-drop): add element owner preparation plumbing * fix(file-drop): preserve feedback acceptance and destination ordering * fix(file-drop): keep path resolution in filesystem namespace |
||
|
|
0bdcaf36ed |
fix(claude): start a Claude chat with its saved options and send the first message at once (#25152)
* fix(claude): end a Claude start that never answers initialize after 120 s * Read the Claude startup deadline inside startup; fix a stale test comment * fix(claude): start a Claude chat on its initialize answer, not on a frame only a SessionStart hook sends Startup waited for system/init or a SessionStart hook frame as well as the initialize answer. Before the first turn only a SessionStart hook sends one, and Orca adds that hook only through its optional status hooks, so with them off the first message was held forever. Startup now lands on the initialize answer; a start frame already seen is still checked, and one naming another session ends a started session. The deadline drops to 90 s so it fires inside the host's 120 s start wait. * fix(claude): time the Claude start by silence, and fail it at once on another session's frame Claude answers initialize only after its SessionStart hooks finish, so a total-time deadline would fail every start behind a slow hook. Each start frame now restarts the clock. A frame naming another session fails a start still waiting on initialize at once, as before. * test(claude): a real Claude chat starts and answers with every hook disabled * Say what the start-frame re-arm covers, and check only start frames in the hook-less real test * fix(native-chat): a Claude chat starts with its saved options and takes its first message at once Saved model, effort, Fast and permission mode are passed as launch options, checked against the account's cached model catalog, instead of restored by control requests after initialize. With nothing left to restore, the host no longer holds a message until the CLI answers initialize, and the 90 s startup deadline is gone. A Stop on a start that never answers ends the child and settles what it was handed as stopped. A failed result that repeats the turn's own API error reply writes no second row. * fix(native-chat): a host stop of a Claude start fails the message it was handed, with one row With no start-hold the delivery loop no longer sees a host stop of a start it waited on. The child's end now rejects what it handed over with the host-stopped words and writes the one row, as an exit of its own would; an idle start the host stops still goes quietly. * test(native-chat): a Claude chat's first message is written before initialize answers Rewrites the tests that encoded the start-hold, the startup deadline and the option restore to the new contract, and adds: saved options at launch (catalog checks, bypass, fresh-session Fast), a message written before initialize answers (adapter and runtime), Stop on a start that never answers (stopped, child closed, nothing working), and an API error said once. * revert(native-chat): keep a failed Claude turn's error row The shared turn fold already shows a failed turn's error once after it settles, as an error; dropping the row left the CLI's synthetic reply looking like something Claude said. * fix(native-chat): pass saved Claude options unchecked, heal a retired model on the CLI's word, and never leave an unrun message in doubt - Saved model, effort and Fast are launched as picked; only values no Claude can parse are left out. The pre-spawn cache check is gone. - A saved Fast on for a new conversation is applied once the settings readback shows no per-session opt-in (dropped when there is one, or when the model is listed without Fast), with nothing waiting on it; the record keeps the pick. - Under an Agent Permissions bypass, a saved narrower mode launches with the allow flag so bypass stays reachable. - A turn whose reply is the CLI's model_not_found for the launched model drops that model from the record; the launch's own row for it is kept out of the account model cache. - A child that ends before it answered initialize, for any reason, settles every message it was handed as not sent (cancelled for a Stop). - A launched effort the CLI reports only as `applied.effort` is confirmed from there. - The untimed-initialize comment is back to main's text. - A real-CLI test for a message written before initialize answers, under saved options. * fix(native-chat): type the close's ended event and the start-exit test fixtures The close's ended event is typed as the adapter event so its optional startupUnanswered spread fits exactOptionalPropertyTypes; two tests guard the fixture's optional generation, and the hung-start fixture records initialize on the fake connection it holds. * fix(native-chat): a Claude model heal keeps a later pick, a refused Fast is dropped, no allow flag - `options-skipped` carries the retired value; the record drops it only while it still holds it. - `started` carries the values a heal retired, and the record does not take them back from the CLI's report of the same value. - A saved Fast on a new conversation is applied before `started`: a refusal drops it and records it skipped, as main's refused restore did; silence keeps it wanted and unconfirmed. - A saved narrower mode under an Agent Permissions bypass launches without any bypass flag again: the allow flag is one older CLIs reject at start. Kept as a known limit. - The real-CLI test asserts the message was written before initialize answered. - The fake reports a launch effort only under `applied`, and a misplaced doc comment moves back. * fix(native-chat): a new Claude chat reports started before its saved Fast is applied The Fast apply on a new conversation now runs after `started`, so a Stop interrupts a running first turn and an option write is not refused while the round trip is out. A refusal drops the pick through `options-skipped`, in order after `started`; silence keeps it unconfirmed. The launch's unreachable skipped-model branch is gone. * fix(native-chat): a healed Claude chat goes back to the default model live; comments match the no-hold design When the CLI says the launched model does not exist, the live child is also put back on the CLI's own default (set_model with no model, fire-and-forget), so later messages in the same chat run; a user's pick sent after it wins, and a refused or unanswered reset only logs. Comments that still described the start-hold or the option restore now describe the launch options and the handed-over, never-echoed rule. * fix(native-chat): a message handed to a Claude start that never answered is kept as main keeps an unsent one A child that ended before it answered initialize ran nothing it was handed, the same fact as a send accepted and never handed over. Its end now settles those sends exactly as the chat settles a queued send for that end: a quit keeps a person's message as a held card (restart words), a close keeps it as a held card (closed words), a person's Stop withdraws it as cancelled, and a host stop fails the start with one row. * fix(native-chat): a quit during a Claude start that never answered offers no resume for the message it keeps as a card The restart snapshot now reads the same never-answered fact the exit does, so a message handed to such a start counts as queued work, not as work to resume. The retired-model reset comment names the default it really applies. * refactor(native-chat): the saved permission-mode launch helpers live with the spawn options that use them Keeps claude-structured-launch-resolution.ts within max-lines once merged with main, and names the hung-start test envelope's field type. * fix(native-chat): a Claude chat's saved options take precedence over the agent Arguments' own flags Main now passes the saved agent Arguments to the Claude child, and the SDK writes them after its own options. An Arguments --model or --effort therefore reached the CLI as a second flag after the chat's saved pick (a commander CLI keeps the last), and a saved Fast's launch settings replaced an Arguments --settings file outright. The saved model and effort now stand in for the Arguments' flags, and a saved Fast beside an Arguments --settings is applied by the start instead of at launch. * test(native-chat): the hand-built Claude session in the options test carries fastModeAtStart * test(native-chat): the queued rig's start spy carries a named Mock type An unannotated vi.fn() inferred @vitest/spy's internal Procedure, which CI's typecheck cannot name in the factories' inferred return types (TS2883). * fix(ci): run the PR's SQLite-backed tests in the Node runtime project Lists the hung-start Stop test, renames the send-during-startup entry from its old name, and carries main's own two entries from #26010 so the boundary test passes before the next merge. |
||
|
|
7a59aa39d0 |
fix(worktrees): decide setup before git worktree add on the host's create (#26006)
* fix(worktrees): decide setup before git worktree add on the host's create, so an undecided ask repo leaves nothing behind The host's create (agent.launch, worktree.create, the phone and the CLI) only checked the setup policy after adding the worktree. An ask repo with no setup decision then threw with the worktree already on disk and no worktrees-changed event. It now refuses before the add, as the desktop create does, and a setup hook the new branch adds (one nobody decided on) is skipped with a warning instead of failing a create that already exists. * fix(worktrees): read the host create's setup decision in one place, after recording its host The pre-add refusal now runs after the create records its execution host, so a refused create's failure telemetry still says where it ran; it still runs before any git work. Both setup checks read the decision from one helper so they cannot disagree, and the post-add skip logs once. The refusal test now also proves no base resolution, fetch or branch naming ran and that the hooks came from the main checkout. * test(worktrees): cover a setup hook only the new branch adds through the managed create An ask repo with no decision, whose setup hook exists only in the new worktree's orca.yaml, now creates, refreshes the worktree list, reports setup as skipped and returns the warning. * fix(cli): name the --setup flag when a worktree create needs a setup decision The host refuses an undecided create in an ask repo with "Setup decision required for this repository", which desktop and phone answer in their own UI. The CLI now adds the next step: pass --setup run or --setup skip. |
||
|
|
b3b6c5dc13 |
Give native chat names one source for tabs, sidebar and AI Vault (list and search) (#25986)
* Give native chat names one renderer source and drop Vault's name repair copies The host's saved conversation name now rides the structured session status feed, which already exists per host, is keyed by the durable session id, and keeps a closed chat's summary. Tab strip, sidebar rows and AI Vault (list and search) read it through one hook and one display order (tab alias, saved name, host label). Vault no longer copies names into its cached results, so the projection, recovery and pending-title modules and their tab-snapshot lanes are removed. Indexed search hits now carry the native owner and saved name from the host that indexed them. * Type the sidebar name test fixture without an assertion * Keep Vault search working when the chat host will not install Naming and owning search hits is bookkeeping: if the native chat host fails to install, return the plain hits instead of failing the search. The runtime RPC only installs the host for clients that will receive the owners. * Publish chat names to the feed independently of the tab retitle A failed feed publication no longer skips retitling the open tab. The publish now lives in the naming deps, where a test covers it. * Note why the status feed must keep closed chats' summaries * Bound names and owner ids that come from a paired host Drop a published chat name the record store would refuse, and cap a search hit's owner workspace id at the same length the list row and record use. * Let native chat search hits from a paired host open their chat A paired host's search hits carry no resume command, so the row disabled every open action even for a native chat it can open through its owner, as its list row does. * Ignore workspace ids that name object members in tab lookups A paired host's row or search hit could carry a workspace id such as "constructor", which read an Object.prototype member as a tab list and broke the render. The shared tab index now reads only own workspace entries. |
||
|
|
9fdd90ffa7 |
feat(orchestration): a native chat can be a dispatch worker, like a terminal agent (#22972)
* feat(native-chat): a queued card can record a Dispatch's task as its source A chat worker's task is held in the chat's queue like any agent message, so the card needs to say which Dispatch it is from: the sender, run, task and Dispatch ids. An older build reads an unknown kind as the person's card and sends it as written. * feat(orchestration): a chat can be a dispatch worker, like a terminal agent `dispatch --to orca_session_id:<id>` and `worker-start --terminal orca_session_id:<id>` now accept a chat on this host instead of refusing it. - The chat is refused only where mail to it would be: unknown, a provider id, another host, or closed. A chat can't be its own coordinator's worker. - Its task goes through sendAgentTurn as a queued send, as mail notices do: an idle chat starts a turn, a busy one holds a card naming the Dispatch. The operation id is derived from the Dispatch, so a resend replays. - worker-start reads that outcome as the terminal path reads its write; a card held behind a running turn is handed over with its start unobserved. - The Dispatch names the chat by its /clear root, with no pane or process, and records its Orca session id, so its own sub-dispatches nest under it. - Mail to dispatch:<id> reaches the chat; worker-show/list/read read the chat's session records; stop and abandon never close the chat; closing the chat fails its Dispatch as closing a terminal does. - A dispatch preamble for a structured session names the CLI the structured mail lane names. * test(orchestration): a chat worker's reach, and its Dispatch across a /clear At rest is live, a closed chat has exited, and another host or nothing to read is unverifiable. A /clear keeps the chat's Dispatch, and the session that continues the chat reports as it. * fix(orchestration): chat worker review round 1 - dispatch --inject is a keepalive-backed wait, so a chat that takes a while to accept its task no longer reads as a dead runtime at the 30 s idle cut. - An injected task whose delivery is unknown keeps its Dispatch open, as an unknown worker-start does, instead of failing it while the chat may run it. - A busy chat's worker-start receipt says the task waits as a card in the chat's queue, and that worker-abandon does not remove that card. - A close that puts the chat's tab back no longer fails its Dispatch: the closed-chat settle runs after the close's outcome, not at its hide. - worker-start adopting a chat says it gave the task to the chat, instead of claiming it started a terminal agent; the mode value is unchanged. - Test: a chat whose agent has not taken its task reads outcome_unknown. * test(native-chat): a close's hide defers its hidden notice until the close settles The rollback suite pinned the hide's arguments; it now expects the deferred notice, and that the notice is sent once, after a rolled-back tab is back. * fix(orchestration): a chat worker whose successor is unknown is unverifiable, not exited A /clear successor this host has no record of is missing evidence, not proof the chat is gone. Only a closed chat reads exited, so only a closed chat settles its Dispatch; a lost successor reads unverifiable with that reason. * fix(orchestration): chat worker review round 2 - A chat's close notice settles only that chat's Dispatch. A scan of every chat Dispatch read another chat whose close was still in flight (its tab hidden, maybe to be put back) as closed and failed its Dispatch. - Only a dispatch --inject into a chat is a long poll; a terminal inject writes and returns, and keeps its short-RPC slot. - Receipt wording: worker-start says it gave the task to the chat, and a queued task is sent when the chat's queue reaches it. * fix(orchestration): chat worker wording, review round 3 - worker-start's mode sentence for a chat states the placement, which is true whether the task is delivered, queued or refused. - Comments and a test title no longer claim a close notice re-derives every chat worker, or that a queued task waits for the current turn. * fix(orchestration): build a structured worker's task sender from the start's own ids The minted worker's preamble names who its task is from with the run, task and Dispatch ids worker-start already holds, so it reads no Dispatch row. * fix(orchestration): point a chat worker at its held Dispatch mail again A chat worker's coordinator mail lands in its Dispatch's mailbox, but a chat's idle edge re-derived the Dispatch mailbox only for a party with a terminal handle. A pointer lost with the provider then waited for new mail after a restart or a /clear. The re-derivation now looks the Dispatch up by the party's address, which names a worker by its handle and a chat by its root. * fix(native-chat): open a task card's sender by address; a task carries no mail Opening an agent message's sender looked up the mail it carried, which only a mail notice has; with the task kind in the union that read no longer typed. A task's coordinator is found by its address alone. |
||
|
|
825d7bd5a9 |
test(vitest): run agent-launch-instant-tab in the SQLite runtime project (#26028)
#25430 added a test that opens a real agent-session record store, but not to the SQLite runtime list, so vitest-sqlite-runtime-boundary fails on main. |
||
|
|
cdd0b7a491 |
fix(native-chat): stop flashing a reconnecting line on a stream drop (#24898)
* fix(native-chat): drop the per-chat reconnecting line on a stream drop A single transcript stream drop flashed 'Reconnecting to this chat…' above the composer for about a second. An unnamed read failure now adds nothing: the transcript and composer stay, and the host's own status shows reachability. Named or final failures keep their line. * fix(native-chat): keep a loaded chat as it is when its stream drops naming nothing A loaded chat with no messages yet still flipped to the full-pane "Could not load conversation" on every stream blip: the read owner stored any read failure as status 'error', and the pane shows that error whenever there is no transcript. Once the chat has loaded, a failure that names no reason is now the transport's own retry and is not stored; a named or final failure, and any failure before the first load, are reported as before. The "names a reason" check moves next to the final-refusal check so the owner and the failure notice share it. The older-page read moves to its own module to keep the read owner within its line budget. * fix(native-chat): word a read failure beside a chat that never loaded A chat that never loaded but already shows the user's own message (a launch prompt or a queued send) said nothing when its first read kept failing with a failure that names no reason. The loaded case is handled at the read owner, so any read failure the pane sees beside messages is named, final, or from before the first load; the status area now says its words in every such case. * fix(native-chat): keep a loaded chat as it is only when contact is lost The read owner skipped any loaded-chat failure without a named reason, but an older host sends every refusal without details, so a refusal it really sent (such as a damaged history) went unshown while the read retried in silence. Skip only failures the host sent no refusal for, which is lost contact; any refusal is the host's answer and is shown. The shared 'named' check is no longer needed and is removed. * docs(native-chat): say what a read failure without a refusal is treated as Comment-only: the read owner treats a failure with no host refusal as lost contact, which covers hosts too old to attach refusals; the test's reasonless case is a reason this build doesn't know. * Say a chat's host outage once, above the composer A loaded remote chat looked live through a long host outage while sends sat in the outbox. Derive a host-scoped notice from the host's connection state: '<host> is reconnecting…' after a 2 s grace, '<host> is offline' at once with the status bar's Connect as Reconnect, and a composer placeholder saying sends go out when the host reconnects. A lost read beside the notice adds no line of its own, and older history waits for the host. * Drop the outage placeholder; no Reconnect for a refused host A send while the host's transport is down fails and waits for its own Retry, so the composer must not promise it will go out on reconnect. A host that refused us (auth, protocol) still reads offline but offers no Reconnect, which would be turned away the same way. Name the host with the shared display-label selector, and keep the notice's live region mounted so it is announced. * test(native-chat): mock the font-size hook main renamed in the host-outage test |
||
|
|
0c96550ee9 |
fix(native-chat): queued messages carry on in order after any turn, and nothing sends by itself after a restart (#24586)
* fix(native-chat): drop the queue-paused header and Resume button
A Stop, a restart or /clear holds the queued cards. The hold stays; only the
header row naming why, and its Resume button, go. A held card shows no
caption, and its own Steer, or any new message, releases the queue.
* test(native-chat): type the unknown hold reason a newer host may publish
* fix(native-chat): a held card offers Send, not Steer, when no turn runs
Steer vs Send now follows whether a turn is running, not the card's hold,
so a card held after a Stop, a restart or /clear reads Send.
* fix(native-chat): the queue sends past held cards instead of stalling behind them
A card queued after a Stop (or written after a restart or /clear) sent only
once the cards held before it were released; with no header to explain or
release the hold, it sat silently. The next sendable card now skips held
cards; a returned card still blocks what is behind it.
* fix(native-chat): the queue's send of a card is the person's turn, so held cards follow it
After a Stop, a card queued later sent past the held cards, but the queue
recorded that send as Orca's own turn. It never ended the Stop's pause, so
the held cards then waited forever with nothing on the card saying why.
A queued card is always something the person wrote: only the client send
RPC may now create one. The queue's send of it is therefore recorded as the
person's turn, which ends the Stop's pause once the agent takes it, and the
held cards then drain in order.
* fix(native-chat): a queued card carries its author, so the queue's send of it is that author's turn
Main now lets Orca's own sends ask to queue (sendAgentTurn's 'queue' delivery),
so "every card is a person's" no longer holds by refusing host sends. Each card
records who wrote it (the submission's client/host vocabulary) in a new nullable
column; the drain records that origin, so a person's card ends a Stop's pause
and Orca's does not. /clear carries the author. Rows from before the column
read as a person's. The userSend-only admission gate is removed.
* docs(native-chat): state why an unrecorded card author reads as a person's
* fix(native-chat): a restart holds only cards written before it, and an idle held queue offers Resume
A restart's pause held every waiting card, including one a person typed after the restart while
Orca's own continuation ran, and nothing released it except a per-card Send. It now holds only
cards another host process wrote, the same way a Stop holds only cards queued before it.
The composer's primary button becomes Resume (Play) while nothing is typed, no turn runs and the
host holds a card Resume would send, whatever held it (Stop, restart or /clear). It calls the
existing agentSession.queuedMessagesResume, guarded against a second press in flight.
A card nothing holds keeps the run going between a turn's end and the queue's send of it, so its
Steer no longer flips to Send for the frame in between.
* fix(native-chat): the host publishes which pause holds each queued card
The host published one pause for the whole queue, so a client held every waiting card while it was
set. Between a turn's end and the queue's send of a card queued after a Stop or restart, the
composer could flash Resume and the cards Send, and a card queued after a Stop lost its
"Waiting for your answer" caption.
Each published card now carries an optional `heldBy`: the pause holding it, or null, derived from
the same rule the drain reads. A client holds only those cards; against a host without the field
it falls back to the queue-level pause.
* test(native-chat): Resume needs the queue capability and is disabled whenever Send is
* docs(native-chat): describe per-card holds in the queue contract and table comments
* fix(native-chat): the composer goes from Resume straight to Stop, and Resume returns focus
After Resume, the host lifts the hold in one update and sends the first card in a later one. In
between nothing was running, so the composer's button flashed a disabled Send. A card nothing
holds now keeps the queue's run going for the button too: an empty composer shows Stop, disabled
until the turn starts. Not when the host refuses every send (a rewind whose outcome is unknown,
read from its status), where nothing is coming. The same fix removes the Stop, Send, Stop flip
between queued turns.
Resume disables the button, which dropped keyboard focus; focus now returns to the composer.
* fix(native-chat): the host names the card its queue sends next, so the chat stays working across the gap
A turn's end, or a Resume, and the queue's send of the next card commit as two host updates. In
between nothing was running, so the working status, timer, pickers and composer button flipped
for one update. The client guessed the drain from its own copy of the host's gates, which missed a
/clear-replaced source and covered only the button.
The queue publication now carries `nextQueuedMessageId`: the drain's own next card through the
drain's own gate (`nextStructuredQueuedMessage`, which the drain step now calls), null whenever the
host would refuse the send. The client derives one fact, the queue is about to send, and every
working reader follows it; Stop stays disabled until a turn can be stopped. The client-side copy of
the gates and the status-feed rewind read are removed.
* test(native-chat): the queue's next card survives the coalescer, the reducer and a history page
* test(native-chat): build the snapshot that names the next card through its helper
* feat(native-chat): a held queue keeps its header row, and a new message asks before passing it
The queue's header row ("Queue paused because you interrupted", or Orca
restarted, or you cleared the conversation) comes back above the cards it
holds, with Resume; it names the oldest held card's pause, as the host
publishes it per card, and hides over cards held only on their own or
returned. The header's Resume and the composer's share one in-flight guard.
A held card reads Steer again whether or not a turn runs; a card held on its
own or returned keeps Send.
Sending a message while the header shows (Enter or the button) first asks
"Send message?": Clear queue deletes every card and then sends (a failed
delete sends nothing), Send message sends and keeps the cards, which follow
the new turn, and dismissing sends nothing and keeps the draft. Host
commands send as they are.
* fix(native-chat): the paused row goes while your own message is on its way to lift it
After "Send message" over a held queue, the row kept saying "Queue paused…"
until the agent accepted the new turn. The chat now reads that gap from the
outbox: while this composer's direct send is recorded by the host and not yet
accepted, the controller shows no paused row (and so no Resume or
confirmation). A refusal settles the entry and the row comes back, since the
hold did not lift. Orca's own sends never enter this outbox, and the queue's
send of a card goes under a fresh id, so neither hides it. Nothing is stored.
* fix(native-chat): a "Send message?" choice is taken once, and a failed Clear queue is one toast
The closing dialog stays mounted and clickable through its exit animation,
and a double-click or a held Enter lands twice before any re-render, so
Send message (or Clear queue) could send the captured message twice. The
pending send now lives in a ref that the first choice takes; a second one
finds nothing.
Clear queue deletes one card at a time and stops at the first failure, so a
failed press shows one toast instead of one per card.
The dialog keeps its compact width at desktop sizes and the primitive's
narrow-window gutter (`max-w-sm sm:max-w-sm`, as the other compact
confirmations).
* fix(native-chat): Clear queue's message goes out once, and keeps text typed while it waits
After Clear queue, the message waited in the composer while the cards were
deleted one by one. A second Enter in that window sent it again, and text
typed meanwhile was wiped when the chained send was accepted.
From the Clear queue choice until its message has gone out, the composer's
structured send does nothing. The chained send (and Send message's) now
carries the composition it was taken from, and the composer is cleared on
acceptance only if it still holds exactly that, as host commands already do.
Also: the v1 contract comment names `nextQueuedMessageId` and its absent-
means-null fallback, and the own-send check returns at once on an empty
outbox.
* fix(native-chat): the queue carries on after any turn, in order, and a restart sends nothing by itself
- Any accepted turn ends a Stop's or a /clear's pause, whoever sent it (a person,
Orca's own messages, or the queue), and so does Resume. The card and submission
author fields that only fed the old person-only rule are gone.
- The queue sends strictly in order: a card never overtakes a held one.
- After a restart nothing sends by itself and no paused row shows: the chat's next
turn (the carry-on, or the person's own message) runs first, then the cards.
- Resume and "Send message?" are offered only while nothing runs and no prompt waits.
* fix(native-chat): after a restart no queue pause shows, and a card written before the next turn waits for it too
* fix(native-chat): a quit hands no queued card off, and the paused row goes while any turn that will lift it is on its way
- The queue stops handing cards off when the host tears down. A card sent during
a quit was refused at close, and that refused send withdrew the chat's restart
offer, so resuming after the relaunch sent nothing.
- The host publishes no pause while a turn sent after it (your message, Steer, or
Orca's own) waits for the agent; a refusal shows it again. This replaces the
client's own-send check.
- A card written after a restart is an ordinary card again: it waits while any
card from before the restart still waits.
* refactor(native-chat): the host's paused-row-while-a-turn-is-on-its-way check in one expression
* refactor(native-chat): the composer's queue Resume rides the structured transport beside the held queue
* fix(native-chat): the "Send message?" choice ends with the pause it asked about; tests follow main's draft props
- The open dialog closes when the queue's pause lifts under it (Orca's mail, another client's
Resume, any accepted turn): nothing is sent, the draft stays, and the next Enter sends as
usual. The pending choice records the hold it was asked under; nothing new is stored.
- The composer-field Resume test passes main's dropScopeKey/draftScopeKey.
- The dialog test expects main's rule: only the sent text leaves the composer.
* test(mobile): a host-kept card's test stands in a Stop's pause, as this host publishes no restart pause
Main's #24660 test published queuePause 'restarted', which this branch's wire
type no longer lists, so the mobile tests typecheck ratchet failed.
|
||
|
|
3fb72d135d | Run Node event-loop measurement after ordinary test suites (#26015) | ||
|
|
c0273b1ff7 | test(e2e): move the Source Control reveal golden into its own spec so older release tags skip it (#26005) | ||
|
|
2d08b3a1da |
Speed up expensive test fixtures (23–82% less time) (#26000)
* Advance Codex fixture deadlines with simulated clocks * Build large file-listing fixtures without promise batches * Compare large binary test results with native byte equality * Speed up Claude stop-note deadline fixtures |
||
|
|
37ff3873a0 | Run combined localization catalog verification on Bun (#25999) | ||
|
|
00e3027762 |
feat(agent-launch): show an agent's tab at once, where the caller asked (#25430)
* feat(agent-launch): host-assigned caller identity and a launch record written when the surface exists Step 1 of the agent-launch unification, on main. - The dispatcher stamps every request's caller from what its connection proved (runtime socket: the local CLI; the desktop's IPC: the desktop; a paired socket: its device). Params never set it. - The launch record is written twice: once when the tab exists (what creation settled: on the launch command, a draft, or a submit still `unconfirmed`), and again once the prompt's fate is known. A restart in between finds the running agent instead of answering "unknown". - A replay re-derives its terminal handle from the pane key in the running host, and shows `unconfirmed` only to callers that advertise agent.launch.prompt-unconfirmed.v1. - The record store opens in its own slot, without building the chat host; the chat host is built on that same store. Rebuilt from this PR's own commits ( |
||
|
|
66c775fbf7 |
refactor(terminal): project main's terminal layout after each save (no reader yet) (#25682)
* refactor(terminal): publish main's terminal topology after each session write Adds a by-value projection of each worktree's persisted terminal topology and a publisher hooked on the two persistence funnels (scheduleSave and durable mutations). Changed slices are pushed to the local window on session:terminal-topology-changed with a monotonic publishSeq; a startup pull (session:get-terminal-topology-slices) and publishSeq on pty:spawn and session:close-terminal-surface replies are in place. No renderer consumer yet, so behavior and saved state are unchanged. Observer failures are counted and logged once and never reach the save. * refactor(terminal): keep the topology publisher dormant until a reader pulls Writes now cost nothing extra until the first session:get-terminal-topology-slices pull (or subscribe()) takes the baseline; markDirty before that is a no-op. * refactor(terminal): narrow topology publishing to an unwired publisher and save hook Drop the window sink, pull handler and publishSeq replies; no listener attaches in production. Filter sleeping records by worktree id only, and guard observer failures once inside the publisher. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
f0520851ab | Keep mobile restore SQLite fixture in the Node test runtime | ||
|
|
1040f673b1 | Keep orchestration SQLite fixtures in the Node test runtime | ||
|
|
14bae2e172 | Route mirrored editor closes through their captured runtime owner | ||
|
|
0c32e80cc8 | test(codex): reuse unproven close fixture sequence | ||
|
|
983098dd00 | test(codex): synchronize fake clock with forced process close | ||
|
|
31d85749b4 | fix(editor): retain placeholders for drafts during bulk close | ||
|
|
8bd567a015 | fix(editor): protect document backing files and complete close cleanup | ||
|
|
91bb636530 | test(editor): align model and cursor custody fixtures with owner-close policy | ||
|
|
c9f4cd7e27 |
fix(editor): close stale duplicate documents safely
Adapt the document-sibling cleanup from Pr1p's #23347 with exact owner and captured provenance checks. Preserve divergent drafts and the backing bytes needed by retained views, and keep one-pane close behavior intact. Co-authored-by: Chen <zwq19980411@gmail.com> |
||
|
|
9c6702bb07 |
test(codex): guard that turning hooks off wins over an in-flight turn-on (#26001)
Claude-Session: codex-hook-e |
||
|
|
3308ff8b26 |
Bound YAML merge conversion and close SQLite routing review gaps (#25998)
* Close test runtime and YAML merge review gaps * Bound YAML conversion inside explicitly tagged pairs * Make completion notification fixture cadence deterministic |
||
|
|
323f312819 |
refactor(terminal): mark layout updates from user gestures (#25680)
* refactor(terminal): mark layout updates from user gestures
Divider drag end, divider double-click reset, pane reorder drop, equalize,
and pane rename/title clear now report themselves as gestures. Their
remote pane-layout push carries an optional intent: 'gesture'; every
other persist is unmarked and byte-identical to before. Saved layouts
are unchanged, and no host reads the field yet: the params schema
accepts it and degrades any other value to unmarked, and the handler
does not forward it.
* refactor(terminal): name the layout gesture marker once, as the wire intent
Gesture sites pass onLayoutChanged('gesture') and persistLayoutSnapshot('gesture')
directly, dropping the per-gesture names and the options bag. The params schema
reads intent as z.literal('gesture') with .catch(undefined) so unknown values
still degrade to unmarked.
|
||
|
|
61bcca9fee |
fix: stop Codex and OpenCode helper servers when Orca quits or crashes (STA-9254, 2 of 2) (#25753)
* fix(supervisor): add a one-shot lifetime that runs past stdin end, and relay a provider's last output A one-shot CLI (claude -p, codex exec -) reads its request until stdin ends, so the supervisor's session rule (stdin end means the owner is closing) would stop it about 1 s into its answer. The one-shot lifetime passes stdin end through and leaves only the owner-death watch and explicit signals to stop it. The supervisor also exited as soon as its provider did, dropping output still in the pipes when the owner reads slowly. It now waits, bounded, for the provider's output to be relayed before exiting. * fix(text-generation): run agent one-shots under the provider supervisor on POSIX The Claude model-list probe, model discovery for every agent, and commit message, pull request and branch name generation all spawned the agent CLI as a plain child of Orca. If Orca quit, crashed or was killed while one was running, nothing stopped it, and a CLI that hung kept running after Orca was gone. They now run under the provider supervisor that native chat already uses, in its one-shot lifetime, so Orca's exit stops the agent's whole process group however Orca exits. A timeout or cancel asks the supervisor to stop (SIGTERM) and only tears the tree down once it has had its full stop time; killing the supervisor first would orphan the agent's group. A missing binary now reaches Orca as the supervisor's exit 127, which is mapped back to the existing not-found message. Windows and WSL keep spawning the agent directly. * fix(supervisor): carry the provider argv on the supervisor's argv, map every spawn error back, and share one stop ladder - The supervisor read the provider's command and arguments from one base64 JSON env string. Linux caps a single env string at 128 KiB, so an argv prompt of about 90-120 KiB, which passes the 120 KiB per-argument guard, failed execve with E2BIG. The provider argv now follows the supervisor script's '--' as real arguments; the env keeps only small fields. - A supervisor that cannot start its provider reports Node's spawn error line and exits 127. That line now becomes the same error a direct spawn emits, so ENOENT still reads as 'not found on PATH' and EACCES or any other spawn error reads as 'failed to start'. - stopSupervisedProvider is the one ask, wait and force ladder: it gives a supervisor its full stop time before forcing. Agent one-shots, the Codex app-server close and the Claude child exit proof now share it, with the same requests, bounds and forced steps as before. * fix(supervisor): report every provider spawn failure on one marked stderr line Node throws most spawn failures (ENOEXEC, ENOTDIR, ELOOP, EPERM, ...) instead of emitting them, and the supervisor had no catch, so it died with exit 1 and a stack trace that the user saw as the agent's failure. A thrown or emitted spawn failure now exits 127 with one marked line carrying whether it was thrown, its code and its message. Orca reads only the last stderr line, so a runtime warning printed earlier cannot hide it, and maps it to the message a direct spawn gave: thrown is 'could not be started', ENOENT is 'not found on PATH', and any other emitted error is 'failed to start'. * fix(supervisor): keep the user's Node options away from the supervisor and hand them to the provider The supervisor runs Electron in Node mode, which honours NODE_OPTIONS and NODE_REPL_EXTERNAL_MODULE. A user value such as a --require of a missing file stopped the supervisor from starting, breaking even native CLIs like Codex that never load it. The launch now takes both out of the supervisor's environment, carries them in its spec, and restores them for the provider only, so a Node-based CLI still gets them. * fix(text-generation): file a forced agent one-shot teardown under its own breadcrumb site The forced tree teardown recorded every self-initiated kill as the Codex app-server's, so a forced commit-message or model-discovery stop read as a Codex teardown in crash breadcrumbs. The teardown now takes the caller's site; source-control stops pass the site their Windows tree kill already uses. * test(text-generation): cover the supervised stop under timeout and output limit; name the direct-child suites The commit-message suites that drive fake children spawn them directly, the unsupervised shape Windows and WSL use, so they now say so. The supervised POSIX stop gets its own compositions: a timed-out Codex generation settles at once but holds the Codex home until its supervisor has stopped (faithful fake, fake timers), and an agent that floods past the output limit is stopped through its real supervisor with no process left behind. * fix(codex): run short-lived app-server sessions under the provider supervisor on POSIX The Codex model-list probe, the hook trust grant and the session index heal each start a short-lived `codex app-server` as a plain child of Orca, and counted on it exiting when its input closes. A wedged Codex (a cold model/list waiting on the network, say) kept running after Orca quit, crashed or was killed. These sessions now run under the provider supervisor that native chat's Codex connection already uses, in its session lifetime, so Orca's exit stops the server's whole process group. The session's end and its deadline share the one stop ladder: end its input (and SIGTERM a session past its deadline), give the supervisor its full stop time, and only then tear the tree down. A missing binary reaches Orca as the supervisor's exit 127 and is mapped back to the spawn error, so the trust-grant telemetry still reads it as a missing binary. Windows and WSL keep spawning the server directly, with the same timings as before. The supervisor now takes only the command line it starts, so the CLI build, which also runs these sessions, no longer pulls in the native chat connection's types. * refactor(supervisor): share the stop of a supervised child process Agent one-shots stop their supervisor with SIGTERM through the shared stop ladder, watching the child's own exit. That adapter moves into one helper so the Codex backfill recovery can use it with its own stop request, instead of a copy. * fix(codex): supervise the app-server that keeps a Codex index backfill alive While Codex rebuilds its session index, Orca keeps a read-only `codex app-server` running for up to an hour so Codex can finish. It was a plain child of Orca with its input held open, so a quit, crash or kill left it running. On POSIX it now runs under the provider supervisor in its session lifetime, so Orca's exit stops its whole process group. Stopping it (done, aborted, or given up) ends its input, as a Codex connection close does; the supervisor then SIGTERMs the group and SIGKILLs it after the grace, and the tree is torn down only if the supervisor outlives its full stop time. Windows and WSL keep the direct spawn and the drain-first probe termination. The spawn and stop of that process move into their own module. * fix(opencode): stop the launch model preflight server with Orca on POSIX Before an OpenCode launch, Orca starts `opencode serve` to read the configured agent and models, then stops it. The server was detached into its own process group and never exits when its input ends, so if Orca quit, crashed or was killed during that preflight (up to 10 s), nothing ever stopped it. On POSIX the server now runs under the provider supervisor in its one-shot lifetime: the preflight's closed input does not stop it, and Orca's exit does. The preflight's teardown asks the supervisor to stop (SIGTERM) and forces the tree only after the supervisor's full stop time; signalling or SIGKILLing the supervisor's own group would orphan the server's. A descendant that ignores SIGTERM is now killed with the group rather than left running once the pipes close. Windows keeps the direct spawn and its existing teardown. * test(codex): pin when supervised session and backfill stops escalate A session past its deadline is SIGTERMed through its supervisor rather than waiting out the stdin-end grace, and its tree is torn down only after the supervisor's full stop time; Windows keeps its deadline kill and 1.5 s close wait. A supervised backfill app-server is stopped by ending its input, with the same full stop time before any teardown. * test(text-generation): run the direct-child suites on the Windows path and cover supervised discovery The commit-message suites that drive fake children mocked the supervisor away on POSIX, so they asserted a direct root SIGKILL that production no longer takes there. They now pin the platform to Windows (with an empty PATH, so host installs cannot answer a bare agent name) and assert the Windows kill, taskkill included. The three tests that check the host's own discovery spawn shape run on the host and read the agent argv past the supervisor's '--'. Model discovery gets its supervised composition: a timed-out Codex discovery settles at once but holds the Codex home until its supervisor has stopped. * fix(supervisor): show a supervised spawn failure in native chat as the spawn error it was Native chat's exit errors carry the provider's stderr tail into Details. Under the supervisor a missing CLI left the supervisor's internal spawn-failure report there instead of Node's own 'spawn <cmd> ENOENT'. The report, its parser and a display formatter now live in one module; the Codex app-server and Claude stream-json exit errors pass the tail through the formatter, which turns a report back into the spawn error and leaves any other stderr unchanged. * fix(supervisor): report a spawn that failed without a pid instead of crashing on its missing pipes When the provider spawn fails outright (EMFILE, ENFILE), Node emits 'error' later and leaves the child with no pid and no stdio. Piping stdin into the missing pipe threw first, so the supervisor died with exit 1 and a stack trace and never wrote its spawn-failure report. The pipes are now wired only for a provider that started. * refactor(supervisor): share the stop of a supervised child process Agent one-shots stop their supervisor with SIGTERM through the shared stop ladder, watching the child's own exit. That adapter moves into one helper beside the ladder, so other supervised children can use it with their own stop request instead of a copy. The caller's breadcrumb site still reaches the forced teardown. Same request, wait and force as before. * fix(codex,opencode): file forced backfill and preflight teardowns under their own breadcrumb sites A forced teardown of the Codex backfill app-server is filed under 'codex-state-db-backfill-recovery', and one of the OpenCode launch model preflight under 'opencode-launch-model-preflight', instead of the generic 'codex-app-server-teardown'. * test(codex): write the stand-in pid report atomically A loaded host let the test read the pid file between its creation and its write (Unexpected end of JSON input); the stand-in now renames it into place. * fix(supervisor): give a session provider its stdin end and grace when its owner dies The owner-death watch went straight to the group SIGTERM and cancelled any stdin-end grace, so when Orca quit or crashed a session provider such as the Codex app-server never saw the EOF that lets it finish writing its state (auth.json, the state database). A session whose owner is gone now closes as an owner's stdin end does: the provider's stdin is ended, it gets the stdin-end grace, and only then the SIGTERM and SIGKILL ladder. A one-shot already had its EOF at the end of its request, so its owner's death still stops it at once. * fix(opencode): stop the preflight server through the shared supervised stop, and trust only a proven stop The preflight's own stop wrapper waited on the supervisor's pipes and counted a forced teardown as proof, though the teardown reports success even when it found no descendants to check, and a supervisor that failed to reap its group exits 1 with its pipes closed. It now uses stopSupervisedChildProcess with its breadcrumb site, and counts the server stopped only when the supervisor ended on its stop signal or relayed the server's own exit. A forced stop, or an exit of 1, returns no context, as an unverified stop did before. * chore(codex): state the supervised stop time in the probe and trust grant deadline comments * test(codex): assert the signals a supervised stop sends itself, not a SIGKILL the mocked teardown never could * fix(supervisor): kill the rest of the provider group once the provider exits on a stop A requested stop waited out the whole SIGTERM grace for the provider's group even after the provider itself had exited, so a SIGTERM-ignoring helper it left behind held every stop for up to 3 s. Under a stop, the rest of the group is now SIGKILLed as soon as the provider has exited, the same rule its own exit already follows. * test(text-generation): cover a supervised Codex discovery past its output limit Model discovery's supervised stop was covered only under timeout. A Codex discovery that floods past the output limit now runs through a real supervisor: it settles with the too-much-data error, its agent is stopped through the supervisor, and the next discovery on the same Codex home starts only after that agent is gone. The direct-child suites' headers now list exactly the supervised cases that are covered. * fix(supervisor): close a provider whose owner is gone the way its owner closes it Owner death gave every session provider the stdin-end grace, so after an Orca crash a Claude session, whose close is a stdin end plus SIGTERM, could keep working on its turn for a second with nobody watching. The spawn spec now names the provider's close request: 'stdin-end' (the Codex app-server drains and exits on EOF, then gets its grace) or 'stdin-end-and-sigterm' (Claude; the default). A gone owner gets that same request. One constant per provider feeds both its spawn spec and its owner-side close, through one requestProviderClose, so the two cannot drift. One-shots still stop at once. * fix(codex): close short-lived sessions and the backfill app-server the way a Codex connection closes Codex finishes its writes and exits on its stdin end, so the Codex connection's supervisor closes it by ending stdin, and an owner that is gone now gets that same close. The short-lived sessions and the backfill app-server now name the same close request, through one shared constant, in their spawn spec and in their own close: Orca quitting or crashing gives them the stdin end and its grace before SIGTERM, instead of an immediate SIGTERM. A session past its deadline still adds a SIGTERM, since it is wedged. * test(codex,opencode): an owner's death drains Codex before SIGTERM, and the preflight stop no longer waits out the grace The stand-in now records when its stdin ended and when SIGTERM arrived. A SIGKILLed owner leaves a Codex session or backfill app-server its stdin-end grace before SIGTERM, and the OpenCode preflight ends within 1 s of its server's SIGTERM even with a descendant that ignores SIGTERM. * test(codex): give the trust-grant deadline tests room for a supervised start on a loaded host At 500 ms a loaded full run hit the deadline before the supervisor had started the stub, which never wrote the pid the test reads. * fix(supervisor): keep the SIGTERM grace for a session's group after its provider exits Killing the rest of the group the moment the provider exited under a stop also reached native chat's closes, so an MCP server, a tool's child or a dev server still in Claude's or Codex's group was SIGKILLed mid-cleanup instead of getting the rest of the SIGTERM grace. The early group kill now applies only to one-shots, where the saved wait was the point; a session's stop is back to waiting out the grace for its group. * test(claude): pin that Claude's spawn passes its close request explicitly The spawn-spec assertion matched the default close request, so dropping Claude's explicit request still passed. The test now checks that the spec is built with the exit-proof ladder's own constant. * refactor(codex): give the Codex app-server close request its own module Other Codex app-server spawns will name the same close request as the connection does. Holding it in its own small module lets them import it without the connection itself. * build(cli): list the Codex close request and the provider supervisor in the CLI project The command-line build runs the short-lived Codex app-server session, which will name the same close request as the Codex connection. Listing the close request, the provider supervisor it takes its type from, and the spawn-failure report the supervisor uses lets the CLI project typecheck that import without pulling the connection in. * test(codex): give a stand-in 20 s to report its pids on a loaded host At a load average near 50 both owner-death tests timed out waiting for the bundled owner's stand-in at 10 s; they pass alone. * test(codex): start the deadline tests' clocks past the stand-in's start, and cover a server that ignores SIGTERM The session and trust-grant deadline tests ran a 2-4 s deadline from spawn, so a loaded host could stop the stand-in before it wrote its pid. The session test's deadline now outlasts its pid-read budget, and the trust-grant deadlines are 8 s. The trust-grant comment said a wedged server may ignore everything but SIGKILL, but that stub dies on its stdin end; a new POSIX case pins that a server ignoring both its stdin end and SIGTERM is SIGKILLed after the SIGTERM grace. * test(text-generation): check the ENOEXEC start failure only where Node reports one On Linux, glibc's execvp hands an executable that is not a program to /bin/sh, so both a direct and a supervised spawn run it and it exits 127; only macOS throws ENOEXEC. The not-a-program case now runs on macOS only; the path-through-a-file case (ENOTDIR) still covers a thrown start failure everywhere. * test(wsl): follow the backfill's wsl.exe spawn into its new process module The WSL invocation boundary lists files that spawn wsl.exe directly. The backfill recovery's spawn moved into codex-state-db-backfill-recovery-process.ts, so the entry moves with it; the count is unchanged. * test(codex): keep factory child_process mocks loadable now that Codex stops reach the process-table reader The backfill recovery and Codex sessions now stop through the shared supervised teardown, whose process-table reader binds execFile when it loads. Five rate-limit fetcher tests mocked node:child_process with only spawn and failed at load: they now mock the backfill recovery, as their sibling fetcher tests already do. The account add-login tests' child_process mocks gain an execFile stub. * fix(opencode): count no forced preflight stop as proof the server is gone A forced teardown walks and group-kills the supervisor's tree, but the server leads its own detached group, so a teardown verdict of 'exited' does not cover members left in the server's group. Only a supervisor that reaped the group itself, by its own stop or relaying the server's exit, now counts as a proven stop. The Windows session-stop test now says why it sees no direct kill. |
||
|
|
86d03ff908 |
fix(native-chat): sending brings the latest into view and follows the reply (#24514)
* fix(native-chat): sending brings the latest into view, with visible jumps to latest and top Sending while scrolled up left the reader parked in old history, and the way back was a faint button. A send now resumes following, opening a row to read it stops following until it is closed, and Jump to latest and Jump to top sit above the composer. Refs #23797. * fix(native-chat): a late or unsent answer no longer moves a reader who scrolled away An answer's reveal is now held from the click and dropped if the reader acts before the host accepts. In the terminal lane, an answer with no terminal to write to reveals nothing. * fix(native-chat): an empty answer no longer moves the reader The terminal lane's reveal now uses the same check the send does, and the send-site guard test ignores comments. * fix(native-chat): preserve upstream retry and question cancellation * refactor(native-chat): leave Jump to top out of this change Jump to top moves to its own PR so this one lands the send reveal and follow behaviour on their own. * refactor(native-chat): keep Jump to latest's existing look in this change The solid, fading Jump to latest button moves to the UI PR (#25702) so this one carries only the scroll behaviour. * Preserve native chat reader intent across submissions and disclosure layout Replace open-row lifetime tracking with bounded position preservation, and tie delayed reveals to the originating reader and session. Keep queued drafts and picker actions from navigating the transcript. Co-authored-by: Kelvin Amoaba <97001695+AmoabaKelvin@users.noreply.github.com> * fix(native-chat): retain reader takeover until its frame ends Pending reader takeover could keep an earlier end target once a duplicate 250 ms gate expired. Let the existing pending-frame check keep the reader's actual offset until that frame ends. Move unchanged interactive reveal and approval projection into their existing modules so both view files meet the line limit. Co-authored-by: Kelvin Amoaba <97001695+AmoabaKelvin@users.noreply.github.com> * fix(native-chat): only a scroll gesture stops following; sends reveal at the press Opening or closing a row used to stop the transcript from following the newest output. A reader at the live end who expanded a tool run then lost the stream. Sends that wait on the host (answers, commands, goals, option changes) only brought the latest into view after the host replied, so the reader watched nothing happen at the press. Now following stops only on a reader gesture: an upward wheel or scroll key when there is content above, a scrollbar press, a touch drag once it has carried the view off the end, or a press on content while already away from the end. Opening, closing, layout changes and the app's own scrolls never stop it, and returning to the end resumes it. Every local send reveals the latest at the press. A message that waits as a queued card, a message from another device, and Stop leave a reader who scrolled up where they are. Queue Resume still reveals only after the host lifts the pause, and only in the pane that pressed it while that pane is shown. Removes the disclosure position hold, the reader-opens wiring, the held reveal and its reader generation tracking. The queued-card decision moves into the outbox send so the session hook stays under the line limit. Also restores the transcript label import and the approval card's verified send that the merge with main dropped. * chore: leave an unrelated fixture as main has it * refactor(native-chat): share the terminal send paths' common options in the composer * test(native-chat): drop a field the main merge declared twice * test(native-chat): the pane's steer still steers the queue, and also reveals --------- Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
faed899cd3 |
fix(ci): run three new SQLite-backed tests in the Node runtime project (#26010)
#25888 and #25766 added tests that import the orchestration database or the structured session runtime without registering them in the Node runtime list, so vitest-sqlite-runtime-boundary fails on main and every PR. |
||
|
|
e717ea6b4c |
feat(native-chat): copy chat images from the right-click menu (#24161)
* feat(native-chat): copy chat images from the right-click menu
Fixes #23904
* test(native-chat): cover Copy image and the open preview in Electron
* fix(native-chat): keep the image preview open by click target, confirm copies, pass PNGs through
The image preview stayed open through a chat-menu click by reading which
element had focus, which depends on the menu's exit-animation timing.
It now ignores an outside interaction whose target is the chat menu, the
same target check other Orca surfaces use for portaled menus. Both
previews drop their color and padding overrides on the dialog surface,
which the design system reserves to the dialog primitive.
A successful Copy image now shows "Image copied"; before, nothing told
the user the copy had finished.
An image that is already a PNG is copied as-is after the size checks,
instead of being decoded and re-encoded, which cost time and could grow
a screenshot past the clipboard size limit.
* fix(native-chat): pass an image through as PNG only when its bytes are PNG
Copy image skipped re-encoding any image whose type said PNG, but a chat
image's type comes from its file extension. A WebP or GIF saved under a
.png name was then sent unconverted, the main process could not decode
it, and the user saw "Image copied" with nothing on the clipboard. The
pass-through now checks the PNG file signature instead.
* test(native-chat): format combined image-copy and session-ID mocks
* Release chat image menu when the retained pane hides
* Validate inherited style directives in their owning scan
* Copy displayed chat images and refuse hidden menu capture
Use full-size sent-image sources, read HTTP images when copying is chosen,
and rasterize browser-decodable formats through the existing converter.
Preserve the current preview surface and prevent a hidden preview from
retaining another copy action.
Co-authored-by: Kelvin Amoaba <97001695+AmoabaKelvin@users.noreply.github.com>
* Revert "Validate inherited style directives in their owning scan"
This reverts commit
|
||
|
|
f6448a1e26 |
fix(native-chat): say which retry a Claude retry is on and what it last failed with (#25818)
* fix(native-chat): say which attempt a Claude retry is on and what it last failed with * fix(native-chat): count Claude's retries as retries, not attempts Claude Code sends its first api_retry frame only after the first request failed, so frame "attempt N of max_retries" is the Nth retry. Say "Retry N of M." and store the bound as maxRetries. |
||
|
|
35e173a5b7 |
fix(native-chat): alert when a structured chat asks for approval or input mid-turn (#25766)
* fix(native-chat): alert when a structured chat asks for approval or input mid-turn A structured Claude or Codex chat that stopped mid-turn to ask for approval or a question raised no OS notification, no phone push and no unread dot: the host's attention feed only fired when a turn settled. The host now announces each newly pending prompt once, from the same feed, and owns the phone push for its own structured sessions; desktops keep presentation only. - host feed: a `prompt` edge per pending approval/question, opt-in on the existing stream; a clean settle that only restates an announced prompt is not re-announced to prompt-aware clients; failures always are - host push: an in-process subscriber pushes the host's paired phones with labels from in-memory metadata, and withdraws a prompt's alert once answered - one attention identity for desktop banner, phone push, dedupe and retirement; delivery dedupes on it instead of the per-workspace burst window - agentSession.acknowledgeAttention routes reading a chat to its owning host, which retires what it pushed for that session via its dismissal record * fix(native-chat): keep prompt alerts under their retirement identity * Fix structured attention read and delivery lifecycle * Allow later reads to retry failed attention retirement * Prevent delayed prompt alerts after accepted reads * Keep attention integration tests across process boundaries * Check structured capability before cursorless attention reads * Keep host retirement fixture in the Node typecheck boundary * Keep attention test seams in the mock boundary * Keep host notification formatting usable without Electron * Declare agents in restored-host attention fixture * Group session attention capability declarations * Retire structured alerts on every read and when a remote prompt ends - Mark read and Mark all read on a chat that was never opened, or is hidden, now cover what lit its row: the read boundary is the later of the transcript cursor and the newest edge the tab was handed, captured at the click. The one boundary closes the banners, retires relayed phone alerts and drives the host acknowledgement; an edge after the click stays live. - An alert from a host older than journal cursors has no position, so a pane read retires it again, banner and relayed phone push, as before this change. - A desktop-relayed prompt alert for a remote chat is withdrawn once that host's live status stops reporting attention (answered elsewhere, cancelled, host restart); cached status after loss of contact never settles it. - The attention acknowledgement is agent-session.attention-ack.v1 with a required journal cursor and strict params; the cursorless no-op is gone. - The dismissal store keeps a record whose origin a newer build wrote, dropping only that origin. - Host reconciliation runs on restore and when a prompt leaves the pending set, not on every journal commit. - Remove unused session-wide retirement (retire, retireMatching, retireMobileNotificationsMatching) and test-only store lookups. * Settle relayed prompts only from host-dated status; reads default to view - A remote chat's status and its prompt edges arrive on separate sockets, so a status row queued before the prompt could land after it and withdraw a relayed alert that was still pending. Settlement now needs a live, non-attention row whose host updatedAt (the journal's latest row time, never decreasing) is later than the prompt's host raisedAt. The owed prompt is cleared only once the settle call succeeds, so a failed call retries on the next qualifying row. - An acknowledgement without captured reads now reads as a view. Mark read, Mark all read, the hover-card jump and both dashboard acks pass 'explicit'; Activity's automatic per-turn read stays a view. * Re-judge relayed prompt settlement after each call; dashboard watching is a view - Settlement is one check of the current mirror, run on each status row, after each settle call returns, and when a prompt edge arrives. A row that lands while a call is out is no longer dropped, and an edge that arrives after its own resolution settles at once. A rejected call is retried only once the mirror holds a different row, so a failing call cannot loop. - The Agent Dashboard's card click acknowledges as explicit; watching the open dialog acknowledges as a view, through both the in-window drawer and the pop-out relay, and no longer re-acks the state the click just read. * Keep a reused read owner findable after its holders release it A chat pane memoizes its read owner. When every holder let go in one commit and the same owner was picked up again (React StrictMode's effect re-run in development, or a pane remounted in one commit, as on a tab move), the release dropped it from the registry and nothing put it back. The attention bridge then found no owner, so a view read carried no cursor: the unread dot cleared but the banner and phone alerts stayed. An owner now names its key again whenever it gains an activation or a subscriber and the key is free, never over a different live owner. * Type the ack test's execution host ids * Keep the turn-completion types in main's turn-completion wire module Main moved these types into agent-session-turn-completion-wire.ts while this branch moved them into agent-session-attention-wire.ts. They now live at main's path with this branch's additions (journal cursor, prompt attention, prompt arm); the unused completion key is dropped. |
||
|
|
53b9ccba6c |
refactor(terminal): add in-place pane layout geometry apply (unused) (#25671)
* refactor(terminal): add in-place pane layout geometry apply (unused) PaneManager.applyLayoutGeometry applies a layout's split orientation and ratios to the mounted pane tree without remounting panes, and never fires onLayoutChanged. Nothing calls it yet; the live-layout reconciler adopts it in a later change. * fix(lint): merge duplicate removal import in delete-worktree failure toast Main (#25668) introduced a duplicate import that fails the focused code-quality lint. * refactor(terminal): share split ratio read; refuse zoomed or dividerless trees in geometry apply |
||
|
|
31f103535f |
fix(orchestration): serve a /clear'd native-chat worker through the session running it (#25875)
* fix(orchestration): serve a /clear'd structured worker through the session running it A structured (native chat) worker the user /clear'd keeps working in a successor session, but orchestration kept acting on the session it was started with, which the clear had closed. Terminal read and @idle status lost the handle, worker-read and the release archive served the pre-clear transcript, worker-show reported the worker exited, worker-stop closed the already-closed session and left the successor running, mail and Dispatch nudges were refused, and the idle sweep could put the successor to rest while its Dispatch was open. One forward walk over the durable session records now answers which session runs a worker's conversation. Worker authority and custody are judged on that session, so every caller holding a worker handle gets it by default; reads, archive, observation, status, group addressing and the incarnation liveness probe use it too. Stop closes the running session and re-resolves, so a clear that commits while the close waits is followed to its successor. The reverse direction (a successor's child env, user takeover, the idle sweep) maps a session to its worker through the lineage root. A successor's idle edge re-derives the worker's Dispatch mailbox, and settlement forgets parked mail on every session of the lineage, derived rather than stored. When the running session cannot be found (host not installed, a successor with no record, a looping chain) observation answers unverifiable, never exited; readers refuse with session_caller_not_live and stop closes nothing. worker-read and the archive warn that earlier conversation from before a /clear is not included. * fix(orchestration): close the review gaps in serving a /clear'd structured worker Review of the first change found places where a /clear'd worker still went wrong, plus a type gap that let the original bug shape compile: - worker-stop refused a structured worker whose old session read exited mid-clear (stopped, but the successor not yet committed); it now goes to the clear-aware stop once identity and ownership are proven. The terminal (PTY) gate is unchanged. - A worker stopped while its /clear successor had never run could never be released: no close writes death evidence for an agent that never started, so the probe read unverifiable forever. A released session that never started anything and whose chat is gone now reads exited; it stays unverifiable while the chat is listed. - worker-read, terminal read and the release archive now read the worker's whole lineage, oldest session first, under the same page limit and byte bound, instead of only the running session. This replaces the "earlier conversation is not included" warning. Cursors key later sessions' items by session, so a cursor taken before a clear stays valid and continues into the successor. - Pointer delivery re-derives a mailbox's target when an attempt ends; if a /clear moved it meanwhile, it delivers once to the new session. - Custody and status reads take the resolved running session, and one typed hold (held, not held, unverifiable) replaces the rebuilt authority in terminal read, worker-show and group addressing. - Smaller fixes: stop keeps an earlier close on its receipt; abandoning a side task no longer forgets the worker's parked mail; a release whose lineage cannot be verified ends release_unknown instead of staying requested forever; mail reach reports an unverifiable lineage as unverifiable, not ended; party resolution scans records once. A /clear typed into a worker's chat is reported as a user takeover, the same as any other message the user types there; a test pins it. * fix(orchestration): tighten the never-ran rule and the follow-ups to /clear handling Round-2 review of the /clear'd structured worker fixes found that the new rule for "a session no agent ever ran is exited" was too loose, and two of the new follow-ups missed a case: - The rule now requires the founding fence, which every reservation moves. A reservation that a restart released without proof, or a failed first start whose exit was never proven, stays unverifiable instead of reading exited. It no longer asks whether the conversation is open (any history read opens it, which stranded a stop, read, release sequence), and it treats only an authoritative tab index without the chat as retired. - Pointer delivery follows a mailbox a /clear moved on a thrown attempt too, keeping the attempt's own failure, and tries each session once. - Releasing a retired worker whose newest session cannot be read keeps the readable earlier sessions' output and says the latest one was lost, instead of freezing an empty archive. - worker-show judges status and addressability on one lineage walk. - The worker-stop comment says plainly that any structured worker that reads exited (a crashed agent too) is closed and its chat hidden; tests cover that and a takeover reported by the successor session. * fix(orchestration): keep a retired worker's earliest readable output across several lost sessions Releasing a worker whose chat was retired and whose newest session could not be read kept the earlier sessions' output only when exactly one later session was lost; after two clears with both later journals gone, release still froze an empty archive although the first session held the worker's answer. The archive now walks back past every unreadable later session, under the same page limit, and its warning says how many later sessions could not be preserved. The rule that reads a never-started session as exited also requires that the record carries no fence floor: a copy restored from backup may hide a reservation that the backup lost, so it stays unverifiable. * test(orchestration): build the takeover test's session record from the shared fixture The changed-code quality gate rejected the test's `as unknown as AgentSessionRecord` stub, which this branch had touched. It now builds a real record with agentSessionRecordFixture and a typed lease, so no type assertion is needed. * test(orchestration): give the cleared-worker pointer harness the sender-name dependency main added |
||
|
|
5345ba34bf |
Fix ownership for terminal and chat file drops (STA-6940, PR 1/6) (#25749)
* Fix terminal and chat file drop destination ownership * fix runtime terminal drop ownership and queued chat retries |
||
|
|
deb737478a |
Simplify reveal and test fixture cleanup (#25978)
* refactor: remove redundant reveal and watcher test wrappers * test: keep closed restore fixtures free of deferred WAL writers |
||
|
|
2137295bb6 |
fix(git): use branch-safe author names as username fallback (#25984)
Co-authored-by: Kyou0203 <kyou12138@gmail.com> |
||
|
|
1eeb8f3f72 | fix(test): drop the duplicate resolveLaunchArgs that two merges both added (#25997) | ||
|
|
7bc9ff947c |
fix(grok): add Grok 4.7 to the fallback model catalog (#25976)
grok 1.0.41's `grok models` now reports `Default model: grok-4.7` and its model cache advertises low/medium/high/xhigh effort for grok-4.7. Seed it as the CLI default with an xhigh ceiling, keep grok-4.6 and grok-4.5, and move the probe spec default in step with the seed. Fixes STA-9652 |
||
|
|
add534475d |
fix(browser): detect Comet in its Windows vendor directory (#25983)
Co-authored-by: Masaki Yamamoto <nnasakick@gmail.com> |
||
|
|
fa42e57659 |
fix(gitlab): select SSH workspace host for merge request and issue writes (#25985)
Co-authored-by: makoto-developer <72484465+makoto-developer@users.noreply.github.com> |
||
|
|
d8c871a1f0 |
Speed up unit tests with cross-runtime duration scheduling (#25967)
* Schedule unit tests across runtimes by measured duration * Keep new SQLite fixture suites on Node after updating main * Inject scheduling timings instead of mocking the module * Keep agent-session database lifecycle contracts on Node |
||
|
|
6352ff3ff9 |
Run nested workspace deletion in the background (#25975)
* Run confirmed nested worktree deletion in the background * Reuse the translated workspace deletion failure title * Supply launch arguments in the Codex session test fixture |
||
|
|
72b84118e0 |
test(e2e): terminal layout parity check against main (#25681)
* test(e2e): add terminal layout parity check for topology refactor PRs Runs fixed terminal-layout journeys in the real app on two builds, captures the renderer topology and the saved workspace session, normalizes volatile values and fails on any difference not declared for a named bug. Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): record quit exit status and report paths main does not reproduce Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): invoke pnpm correctly under corepack and silence the typeless-module warning Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): accept pnpm's forwarded -- in the layout parity runner Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): close parity panes through the user's chord and treat absent maps as empty Driving PaneManager.closePane directly left main to learn of the close from the PTY exit, which raced quit; the keyboard path commits the close in main first. Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): build parity checkouts before running, forward -g, and add a drag-out scenario Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |