mirror of
https://github.com/stablyai/orca.git
synced 2026-10-02 00:02:05 +00:00
1069bb053fadd3033096d1aec394be0e2784cb80
2288
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1069bb053f |
fix(agent-launch): a cwd at the workspace root no longer forces a terminal (#22729)
"Continue in New Session…" always names a cwd, and both the renderer route input and the host launch-mode decision read any cwd as a custom start directory, so the continuation opened a terminal agent even when chat was the user's default. Both now share one rule: only a cwd outside the workspace root (after normalising slashes, Windows case, WSL aliases and the distro's Linux spelling) requires a terminal. A subdirectory still does, because a structured session cannot start there. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
5e6fcca0b3 |
fix(native-chat): date a session by its own lifecycle, not its subagents' work (#22520)
* fix(native-chat): date a session by its own agent's rows, not its subagents' The journal reducer's lastActivityAt is the structured status summary's updatedAt, which the status row uses as its completion stamp and acknowledgement clock. It took the max over every journal row, and a session's subagents write into the same journal after its own agent has settled, so an idle parent was re-dated and marked unread on child work. A row now dates the session only when the session's own agent produced it: not a row whose producer linkage names a subagent, and not a subagent roster row (a subagent-group block), which the session writes but revises on every child transition. The roster rule is derived from the row body; no new persisted field. Replay folds through the same rule, so existing journals are re-dated to their own last row on reopen. Claude: a backgrounded subagent emits no child frames, so its re-dating came entirely from roster revisions (task_updated, task_notification) and from the stale-roster revision written when a journal reopens. Codex: the roster is revised on every child token-usage report; child-thread rows carry no producer linkage yet, and read as the session's own until they do. * fix(native-chat): a reopened journal's verdict on stale work does not date the session Reopening a journal settles rows the previous host left live (a working subagent roster, a live background task) to unverifiable. Those revisions were appended at the reopen moment, and a background-task row is the session's own non-roster row, so a crash-restarted session with a live shell was re-dated to the restart although no agent acted. The reconciler now writes each verdict revision with the row's own observed time. That is one rule at the one writer, covering both settle shapes; the render item's observedAt was already pinned to the row's first write, so nothing the transcript shows changes. The live-transition roster exclusion stays: live roster revisions are written by the providers, not here. The `recovered` row flag is not used as the discriminator: the live unexpected-exit settlement also writes recovered rows, and a clock rule keyed on it would stop dating a provider crash the host just observed. * fix(native-chat): date a session by the reducer's attribution of what a row wrote The clock read producer linkage off the raw row. A lifecycle batch names no row-level producer, a tombstone names none, and a revision may name none while the reducer still attributes the item to a subagent, so each of those dated an idle parent. The clock now asks the reducer: after a row applies, whether any item it wrote is the session's own work; before a removal, whether the item it removes was. * fix(native-chat): date a session's status by its own lifecycle edges A subagent writes into its parent's journal and keeps going after the parent settles. Every one of its rows advanced the summary's updatedAt, and the status row re-dated a done parent to it, so an idle parent read as newly finished and unread on each child step. The host now publishes statusStartedAt beside updatedAt: when the session's own agent entered its status, read off edges only it writes. Idle is when its newest turn ended; working is when the running turn was requested, or the earliest send still unanswered; attention is its own oldest pending ask, or a subagent's when that alone holds it. A turn that recovery settled after its host went away ended when that settle was written, so it reads as a completion the user has not seen; it carries no outcome, so no completion event or notification calls it a success. Render items carry recoveredAt, the recovered row's own write time, so nothing new is persisted. The sidebar bridge and the host ingest date the row and the main agent's clock from it whenever the row shows the main agent's own state, and keep their existing rules for a row child work holds open or a summary from an older host. The status feed republishes when the clock moves instead of on every idle row. * revert(native-chat): keep the journal clock over every row The row filter this branch put on the reducer's lastActivityAt decided which rows could date a session: a list of exclusions that each new row kind could slip past. The session's state is now dated by its own lifecycle edges, so the filter, its attribution helper and the backdated reopen verdicts go back to main. lastActivityAt, and the summary's updatedAt it feeds, is again the evidence clock over every row, including a subagent's. * fix(native-chat): keep republishing an idle session its live child work holds open A row held open by live child work is dated by when each reader saw the publish, and mobile decays a working row whose evidence is older than the staleness window. Suppressing row-activity republishes for every dated idle session froze that evidence, so a subagent running more than 30 minutes past its parent's turn made the row read idle on mobile. Only a session nothing holds open stays quiet on row activity now; its state clock is unchanged. * fix(activity): order an agent's timeline by when each state was seen An answered ask returns a settled parent to its own turn's end, so its done repeats the time of the done before the ask. Activity keyed and ordered events by that time: the new done collided with the old one and was dropped, and the row took its state from the newest-dated event, the blocked ask, so a done parent read Blocked and needed attention. Each state switch now records the `updatedAt` it was seen at. Events are keyed and ordered by that, while unread and "Clear completed" still compare the state's own time, so the answer neither re-lights unread nor revives a cleared done. The row's state comes from the pane's own status entry, so a clear that hid the done cannot leave it reading Blocked either. * test(activity): pin the timeline across repeated asks, a clear, and a stale turn Three parts of ordering the timeline by when each state was seen had no test that failed without them: - A second ask moves the answered done into history. Both dones share the turn's end, so only the history entry's own seen time keeps them apart; without it one done collided with the other and the timeline showed two Blocked events in a row. Three asks also exceed the per-pane cap, which must keep the most recently seen events, not the most recently started. - "Clear completed" on an answered row must cut off past the ask, which is dated after the done, or the cleared row stays listed. A done that the user cleared must also stay hidden once a later ask moves it into history. - A stale working row must not read as running just because the pane's own status says working. |
||
|
|
6ae6ed08bb |
fix(claude): open structured chat without a startup deadline, and make Retry start fresh (#22364)
* fix(claude): open structured chat without a startup deadline, and make Retry start fresh
Publish the Claude session as soon as its process is spawned instead of racing
initialize against a fixed 10s deadline. Prompts sent before startup lands are
held and written in order once it does. An exit or sign-in failure before startup
ends the session with the reason and the CLI's stderr.
A create that failed because the process provably exited now carries
ownerVerdict 'exited', so the client marks the launch failed and Retry mints a
new operation instead of replaying the stored failure.
* fix(native-chat): sending into a chat that failed to start restarts it
* fix(native-chat): a send with no live owner restarts it once
A provider child that timed out or exited hands its lease back, and every
later send was refused agent_session_ownership_unknown. Clients read that
code as "not admitted yet" and resend forever, while only a surface hold
could make a new child, once per mount, with its failure swallowed.
The send now routes to a live owner, otherwise restarts one from the
persisted resume state where resume eligibility allows it (single-flight
per session), otherwise refuses with the new settled
agent_session_owner_unrecoverable. Unverifiable, reserved and handed-off
leases are left alone. The desktop hold now logs its failure.
* test(native-chat): pin the unrecoverable refusal as settled in the outbox
* test(native-chat): pin the release clock after a send restarts an unheld owner
* test: read the sent operation id without a cast
* fix(native-chat): type the send-recovery record lookup as the store returns it
* fix(native-chat): a send ensures its owner before admission, and an unheld owner idles for 30 minutes
* fix(native-chat): a create that throws releases its event sink
A child that dies between spawn and journal attach can still write through
the host's event sink, which attach unbound in onAcquiring and never re-bound
because onAttached never ran. The orchestration released that sink only when
performAttach returned a refusal; a thrown failure (the root-exit path) kept
the sink cached with its queued write, so the next attach's drain barrier and
runtime shutdown's flush waited forever.
Also pins the publish-on-root-exit clause for a start that never proved:
deleting it reddened nothing before.
* fix(native-chat): a resend the journal answers restarts nothing, and a send joining a restart rebases from the fence it replaced
* fix(native-chat): the host learns a Claude start positively, and persists only proven options
A publish-first create used to read the session's options before Claude had
answered initialize. With startup pending that read fell back to the built-in
catalog's default, so `record.options.model` was persisted as `sonnet` for
every user whose CLI default is something else; an owner handoff or a reopen
then replayed `set_model('sonnet')` and silently switched their model.
The adapter now reports `started` once startup facts are applied and saved
options restored. The host keeps a `providerChildPhase` on the session it
owns: a starting child hands over nothing but the saved options as intent,
and the `started` event re-reads the options as fact and persists them through
the same record write a user's option change takes. The status summary carries
`hostExecutionPhase` (optional, wire-safe), and the chat pane says the agent is
still starting instead of showing nothing.
A child whose exit already reached the adapter before acquire returns is no
longer handed over as live; the create fails with the CLI's diagnostic.
* fix(native-chat): a hold and a send that find the owner gone share one restart, and a send the ledger already holds restarts nothing
* fix(native-chat): a failed create answers one refusal shape, stamped once at the boundary
A create whose Claude process was seen to exit answered twice in two shapes:
the first call threw a generic runtime error, and only the replay of the same
operation carried the `ownerVerdict: 'exited'` refusal that lets a client
retry under a new operation. Three sites stamped the verdict and the store
failure path stamped nothing.
The first-hand root exit is now returned as the refusal on the first call,
with the provider's own diagnostic as its message. The verdict is stamped in
one place, at the boundary of the attach, from the durable row the operation
settled to, so every refusal shape answers the same fact and no site can
forget it. The per-site stamps are gone.
* fix(native-chat): a send into a session whose child ended restarts it before admission
A session that published and then lost its Claude child before startup (not
signed in, for one) keeps a released lease and a chat the user can still type
into. The send was refused as ownership-unknown, the outbox parked it as
pending admission, and nothing ever restarted the child: the message sat
there until the user closed and reopened the tab.
A send reaching a session with no provider child now runs the same resume a
surface's first hold runs, before the write is admitted. The resume reserves
a new fence, so that send is answered stale with the published fence and the
client's outbox re-drives under it, as after any fence change. A resume that
fails is not this send's answer; admission reports the lease as it stands.
* chore: restore pnpm-lock.yaml to origin/main (local pnpm rewrote it)
* test(native-chat): pin the pre-handover exit as a failed acquire; stub the status feed in the delivery test
An exit the adapter observes before acquire returns now fails the acquire
with the CLI's diagnostic instead of handing over a dead child; the
published-then-ended path stays pinned by the slow-init startup case. The
delivery test renders the pane, which now activates the host status feed.
* test(native-chat): a same-ID re-hold over the wire joins the one resume, and a replay reopen goes on the idle clock
* test(native-chat): a re-hold that joins a failing resume proves one resume ran
* fix(native-chat): a create whose child was proven gone answers the refusal on the first call
The previous change answered a first-hand root exit as the exited refusal on the
first call, but the common failed start never took that path: when the close
ladder proves the whole tree dead the acquisition error is a plain one, the
store-failure classifier rethrows it, and the client still saw a runtime error
first and the refusal only on replay.
The cleanup that proves the child gone now names such a failure
`AgentSessionAcquisitionExitProvenError`, carrying the provider's diagnostic,
unless it already names its own verdict (a refusal, a typed exit proof, a host
store code). The attach answers both proven-exit kinds as the refusal its replay
gives. How a failed acquisition settles and how it is first answered now live
beside the verdict stamp, in the failed-create module.
* test(native-chat): pin the outbox re-driving a stale-refused send under the resumed fence
A send into a session whose child ended is answered stale once the host has
restarted the child. The outbox keeps that operation queued and blocked, and the
fence change the resume publishes re-drives the same operation under the new
fence; the host admits it.
* fix(native-chat): a child restarted for a send nobody holds is still released
The restart a send runs for a childless session takes no holder, on the premise
that the sending surface already holds one. A one-shot writer holds nothing, so
the child it restarted had no release clock and lived until the app quit. The
write resume now arms the clock when no holder is present, as the first-hold
resume already does. The send-after-failed-start cases also pin that the stale
answer's operation is admitted when re-sent under the new fence, and that two
racing sends restart the child once.
* test(native-chat): pin the picked Claude model across a resume whose child starts on its own default
The started event re-reads and persists what the child reports. A resumed child
answers initialize with its CLI default before the saved pick is restored over
it; the record must hold the pick while starting and after started.
* Revert "fix(native-chat): a child restarted for a send nobody holds is still released"
This reverts commit
|
||
|
|
f0b3f44f10 |
feat(agent-session): let the host own a chat's tab id and let a create reserve it (#22616)
* feat(agent-session): let the host own a chat's tab id and let a create reserve it A structured chat's tab id was derived from its session id by every layer that needed one: the renderer, the host snapshot and the status address each built their own spelling. The join between a conversation and the tab that shows it must be a pointer the host owns, not a derivation each client repeats. The session record now carries surfaceTabId. A create pins it: the tab half of the pane agent.launch reserved, an optional tabId on agentSession.create, or a host-minted UUID. Records written before the field existed are backfilled at open with the string clients derived, in memory at once and on disk with the store's first transaction, so nothing keyed by it (read state, notification ids, worker rows) moves on upgrade. A second record under a held id is refused. Only the record and the two create wires change here. The snapshot still publishes agent-session:<sid> and the renderer still derives its local id; those move in the next two changes. agentSession.create is a strict object, so the field is advertised as a capability a client checks before sending it. * fix(agent-session): record the derived tab id for an unreserved create A create that reserved no tab minted a random UUID that no reader uses: the renderer, status address, worker rows and host-shared read state all still key by structured-agent-session-<sid>. Persisted, that id would move every chat created before readers switch to the recorded one, orphaning its read state and worker rows the way the backfill exists to prevent. An unreserved create now records the derived id, the same rule the backfill applies, so the record always matches the prefix every existing key uses; an opaque mint belongs with the change that moves the last reader. Also: - a chat tab id must be a host tab id on the record, the create wire and in admission, matching what agent.launch already requires of paneKey; a web-surface id would decode as another tab - the stored launch-result guard checks the structured outcome's tabId - comments no longer claim a retry naming another tab conflicts; replay keys on the attach fingerprint and answers with the recorded id (now pinned) - the wire refusal test used a non-hex digest, so the schema refused it for that reason; it now reaches the tab id rule - pin that the reload path refills the id without forcing a save * test(agent-session): correct the tab-id fingerprint comment to match replay |
||
|
|
98584332a3 |
fix(native-chat): record which Codex agent produced each journal row (#22532)
* fix(journal): a batch revision restates the producer of each row it revises The reducer rebuilds a row's producer linkage from its NEWEST revision, and absence is a positive claim: no agent id means the session's own agent wrote the row. So any revision written without the stamp hands a subagent's row back to its parent, permanently. Three host paths revise rows they did not write, from the render item they already hold, and all three dropped the stamp: - answering a prompt re-appended the asker's row with the fence only; - dead-generation settlement failed running tool calls and cancelled pending prompts through a lifecycle batch; - stale-session settlement on acquire cancelled lost prompts the same way. The batch path could not carry a producer at all: linkage was removed from the batch row because one row covers N mutations, with a note that a mixed batch would have to stamp per mutation. Dead-generation settlement is such a batch already, and Codex settlement is about to become one. So each item mutation now names its own producer, inline like the row base. A mutation that names none falls back to the row's linkage, which is what a batch read before. Parse sanitizes a bad per-mutation id the same way it does a row's: the field is dropped and the mutation kept. No schema version bump. An older host's mutation validator ignores unknown keys, so it reads a stamped mutation as the session's own, which is exactly what it shows today. Old journals carry no stamp and read as before. Turn revisions still carry nothing: a turn is the session's unit of work, and the live-turn scans rely on a turn row never carrying linkage. The note recording that invariant is updated to the new write sites. * fix(native-chat): attribute a Codex subagent's journal rows to the subagent Codex journals every thread on its app-server connection into the session's journal, and a spawned subagent's items arrive on the child's own thread. None of those rows carried producer linkage, so under the journal's rule that absence means the session's own agent wrote a row, every child's command, message, reasoning, prompt and status row read as the PARENT's: the parent could show its child's running command, its child's reasoning as "thinking", and its child's prose as its own latest line. The Claude lane's model is reused, not reinvented: the same fields and the same absence rule. What differs is how the producer is known. Orca opens exactly one thread per app-server, so any other thread is one Codex spawned. That decides WHETHER a row is a child's from its first frame, announced or not, and the thread id is final at once: it is never re-minted the way a tool-call reference is, so no correction ledger is needed for identity. - agentId: the child thread id, the same id the status side keys a Codex child on. - parentAgentId: the thread whose stream carried the child's `started` activity. Codex emits that item on the spawning agent's own session, so a child that spawned a grandchild is named; the session's own thread is not. Other activity kinds ride whichever agent acted and are not used. - producerKind: 'agent'. - attempt: which run of the child the row's own turn was, counted from the child turns the roster already observes; absent on the first run. Taken from the row's turn rather than the child's latest, so a persistent shell that outlives its turn keeps its run across revisions. - providerParentRef is omitted: a Codex child's frames carry no parent reference of their own beyond the thread id, which is already agentId. One resolver, owned by the roster (which already owns what is known about each child thread), is handed to every writer: items, streams, generic and summary rows, prompts, compactions, goals, and the three settlement batches. The session-end settlement mixes every thread's rows in one batch, so each mutation names its own producer. Turn rows stay unstamped: Codex writes them only for the primary thread. The spawn-group roster row stays unstamped on purpose: a child's frame can trigger its write, but it is the parent's list of its children. The translator's construction moves to a parts module so the translator stays a router under the line cap, and the item streams reuse one append-and-publish helper instead of two copies. Children are never swept at turn end; nothing here changes that. * test(native-chat): pin Codex subagent attribution at every writer and every parent reader Two layers, so a stamp that is correct in the store and never read, or read and never persisted, cannot pass. The readers, through the real path: translator, deferred sink, on-disk journal, snapshot. Each is a defect on main: the parent named its child's running command as its own tool, read its child's reasoning as itself thinking, showed its child's compaction as its activity line, and quoted its child's prose as its latest line (checked after closing and reopening the journal, so the stamp is read back from disk). The transcript still renders the child's rows. The writers, through a sink that records the linkage of every plain append, batch mutation and lifecycle transition: start, streamed checkpoint and completion of one command all restate the child; a row that beats the spawn announcement is still the child's; a grandchild names the child that announced it, while an `interacted` activity names no parent; a follow-up turn is the child's second run, and a shell that outlives its turn keeps its own; the exit batch settles each thread's rows under its own producer and the turn row under none; a child's provider frames, approval and goal rows are its own; nothing is stamped while the session thread is still opening; and the spawn-group row stays the parent's. * test(native-chat): pin linkage forwarding on the sink's lifecycle-transition path A Codex child's goal row is written through a lifecycle transition, so a sink that forwarded only the fence there would file the child's goal as the session's own. * test(native-chat): type the Codex item fixtures as thread items * refactor(journal): keep a row's producer across revisions that name none The reducer took a row's producer linkage from its newest revision, so every writer that revised a row it did not write - a prompt answer, a dead-generation or stale-session settlement, the reopen sweep of stale subagent rosters - had to restate the producer or silently hand a subagent's row to the session's own agent. Three of those writers had been patched to restate it; the next one to forget would reintroduce the bug. Attribution is now fixed by a row's first write. A revision that names no producer keeps the row's existing linkage; one that names any replaces the whole bundle, which is how a provisional stamp is still corrected in place. A row re-created after a tombstone starts with nothing. The reducer runs the same fold on replay, so the kept producer survives a reopen. The three restatements are removed. Per-mutation linkage on lifecycle batches stays: a batch can create a row (a Codex child's prompt, or a child's item settled before any checkpoint landed) and one batch can mix producers. * test(journal): pin producer inheritance in the reducer and across a reopen A revision naming no producer keeps the row's, on the plain item path and in a batch settling a child's row beside the session's own; one naming any replaces the bundle wholesale; a tombstone clears it; a stale revision cannot touch it; and a reopened journal replays it exactly as it was folded live. * refactor(codex): name the translator's writer factory for what it builds * docs(codex): say why a settled row names its producer * test(journal): drop a producer test the stale-revision guards make unreachable The stale revision is dropped whole by two independent guards before the inheritance rule runs, so its producer assertion could never fail; the reducer's own stale-revision tests already cover the drop. Also say what the batch sink does forward: each mutation's own producer. |
||
|
|
5610b11703 |
feat(ipynb): create a .venv when pip is locked out, and show ipykernel setup progress (#22710)
* feat(ipynb): set up ipykernel in a new .venv when pip is locked out, and show install progress in the dialog * fix(ipynb): drop the retired installFailed string from translated catalogs * refactor(ipynb): drive the setup dialog from one setup state; fix review findings - Kernel status now only describes the kernel; a single `setup` object (base, offer, phase, error) drives the dialog, replacing the extra statuses and the externallyManaged/setupError fields. - The picker's 'Create virtual environment…' opens the same dialog; success switches through selectEnvironment, so a running kernel is only replaced once the venv exists. - Async setup results are dropped when the dialog they belong to is gone (tab closed/reopened). - main verifies ipykernel imports after pip, reuses an existing .venv instead of re-running venv over it, and explains failures that printed nothing. - Windows copy command guards the install with if ($?); notebooks at a filesystem root get a correct .venv parent; 'Try again' shows for both retry paths. * fix(ipynb): close the setup prompt when another Python is picked |
||
|
|
85642d0d88 |
fix(agent-status): count only agent work in stats, and read a Grok background subagent as working (#22474)
* fix(agent-status): renderer and recorder consumers read the question they mean Since #22295 a row's combined `state` reads `working` both while the lead's turn runs and while a subagent or background shell outlives a settled lead. The lead's own state now rides beside it (`lead`); each consumer in this slice reads the question it actually asks. - Smart sort and the Activity unread badge keep reading the combined state: their classes and rows are what the sidebar shows. Pinned with tests, including a restored `lead.state: 'working'` row that must never read live. - The stats recorder asks "was an agent executing" and now reads a shared derivation (`isAgentExecutionOwed`): the lead's turn, or live agent child work holding a settled lead's row open. A settled lead's background shell no longer accrues "Time agents worked". Old hosts without `lead` fall back to today's read; restored and replayed rows still never open a session. - The `agent.status.changed` plugin event gains `lead` as an optional field through one tested projection; `state` keeps its meaning and restored rows still project to nothing. - A Codex root Stop that follows an inferred interrupt keeps the `cancellation` verdict, as the Claude lane already does at its turn boundary, on both the hook and relay paths. * test(agent-status): pin that a child's approval wait no longer splits the recorded span The recorder's move to the lead fact quietly changed one more story: a Codex child's PermissionRequest turns the combined row waiting while the root's own turn keeps running. The old state read closed the span there and minted a second spawn on resume; the new read keeps one span, because the lead never stopped. Pin it at both boundaries (the shared derivation and the recorder) so the change is deliberate, not incidental. * fix(agent-status): date stats edges by this host's clocks and scope the accrual predicate to stats The recorder dated a start by the producer's mainAgent.stateStartedAt. An SSH host stamps that with its own clock while every stop is stamped locally, so each span gained or lost the clock skew. The same clock also survives a row that briefly lost the fact (an OSC repaint to another state), dating the reopen before the close already sent, and a subagent reopening a monitoring row took the row clock the hook lane pins to the main agent's turn start, re-billing the whole watch-loop window. Edges now use the row clock when the row settles or pauses and the evidence clock otherwise. Rename isAgentExecutionOwed to isAgentTimeAccruing and state that it is the stats question, not a liveness gate: it excludes watch loops, which lifecycle gates must keep treating as live. Note on the plugin schema that mainAgent.stateStartedAt is the execution host's clock. * fix(agent-status): pause agent time while the row waits on the user, whoever asked Time agents worked now accrues only while the combined row reads working. A child's approval or question wait pauses the clock exactly like the main agent's own prompt, and the pause edge is dated by the row's own clock. * refactor(agent-status): read agent time from the combined row and date edges by the row's own state Time agents worked now accrues while the combined row reads working and is not a watch loop. The shared fold emits monitoring only for a settled main agent, and hosts that predate the main agent fact did the same, so this is the same answer on every new-host row without reading mainAgent, and it applies the watch-loop rule to older hosts too instead of billing their monitoring windows. An edge that leaves working is dated by the row's state clock; an edge inside working is dated by the evidence clock. This also stops a live repeat of a hydrated working row (any row without the main agent fact, such as an OSC row) from dating its start at the persisted state clock from the earlier runtime. * fix(grok): read a background subagent as agent work, not a watch loop Grok's end-of-turn Stop lists each in-flight background task with its type (shell, monitor or subagent). Orca filed a running subagent with the shells, so a Grok subagent that outlived the main agent read "Monitoring background tasks" and, with the stats recorder now skipping watch loops, stopped the "Time agents worked" clock. Map shell and subagent entries to the shared child-work kinds and let the shared liveness classifier decide: any live subagent keeps the pane working, a shell alone or an active stop hook stays monitoring, monitors stay excluded. * test(agent-status): cover a waiting child in the fold's every-input accrual check Since the shared fold learned a child's human wait, a waiting child makes the row wait, so it must not accrue agent time whatever the main agent is doing. The exhaustive check now includes that input. |
||
|
|
ad6cb0e05c |
fix(worktrees): version every catalog publication so a stale listing cannot undo a create (#22507)
* fix(worktrees): version every catalog publication so a stale listing cannot undo a create
A worktree listing is a snapshot from when its scan began. The renderer treated any
authoritative listing that lacked a known worktree as proof of deletion, judged at apply
time against the live store, so a listing delayed past a create reply purged the new
workspace: tabs wiped, selection cleared to the landing, pending structured launch
tombstoned so the host session was closed the moment it published. #22311 re-runs a scan a
mutation overtakes, which covers a bump during the scan but not a reply that is simply
applied late, on the host or in the renderer, or a refresh that joined an older one.
The host now stamps every listing with the catalog version its scan began at (the existing
per-repo scan generation, scoped by a per-process epoch) and every create and remove reply
with the version the mutation produced. Clients keep the newest version applied per repo
and host; a listing older than that is not applied at all, not its rows, not its purge, not
the pre-merge terminal teardown. Coalesced joiners inherit the reply and therefore the rule.
Fields are optional on the wire; older hosts and clients keep today's behavior.
* fix(worktrees): relist after a refused stale listing and stop version churn
- A refused listing can be the only answer a caller gets (a change-event
refresh that joined an older in-flight listing), so fetchWorktrees lists
once more; that listing scans at or past the applied version.
- An equal catalog version keeps the held object, so a no-op listing no
longer writes store state on every refresh.
- Versions the client cannot order are treated as unstamped at ingest.
- worktree.rm takes the repo from its id selector instead of resolving the
worktree a second time, which also stamped nothing for an id two hosts share.
- A removal on one of two hosts sharing a worktree id records its version.
* test(worktrees): pin the scan-generation bump right after git worktree add on every create path
A listing is stamped with the generation its scan began at, so a create must
advance it before any post-add work can yield. Pins the local desktop, SSH and
runtime local create paths.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(worktrees): bump the scan generation right after git worktree remove
A listing is stamped with the generation its scan began at. Removals bumped it
only at the end, after watcher release, push-target cleanup and the metadata
purge, so a listing scanned before the git removal and one scanned after it
could carry the same sequence. Applied out of order, the older one restored the
removed row until the removal reply. Bump right after the git removal on the
desktop local, desktop SSH and runtime paths, as creates already do.
The ordering pins now witness the generation at the first step after the git
mutation rather than at the re-list, so moving a bump past any intervening
await fails them.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(runtime): stop leftover worker-recovery retries from firing into later tests
The legacy worker terminal recovery retry timer re-arms itself and had no
way to end, so a runtime from one aggregator test kept rescanning repo-1
through the shared listing mock during later tests, consuming the listing a
lineage create expected ("Worktree created but not found in listing").
Give the controller a dispose() that cancels pending retries and refuses to
re-arm, and have the runtime test harness dispose every controller it
constructed after each test.
* Revert "fix(runtime): stop leftover worker-recovery retries from firing into later tests"
This reverts commit
|
||
|
|
25d7c21fcb |
feat(native-chat): show context window usage in the composer (#22301)
* refactor(native-chat): move the composer's stop action into its own hook The composer is at its line budget; lifting the stop action out makes room for the context usage ring without changing what Stop does. * feat(native-chat): record the Claude CLI's context window facts on the structured journal A structured Claude session now keeps what the CLI says about its context window on journal rows, so every client reads the same answer and a restart replays it: - Each main-thread assistant response keeps its API usage. A subagent's response measures its own window, so it carries none. - The turn a result settles records the session's window: the largest contextWindow across the result's per-model usage, since side calls to a smaller model report their own smaller window. - After a result and after a compaction boundary the host asks the CLI for its /context breakdown (5s bound) and records the answer on the current or last turn. An answer is dropped when the main conversation moved, a send was accepted, a newer request was issued, or the session was released while it was in flight; a failure or an older CLI leaves the row unchanged. - A compaction boundary or conversation reset records that the used count is unknown until the next response or report, so the pre-compaction size is never shown as current. Every part carries its own host clock, and the reader takes the newest, since a revised turn row keeps its place in the transcript. The persisted validator admits every value the writer can write, including a zero auto-compact threshold: a row replay rejects truncates the journal from that row. * feat(native-chat): show context window usage in the composer A structured Claude chat shows a ring beside send once the journal can state the session's context usage. Hovering shows used/window with a bar and, when the CLI has reported its breakdown, one row per CLI category as a share of the window, largest first. Between reports the ring shows the newest response's usage against the newest window the CLI reported, marked as estimated. Before the CLI has reported any window, and after a compaction or reset until the next response, there is no ring. A terminal-backed chat shows none. * fix(native-chat): measure the context ring against the main thread's model window The result's per-model usage is cumulative across the session and includes subagents and side calls, so the largest window was often not the one the main conversation runs in: after switching from a 1M model to a 200k one, or when a subagent ran on a larger-window model, the ring read against the wrong window after every turn. Pick the entry named by the model that served the newest main-thread response, and among its [1m]/non-[1m] entries the one the result moved; fall back to the largest only when nothing names it. * fix(native-chat): keep the context ring moving through tool-only responses The live estimate lived on assistant message rows, and a response with only tool calls or thinking writes no message row, so the ring froze through long tool loops and stayed hidden after a mid-turn auto-compaction until the next text reply. Record every main-thread response's usage on its turn row instead, once per response, so the selector sees each one. * fix(native-chat): read the ring's window from the model the turn's init names The CLI keys per-model usage by the main loop's model string, [1m] included, and every turn's system/init frame carries that exact string, while a response drops the suffix. Match the init's key first, so a session that switched between the 1M and 200k variants of one model reads the right window; fall back to the newest response's model, then the largest entry. * fix(native-chat): correct the composer's control-order note for the context ring * fix(native-chat): keep a running turn's context facts when the host settles it A turn row now carries the live context estimate while it runs. When the host settles a running row itself (a crashed or stale generation, a close the translator never saw), it rebuilt the record field by field and dropped those facts, so after a crash the ring fell back to an older turn's size, or to a pre-compaction size the dropped reset had superseded. * refactor(native-chat): revise Claude turn rows from the journal so the ring survives a restart The context ring's facts were written to turn rows through an in-memory list of recent turns. A new translator is built on every acquisition, so after a restart or reattach that list was empty and every fact for a turn that was not open was dropped: a /compact as the first action after a restart never cleared the ring and never showed the fresh breakdown. Every Claude turn-row write is now a revision of the row as the bound journal holds it when the write runs. The sink gains a resolved revision that reads the target row and its body at execution; the queue runs one operation at a time, so the read-modify-write cannot interleave, and revisions are never coalesced. Lifecycle writes own the lifecycle fields and context writes own contextUsage; each keeps every other field. Only the open turn is kept in memory. A fact with no open turn lands on the newest turn row, and a report lands on the turn it was requested for. The persisted facts are simplified to a window, which now names the model it was measured for, and a single used part (report, estimate or unknown) that each write replaces. The ring reads the newest turn row carrying each part, and hides an estimate whose model the window was not measured for instead of dividing by another model's window. A turn opening, and a reset, count as activity, so a late report can never land behind a newer turn. Host settlement of a stale running turn now drops only the fields its verdict owns, so context facts and any field a newer build wrote survive it. * fix(native-chat): keep a turn row whose context facts this build cannot read Context facts are validated deeply, so one malformed or future-shaped fact made the whole turn row malformed, and replay truncates the journal from that row on. Replay now drops unreadable facts from a turn row, in item rows and in settlement batches, and keeps the row, the same way it already drops producer linkage it cannot trust. The ring shows nothing for that turn instead of the session losing its history. * test(native-chat): pin that a child exit mid-turn keeps the ring's last size A lifecycle-only revision, the end a turn gets when its child exits without a result, must keep the context facts the row already carries. * perf(native-chat): revise a named Claude turn row by key instead of scanning the journal Every Claude turn-row write walked every reduced journal item to find its row, even when it already knew the row's identity, so a long session paid O(items) per write on the main process. The journal now answers a keyed read, and a context report names its turn by row identity rather than turn id, so only a write made while no turn is open still scans. * fix(native-chat): tell a 1M window from a 200k one of the same model Responses drop the [1m] suffix, so after a switch between the 1M and 200k windows of one model the running turn was measured against the previous turn's window until its result arrived. An estimate now records the turn's init model, which keys the window exactly, and the reader requires the full model id to match. * fix(native-chat): show no ring for a context kind a newer host writes A paired client reads turn rows from the host unvalidated, so a used-count kind this build does not know fell through to the estimate branch and threw reading its missing usage. Only the kinds this build can measure now produce a ring. * fix(native-chat): keep the context ring through plan-mode turns on another model Plan mode can run a turn on a model the turn's init does not name (opusplan upgrades to Opus's 1M window). The estimate then carried only the response's id, which drops [1m], and the exact comparison against the window hid the ring for every plan-mode turn. The estimate now records the response's id beside the init's exact key, and the reader matches the base model only when no exact key was recorded. * refactor(native-chat): pair the context ring's window by model change, not by model id The ring divided the newest response's size by the newest window only when their model ids matched, which meant comparing ids from the init frame, the response, per-model usage keys and canonical ids. Those disagree in plan mode and across 1M and 200k windows of one model. The writer now knows when the model may have changed: after a model or permission-mode write that changes the value, when a restore cannot put the stored model back, and when a main-thread response comes from a different model than the one the window serves (an approved plan). It then marks the size unknown, holds estimates, and asks the CLI for its context report, which states the new model's window. Any new window, from a report or a turn result, releases the hold. The reader compares nothing: a report, or the newest estimate over the newest window. Turn rows no longer store window.model, window.canonicalModel, estimate.model or estimate.responseModel. * fix(native-chat): keep a late context report's window when only its count went stale * test(native-chat): pin that each turn's init lets its result restate the context window * fix(native-chat): open the context card on click and tap * fix(native-chat): wait a beat before a mouse hover opens the context card * fix(native-chat): publish each context write in the operation that makes it A context report answers after the turn's last frame, so a revision that waited for the next frame's publish reached live clients only on the next turn. Context writes now queue their revision and its publication as one operation. * fix(native-chat): write million-token counts with a capital M A lowercase m read as minutes on the context card. * fix(native-chat): keep the context card open while the pointer crosses into it The card closed the moment a mouse left the ring, so the pointer could not cross the gap into the card. Leaving now waits a beat, and entering the card cancels the close. * fix(native-chat): let Escape close the context card without stopping the agent The card keeps focus in the composer, so the Escape that closed it also reached the composer and interrupted the running turn. The composer now skips an Escape an open layer already handled. * fix(native-chat): show the context ring when the chat has not loaded the turn it belongs to The ring read context facts only from the rows the chat had loaded, so a reopened chat whose recent page started after the last turn row, or a live turn longer than the retained window, showed no ring until the next turn. The host now derives the newest context facts from its whole journal with the same selector the chat uses, and returns them on agentSession.options for sessions that write them. The chat prefers each fact its loaded rows carry and takes the host answer for a fact they lack. When a live batch revises a turn row older than the loaded window, the chat asks for options again so that answer stays current. * fix(native-chat): bound context refresh reads and refresh when the turn row is trimmed Each turn-row revision the loaded window missed started its own options read. Those reads share the session's host queue with sends and interrupts, and each asks the CLI for its settings, so a burst could pile reads in front of a user action and discard every answer before it landed. The chat now keeps one options read in flight and at most one behind it. A live turn longer than the retained window also lost its turn row to the trim without asking for a fresh host answer, so the ring fell back to the answer read at turn start until the next response. Trimming a turn row now asks again, like a dropped revision does. * fix(native-chat): show the context ring from the first response, sized from the session's model A new session has no measured window until its first result, so the ring stayed hidden for the whole first turn. The host now keeps the window the applied model's name implies (1M for a [1m] name, unknown for default, 200k otherwise) and writes it beside an estimate when the journal holds no window, or after a model write, until the result or the CLI's report replaces it. * fix(native-chat): imply a context window only from a [1m] model name A bare model name does not fix the window: first-party runs today's opus, sonnet and fable models natively at 1M while a gateway or cloud provider runs them at 200k, and opusplan and haiku run another model in plan mode. Sizing their first response at 200k read the ring about five times too full, so only a [1m] name implies a window now; any other name waits for the result. * fix(native-chat): size the first response from a report taken before any turn A model picked in a chat with no turn yet asks the CLI for its context report, but with no turn row the report's write lands nowhere. Recording it still marked the journal as holding a window, so the first response wrote none and the ring stayed hidden until the turn's result. The report's window now serves as the fallback a response writes while the journal holds no window, and recording a report no longer assumes its write landed. * test(native-chat): move the fake Claude connection out of the structured integration suite The context-report delivery case pushed the suite past the 800-line limit, failing repo-wide lint. The fake child now lives in its own fixture. |
||
|
|
ca75bc4c8d |
fix(orchestration): type a request ahead of pasted dispatch briefs so Claude workers follow them (#22582)
* fix(orchestration): type a request ahead of pasted dispatch briefs so Claude workers follow them Claude Code wraps a bracketed paste in <pasted_content> and tells the model to follow instructions inside it only where the user's own message asks. Orca sent the whole dispatch brief as a bare paste, so Claude workers (Opus 5.5, Sonnet 5) refused it as suspected prompt injection. Every dispatch path now types a short lead line in the same PTY write as the paste frame, the preamble drops shouted rules, and dispatch detection accepts the lead line and pasted_content wrapper. Fixes STA-8200 * refactor(orchestration): tidy dispatch lead-line delivery after review - Share one dispatchPreambleSendOptions() across the four dispatch paths. - Fold every C0 control and DEL out of the typed lead line. - Bound the <pasted_content> tag scan and let compaction return null for non-dispatch prompts, removing the separate detector. - Restore the stay-off-other-channels rule in plain wording. - Test through the real status normalizer and trim duplicated assertions. Refs STA-8200 * test(orchestration): guard coordinator auto-dispatch lead line - Capture send options in the coordinator runtime fake and assert the auto-dispatch send uses dispatchPreambleSendOptions. - Drop the helper test that only restated its literal. - Share DispatchPreambleSendOptions with the coordinator runtime contract. Refs STA-8200 * docs(orchestration): fit the pasted-spec note inside the kernel line budget Refs STA-8200 * docs(orchestration): drop the pasted-spec note from the coordinator guide The typed lead line is the fix; the advisory note cost always-loaded context. Refs STA-8200 * fix(orchestration): type the dispatch lead line only for Claude agents Codex discards typed text that shares a PTY write with a bracketed paste, so the lead line never reached it. Only Claude Code needs the lead to follow a pasted brief, so known non-Claude agents now get the pre-lead bytes and unidentified agents keep the lead in case they are Claude. Refs STA-8200 |
||
|
|
b4d732685c |
feat(agent-status): combine Codex child work through the shared main-agent status fold (#22475)
* feat(agent-status): combine Codex child work through the shared main-agent status fold * docs(agent-status): correct two comments the waiting child-work arm made stale A child failure reported in place as `blocked` now pins the row `waiting`, not `working`; and no relay ever sent an unfolded `working` beside a waiting child. * fix(agent-status): only a waiting child asks for a human A child's `blocked` state means its task failed (the only producer maps a failed background task to it, and the background-task view labels it "failed"), not that a human must act. Folding it into the waiting arm would surface a failed child as needs-you. It stays live work, as before this series. * docs(agent-status): say a waiting child, not a blocked one, makes the row wait A child's blocked state means it failed; only its waiting state feeds the waiting arm. Two fold comments, a test describe and two parity story names still called the waiting child blocked. * docs(agent-status): name where a child's wait is still lost, and pin the structured lane's real input The doc said the Claude hook lane's rows match Codex and that every lane feeds a child's wait into the fold. Neither holds: Claude keeps the wait in one slot the next main agent event overwrites, the structured lane turns a child's prompt into the main agent's own attention, and Codex drops its roster on a root Stop when it tracks no child transcripts. The parity story now drives the structured lane with the input it actually receives. |
||
|
|
7a4f080086 |
revert: #18790 (orchestration incarnation reap fallback and bundled Freebuff agent) (#22601)
This reverts commit
|
||
|
|
3ea15dd0a2 |
fix(native-chat): keep chats that failed to resume in the status bar and say what to do (#22448)
* fix(native-chat): keep chats that failed to resume in the status bar and say what to do After a restart, a chat whose resume did not carry on was reported only by a four-second toast that named nothing, and the status bar entry vanished because the reattach had already spent the offer. The host now files the outcome as a durable `failed` entry in the recovery capsule, with the refusal code and the prompt the offer quoted, and returns it from the restart-resume RPCs. The renderer shows a "N chats failed to resume" status bar entry, a count-only toast with Show and Dismiss, and keeps the resume dialog open with a status icon per row and a "To resume" line whose action is chosen from the reason. A failure dies on dismiss, on a successful retry, on the user's own send in that chat, or with the marker's 24h expiry. * fix(native-chat): keep the resume dialog unchanged and add failed chats as rows The failure view had replaced the resume dialog's title, checkboxes, preference box, and footer. The dialog is back to main's layout. A chat an earlier resume could not carry on is now an ordinary selectable row there, plus a status icon whose tooltip carries the reason, a dismiss control, and a "To resume" line. Selecting it and pressing Resume retries it; it is pre-selected only when a retry can succeed. Resumed chats leave the list as before. Also stubs the new failure listing on the cross-version wire host fixture, which the restart-resume RPC now reaches. * refactor(native-chat): release a failed-resume record where a send is admitted Keeps the host file at main's size, and only releases the record for a send the controller actually lets through. * fix(native-chat): note a failed restart resume in the chat and derive when it is settled The chat itself now says when Orca could not continue it after a restart, with an error (refused) or warning (unconfirmed) status row, so the failure survives the toast, a dismissed record, and another restart. A recorded failure is current only while the chat's newest user message is the one it had when the failure was filed. Listing re-derives that from the journal and prunes superseded records, replacing the in-memory set and the hook on every send. Failures move to their own optional top-level key in the recovery file, so an older build that rejects unknown entry states keeps reading its offers. A retried failure stays a failure through a rollback or a lapsed lease instead of returning as a pending offer, and the toast's Dismiss names only the chats the host listed as failed. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): check the older reader against a filed failure before any retry Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): file a failed restart resume against the chat as its attempt ended A failure was filed against the chat's newest user message read at settlement, after every chat in the action had finished. The chat's note asks the user to send a message, and one sent while other chats were still being continued became part of the filed state, so the failure stayed listed after the user had done what it asked. Each chat's newest message is now observed as its own attempt ends. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): let a reply that made a chat ineligible retire its failure When a restart resume reached a chat the user had already replied in, the attempt was refused as no longer eligible, but the failure was filed against that very reply. It then stayed listed as "finished on its own" until the user sent yet another message. An ineligible chat is no longer observed at the attempt, so its failure falls back to the reserved marker and the reply that made it ineligible supersedes it at the next listing. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): offer Retry on a failed resume only when a retry would run When the agent refused Orca's "continue" message with its own reason, the failed row fell to the generic guidance, which offers Retry and pre-selects the chat in the resume dialog. The refused message is already the chat's newest user message, so a retry is never eligible: it did nothing and the same "couldn't be resumed" toast came back. The host now reports whether a retry would run, derived at list time from the same predicate the retry applies to the failure's marker. Where it would not, the row offers Open chat and Dismiss and is not pre-selected. An older host omits the flag and the reason alone decides, as before. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): file only restart failures the user must act on A chat that stopped being resumable between listing and acting (it finished on its own or is waiting on the user) was filed as a failure with no note in the chat. It now just spends the offer. A failed reattach now writes the same in-chat note as a refused continuation, so every filed failure explains itself in the chat. An unconfirmed continuation's failure retires once the chat shows the continuation's own message opened the newest turn. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): clear a failed resume from the status bar once its chat moves on While a failed resume is listed, the restart store watches the host's status feed; when a failed chat's status or latest prompt changes after the list was read, it re-reads the host once. The host still decides whether the failure stands. Nothing is watched while nothing failed. The failure toast now counts only the requested chats the host still lists as failed, keeping the old count for a host that sends no list. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that an unjournaled continuation never retires its unconfirmed failure Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop telling the user to send a message into a chat Orca couldn't reconnect A failed reattach, or a continuation refused because another window or terminal owns the session, now leaves a note saying Orca couldn't reconnect the chat instead of advising a send that would meet the same refusal. The restart list keeps the reason-specific advice. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the failed-chat re-read from undoing an action or missing a reply The status-bar re-read no longer runs while a resume or dismiss is in flight, and its answer is dropped if one settled meanwhile, so a dismissed failure cannot come back. A change to a failed chat already seen always triggers it, whatever the host's timestamp says. A chat whose unconfirmed continuation the host already retired is now reported as resumed instead of saying nothing. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that no failed-chat re-read runs under a resume in flight Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): typecheck the unjournaled-continuation case against a nullable marker Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): drop the stale "no arguments" note on the restart-offer params The dismiss call now names sessions, so the older comment contradicted the schema below it. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): count unconfirmed resumes apart from refused ones in the action toast The post-action toast said "N chats couldn't be resumed" for chats whose continuation may well have gone out, while the list and the chat itself say Orca couldn't confirm it. That wording invites a duplicate "continue" send. Unconfirmed chats now get their own count, classified by the outcome the host filed, so the toast matches the row it points to. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop offering Resume on a failure the host says cannot retry A failed row the host marks unretryable could still be ticked, sending a resume that could only fail again; its checkbox is now disabled and it never joins the action. An older host that omits the flag keeps today's selectable row. The status bar no longer calls a chat "failed to resume" when the host only couldn't confirm the resume, matching the dialog's own wording, and the mixed toast's second line now says "other" so it cannot read as the same chat. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): skip unreadable failure records instead of rejecting the capsule A failure record this build cannot parse (a newer outcome, say) made the whole recovery file unreadable, so a downgraded build listed no restart offers and could not record new teardowns. Failures are advisory: parse each one on its own and drop what does not parse. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a failure record with an unreadable marker is skipped The skip-unreadable-failure test only covered an unknown outcome, so going back to the throwing marker parser for failure records still passed. A failure record usually outlives its offer entry, so a newer marker shape can appear only there. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
8ba7f829ac |
feat(ipynb): run notebook cells in a persistent Jupyter kernel (#22581)
* feat(ipynb): run notebook cells in a persistent Jupyter kernel Replaces the fresh-process runner (which silently re-ran every earlier cell) with the user's own ipykernel, driven by a small bundled Python bridge over line-framed JSON. One kernel per open notebook: started on first Run, shut down when its tab closes or Orca exits (stdin EOF), and the kernel's own parent poller reaps it if the bridge dies. The header gains a kernel pill (workspace .venv/.conda recommended, PATH interpreters, Browse), Interrupt/Restart/Run all/Clear all, and a one-time missing-ipykernel dialog with Install. Outputs stream live per cell and are written into the document when the run finishes. * test(ipynb): cover the stalled-interrupt restart offer * refactor(ipynb): disable the kernel pill while settling; merge its classes with cn * refactor(ipynb): tie kernels to their renderer document and simplify the run flow - Main keys kernels per renderer, so a reload, renderer crash or closed window shuts them down, and two windows never share one notebook kernel. One close-driven cleanup replaces the separate exit and start-failure deletions. - The first run uses the nearest workspace env, else the first Python on PATH; the picker no longer opens itself, so its open state stays in the toolbar. Closing the picker brings the missing-ipykernel dialog back instead of dropping the queue, which also keeps a Browse pick's run. - Discovery marks the kernel starting, so a second run during it queues instead of starting a second kernel, and a tab closed mid-discovery no longer leaks one. - Running an nbformat 4.4 notebook gives its cells ids (upgrading to 4.5), so moving a cell mid-run cannot misroute its output. - The death notice drops stderr from before the kernel was ready (the unencrypted-TCP warning). - SSH and non-Python runs toast instead of writing a notice into the cell. - Windows conda envs are named after their folder. * fix(ipynb): install ipykernel into envs without pip uv-created venvs ship without pip, so Install failed with 'No module named pip' there. When pip is missing, bootstrap it with the stdlib's ensurepip and retry. The install moves beside the other interpreter probes, and the copyable install command comes from one helper. * fix(ipynb): address PR review comments on stream errors, old jupyter_client and the Windows venv hint - Swallow stdout/stderr stream errors on the bridge child, as spawnProcess requires, so a broken pipe cannot crash main. - The bridge exits (reporting the death) even when cleanup_resources is missing (jupyter_client < 6.1.5) or raises. - The install-failure hint suggests `py -m venv .venv` on Windows. * feat(ipynb): add Cancel to the missing-ipykernel dialog It does what Esc does: drops the cells waiting on the kernel. * fix(ipynb): recover from a rejected kernel start; quote the install command per shell - Discovery moves into start, so one catch turns a rejected listPythonEnvironments or startKernel into the usual failed start: the session returns to off with the error in the cell, instead of sticking at starting. - The copyable ipykernel command quotes the interpreter only when its path has whitespace, prefixing PowerShell's call operator on Windows. Install itself still spawns without a shell. * fix(ipynb): always shell-quote the copyable ipykernel install command Quote the interpreter path for every path, not only ones with whitespace, so paths with shell metacharacters like & copy as a working command. Single quotes are literal in POSIX shells and PowerShell; embedded quotes are escaped per shell, and PowerShell keeps its & call operator. |
||
|
|
795b64b9a6 |
docs(tui-agent-config): correct the OpenCode readiness-budget rationale (#22593)
The comment merged with #22546 claimed ConPTY never forwards DECSET 2004 and that the signal therefore cannot fire on Windows. Verification on two real Windows hosts refuted that: the sequence arrives in order on both ConPTY backends, and the readiness signal fired in every run. The budget was the actual problem — opencode does not enable bracketed paste until ~4.8s and its composer is not ready until ~10s, so the 8s default expired first and the draft was pasted blind. Same fix, accurate reason. Note the claim that seeded this: terminal-agent-paste-bracketing.ts says 2004 "can be lost by remote replay or ConPTY", which is careful and not contradicted here; the absolutism was mine. |
||
|
|
5e3effc32f |
fix(native-chat): show every user message on the message rail, not just loaded ones (#22558)
* fix(native-chat): show every user message on the message rail, not just loaded ones The message rail was built only from transcript rows the renderer had loaded, so any prompt above the loaded page had no tick, and the rail lost ticks when a long live session trimmed its retained window. The host now answers `agentSession.conversationOutline`: every user message in a structured session's journal (item id, creation sequence, a preview cut to 200 characters, image count) plus the journal position it is current through. It is derived from the reduced journal on each request with the same projection the transcript runs, so an entry's id and preview are what the loaded row shows. The reply is bounded like a history page: previews shorten, then drop, and only then do the oldest entries go. The renderer asks only while the pane is visible and older history is unloaded, uses outline entries only for messages older than its loaded window (the window is authoritative for the rest), and falls back to loaded messages while the outline is stale (epoch change, or the window trimmed past what it covers). Selecting a tick with no row pages older history in until the row exists, then uses the existing rail jump. The method is negotiated with `agent-session.conversation-outline.v1`; a client never calls a host that does not advertise it, and any failure leaves the rail on loaded messages. * fix(native-chat): keep a rail jump from the bottom from re-arming follow and cancelling itself A rail jump started by a reader following the end stopped a few pixels above the bottom instead of reaching the message. Paging older history in for a jump always leaves the reader following at the very end, so jumps to unloaded messages hit it every time; a jump to a loaded message from the end did too. The jump scrolls smoothly, and only its landing is marked as the application's own scroll. Its first frames move a pixel or two, still inside the band where a reader event re-arms follow, so the list read the jump leaving the end as the reader arriving at it. The next frame, just outside the band, then read as the reader taking over and rebased the view with an instant write, which cancels the smooth scroll. Re-arming follow now needs the reader to be arriving at the end: an unmarked event that moved the view up never reattaches a detached reader. * test(native-chat): check a rail jump left the end before reading where it landed * fix(native-chat): keep the rail's message list still while a press selects an item With the whole conversation in the rail, its hover list overflows and opens scrolled to the message being read. Clicking an older item did nothing: pressing it focuses it, which turns the hover preview interactive, and that switch re-ran the effect that scrolls the lit row into view and focuses it. The list moved under the pointer between press and release, so the click landed on the list instead of the item, and focus jumped to the lit row. Revealing the lit row now follows the list opening (and its rows shifting), not the switch between hover and interactive. Entering interactive moves focus into the list only when focus is not already on one of its items. * perf(native-chat): keep the rail's outline entries stable while a trimmed window slides A long live session holds a head-trimmed window, so every new row moved the oldest-loaded edge and rebuilt the outline view even when no user message crossed it. The rail then re-merged, re-rendered and re-read the scroll geometry on each new row. The view is now reused while the set of entries older than the edge is unchanged. * fix(native-chat): let a rail jump wait out an older page already loading Scrolling to the top of the loaded window asks for the next older page. A rail jump made while that page was in flight asked again, got the lane's immediate no-op return, read it as a page with no progress, and dropped the click. The jump now waits for the in-flight page to land before deciding. * fix(native-chat): reattach follow when content shrinking clamps a reader onto the end The rule that stops a smooth scroll leaving the end from re-arming follow compared offsets, so it also refused a reader whose offset dropped because settled content folded away beneath them and the browser clamped them onto the end. They sat at the bottom without following, and the next reply grew out of view. Re-arming now requires closing on the end rather than moving down, which still rejects a scroll leaving it. * fix(native-chat): keep the rail hooks' ref writes out of render Both hooks wrote a ref while rendering, which React may replay or discard. The rail's structural-sharing baseline is now recorded after commit, and the history jump calls the lane's page loader from its effect instead of through a render-updated ref. * fix(native-chat): let the latest rail pick win over a jump still paging A jump to an unloaded message keeps paging older history until it lands. A pick made meanwhile lost to it: a loaded message scrolled into view, then the earlier jump finished and pulled the reader away; another unloaded message was ignored. Picking a loaded message now cancels the paging jump, and picking an unloaded one retargets it without asking for a second page. * fix(native-chat): step a rail history jump with a functional update The step that requests the next page wrote the pending jump from the value its effect closed over, so a pick or cancel queued since that commit would be overwritten. * refactor(native-chat): run a rail history jump as one abortable awaited loop The jump through unloaded history was an effect-driven state machine that guessed "no progress" from the message list's identity and could not be cancelled by anything but another rail pick. A diff reveal, "Jump to latest" or the reader scrolling left it paging, and when its page landed it pulled the reader away; a history read that kept failing during a live turn could repeat back to back. Loading an older page now reports how it ended, and a second request while a page is in flight joins it instead of being refused. The jump is an awaited loop that reads the rail from a commit after each page, stops on anything but a page that moved the window, and is aborted by any other navigation, reader input (wheel, touch, scroll keys, scrollbar), a session switch or unmount. The latest pick wins. * fix(native-chat): derive the rail outline from the transcript's own projection The host built the outline from user items alone, while the transcript orders every message by when it was observed, folds tool results into the turn above and then drops harness turns. An imported user row carrying a tool result beside harness text therefore got a rail tick previewing the harness text, and clicking it paged history for a row that never draws. The transcript's order-fold-strip projection now lives in one shared function. The renderer's list projection wraps it with its own tail-row order, and the host runs it over the whole journal and keeps the user rows that draw content, so the outline lists the same messages in the same order. * fix(native-chat): retry a failed rail outline read a few times A failed outline read left the rail on loaded messages until a new gap opened, the pane was shown again or the epoch changed. The client now rejects a failed read (a host without the outline still resolves to nothing, without being called), and the rail retries up to three times with doubling backoff. * test(native-chat): give the rendered-transcript fixture the older-page result contract * perf(native-chat): sort the shared transcript projection without a spread copy * fix(native-chat): let a wheel over the rail cancel a rail jump still paging The rail forwards its wheel to the transcript, so a reader scrolling there is scrolling the transcript. That wheel never reached the scroller's reader-input handlers, so the jump kept paging and later pulled the reader to its target. * fix(native-chat): keep the rail's list open while a picked message pages in Picking a message that is not loaded yet can take several pages of older history. The list closed on the pick, so its busy item was never seen and the click looked ignored. The list now stays open with that item pulsing until the jump lands or is abandoned, and stops revealing the lit row meanwhile so the rows do not move under the pointer. * fix(native-chat): keep the shared transcript projection loadable on mobile The projection moved to src/shared, which mobile's Hermes engine also loads, and switching its sort to toSorted broke the Hermes compatibility guard. Sort a copy made with Array.from instead, as other shared code does. |
||
|
|
5802b54579 |
fix(rate-limits): read OpenCode Go usage with the Go API key (#22551)
* fix(rate-limits): read OpenCode Go usage with the account API key Since OpenCode's console migration (upstream fe51b0b19a, "fix(console): restrict legacy access to Black"), an account with no Black subscription is redirected from the legacy console to /console/login, so Orca's cookie-based workspace lookup returns nothing and the Go bar stays empty. Fetch usage from GET https://opencode.ai/zen/go/v1/usage instead, which authenticates with `Authorization: Bearer <key>` and needs no console session. The key resolves in order: Orca settings override, OPENCODE_API_KEY, then whatever OpenCode itself stored on /connect -- auth.json for 1.x, the credential table for 2.x. The cookie path stays as the fallback so Black/legacy accounts keep working. A 403 EntitlementError now reads as "no OpenCode Go subscription" in the status bar instead of a generic refresh failure (#22257's reporter was misled by exactly that). * fix(rate-limits): prefer OpenCode's stored Go key over OPENCODE_API_KEY OpenCode applies the key saved on /connect after the environment, so the stored key is the one its own Go requests use. OPENCODE_API_KEY is also the Zen provider's variable, so ranking it first could read a key that OpenCode itself is not using for Go. Co-Authored-By: Claude <noreply@anthropic.com> * fix(rate-limits): name the API key when OpenCode Go usage lands on sign-in A redirected usage request arrives as a 200 sign-in page because Electron follows redirects; report it as a rejected key instead of a parse failure. The cookie path's empty workspace lookup is what non-Black accounts now hit after the console migration, so its message points at the API key rather than only the workspace override. Co-Authored-By: Claude <noreply@anthropic.com> * chore(i18n): add the OpenCode Go API key strings to the English catalog Co-Authored-By: Claude <noreply@anthropic.com> * docs(rate-limits): stop calling the credential table an OpenCode 2 marker Verified on two real Windows hosts running OpenCode 1.18.16: the `credential` table exists there too (empty, same columns), so its presence does not identify a 2.x install. Neither host had an `auth.json` at all. The resolution already probes both stores on every version, so only the comments were wrong. Says so now, and records that a 2.x install which never ran the legacy import has no `auth.json` either — which is why both tiers exist. * refactor(shared): move GhosttyImportPreview out of global-settings-types Adding `opencodeGoApiKey` pushed global-settings-types.ts one line past the 300-line ceiling, failing `oxlint` in CI. AGENTS.md forbids a max-lines suppression, so split instead: the Ghostty import preview is a distinct concern that never belonged in the settings-shape file. Re-exported from the original module so no importer changes. 293 code lines now. --------- Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
b0ae7d18a0 |
fix(opencode2): resolve subagent session lineage so child work stops taking over the pane (#22444)
OpenCode 2's plugin adapter unwraps a single-property `{ data }` success schema,
so `ctx.session.get` resolves to the bare session record. The shared lineage
lookup only accepts `result?.data?.id === sessionID`, and OpenCode 2 has no
`session.list` fallback, so `resolveRootSessionID` returned null for every
session and `childState` was permanently null.
With unknown lineage `canFailOpen` is true for attention events, so a subagent's
`permission.asked`/`question.asked` fell through and pinned an un-evictable
blocker keyed to the child's own session id — publishing a subagent as if it
were a root. Observed in hook posts: SessionBusy for a child session id whose
`session_v2` row carries a parent.
Envelope the result in the OC2 client shim so the shared lineage module works
unchanged; OpenCode 1 already receives enveloped results and is untouched.
Also adds `opencode2` to the double-Escape interrupt list, extracted into one
shared helper so the server inference and renderer gate cannot drift. A single
Escape was inferring an interrupt, and Escape is how the Subagents dock closes.
7 of 11 new lineage tests fail without the shim.
|
||
|
|
b7a4fee700 |
fix(agents): stop claiming an unconfirmed OpenCode handoff succeeded (#22546)
* fix(opencode): stop claiming a handoff prompt was delivered when it was written blind "Continue in New Session…" to OpenCode reported success even when the prompt never reached the TUI (#22479). The paste-after-ready helper falls back to a blind write when the composer-ready signal never arrives and only the agent process is known to exist; that write was indistinguishable from a real delivery, so the continuation showed its success toast. - pasteDraftWhenAgentReady / pasteDraftToAgentPtyWhenReady report the blind fallback via onUnconfirmedDelivery, plumbed to launchAgentInNewTab as onPromptDeliveryUnconfirmed. - The session continuation hedges instead of claiming success, and both the failure and the hedged notice offer "Copy prompt". - OpenCode gets Codex's 20s composer budget. Both are quiet-window-less signals anchored on DECSET 2004, which ConPTY never forwards, so on Windows that budget is the settle delay before the blind paste. * test(runtime): retarget the 8s startup budget test off OpenCode The main-runtime startup-draft budget test used `opencode` as its stand-in for "an agent without an override", which this branch invalidates by giving OpenCode 20s. It failed with "expected vi.fn() to not be called at all, but actually been called 1 times" — the readiness signal now legitimately arrives inside budget. Point it at `claude`, which still takes the 8s default, and add a companion pinning OpenCode's 20s: a readiness signal at t+10s, past the old default, must now deliver the draft. Removing the override makes that companion fail. |
||
|
|
80f5aae0f9 |
feat(agent-status): publish the main agent's own state beside the combined row state (#22452)
* feat(agent-status): publish the lead agent's own state beside the combined row state
Every status producer folded the main agent's state together with live child
work into one `state`, so a lead that had finished while a subagent still ran
read `working` and its own state was lost. The row now also carries
`lead: { state, outcome?, stateStartedAt }`, admitted by the one payload
normalizer on the relay wire, IPC and disk, and published from the Claude hook
lane, the structured host ingest and renderer bridge, Grok (now on the shared
fold) and Codex (own combine kept). The persisted child-only boundary flag is
derived from `lead` plus child evidence and no longer written; old rows map
onto `lead` at hydrate. Combined `state` and `workingMode` are unchanged for
every reader; a cross-lane parity table pins that, with the cancelled-turn
watch-loop story recorded as a known divergence.
* fix(agent-status): make Orca's inferred interrupt the primary source of a Claude lead cancellation
Current Claude Code sends no hook at all on a cancel and no is_interrupt on
Stop, so the cancellation enters the lead record from the server's inferred
interrupt and rides into the next real Stop; is_interrupt on a turn boundary
stays as the secondary source for builds that send it. Comments, the store
reference and the parity table say so; no suppression changes.
* docs(agent-status): the child-only boundary comment now describes the persisted shell fact
The old sentence said a hydrated row no longer carries the shell fact, which is
the opposite of the mechanism: claudeRunningNonAgentTask is persisted precisely
so hydration can read it, and only a pre-lead row lacks it — reading as
shell-free, the same assertion its legacy flag made at write time.
* rename the lead fact to mainAgent: the main agent's own state
* docs(agent-status): the inferred cancel comes from Ctrl+C, not Esc
* fix(agent-status): an inferred interrupt keeps an already settled main agent, and the row verdict docs name its inferred source
* fix(agent-status): a child-induced wait publishes the main agent state it displaced
* fix(agent-status): decide child-held Claude rows from the saved main agent fact
Restart seeds the Claude main agent from the row's saved mainAgent whenever it
settled and no shell held the row, instead of re-deriving a child-only shape.
OSC cannot settle or repaint a row child agents hold open, including a row
waiting on a child's permission prompt. A sticky child permission prompt still
records the main agent's own progress, and OSC repaints and inferred answers
keep the shell fact beside the main agent they preserve.
* fix(agent-status): keep a finished turn's main agent verdict and clock with that turn
A Claude SessionStart restarts the main agent's clock instead of inheriting the
previous session's last Stop. A Grok idle prompt or session end, and a late
Codex root Stop after an inferred cancel, restate the same finished turn, so
they keep its recorded verdict; only a new turn clears it.
* test(agent-status): publish the Grok verdict restatement past the late-event window
* docs(agent-status): describe hydrate seeding and the OSC refusal from the saved main agent fact
* fix(agent-status): push a held child permission row when its main agent changes
* fix(agent-status): keep the shell fact on a held child permission row so restart does not settle it
* docs(agent-status): note the held child permission row carries the shell fact and is pushed
* fix(agent-status): pair the Claude shell fact with the main agent at the one row-build point
Every non-hook rewrite (terminal-title repaint, inferred answer, held child
permission) had to re-carry the shell fact beside `mainAgent`, and each one that
forgot let a restart settle a row while a shell still ran. The row builder now
pairs the fact once: a listener event restates it, any other write keeps it only
while `mainAgent` is unchanged. Restart seeds a settled main agent only when the
row says no shell ran, and legacy child-only rows map to that explicitly.
A held child permission now also accepts the main agent event's background
evidence, as it already accepts its `mainAgent`, so the child's drain no longer
settles a row a shell still holds. The renderer keeps a previous `mainAgent`
only for writers that never carry one, so a hook row without it matches the
host snapshot.
* test(agent-status): pin that restart never seeds a main agent from a row silent about its shell
* docs(agent-status): the row builder pairs the shell fact with the main agent, and restart seeds only on an explicit no-shell
* test(agent-status): name the legacy-row case parameter for what it holds
* docs(agent-status): name which rows carry the main agent fact
|
||
|
|
800d33e5c9 |
feat: name runtime machines (#22094)
* feat: name runtime machines
* fix: preserve pairing address optionality
* fix(cli): keep host and environment listings local
Listing paired servers read each one's machine name by dialing it, so both listings made a network
round trip per server and waited out a timeout on any that were offline. They answer from this
machine's own pairing store; `orca host name --environment <name>` reads one server's name.
* fix(settings): caption the machine name paired devices actually receive
The caption read the runtime's published name once, when the pane opened, so saving an override
left it naming the old computer while phones already showed the new one. It now re-reads whenever
the saved override changes; the settings write lands in the main process before the store publishes
it, so that read already sees the new name. The name is interpolated rather than baked into the
fallback, and the caption, label and placeholder are in the English catalog.
* refactor(settings): normalize the machine name in one place
The trim and length rules for `machineName` were spelled out separately at the
renderer IPC (trim + 255), the settings load path (trim only, no cap), the RPC
schema (zod trim + 255) and the runtime reader (trim). A hand-edited or legacy
profile could therefore load a longer name than any writer accepts.
`src/shared/machine-name.ts` now owns `MACHINE_NAME_MAX_LENGTH` and
`normalizeMachineName`, and every writer and the load path use it. The RPC
schema keeps rejecting over-long names but derives its cap from the constant,
and the runtime settings controller normalizes an RPC write before storing it.
* fix(runtime): detect the machine name once and label handoffs with it
Every runtime constructed in a process (the app, plus each one a test builds)
ran its own `scutil` lookup. The friendly name is a property of the host, so the
lookup is now a single shared promise; construction still never blocks on it,
and a rejected lookup can no longer surface as an unhandled rejection.
The structured-chat handoff banner ("Agent is open in terminal on X") named this
host with the bare `os.hostname()` while paired devices saw the published name.
The transport now reads the same `RuntimeMachineName`, through a getter so a
rename in Settings is reflected without rebuilding the transport.
* fix(cli): print the name the runtime publishes and keep its envelope
`orca host name --name X` printed `undefined`: `settings.update` replies with
`{ settings }`, but the handler read a bare `machineName` off the reply, and the
test fixture mirrored the wrong shape so it passed. After a write the command
now re-reads `status.get` and prints what the runtime publishes, so a blank
`--name` prints the detected name it returned to rather than an empty string.
The read path wrapped a possibly routed answer in a local envelope, stamping
`_meta.runtimeId: "local"` on a reply from another server. It now returns the
`status.get` envelope itself, and an unreachable runtime is reported as the
usual error instead of an invented "unknown" name.
`environment list` had gained machine-name and platform columns that no caller
populated, so every row printed "platform unknown"; the columns are removed.
* refactor(settings): give the machine name field its own component
The caption under the field re-read runtime status every time the saved value
changed, relying on a comment about write ordering to show the new name. A saved
override already is what paired devices see, so the hook now derives the caption
from it and asks the runtime only for the detected name; a stale status read can
no longer show the previous name.
`MobileMachineNameField` owns the store read, the published-name hook and the
debounced input, so `MobilePairingSetupSection` returns to its prop shape and
the pass-through `MobilePanePairingOutput` wrapper is gone. Paired-device
revocation moves into `useMobilePairedDeviceRevocation`, which keeps
`MobilePane` within its line budget with an extraction that carries behavior.
The web client mounts this pane too, but its settings store kept the name
locally where nothing published it. `machineName` now rides the existing
runtime-backed settings sync so the field renames the paired runtime.
* refactor(settings): normalize the machine name at the store boundary
Every writer (desktop IPC, web RPC, CLI) reaches the store through
updateSettings, which already normalizes the other free-text settings
there. Trim and bound the machine name in that one place instead of at
two upstream edges, so a future main-process writer is covered too.
* test(settings): pin machine-name routing and detection, and make the field searchable
The shared machine-name lookup test spawned the real `scutil` twice and compared the answers, so a
slow runner could time one spawn out to the hostname and fail. It now mocks the subprocess, proves
the hostname answers until the one shared lookup lands, and that a second runtime does not spawn
again.
`host name` is no longer pinned local, but only the explicit `--environment` route was covered; an
ambient `ORCA_ENVIRONMENT` now has its own test so the pin cannot silently grow back.
The Machine name field is added to the Mobile pane's search catalog at the tail, keeping every
existing row's tie-break index.
* fix(runtime): wait for the machine-name lookup before publishing status
A status read answered in the first few milliseconds after launch published the bare
hostname because the friendly-name lookup had not landed yet, and a caption fetched in
that window never corrected itself. RuntimeMachineName now exposes the settled lookup
as a promise, and both status publishers (the status.get RPC and the desktop
runtime:getStatus IPC) await it before reading. Construction, listen, and every other
method stay unblocked; the worst case is one wait of at most a second on the first read.
* fix(cli): refuse to rename a runtime that does not publish a machine name
An older Orca runtime rejects the unknown settings field with a bare invalid_params, so
'orca host name --name' routed at one failed with no explanation. The runtime that does
not publish machineName on status cannot store one either, so the CLI reads status first
and refuses with incompatible_runtime and a message that says to update that host,
before writing anything.
* fix(ipc): introduce this desktop to remote hosts by its machine name
When this desktop connected to a remote workspace host it announced itself under a
hostname captured once at module load, so a renamed machine kept its old name on every
other device's connected-clients list. The client name is now read at send time from
the runtime's machine name (the configured override, else the detected one), passed in
where the remote workspace handlers are registered, so a rename reaches the next
presence frame without a relaunch.
* fix(runtime): keep the machine-name lookup under the status probe budget
Status publishers now wait for the one-time name lookup, and `orca status`
probes them with a one-second budget. scutil answers in milliseconds, so a
half-second cap keeps a stalled lookup from making a healthy runtime read as
"starting" while still preferring the friendly name.
* refactor(web): drop the unreachable machine-name write path
The Mobile settings section is desktop-only, so the paired web client can
never render the field. Forwarding the name through the web settings sync was
dead code, and against an older host the strict update contract would have
rejected it while the local mirror kept the value. Remove it until a web
surface exists.
* chore(i18n): translate the machine-name strings and document paired-server rows
Add the Machine name field and its Settings search entry to the five non-English
catalogs, explain in the host list spec why paired-server rows report an unknown
platform, and drop a stale timeout figure from a test comment.
* refactor(settings): make the machine name a machine-wide setting with a General home
The name other devices and hosts list this computer under is not a mobile
setting. Rename MobileMachineNameField to MachineNameField, give it a per-mount
id, and put its primary home in Settings > General under "This computer". The
Mobile pane keeps the same field. One shared search entry feeds General, the
Mobile pane, and the copy now says "other devices and hosts" in all six locales.
The web client has no machine of its own to name and its settings mirror cannot
persist one, so the field renders nothing there and General omits the section.
* feat(mobile): name this computer in the Orca Mobile pairing step
The "Pair this computer" step now shows the same machine name field above the
connection choice and code, so a user pairing a phone from the sidebar page can
name the computer right there.
* feat(settings): name this host when sharing it with other devices
Share this host produces the access link other devices use to reach this
machine, so it mounts the machine name field first. The pane's search entry
takes the shared machine-name keywords so a search lands there.
* feat(sidebar): name this desktop when adding a remote host
This desktop introduces itself to a new SSH host or remote server under its
machine name, so the Add Remote Host dialog mounts the field once, between the
header and the host fields, in both modes. Submit logic is unchanged.
* feat(settings): name this computer in the SSH pane add form
The SSH pane's add form mounts the machine name field above the host fields.
Editing a saved host leaves it out; that host already met this computer.
* fix(mobile): drop the empty machine-name grid row on the web client
The pairing step wrapped MachineNameField in its own grid-area div. On the
web client the field renders nothing, so the wrapper left an empty row and
an extra row gap between the copy and the connection options. The field now
takes a className for its root, so the grid slot disappears with it.
* fix(settings): let Enter in the machine name field submit its form like sibling inputs
The field intercepted Enter to blur and commit instead of submitting the enclosing
SSH add form. The draft is already flushed on blur and on unmount, and the name is
read from the store whenever a peer asks, so nothing is lost when the form submits
first. Enter now behaves like the neighbouring inputs; the test proves the submit
fires and the name still commits when the form closes.
* fix(mobile): keep the machine name inside the pairing copy cell
A dedicated grid row stayed in the template on the web client, where the field
renders nothing, adding an empty track and a second row gap between the copy and
the connection options. The field now sits at the end of the copy cell with the
same 18px rhythm, so an absent field leaves nothing behind.
* fix(runtime): retry a failed machine-name lookup instead of latching the hostname
On a loaded Mac the scutil lookup missed its 500 ms cap during app boot, and
because the fallback was memoized for the process, every status read and the
Settings caption showed the bare hostname for the rest of the session.
The lookup now gets a 5 s timeout, a failed attempt (timeout, spawn error,
non-zero exit, empty output) clears the shared memo so a later ready() retries
after a 30 s interval, and status publishers wait only up to a 750 ms publish
budget before answering with what read() has now. A friendly name and the
non-darwin hostname stay final.
* refactor(settings): show the machine name only where other devices join this computer
The Add Remote Host dialog, the SSH pane add form, and General all describe
another machine, so a field about this computer's own name read as a third
kind of label there. The field now mounts only where other devices pair with
or connect to this computer: the Mobile pane, the Orca Mobile pairing step,
and Remote Servers > Share this host.
|
||
|
|
98e5ea3d5f |
refactor(tabs): delete the terminal tab's dead adopted-session field (#22557)
* refactor(tabs): delete the terminal tab's dead adopted-session field and every branch that read it * refactor(tabs): drop the stale adopted-agent comment and pin legacy load on a chat terminal The comment above the terminal chat-eligibility agent fallback described the removed adopted-session agent fallback. The legacy load test now puts the retired key on a chat-mode terminal tab, the only shape that ever carried it. |
||
|
|
069dc8a1d8 |
feat(agent-launch): let a caller reserve the chat session, and start terminal launches with the session picks (#22523)
* feat(agent-launch): let a caller reserve the chat session and carry session picks to a terminal launch * fix(agent-launch): keep a caller-minted session id named for its agent, and mint the fallback the same way * test(mobile): model the older host from the launch fields, not the refined schema * docs(agent-launch): describe the reserved session id as conversation identity, not placement The caller mints the session id so it knows which conversation it started; tab placement is not keyed on it. Also puts the terminal surface's doc comment back on createTerminalSurface. * docs(agent-launch): say a terminal launch reads the session picks on the wire contract The `sessionOptions` field doc still said a terminal launch ignores them, which this branch changed. * fix(agent-launch): check a reserved session id's token after the agent name, not the whole id A hyphenated agent name failed the one-token check, so any session id for such an agent was refused at the wire, while every other agent without a chat has its id ignored on the terminal. |
||
|
|
60bd1dfdea |
feat(native-chat): one shell-environment setting for every structured chat (#22387)
* feat(native-chat): one shell-environment setting for every structured chat Structured Codex chats started from the login-shell environment, while structured Claude chats started from Orca's own process environment, so a variable exported in .zshrc reached one and not the other. Both now start from the same base, chosen by a new setting: - on (default): the whole login-shell environment, as a terminal gets - off: Orca's environment plus PATH, locale, SSH_AUTH_SOCK, and the variable names the user lists The setting is re-read each time a chat starts or resumes. It is shown only when Chat UI, the Chat UI default view, and structured native chat are all on. Terminal-backed chat is unchanged. * fix(native-chat): normalize the shell-environment settings when a profile loads A hand-edited settings file could store the variable list as something other than an array, and the structured runtime called `.filter` on it per launch, so a malformed value failed every structured chat create and resume, and the settings pane render. Normalize both keys where the profile loads, the same way the other array settings are, through one shared normalizer the runtime policy also uses. Also pin that an uncommitted name draft survives an unrelated settings re-render. * fix(native-chat): keep the pinned account as the only source of a structured chat's Claude home The session record owns which Claude home a structured chat uses, and the acquisition pin (claudeConfigDirEnvPatch) is the only emitter of CLAUDE_CONFIG_DIR, compared against what the child would otherwise inherit. With the login-shell snapshot as the inherited base, a CLAUDE_CONFIG_DIR exported only in a shell rc flipped that comparison and produced an explicit pin to the CLI default home, which moves the CLI off its default Keychain item. Drop the inherited CLAUDE_CONFIG_DIR in the Claude launch resolver before the pin runs, as Codex already does for an inherited CODEX_HOME. A configured per-agent overlay still passes through, since the record already honors it. * fix(native-chat): drop Orca's own CLAUDE_CONFIG_DIR from a structured Claude child too The process spawner merges Orca's process env under the launch env, so a CLAUDE_CONFIG_DIR exported to Orca itself reached the child around the launch resolver's drop and unseen by the account pin. One helper now strips it from both inherited bases. Also declare the two shell-environment settings on the runtime store contract and add the six new strings to every locale catalog. * feat(native-chat): add shell variables one at a time with a removable list * fix(native-chat): return focus to the name input after removing a shell variable * fix(native-chat): use a neutral placeholder for the shell variable input The empty input showed a grey HTTPS_PROXY as its placeholder, which reads as a saved value, especially right after that exact entry is removed from the list. Use "Variable name" instead, in every locale catalog. |
||
|
|
845db9e5e2 |
fix(native-chat): underline only file links a click can act on (#22370)
* fix(native-chat): underline only file links a click can act on A chat message could underline a bare file name such as `deck.md` that resolved nowhere, and clicking it did nothing, so it read as a broken link. - Inline code and quoted text become file links only when they name a path (contain a `/` or `\`), matching plain prose; a bare file name stays plain code. - Every file link click now answers: it opens, or says the file was not found, that the host could not be checked, or that the path could not be resolved. - Explicit links like [x](README.md:5) route as files, and linked text keeps `#`, `?` and `%XX` literally instead of re-parsing them as URL syntax. * fix(native-chat): wrap the parsed file location so file URIs in chat text still open Linkified prose, quoted text and inline code wrapped their display text, which the literal wrapped-href route no longer URL-parses, so file:///... resolved as a relative path under the worktree. Wrap pathText[:line[:col]] from the parsed link instead. |
||
|
|
641a7f36d9 |
fix(native-chat): keep one live tool-run header from a call's start to the turn's end (#22432)
* fix(native-chat): keep one live tool-run header from a call's start to the turn's end
The collapsed tool run's header was two elements, one for "a call is running"
and one for "nothing is", chosen call by call. Every call start and end
remounted it, the count disappeared while a call ran and came back one
higher, and a call that finished inside a frame still bought the whole swap.
That is the 42→43 flicker in the report.
The header is now one element whose live state belongs to the turn, not to
any call: it stays live from the run's first call until the agent moves past
it (prose, a further run, or the turn's end), and settles in place. While
live the sentence speaks in the present tense and counts the call in flight
("Running 3 commands"), with the latest call's command beside it as a muted
preview; once settled it reads as before ("Ran 3 commands ✓"). The category
glyph is the run's in both states, and the completion mark only appears once
settled, so nothing pops between calls.
Which run is live is derived where the transcript is sliced into rows: the
last row that speaks or acts is the trailing one. A reasoning aside after it
leaves it live; an answer or a further run settles it.
Present-tense forms for the ten sentence categories are added to the shared
copy and the English catalog. The transcript-file lane, which renders with
the structured activity UI off, is unchanged.
* fix(native-chat): settle a run blocked on the reader, keep it live past an approval
- A run whose question is awaiting the reader's answer no longer pulses
"Reading 1 file" while the agent is blocked; it falls back to its calls.
- An approval's receipt no longer moves past the run above it, so the call
it just approved reads as running while it runs.
- The header button is the live region, so the count is announced too.
- Drop the unused live option and record from the shared English sentence;
nothing renders it yet.
* fix(native-chat): stop the settled run's check from fading in on every mount
Windowing remounts settled rows as the reader scrolls, and a restored transcript
mounts them all at once, so the fade replayed where nothing had changed. Also
pin that the live header counts the next call on the same element.
|
||
|
|
563dd5487f |
feat(native-chat): show a Codex chat's goal above the composer, and set it from goal mode (#22377)
* feat(native-chat): show a Codex chat's goal above the composer and set it from goal mode Structured Codex chat now treats the thread goal as session state: a banner above the composer shows the current goal (pursuing / paused) with clear, pause/resume and expand; /goal enters a goal mode whose send calls thread/goal/set; the objective is journaled as a user message marked as sent as a goal. The banner is derived from the journaled goal rows, which Codex's resume snapshot refreshes, so a reopened or adopted chat shows its goal. Fixes STA-8159 * fix(native-chat): replace a recorded goal by clearing first, and recover a lost goal-change response - A set while the journal records a goal (any status) clears it before setting, so the new goal starts with its own time and token counters instead of rewriting the old goal's objective in place. - The threadGoal plan answers an unknown outcome from the goal the journal records and reruns otherwise, so one request timeout no longer refuses every later Clear/Pause/Resume as unknown for the mounted session. - The goal-mode chip says "Exit goal mode"; "Clear goal" stays the banner's action on the provider goal. - A typed bare /goal on Enter enters goal mode, the same as picking it. - The renderer reads the goal off the tail of its ordered snapshot; the host's unordered map keeps the by-sequence reader. - Drop the composer's duplicate in-flight guard; the goal controller already serializes changes. - Pin that a counter-only revision reaches a subscriber's live page under its original sequence. * fix(native-chat): keep a bare /goal inside goal mode as the entrance, and pin goal delivery and serialization - A bare `/goal` submitted while already in goal mode re-enters the mode instead of setting a goal whose objective is the literal text "/goal". - The counter-only revision pin now drives the host's own event sink bound to a real journal, so it goes red when the publish after a lifecycle transition is dropped; the previous fake sink never published. - Pin that a set which threw after journaling its objective puts that objective back exactly once when the ledger reruns the same operation id. - Cover the goal controller hook: absent without host support, the loaded window wins over the host's answer, a second change while one is unsettled answers false without a request, and a refused change frees the next one. * fix(native-chat): resume a blocked or usage-limited goal, and keep goal-mode drafts honest - The goal bar offers Resume on a blocked or usage-limited goal, which the provider resumes exactly as it resumes a paused one; a goal whose token budget is spent still offers only Clear. The rule lives beside the other goal facts in shared code so every reader answers it the same way. - A `/goal <text>` typed inside goal mode sets the objective `<text>`, as it does outside goal mode, instead of a goal whose objective is the literal command. - Setting a goal is a host round trip; a draft edited while it was in flight is no longer wiped when the goal lands, matching every other host command. - Pin that a lost status-change response is read as applied only when the recorded goal is in that status, that a cleared row in the loaded window outranks the host's earlier answer, and that the PTY lane is untouched. * fix(native-chat): keep the load-older anchor on the loaded window when a live revision lands below it A live revision of a row keeps that row's original sequence. When the row is older than the client's loaded window, the shared reducer merged it in and it became the load-older anchor, so paging `before` it skipped every row between. A goal's counter-only revisions during a long goal turn reach any client that attached after the goal row left its window, so a reopened chat lost rows on scroll-back. The reducer now admits live rows only at or above the window's oldest row while older rows remain on the host; the journal keeps the revision and the page reader serves it once the window reaches the row. With nothing older on the host the window is the whole journal, so a row below the head is admitted as before. Also drain accepted provider events before a goal set reads the journal to decide whether it replaces a recorded goal. |
||
|
|
a375936c04 |
feat(agent-launch): let a caller reserve the pane its terminal launch creates (#22291)
* feat(agent-launch): let a caller reserve the pane its terminal launch creates * fix(agent-launch): refuse a launch whose reserved pane is already live * fix(agent-launch): refuse a live reserved pane before it is revealed The live-pane refusal used to fire in the executor, after createTerminal had already issued a handle, published the mobile snapshot and revealed the tab. The reveal re-registered a fresh launch config over the running agent's. agent.launch now passes requireFreshPane with a reserved pane, and createTerminal throws AgentLaunchPaneAlreadyLiveError as soon as spawn reports it attached to a live pane. That is before any handle, snapshot or reveal. The spawn reattach itself is the one terminal.create already uses, so the live PTY is never killed, and the stable-pane create claim is still released in finally. The isReattach plumbing added to the launch factory for the old check is gone. A replay-safe launch refused this way on an existing workspace now records a failed ledger row, the same way a name collision does. Before, the row stayed claimed, so every retry got agent_session_operation_unknown. agent.launchReplay passes the code through. On create-worktree the workspace already exists when the terminal is refused, so the row stays unknown. The code is added to the runtime passthrough list so callers can branch on it. The pane key is now in the replay fingerprint, deliberately. It is not placement: group, anchor and focus still stay out of the request and out of the ledger. It is identity. It is written into the pane's PTY environment and names the tab the caller has placed. A retry that reserved a different pane is therefore a different request. Replaying the first answer would return a key the new reservation can never find. This matches terminal.createAgentSession, which also fingerprints its tab and leaf ids. The key is only folded in when present, so every existing digest is unchanged, and a test pins that. The wire schema now refuses a tab id the runtime would not adopt as sent: one with surrounding whitespace, which the runtime trims, and one longer than 512 characters, which the spawn reservation does not key on. It reuses the tab-id schema that Placement uses. * test(agent-launch): pin that a refused live pane issues no handle The refusal test named handle issuance but only asserted the reveal, so a throw moved to just before the reveal would still pass. Assert no terminal is registered, with the attach test as the positive control. |
||
|
|
eb18eaf2b6 |
feat(usage): add Muse Code local usage provider (#22379)
* feat(usage): add Muse Code local usage provider Scan Muse session logs (including subagent logs, which hold usage the parent log does not) for model_completed token events and surface them as a fourth local usage provider: shared scan worker, persisted per-file cache reused by mtime/size, cross-log dedupe, Stats tab, and Usage Overview integration. Muse logs carry no price, so the provider reports tokens only. * fix(usage): name Muse in Stats & Usage copy; skip partial-cost warning when nothing is priced * fix(usage): surface unreadable Muse sessions root; name Muse in remaining Stats & Usage copy * fix(usage): count distinct same-content Muse records within one log |
||
|
|
8757e40063 |
fix(native-chat): keep a structured agent's tool line between tool calls (#22349)
* fix(native-chat): keep a structured agent's tool line between tool calls A structured session's status named a tool only while the call was still running, so the sidebar's tool line went blank whenever the agent was thinking or writing between calls. Terminal agents keep naming the finished tool until the next one starts, and clear it after a failure. The structured status projection now does the same: a running call wins, otherwise the turn's newest root call if it completed. * fix(native-chat): bound the structured tool line by the running turn, not the user row A send made while a turn is running writes its user row into the journal straight away, and the turn keeps going. Stopping the scan at that row blanked the tool line while a tool was still running. The scan now runs to the turn record and names a call only when that record is still running, so a turn that already ended never lends its last tool to a pending follow-up. This lookup was the running-only lookup's only production caller, so it replaces that lookup instead of sitting beside it. * fix(native-chat): keep naming a structured agent's failed tool until the next one Clearing the tool line after a failed call brought the blank gap back for much of a turn: Codex marks any nonzero exit as failed, so a search with no match or a red test run is enough. The failure already shows on the tool's own row in the transcript. The running turn's newest running call still wins; otherwise its newest root call is named whatever it settled to. * fix(native-chat): name a structured Codex edit on the tool line as the chat draws it Once a Codex edit's changes exist, its apply_patch call becomes a diff row, which the status lookup skipped, so the row named the command before the edit. The chat's tool-call block for a journal row now comes from one shared builder, and the status lookup reads the same definition: a diff is named as Diff with its path, and counts as settled since it carries no lifecycle. * docs(native-chat): describe the structured tool field as running-or-latest The status summary's toolName/toolInput now name the running turn's latest tool between calls, not only a running one. Update the wire type and status bridge comments that still said "the running tool". |
||
|
|
996f9cc306 |
feat(mobile-web-bundle): gzipped 384 KiB ranges over a capability-negotiated mobileWeb.bundle.range (OTA phase C follow-up) (#22381)
* feat(mobile-web): serve gzipped 384 KiB bundle ranges behind a capability Adds mobileWeb.bundle.range with its own strict params and result, so shipped chunk readers see no reply change. The host gzips each range at level 6 and sends identity when gzip does not shrink it, sharing the chunk method's read-slot budget and per-asset verification. status.get advertises mobileWeb.bundle.range.v1 beside mobileWeb.bundle.v1, and the method is allowlisted for paired phones. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * feat(mobile): read the bundle range capability and range replies Adds the range reply reader and operation, and picks range or chunk from the status.get capabilities the connection already proved, so an older desktop keeps being paged in chunks with no probe round trip. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * perf(mobile): keep four bundle chunk reads in flight across the whole manifest The fetch ran one worker per asset and paged inside an asset sequentially, so the largest script's 71 chunks were 71 serial round trips while the other readers idled. One window of four chunk reads now covers every (asset, offset) on the host's chunk grid, largest asset first. A read_limited refusal narrows the window and retries the read; eof is still read from the reply. Synthetic manifest (one 71-chunk asset, five small): 72 round trips -> 19. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * chore(mobile): add fflate 0.8.2 for gzip bundle ranges Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * feat(mobile): decode gzip bundle ranges into a bounded buffer Inflates each range into a buffer one byte past its window, so a gzip bomb costs at most that allocation and an overlong body is visible. A corrupt, truncated or unknown-encoding body refuses as range-undecodable; a body of the wrong decoded length refuses as range-length-mismatch. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * feat(mobile): pass the bundle read method from the session to the fetch The download reads the capabilities of the gates the reducer decided under and hands the fetch range or chunk. The fetch does not act on it yet; the range read lands on the pipelined window. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): re-record the bundle-fetch family under pipelined reads Baseline moves to |
||
|
|
51d3cafc4f |
feat(editor): add a setting to turn off preview tabs (#22398)
Single-clicking a file in the Explorer, or following a link in Markdown source, opens it as a preview tab that the next preview open replaces. There was no way to turn that off, so browsing files kept swapping one tab. Adds `editorPreviewTabsEnabled` (General -> Navigation, on by default). A caller's `preview` flag is now an intent that `resolveEditorPreviewIntent` resolves against the setting, covering every open path - files, diffs, history diffs, conflicts - in one place. Preview-ness is derived rather than reconciled: readers treat a tab as a preview only when the stored flag and the setting agree, so a flag left over from a saved session, another window, or a host switch is inert while previews are off. Nothing rewrites stored flags when the setting changes, so no settings-landing path has to remember to clean up. Fixes #22397 |
||
|
|
52a1e2875b |
feat(orchestration): accept Muse model and effort for supervised workers (#22383)
* feat(orchestration): accept Muse model and effort for supervised workers `worker-start --agent muse` already launched, but `--model` was refused because Muse had no session-option catalog. Add one that maps worker preferences to `muse --model <id>` and `--reasoning-effort <level>`; it seeds no models, so native-chat surfaces show no picker. opencode stays without `--model`: the opencode 2 TUI (now shipped as `opencode`) rejects the flag, so the refusal now tells callers to rely on the agent's own config. Help, skill guide, and docs list valid `--agent` ids and the agents that accept `--model`. Refs #19823 * test(mobile): repin session route closure for the Muse option catalog |
||
|
|
83dd047fd9 |
fix(explorer): find files by name in large local workspaces (#22369)
* fix(explorer): search local workspaces by file name across every file The Explorer name filter only searched remote workspaces directly; local workspaces still filtered the first 20,001 listed files, so files beyond that cap never matched in large repos. Local name queries now rank the whole workspace on the host, and fall back to an uncapped git listing when ripgrep is not installed. * fix(explorer): filter capped local listings on the host with the Explorer word rule Replaces the Quick Open fuzzy top-32 routing, which dropped multi-word matches and capped visible results. The Explorer keeps its instant renderer-side filter; only when the local listing hits its cap does it re-list on the host with the same word rule applied before the cap. * fix(explorer): keep capped matches when host name filtering fails - Fall back to the capped listing (and stop re-listing) if the host scan fails - Keep primary matches when the ignored-file pass fails during a filtered scan - Key host scans on normalized filter words; reset capped state per filter session - Bound nameFilter size at the IPC boundary; drop the double readdir walk * fix(explorer): match name filters without locale-dependent lowercasing * fix(explorer): avoid render-time ref writes in the host name filter fallback |
||
|
|
ebed0964a2 |
feat(agents): add first-class Muse Code harness (#22216)
* feat(agents): add first-class Muse Code harness Add Muse as a supervised Orca agent across desktop, mobile, session history, source control, local hooks, SSH, WSL, and native Windows. Preserve user settings, support Muse 1.3 hook environment allowlists, and recognize versioned foreground processes. Include question, waiting, completion, resume, and readiness coverage. Co-authored-by: homesh-dev <300847526+homesh-dev@users.noreply.github.com> Co-authored-by: jeffhuen <32542276+jeffhuen@users.noreply.github.com> Co-authored-by: John Cusack <johncusackccm@gmail.com> Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com> * test(agents): cover Muse remote hook registration * test(agents): cover Muse hook and source-control contracts * test(agents): exclude Muse hook metadata from script mode check * test(agents): keep Muse skill picker coverage stable * test(ai-vault): include Muse in every-agent fixture * test(mobile): repin Muse agent icon closure * fix(muse): detect questions and approvals from structured Muse signals Muse 1.3 fires no hook for request_user_input, so a pending question left the pane "working". Its internal reminder subagents also post hooks with their own session ids (even after Stop), which surfaced "tool failed" rows and flipped finished panes back to working. - Read pending questions from Muse's session log (user_input_prompt_requested/settled) via the existing transcript poll, now generalized from Codex subagents to Muse on main and relay. - Drop child-session hooks (SubagentStart ids, or turn_id === session_id). - Treat Notification permission_prompt as the approval wait; PermissionRequest also fires for auto-approved calls, so it only caches the approval card. - Ignore Notification copy as the prompt; poll replays are not new prompts or turn boundaries. - Allowlist USERPROFILE so Windows cmd AutoRun doesn't fail every hook. * perf(muse): parse only question events from the session log Most Muse session-log lines are large model/tool records. Filter raw lines by the user_input_prompt_ marker before JSON.parse via an optional readJsonlCursor line filter. * fix(muse): unwrap batched log records and scope questions to the live turn Review follow-ups: question events inside retained_frame batches were skipped, and a question left open by a crash or interrupt stayed pending for the pane's life. Share the history scanner's retained_frame unwrapper, and only report a pending question whose run_id matches the hook turn_id. * refactor(muse): drop type assertion in retained_frame unwrap * fix(agent-hooks): satisfy exhaustive-switch lint in transcript poll policy --------- Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com> |
||
|
|
7c46a69049 |
feat(telemetry): report the macOS daemon's code identity on adoption and folder-denial events (#22171)
* feat(daemon): import the macOS process code-identity probe from PR #21826 Takes `daemon-mac-code-identity.ts` and its test verbatim from David Bebawy's community PR #21826 (stablyai/orca). The probe asks Security.framework, via `codesign --display --verbose=1 +<pid>`, where a live process's code lives on disk — the question Node cannot answer, and the one that decides whether tccd can still resolve a running daemon's code identity after an app update. Imported unchanged here so the adaptation that follows is reviewable as a diff against the author's original. Co-authored-by: David Bebawy <david.ayad2@gmail.com> * feat(telemetry): report the daemon pid's macOS code identity on the two adoption events Community PR #21826 argues that macOS terminal daemons lose Documents/Desktop/ Downloads access after an update because the daemon's own executable is unlinked — Squirrel parks the outgoing bundle under a ShipIt staging directory and later deletes it — so tccd can no longer map the daemon pid to on-disk code. Today's `spawner_path_class` and `tcc_attribution` read the binary that forked the daemon, which an in-place update deletes and recreates, so neither can see that state. This adds the detector as a measurement only. `code_identity` rides on `daemon_adopted` and `daemon_pty_cwd_denied`, the two events that already describe an adopted daemon, so denied daemons can be cross-tabbed against healthy ones. Nothing reads the verdict: no replacement, no notice, no UI. The probe is David Bebawy's, narrowed from a path-carrying union to the closed enum the wire allows, and memoised per pid so one codesign spawn answers for a whole daemon generation. Off macOS, or with no pid, it reports `probe-failed`, which keeps both schemas strict and non-optional. Co-authored-by: David Bebawy <david.ayad2@gmail.com> * fix(telemetry): read the daemon's code identity fresh on every adoption event The probe memoised its verdict per pid and never expired it, so `daemon_pty_cwd_denied` reported whatever the probe saw at adoption rather than what was true at the denial. That breaks the measurement in both directions: a transient codesign failure during startup pinned `probe-failed` for the rest of the run, and the `parked` to `unresolvable` transition became invisible. Squirrel leaves the parked bundle in place until the next update, which can be days, so a daemon adopted as `parked` and denied as `unresolvable` is the exact crossover this study exists to catch, and the cache hid it. Now every ask runs its own codesign. Only concurrent asks about the same pid share a probe, and that entry is cleared as soon as it settles, so nothing survives to be reported later. Both events are rare enough that one spawn each is not worth a cache. * fix(telemetry): drop the dead existence check from the code-identity probe The classifier stat'd the path codesign displayed and called a missing one unresolvable. That path is unreachable: once the executable is unlinked, `codesign --display` prints no `Executable=` line at all and exits 1 with "No such file or directory", which the fallback below already classifies as unresolvable. Verified directly on Darwin 25.5 against a signed binary deleted out from under a running pid. All the branch actually covered was the window between codesign reading the path and this process stat'ing it, and it paid for that with a synchronous stat on the main thread. * fix(telemetry): never classify a timed-out codesign probe as a verdict `runProcess` kills the child at the deadline and reports `timedOut`, but the runner type dropped that field, so a codesign killed mid-display could still have printed an `Executable=` line and been read as `resolved` or `parked`. A half-written display proves nothing about where the daemon's code lives. The runner result now carries `timedOut`, and a timed-out probe returns `probe-failed` before the output is looked at. * docs(telemetry): state what each code-identity verdict actually asserts A reviewer read `resolved` as a claim that the executable sits inside the installed app and asked for that to be validated. It is not that claim, and we are not making it: proving containment needs the pid record's spawner path, and deciding anything from where the code lives is #21826's proposed behaviour rather than this measurement. The enum doc now spells out all four verdicts in the terms the probe can actually support, and says plainly why `resolved` stops at "exists and is not parked". A matching note sits beside the parked-path pattern. * docs(telemetry): stop asserting how long a parked bundle survives The probe's rationale claimed Squirrel keeps the parked bundle "until the next update". A reviewer claimed the opposite, that it is deleted at the end of the same install. Neither holds up against this Mac's ShipIt log: the install moves the outgoing bundle to a TMPDIR ShipIt directory and logs no removal of it at all, and the one "Couldn't remove owned bundle" line names the incoming download staging copy, not the parked one. Every parked bundle from the last two days is nevertheless gone now. So the rationale in the probe doc, the enum doc, and the reprobe test comment now assert only what is established: the outgoing bundle is moved aside at install and disappears later on a schedule we have not pinned down. That is already enough to justify the design, since one pid's verdict can change within an app run, which is exactly why every ask reads fresh. * feat(telemetry): report readable TCC-gated spawns as the code-identity control `daemon_pty_cwd_denied` gives code_identity's hit rate on denials, but a readable spawn emitted nothing, so an `unresolvable` adoption with no denial could not be told apart from a user who never opened a terminal in Documents, Desktop, or Downloads. The false-positive rate that gates #21826's auto-replacement was unmeasurable. `daemon_pty_cwd_readable` now fires when a daemon reads a TCC-gated cwd, once per daemon and folder class per app run, with the same origin properties as the denial event. The read-out becomes a 2x2 of code_identity against readable/denied on protected-folder spawns. Fire-and-forget on the spawn path like the denial emit, and no app-side directory read. * refactor(telemetry): one emitter and schema for both cwd verdicts, no dedupe state The once-per-daemon dedupe on `daemon_pty_cwd_readable` was keyed before the probe ran, so a daemon first seen readable while `parked` never reported again once it turned `unresolvable` — the one cell that would count most against #21826. It also counted per daemon while denials count per spawn, so the 2x2 mixed units. Readable now reports every spawn, like denied, and both events share one emitter (`trackDaemonPtyCwdVerdict`) and one schema. The TCC-folder gate lives in the verdict branch. The origin fields are one shape spread into both schemas. The codesign probe calls `runProcess` directly and tests mock it, replacing a test-only runner parameter. The repeated "never cached" rationale is now said once. * fix(telemetry): rename the shared origin schema fields for the anti-slop gate no-shape-in-symbol-names rejects daemonOriginShape; the fields are event props. --------- Co-authored-by: David Bebawy <david.ayad2@gmail.com> |
||
|
|
1b85be67d8 |
feat(native-chat): notify on every settled structured turn (#22105)
* feat(native-chat): notify on every settled structured turn
A structured chat that finished while you were elsewhere lit the sidebar
but never raised an OS notification, and a notification that did arrive
for one could not open the chat it came from.
Unread and delivery now come out of the single resolveAgentAttention
decision the terminal lane already uses: the structured dispatcher calls
applyAgentAttention instead of applyAgentAttentionUnread, so the same
policy that decides what to light also decides what to deliver, through
the same sound and blocked-permission tail.
Every settled turn notifies, as the CLI lane does. Success says
"finished"; failure and cancellation say "stopped" through the shipped
agentInterrupted flag rather than a second vocabulary. A turn whose
outcome the host never stated stays unknown and lights nothing.
The host now dedupes mobile fan-out by event identity (scope, session,
turn) beside the existing per-workspace burst cooldown, so a completion
two windows both saw reaches the phone once while each window still
decides its own banner. Clicking a structured notification reveals the
chat tab: its pane key's leaf is synthetic, so focusTerminal would hunt
a split-layout leaf that does not exist.
* fix(notifications): spend each mobile gate only when it actually notifies
Two review findings on the structured-chat notification lane, both real.
The mobile event gate consumed its reservation before the per-workspace
burst cooldown ran. Two chats in one workspace share that cooldown key,
so the second chat's completion could burn its event key and then lose
the cooldown to the first chat — never announced, yet permanently marked
as announced, so a later window dispatching it could no longer reach the
phone. The gate now peeks first and records the event at dispatch, which
also keeps a known duplicate from burning the cooldown slot.
A notification id is minted from the status row's stateStartedAt, and the
row re-projects that field as the turn settles: the working episode's
start moves into stateHistory and the settled start takes its place. A
banner raised in the window before that re-projection therefore carried
an id acknowledgement never rebuilt, leaving it on screen for good.
Acknowledgement now collects ids for the row's left episodes too — the
same episodes the unread check beside it already scanned, so the two
halves finally read the same turns. Lane-neutral: the terminal lane
mints its ids the same way and had the same gap.
* fix(notifications): drop the mobile event gate and reveal chats in folder workspaces
The per-event mobile dedupe defended against one completion being
dispatched by several Orca windows. Only one renderer mounts the
structured attention bridge, the completion feed is live-only with no
replay, and any in-process duplicate lands inside the existing 5s
per-workspace burst cooldown, which already collapses mobile and
desktop alike. The gate never acted on a real sequence, so the wire
field, the shared ledger and its tests go; mobile delivery is back to
main's behavior.
A folder workspace id ("folder:<id>") has no "repoId::" prefix, so the
click binding was skipped and clicking a chat notification there did
nothing. The chat route selects its workspace itself through
ui:focusEditorTab, so it now binds without a repoId; the terminal
route is unchanged.
* fix(notifications): retire the banner ids actually dispatched, not ids rebuilt from a moved row
A banner's id is minted from the status row's stateStartedAt at dispatch, and that field moves
afterwards: a completion can outrun the settled re-projection, and a settled structured row is
re-stamped with no history entry by any later journal row (a cancel appends a status note after
the turn settles). Rebuilding ids from the row's episodes at acknowledgement missed the second
case and fanned out up to 21 mobile dismissals per pane for ids never raised.
The shared delivery tail now records each dispatched id per subject; acknowledgement retires
those plus the current-row rebuild it always had. The acknowledgement collector is back to
main's single-field form.
* refactor(notifications): retire announced notifications by subject in main
Main now records, per pane, the ids it actually announced (a desktop banner
shown or a phone alert sent) and an acknowledgement passes the acknowledged
pane keys so main retires all of them. This replaces the renderer-side record
of dispatched ids: main is where the announcement happens, so it records only
real announcements, including phone alerts whose desktop banner focus
suppressed. The id rebuilt from the current row stays as the fallback after
a restart empties the in-memory record.
|
||
|
|
dfff3915c4 |
fix(browser): scope back/forward/reload/zoom/grab shortcuts to the originating split (#22340)
* fix(browser): scope back/forward/reload/zoom/grab shortcuts to the originating split With two browser panes visible in a split, Back, Forward, Reload, Hard Reload, page zoom and Focus Address Bar fired in every visible pane. Main forwarded these guest chords without the page id, and each split's active pane subscribed. The renderer-side listeners for the same chords were also window-wide per pane, so a key pressed in the toolbar (or in a terminal in another split) reached every active browser pane. Guest-forwarded chords now carry the originating browserPageId; preload admits only well-formed payloads and each pane ignores ids that aren't its own. Toolbar-path listeners use the same focused-split scope Find already uses. The streamed remote pane's history chord moves onto that scoped hook. Cmd/Ctrl+C grab (STA-3319) gets the same scope and no longer arms while a text selection exists outside the browser pane, so copying from the native chat transcript works again. * refactor(browser): simplify split shortcut scoping per review Drop the preload payload admission (main and preload ship together), fold the three inline scope checks into browserChromeShortcutOwnsEvent, and replace the outside-overlay selection check with a plain live-selection rule so Cmd+C copies from surfaces that do not move split focus. * refactor(browser): share one zoom command type and tidy shortcut comments BrowserPageZoomEventDetail and BrowserPageZoomCommand were the same shape; keep one in shared/browser-page-zoom.ts and route guest and local zoom through a single handler. * refactor(browser): narrow the zoom event with instanceof instead of a cast * test(e2e): pin split-scoped browser shortcuts Two browser splits (and a terminal beside a browser) now prove that Back, Forward, Reload, Hard Reload, page zoom, Focus Address Bar, and the element grab chord act only on the split that sent them, from both the guest page and the browser toolbar. A native chat selection proves Cmd/Ctrl+C copies instead of arming grab. Split fixtures move to a shared helper so both specs reuse them. |
||
|
|
9ece273056 |
fix(native-chat): journal rows name the agent that produced them (#22299)
* feat(native-chat): carry producer linkage on every journal row A journal is the durable record of one agent SESSION, and a session that runs subagents journals their rows into the same timeline with nothing on the row saying which agent wrote it. Add that: a per-row linkage bundle naming the producing agent, its parent, the provider's raw parent reference as provenance, the kind of work, and which run of the agent produced the row. The bundle rides the row BASE, not the body: two nested prompt shapes are strict, so an unknown key on a body makes the whole row parse as malformed. It is deliberately not a schema-version bump either — an unknown `v` makes a row unreadable and latches the host read-only, while an unknown key is ignored, so an older host reads a stamped row and behaves exactly as it does today. One reader predicate interprets absence, by presence and not by truthiness: an id that failed to resolve is still an id, and a truthy test would read it as root and put the child's content back on the parent. The parent-facing status scans — thinking, the running tool call, the latest assistant line and the quoted prompt — now skip rows a subagent produced. The transcript is left unscoped on purpose: it shows every agent's output. No producer stamps anything yet; this is the carrier and the reader. * fix(claude): attribute a subagent's journal rows to the subagent The Claude translator already parsed `parent_tool_use_id` on every envelope and threw it away. It now resolves that reference to the producing agent's canonical task id — never to the reference itself, which names the tool CALL and is re-minted on every resume, so a row stamped with it would split one child into two the moment it resumed. The raw reference is kept beside it as provenance. Resolution is its own module rather than more roster: the roster maintains the spawn-group row a user reads, while this answers, for one frame's parent reference, whether the rows it produces are the session's own agent's, some child's, or nobody's yet. It reads the alias table directly to tell "an announcement named this spawn call" from "this id is simply unknown", which comparing the canonical id against the raw one cannot do when the two match. Where the identity is not final the row waits rather than guessing. A top-level spawn whose `task_started` has not landed is the one case that can still resolve, so its rows are held — bounded at 64, oldest written first — and released when the announcement arrives or when nothing can name the producer any more. Nothing is dropped and nothing is written as the parent's. A release that announces no tasks at all is decided immediately instead of held: nothing stable is ever reachable for its children, and an id that rotates is worse than no id because it is silently wrong rather than visibly absent. Those rows read as the session's own, exactly as they do today. Holding them instead would strand every row that REVISES an earlier one — a tool result would leave its tool row reading "running" for the rest of the turn. The attempt counter moves in exactly one place, the existing reactivation branch where a new spawn alias reopens an entry, and is gated on that observed alias change rather than on the counter, so a late duplicate cannot advance a settled run. The first run carries no attempt at all. The group row keeps its own module's write path, now named there, because it is the one row written from a child's frame that is deliberately the parent's. * test(native-chat): pin producer linkage end to end, and fix the harnesses first Four harnesses in this area silently discarded the append options they were handed, so every assertion about attribution would have passed against `undefined`. Two are fixed here — the journal double behind the real deferred sink, and the Claude subagent translator's sink — and each records the options beside its existing call log rather than on it, so the assertions about call order stay about call order. The read side is pinned first, because linkage correct in the store and never read by the projection is the way this ships looking finished and fixing nothing. The three defects are asserted through the live-turn and projection readers: a parent no longer reads as thinking because its child is reasoning, no longer shows its child's running tool, and no longer quotes its child's prose or prompt. The opposite direction is pinned too — the transcript still renders the child's output, and a parent's own line is never suppressed. Also covered: a resumed child keeps one identity while its spawn call id rotates; a child row arriving before its announcement is held and then written linked rather than dropped; a held row is written under the raw reference when no announcement ever comes; the buffer's bound writes the oldest row rather than losing it; an unresolvable id reads as a child rather than as the parent; a row with no linkage reads as root; the schema version is unchanged; and a strict prompt shape still parses, with a positive control proving that strictness is real and is why linkage rides the row rather than a body. * test(native-chat): give the journal double's cast its SAFETY rationale Editing inside the object literal re-attributes the pre-existing assertion to changed lines, and the changed-code gate requires a line-specific rationale. * fix(claude): name the agent that spawned a nested subagent A grandchild's rows carried an agent id but no parent, and under this journal's semantics an absent parent is not silence — it is the claim that the session's own agent spawned the row's producer. For a task spawned from inside another subagent's sidechain that claim was simply false. The frames from such a task carry exactly one handle: the nested call's tool id. That id was journaled once already, as a tool-use block on the row of the child that made the call, so the child is recoverable from it — but only if something remembers which row carried it. The registry that already tracks which tool calls reached the top-level transcript now records the sidechain ones too, against the reference naming their owner, and the resolver follows that reference to name a row's parent. The reference is recorded, not an identity, and it is resolved through the same path the owner's own rows resolve through, so a parent id always matches the agent id the parent's rows carry however either was settled. A row now persists only once BOTH its producer and its parent are final; a grandchild whose child has not been announced yet waits in the same buffer, and leaves it through the same three doors. * test(claude): pin the streamed-text lane's attribution instead of only capturing it The checkpoint harness was fixed to record the append options, and then nothing asserted them: all five of its tests passed unchanged against an implementation that resolves no producer at all, so the lane's attribution was covered by a capture and no claim. These assert it: a block streamed inside a child carries that child's linkage, the session's own carries no keys at all, a checkpoint is held rather than written while the producing agent is provisional, the flush before settlement writes a held block under the raw reference rather than losing it or filing it as the parent's, and a block's producer is resolved once and kept — every checkpoint rewrites the same row, so a producer that moved would file one agent's prose under two identities. * refactor(journal): name the linkage row fields for their role, not their shape The anti-slop gate rejects "Shape" in a symbol name. These are the linkage fields a render item carries, and the name now matches the sibling helper that builds them. * fix(claude): hold the session's first subagent instead of filing it as the parent A child's first frames can arrive before the `task_started` that names it, and the resolver treated "this release has announced no task" as a settled fact about the CLI. Before its own first announcement every session looks exactly like that, so the FIRST subagent's pre-announcement rows were written straight out as the session's own — the whole defect, for the first child of every session, persisted with no backfill to repair it. A spawn call the session forwarded at top level is positive evidence that an announcement is still expected, so it now outranks the release check. A release that genuinely announces nothing is unchanged: its rows reach the release verdict at settle and still read as root, just written a little later. Also stops a malformed owner chain that loops back from naming an agent its own parent; the depth guard bounded that walk but could not make its answer mean anything, and absence is the truthful claim. * fix(claude): read a tool result as its caller's row, not a child's Every top-level tool call is a forwarded tool id, not just a spawn, so a result frame naming its own call as parent resolved as a child awaiting an announcement that is never coming. The row was parked until the turn settled and the tool sat `running` in the meantime. A frame delivering the result of the very call it names is the caller consuming its own output; only a spawn call ever gets a sidechain. Pins added for that and for the first-subagent hold, and the never-announced case now asserts the row was WRITTEN as root rather than that it carries no agent id, which an absent row also satisfied. * fix(journal): refuse an empty producer id, and drop a bad one without losing the row The reader that scopes a parent's surfaces tests PRESENCE, so `agentId: ''` is present: a row carrying it reads as a subagent's and disappears from its own author's surfaces for good. Neither validator caught it — the wire schema accepted any string, and the persisted-row guard type-checked nothing in the bundle at all, against that file's own stated policy. The wire schema now requires a non-empty id. The persisted side sanitises instead: a bad linkage field is DROPPED and the row is kept. Rejecting there would turn a tightened validator into a whole-store kill switch, and degrading a row to the session's own agent is what every row said before linkage existed. * fix(native-chat): answer the turn activity line for the session's own agent `selectStructuredAgentTurnActivity` is a "what is this agent doing right now" reader and was not scoped by producer. It builds a label set from every tool-call row in the turn — a subagent's included — and both readers below it use that set to suppress a line that repeats it. So a CHILD's tool label could blank the PARENT's activity line: child data deciding the parent's surface. Live, not latent: the provider-activity branch is populated in this lane, and it consults the label set without ever consulting `providerFrame`, which is what the status fallback loop relies on. Scoped once at the top, so both readers share one interpretation point. Renderer and mobile share this function, so both are covered. * fix(journal): stop a lifecycle batch stamping one producer onto N mutations A lifecycle-batch row carries N mutations but stamped linkage at ROW level, so a future mixed-producer batch would silently attribute every mutation to whoever opened it. Both callers are single-producer today, so this was latent. The write path no longer accepts linkage for a batch, which removes the failure mode by construction rather than guarding it. The reducer still READS linkage off a batch row — a row may arrive from a host that writes one — and a genuinely mixed batch would have to stamp per mutation, which nothing needs yet. Chosen over adding a per-mutation field because that would persist a new key forever with no writer and no reader. * docs(journal): say why each unlinked write site is unlinked, and drop two false claims Completes the write-site audit the PR claims. Prompt rows carry no linkage and CANNOT: a prompt arrives through the SDK's permission callback, whose options carry a request id and the tool awaiting approval and no parent reference of any kind — unattributable at that site, not deliberately root. Turn rows are deliberately root and now say so. Two comments justified decisions by mechanisms this store does not have. The linkage docblock cited compaction dropping a start row and a pagination boundary; there is no compaction, and pagination is complete-or-reset. Per-row repetition is still right, for the reason that is actually true: every reader scans back from the tail and stops at the turn. A test carried the same false framing. `claudeFrameParentRef` claimed to read the field by the same rule as `isRootClaudeFrame`; it is deliberately stricter on the empty string. * refactor(claude): write a child's rows through, then correct the attribution Four misattribution paths shared one cause: the lane committed to an attribution verdict at write time and could never revise it. That followed from "the journal has no backfill", which is false — re-appending an `itemId` bumps its revision, the reducer rebuilds linkage from the newest row, and it pins `sequence`/`observedAt` so a correction does not move the bubble. This lane already relied on that twice. So the order inverts. A row whose producer is still provisional is written immediately, stamped with the spawn call's own id, and re-attributed in place when the announcement names it. Bookkeeping no longer gates a user's view of what an agent said. The hold buffer is deleted rather than left as a pass-through. Corrections are bounded and die four ways: the announcement, turn settle, teardown, or the bound. Passing the bound gives up on that producer WHOLESALE — correcting some of a child's rows and not the rest splits one child across two ids, which is worse than correcting none. A correction that would change nothing is dropped rather than burning a revision. The streamed lane loses its producer latch, which pinned the first verdict permanently and is why an announcement one frame later could never reach the row. Every checkpoint rewrites the same identity, so there is one row per block and re-resolving can only revise it; the latch was guarding against a split that cannot happen on this path. A block that stops streaming before its announcement is re-attributed explicitly, since nothing else revisits it, and the announcement is now observed BEFORE the forced flush that would otherwise stamp it a line too early. Also narrows the no-announcements-at-all escape so it no longer swallows a forwarded spawn call. That escape now applies only to a sidechain id no spawn call ever forwarded, where there is genuinely no handle to stamp. * fix(journal): move two test doubles onto the signatures they pin Both failed typecheck while passing at runtime, which is what a test double gets to do: vitest never typechecks them. The sink's lifecycle-batch double still read producer linkage off the batch input after that input stopped carrying any, so it had no property in common with the linkage type. The fence is now all it records, which is what the narrowed contract actually forwards — and what the test beside it already asserts. The row-schema helper returned the whole six-arm `JournalRow` union while every caller reads `body`. It now narrows to the item arm it always builds, so the assertions read it directly rather than through a cast. * fix(claude): resolve a tool result to its real caller, not to the session root A nested tool row could end stuck `running` with its result content dropped. Cause was in the result-frame attribution, not in the correction ledger. A frame delivering the result of the call it names as parent is the CALLER consuming its own output — but the code read "the caller" as "the session's own agent", which is only true when the caller is the root. A call a child made is owned by that child. Collapsing it to root both misattributed the row and made the result's write resolve through a different reference than the call's, so the correction owed to that row was left holding the body it had BEFORE the result landed, and re-attribution then reverted the row. The caller is now resolved through the registry that already records which agent journaled a tool call, so both writes to one row resolve through the same reference and the newest body wins. A settled write also supersedes any correction owed to its row. One `itemId` is legitimately written under two references — `claudeToolIdentity` is keyed on the tool id alone — and a settled write already carries a final verdict, so an outstanding correction could only restamp it from a reference that write did not use. Dropped rather than re-bodied for that reason. Adds the ledger's first unit tests, including the invariant this defect broke: a correction changes a row's attribution and never its content. * fix(claude): keep a correction owed when the sink refuses it A correction went out through the plain append, which discards the queue's admission. Under backpressure the write was refused and `retry` had already dropped the entry, so the obligation died with nothing re-deriving it — the failure class this work exists to refuse. It is self-feeding too: a correction costs a commit on the same serialized writer that carries live rows, so the burst that generates many corrections is what builds the backlog that drops them. It now uses the admission-returning path the sink already exposes, keeps the entry outstanding on a refusal, and lets `abandon` try once more. A refusal there ends it: the row keeps the spawn call's own id, which is usable, and an obligation with no exit is worse than one that settles for less. `settle` also reports what actually happened instead of always claiming it wrote, so publish no longer fires for a write nobody accepted. Also records why the live-turn scans may read the turn record before checking the producer. A turn is the session's unit of work and no producer of a turn-bearing body stamps linkage: Claude's turn rows carry none, Codex has no linkage concept, the compact row passes only a fence, and the stale-turn sweep goes through the lifecycle-batch path, which cannot carry linkage by type. The ordering is safe by construction rather than by accident, and the comment says so, so a future producer knows what it would break. * fix(claude): never read a row naming a parent as the session's own A non-null `parent_tool_use_id` names a child, always. The resolver still had one branch that read such rows as the session's own agent's — a release that had announced no task, where the comment claimed "there is no handle to stamp". There is one: the reference itself. The branch was buying a false attribution to avoid an id nothing joins on, which is the trade already reversed once for forwarded spawn calls, and every reader of this field is a presence test. So the branch goes, and with it the `root` arm of the verdict and the resolver's whole dependency on whether the release announces tasks. Two states remain: linked now, or linked now and owed a correction. The type deleted a stale test double on sight, which is the argument for removing the arm rather than the branch alone. This also closes the severe half of the tool-origin eviction exposure. A spawn id evicted from the bounded top-level set used to flip the release check on and stamp a child's rows as the parent's; with nothing returning root that cannot happen. What remains is a missed correction, which splits one child across two ids — the same end state as passing the correction bound, benign in kind and disclosed. `isForwardedParentTool` stays where it gates PENDINGNESS. It now decides only whether a correction is owed, never whether a row is a child's, so a stale answer costs precision rather than correctness. |
||
|
|
4c696a1e2a |
fix(agent-status): a structured session with live child work reads as working (#22295)
* fix(agent-status): a structured session with live child work reads as working An idle native-chat session whose subagent was still running showed a green check in the sidebar, the collapsed worktree pill, and worktree ps, while a terminal Claude session in the same situation showed working. The two lanes folded child work into the parent's status with different code: the hook listener did, the structured lane did not. Both lanes now share one child-work liveness vocabulary and one lead-status fold. Live agent work makes a settled lead working; shells and monitors alone make it monitoring. The structured lane derives liveness from the background task list already on the wire, in both its readers, so the sidebar, the CLI, the dashboard and mobile agree. The Claude task-kind table is one shared file covering both the hook inventory and SDK stream names, and the renderer bridge reuses the shared child-work projection instead of carrying its own copy. * fix(agent-status): a blocked or out-of-contact subagent still holds its session working Child-work liveness retired an agent-kind child on any state but working/monitoring, while the shell beside it stayed live on everything except done/idle. A subagent waiting on a permission prompt, or one whose host lost contact, therefore counted for less than a backgrounded sleep and let the session read done. Both kinds now share the settlement rule `resolveAgentChildWorkFreshness` already reads rows by: only an explicit done/idle retires child work. Also keep empty task labels out of the shared background-task projection candidate, so a host that publishes `name: ''` cannot beat the child-row fallbacks. * test(agent-status): pin the widened hook-inventory agent names, and correct two stale claims The hook inventory now classifies through the shared kind table, which also maps the SDK stream's `local_agent` / `local_subagent`. Nothing pinned that widening, so add cases for all four agent names — including `teammate`, whose pane state stays `done` under the #8825 idle-squat rule. Two comments the fold made false: - the teardown marker rule's comment claimed it could not disagree with what the UI calls working; it is deliberately lead-only, so now it says that and why; - the agent-status store reference still described the structured row's `state` as the deleted `structuredAgentSessionStatusState`, and omitted the `workingMode` the ingest now writes. * fix(agent-status): the state clock restarts when monitoring becomes a real turn `stateStartedAt` carried forward whenever the prior `state` matched, which was sound while `state` meant "a turn is running". Now that it folds in child work, an idle lead watching a `sleep 3600` publishes `working`/`monitoring`; the user's prompt 45 minutes later keeps `state: 'working'`, so the row inherited the watch loop's clock and read "Working for 45m" the instant the turn began. Monitoring is its own displayed label (`worktree-card-compact-agent-row.tsx:40`), so the continuity key is now the whole published work identity — state AND workingMode — in both writers. Also record two facts the code stated wrongly: the structured lane's `interrupted: false` is inert (a projected session status has no interrupted member) rather than a decision, and the child-work liveness rule's escape hatch is the roster's session lifetime, not a settled state. * fix(agent-status): a workflow is watch work, and child work dates itself Two defects the fold introduced. `isAgentChildWorkKind` counted `workflow` as agent work, so a structured session whose only live task was a backgrounded `local_workflow` published a full working spinner while the children projection — which admits `kind === 'agent'` only — rendered nothing to expand, and the same workflow in a terminal pane showed the monitoring badge instead. The repo already decides this: `isClaudeSubagentTask` excludes workflows by name, and MATERIALIZED_TASK_KINDS leaves "the backgrounded shell command and the workflow" to the non-agent owner. The predicate is now `kind === 'agent'`, and the three sites that restated the same test route through it, so a new kind is decided in one place instead of three that merely agree. `evidenceObservedAt` dated every row by `summary.updatedAt`, the journal's last activity. The journal cannot date child work: its clock stopped when the lead's turn did, so a genuinely live roster aged past the 30-minute staleness window and mobile's dot decayed a running session to idle. The fold now reports whether child work alone holds the row open, and only then does the host's observation clock stand in — keeping "a restart's republish is not new evidence" for lead turns. * fix(agent-status): the sidebar dates child work the same way the host does `fromChildWork` reached the host ingest but not the renderer bridge, so after ~30 minutes of live child work with no journal activity the sidebar's row aged into staleness while `worktree ps` and mobile stayed fresh — two writers for one session answering differently, which is the defect this PR exists to remove. For a remote host the client's own receipt time is also the more honest clock, since the journal stamp is the host's and is never comparable against this machine's now. |
||
|
|
60c43695e5 |
feat(agent-launch): report the pane a terminal launch created (#22108)
* feat(agent-launch): report the pane a terminal launch created A `term_*` handle is a main-side mapping the renderer cannot resolve (terminal-handle-links.ts:309), so a client that draws its own tabs had no way to name the tab it had just asked `agent.launch` to build. The runtime already mints that pane, bakes it into the PTY's environment and hands it to its own reveal; the surface factory then dropped it on the floor. Carry it through as `paneKey` on the terminal outcome. Identity, not placement: where the pane goes — which group, what order, whether it takes focus — stays with whichever client is drawing, and nothing here rides the wire for it. One field rather than a tabId/leafId pair, because the key already holds both and two copies of one fact can disagree. Absent when this launch minted no pane: a reused terminal was already running, and a worktree-create startup terminal is built by the create, which reports only a handle. Naming the wrong surface is worse than naming none. Optional on the wire and optional on the read side. Mobile parses the receipt with a loose object and is deliberately mode-blind, so it ignores the field; the persisted-row guard checks it when present and accepts a row written before it existed, because a read rule stricter than the write side turns one odd row into a refused replay. * fix(agent-launch): retain startup terminal pane identity |
||
|
|
ba742a86bb |
fix(linux): release orphaned processes when their owner exits (#22247)
* fix(linux): release orphaned processes when their owner exits * fix(linux): handle inhibitor errors until streams close --------- Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local> |
||
|
|
90d363afc9 |
fix(renderer): contain Monaco initialization failures (#21555)
* fix(renderer): add defensive error handling for Monaco editor crashes Analyzed 34 crash reports for v1.4.205 released 2026-09-17. Identified and added defensive fixes for React error boundary crashes in Monaco editor setup. - Error: ReferenceError: thũs is not defined - Location: Monaco editor initialization (editor.api2 bundle) - Platforms: Linux, Windows, macOS - Root cause: Undefined variable in Monaco setup or language registration - Status: Added try-catch to prevent cascade crash - Pattern: Cascading process deaths (network service + GPU service) - Platforms: Primarily Windows - Root cause: Infrastructure/concurrent process failure (not code defect) - Status: Documented, requires Electron/Chrome infrastructure review - Pattern: Renderer memory grows to 851MB on low-RAM Windows systems - Root cause: Memory exhaustion on systems with <2GB free RAM - Status: Existing memory monitoring detected; needs leak investigation - Status: Requires minidump analysis with source maps 1. Added try-catch to Monaco editor mount callback (use-monaco-editor-mount.ts) - Catches errors during editor initialization - Logs file path and error for better diagnostics - Prevents crash cascade to React error boundary 2. Added try-catch to Monaco language registration (monaco-setup.ts) - Catches errors during Vue/Svelte/Astro/Nim language registration - Logs failures without crashing Monaco setup - Allows app to continue even if optional features fail - Analyzed 34 crash reports across 3 categories - Examined crash dumps, diagnostics, and memory profiles - Reviewed Monaco setup and editor component code - Checked git history for recent changes - Crash breadcrumbs (memory, process state, user actions) - Process metrics (heap, private memory, system memory) - Component stacks (React error boundaries) - Exit codes and system signals - Error silently continues instead of crashing: Users get degraded experience instead of app crash, can still use editor in most cases - May hide underlying issues: Errors are logged for crash reports, but won't be surfaced as prominently - Type checking: pnpm tc:renderer (passed) - Changes preserve existing error reporting through crash breadcrumbs - Defensive coding only adds try-catch, no behavior change for success path * fix(renderer): keep Monaco mount failures inside error boundary * fix(renderer): isolate Monaco setup failures * fix(renderer): contain Monaco mount failures at the editor surface The try/catch around the onMount body did the opposite of containment: React already routed that throw to the page boundary, so swallowing it left a half-wired editor and hid the crash from the reporting pipeline. It also never saw the reported failure, which is raised inside @monaco-editor/react's own create effect before onMount runs. Revert the hook to main and wrap the editor element in RecoverableRenderErrorBoundary instead, so either throw degrades the file pane only, still files a crash report, and retries by remounting on the existing pane+path key. * refactor(renderer): drive Monaco setup steps from one guarded table Ten near-identical guarded calls, each repeating its own function name as a label, become one [label, step] table run by a single loop. Same behaviour: an optional registration that throws is logged and the rest still run. loader.config and the editor model registry stay unguarded — they are load-bearing, so catching there would only move the failure later. * fix(renderer): breadcrumb swallowed Monaco setup-step failures A guarded registration that throws was console-only, so a lost language or behaviour guard never reached crash reports. Record a breadcrumb so the containment stays visible in the field. Claude-Session: ab8ff806-4870-4ea8-bbf5-bbd123b1166e --------- Co-authored-by: m4air <m4air@Mac.localdomain> Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> |
||
|
|
e47ef8cc28 |
feat(mobile): the shell tells a page which optional capabilities it has (OTA phase D, C8.1) (#22141)
* fix(mobile): publish page-route pairs the strict host schema accepts (OTA phase D, C8.1) `routeViewOf` handed the manifest's own route entries to the host as `pageRouteGrants`. The phone reads a manifest route loosely, so an entry arrives carrying whatever field the desktop that wrote it knew about, and `BridgePageRouteGrantsSchema` is `.strict()`: one unread key refuses the pairs, `createBridgeHost` refuses the route with them, and the page gets no `init` at all rather than losing one field. Fixed before any route carries an optional grant (ruling 37.4), so the manifest field the next commits add costs an installed shell nothing. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * chore: drop the closure and bundle probe scripts from the tree Scratch measurements for C8.1 (which route closures reach the HTML preview, and what the preview render rig costs to bundle with a client provider). They belong outside the repository and were swept in by the previous commit's `git add -A`. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * feat(mobile): a manifest route may declare optional grants (OTA phase D, C8.1) Design B of design-ota-c8-1.md, ruling 37. `MobileWebBundleRouteSchema` grows `optionalGrants` under the required lane's own grammar, with the 16-name ceiling applied over the union of the two lists rather than to each. Serving a route still reads `grants` alone, so a capability a screen cannot work without stays required and takes the route native; a session's granted list is `[...grants, ...optionalGrants]` narrowed to what this shell implements, from one helper that both `grantsForRoute` and the `pageRouteGrants` publish read. The ruling's compatibility rationale is corrected in place. `z.looseObject` passes unknown members through rather than dropping them (measured, zod 4.4.3), so a shell older than the field still receives the key; what it lacks is a policy that reads one. What makes the lane safe against such a shell is therefore the previous commit's publish fix, not the reader. BRIDGE_PROTOCOL_VERSION stays 1. No new notify, verb or frame field. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * feat(mobile): name the shell's cancelled-navigation behaviour as a grant (OTA phase D, C8.1) `externalNavigation` joins `MOBILE_WEB_SHELL_GRANTS` beside `screencastBinary` and `haptics`, declared in `cancelled-navigation-target.ts` because that is the module holding the rule which acts on it. A third token that is neither a verb nor a notify: the page posts nothing to make a cancelled top-frame navigation happen, so this list is the only thing that can tell a page whether a tap inside the sealed HTML-preview frame escapes at all. A constant and not a platform read (ruling 37.1): both engines dispatch the event, `ios/MobileWebShellView.swift:481` and Android's `MobileWebShellView.kt:382`, so an app build carries the behaviour on both or on neither. The policy census grows the half that was only pinned by the verb table: the implemented set is that table plus exactly three non-verb tokens, each read off the module that declares it. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * feat(mobile): the bundle builder carries a route's optional grants (OTA phase D, C8.1) `resolveMobileWebPageRoutes` maps each declaration member by member, so a field the declaration grows reaches a phone only once the map names it: until now `optionalGrants` would have been dropped in silence and every route would have declared nothing optional. Omitted when the route declares none, because absent and empty are the same answer to a shell. The declaration suite grows the rule rather than a row: the map carries the lane through and writes no key without one, and the lane is held to the manifest's own grammar and to the ceiling over the union of the two lists. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * feat(mobile): the HTML preview hides its links on a shell that cannot open one (OTA phase D, C8.1) The session route declares `externalNavigation` on the optional lane, and the preview asks for it before it renders an artifact's links as links. Ruling 37.2's three readings are what "hide" means here, and removing `href` is what delivers all three at once: `a:any-link` stops matching, so the UA stylesheet stops underlining, the element leaves the tab order, and there is no dead anchor a tap does nothing on. The text the author wrote stays where it was, the artifact paints, and the Preview/Source toggle is untouched. Done with the browser's own parser rather than over the source text: an `href` inside a comment or a `<template>` is text to a browser, and a pass that rewrote either would be editing the artifact instead of its links. The frame also loses `allow-top-navigation-by-user-activation` on that path, so a link the pass somehow missed is refused by the browsing context as well. One route, measured rather than assumed: the design said two, and the file preview route's closure does not reach the HTML preview at all - it renders `MobileFilePreviewScreen`. The new closure census derives that list from the hook's callers. The render rig grows the case on both engines and the readings it needs, and `mobile-web-app-preview-frame-readings.mjs` is split out of it at the readings/arms boundary, because the two were over the 600-line cap together. Two engine findings are recorded in the rig: an `<a>` with no `href` still answers `tabIndex` 0 on both, so focusability is asked by focusing; and WebKit computes `cursor: auto` for a real link, so that reading is pinned where it discriminates and its blindness pinned where it does not. The hop-coverage census now reads the effective set, because that is what the running rule compares. Inert today: the session route is the only declarer and an opener into every other route. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): pin the preview's hidden link path where the unit suite can reach it (OTA phase D, C8.1) The mobile suite runs in a `node` environment whose resolver has no `.web` precedence, so `MobileHtmlPreview.web.tsx`'s import of the grant hook lands on the native sibling, which answers yes unconditionally. That is why the existing component suite still measured the granted frame without knowing a grant exists, and it means the hidden path had no coverage in the sharded `test` job, where the render rig is skipped for want of the bundler's dependencies. So the wiring gets its own file with the module replaced: that the component asks, and that both the frame's sandbox and the document it is handed follow the one answer. happy-dom rather than the suite default, because the inerting pass parses with the browser's own `DOMParser`. `String(node.type)` rather than a literal comparison: `node.type` is `ElementType`, which overlaps a real intrinsic tag and not the host strings these mocks render, so `=== 'Pressable'` is a no-overlap error under `tsconfig.test.json` and the tests-typecheck ratchet reds on it. Also replaces a `Reflect.get` the anti-slop gate refuses with an `in` check. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): repin the session page closure at 4,362 for C8.1's three modules Measured on both sides with `mobileWebAppRouteClosure(SESSION_ROUTE)` at base `841d06a969` with all five postinstall generators run first, and the two `local` lists diffed rather than the total inferred: 4,359 -> 4,362 modules, 1,017 -> 1,020 local. All three are local source modules and none is vendored: the page's read of `init.grants.native`, the pass that turns an artifact's links back into text without the grant, and the module declaring the token beside the rule that acts on it - reached both by that hook and by `page-route-policy.ts`. The `bridge-caps.ts` it imports was already in this closure, and the hook's native sibling is replaced rather than joined. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): allowlist the preview's grant sibling among the .web.* overrides `mobile-web-app-web-overrides.test.mjs` pins the allowlist against the `.web.*` files on disk, so a new web sibling reds it until the file says why the page needs one. Red before: `expected [ …(36) ] to deeply equal [ …(37) ]`, naming `src/components/use-html-preview-link-grant.web.ts`. The preview's own entry is corrected with it: its reason said `allow-top-navigation-by-user-activation` is granted, and that token is now conditional on the shell answering that it can open such a navigation. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): the hidden-link render case waits on the frame's own reading (round 1) CI's chromium arm timed out at the full 240 s on this case alone while the WebKit sibling passed in 1.5 s and it passed 26/26 locally. The cause is the third arm: it tapped the granted link and waited through `expectNavigation: 'main-frame'`, and `waitForRecordedNavigation` has no bound but the case's own timeout. Under CI load the click missed its 2 s actionability window, no navigation was ever recorded, and the arm sat in that wait until vitest gave up - `recorded []`, with the frame attached only at 38.9 s. Three arms sharing one budget is what made this the case to find it. The arm is dropped rather than its wait lengthened or retried. Every verdict left is a reading the frame itself publishes: the anchors its document holds, the style the engine computed for one, whether focus lands on it, and now whether the tap this arm made landed at all - `actError` is asserted null, so a click that never reached its target is no longer the same three zeros as a tap that did nothing. Nothing is lost. The tap's outcome on a granted shell is the next case, on these same counters from this same rig and with a budget of its own, which is the presence precondition this file already uses elsewhere for the same reason. The WebKit sibling's discriminating reads are untouched. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): the inert-link pass changes nothing an engine renders but the links (round 2) Round 2's ruling: the hidden-link path may change nothing about the artifact's rendering except that links are not links. A parse and a reserialise is not free of that by default, and all four findings reproduced on Chromium 147 and WebKit 26.4. A same-document fragment link is kept. It starts no navigation at all, so it goes on working inside the sealed frame whatever the shell can do, and taking it away would be degradation over a capability it never needed - an artifact's own table of contents is the case. Its `target` still goes, because a fragment aimed at another frame is a navigation rather than a scroll, and `href=""` is not a fragment: it resolves to the frame's own URL. Links inside `template.content` are reached, recursively. `<template shadowrootmode>` is a declarative shadow root the frame's parser attaches and renders, and `querySelectorAll` does not walk into template content, so those links arrived live inside a sandbox that refuses their navigation - the dead anchor ruling 37.2 forbids. Measured: `parseFromString` attaches no such root on either engine or in happy-dom, so the pass can reach them. The leading newline of a `pre`, `listing` or `textarea` is written back. A parser drops one after the start tag and the serialiser is specified to put it back; measured, neither engine's does, so a round trip lost a blank line from every such block. The doctype is carried whole, and the reason is corrected from the one the finding gave. It cannot move this frame between layout modes: a `srcdoc` document takes its mode from its embedder, and measured, a quirks doctype, the bare name and no doctype at all all read `CSS1Compat` inside the frame. What rewriting it does is change the document the author wrote for no reason, with `document.doctype` observable beside a Source tab showing the original. The render case pins `compatMode` as the blind reading it is and reads the frame's own doctype identifiers as the one that discriminates. Option B was not available: the frame has no `allow-scripts` and inherits `script-src 'self'`, so nothing runs inside it and there is no injection to carry the work. Also drops a vacuous half of the affordance test. `renderSource()` is called with no argument, so the markup a Source view shows is the caller's own closure and asserting it equals the fixture passed whatever the component did. What the component decides is whether the rewritten frame stays mounted underneath, and that is what is read now. `mobile-web-app-preview-arm-driver.mjs` is split out of the render rig at the boundary the readings module already names - the rig holds what each case claims, the driver how an arm is driven, the readings what it reports - since the three were over the 600-line cap together. No max-lines disable or bump. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): a fragment link is a frame navigation in this preview, so it is inerted too (round 3) pullfrog is right, reproduced on both engines before believing it. Round 2 kept `#`-prefixed hrefs on the theory that they are same-document scrolls. In this frame they are not: the document's URL is `about:srcdoc` while its base URL is inherited from the embedder, so `#section` resolves against the shell's own URL and the destination differs from the document's by more than a fragment - which makes activating it a frame navigation, and the shipped `frame-src 'none'` refuses it. Measured under the shipped policy, one tap, with something to scroll: Chromium 147 scrollY 0, frame becomes chrome-error://chromewebdata/, artifact gone, embedder reports frame-src <origin>/preview WebKit 26.4 scrollY 0, frame stays about:srcdoc and intact, same report So the destruction is Chromium-only but the absence of a scroll is not: there was no working affordance to carve out for, and the carve-out left a live link that destroys the preview - worse than the inert text it was meant to avoid. Both sandbox values behave the same, so this is the base URL and the policy rather than the sandbox. The same tap does the same thing on the granted path, where this pass does not run, so an artifact's internal links have never worked in the preview. That is not this change's to fix; it is recorded in `followup-html-preview-fragment-links.md`, and the render case reads the granted arm's violation as its presence precondition so the behaviour is pinned rather than merely known. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
7650abe224 |
fix(macos): tell the user when Orca's terminal service can't read their folder, and walk them through the fix (#21923)
* fix(macos): tell the user when Orca's terminal service can't read their folder
On macOS, a terminal daemon that survived an app update can be refused access to
a workspace under Documents, Desktop, or Downloads while the Orca app itself can
still read it. Terminals opened there die with "Operation not permitted" and
nothing on screen explains why. The daemon has reported `cwdReadableByDaemon` on
every create since #18043 and main has emitted `daemon_pty_cwd_denied` on proven
divergence since then; the field data says 1,438 users hit it in 21 days. What
was missing was the notice.
The verdict itself moves off `access()`. A grant-less probe on an affected
machine showed a TCC mode where `access(R_OK|X_OK)` passes on `~/Documents` and
`opendir` still fails, so the check now does what a shell listing its cwd does:
`opendirSync`, one `readSync`, `closeSync`. Only EPERM/EACCES reads as denial —
a missing path, a non-directory, or an unexpected error still reads as readable,
so a non-permission failure can never masquerade as one. The same probe is what
the app side compares with, through one oracle shared by the telemetry emitter
and the notice, so the spawn path reads the directory once.
Proven divergence now also records evidence in main: one entry, keyed by the
daemon's pid, start time and launch nonce, carrying an opaque digest of that
identity and the folder class. No path leaves main. The existing focus-time
`macTccAttribution` poll carries it to the renderer, which raises a second toast
latched per daemon scope: dismissed stays dismissed, and a restart mints a new
identity so the poll returns null and the toast clears with no post-restart
probe. If the replacement daemon is denied too, about 31% of cases, the next
spawn re-records under the new scope and the notice returns, now with the
re-allow sentence doing the work.
No new IPC channel, no daemon protocol field, no polling change, and nothing new
on the spawn path beyond one `opendir`. `daemon_folder_access_notice` counts
shown, dismissed and open_manage_sessions against `daemon_pty_cwd_denied` as the
denominator; `shown` is emitted from main the first time a scope leaves the IPC
handler, so the renderer carries no telemetry plumbing for it.
* fix(macos): clear folder-access evidence only when the same folder class reads back
A readable spawn in ~/code said nothing about a Documents denial but was
hiding the notice; retire the evidence only when the daemon reads a folder
of the class it was denied on.
* fix(macos): say what a terminal-service restart actually does
The Manage Sessions restart confirmation still described the product as it was
before agents resumed themselves: it promised panes showing "Process exited"
that the user reopens by hand, and mentioned legacy-protocol sessions nobody
outside the daemon code can act on. Open terminals and agents come back on
their own now, so the old copy made a routine remedy sound like data loss.
It also called the thing a "daemon". The same restart is about to be offered
from a user-facing fix dialog, so both surfaces now say "terminal service", and
the confirm button is just "Restart".
The new body adds the one fact the old one never stated: terminals on remote
hosts are not affected. Translations of the two changed strings are dropped so
the five non-English locales fall back to English rather than keep showing copy
that is now wrong.
* feat(macos): give the denied-folder notice a fix the user can follow
The folder-access toast told the user their terminal service could not read
Documents and then handed them a paragraph: restart from Manage Sessions, and
if that does not work, re-allow Orca in System Settings. Both halves were
guesses. Roughly a third of restarts do not fix it, and the user had no way to
know which case they were in before spending every open terminal on finding
out.
Main can now answer that. `daemon-folder-access-probe.ts` forks a short-lived
child of the app binary the same way the daemon itself is forked, runs one
opendir/readdir/closedir against the denied path, and prints a single JSON
line. macOS attributes a TCC grant to the process that forked the child, so a
child of the app running now answers exactly the question the running daemon
cannot: would a replacement daemon get in? The child goes through the shared
child-process wrapper, never a shell, with a 3s deadline, a 1KB output cap and
an environment scrubbed to PATH/HOME/TMPDIR. Every failure — timeout, bad
output, spawn error — reads as `unknown`, never as a verdict.
That answer rides out as `restartWillHelp` on the evidence the existing
focus-time poll already carries, and the toast becomes a title and two buttons:
Fix… and Not now. Fix opens a dialog with the two real steps. When the grant is
already in place, step one is shown as done and Restart is live. When it is
not, step one is open and Restart is disabled until it completes — which it
does by itself, because the poll re-probes while the answer is still no, and
returning from System Settings is the moment that lands. An unanswered probe
never accuses the user of a missing grant; it leaves both steps open.
Restart calls the management API directly rather than stacking the Manage
Sessions confirmation on top, since the dialog already states the consequence.
Success replaces the steps with a done line and takes the toast down; failure
says so inline and leaves the button usable.
System Settings opens through the existing developer-permissions pane opener,
which takes an id rather than a URL, with Files and Folders added to it. The
event's action enum now also counts fix_opened, settings_opened,
restart_clicked and — emitted from main when a replacement daemon's first spawn
lands in the folder class the previous one was denied on — whether the restart
actually worked.
* fix(macos): let the folder-access notice return after a poll that read no daemon
A daemon identity reads as null during any reconnect blip, and the poll reports that as
"no mismatch". The notice dismissed itself and then never showed again for that daemon,
because the once-per-daemon latch still held its scope. Only "Not now" should latch.
* fix(macos): say what the folder-access notice costs the user
One line read like a stray warning. The toast now says who is blocked and what fails,
and still leaves the steps to the fix dialog.
* fix(macos): give the folder-access toast one action and the X, like every other toast
"Fix" is the only button; the X dismisses. Sonner fires onDismiss for programmatic
dismissals too, so the post-restart takedown now goes through the store and the hook,
and only a user's X is counted as dismissed.
* fix(macos): keep the fix dialog's steps a checklist and put the one action in the footer
Buttons inside each step made the list look like a form, and a footer Close duplicated
the X. The footer now carries the active step's action, with a ghost Cancel; a probe
that could not answer says so under step 1 instead of showing a check.
* fix(macos): let the checklist show the fix landed instead of saying so
A hedged sentence addressed to the user read like chat. On success both steps check
off and the footer offers Done; the unanswered-probe helper is a status, not advice.
* chore(i18n): drop the fix dialog's unused close key
* Revert "chore(i18n): drop the fix dialog's unused close key"
This reverts commit
|
||
|
|
eb92222e7f |
feat: support Antigravity as supervised worker (#21705)
* feat: add supervised Antigravity worker support * fix: address Antigravity worker review findings * fix: stabilize Antigravity readiness detection * fix: allow Antigravity resume footer after readiness * fix(antigravity): make agy reach worker_done as a supervised worker Three defects each blocked `orchestration worker-start --agent antigravity --worktree new-child` at the agent_readiness stage. 1. Readiness never fired. The composer check required the trimmed line to be exactly one character, but agy 1.2.7 launches in accept-edits mode and paints it into the caret row (`> Accept-edits mode: ...`). Widened narrowly to a bare `>` or `> <name> mode:`; matching any `> <text>` would make every menu dialog read as ready, since they all prefix their highlighted row the same way. 2. No trust artifact for agy. Added markAntigravityWorkspaceTrusted, writing ~/.gemini/antigravity-cli/settings.json under `trustedWorkspaces` — verified empirically against agy 1.2.7, and distinct from the Gemini CLI's trustedFolders.json, which agy does not consult. Trust is exact-path and not inherited by subdirectories, so each child worktree needs its own entry. 3. The orchestration path skipped the preset. Orca has two trust dispatch chains: the renderer's preflightAgentTrust and the main-process markLocalWorktreeTrusted. worker-start only takes the second, which matched cursor/copilot/codex and fell through for antigravity, so the trust write never happened while renderer-side tests passed. Verified live end to end: the dispatch settles `succeeded` with worker_done carrying the right task and dispatch ids, and the worktree is appended to agy's settings with sibling keys untouched. Known gap: remote-agent-trust-presets.ts has no antigravity branch. The SSH artifact path is unverified, so agy over SSH still stalls at agent_readiness. Recorded in a comment there rather than guessed at. * fix(antigravity): wire trust preset through preload safely * fix: preserve Antigravity readiness across transcript tails --------- Co-authored-by: Neil <neil@stably.ai> Co-authored-by: LielinaH <lielinah@gmail.com> |
||
|
|
0677271709 |
fix(orchestration): reap leaked worker terminals via process-incarnation fallback — stops an unbounded PTY/process leak on Remote Server (OOM / cgroup PID exhaustion) (#18790)
* fix(orchestration): remint live handle from process incarnation on worker release When a durable terminal handle goes stale (rendererGraphEpoch fence), inspectWorkerTerminal re-mints a live handle via resolveTerminalHandleByProcessIncarnation + matchesProcessIncarnation so release/stop/read act on the still-running PTY instead of reporting missing and leaking the agent process tree. - keep main shared host-scope re-exports; add matchesProcessIncarnation - wire observation.terminalHandle through control/stop/release - rebuild release-completion on main structured paths - on missing/unattached + provably exited: settleDead fence first, then same-incarnation settleWorker fall back (archive may block settleDead mid-request); settle before recovery defer * fix(orchestration): derive SSH host scope from the reminted handle; reuse fresh-request recovery guidance for structured workers Addresses two open CodeRabbit review comments on PR #18790. inspectWorkerTerminal read the dispatch authority with the stale durable terminalHandle, so after a remint the lookup resolved nowhere and currentHostScope was always undefined — an SSH worker with no liveness verdict and no persisted host_scope got classified from terminal.connected instead of unverifiable. It now reads the same effectiveHandle every other observation in the function uses. stopStructuredWorkerForRelease told the caller to repeat the release with the same --retry-request, which only replays the stale release_unknown receipt and made a structured-worker close failure permanently unretryable. It now sources releaseUnknownRecovery from worker-release-completion so the fresh-request-ID guidance lives in one place. Pre-commit lint-staged (oxlint + oxfmt) run manually: clean. * test(orchestration): exercise incarnation recovery through runtime paths * test(orchestration): pin the incarnation read scenario to the reminted terminal The read scenario only asserted that the call resolved, so it documented nothing about which handle the read reached. Assert that the handle readTerminal received resolves to the registered pane and incarnation, so the scenario proves the read went through the reminted terminal instead of passing on the incarnation fence's throw. * refactor(orchestration): drop redundant incarnation prefix check; require liveTerminalHandle * feat: add freebuff as a first-class TUI agent (#42) <!-- orca-pr-loc --> <!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. --> | | Files | Added | Deleted | Net | | :--- | ---: | ---: | ---: | ---: | | Test | 0 | 0 | 0 | 0 | | Prod | 28 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$37 | 0 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$37 | <!-- /orca-pr-loc --> ## ELI5 Add Freebuff (`freebuff`) as a recognized first-class TUI coding agent in Orca alongside Codebuff and other supported agents. ## What Changed - Registered `freebuff` across shared TUI agent definitions, configuration catalogs, display names, and telemetry schemas. - Added agent icons, favicons, status mappings, and mobile asset references for Freebuff. - Added localization strings across supported language packs (`en`, `es`, `fr`, `ja`, `ko`, `zh`) and updated locale translation policy. - Documented Freebuff CLI in README agent table (`npm i -g freebuff`). ## Why Freebuff is a CLI coding agent twin of Codebuff (`npm i -g freebuff`). Adding it to the catalog enables users to launch worktrees, run automated sessions, and pick Freebuff directly within Orca. ## Linked Issue N/A ## Visual Proof `N/A` - Catalog registration and metadata definition for CLI agent launch; UI rendering uses existing TUI agent picker and status components. ## Testing - Verified TypeScript contracts, schemas, and catalog configurations. - Tested CLI detection / agent picker integration locally on Linux (`worktree create --agent freebuff`). ## AI Disclosure Assisted by AI coding tooling. ## Checklist - [x] This PR is small and focused - [x] I explained what changed and why (including ELI5) - [x] Before/after screenshots or videos attached for UI changes, or `N/A` with reason - [x] Self-reviewed for correctness, security, and performance - [x] Cross-platform, SSH/remote, and path/shortcut impact considered (or N/A) --------- Co-authored-by: Lesley Murfin <lesley@revivebusiness.ca> * test(orchestration): erase method overloads in worker reap fixtures * test: document worker fixture type boundaries * test: simplify worker fixture typing --------- Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: svc-orca[bot] <313947298+svc-orca[bot]@users.noreply.github.com> Co-authored-by: m4air <m4air@Mac.localdomain> |
||
|
|
88f2f01061 |
fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY (#19430)
* fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY Root cause: daemon-launched-child.ts forks the detached terminal daemon with detached: true, which escapes the POSIX process group (setsid) but never the systemd cgroup. Every PTY the daemon owns is itself an undetached direct child of the daemon (native-pty-spawn.ts). Under a combined systemd unit (Type=simple, KillMode=mixed, per docs/reference/headless-linux-server.md), a systemctl restart/stop SIGKILLs every process still in the cgroup at the stop timeout -- the daemon and every live terminal -- even though the codebase already has a fully-built adoption/reattachment path for a surviving daemon (orcad-entry.ts's refreshRestoredOrchestrationAuthority + reconcileLegacyWorkerTerminals, gated on daemonOwnsFreshPersistentPtys()). That path never fires today because the daemon never survives long enough. Fix: when systemd is actually supervising the process and the OS user has a reachable systemd --user manager (isDurableDaemonScopeSupported(), Linux only), launch the daemon via systemd-run --user --scope so it lands in a cgroup that is a sibling of the service unit's cgroup, not a descendant of it. A systemctl restart of the combined unit then never reaches it. Any failure of the scoped launch (no reachable bus, D-Bus policy rejection, etc.) falls back transparently to the existing plain fork() launch, so every platform/environment without this capability is unaffected. The daemon self-detects its own resulting cgroup scope via /proc/self/cgroup (detectOwnCgroupScopeUnit()) rather than trusting the launcher's intent, and publishes it as cgroupUnit in its pid record and orcad's health/readiness payload (health.terminalDaemon.cgroupUnit), so a running deployment can be observed to confirm the fix actually engaged. No new session registry is added: the existing daemon pid-record + adoption protocol (publishDaemonPidFile, daemon-pid-record-quarantine.ts's dead-record reclaim, refreshRestoredOrchestrationAuthority) already implements durable, crash-safe reattachment for a surviving daemon -- it was simply never exercised against a full unit restart before now. Proven via a systemd-in-Docker recovery test: a live PTY session's shell process, its daemon, and the daemon's cgroup scope were all confirmed unchanged across a real systemctl restart of a Type=simple/KillMode=mixed unit, while the main process pid changed (confirming the unit actually restarted) and the new process's health payload recognized the surviving daemon as adopted and live. A fresh write into the same PTY post-restart reached the same running shell. Ordinary terminal create/work/release and the #18789/#18790 worker-release reap-fix regression tests are unaffected. Fixes stablyai/orca#19408 * fix(daemon): probe the real per-UID XDG_RUNTIME_DIR before trusting the process's own env isDurableDaemonScopeSupported()/buildDurableDaemonScopeCommand() trusted the current process's own XDG_RUNTIME_DIR env var first, falling back to /run/user/<uid> only when that var was unset entirely. On mtl-02, orca-serve@factory.service's RuntimeDirectory= hardening directive makes systemd export XDG_RUNTIME_DIR=/run/orca_serve/factory into the unit's process -- a private scratch dir that shares the env var's name but has nothing to do with the user session bus. /proc/<pid>/environ on that host confirmed exactly that path plus DBUS_SESSION_BUS_ADDRESS=disabled:, while the real bus was reachable the whole time at /run/user/985 (confirmed via systemctl --user is-system-running with that dir exported by hand). The probe treated the hardened override as authoritative, found no bus socket there, and reported unsupported on every launch -- so the cgroup-escape fix from #19408/#19430 never actually engaged on real hardware, even though tonight's factory deployment picked it up. Fix: resolveUserRuntimeDir() now always tries the conventional /run/user/<uid> path first (computed independently via getuid(), never trusted from env), checking for a genuinely connectable bus socket via statSync(...).isSocket() rather than a bare existsSync. It falls back to the process's own XDG_RUNTIME_DIR only when that canonical path has no reachable bus -- covering hosts that legitimately have no /run/user/<uid> at all but do have a working bus wherever their own environment points. buildDurableDaemonScopeCommand() now explicitly sets XDG_RUNTIME_DIR to whichever path this resolution picked, rather than inheriting the spread env's (possibly hardened-wrong) value. Both isDurableDaemonScopeSupported() and buildDurableDaemonScopeCommand() gained an injectable canonicalRuntimeDir parameter (defaulting to the real computed path) so tests can exercise the hardened-override scenario deterministically with a real, connectable AF_UNIX socket fixture instead of the live host's actual runtime directory. Docker's stock jrei/systemd-ubuntu test container never had this hardening directive, so this gap was structurally invisible to the container-based verification in #19430 -- only caught against real mtl-02 hardware. * fix(daemon): report the daemon's own pid over the ready handshake, not systemd-run's The launcher used to infer the daemon's identity pid from the immediate spawned child (`child.pid`). On the durable-scope path that child is `systemd-run --user --scope`, not the daemon, so the launcher was asserting an identity it had no authority over. `DaemonReadyIdentity` now carries a required `pid` populated from `process.pid` inside the daemon itself, and `daemon-launched-child.ts` takes `launchedIdentity.pid` from that self-report. Both sides of the `holdDaemonAdoptionLease` pid comparison therefore originate inside the daemon process, which is the idiom this branch already uses for cgroup membership (`detectOwnCgroupScopeUnit` reads `/proc/self/cgroup` rather than trusting what the launcher intended). Note on the reported consequence: `systemd-run --scope` registers its *own* pid on the transient scope unit and then `execvpe()`s the target command -- same pid, no intermediate process -- so adoption did not in fact fail on systemd >= 206 (verified against systemd 255.4-1ubuntu8.17 and current main, `src/run/run.c` `start_transient_scope()`). The fix stands on its own merits: it removes a silent dependency on that exec-vs-fork implementation detail, which a `systemd-run` shim earlier in PATH or any future systemd change would have broken with no diagnostic. `terminateLaunchedDaemonChild` was audited and deliberately left on `child.pid`: for the same execve-preserves-pid reason that pid is either still systemd-run mid-scope-setup (killing it correctly aborts the launch) or already the daemon, so it targets the right process either way. Regression coverage: `daemon-launched-child-identity.test.ts` pins the identity source, and `daemon-ready-identity.test.ts` gains pid-validation cases. Ready-message fixtures across the `daemon-init-*` suites were updated for the now-mandatory field. Addresses: https://github.com/stablyai/orca/pull/19430#discussion_r3953722704 https://github.com/stablyai/orca/pull/19430#discussion_r3954346518 * test(daemon): assert cgroupUnit in the pid-file parse contract `parseDaemonPidFile` returns `cgroupUnit` on every branch as of the durable-scope commit on this branch, but five exhaustive `toEqual` assertions in daemon-health.test.ts still described the pre-scope shape, so they failed on the branch independently of any later change. Adds the field to those expectations. Deliberately not relaxed to `toMatchObject`: asserting the full parsed shape is what makes these tests catch a field silently dropped from the pid-file contract. * refactor(daemon): resolve the canonical user runtime dir at one point The per-UID path cannot change for a live process, so compute it once into a module const instead of threading the same default call through three signatures, and drop the try/catch around a getuid() that cannot throw once it exists. Trims the module prose to the non-obvious facts and corrects the pid-file record comment: an unscoped daemon writes null; only records no daemon wrote are absent. * test(daemon): clean up the cgroup-scope fixtures and assert a verdict The cgroup fixture tracked only the file it wrote, leaking one temp dir per case. Drains both fixture lists with splice so the pop-may-be-undefined guards go away, and replaces a not-throw/typeof-boolean pair with the verdict it was circling: no resolvable runtime dir means unsupported. * refactor(daemon): share the detached child options across both launch paths cwd, detached and stdio were repeated in the fork and systemd-run branches, which left the two comments explaining them hovering over the env block instead. Names them once so each branch carries only its own delta. * refactor(daemon): validate the ready pid like every other field typeof-first narrows the value, so the two 'as number' casts the isSafeInteger check needed disappear and the pid guard reads like the startedAtMs guard below it. * fix(daemon): don't retry the launch unscoped after losing the endpoint race A scoped attempt that lost the endpoint to another daemon was retried unscoped: a second doomed fork, a misleading 'cgroup-scope launch failed' warning, and the same DaemonEndpointUnavailableError the caller was already going to adopt on. Rethrows it instead, since no launch mode can win a race that is already lost. Also drops a private alias for DaemonChildSpawnOptions and the two 'as number' casts on child.pid in the startup-failure cleanup. * fix(daemon): unlink the pid record by the pid the daemon published The record holds the daemon's self-reported pid, so match on that rather than on the immediate child's, which is the systemd-run wrapper's until it execs. * fix(daemon): route the scope launch through the child-process chokepoint The two files this PR added imported `node:child_process` directly, which `child-process-import-boundary.test.ts` fails on deterministically: the offender count went 155 -> 157 against a pin of exactly 155. Raising the pin or listing the files is what that test explicitly forbids, and the allowlist's own note says a split "moved the import, it did not add one" -- so the fix is to get both new files off the module and put the count back at 155. - `daemon-cgroup-scope.ts`: the `systemd-run --version` probe now uses `runProcessSync` instead of `execFileSync`, so it gets the shared spawn decisions. Kept synchronous deliberately: `launchDaemonChild` attaches the readiness listener in the same tick it is called, and an await before the spawn moves the child past that tick. A non-zero exit is data rather than a throw here, so the verdict now checks `code === 0 && !timedOut`. - `daemon-launched-child-spawn.ts`: the scoped launch uses `spawnProcess`, and the long-standing unscoped launch keeps `fork` semantics through a new `forkProcess`. - `src/shared/child-process/fork-process.ts`: the fork arm of the chokepoint. `spawnProcess` cannot express a Node child with an IPC channel started from a module path under an overridden `execPath`, and the existing launch tests are written against `fork`'s contract, so a spawn rewrite would have changed module resolution, `execPath` and `execArgv` at once. It passes `windowsHide: true` -- the flag every other call site in that directory sets, reachable via an assertion because `ForkOptions` omits it -- which keeps `windows-console-visibility.test.ts` at its pin of 65 too. Both ratchets pass with both pins and both allowlists untouched. Docs: `orcad-operations.md` and `headless-linux-server.md` still described the limitation this PR removes as permanent. Both now describe the durable-scope survival path and its preconditions (systemd as PID 1, a reachable user bus / `loginctl enable-linger`, `systemd-run` on PATH), and scope the old text to the unscoped-fallback case, pointing at `health.terminalDaemon.cgroupUnit` as the way to tell the two apart on a running host. * fix(daemon): seal the cgroup capability probe from the host and correct KillMode=mixed docs The capability probe consulted the host's own /run/systemd/system marker and spawned the real systemd-run binary, so the hermetic unit tests could only pass on a systemd host (and fail closed otherwise, even with faked bus sockets). - Thread systemdBootPath and runVersionProbe as test seams through isDurableDaemonScopeSupported, defaulting to the real boot marker and systemd-run --version probe in production. - Narrow the injected probe to the ProcessResult slice it consumes. - Cover: no-systemd-boot, non-zero probe exit, and probe-timeout cases. - Correct KillMode=mixed semantics in the docs: the cgroup-wide SIGKILL fires the instant the main process exits, not after TimeoutStopSec; document the Docker-container caveat and add KillMode=mixed to the multi-service template. * fix(daemon): satisfy assertion checks in scoped launch * fix(daemon): satisfy anti-slop and console guards * test(serve): update shutdown docs assertions for daemon scope * fix(daemon): migrate adopted legacy scopes * docs: qualify restart safety by daemon scope * docs(daemon): qualify Upgrade restart prose with durable scope caveat Align the Upgrade section in docs/reference/headless-linux-server.md with the earlier preservation section and docs/reference/orcad-operations.md: a service restart terminates live processes only when running under the unscoped fallback, and stops should be treated as destructive unless health.terminalDaemon.cgroupUnit names an orca-daemon-*.scope. Update the shutdown workflow test assertion in config/scripts/headless-serve-shutdown-workflow.test.mjs to match. * fix(daemon): harden legacy scope migration --------- Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com> Co-authored-by: m4air <m4air@Mac.localdomain> |
||
|
|
15472cd4c6 |
feat(native-chat): keep restart recovery available in status bar (#21397)
* feat(native-chat): keep restart recovery available in status bar * fix(native-chat): source the restart offer from the host and retire it on recovery Closing the reconnect dialog spent the durable recovery offer, so looking around before deciding lost the recovery for good. The offer now survives a close, and the status bar carries it — but a durable offer needs a way to die, and it only had a reconnect, an explicit dismiss, and a 24h expiry. The claim's launch-scoped lifecycle moves into its own collaborator, which splits what the host ADVERTISES from the evidence it holds. A resume-capable hold that hands a marked chat its provider child back is the recovery the offer existed to perform, so it stops being advertised and stops being written back at quit, while the marker stays valid evidence — a user who reopened a chat can still ask the agent to carry on. Teardown re-derives the snoozed offer rather than round-tripping raw markers, and this teardown's own witness now outranks the stale claim for the same chat instead of being overwritten by it, which was silently persisting an old turn id and making the next launch refuse the chat that was actually mid-turn. On the renderer the candidate list gets its own producer against agentSession.restartResumable, so the status entry and the dialog read one host-owned answer instead of the dialog pushing its local state at a sibling. The entry re-reads the host before reopening, so a reopened list can never name a chat the host would now refuse; dialog open becomes the external one-shot request rather than a flag mirrored into render state, which is what let a reopen replay the launch answer and re-offer chats already reconnected. Dismiss all is quiet rather than destructive, saves the preference like every other exit, and reports a write the host never confirmed instead of trapping the dialog open. * fix(native-chat): keep a durable offer a launch never read, and settle the one a continuation spent Teardown replaced the recovery capsule with whatever this launch still owed, and a launch that never read the offer owes nothing — so a quit after a failed first read, a disabled flag, or a window that never mounted deleted a recovery the user was never shown. The write-back now distinguishes "claimed and still owed" from "never claimed": the first is re-derived as before, the second carries forward verbatim, because nothing revealed those sessions and the predicate would refuse every one for want of a journal nobody opened. Reconnect and continue spent the same claims Reconnect does but never shrank the offer, leaving the status bar counting chats the host had already handed back and sending the user to an entry that re-reads, finds nothing and does nothing. * fix(native-chat): stop a teardown answering for an offer it could not read Two ways the write-back deleted a durable recovery offer nobody had seen. A take that FAILED left the claim holding an empty list and reporting that this launch had answered for the offer. The markers were still on disk, unread and unknowable, and teardown then overwrote them with its own empty list. It now writes nothing at all unless it has a witness of its own. `owed()` read "has the capsule been touched" where it meant "did anything here LOOK at the offer" — and its own write-back read counted. Teardown is retried when a phase fails, so the second attempt re-derived carried markers against a session map eviction had already emptied, refused every one, and wiped what the first attempt had just carried forward. The flag is now set only by the paths that actually read or act on the offer. The mock guard for the carry could not fail: it indexed the session it claimed nothing had revealed, so re-deriving passed and the verbatim carry was never the reason it went green. It now runs against no indexed session, which is what an unread offer looks like. Also drops the `Not now` row from the preference table, where it was paired with a dismiss method it no longer calls, and asserts the same thing where the snooze is already covered. Splits the marker predicate's journal reader out of the resume host, which was at its line ceiling. * fix(native-chat): clear the corrupt recovery capsule the take refused A capsule whose contents no longer parse made take() throw before it ever reached the clear, so the bad file survived every launch. Nothing else rewrites it now that a teardown owing nothing readable declines to write, and the freshness filter runs after the parse, so the 24h window could not release it either: one corrupt file refused recovery forever. Clear it inside the same transaction that failed to read it, then rethrow, so the poison dies on the next launch while callers still see why the take failed. A clear that fails is swallowed rather than allowed to mask the parse error. Refusing to expose partial candidates is unchanged, and a read that fails for any other reason still writes nothing. * feat(native-chat): make resume the one restart action, and make it actually resume The restart prompt offered two actions: "Reconnect all", which reattached and sent nothing — exactly what opening the chat already does — and "Reconnect and continue", which reattached and asked the agent to carry on. The vacuous one is gone, the "Not now" button and the info popover with it, and the feature is now called resume throughout. "Don't ask again (resume automatically)" now runs the action the button runs: the launch calls agentSession.restartContinue instead of agentSession.restartResume, so the preference means what it says. Several comments asserted the opposite as a structural guarantee and are corrected. agentSession.restartResume stays: no in-app caller is left, but it is a published wire method a non-desktop or older client can call. * fix(native-chat): label the resume button with the number of chats selected The button read "Resume all" whenever every chat happened to be ticked, which described the selection rather than the action. It always acted on the selected chats only. Now it always names that count, with a singular variant so one chat does not read "1 chats". * refactor(native-chat): drop the reconnect vocabulary the resume action left behind Resuming became one action — reattach and ask the agent to carry on — so the notification helpers no longer need to be told which action they are reporting. Every caller passed `continue`; the `reconnect` branch, its helper and its catalog keys are gone. The dialog and the launch path had grown two copies of the same call: same RPC, same response shape, same announce-and-settle. That now lives once in the store module that owns the offer, which also takes the dismiss call, leaving the modal presentational. The two copies had drifted — only the dialog's caught a malformed payload — and the unified one keeps the defensive reading. No behaviour change. `agentSession.restartResume` stays: it is a published wire method even though nothing in the app calls it. * refactor(native-chat): derive the resume selection instead of intersecting it The modal's selection was intersected back against the host's candidate list before every action, as a guard against naming a chat the host never offered. That guard could never fire: the selection was already derived from that same list, so the intersection was the identity. The array of chosen ids is now the derived value and the lookup set falls out of it, which makes the property structural rather than checked. The helper had no other caller and is gone, along with its three tests. Three tests mocked the resume response in the shape the old API returned. Two never reached that branch at all; the third only passed because the unreadable shape happened to exercise the malformed-payload path. All three now use the real shape, and the malformed-payload behaviour — report an unconfirmed delivery, leave the offer standing — gets a test that says so. Also: the candidate reader took two trailing optional parameters, so one caller passed a placeholder `false` to reach the second; they are an options object now. `isFolderWorkspaceId` had no caller outside its own module and is no longer exported. `RestartActionOutcome` only ever describes a continuation row, so it is named for that. `dismissAll` set a busy flag that nothing could render, since it closes the dialog first. Several comments repeated an argument already made in the module they point at. Settings: the automatic-resume description is one sentence again. No behaviour change. * fix(native-chat): make restart recovery explicitly durable * fix(native-chat): preserve dismissal fence across new interruptions |
||
|
|
cd59678394 |
refactor(agent-launch): assemble host startup-plan inputs in one resolver (#22082)
* refactor(agent-launch): assemble host startup-plan inputs in one resolver buildAgentStartupPlan was already one shared implementation, but every host re-derived its argument object by hand from the same four settings (agentCmdOverrides, agentDefaultArgs, agentDefaultEnv, terminalWindowsShell), and the copies had drifted. resolveAgentStartupPlanInputs owns that assembly. What genuinely varies per launch stays a parameter: the host (platform, isRemote), a requested shell, the per-launch agentArgs override, and the picked session options. Fixes a live divergence on the agent.launch path: orca-runtime-create-agent-session passed sessionOptions without sessionOptionsOverrideAgentArgs, so a configured `--model` in agentDefaultArgs reached argv alongside the picked model and won on argv order, while the same launch through worktree.create honored the pick. The plan also reported no applied sessionOptions, so the chat surface could not name the model the user chose. Migrates the four host sites; the eleven renderer sites are unmigrated and still assemble their own inputs. * fix(agent-launch): preserve picked options in draft launches * test(agent-launch): assert draft option precedence |