mirror of
https://github.com/stablyai/orca.git
synced 2026-10-02 16:02:15 +00:00
e899809ff88a29af6b00a9dff38d54f1df1f120f
19
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a5ce8251e3 |
Agent launches carry the surface that started them (#23697)
* feat(agent-launch): every launch carries the surface that started it The host now attributes every agent it builds to the surface that asked for it, resolving a missing or unrecognized surface to 'unknown' in one place instead of silently skipping it. The CLI names itself on worktree.create and orchestration workers name themselves host-side. * fix(agent-launch): attribute the agent a startup-draft create launches The host builds a third kind of agent launch: a worktree.create with a startupDraft and no startupAgent, where the host picks the agent itself. It carried no launch record at all and ignored the caller's launchSource. Route it through the same resolver as the other two builders, and derive the startupAgent terminal record only from the resolver so no prebuilt record can stand in for it. * fix(agent-launch): attribute the agent a host-built agent session launches terminal.createAgentSession builds a fresh agent's launch on the host, like the other startup builders, but spawned it with no launch record, so those launches were never counted. Record them through the same resolver; the request names no surface, so they count as unknown. * test(agent-launch): require an attribution decision for every host-built agent startup |
||
|
|
7a24d3d335 |
fix(native-chat): the conversation outlives its agent (#22835)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * fix(native-chat): name a chat write by its target, not the owner generation A write carried the fence of the last frame the pane read, and the host refused it unless that fence was still current. An idle release and the restart after it each move the fence, and the release publishes nothing, so a send after a release was refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a cold start was refused as stale. Every write already names what it acts on: a send its conversation, a cancel its turn, a prompt answer its item revision, a rewind its epoch; an option is last-writer-wins. So admission stops comparing the client's fence, and the rebase that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it. The writer-lease check stays, and so does the attach's compare-and-swap. Frames now stamp the fence read when each frame is sent instead of a copy each subscriber kept, which went stale on the same release. * fix(native-chat): every journal append reaches the chats that are open A journal write and its delivery to open readers were two calls, and some writers made only the first. A failed start whose lease could not be handed back, a provider revision with no frame behind it, and eviction's settlement were all journaled without reaching an open chat. A journal handle now reports every durable change, and the host's session map binds that report to the session's readers when the handle is set. Writers no longer publish what they append; the per-writer publish calls are deleted. * test(native-chat): an epoch replacement reaches the open chat * test(native-chat): each row reaches an open chat once, and a live handle enters only through the map * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. * perf(native-chat): a publish behind a delivered commit reads nothing Each commit now delivers itself, so the publish a provider frame still sends afterwards found every reader caught up but still read rows and rebuilt the timeline for each one. A caught-up reader now skips the read. * test(native-chat): state why the teardown test's fake journal is safe to cast * docs(native-chat): say mutation admission checks only the writer lease * docs(native-chat): drop the send rebase from comments that still described it * fix(native-chat): a message is accepted, then delivered A send to a chat with no running agent restarted the agent inside the send call, before the message was recorded, so the client waited for the whole start and a failed restart refused the message. Claude held prompts sent during startup, and those could settle as "unconfirmed". A send is now accepted inside the session's serialized queue: one ledger row and one submission row marked handoverRecorded, published, answered pending. A per-session delivery loop exists while a message is queued. It starts the agent through the same serialized attach a hold uses, waits outside the queue for a Claude child to prove its start, and hands the oldest queued message over as its own serialized step, writing dispatch{pending} before the adapter call. A start it needed and did not get writes one error-tone row and rejects every queued message with the same words; a start Stop cancelled writes none. Settlement follows from the rows. A queued message is provably unwritten, so a close, an eviction or an exit rejects it. A handed-over message stays in doubt. A queued row at or below the sequence a handle found when it opened was left by an earlier process and is rejected at open, with no latch. Stop withdraws queued messages with no writer lease and no fence. An attach failure keeps the conversation open, and the attach adopts its journal. Owed work counts the loop and queued rows. A compaction or rewind found prepared when a conversation opens was started under a child this process no longer has, so the open settles it rather than leaving it to refuse every send until a view attaches. The open cursor is scoped to its epoch, because sequences restart when an epoch is replaced. Deleted: restart-before-admission, recordFailedRestart, the fence rebase, Claude's startup gate, the attach's forget on failure and its own crash boundary. Clients without agent-session.accepted-send.v1 get their reply held until the handover; the desktop and paired desktop lists advertise it. * fix(native-chat): settle queued messages only for the child that ended A child that proved its start and then exited before its message was handed over left the message queued: the exit settlement returned early when nothing else was in flight. Delivery then started another child for it, and a child that died the same way started another, without end and without a row. A retried settlement for an earlier generation, run by the attach that delivery started, did the opposite: with that generation's turn unfinished it rejected the message queued for the child being attached. The settlement now takes the rejection for queued messages from its caller. The unexpected exit and the eviction pass one, and it applies even with no other work in flight; the retry for an earlier generation passes none. * fix(native-chat): an adoption that fails to import keeps the conversation open The attach now writes into the conversation's own open journal, but a failed transcript import still closed it as if it were the attach's provisional one. The conversation stayed indexed with a closed journal, so every later send answered "could not be recorded" and every attach failed again until the app restarted. The import now closes only a journal the attach opened for itself. * perf(native-chat): the recovering open reads the journal once Every conversation open now goes through the recovering open, including the read restore of every chat at startup, which used to replay its journal once. The recovering open replayed it twice: once to probe it and again inside the open. The probe is now handed to the open as its load. * fix(native-chat): an attach that fails after indexing its child leaves no child behind A failed attach now keeps the conversation open, but a failure after `onAttached` indexed the child (the rewind or compaction recovery, or the attach's own success record) left that entry claiming a child the failure path had already released. The next send found the phantom, skipped the start, and wrote at a fence the journal had moved past, so the message stayed queued for good. The entry now drops the released child and its event sink, and follows the record's fence, as a failure before indexing already did. * fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer The error strip for a message the host accepted and then did not deliver matched the entry before the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not send your message" with nothing to retry. It now reads the reconciled entry. A rejection the journal records before the send's own pending answer lands is final as well: that answer no longer puts the entry back to dispatching with no Retry. * fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down The preamble waits for its submission to be delivered while the worker's agent starts. When that wait ran out it threw operation_unknown, and the failed-start teardown then closed the session, which rejected the very preamble the host was about to deliver. It now reports a turn start nobody observed yet: the worker is start-unknown with its session kept, the host delivers the preamble when the agent starts, and the worker's report settles the dispatch as for any unobserved start. The receipt no longer suggests reading a screen a structured worker lacks. * fix(native-chat): a message rejected while its chat was closed reads as not sent A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked every later message behind a Retry and no reason, and the delivery probe, seeing the journal already answered, never ran. The reconcile now settles it as rejected like a dispatching one. * test(orchestration): name why the readiness settlement fakes are cast * fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent * docs(native-chat): drop the fence from the admission the send effects run behind * docs(native-chat): give the fence move on release the reason that still holds * docs(native-chat): stop citing a write fence check in launch and mailbox comments Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease. * refactor(native-chat): the provider child is its own record A conversation now outlives any number of provider children, so the child is one record on the conversation's entry instead of five loose fields beside its journal. It is written in one place: indexed only once an attach has fully succeeded, and ended through one function that an exit, a failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence. - A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence patch after it are gone. - Conversation writes read the record's fence, the way mutation admission already does; a child's own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of the conversation's fence, are gone. - The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer dropped when an attach replaced the whole entry. - Stop on a child still proving its start stops only the child: its lease goes back and the chat is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus the conversation's close. - The settlement retry uses the conversation's own journal, opened through the host's one open. * fix(native-chat): the delivery loop alone settles a message its start or child failed A queued message was settled by whichever path happened to end the child first: the loop, the unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that rejected every pending row. That gave two failure rows with different tones for one start, a loop that could hand over to a different child than the one it waited on, and a Claude start that died while starting reading unlike every other failed start. - The loop remembers the child it waited on. At handover, if that child is gone or replaced, it reads how it ended: a Stop continues; anything else writes one failure row and rejects every queued message with the same words, then stops. A child still starting whose start the adapter says did not land fails the same way. The exit, eviction and the settlement retry only settle the handed-over and legacy rows of the child that ended. - One failure row, always an error, keyed by the start. A start a view began that dies with nothing queued writes the same row through the same builder, so a second report revises it. - The open no longer rejects leftovers; the loop's first step does, and the open wakes it. - `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the failure before the exit is processed. - Quit closes every conversation the way closing a chat does: what is still queued is rejected as closed, with or without a child, and a start the loop already has in flight is waited for so the child it produces is stopped rather than left behind. * refactor(native-chat): a stopped child ends on the one reading of its stop The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that verdict to the child's ending, so the host never forms a second view of whether the root is gone. Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition, and a failed re-attach passes what its release saw. The end-of-child record can therefore also carry a stop whose root was not seen to go, which nothing ends on yet. * feat(native-chat): the host says it accepts a send before any agent has it The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same string capable clients already send. A client can then tell a host that answers a send at acceptance, and admits a Stop with no writer before a turn starts, from an older one that still restarts the agent inside the send. Additive: an older client ignores a capability it does not know. * refactor(native-chat): an attach never opens a journal of its own The attach adopts the conversation's open journal, which outlives it, so it no longer opens one for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag that told the two cases apart is gone. Tests that attach without a host open the conversation the way a host does. * fix(native-chat): a moved fence resends nothing on a host that accepts first The outbox treated any fence change as a new owner: it dropped the answer of a send in flight, queued that send to go out again under the same id, and unblocked a refused head. On an older host that is how a send the restart refused, unrecorded, gets another try. On a host that records every send before it starts an agent, a fence moves because that start ran, so the same rule resent into every failed start. With a fence stamped on every frame, that became a loop. The outbox now reacts to a fence change only when the host has not advertised that it accepts a send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed start reaches the client as a rejected message it keeps with its Retry. Against an older host, or before one has answered, the outbox behaves as it did. Desktop and paired web share this hook. * refactor(native-chat): a child's end says whether the user or the host stopped it The end-of-child record's cause now tells a user's Stop from the host stopping the child for a cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a user's Stop, as before, and fails the start it was waiting on after a host stop, with the one error row and every queued message rejected, in the stop's reason when it gave one. The reason stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet. * fix(native-chat): a chat whose only work is a queued message is not offered for resume A message accepted while the agent was starting counts as working in the chat, and quit rejects it as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a chat whose agent never had the message. The snapshot now reads only what was handed over. * fix(native-chat): the conversation outlives its agent Opening a chat no longer starts its agent. A conversation is reached through one host accessor that opens its journal at rest, and a send is what starts the agent, through the delivery loop. One idle sweep, every five minutes, stops an agent that has been quiet for thirty minutes and owes no work, then drops an open journal handle that is only a cache. Its record, tab, status row and readers stay. - hold and release are no-ops; hold still builds the host for shipped mobile builds. - The holders, the holds, the release clock and the exit respawn are deleted. - Options, the model list, the goal and the context meter answer at rest; a model pick at rest is recorded as intent for the next start. - Compact, rewind, clear and goal changes start the agent first. A send does too when a rewind is still in doubt after the conversation opens. - Orchestration routes mail and group addresses on ownership (the record plus the chat tab), not on whether the process runs. An open dispatch keeps its worker running. - The restart continuation is a send; Resume all holds each slot until the message is handed over or rejected. - A read error never replaces a loaded transcript, and shows the host's own words. * test(native-chat): type the queued-message fixtures in the resume-offer tests * fix(native-chat): a start that dies while a message waits on it is that message's failed start Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When that start died, its exit wrote the start's error row and left the message queued, so the delivery loop started a second agent into the same failure and wrote a second row. A child's end now records where the conversation's journal stood, and the loop settles a message accepted before a failed start ended with that start: one row, under its key, and no second start. A message sent after the failure still gets a fresh start. * docs(native-chat): say what an attach's open conversation and unconfirmed ids are now * test(native-chat): pin what a failed start settles, and what a resume offer names A view's child that dies while a sent message waits settles that message only when it died starting and no child has taken its place: a proven child's crash, or a second start since, gets the message delivered. The resume offer names the handed-over message, never a newer one still queued. * test(native-chat): the failed-start pins fail on what the message became, not on a timeout * fix(native-chat): a restart offer ends when the chat's agent starts again The offer used to end only when the chat's newest user message changed, because opening a chat started its agent and that start could not be told apart from real activity. Opening a chat starts nothing now, so the host reads the fact it already publishes: a chat's status row goes from not host-owned to host-owned exactly when its agent is started. At that edge the offer and any failure record for the chat are withdrawn, unless the start is a resume action's own (its continuation is the oldest undelivered message). A continuation and a message racing to be first are decided at acceptance: the continuation is refused, quietly and with nothing filed, when any other message was accepted since the restart. A failed continuation start leaves the offer retryable, and each resume action sends its own message id. Deleted: the newest-user-message comparison, its journal reader, the continuation filter, and the failure ledger's own "answered by the chat" check. The marker still carries its message id for one release, so the previous build can read it. * fix(runtime): end a transcript stream when its client unsubscribes Desktop: the IPC subscription controller was dropped as soon as the streaming handler returned, which for most streams is right after it binds. A later runtime:unsubscribe then found nothing to abort, so the host kept the subscriber and derived and sent every publish to a channel no one listened to. The controller now lives until the renderer unsubscribes, resubscribes the same id, or goes away. Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe with the stream's frame id, so the host ends that subscriber and leaves a sibling stream on the same socket running. The direct path now passes the frame id the relay path already passed. * test(orchestration): the preamble's host stub is typed, not cast The preamble send now takes only what it reads of the host, the send, the settlement wait and the record's fence, so its test builds that host with real types instead of `as never`. * fix(native-chat): one fact ends a restart offer: the chat moved on since the restart The offer is live while no other message has been accepted in the chat since the restart and its agent has not proved a start since. The offer list, the resume's reservation check and the continuation's acceptance check all read that one fact, so a message whose start then failed withdraws the offer too, and a stale click finds nothing to act on. The fact is read off the conversation's open handle, which the restart closed, so it is retired durably whenever it may have changed: a message accepted, a start proven. A close and reopen within the same run therefore cannot bring the offer back. A continuation rejected before it reached the agent does not count, so a retry after a failed start still runs. Deleted: the quit-time gate on withdrawal, which changed nothing because the withdrawal and the quit's own offer write share one queue; the per-action "withdrawn" flag and the separate acceptance check it paired with. * test(native-chat): an older build reads the restart offer this build records The offer lives in a file the previous release reads after a downgrade. Pin that against the pinned release's own capsule, and run the lane when the marker or the capsule changes. * fix(native-chat): read a restart offer against where the journal stood when it was taken "Since the restart" was read off the conversation's open handle, which the idle sweep closes: after a reopen, a message the user had already sent looked older than the handle and the withdrawn offer came back. The offer now records the journal position (epoch and sequence) at the moment it is taken, and a message accepted after that position, or a journal on another epoch, means the chat moved on. That is derived from the journal, so it holds across any number of closes and reopens. An older build's offer has no position; only a start withdraws it. Because the message half is now durable, the offer is no longer rewritten in the recovery file on every accepted message; a proven start still writes it, since only the host that saw the start knows of it. * test(native-chat): wait for the listing's retire write before reading the recovery file * fix(native-chat): keep the terminal-backed chat's read error over its local echoes Messages winning over a read error is right for the structured chat, whose read retries and whose messages came from the transcript. The terminal-backed view assembles its list from local echoes too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no error. Only the structured pane now keeps messages over an error. * fix(native-chat): a start retries the exit settlement a failed journal write left owed An agent exit whose journal settlement write failed releases the lease latched until a retry lands. Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried it before the next app launch, and every send was refused. The start the send needs now runs the retry first, where the attach would. * perf(native-chat): answer the owner check without opening the chat Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the answer comes from the session record alone. Reaching it through the accessor opened each resting chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it open for the idle window. It now checks the record and the adapter's support, as before this series, and opens nothing. * fix(native-chat): a read waiting on the session lock opens nothing once quit began The accessor checked for quit before queueing the open, so a read queued behind a session task ran its open after teardown had begun and indexed a journal no teardown step would close. The check now runs at the open itself. * fix(native-chat): read a failed resume's chat before calling it retryable Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the restart, read from its journal. The failure list read it only for a chat already open, so once the idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did nothing. The list now opens the failed chats first, as the offer list does. * test(native-chat): type the provider event sink the settlement test reaches for * fix(native-chat): say the structured read keeps trying only where it does The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an untranslated fallback whenever the read error had no text, and the empty state prefers any message. The view state now leaves the message out, so the structured pane shows that line and the terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an error frame, so it no longer makes the claim. * test(native-chat): await the send's settlement instead of polling for the start The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a loaded machine outran. They now await the host's own settlement of the message. * fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now dated by the resume action. Telling a rejected continuation from the user's own message read the operation ledger, whose rows expire after about a day; after that a failed resume stopped being retryable. The offer now records the continuation each action sends on its own capsule entry, bounded to the newest 16, so the ids end with the offer. The ledger read is deleted. * fix(orchestration): route no mail to a structured worker its orchestration released A structured worker is routed on ownership, and a resting worker's lease is released, so ownership held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one restarted its agent. Routing now also reads the orchestration's own resource row: once it is released, direct mail, group addressing and worker-show's addressable answer drop the worker, as they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored. * fix(native-chat): a failed retry names the user's prompt, not Orca's continuation A resume's continuation is written to the chat before its start, so after a failed attempt the chat's newest user message is that rejected continuation. A second failure then showed Orca's own restart text as the chat's prompt. A retry now keeps the prompt its first failure named. * fix(orchestration): read the released row optionally, as the authority does worker-show's observation called the row lookup directly, which a runtime double without it threw on and failed the structured tab-retirement release. * fix(native-chat): the status bar drops a restart offer the chat moved on from The renderer re-read the host's restart offer only when a failed chat showed activity, so after a message withdrew a pending offer the host answered no chats while the status bar kept counting one, and clicking it opened nothing. The same watch now covers pending offers: a status change in an offered chat asks the host again, once. * test(native-chat): a roster of idle or finished children does not keep an agent awake The sweep reads owed background work through the shared child-work liveness that upstream's release clock adopted; a child that went idle or finished is not work the agent still owes. * fix(orchestration): a task dispatched into a resting structured worker keeps it running The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's process incarnation now counts, derived from the existing rows. * docs(native-chat): comments stop describing the hold this PR removed Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it. Comment-only. * fix(native-chat): a restart offer keeps the start its own continuation made Whose start ended an offer was decided at read time, from whether the offer's continuation was still the queued message. Once the provider refused that continuation, the child it had started read as someone else's start, so the offer ended and its failure showed no Retry. The delivery loop now records which queued message a start is for on the in-memory child, and the child's end carries it; the offer counts a start as its own when that message is one of its continuations. * fix(native-chat): an agent gets a full idle window after its owed work ends The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can read done before the lead's wake-up turn writes anything, and stopping in that gap loses the wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full window afterwards, as the release clock it replaced did. * test(claude): the options-read fixture runs a live child The fixture marked its conversation running with a hasProviderChild field the session type does not have, so the read took the at-rest path and refused a session with no record. It now carries a child, which is what the read checks. * test(native-chat): host tests reach its collaborators through a typed seam The rest-test rig and three test files read the host's private members with Reflect.get and cast the result. The host now exposes one test-only accessor, collaboratorsForTests(), and the subscribers class a subscriberCountForTests() beside its existing retainedActivityCountForTests(), so the tests are checked against the real types and the casts are gone. * refactor(orchestration): one owner answers a structured worker's custody Routing, group addressing, worker-show and the idle sweep each composed their own reading of whether orchestration still holds a structured worker, so each new obligation or retirement state had to be added to every reader. structured-worker-custody now derives both answers from the worker-terminal list state coordinators see in worker-list: addressable is owned and not released, and owed work is an active custody or an unsettled task dispatched to the same incarnation. The owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests. * refactor(orchestration): owed work is an open dispatch on the worker's incarnation A supervised worker's own dispatch context stays open exactly while the worker is active, so the separate active-custody branch only repeated it. Owed work is now one fact, which also states the policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are written once at the top of the module. * fix(native-chat): a restart offer knows its continuations by a tag in their id The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running action's id in memory. Both could disagree with the journal: past the cap an old rejected continuation read as the chat moving on, and a crash during a retry restored the failure's older entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer (its teardown and chat), then the action's own part, so any continuation of this offer, queued or rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own continuation the start was for, read against the stored marker. * test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait The tui-idle probe reads through readTerminal, which now awaits the structured worker check before the PTY read, so the probe's snapshot request starts a microtask later. vi.waitFor missed it on its first check and polled again at 50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot then resolved after the wait had already timed out, so the test passed without judging it, and the rejection landed before any handler was attached. Vitest reported that as an unhandled error and failed the shard. Polling every 1 ms sees the request within a few ms, so the snapshot is judged while the wait is still pending. * fix(native-chat): a view never restarts a chat whose last start failed A Claude chat whose CLI exits during startup left one red row per start, and every time a view bound to it (the chat opening right after its create died, or the user switching back to it) the hold started the CLI again, so the same launch-failure row repeated. Only a send retries a failed start now, the same rule provider-exit recovery already applied; the rule lives in one predicate the hold, exit recovery and the delivery loop share. * fix(native-chat): the idle sweep reads owed work every tick Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it refreshed the clock at most once a window. Work that ended just before the next read left the agent to be stopped at that read, moments after the work ended, which is the gap the refresh was meant to cover. The sweep now reads owed work on every tick for a started agent, so the window always runs from the last tick that saw work owed. * fix(native-chat): a continuation handed to the agent stays sent The offer read its own continuation as not reaching the agent while its dispatch was pending, which also covered one already handed over and still unanswered. When the wait for that answer ended first, the failure it filed read as retryable, and a retry sent a second continuation to an agent that may have acted on the first. Only a continuation still queued, or rejected, is now read as unsent. * test(native-chat): start the child the loop waits on with an attach, not a second view A view no longer starts a child whose last start failed, so the R2 case that waits on a child started since the failure now gets that child from a client attach, the one non-send starter left. * fix(native-chat): settle a gone generation's turn wherever a conversation opens A send that opens a chat this process had not read yet (after a crash, from a phone or the CLI) went through the delivery open, which never settled what the dead generation left running; only the read restore and a successful acquire did. When the send's start then failed, the turn stayed running for every reader. The settlement now runs in the one journal open, at the crash boundary, for every opener except an acquisition, which settles from the evidence it read before its reserve; the read restore's separate step is gone. * test(native-chat): prove the next child's start settles the turn an earlier child left The R1 case lost its only settlement assertion when the latch it checked was deleted. It now seeds the running turn the earlier child left and asserts it ends at the exit's receipt, with the exit's row, before the message is handed to the new child. * test(native-chat): count a failed start's rows by row, not by text Comparing the set of texts passed when two different rows carried the same words, which is the duplicate the test exists to catch. * test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget * test(native-chat): the interrupted create's own retry continues again The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its own operation id, with a fresh start whose result nothing read. That fresh start passes with the released-reservation continuation deleted, so the case the fix exists for went untested. The retry and its assertion are main's again. * docs(native-chat): three comments that still had views starting agents A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction left alone would refuse every send, so no agent would ever start to finish it; and a current host raises the unattached read refusal only once quit began, with the attach window belonging to an older host. * test(native-chat): pin the open's and the send's start and row counts, however the view binds Opening a fresh chat whose starts fail makes one start and one row, with two views bound before or after the create's child died; one send makes one more of each. * test(native-chat): a reader's open settles the turn a failed exit settlement left running An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle sweep closed, and a read that opens the chat before the restart restore reaches it. * test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner A subscription reads the conversation before it returns, so under load the two views took longer than the create child's 300 ms start, which then exited before the test checked that it had not. The child now takes a second to fail. * fix(native-chat): settle a gone generation's turn at every open but an acquisition's The journal open skipped the settlement whenever the lease read reserved or live, to leave an acquisition's own open to the acquisition. But a lease a crashed process left in recovery also reads live, until the next acquire resolves it. A send that opened such a chat, from a phone or the CLI after a crash on a host that could not prove the old owner gone, skipped the settlement; when its start then failed, the dead turn stayed running for every reader. The acquisition now says it is the opener, and every other open settles, whatever the lease still claims. * test(native-chat): hold the create's start open until the views bind The "view binds while the create is still starting" case gave the create a 300 ms head start and asserted the views bound before it died. On a loaded runner the holds took longer, the create's exit landed first, and the case failed its own precondition. The create's initialize now waits on a gate the test releases once the views are bound. * test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the dead process's lease still reads live. The open settles the turn it left running anyway, and the restore that follows finds it settled. * refactor(native-chat): drop the composer's second error formatter After the merge with main, every chat write in the composer path reports its failure as a typed outcome worded by the refusal-notice table, so the send's catch sees only a local throw. The {code, message} formatter this branch added for it has no payload left to format, and its claim to be the one way a chat words a failure is no longer true. The composer send is main's again. * test(native-chat): pin the reason on a message rejected while its chat was closed The reopen test checked only that the message reads as not sent; it now also checks the Retry row carries the host's reason. * docs(native-chat): drop the removed dispatch hold from six comments A worker's session no longer takes a dispatch hold, and no release clock rests a chat by visibility; the agent-launch comments, the abandon test, the teardown test and the refusal census still said so. * test(native-chat): rest the owner-status chat through the idle sweep, not a hold The activation-gate test from #22808 put its chat at rest by holding and releasing it, and passed the release-clock grace. This branch deleted both, so the case threw before it reached its assertions. It now moves the host's clock past the idle window and lets the sweep stop the agent and close the conversation, then asserts the same owner answer and activation gate. * fix(native-chat): show the structured pane's retrying line when a read fails The read transport always hands the pane the host's words, so the error state's "Orca keeps trying to load it" line, which showed only when there were none, was never seen: the pane showed the host's text twice, as its subtitle and on the status line under it. The structured pane now always says its read keeps retrying, and the host's text stays on the status line. The terminal-backed chat is unchanged. * test(native-chat): wait for a send's background start before the refusal oracle removes its store An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals. * fix(native-chat): a start a message waited on gets one failure row, the delivery loop's When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice. The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
89cf55dfc8 |
fix(agent-launch): report a launch that failed before spawning as failed, with its cause (#22913)
* fix(agent-launch): settle a launch that failed before spawning as failed, with its cause A launch into an existing workspace whose terminal create threw before the spawn request left this process (agent disabled, no launch command, runtime unavailable) created nothing, yet agent.launchReplay recorded and answered it as agent_session_operation_unknown. createTerminal now reports when it hands the spawn to the pty controller; a failure before that point settles the ledger row as failed and returns the original error. After the request leaves, the outcome stays unknown: an SSH or daemon spawn whose reply was lost may still have started. * test(agent-launch): expect the spawn-dispatch hook on the launch's terminal create * test(agent-launch): drive the pre-spawn failure with a missing launch command A disabled agent is moving to a check made before either launch route runs, so the tests use a failure that stays inside the terminal build. * fix(agent-launch): keep the not-started verdict on the launch, not the shared error A failed pane spawn rejects the same error object into the spawner (after its request left) and into a concurrent create waiting on that pane (before its own). Marking the error object globally let the waiting launch's verdict clear the spawner's, recording a launch that may have started an agent as failed. The launch now owns its dispatch tracker and carries the decision on its execution error instead of re-deriving it from the error. Also names the test's launch parameter type for the anti-slop audit. * refactor(agent-launch): move launch failure classification into its own module Rebasing onto the caller-selection change took agent-launch.ts past its line limit; the failure-code helpers are a self-contained concern. |
||
|
|
16784c1a67 |
fix(native-chat): name a chat write by its target, not the owner generation (#22812)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * fix(native-chat): name a chat write by its target, not the owner generation A write carried the fence of the last frame the pane read, and the host refused it unless that fence was still current. An idle release and the restart after it each move the fence, and the release publishes nothing, so a send after a release was refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a cold start was refused as stale. Every write already names what it acts on: a send its conversation, a cancel its turn, a prompt answer its item revision, a rewind its epoch; an option is last-writer-wins. So admission stops comparing the client's fence, and the rebase that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it. The writer-lease check stays, and so does the attach's compare-and-swap. Frames now stamp the fence read when each frame is sent instead of a copy each subscriber kept, which went stale on the same release. * docs(native-chat): say mutation admission checks only the writer lease * docs(native-chat): drop the send rebase from comments that still described it * fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent * docs(native-chat): drop the fence from the admission the send effects run behind * docs(native-chat): give the fence move on release the reason that still holds * docs(native-chat): stop citing a write fence check in launch and mailbox comments Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
1069bb053f |
fix(agent-launch): a cwd at the workspace root no longer forces a terminal (#22729)
"Continue in New Session…" always names a cwd, and both the renderer route input and the host launch-mode decision read any cwd as a custom start directory, so the continuation opened a terminal agent even when chat was the user's default. Both now share one rule: only a cwd outside the workspace root (after normalising slashes, Windows case, WSL aliases and the distro's Linux spelling) requires a terminal. A subdirectory still does, because a structured session cannot start there. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
f0b3f44f10 |
feat(agent-session): let the host own a chat's tab id and let a create reserve it (#22616)
* feat(agent-session): let the host own a chat's tab id and let a create reserve it A structured chat's tab id was derived from its session id by every layer that needed one: the renderer, the host snapshot and the status address each built their own spelling. The join between a conversation and the tab that shows it must be a pointer the host owns, not a derivation each client repeats. The session record now carries surfaceTabId. A create pins it: the tab half of the pane agent.launch reserved, an optional tabId on agentSession.create, or a host-minted UUID. Records written before the field existed are backfilled at open with the string clients derived, in memory at once and on disk with the store's first transaction, so nothing keyed by it (read state, notification ids, worker rows) moves on upgrade. A second record under a held id is refused. Only the record and the two create wires change here. The snapshot still publishes agent-session:<sid> and the renderer still derives its local id; those move in the next two changes. agentSession.create is a strict object, so the field is advertised as a capability a client checks before sending it. * fix(agent-session): record the derived tab id for an unreserved create A create that reserved no tab minted a random UUID that no reader uses: the renderer, status address, worker rows and host-shared read state all still key by structured-agent-session-<sid>. Persisted, that id would move every chat created before readers switch to the recorded one, orphaning its read state and worker rows the way the backfill exists to prevent. An unreserved create now records the derived id, the same rule the backfill applies, so the record always matches the prefix every existing key uses; an opaque mint belongs with the change that moves the last reader. Also: - a chat tab id must be a host tab id on the record, the create wire and in admission, matching what agent.launch already requires of paneKey; a web-surface id would decode as another tab - the stored launch-result guard checks the structured outcome's tabId - comments no longer claim a retry naming another tab conflicts; replay keys on the attach fingerprint and answers with the recorded id (now pinned) - the wire refusal test used a non-hex digest, so the schema refused it for that reason; it now reaches the tab id rule - pin that the reload path refills the id without forcing a save * test(agent-session): correct the tab-id fingerprint comment to match replay |
||
|
|
069dc8a1d8 |
feat(agent-launch): let a caller reserve the chat session, and start terminal launches with the session picks (#22523)
* feat(agent-launch): let a caller reserve the chat session and carry session picks to a terminal launch * fix(agent-launch): keep a caller-minted session id named for its agent, and mint the fallback the same way * test(mobile): model the older host from the launch fields, not the refined schema * docs(agent-launch): describe the reserved session id as conversation identity, not placement The caller mints the session id so it knows which conversation it started; tab placement is not keyed on it. Also puts the terminal surface's doc comment back on createTerminalSurface. * docs(agent-launch): say a terminal launch reads the session picks on the wire contract The `sessionOptions` field doc still said a terminal launch ignores them, which this branch changed. * fix(agent-launch): check a reserved session id's token after the agent name, not the whole id A hyphenated agent name failed the one-token check, so any session id for such an agent was refused at the wire, while every other agent without a chat has its id ignored on the terminal. |
||
|
|
a375936c04 |
feat(agent-launch): let a caller reserve the pane its terminal launch creates (#22291)
* feat(agent-launch): let a caller reserve the pane its terminal launch creates * fix(agent-launch): refuse a launch whose reserved pane is already live * fix(agent-launch): refuse a live reserved pane before it is revealed The live-pane refusal used to fire in the executor, after createTerminal had already issued a handle, published the mobile snapshot and revealed the tab. The reveal re-registered a fresh launch config over the running agent's. agent.launch now passes requireFreshPane with a reserved pane, and createTerminal throws AgentLaunchPaneAlreadyLiveError as soon as spawn reports it attached to a live pane. That is before any handle, snapshot or reveal. The spawn reattach itself is the one terminal.create already uses, so the live PTY is never killed, and the stable-pane create claim is still released in finally. The isReattach plumbing added to the launch factory for the old check is gone. A replay-safe launch refused this way on an existing workspace now records a failed ledger row, the same way a name collision does. Before, the row stayed claimed, so every retry got agent_session_operation_unknown. agent.launchReplay passes the code through. On create-worktree the workspace already exists when the terminal is refused, so the row stays unknown. The code is added to the runtime passthrough list so callers can branch on it. The pane key is now in the replay fingerprint, deliberately. It is not placement: group, anchor and focus still stay out of the request and out of the ledger. It is identity. It is written into the pane's PTY environment and names the tab the caller has placed. A retry that reserved a different pane is therefore a different request. Replaying the first answer would return a key the new reservation can never find. This matches terminal.createAgentSession, which also fingerprints its tab and leaf ids. The key is only folded in when present, so every existing digest is unchanged, and a test pins that. The wire schema now refuses a tab id the runtime would not adopt as sent: one with surrounding whitespace, which the runtime trims, and one longer than 512 characters, which the spawn reservation does not key on. It reuses the tab-id schema that Placement uses. * test(agent-launch): pin that a refused live pane issues no handle The refusal test named handle issuance but only asserted the reveal, so a throw moved to just before the reveal would still pass. Assert no terminal is registered, with the attach test as the positive control. |
||
|
|
60c43695e5 |
feat(agent-launch): report the pane a terminal launch created (#22108)
* feat(agent-launch): report the pane a terminal launch created A `term_*` handle is a main-side mapping the renderer cannot resolve (terminal-handle-links.ts:309), so a client that draws its own tabs had no way to name the tab it had just asked `agent.launch` to build. The runtime already mints that pane, bakes it into the PTY's environment and hands it to its own reveal; the surface factory then dropped it on the floor. Carry it through as `paneKey` on the terminal outcome. Identity, not placement: where the pane goes — which group, what order, whether it takes focus — stays with whichever client is drawing, and nothing here rides the wire for it. One field rather than a tabId/leafId pair, because the key already holds both and two copies of one fact can disagree. Absent when this launch minted no pane: a reused terminal was already running, and a worktree-create startup terminal is built by the create, which reports only a handle. Naming the wrong surface is worse than naming none. Optional on the wire and optional on the read side. Mobile parses the receipt with a loose object and is deliberately mode-blind, so it ignores the field; the persisted-row guard checks it when present and accepts a row written before it existed, because a read rule stricter than the write side turns one odd row into a refused replay. * fix(agent-launch): retain startup terminal pane identity |
||
|
|
d60043787b |
feat(agent-launch): carry the launch inputs the host cannot derive (#22037)
* feat(agent-launch): carry the launch inputs the host cannot derive Desktop's launch call sites cannot move onto `agent.launch` while the wire drops inputs they depend on. This adds the three the host genuinely cannot work out for itself, and deliberately adds nothing the host can. - `agentArgs` — the host read only `settings.agentDefaultArgs`, so a saved launch recipe's arguments had no way across. Tri-state is preserved: `null` is "no arguments", absent is "use the settings default". - `cwd` — `TerminalCreateOptions.cwd` already reached the spawn, but nothing on the wire filled it. It also decides the route: only a terminal can start somewhere other than its workspace, so the host now feeds it to `requiresTuiLaunchCommand` and downgrades with `tui_launch_command` rather than running a structured session in the wrong directory. - `launchSource` — telemetry, and the only member of the `agent_started` triple the host cannot derive; `agent_kind` and `request_kind` are computed host-side. Typed `z.string()`, not the closed enum: params are validated by the HOST, so a closed arm set would let an older host refuse a newer client's launch over a label. Attribution must not gate a user action. Not added, because the host already derives them: `launchPlatform` (`getAgentLaunchPlatformForWorkspace`, from the same connectionId/path/ projectRuntime the renderer uses) and `startupCommandDelivery` (a pure function of the agent inside `buildAgentStartupPlan`). Fingerprint: `agentArgs` and `cwd` are in — they change what the call does, so a retry carrying different ones must conflict rather than replay. `launchSource` is out — two buttons producing the same launch are one operation, and folding it in would refuse an honest re-attributed retry. A caller sending none of the new fields digests exactly as before, because the canonicalizer drops undefined keys, so launches admitted by an older build still replay across the upgrade. Arguments reaching a structured route are ignored by an existing deliberate decision (the Agent SDK and app-server version their option sets separately from the interactive CLI), so the host reports it in `warning` instead of overriding the user's preference on the strength of a field that is not evidence about the surface. * fix(agent-launch): forward create-target launch inputs |
||
|
|
663d670878 |
feat(agent-launch): deliver a launch prompt to a terminal agent (#21891)
`agent.launch` could hand its initial text to a structured session but not to
a terminal. The contract already anticipated the terminal half — the
`handed-to-terminal` arm has been declared in agent-launch-intent.ts since the
receipt was written and had zero producers — and the executor's own docstring
recorded the assumption behind the gap: that a terminal's paste belongs to the
pane owner. That assumption is what this overturns. The host owns the PTY, so
it can write into one whether or not any window is open on it, which is why
mobile and the CLI got an agent and no prompt.
A terminal takes its prompt one of two ways, and which one is not a
preference. `argv` exists so multi-line and special-character text reaches a
CLI as one argument rather than keystrokes, and it has no readiness race
because the text is in the process's arguments at exec time. So an agent whose
CLI accepts a prompt argument gets it on the launch command, and only a
`stdin-after-start` agent — plus any reused terminal, whose process started
before the launch existed — is written to as a bracketed paste.
That fork is asked once. `agentPromptRidesLaunchCommand` is derived from the
same injection table `buildAgentStartupPlan` branches on, and
tui-agent-prompt-transport.test.ts pins the two against each other for all 37
agents, so adding an agent or changing its mode fails loudly instead of
silently dropping that agent's prompt.
Reused rather than rebuilt: `sendTerminalAgentPrompt` is the runtime's one
agent-prompt writer (bracketed paste, per-PTY serialization, lifecycle
generation pinning, per-agent submit timing, and local/WSL/SSH routing), gated
by `waitForTerminal('tui-idle')` — the same pair orchestration's worker
dispatch already delivers a preamble through. The agent-first create path
needed no new mechanism at all: `startupPrompt` already flows to
`buildWorktreeStartupForAgent`, and the launch had simply been stripping it as
a reserved field without re-supplying its own.
Receipts stay consequences of the act they name. `handed-to-terminal` is
reported only from a launch command that carried the text or a PTY write that
returned; everything unproven under-claims as `not-delivered`. No fourth arm.
The one inversion is a stalled submission, which the verifier raises after the
write: that is reported as delivered, because a resend would paste the whole
prompt a second time into an agent already working on it.
A prompt the launch command cannot carry is refused at the terminal-create
resolver rather than dropped, since that path returns options and has no PTY
to fall back to.
`delivery: 'draft'` remains `not-delivered` for both surfaces. The host could
paste a terminal draft without submitting it, but it cannot observe that the
composer accepted it, so a receipt claiming delivery would be a guess.
No call site is migrated, nothing is added to the wire, and placement and tab
creation are untouched.
|
||
|
|
754134fd67 |
feat(agent-launch): deliver a launch prompt from the host (#21155)
* feat(agent-launch): deliver a launch prompt from the host `agent.launch` created the surface and then reported the caller's text as `not-delivered`, always: delivery lived in the renderer, so mobile and any other caller got an agent and no prompt. The host now commits a `submit` prompt to the structured session it just created, through the same send path `agentSession.send` runs, and reports `journaled` with the transcript row's id. Nothing is queued — the durable record that the text is owed is the journal's own submission row, which the send appends before dispatching, so a host-side copy could only disagree with it. The outbox's entry and envelope builders are reused so this send is shaped exactly like a client's, fingerprint included. Everything else under-claims as `not-delivered`: a terminal's paste is observed by whoever owns the pane, a `draft` has no host-side home, and a refused or thrown send commits nothing. There is no fourth "maybe" arm — a caller holding one could neither resend nor drop the text — and dispatch doubt stays on the submission row where it already lives. * fix(agent-launch): recover committed prompt after send errors |
||
|
|
b66ef2e8a8 |
fix(agent-launch): resolve a launch scope, not a git worktree record (#21193)
* fix(agent-launch): resolve a launch scope, not a git worktree record `agent.launch` asked the runtime for a managed worktree record and then read exactly one field off it, `.id`. That record does not exist for every workspace a launch can run in, so the request refused launches the method could otherwise run: the floating workspace resolves to a scope with an id and a path but no worktree row, and `showManagedTerminalWorkspace` throws `selector_not_found` rather than hand back the id it had already resolved. A folder workspace survived that only because the resolver fabricates a worktree row for it. The scope is the answer that is real for all three kinds, so the launch asks for that instead. `showManagedTerminalWorkspace` is unchanged - callers that genuinely need the git record still get it, and still get the refusal. With floating now reaching the mode decision, the host must know which kind of workspace it resolved. The kind is derived from the id it resolved itself, never accepted from a caller, and the route module's existing `floating` blocker does the rest: a workspace with nowhere to keep a session runs a terminal agent. Behaviour change, deliberate: a floating-workspace `agent.launch` used to fail with `selector_not_found` and now succeeds as a terminal agent. That is what lets the floating titlebar agent button move onto the shared launch command instead of driving tab startup itself. No wire change: `AgentLaunchTarget` is untouched. * test(agent-launch): cover floating RPC workspace resolution |
||
|
|
4b87bc718e |
refactor(agent-launch): redefine the agent.launch contract (#20999)
* refactor(agent-launch): redefine the agent.launch contract
`agent.launch` has no clients yet, so the contract is redefined in place
rather than versioned.
- params require `operation.id`, pinned to the shipped operation-id mint so
the host can read the embedded timestamp back. No caller-supplied
fingerprint: the host derives its own.
- the result carries `disposition` ('created' | 'replayed', the same
vocabulary `RuntimeCreateAgentSessionResult` already uses) and a single
top-level `warning` instead of one on the terminal arm only.
- the prompt receipt becomes an outcome enum, so a receipt can under-claim
instead of reporting a bare `delivered: false`.
- the dead `customization` field is deleted, and the mode-reason union and
receipt are declared once in shared with main re-exporting.
- `clientMutationId` joins the reserved create fields, with a test pinning
the list to the create schema in both directions.
Contract only; no behaviour change and no ledger wiring.
* docs(agent-launch): stop calling the stripped set "agent fields"
`clientMutationId` joined AGENT_LAUNCH_RESERVED_CREATE_FIELDS, so three
comments describing the stripped set as agent fields now teach the wrong
model — including a SAFETY rationale, where a reader is trusting it most.
The rationale's claim is unchanged and still sound: deleting keys from a
parsed object leaves the rest the parsed shape.
* refactor(agent-launch): make the attempt id the launch's only idempotency key
Review follow-ups on the contract redefinition.
`operation: { id }` becomes a flat `clientOperationId`, spelled the way
`terminal.createAgentSession` and the structured mutation envelope already
spell the same concept, and admitted by the shipped
`parseAgentSessionOperationTimestamp` rather than a second copy of its
pattern — so `agent-session-host-authority` keeps the regex private.
The handler now dedupes on that id instead of the create payload's
`clientMutationId`. That field is optional, so keying on it left any launch
that omitted one with no idempotency at all, while the required attempt id
did nothing. Reserving `clientMutationId` is still right, but for the reason
the comments now give: `createManagedWorktree` never reads it, so a copy left
in the forwarded payload is inert while still reading as a guarantee. The
previous rationale — that it was a second live dedupe key — was not true.
`messageId` moves onto the prompt receipt's `journaled` arm so a producer
cannot report the text as committed without saying where, and `rpcCallerKey`
picks up the `terminal.create` call site it was lifted from instead of
shipping with no callers.
* docs(agent-launch): record why disposition is two-valued only for now
The ledger admits attempts whose outcome was never recorded, and neither
`created` nor `replayed` can say "I cannot tell you" — a caller handed
`created` for an unresolved attempt starts a second agent. Noted at the type
rather than in review, so whoever wires the ledger reads it where they edit.
* fix(agent-launch): keep contract within implemented guarantees
|
||
|
|
97aa5ff19b |
fix(mobile): open native chat when a new worktree launches a default agent (#19850)
* refactor(agent-launch): make the launch-mode decision surface-neutral
`decideWorkerStartMode` was the only shared answer to "structured chat session
or terminal agent?", but it lived in an orchestration-named module and spoke
orchestration's vocabulary, so the other launch surfaces could not call it.
Move the decision to `main/agent-launch/agent-launch-mode` unchanged and leave
`orchestration-worker-start-mode` as the adapter that supplies the noun.
A worker is not a special kind of launch; it is the same launch with a dispatch
attached. Naming the receipt's subject is the only thing orchestration actually
contributed, so that is the only thing the adapter keeps: "worker" in both
sentences, plus the `--terminal` wording, which reads as nonsense anywhere a
`--terminal` flag does not exist. Both are pinned, because they are asserted.
No behavior change. The receipts are byte-identical for every reachable case,
proven by running the new pin against both implementations.
Also pins the wording, which nothing was holding. The existing suites assert
`toContain` fragments ('terminal agent', 'cannot create') and the CLI suite
asserts a receipt handed to it by a mock rather than one this code produced;
all six files stayed green against a deliberately corrupted vocabulary. A
dispatch receipt is the only place a structured-to-terminal downgrade explains
itself, so the whole sentence is the contract, not a fragment of it.
* feat(agent-launch): add the launch intent and the one executor that runs it
The sequencing around the launch decision was duplicated per surface, and the
duplicate is where the bug lives. A new worktree was created agent-first, so
its startup terminal WAS the agent and the structured branch below it could
never be reached — every new-worktree launch was a PTY regardless of the user's
default. Orchestration fixed that for itself in #19431; mobile and the CLI
still have it.
`executeAgentLaunch` inverts the order once, for everyone. When the preference
is structured the worktree is created with NO startup agent, the executing host
is then asked whether it can host a session for the workspace that now exists,
and only then is a surface created. The host verdict cannot be hoisted above
creation: `agentSession.createSupport` only answers for a workspace it can
resolve, which is why the decision stays in two halves.
Agent-first creation is deliberately preserved for PTY launches — it is what
sequences the agent's startup command behind the setup runner, so wait-for-setup
comes for free there.
What actually differs per surface is only how a surface is built (an
orchestration worker's session takes a dispatch hold and a mailbox a plain
launch must not take), so that is injected as a factory rather than branched on.
The intent also strips the reserved agent fields from a migrated create payload:
a caller moving off `worktree.create` passes its existing params, and a stale
`startupAgent` in there would re-create the very path this replaces.
Tests assert order and arguments, not just the resulting mode. Reintroducing
agent-first creation reddens 4 of 11.
* feat(agent-launch): expose the launch executor as the agent.launch RPC
Adds `agent.launch` — one host-side method that decides structured-vs-terminal and
creates the surface — wired to the real runtime factories: `createManagedWorktree`
for the workspace, forking on `startupAgent` exactly as the orchestration worker
path does; `createStructuredAgentSessionForWorktree` for a chat session; and
`createTerminal` for a PTY agent. Allowlisted for mobile, which is the surface the
routing gap was reported on.
`worktree.create` is untouched. Its `startupAgent` keeps meaning "spawn a PTY agent"
verbatim, because it answers with `agentTerminalHandle` only on that path: a host
that quietly routed it to a structured session would hand every older client a
response with no handle and no error. All new behaviour sits behind
`agent.launch.v1`, which the host now advertises and a remote client must negotiate,
so a client that does not gets today's behaviour unchanged.
* feat(mobile): route workspace creates through agent.launch
Picking an agent on the mobile create sheet always produced a terminal, even
when the user's default was native chat, because all three create paths put
`startupAgent` on `worktree.create`. That means "create the worktree
agent-first", so its startup terminal IS the agent and the structured branch
below it is unreachable — while the same phone's in-workspace "+" button opened
a chat.
The blank, branch and new-branch creates now send the same payload through
`agent.launch` and let the host settle the surface. `worktree.create` is
untouched, and a host that does not advertise `agent.launch.v1` (read from the
existing `status.get` probe) keeps today's path exactly.
Work-item creates stay on `worktree.create`: they pre-fill the issue/PR URL as
an unsent `startupDraft`, which a structured session cannot hold yet, so routing
them would submit the URL as a first turn.
* fix(agent-launch): drop the deleted draft-prompt blocker from the reason map
main removed the draft-prompt blocker in #19681 (a structured session now holds
an unsent draft), so the exhaustive Record no longer typechecks.
* chore(agent-launch): carry a SAFETY rationale on the agent placement cast
The type-assertion gate landed after this branch's base, so the new file's
copy of the worker-start cast is now a changed-code finding.
* chore(agent-launch): carry agent.launch through main's RPC typing and casting gates
The typed-method contract, the generated params catalog and the
`assertionStyle: never` casting scan all landed after this branch's base.
- AGENT_LAUNCH_METHODS kept an `RpcMethod[]` annotation, which widened its
method name to `string` and broke assignability; every sibling infers instead.
- `agent.launch` binds a schema under src/main, so it joins the catalog's
RPC_METHODS_WITHOUT_SHARED_PARAMS and the parity gate's hand-listed twin.
- The now-typed methods make most test casts unnecessary; the few that remain
carry the line-specific SAFETY rationale the casting gate requires.
* test(mobile): supply the agent-launch fixture the create-submit recording needs
The golden RPC recordings landed upstream while this branch was out, so they
first met agent.launch here. Three things had to happen, and only one of them is
a fixture bump.
1. workspace-settings-mounts.ts mounts useNewWorkspaceCreateSubmit against a
fixture model that throws on any member it was not given. This PR added a
required getAgentLaunchSupport, so the submit aborted with "Missing model
fixture" before it ever issued the create, and three cleanup checkpoints
vanished. That read like a product regression and was not one. Supplying the
member restores the recording byte-for-byte; it is pinned false for the same
reason the cutover probe is, so the baseline stays on worktree.create.
2. Editing that adapter moves adapterSha256 for the twelve settings goldens it
mounts. Their recordings are unchanged - header only, by design: the digest
is per-golden so editing a module fails exactly the goldens that mounted it.
3. Five goldens changed behaviourally, and both changes are this PR's:
the capability probe now reports agentLaunch, and a create whose reply
carries no worktree returns "Failed to create workspace" instead of throwing
a TypeError off an unguarded result.worktree read. The launch route needs
that guard, since a receipt can arrive without a worktreeId.
* refactor(mobile): decode the launch receipt instead of asserting its shape
The changed-code quality gate refuses type assertions, and the eight it flagged
were worth removing rather than suppressing.
The production one was the point. readAgentLaunchCreateOutcome asserted the RPC
payload into Partial<AgentLaunchResult> and then runtime-checked it anyway, so
the assertion bought nothing and claimed a contract the host had not proven. It
now narrows with `in` and validates each hop, which is the same nullability
question readCreateResult already answers on the sibling path - a launch receipt
can legitimately arrive without a worktreeId. AgentLaunchCreateOutcome ties
worktreeId to the shared contract so a change there fails this reader's
typecheck rather than passing a differently-typed field through.
The test fakes claimed a whole RpcClient via `as unknown as RpcClient` while
implementing one member. They now build a typed literal, matching the pattern in
use-mobile-structured-agent-options.test.ts. The read sites cast params and then
read one field; they now assert the payload with toMatchObject, which removes
the cast and pins more of the shape than the cast did.
Also pins the warning passthrough, which nothing covered: a terminal launch that
seats the workspace but cannot start the pty reports why, and the absent, blank,
non-string and structured-surface cases report nothing. Writing that test caught
a real drop I had introduced in the reader.
* ci(mobile): re-run Mobile Checks when a shared capability changes
Mobile Checks is path-filtered to mobile/**, but mobile imports the negotiated
capability names straight from src/shared/protocol-version.ts and records the
whole capability read verbatim in its goldens. So a capability added desktop-side
rewrites a mobile fixture while never triggering the suite that would catch it.
That is what happened here: #19849 introduced agent.launch.v1 and Mobile Checks
never ran on it. Verified at the run level rather than by check name - the
window-free check-runs API on
|
||
|
|
c702e77bc7 |
Stop reading the terminal arguments field on the structured chat route (#20944)
* fix(native-chat): stop reading the terminal arguments field on the structured chat route Setting Claude's Arguments to "--dangerously-skip-permissions --model Opus" made every new Claude tab open in the old terminal-backed chat instead of the new structured one, with nothing on screen to explain why. Removing "--model Opus" fixed it. The cause was a whole-string comparison: the configured arguments were checked against a single blessed value per agent, so any added token at all — including one the agent supports — stopped the string matching and the launch was demoted. Structured chat does not run the interactive CLI. It drives Claude through the Agent SDK and Codex through app-server, and those take narrower option sets that are versioned separately from the CLI's, so one free-text field cannot have a guaranteed meaning for all three. The structured route now reads only what it can actually honour: a replaced launch command, or a launch that names its own working directory. Terminal launches still apply the field exactly as before. Permission posture no longer travels as a raw flag. It is derived from the resolved launch arguments, which is the same fact a terminal launch acts on and which falls back to the default Orca ships when the field was never touched, so bypass stays on by default and Manual is still honoured. Claude gets the SDK's typed permissionMode and allowDangerouslySkipPermissions at query start; Codex gets its bypass flag placed before the app-server subcommand. Both are re-derived per acquisition beside the auth policy and environment overlay rather than stored in the session record, so nothing can disagree with the setting. Codex also loses the --profile, --add-dir and -c passthrough that reached app-server through that field. Only the permission posture comes back. * test(native-chat): pin routing authority on the narrowed feasibility input The routing-authority pin still named the old bundled blocker and built its "customized" fixture out of the arguments field, which is no longer a feasibility input. Both are now the launch command, and arguments and environment are customized on both passes of the loop, so the flag handed to the shared resolver tracks the command alone — a caller that resumed reading either one fails here. No case is dropped and no assertion is relaxed: the blocker list is still exhaustive and every caller must still honour a refusal from the shared resolver. |
||
|
|
357c9780f8 |
refactor(agent-launch): retire the duplicate worker-start mode decision (#20911)
`orchestration-worker-start-mode` becomes a thin adapter over
`agent-launch/agent-launch-mode`, which already owns the same decision.
Orchestration keeps its receipt vocabulary via WORKER_START_VOCABULARY, so
every sentence a dispatch receipt prints is unchanged.
Recovers the cutover written in
|
||
|
|
6da72383df |
feat(agent-launch): one executor for agent launches, exposed as agent.launch (#19849)
* refactor(agent-launch): make the launch-mode decision surface-neutral
`decideWorkerStartMode` was the only shared answer to "structured chat session
or terminal agent?", but it lived in an orchestration-named module and spoke
orchestration's vocabulary, so the other launch surfaces could not call it.
Move the decision to `main/agent-launch/agent-launch-mode` unchanged and leave
`orchestration-worker-start-mode` as the adapter that supplies the noun.
A worker is not a special kind of launch; it is the same launch with a dispatch
attached. Naming the receipt's subject is the only thing orchestration actually
contributed, so that is the only thing the adapter keeps: "worker" in both
sentences, plus the `--terminal` wording, which reads as nonsense anywhere a
`--terminal` flag does not exist. Both are pinned, because they are asserted.
No behavior change. The receipts are byte-identical for every reachable case,
proven by running the new pin against both implementations.
Also pins the wording, which nothing was holding. The existing suites assert
`toContain` fragments ('terminal agent', 'cannot create') and the CLI suite
asserts a receipt handed to it by a mock rather than one this code produced;
all six files stayed green against a deliberately corrupted vocabulary. A
dispatch receipt is the only place a structured-to-terminal downgrade explains
itself, so the whole sentence is the contract, not a fragment of it.
* feat(agent-launch): add the launch intent and the one executor that runs it
The sequencing around the launch decision was duplicated per surface, and the
duplicate is where the bug lives. A new worktree was created agent-first, so
its startup terminal WAS the agent and the structured branch below it could
never be reached — every new-worktree launch was a PTY regardless of the user's
default. Orchestration fixed that for itself in #19431; mobile and the CLI
still have it.
`executeAgentLaunch` inverts the order once, for everyone. When the preference
is structured the worktree is created with NO startup agent, the executing host
is then asked whether it can host a session for the workspace that now exists,
and only then is a surface created. The host verdict cannot be hoisted above
creation: `agentSession.createSupport` only answers for a workspace it can
resolve, which is why the decision stays in two halves.
Agent-first creation is deliberately preserved for PTY launches — it is what
sequences the agent's startup command behind the setup runner, so wait-for-setup
comes for free there.
What actually differs per surface is only how a surface is built (an
orchestration worker's session takes a dispatch hold and a mailbox a plain
launch must not take), so that is injected as a factory rather than branched on.
The intent also strips the reserved agent fields from a migrated create payload:
a caller moving off `worktree.create` passes its existing params, and a stale
`startupAgent` in there would re-create the very path this replaces.
Tests assert order and arguments, not just the resulting mode. Reintroducing
agent-first creation reddens 4 of 11.
* feat(agent-launch): expose the launch executor as the agent.launch RPC
Adds `agent.launch` — one host-side method that decides structured-vs-terminal and
creates the surface — wired to the real runtime factories: `createManagedWorktree`
for the workspace, forking on `startupAgent` exactly as the orchestration worker
path does; `createStructuredAgentSessionForWorktree` for a chat session; and
`createTerminal` for a PTY agent. Allowlisted for mobile, which is the surface the
routing gap was reported on.
`worktree.create` is untouched. Its `startupAgent` keeps meaning "spawn a PTY agent"
verbatim, because it answers with `agentTerminalHandle` only on that path: a host
that quietly routed it to a structured session would hand every older client a
response with no handle and no error. All new behaviour sits behind
`agent.launch.v1`, which the host now advertises and a remote client must negotiate,
so a client that does not gets today's behaviour unchanged.
* fix(agent-launch): drop the deleted draft-prompt blocker from the reason map
main removed the draft-prompt blocker in #19681 (a structured session now holds
an unsent draft), so the exhaustive Record no longer typechecks.
* chore(agent-launch): carry a SAFETY rationale on the agent placement cast
The type-assertion gate landed after this branch's base, so the new file's
copy of the worker-start cast is now a changed-code finding.
* chore(agent-launch): carry agent.launch through main's RPC typing and casting gates
The typed-method contract, the generated params catalog and the
`assertionStyle: never` casting scan all landed after this branch's base.
- AGENT_LAUNCH_METHODS kept an `RpcMethod[]` annotation, which widened its
method name to `string` and broke assignability; every sibling infers instead.
- `agent.launch` binds a schema under src/main, so it joins the catalog's
RPC_METHODS_WITHOUT_SHARED_PARAMS and the parity gate's hand-listed twin.
- The now-typed methods make most test casts unnecessary; the few that remain
carry the line-specific SAFETY rationale the casting gate requires.
* docs(agent-launch): stop the receipt-wording comment claiming a migration
The decision was never moved out of orchestration-worker-start-mode; this PR
adds a second copy beside it. Say so, and name the unenforced agreement.
* docs(agent-launch): stop the executor comment claiming a migration that has not happened
The header asserted two things the tree does not support: that every launch
surface routes through the executor, and that the mode decision "already lived"
in `agent-launch-mode`. `agent.launch` is the executor's only consumer, and
`orchestration-worker-start-mode.ts` is byte-identical (blob
|
||
|
|
2eb93206c8 |
refactor(agent-launch): make the launch-mode decision surface-neutral (#19848)
* refactor(agent-launch): make the launch-mode decision surface-neutral
`decideWorkerStartMode` was the only shared answer to "structured chat session
or terminal agent?", but it lived in an orchestration-named module and spoke
orchestration's vocabulary, so the other launch surfaces could not call it.
Move the decision to `main/agent-launch/agent-launch-mode` unchanged and leave
`orchestration-worker-start-mode` as the adapter that supplies the noun.
A worker is not a special kind of launch; it is the same launch with a dispatch
attached. Naming the receipt's subject is the only thing orchestration actually
contributed, so that is the only thing the adapter keeps: "worker" in both
sentences, plus the `--terminal` wording, which reads as nonsense anywhere a
`--terminal` flag does not exist. Both are pinned, because they are asserted.
No behavior change. The receipts are byte-identical for every reachable case,
proven by running the new pin against both implementations.
Also pins the wording, which nothing was holding. The existing suites assert
`toContain` fragments ('terminal agent', 'cannot create') and the CLI suite
asserts a receipt handed to it by a mock rather than one this code produced;
all six files stayed green against a deliberately corrupted vocabulary. A
dispatch receipt is the only place a structured-to-terminal downgrade explains
itself, so the whole sentence is the contract, not a fragment of it.
* fix(agent-launch): drop the deleted draft-prompt blocker from the reason map
main removed the draft-prompt blocker in #19681 (a structured session now holds
an unsent draft), so the exhaustive Record no longer typechecks.
* chore(agent-launch): carry a SAFETY rationale on the agent placement cast
The type-assertion gate landed after this branch's base, so the new file's
copy of the worker-start cast is now a changed-code finding.
* docs(agent-launch): stop the receipt-wording comment claiming a migration
The decision was never moved out of orchestration-worker-start-mode; this PR
adds a second copy beside it. Say so, and name the unenforced agreement.
---------
Co-authored-by: Merge Sim <sim@local>
|