mirror of
https://github.com/stablyai/orca.git
synced 2026-10-09 00:02:39 +00:00
a747c1c013fae3b04ff76e87587ae44bf9ee36ca
1080
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a747c1c013 |
Allow another phone send after delivery is unconfirmed (#26392)
* Allow another phone send after delivery is unconfirmed * Update phone image recordings after journal removal |
||
|
|
e4cd14d914 |
Remove legacy OS file-drop routing (STA-6940 PR6/6) (#26385)
* feat(file-drop): add element owner preparation plumbing * Fix terminal and chat file drop destination ownership * fix runtime terminal drop ownership and queued chat retries * fix(file-drop): preserve feedback acceptance and destination ordering * fix(file-drop): attach chat and composer files at the drop surface * fix(file-drop): preserve project destinations and live chat availability * fix(file-drop): keep path resolution in filesystem namespace * test(file-drop): use filesystem path bridge in PR3 fixtures * fix(file-drop): bind terminal drops to pane elements * fix(file-drop): compare terminal pane identity across public views * test(file-drop): align composer lifecycle with element-owned drops * Fix interrupted chat composition and refuse unused quick-create drops * Type the quick-create drop regression transfer * fix(file-drop): let the explorer, project sidebar, tab strip and editor own OS drops STA-6940 PR5. The file explorer tree, the project sidebar, each editor group's tab strip and its editor area now take OS file drops on their own element through the shared owner hook, with identity from their own render (the explorer's shown workspace and target row, the tab strip's and editor group's worktree and group) instead of the active worktree. The explorer and sidebar broadcast subscribers and their legacy drop markers are deleted; editor-open moves out of useGlobalFileDrop into editor-dropped-file-open, which the legacy unmarked-chrome route still uses until PR6. * fix(file-drop): keep editor delivery tied to its destination group * fix(file-drop): guard delayed opens and share floating ownership * fix(file-drop): remove legacy OS drop routing (STA-6940 PR6/6) * test(file-drop): finish owner test and comment cleanup * Clarify dropped-file default editor destination * Retain UI test namespace for terminal resume coverage |
||
|
|
b1e0092d5d |
feat(opencode): open OpenCode 2.x in the structured chat too (#26395)
* feat(opencode): run OpenCode 2.x in the structured chat too `opencode acp` on stable 2.x starts its own private `opencode serve --stdio` child with the chat's environment and ends it when stdin closes, so the chat's account pin reaches it just as it does on 1.x. Admit stable 2.x from 2.0.14 beside stable 1.x from 1.18.31; pre-releases, older releases, and other major lines keep the terminal chat. The `opencode2` agent is unchanged. Restart recovery already reads a 2.x session's own tables first; tests now cover a database an upgrade left with both generations, and one whose 2.x message table has an unknown shape (nothing found, no error). * fix(opencode): show OpenCode's own permission option names * test(opencode): run the 2.x recovery tests with the real-SQLite Node project |
||
|
|
b41e2730db | Persist multiple workspace references with atomic edits | ||
|
|
8fdad2a3af |
fix(claude): activate account profiles and remove credential replay (Step 4 of 4) (#24434)
Each saved Claude account now runs in its own CLAUDE_CONFIG_DIR folder, so Claude renews each login itself and switching no longer replays a copied credential. A claude shell function routes every launch to the selected account; the terminal daemon protocol moves to 42 so new terminals get it. After updating, each saved account signs in once: typed claude shows Claude's own sign-in, old terminals show an in-app banner, chats and the status-bar menu offer Sign in, and a one-time toast explains it. Orca prints nothing into terminals. |
||
|
|
316822f2eb | ci: skip cache warming for cache-test-only changes (#26399) | ||
|
|
acd298a263 |
feat(runtime): isolate the server behind a compatibility launcher (#26374)
* feat(runtime): isolate the server behind a compatibility launcher * fix(runtime): preserve server crash exit semantics |
||
|
|
c71601f51c |
Add Pi structured native chat through its RPC mode (#25851)
* End a running call as its turn's journal row ends
A call still running when its turn ends takes the state of that turn's
row: a row another writer settled first (a person's Stop) stands, so its
calls read interrupted whatever the provider's later end reports. The
no-ending path that settled calls from the Stop row is gone, since a Stop
now leaves running calls to the provider. Adds the two Spanish strings.
* Say why a Grok turn failed, and keep task rows in Grok's own words
A failed Grok turn ended with no reason on screen: the translator dropped
every copy of Grok's message. The failed turn now gets one status row in
Orca's existing "provider did not accept this message" words with Grok's
reason, read from whichever copy arrives first (the given-up retry, the
turn's end, the prompt's completion notice, or the prompt's error answer);
later copies only fill a reason the row still lacks.
A running background command no longer reads "Background task <id>
started": a task's summary is mapped only once it has settled. A monitor
stays a monitor when the agent reads its output: a frame that names no
kind keeps the known one, and a "[monitor" command is a monitor.
A prompt's turn is marked started, so a late frame for an ended prompt
neither reopens it nor becomes the active turn. A tool's turn is held in
one place at a time.
* Read a monitor from Grok's exact output prefix
* Word a failed Grok turn in Grok's own text, not as a refused message
A turn that started and then failed was told "The provider did not accept
this message", Orca's sentence for a message refused before its turn. The
row now reads as a Codex turn-ending error does: an error status row with the
provider's own words. With no words, the dialect names the failure ("Grok
ended this turn with an error." / "Grok usage limit reached."), else the
agent's display name does.
* Settle a stopped turn's running call as its turn row ended after a restart too
The restart sweep ended every running call by the death evidence alone, so after
a person's Stop with no proof the child died the call read failed under a turn
that read interrupted. The sweep and the live dead-generation settlement now ask
the same rule the assembler does: a call in a turn already settled ends as that
row ended; only a turn still running leaves its calls to the evidence.
* Keep the dead-generation settlement under the line cap
* Register the ACP schema verify step in the PR preflight phase test
* refactor(agent-session): one required agent registry; declarations admit what they claim
A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.
/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.
Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).
* Read ACP permissions, session events and prompt errors through the protocol client's own types
The translator now reads a permission request with the client's lenient reader, a session update
with its session-event reader, and takes only the agent's own error answer as a failed prompt's
reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values
newer than this build.
* refactor(agent-session): the router applies the declared rewind itself
The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.
* test(agent-session): register the agents the merged-in tests now need
The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.
* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record
* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop
The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.
The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.
One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.
* fix(agent-session): a changed agent definition never hides that agent's chats
A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.
Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.
* refactor(agent-session): each agent's registration says where it runs and which account it pins
createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.
Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.
* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state
A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.
* fix(agent-session): a scoped dismiss-all persists no per-session fence
The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.
* fix(agent-session): refuse an attach whose agent is not the session's own
The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.
* fix(agent-session): offer to start a chat only when the start would accept it
The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.
* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it
A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.
* refactor(agent-session): the record store admits agent ids; comments say where transport is checked
The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.
* docs(agent-session): the record store admits the registered agents' ids
* refactor(native-chat): Grok's registration declares where it runs; ACP no longer borrows Codex's location rule
The rule a self-supervised agent child runs under (this machine, no WSL, Windows only with process
start-time proof) is its own module that Codex and the ACP adapter both use. Grok's registration
takes the full account-home resolver signature, and D3's tests build hosts with the agent registry.
* fix(native-chat): Grok follows the ACP runtime's request contract and the managed process's close
A request the agent or a Stop cancels is answered with the agent's own cancelled reply by the code that
owns it (the runtime no longer answers a silent handler), so a Stop needs no separate decline pass. A
permission answer still being saved when the agent stopped waiting is reported unconfirmed, since the
protocol already answered it cancelled. Cancelling the agent's own turn is the plain cancel. Request
rows are matched under their generation-scoped ids. A refusal's reason comes from the dialect's wording
path. The child drops its own stderr tail and close policy for the managed process's, and a close
whose process tree was not proven gone is reported as the adapter contract asks.
* fix(native-chat): a Grok chat Orca already holds resumes without writing what Grok replays
A chat with a saved Grok session reattaches with session/resume where the agent offers it, else
session/load. Either way the call runs inside the translator's load window, so what Grok sends while
it reattaches (its saved exchange, a task the dead process left running, ended by the restart) opens
no turn and writes no row; only context usage reads on. A reply an Orca or Grok crash cut short is no
longer completed from Grok's saved history: it reads like a Claude or Codex chat's, with the existing
notice. The attach window also closes after a failed attach, and a created session that session/resume
reports missing is replaced like one session/load reports missing.
The replay reconciliation is removed: the lane no longer reads the journal, and D3's replayed-input
grammar test and completed-turn check in the assembler go with it.
* refactor(native-chat): a failed Grok reattach needs no window close of its own; its lane is replaced
* test(native-chat): D3's merged tests use the shipped declarations and the launch options main requires
* fix(native-chat): typecheck fallout of the base merges; any agent's empty chat is reusable
Main's idle-empty-chat lookup and launch join now take any registered agent, as the rest of the
launch path does. The refusal check moved into the prompt turns and the prompt-block conversion beside
the turns that send it, keeping both files in their line limit.
* fix(native-chat): a Grok Stop ends the process once Grok settles its turn; the next send resumes
Grok's session/cancel ends only the running turn: work it already moved to the background keeps
running and can begin a turn of its own after the person pressed Stop. Stop is now a session
boundary, as it is for Claude: the cancel answers open requests and lets Grok end the turn, the host
waits a bounded grace for that, then ends the process; the next send relaunches and resumes.
The adapter's own bounded close of a turn Grok began is gone. Its named-turn check stays: the host
ends the session unless the provider declines a Stop naming a turn that has since ended.
* test(native-chat): a Grok Stop ends the process only after Grok answered the cancel
* fix(native-chat): Steer on a Grok card cancels the running prompt, then sends it
A send that reached Grok while a prompt ran was held in the adapter until that turn ended: Steer
on a queued card took the card out of the host's editable queue and meant 'send after this turn'.
It now cancels the running prompt (session/cancel; the session stays) and sends as the next prompt
once Grok answers the cancel, as the common pattern does; a steer behind another cancels it in
turn, so the last one runs. The adapter holds a send only while that cancel lands, so its general
held-send queue and its holdsDispatch report are gone (every send it holds has its turn open in
the journal). An older client's mid-turn send takes the same path. capabilities.steering is
unchanged and still unread.
* refactor(native-chat): a close or Stop cancels a start through the acquire's own abort signal
The host owns the acquire it runs, so it now owns its cancellation: each attach's acquire gets an
AbortSignal, aborted from outside the session's queue by a close and by a Stop admitted now (the
same admission rule as before). The optional abandonStart adapter hook, the router's fan-out to
every adapter and the ACP adapter's session-keyed start map are gone; the ACP adapter keeps an
unkeyed set of starts only so quit can prove their children gone, and keeps a failed start's
unproven child until its exit is proven.
The hook also let a later close ask that child again. The host now does that from state it holds:
a close of a chat with no live child whose record still names an owner process with no death
evidence asks the adapter to release it. The answer is not recorded as proof (the lease probe
does that), so an owner pid an earlier Orca left is never killed or marked gone. Claude and Codex
ignore the signal and hold no such child; their release is a no-op (tested).
* fix(native-chat): a Grok crash that closes stdout before its exit still ends with Grok's last words
On macOS and Linux the agent's stdout ends before its exit is observed, with or without the
supervisor's EOF forwarding, so the connection's loss closed the journal first and its error text
became the session's ended reason, dropping Grok's stderr. The reason is now read at the proven
exit: the agent's last words when it left any, else why the connection closed. The failure already
carried them. Comments that assumed the exit comes first, that early frames past the cap refuse the
start, and that dispatch re-checks image support are corrected.
* fix(native-chat): nothing Grok sends while a held chat reattaches is written, marked as replay or not
The reattach window relied on the dialect's replay verdict, and Grok's frames read as live unless
they carry isReplay, so an unmarked chat frame during session/resume opened a turn that never
ended. D3 now marks every frame inside the window as replay before the translator reads it, so the
translator keeps only context usage whatever the agent marked; options and commands are still
adopted. The translator's load semantics are unchanged.
* test(native-chat): a Stop after a resume finds no turn an unmarked old reply opened
* test(native-chat): a resumed Grok chat keeps its last context reading; the resume refreshes only the window
* test(native-chat): a Grok background task a Stop ended reads as stopped reporting
* refactor(native-chat): quit's stop of each start answers through one promise kind
* fix(native-chat): quit aborts every start the host has in flight before draining attaches
A Grok that never answered its handshake held quit until the start's own 60 s bound, past the
20 s quit deadline. The host's teardown now aborts each in-flight acquire (and any the drain
still begins), so the adapter's own quit controller and its map of starts are gone: a start
has one canceller, the host's signal.
* fix(native-chat): a Grok start's abort stops reaching its child once the start has returned
The listener stayed on the host's signal until the attach finished committing, so a Close in that
window killed the now-live child behind the host's back and it read as Grok crashing. The start
now detaches it when it ends; a later Close goes through the session's own stop.
* fix(native-chat): a close or Stop during any attach phase stops the start before it launches
The attach began its abort controller only after reconciling leases, resolving recovery and
probing the previous owner, so a close or admitted Stop in those phases reached nothing and Grok
launched anyway. The controller now begins first, and the acquisition checks it before asking the
adapter to start.
* test(native-chat): a close during the attach's owner probe asks no adapter to start
Also renames the close test after the hook it no longer exercises.
* test(native-chat): a close's re-ask closes a Claude or Codex child a failed cleanup left
The re-ask is not a no-op for them: when the adapter still holds the child its cleanup could not
prove gone, the close stops it again as a requested close, and Claude persists the handle of the
conversation it ran so the next send resumes it. Corrects the tests' and comment's wording; the
close awaits the re-ask, bounded by each adapter's kill ladder.
* fix(native-chat): Steer during a turn Grok began itself cancels it and sends once it ends
A send while Grok ran a turn of its own (a background task waking it) went straight to Grok, which
queued it behind that turn where Orca could no longer withdraw it, while Stop treated the same turn
as the running reply. The send now waits as a steer, the turn is cancelled once, and the message
goes when the turn ends; a Stop withdraws it and an exit rejects it as never sent.
* test(native-chat): a steer whose cancel Grok never answers ends Grok and is rejected as never sent
Pins the bounded steer cancel kept from the runtime: past the bound the connection closes, the
running reply reads unverifiable, Grok's end reads as its exit, and the waiting steer is rejected
as never sent.
* fix(native-chat): a Grok crash stays a crash when a stop lands before its exit is proven
After the connection broke and the close could not prove Grok's exit, any later stop Orca asked
for (the next start, a Stop, a Close) marked the child as closed by Orca, so the crash read as a
requested close and Grok's last words were dropped; a send meanwhile was recorded unconfirmed.
The connection loss now decides the cause, and a send on that session is rejected as never sent.
* test(native-chat): fixtures this PR's registered Grok and desktop capability made stale
CI's unit shards failed on tests outside the PR's own lists. Each encodes something this PR changes
on purpose: Grok is now a registered agent (the seam test's unregistered agent is now Cursor); the
desktop now advertises registered agents (the restart-offer tests' older client drops that
capability explicitly); the attach context carries the start's abort controllers (the forget-status
double gains them); and the ACP real-host test rig sends to the host directly (listed beside the
other real-host rig in the send ratchet).
* fix(native-chat): a start quit stops is not the queued message's start failure
With quit now aborting a start it would have waited for, the delivery step recorded the aborted
start as the message's failure ("couldn't restart"). After quit has stopped delivery, the step
leaves the message to quit, which settles it as a close does ("The chat closed before this message
was sent."). The test that pinned quit waiting for that start and stopping its child now pins that
nothing is launched behind quit.
* fix(native-chat): a message sent after a Stop or close aborted a start gets its own start
A start the host aborts (an admitted Stop, a close, or quit) returned its refusal to the delivery
loop, which then rejected whatever was queued at that moment with "couldn't restart", including a
message the user sent after the Stop. The attach now reports that the host aborted it, and the loop
re-derives from the journal instead: what the Stop or close withdrew is already settled, a message
accepted since gets a start of its own, and quit's next step stops the loop. This replaces the
quit-only carve-out with the same rule for every abort and every agent.
* test(native-chat): the message sent after an aborted start is answered, so no settlement outlives the test
* fix(native-chat): a Grok model pick Grok never answers no longer holds Stop or Close
The pick runs on the session's queue. It now registers in the host's out-of-queue
abort registry beside a start, so a close, an admitted Stop or quit abandons it, and
the ACP adapter bounds it at 30 s like Claude and Codex. A late answer is still adopted.
* fix(agent-launch): a phone's launch opens a terminal for an agent whose chat it cannot show
agent.launch now reads the caller's capabilities by the rule tabs and restart offers
use (clientRendersStructuredAgent). A phone without registered-agents.v1 gets Grok as
a terminal again, as on main; the host's own callers and desktop clients are unchanged.
* fix(acp): strip every agent hook variable from the ACP child, from the shared list
ACP_CHILD_ENV_TO_DELETE was a second copy of the hook runtime keys that missed
ORCA_AGENT_HOOK_TRANSPORT; it now spreads AGENT_HOOK_RUNTIME_ENV_KEYS beside the pane
identity keys.
* refactor(native-chat): the mutation context carries the provider-wait registry itself
Keeps the host file within its line limit; one field instead of two closures over it.
* fix(agent-launch): agent.launch.v2 still vouches for Claude and Codex chats
The caller rule from the previous commit also turned Claude and Codex into terminals
for a client advertising only agent.launch.v2, whose contract says it opens a chat
(mobile retry-authority tests). Only an agent beyond those two now needs the client to
read it (clientRendersStructuredAgent); the test fixtures go back to what they were.
* refactor(native-chat): drop saved-history adoption from the timeline assembler
The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.
* refactor(acp): drop session/load history adoption from the translator
The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.
* refactor(native-chat): a pending input is only Orca's send now
Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.
* test(acp): keep the task-result status table on live frames
Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.
* test(acp): a frame helper for a shell command Grok is running
* fix(acp): a Grok crash settles through the host's provider-exit batch, scoped to the turn it ended
A Grok crash ended the journal unverifiable before the adapter reported the exit, so the host's
provider-exit settlement found no running turn and wrote nothing: the adapter's failure (with
Grok's last words) never reached the journal, and a later stale-session pass wrote a bare,
thread-scoped cut-short row, so the partial reply was not folded as Claude's and Codex's are.
At a proven exit the ACP lane now ends its running turn interrupted at the exit instant, as the
host's exit contract expects of a child's own translator (Codex's does the same). When Grok's
stdout closed first (every POSIX crash), the turn is unverifiable only until the exit is proven:
the host's provider-exit settlement now takes the exit as proof naming the child's fence and
revises what that child left unverifiable in the same batch, with the turn-scoped row and the
adapter's failure. Claude and Codex write no unverifiable turn of a live child except a command
whose hand-off is in doubt; that turn is now revised at the exit instead of at the next open.
* test(acp): a crash seen first leaves the host no Grok turn to revise
* refactor(native-chat): what a gone generation left unfinished gets its own module
The settlement file passed 300 lines with the exit-proof revision. The unfinished-work reads
(capture, interrupted-by-the-exit, in-progress) are their own concept and move out unchanged,
apart from the exit proof they now take.
* refactor(native-chat): a watched exit revises what its child left unverifiable without reading Stop marks
An exit's own instant is the turn's end, so the revision needs only each row's fence: the
settlement's journal type gains itemFence alone, and the host test fakes say so.
* test(native-chat): drop the duplicate itemFence on the fake that already had one
* test(claude, codex): an exit whose stdout ended first still reports as it always did
The provider supervisor now ends Orca's stdout when the agent's ends, so on every crash EOF
arrives before the exit is seen. Claude's and Codex's connections report nothing at EOF and
report the exit, with its usual reason, once it is seen.
* fix(acp): reopen a chat with session/load, as the common pattern does
An agent that offers both now reloads its session instead of resuming it; the
reattach window still discards what it replays except context usage.
* fix(acp): drop the 60 s handshake bound; an abort fails the start's waits at once
Neither common design bounds an ACP handshake: Close, Stop and quit end a start
that never answers. The abort now also closes the connection, as a kill there
does, so the start settles even before the child's exit is proven. The
host-stopped start refusal only this bound produced goes with it; the idle
sweep keeps its words.
* fix(acp): a Stop naming an ended turn follows Claude's rule
It still stops nothing while another turn is live, but in the gap before a
follow-up's turn opens, which no client can name, it now stops what is in
flight and the session ends, as a Claude Stop does.
* fix(native-chat): a close no longer re-asks a failed start's unproven child
Neither common design retries that stop at Close, and Orca's Claude contract
re-asks only at the next start and at quit. The ACP adapter keeps the child
until its exit is proven and asks it again there, as Claude does.
* fix(acp): a message sent during a turn the agent began itself goes at once
Both common designs send it straight to the agent with no cancel; only Orca's
own running prompt is steered (cancelled, then re-prompted).
* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer
Uses an audience production sends (one that cannot show every agent), per review.
* fix(acp): launch Grok as `grok agent stdio`, without the update and leader flags
The common pattern passes neither --no-auto-update, --no-leader nor
GROK_DISABLE_AUTOUPDATER; full access still adds --always-approve.
* fix(acp): an agent that ends its stdout, or answers unreadably, is not a lost connection
As in the common pattern, only a broken stdin (or Orca's own close) ends the
agent; one that closed its output but can still be written to stays until a
Stop, a close or its exit. The provider supervisor goes back to its base
content, so Claude and Codex no longer get the forwarded stdout end either.
* fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts
On main every close of the chat writes its own Stop and settle. Here a later close joins the
first and writes no row, and the first's settle closed when its kill failed, so a turn that opened
in between and was cut by the next close read as failed. A person's close joining a person's close
whose Stop opened a settle now reopens that settle until its attempt is done.
Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and
the next reads as the person's cancellation (each fails without its half of the fix).
* test(native-chat): Grok opens as a chat only behind the structured-chat setting
agent.launch and orchestration worker-start read the same setting as the
renderer route; pin both states for Grok on each. The setting's description no
longer names only Codex and Claude, in every catalog.
* docs(acp): generic ACP comments say what holds for every agent, not Grok
Stop ends the session for every ACP agent, as in the common pattern; the
adoption hook comment goes (adoption is not planned); a failed start's child is
retried at the next start or quit.
* test(claude, codex): type the EOF-before-exit test's streams; the supervisor no longer forwards EOF
The Claude test wrote to the child's stdout and stderr through their Readable
type, which the node typecheck rejects; it now holds its own PassThrough
streams. The comments no longer credit the reverted supervisor change.
* feat(acp): a steer's cancel asks once and never ends the agent
The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.
* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge
* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns
Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.
* fix(native-chat): drop the stopDelivery the A3 merge doubled
* fix(acp): a steer's cancel asks Grok once and never ends it
A steer now uses D1's notify-only cancel. Two messages sent during a reply Grok began itself
cut that reply, as the common pattern does, and then both run; before, the queued first
message could not answer the bounded cancel and Orca ended Grok although Grok answered.
A Stop keeps the bounded cancel and its 4 s grace.
* fix(acp): a permission Grok asks with no prompt of Orca's running is declined
During a turn Grok began itself nobody asked it to act, so the request is answered
cancelled at once instead of opening a card that waits, as the common pattern does.
* fix(acp): a Grok that dies while starting is reported with its own last words
A dying process's stdout ends before its exit is seen, so the start failed as a closed
connection and Grok's stderr was lost. A start whose connection closed now waits, bounded by
the Stop grace (or a Close/Stop), for the exit before it is told.
* test: a Stop after a steer sends its own cancel; drop the import the A3 merge doubled
* test(native-chat): main's Stop-note test builds its turn context with the agent registry
* test(claude): say why the close test's fake child cast is safe
* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way
The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.
* test(native-chat): build the Stop-opened-turn test's identity with the opaque handle
The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle.
* test(ratchet): require src/main/provider-process now that it has landed
* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test
Main's #25706 and this branch both added the import at different lines; the merge kept both.
* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test
Main's #25706 made the same fix as this branch at a different line; the merge kept both imports.
* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore
A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.
Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.
* test(native-chat): build this stack's journal identities with main's opaque provider handle
Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test
files from this stack still wrote the old shape. Same lines the downstream ACP branch uses.
* docs(acp): every reattach drops the agent's replay, not only for a chat the journal holds
* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row
Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.
* refactor(native-chat): read hosts' structured agents from the app-shell services
Main grew the startup hydration hook to its line limit; the host agents sync is an app-lifetime subscription like the structured session tabs sync beside it, so it moves there.
* Use current provider handles in transition tests
* Use current provider handles in timeline fixtures
* test(native-chat): prove replacement rows survive downgrade and re-upgrade
* Require the ACP directory in the runtime import check
* test(ratchet): require src/main/acp now that this PR lands it
* feat(acp): a saved session the agent cannot reopen continues in a new one, with one warning row
When session/load (or session/resume) of a saved ACP session fails, the chat starts a new session and records it as a creation that replaces the lost one (#25747's 'replaces' link), and writes one warning row that the agent no longer remembers the earlier messages. A created session the agent reports missing is still superseded silently; a signed-out agent or a start that is over (Close, Stop, a lost agent) still fails the start.
* chore(acp): rewrap the acquire header comment
* test(acp): a start closed while the agent reopens fails without opening or announcing a new session
* Let ACP connections own their supervised agent process
* Preserve ACP cleanup evidence and isolate exit observers
* Expose ACP cleanup observations and type the permission fixture
* refactor(native-chat): the registered-agents capability lives in its own module
Main's growth put protocol-version.ts one counted line over its 300-line limit once the capability
was added; like main's other per-feature capabilities, it now has its own module, and importers read
it from there.
* refactor(acp): one connection owns the Grok process and its protocol
D3 now opens each ACP agent through createAcpAgentConnection (ACP-ALIGN #25810): one object spawns the
process on the execution host, owns its stdio and protocol, and reports its proven exit. It is built and
tracked before the handshake, so a start's abort (Close, Stop, quit) still reaches it, and a failed start
keeps that same connection for the next close to retry rather than spawning another process.
Deleted: the spawnAcpStructuredChild wrapper and its test, the raw-stream runtime assembly, the caller's
exit -> runtime.close wiring, the stdout-EOF heuristic (the connection no longer treats stdout EOF as
exit), and the 10 s steer/Stop cancel bound with requestSteerCancel. Reader control maps to
pauseReading/resumeReading; a close is connection.close after the host's existing 4 s Stop grace.
The adapter owns what the protocol no longer does: one session/cancel per running prompt however many
steers arrive (cleared with that send's settlement, retried after a failed write), and a Stop or steer
answers every open agent request the person has not already answered with the agent's own cancelled
reply. An answer already being saved when the Stop lands is sent.
Tests: blocked cancel write never holds Stop's grace, two quick steers send one cancel, a failed cancel
write is retried, a real process exiting while a child holds its stdout ends the session, and the
existing start-abort, retention, crash, connection-loss and reload-failure suites on the new rig.
* fix(acp): Grok signs in on its own machine with its API key or cached sign-in
When Grok reports that it needs authentication, Orca now names a sign-in method on the machine Grok runs
on, read from the same environment Grok was launched with: xai.api_key when XAI_API_KEY is set there and
Grok offers that method, else cached_token when Grok offers it, else none and the chat keeps the existing
not-signed-in refusal. The rule lives in Grok's launch spec; the adapter applies any agent's rule for new
and reopened sessions through the protocol client's caller-named method (authenticate, then retry once).
No new sign-in UI; interactive methods are never chosen.
* fix(acp): the adapter decides which of Grok's requests reach the person
The turn owner now admits every agent request, permission or question, from its own turn state: a
request reaches the person only while Orca's prompt runs and no steer or Stop is cutting it short (a
question may also come from a turn Grok began itself, until a Stop). Anything else gets the agent's
own cancelled reply and opens no card, so a question arriving after Stop or during a steer never
appears. A steer, like a Stop, withdraws the requests already open; an answer already being saved is
still sent. The protocol client's abort-on-cancel path is no longer used: after the connection
change its request signal aborts only when the connection closes.
* feat(native-chat): add inactive Pi RPC transport foundation
* fix(acp): a plan Grok proposes shows as a plan, with no approval card
When Grok leaves plan mode it asks the client to approve its plan (x.ai/exit_plan_mode). Orca showed a
blocking 'Approve plan / Request changes' card for it; the common pattern has no such gate. Now the
plan goes into the chat's existing Plan row (the plan-document status row Codex and ACP plan updates
already use) and the request is answered at once with 'abandoned' plus feedback telling Grok to stop and
wait for the person's feedback or a request to implement it in a later turn, so nothing is approved on
the person's behalf. Dialects gain settleRequest for requests answered without asking anyone.
* fix(orchestration): worker-start opens a Grok worker in a terminal, as before
With the structured chat setting on, worker-start decided 'structured' for Grok and then the structured
worker factory (Claude and Codex only) refused it, so the start failed; main opened a terminal Grok
worker. Worker-start now decides with no registered agents beyond Claude and Codex, so Grok gets a
terminal worker as before. agent.launch and the app's own launches still open Grok as a structured
chat. Temporary until structured workers take registered agents.
* fix(acp): a prompt answer Orca can't read ends the turn instead of hanging it
A session/prompt rejection that was not the agent's own error answer (an answer that fails Orca's
schema, or one too large to read) left the turn running: the next message became a steer with nothing
to cancel and was never sent or settled, and Stop waited its full grace. As in the common pattern, any
prompt failure now ends the turn as failed (a failed-turn row without words, since none are the
agent's) and settles the send, so the next message goes. Only a closed connection keeps the send
running, for the connection-loss path to settle.
* feat(native-chat): add bounded RPC reading control
* fix(acp): send Grok's prompt-identity extension only to agents that echo it
session/prompt carried _meta {promptId, requestId} for every ACP agent, though only Grok's dialect
echoes it (injectedPromptIdentity). Now only an agent whose dialect declares it gets the extension;
other ACP agents get a plain prompt.
* refactor(native-chat): the registered-agents capability lives in protocol-version again, as on main
This reverts
|
||
|
|
b1e09a4c8e |
fix(native-chat): show an SSH session's chat history on the phone and desktop (#26334)
Fixes #26057. The hook reports the SSH host's transcript path, but native chat read it on the desktop's own disk, so SSH sessions showed an empty chat. The desktop now reads the transcript on the SSH host through its existing SSH filesystem provider, routed by the path layer that already handles WSL. A local row attesting the requested path keeps the session local; a transcript not yet written keeps the chat waiting. Desktop-only, no wire change. |
||
|
|
c785c986c2 |
feat(native-chat): teach chat agents to show inline visuals in their own folder (#26099)
* feat(native-chat): visual directive grammar and host read for a chat's visuals folder
A shared grammar for the ::orca-visual{file="..." title="..."} reply line,
the per-chat visuals folder location on the owning host, and the
agentSession.readVisual runtime method that reads one visual with lexical and
canonical containment, a 512 KiB bounded read and UTF-8 refusal.
* feat(native-chat): render chat visuals inline and in the right sidebar
Native-chat assistant replies render a ::orca-visual{...} line as the chat's
HTML visual in an opaque, scripts-only sandboxed frame: CSP first, the host
frame navigation guard registered before content runs, live theme without a
reload, fitted height, links opened in the viewer's browser only from a real
gesture, lazy mount, and one muted line when the visual cannot be shown.
Open in sidebar shows the same frame in the right sidebar, widened while it
is open and restored after.
* feat(native-chat): teach chat agents to show inline visuals in their own folder
Native-chat Claude and Codex sessions now get a per-chat visuals folder on
the host that runs them, write access to exactly that folder, its path in
ORCA_CHAT_VISUALS_DIR, and an Orca skill that teaches the ::orca-visual line.
- Skill ships as an unpacked plugin folder in desktop and headless builds.
- Claude: --plugin-dir via SDK plugins, behind the CLI version probe (now one
shared probe for every version-gated flag); folder added to
additionalDirectories beside the user's own.
- Codex: skills/extraRoots/set and the folder appended to the user's own
writable roots, between initialize and the thread open, under a 2 s budget;
unsupported, failed or hung setup opens the chat without visuals.
- A host sweep removes folders no chat record maps to, and folders whose
local workspace is provably removed; anything unproven is kept.
* test(native-chat): the visuals skill names only theme variables the frame sets
* fix(native-chat): visual CI fixes, shared height governor, live-turn streaming hold
Registers agentSession.readVisual from the methods index so the structured
method file stays under its line budget, replaces reflective reads with checked
narrowing, moves the pure height governor to src/shared for mobile, and holds a
half-written directive tail while the turn works (structured text rows carry no
running state).
* fix(native-chat): harden the visual read and link opening
Re-checks after the open that the chat's visuals folder is still the real
directory at Orca's path, reports unexpected filesystem faults by code without
host paths, and lets one click in a visual open at most one page.
* fix(native-chat): keep visual lines out of plain-text reply surfaces; review fixes
One shared helper drops visual lines (outside fenced code) from reply text where
it becomes plain text: the structured status summary that feeds the sidebar row,
dashboard, notifications, phone rows and handoffs, and AI Vault reply previews.
Review fixes: height also counts a pinned body's overflow, only the live
frontier row holds a half-written visual line, the runaway-height stop needs the
same step repeated, and any host refusal evicts the cached revision.
* fix(native-chat): review round 1 for chat visuals delivery and cleanup
- Visuals sweep: a workspace counts as removed only when no profile's
catalog holds it (chats and visuals are shared by every profile, catalogs
are not); a worktree in a known project is removed only when git no
longer records it; any unreadable profile decides nothing; one catalog
snapshot per run; a symlinked visuals root is never walked.
- The other-profile catalog reader moves out of window/ and also returns
project ids.
- A folder Orca itself inherited is stripped at the spawn layer for both
agents, not only from the launch overlay.
- Claude version probe: a probe that gave no version is never remembered;
the plugin check waits up to the probe's kill time so a slow first probe
no longer costs a chat its skill.
- Skill: kept out of Claude's / menu, filename and theme guidance matched to
the renderer, refused writes are not retried.
- Opt-in real Codex test for skill discovery and the writable root.
* feat(native-chat): ask the Claude CLI its version when native chat starts
The first chat after Orca starts usually finds the version known, so the
plugin and thinking-display checks answer at once instead of probing a
cold binary while the chat waits.
* fix(native-chat): resolve the visuals folder without the removed journal-paths helper
Main removed the per-chat journal paths and the journal database's state
directory; the visuals folder keeps the same sha256 layout on its own and the
read method uses the profile state directory the chat host is opened in.
* fix(native-chat): review round 2 fixes; copy a reply without its visual lines
Reply previews in Agent Session History drop visual lines per text part before
lines are folded; the frame adds a body's overflow only when the body really
overflows; fence tracking follows CommonMark closers and openers; the copy
button copies a reply without visual lines; a coded read fault keeps its cause.
* fix(native-chat): review round 2 for chat visuals delivery and cleanup
- Read git worktree records written relative to their own folder (git 2.48+),
so a worktree on an unmounted drive in such a repo is still kept.
- Skip the running profile by its own storage folder, not the profile index a
switch rewrites first; a profile never written to counts as empty, so the
workspace rule is not switched off by a profile that was never opened.
- Ask for readable thinking again once the plugin check has waited for the
version, so a slow first probe no longer drops it for the chat's life.
- Launch flag decisions move to their own module; tests use a typed record
fixture instead of casts.
* fix(native-chat): review round 3 for the visuals sweep and version check
- Resolve a relative git worktree record against its folder's real path, so a
project added through a link still matches and its worktree is kept.
- A profile with only backups of its data file is a lost file, not a fresh
profile: it still stops the workspace rule.
- The readable-thinking re-check reads what is known and never starts a
second version probe.
* fix(native-chat): update the frame's theme ref after render; read the visuals folder pair at once
* test(native-chat): declare agentSession.readVisual on the cross-version agent-session surface
* fix(native-chat): copying a reply keeps its code blocks and indentation
Removing visual lines now closes only the gap each removal leaves, instead of
collapsing blank lines across the whole reply and trimming its indentation; the
visuals folder is checked parent first again so a broken path answers the same
way every time.
|
||
|
|
995ef11ce7 |
feat(native-chat): say in the chat why Orca stopped a reply, and offer Continue (#25675)
* feat(native-chat): say in the chat why Orca stopped a reply, and offer Continue When the Orca that runs a structured chat (this computer or a paired server) quits, updates or crashes mid-reply, the chat's stopped row now names the cause and the machine, and a Continue button sends the existing restart continuation for that cut turn, with or without a restart offer. - Host: a quit/update writes one turn-scoped row for the turn its stop cut, in today's words, with an optional `orcaStop` cause on the providerExited fact; restart adjudication stamps how the previous runtime ended on the deaths it proves (crash, or the quit/update it began), so the crash row names it too. Older clients keep their single row. - Host: agentSession.continueInterrupted, capability-gated, rechecks under the session lock that the chat still sits on that cut, so a second click or a retry sends nothing. - Client: the row's copy names the cause and machine; Continue sits above the composer. * test(native-chat): Continue is not held by a recovery file that never answers * fix(native-chat): bind an Orca stop's cause to the runtime that held the agent; neutral row, Continue explains itself - The cause now rides on the host's row itself (`orcaStop` on the status row, beside today's words), the same row family and id scheme as the reopen's death row. - Each recorded owner is stamped with the Orca runtime that holds it; a death proven later (at restart, or when recovery stops a survivor) names how that runtime ended: the quit or update it began, else a crash. Owners an older build recorded, a terminal's claim, an agent that died while its Orca ran, and unreadable quit records all keep the generic words. - The quitting runtime's word is written first in teardown, before the recovery wait, through a bounded asynchronous writer apart from the chat database. - Row copy: one neutral sentence naming the machine and cause; it drops "You can continue in this conversation." while Continue is offered, and Continue's tooltip says what it does. * test(native-chat): type the Orca-stop test fixtures so the typecheck passes The cut turn's outcome takes the journal's outcome type, and the stand-in close reads the provider sink through a checked lookup instead of an index that may be absent. * fix(native-chat): call a cut a crash only when Orca's runtime started and never ended A chat said "Orca stopped unexpectedly" whenever its runtime left no quit record, and only the desktop quit wrote one, so a headless server's restart or update, the Settings relaunch, and a Windows logoff all read as crashes. Each runtime now records its own start when its chat store opens, and every graceful exit records its end through one synchronous entry point: the desktop quit's teardown, the headless server's stop, the in-app relaunch, the GPU-fallback restarts, the update-install watchdog, and Windows session end. A crash is a runtime that started and never ended; a runtime with no readable record (never written, pruned, unreadable) names no cause, so the chat keeps its generic words. One file per runtime, written durably and only by that runtime, so a damaged file never blocks a later write and two processes never lose each other's record. * fix(native-chat): Continue answers once Orca accepts it, not once the agent has started On a paired server, Continue waited for the agent to start before answering, and the client gives a paired call 15 s. A slow start (account switch, login shell, a long resume) showed "Couldn't continue this chat" while the agent was in fact continuing. Continue now answers when Orca has accepted the message, as a send does. The agent's start and answer settle afterwards, and a start that fails is the chat's own note, as before. The restart dialog's batch still waits for the handover, which is where it counts a start as done. The verdict helpers move to their own module to keep the continuation file within its size limit. * fix(native-chat): a reply the user steered, or a command run after the cut, still offers Continue The cut detector stopped at the first user message after the cut turn, so a steer the turn had taken, or a conversation command such as /context run after the cut, removed Continue while the row still named the cause. The rule for what is no request of its own (a conversation command, a row its turn produced, a send handed into a running turn) moves out of the latest-request reader into one shared predicate, which both that reader and the cut detector use. The host's "still wanted?" check reads the same detector, so the client and host agree. * fix(native-chat): a death proven after an earlier settle explains the turn it ends When a chat was read before the restart proved its old agent dead, the read could only call the turn unverifiable. The proof then revised the turn to interrupted, but the death row was scoped by the turn still marked running, and none was, so it landed on the conversation instead of the turn. That cut never offered Continue, and the row did not name it as the turn's explanation. The row is now scoped to the newest root turn the settle actually ends, running or revised. * fix(native-chat): the cause row's words stay put, and Continue waits out a resume already running The row's "You can continue in this conversation." came and went with the button: it showed while a paired host's answer was still on its way, vanished when the button appeared, and came back the moment Continue was clicked. Continue also appeared on chats the restart prompt or the launch's own resume was already carrying on. The row now drops that sentence wherever the chat's host can continue a cut, counting a host that has not answered yet as able (a host that writes cause rows has Continue), so its words never change on screen. Continue is hidden while a resume is carrying that chat on. * fix(native-chat): a refused Continue says so once, in the composer A Continue the host refused before accepting anything wrote a red note into the chat and brought the button back, so each retry added another identical note; a chat the host had no record of was refused with no word at all. Continue now reports every refusal the same way as a failed request: the existing composer line "Couldn't continue this chat. Try again, or send a message.", which a retry replaces rather than repeats. A refusal before acceptance writes no note. A failure after the message was accepted (the agent could not start) is still the chat's own note, as for any send. * fix(native-chat): the row naming Orca's stop carries its own presentation and never folds A client that re-words host rows it cannot name (the draft that makes these cuts read as interruptions) treated the cause row as an older red row and replaced its words, so the cause never showed there. The row was also folded away under its collapsed turn once shown muted. The host's cause row now names the presentation 'orca-stop' beside today's words, failure fact and red tone, so a client that predates both changes still prints exactly today's row, red and on screen, and a client that re-words unnamed rows passes it through. This build shows it muted, counts it as no failure (the reply it cut stays the turn's answer), and never folds it; the fold field becomes `explainsTurn`, as the other change names it. * test(native-chat): pass the session-end event without a type assertion * test(native-chat): the row naming Orca's stop renders neutral, whoever re-presented it Pins the rendered tone on this build: the stored red row, and the same row after a reader re-presents it in the neutral tone with its presentation and cause kept, both render muted and never fold. The phone draws chat rows without tone styling, so it needs no change. * fix(native-chat): the "Couldn't continue" line goes once the chat is continued The composer line a failed or refused Continue set stayed on screen while the agent carried on: after an answer lost in transit, or once another client or the restart prompt continued the chat. Only the next Continue click or the user's own send cleared it, and a click also wiped an unrelated composer error. The line is now derived: shown only while the chat still sits on the cut that Continue failed on, so it goes as soon as the journal shows the chat continued, from anywhere. A Continue click clears only its own line, and a retry answered "already continued" leaves none. * fix(native-chat): Continue waits while an opted-in launch may still resume the chat With "resume automatically" on, Continue showed on a quit or update cut while the launch was still waiting for its settings and reading the restart offer, then vanished when the launch's own resume began; a click in between sent a competing continuation. The launch's one decision (nothing offered, ask, or resume) is now published, and the chats it resumes are named the moment it decides, with no gap. Until it decides, and while the setting has not loaded or is on, Continue stays hidden on this machine's chats; a paired server's chats are not the launch's to resume and keep it. * fix(native-chat): a runtime's end survives a late reinstall, a failed write and any clean exit Three ways the runtime record could still read a graceful stop as a crash: - A chat host reinstalled during the quit (a request landing after teardown began) recorded the runtime's start again and erased the end it had just written. A second start of the same runtime now keeps that end. - When the end could not be written (a full disk), the start alone stayed and read as a crash. The runtime now removes its record, so its chats name no cause. - Each `app.exit(0)` had to remember to record the end. A process 'exit' with code 0 now records a quit when nothing else did: Electron emits it on every quit and exit once its loop runs (`app.exit` -> Browser::Shutdown -> the app's 'quit' -> process 'exit'), and Node on every `process.exit`. The relaunch and GPU-fallback calls it covers are dropped; the quit teardown, the headless server's stop, the update watchdog and Windows session end keep theirs, which run earlier or say more. * test(native-chat): build the re-presented row as the plain status item it is * fix(native-chat): a Continue click clears the composer's old error, so its own failure shows Since the "Couldn't continue" line became derived, an older composer error (such as "Remove attachments before using a chat-session command.") outranked it: a failed Continue showed the old error instead, and a Continue that went through left the old error on screen. A Continue click is the user's newer action, so it clears the composer's error again, as before; the line then shows the Continue's own failure, if any. That failure still goes away by itself once the chat is continued, and nothing but the user's own Continue click clears an unrelated composer error. * fix(native-chat): a chat start compares the owner process, not the runtime stamped on it A chat start checks that the process it just started is the one the record names, by a deep comparison of the stored owner with the adapter's process. The store stamps that owner with the Orca runtime holding it, so the check passed only because the store happened to return the record from before the stamp; returning the published record would have refused every chat start with agent_session_ownership_unknown. The start now compares the process identity without the runtime stamp, which says who holds the process rather than which process it is. * test(native-chat): count agentSession.continueInterrupted among the structured methods * refactor(native-chat): derive the structured chat's transcript session in its own hook Main's appearance work and this branch's Continue wiring together put NativeChatStructuredSession past the 400-line limit for components. The session the transcript reads moves, unchanged, to use-structured-chat-live-session.ts. * refactor(native-chat): keep the Continue capability in its own module Main grew protocol-version.ts to its line limit; the Continue capability moves to its own module, as other capability groups have, and the runtime list still names it. * refactor(native-chat): keep two shared files within their line limit after the main merge Main left agent-session-record.ts and structured-agent-session-params.ts just under 300 lines, and this branch's additions put them over. The account-home shape check moves next to the account-home type it checks (written without a type assertion), and the Continue params move to their own contract module; the params catalog is regenerated. No behavior change. * test(native-chat): compare the store directory's files without depending on listing order The corruption test checks that no file was created or removed by comparing two recursive listings. Their order is the runtime's: with the per-runtime record directory nested under the store, Bun returns the same entries in a different order than Node. Both listings are now sorted. * fix: share the path bound main's launch-directory check needs * test: give the stop-row fold rows the draws flag main's fold now reads * refactor: mark the launch's resume decision where the resume begins * test: count main's new structured method alongside agentSession.continueInterrupted * fix: the journal database keeps its folder, where runtime end records live Main's #26038 dropped stateDirectory from JournalHostDatabase; this PR's runtime end records are read from and written beside it. |
||
|
|
0f9f199322 |
fix(native-chat): no saved outbox on the desktop; one send at a time, and the host owns what it accepted (#25959)
* fix(native-chat): the host owns the send queue; the window keeps no saved outbox The desktop kept each structured chat's unsent messages in localStorage and sent them in order, so one message whose fate was unknown froze every later send, Retry dropped it silently, failed sends could not be discarded, and an offline chat could send hours later. The host already records every message and owns the queue; the window now only sends. - One in-memory sender for composer, launch prompts and messages sent from outside the chat. One send in flight per chat; a transport failure resends the same id for up to 30 s; a refusal that proves nothing was recorded puts the text back in the composer with the reason; a send that went out and was never answered shows an in-doubt line with Send again and holds nothing up. - The host's "unknown" rows get the same in-doubt line and Send again, from the journal, in every window. - A message an older build left in localStorage is never sent: the host's conversation outline decides what goes back to the composer, and the copy is deleted once that is saved. * fix(native-chat): send through the structured chat RPC wrapper, with its timeouts The sender called the runtime RPC directly, skipping the per-method timeouts every other structured chat call gets. Only a remote host's request skips the compatibility check the sender already ran. * fix(native-chat): an unconfirmed send goes back to the composer, with no new row line The common pattern draws nothing extra on a message whose delivery is in doubt once its turn is over, and puts a failed send's text back in the composer with the reason. So a send nobody answered in time, or one a host answers in a way that proves nothing, goes back to the composer worded as unconfirmed, and a host "unknown" row shows nothing extra. Removes the in-doubt phase, Send again, and its strings. A host's made-up row for an id its journal lost now reads as unconfirmed, not as recorded, so that message comes back instead of vanishing. * test(native-chat): pin that nothing resends a send given back as unconfirmed * fix(native-chat): sends survive a tab close, never resend after a Stop, and keep remote images - A normal tab close lets sends on their way settle; what the host never took goes back to the conversation's draft. Only a cancelled launch or a worktree purge drops them. - After any Stop, a send already on its way is never sent again under its id: a doubtful answer, or a resend that was due, hands its text back worded as unconfirmed. - A returned image keeps the SSH connection it lives on; the connection never goes to the host. - An older build's saved message the host recorded and then rejected is left to the host's own row, never handed back. * fix(native-chat): a Stop or a tab close never resends a send already out, and a kept card is the card's A Stop that landed while a same-id resend was being readied (its timer fired, its request not out yet) let that resend go out after the Stop. The sender now tracks whether a request is awaiting its answer: a Stop hands back every send between attempts as unconfirmed, lets one whose request is out settle from its answer, and never issues a request after it. A normal tab close withdraws the same way instead of doing nothing, so nothing more goes out and nothing is dropped. A send the host rejected but kept as a card (keptAsQueuedMessageId) is the card's, from its reply or the journal, and never goes back to the composer. Adds the freeze tests: a send whose fate is unknown holds later sends no longer than its deadline, and one the host can neither confirm nor deny releases the next at once. * fix(native-chat): read a resend's turned-away call or reused id as unproven Ports the send-answer proof contract. A call the host turned away before running it (method_not_found, invalid_argument, unauthorized) proves only that this request wrote nothing, so it reads as never sent on a first attempt only; on a resend an earlier attempt may have landed, and it goes again under the same id. A resent id the host says was already used for other content (messageIdReused) proves nothing either, like an expired or conflicting id. Pins the rest of the contract: any row the host returns is its own, a thrown error is no answer whatever its code, and an older build's refused, held or outlived-Stop copy is handed back once and never sent. * refactor(native-chat): type the structured composer's send with its attachment type Keeps NativeChatStructuredSession.tsx within the file length limit. * refactor(native-chat): drop the kept-card guards the sender never needed A kept send's row is a rejection that was never a Stop's, so the sender already reads it as recorded, from its reply or the journal. The test that pins it stays; the two extra checks only covered a kept row that is also a withdrawal, which the host never writes. * test(native-chat): route launch tests' sends through the client wrapper the sender calls The sender sends through callStructuredAgentSession, but these launch tests replaced that module with a factory that answered nothing (and still named a probe that no longer exists), so every launch prompt resent until its 30 s deadline and the tests timed out. Each factory now hands sends to the runtime RPC mock the tests already answer, and expectations of a local send no longer ask for the remote-only compatibility option. * fix(native-chat): hand a message back without importing the composer's attachment hook The worktree purge reaches the launch prompt, which hands text back, and the attachment hook's imports reach the store. A test that builds the real store behind a mocked one then waited on itself and hung. Handing back now writes images to the draft store directly, as the hook's helper did, and a test pins that each returned image keeps its SSH connection. * fix(native-chat): one send per chat, with no line of sends behind it A chat with a send out took further messages into an in-memory line and sent them one by one. A message typed behind one in doubt then hit its own 30 s deadline and came back as not sent without ever going out. Now, as the common pattern does, a chat takes one send at a time: while it is out, Send is disabled and Enter leaves the text in the box. The sender refuses a second send instead of lining it up, so the 30 s deadline always runs from the send itself. A Stop or a tab close stops the one send: settled from its answer if its request is out, handed back as unconfirmed between attempts, or silently if it never went out. Notes sent from outside the chat while its send is out stay with their sender (not ready); a launch prompt that meets the person's own first message waits in the composer instead of being lost. * fix(native-chat): give a Stop-withdrawn note back to its chat once its notes were cleared Notes sent from outside a chat clear once their message is recorded. A message the host recorded as pending and a Stop then withdrew came back only to its sender, which had already let go of it, so the text was lost. The chat's draft now takes it, as an earlier build's outbox did. * test(native-chat): type the composer-actions probe without a cast * chore(native-chat): drop outbox wording left in comments and an empty locale group * fix(native-chat): pace failed checks, keep the host's reason, and hold the chat for its launch prompt - A send whose checks failed before its request went out (an unreachable or incompatible host, an unreadable history) was tried again at once, about a thousand times a second for 30 s. Attempts are now paced by the attempts made, whether or not their request went out. - A host that refuses every resend by throwing (native chat turned off, a journal that won't open) gave back only "couldn't confirm". The host's reason now comes first, still without claiming not sent. - A launch's prompt now holds the chat's one send from the click, drawn as sending, so a message typed while the chat starts can't overtake it; it goes out once the chat exists, and a cancelled launch frees it. - Notes whose send nobody could confirm say so instead of "did not accept", and notes launched into a new chat let go of their text once that chat's composer holds it. - An open chat keeps drawing a recorded send until its row arrives, so a reply that beats the history no longer makes the message flicker out. * fix(native-chat): hold a chat's sends until it has started, and send a failed chat's message with its restart A message sent to a chat still starting went out at once to a host that had no record of the chat yet, read for a fence it could not get, and came back as not sent. A message sent to a chat whose start failed restarted it but no longer went with the restart. While a chat starts, Send stays off and Enter leaves the text in the box, as with a send already out. A text message sent to a chat whose start failed restarts it and goes as the restart's first message, through the same staged-prompt path a launch prompt takes: it holds the chat's one send slot, drawn as sending, until the chat exists, and comes back to the composer if the restart fails again. Notes wait while a chat starts and ride a failed chat's restart, keeping their text until it is sent. * chore(native-chat): test the queue request, rejection words and card hand-offs; drop an unused clear - Pin which sends ask the host to queue, the moved rejection wording, and that a queue send whose card was handed off and then refused or withdrawn, or whose replay names a withdrawn card, is never handed back. - The send-at-most-once gate now says what the desktop promises after a reload: it never sends the id again, and hands the text back. - The legacy read names when it goes, and drops the notice clear nothing called. * fix(native-chat): keep a sent launch prompt's entry, and give back at once what can't go out - Cleaning up a launch prompt released its send slot even after the prompt had gone out, which deleted the entry the sender keeps once the host records it. A launch prompt recorded as pending and then withdrawn by a Stop was lost from both the chat and the box, and an open chat dropped the new chat's first message before its row arrived. The slot now gives back only a reservation that was never sent. - A send stopped before its request went out by something trying again won't clear (this client and the server can't talk, or the host refused the history read) kept Send off for 30 s and then said Orca couldn't reach the agent. It now comes back at once with its own cause: not sent, since nothing went out, or unconfirmed if an earlier attempt did. Transport errors keep the paced retries. * fix(native-chat): notes keep their own text through a new agent's launch Notes sent to a New agent whose start failed came back to the notes and also sat in the new chat's composer, so sending both delivered the text twice. The notes keep their text until it goes out (they hold it from the click), so the launch now says so and no composer gets a copy, on a failed start or a refused prompt alike. The tests that asserted the copy in the composer pinned the old double ownership and now assert the notes are its only owner. Notes that rode a failed chat's restart and ended unconfirmed now say Orca couldn't confirm them, as notes sent directly do, instead of that the agent did not accept them. * chore(native-chat): the send-once gate says what the desktop keeps across a reload The desktop keeps a send's id in memory only: a send still unsettled at a reload or crash is not resent and not handed back. The coverage notes no longer credit the renderer tests with durable identity across a remount. * fix(native-chat): say once why notes sent to a new agent did not go Since notes keep their own text through a new agent's launch, a prompt the host refused, nobody could confirm, or that found the chat's send taken came back to the notes with nothing said anywhere: the chat shows no notice for text its caller keeps. The notes menu now reports it once, with the toast a send to an existing chat already uses: not accepted, couldn't confirm, or not ready. A start that failed still says so in the new chat instead. * fix(native-chat): retry a history refusal the host says clears, and name it at the deadline A send stopped before its request went out came back at once for any refusal of the history read, including ones the host names as clearing (its journal briefly unavailable, a chat detached while the host quits), so the automatic retry was lost. Only a version mismatch and refusals that won't clear come back at once now; the rest go again on the paced schedule, and if the deadline still finds nothing sent, the words are the host's refusal rather than Orca couldn't reach the agent. * feat(native-chat): a send makes one request, and nothing ever sends it again The desktop resent a message under its own id for up to 30 s when the answer was lost. A send now makes exactly one request, as the common pattern's clients do: - An answer that is lost, dropped or proves nothing hands the text back at once with "couldn't confirm… check the chat". - A failure before the request goes out (the environment check, the history read for the fence, a refused connection) means nothing went out: the text comes back at once with its own reason, or as not sent. - The 30 s cap stays on the one request, so a host that never answers can't hold Send. Nothing ever resends, so nothing can go out after a Stop: a Stop only takes back a send still in its pre-send checks. The resend state goes with it (tries, generation, awaiting, stopped, the resend timer, the send-answers-proof probe, and the first-attempt/resend split in the evidence), and the send-once gate says the desktop never resends. * fix(native-chat): notes keep the chat's line, and a send held behind a rewind says it was not sent - Only a send whose text belongs to the chat's composer clears the chat's line. Notes sent from outside the chat no longer wipe a "couldn't confirm... Check the chat" line that explains text already back in the box. - A send the host turns away behind a rewind it could not confirm went back as "not sent" but said "couldn't confirm what happened. Check the chat". It now gives the rewind's reason and says the message was not sent. - Reliability gate names the single-request test and records a fresh evidence run; the sender test drops its leftover resend mocks and a duplicate Stop test. * fix(native-chat): a send ends even when putting its text back fails If writing the returned text into the chat's draft threw, the send never settled: it stayed "sending", the chat refused every later send until a reload. The send now always ends after a hand-back, the failure is logged, and the chat's line adds "Couldn't save your message." * fix(native-chat): a message typed during /clear stays in the box and follows the chat A /clear moves the chat to a new conversation, and its host refuses any send while it runs. A message sent then went out, came back refused into the old conversation's draft with its line, and vanished when the chat moved on: the box and the line now belong to the new conversation. - A /clear holds the chat's one send slot while it runs: Enter does nothing, Send shows busy and the text stays in the box, as for a send that is out. Notes sent from outside get "busy". - Once it moves the chat, the old conversation keeps taking no send until the view leaves it, and what is left of its draft moves into the new conversation's draft, after anything there: when the composer's /clear settles, and again when the view moves. * fix(native-chat): no "Send message?" or queue clear while the chat's send is out While a send is out (or a /clear runs) the chat takes no message, yet Enter over a held queue still opened "Send message?", and Clear queue deleted every card before its message was refused. Enter now does nothing there and the text stays; Clear queue re-checks and deletes nothing if a send went out after the question opened. * refactor(native-chat): take the /clear draft carry out of this change Moving the old conversation's draft into the one a /clear replaces it with fixes a bug main has too (text left in the box during a /clear stays with the old conversation), so it goes in its own change. Kept here: a /clear holds the chat's send slot while it runs, and the conversation it moved away from takes no send until the view leaves it. * fix(native-chat): a /clear's hold on sends always ends - A view that unmounted while a /clear that moves the chat was out left the old conversation holding its sends until a reload: the late reply kept the hold for a view that was gone. The reply now releases it. - The local call for a conversation command has no deadline of its own, so a /clear that never answers kept Send off for good. The hold now also ends at the command's deadline (195 s, the one the remote call already uses), whichever comes first. * test(native-chat): opening a chat an older build left stuck; hand back its copy in send order Pins, through the real chat hook, sends, legacy recovery and draft store, what opening such a chat does: the queued messages behind a message the host recorded in doubt come back to the composer once, in order, with the "not sent" or "couldn't confirm" wording; the chat is free to send under its read's fence; nothing happens while the chat has no fence here. The legacy reader now hands entries back in the order they were sent (queuedAt), as the older build's reader did, instead of array order. * fix(native-chat): a send this window turned away for a re-paired server comes back as not sent A managed server's update rotates its pairing, and this window's main process then answers the next call itself, before forwarding anything, with runtime_environment_changed. The send read that thrown answer as proving nothing, so the person was told Orca couldn't confirm a message that never left. It is now handed back as not sent, with the reason. Every other thrown answer still reads as unconfirmed once the request may have gone out. * test(native-chat): a message refused for an expired attachment comes back with its file and why A paired server checks every stored file a message names when it admits it, and refuses the whole message before recording it when one has expired. The sender hands such a message back to the chat's draft, file included, with the refusal's words, whether or not a view shows the chat, so it can be removed and attached again. Ported from the saved-outbox test that came with attaching files to a structured chat on a paired server. |
||
|
|
e05c69fe8f |
fix(session): at startup, a local copy never overrides an SSH-owned workspace's own copy (#26098)
* fix(session): at startup a local copy never overrides an SSH-owned workspace's own copy Rows the local partition holds for a workspace whose repo the catalog places on an SSH target are residue (pre-#19572 builds, relay reattach). Startup used to keep them whenever they held a tab and skip the SSH partition for that workspace, dropping live tabs, restoring closed ones and erasing agent-resume records on the first save. Boot hydration now drops those local rows (keeping open files and visit recency) so adoption takes the SSH partition whole. * refactor(session): let adoption take a catalog-owned SSH workspace instead of pre-filtering local Replaces the sentinel-partition split with one rule in adoption: the base's terminal tabs no longer keep out the partition the repo catalog places the workspace on (contested ids keep today's behavior). * fix(session): an SSH-owned workspace's records win under shared keys; unsaved local drafts survive Addresses review: a stale local tab sharing an id with the SSH tab kept its layout and resume records (fill-only), and a replaced open-files row dropped local-only unsaved drafts. * fix(session): an SSH copy with no tabs never replaces local tabs Addresses review: an owned workspace the SSH partition holds no tabs for keeps today's rule, so its empty tab row cannot wipe the local tabs. * refactor(session): drop a superseded local copy before adoption; publish path passes owned ids Simplifies the rule: for a workspace the catalog places on this SSH host, uncontested and with host tabs, the base's rows are dropped and the existing gap-fill adoption runs unchanged; unsaved local drafts survive. One resolver says which workspace each scoped entry belongs to, shared by the census, the drop and adoption. The upload path (persistedSessionForTarget) now passes owned ids too, so main cannot publish the stale local copy to the host. * fix(session): match host tab rows by workspace id; a local draft beats a clean host entry Addresses review: a host tab row under a workspace key now counts toward superseding the local copy, and a local unsaved draft for a path the host holds clean is kept instead of dropped. * fix(session): scope catalog-owned ids to the ssh partition the catalog names A workspace homed on one SSH target with leftover rows in another target's partition no longer has the owner's adopted rows removed by the leftover's pass. |
||
|
|
7b26725ff4 |
Keep native watcher contracts in Node and reset runtime test caches (#26315)
* test: keep native filesystem watcher contract under Node * test: reset canonical repository keys with runtime mocks |
||
|
|
00984ccb25 |
ci: run full pull-request unit tests across ten shards (#26295)
* ci: run full pull-request unit tests across ten shards * Bound required Linux package tooling setup |
||
|
|
990c62e6c6 |
feat(native-chat): attach files to a structured chat on a paired server (#25146)
* feat(native-chat): a paired server keeps a store for chat attachments A structured chat on a paired Orca server had nowhere to put a file the client attached. The server now keeps a per-chat attachment store under its userData, filled through agentSessionAttachment.uploadStart/Append/Commit/Abort and read back for previews through agentSessionAttachment.read, behind the structured session gate and advertised as agent-session.attachments.v1. Nothing records cleanup as owed: a sweep re-derives what may go from the host's chat records and journal (unfinished uploads after an hour, uploads for a chat that never existed or that its journal never mentions after a day). The clipboard RPC's in-flight bookkeeping moves into a shared ChunkedUploadRegistry both use. * feat(native-chat): upload attached files into a paired server's chat store Main streams dropped or picked files (fs:uploadPathsToAgentSessionAttachments) and pasted images (clipboard:saveImageAsTempFile with agentSessionAttachment) into the chat's store on its paired server, reusing the file-import slice streamer and pinning every call to the pairing revision and server process. The browser client's paste takes the same route. * feat(native-chat): attach files to a structured chat on a paired server Pasted images, files dropped from Finder/Explorer and the file picker no longer refuse with "Local attachments are not available for remote sessions." in a structured chat on a paired server: they upload into that server's chat store and the chat gets only the stored path. Dropped files show as pending chips at once so Send waits for them; every attached item carries the server, pairing and chat it was stored for, and a send to anywhere else drops it with the existing "changed hosts" notice. Chips and transcript images read stored files back through the server, never from this machine's disk. An older server gets an update notice instead of an upload. The terminal-backed chat keeps refusing. * test(native-chat): type the attachment test doubles without bare casts * refactor(native-chat): keep the clipboard RPC on its own upload bookkeeping The chunked upload registry stays for the chat attachment store only, beside it. * feat(native-chat): claim chat attachments when the host admits a message A message's references into the host's attachment store are claimed in the same journal transaction that makes it durable (a direct send, a queued draft, a /clear carry). The sweep deletes an upload only after marking it in that database while no claim exists, so a send and a sweep can never both win: a client's send that names an expired or foreign upload is refused whole with the new attachmentExpired reason, before anything is recorded. This replaces scanning chat transcripts for paths. The store is now <root>/<upload id>/<name>, refuses uploads for chats the host does not hold, caps stored names in UTF-8 bytes and avoids Windows device names. The preview read goes through the protected bounded read and the request's reply budget, so an image too large for one reply is refused instead of closing the connection. * feat(native-chat): let structured Claude read the host's chat attachment store Attached non-image files live outside the workspace; the store root is passed as an added directory so the agent reads them without asking, the common pattern. * fix(native-chat): skip a dropped folder before staging walks into it * fix(native-chat): pin the browser client's chat paste to its server Every upload call goes to the environment the paste was meant for and checks its pairing and server process, as the desktop upload does; a re-pair ends the upload instead of storing the image on another server. * refactor(native-chat): let the host's claim be the only attachment check The composer no longer records which server and pairing each attachment came from, and Send no longer strips attachments it judges foreign: the host refuses an expired or unknown stored path when it admits the message, which also covers retries, Stop-restores and queued edits the client check missed. Previews of stored files route through the chat's own server by path. Removing a file's chip while it uploads now keeps its @path out of the draft. A dropped or picked file shows its name and kind on its chip while it uploads, with no separate progress toast, and files that did not attach are named in one notice. A rich-text paste into a chat on an older server no longer shows the update notice beside the pasted text. * chore(native-chat): keep the claim hook beside the draft consume and fit the line budgets The submission's claim runs in the same append hook as a draft's consume, owned by the queued-message collaborator, so the journal store stays within its size limit. Formats the new files and regenerates the runtime-required English catalog. * test(native-chat): pin the attachment store grant in the Claude launch, restore upload result types * fix(native-chat): create the attachment store at install and claim only exact store paths Claude drops an added directory that does not exist when it starts, so the store root is created when the host installs it, best effort. A commit keeps its upload in flight until the rename lands, so a sweep never takes a slow upload's part file. Only a path under this host's exact store root is claimed and required; any other mention of a store path is plain text and never refuses the message. * refactor(native-chat): reuse the newer-Orca notice and drop leftover attach options An older server's refusal to store attachments uses the notice every other write it refuses already shows, so one string fewer in every catalog; the attach callbacks lose an options parameter no caller passes since the provenance check went. * fix(native-chat): give a message refused for an expired attachment back to the composer Sending the same message again can never bring the attachment back, so instead of a Retry that always fails the message's text and images return to the composer, as a Stop's withdrawn message does, and the notice says what to remove. * fix(native-chat): keep uploads across a prompt, say why a file did not attach An upload that finishes while a prompt card has replaced the composer lands in the scope's attachment and draft caches, which the composer reads back when it returns, instead of the unmounted composer. The single failure notice names the cause the files share, such as the size limit. Rich text pasted into a chat on an older server no longer flashes an image chip: the chip waits for the server to take the image. A picked command waits, as Send does, while an attachment is still uploading. * fix(native-chat): keep a paste's upload across a prompt; pin the Claude grant hand-off A paste uploading into a paired server's store keeps its chip live when a prompt card unmounts the composer, and its result is filed for the composer's return, as a drop's is. A failure cause ending in full-width punctuation gets no extra full stop. A runtime test pins that the store root reaches Claude's launch resolver through the adapter. * test(native-chat): type the grant hand-off captures without casts * test(native-chat): pin that the composer files a paste's upload across a prompt * fix(native-chat): retry a held rename at commit and keep cut names Windows-safe A commit renames the part file with the Windows retry the app uses elsewhere, so an antivirus or indexer holding it briefly no longer loses the upload. A name cut to the byte limit is stripped of trailing dots and spaces again. The upload pump's result cast states why it holds. * fix(native-chat): keep attachments still on their way in the pane's attachment cache A prompt card unmounts the composer, and an attachment still saving or uploading lived only in that composer: one that came back showed no pending chip, so Send went out without the file and the file then landed in the next draft. Pending chips now live in the pane's attachment cache beside settled ones and settle there whichever composer is showing, so Send waits for them after a remount too. A stored file's @path goes into the pane's draft cache, which keeps it through an input-method composition and a remount. A rich-text paste's image is registered as a hidden pending chip before the server is asked, so Send waits for it from the start, and it shows only once the server takes it. * test(native-chat): a dropped file still uploading stays pending across a remount * fix(native-chat): reveal a rich-text image in a composer that came back; keep caret inserts A rich-text paste's held image is revealed through the pane's attachment cache, so a composer a prompt card remounted during the server's answer shows it rather than holding Send on an invisible chip. A stored file's @path goes in at the caret while the composer is showing and not composing, as every other attach does; only mid-composition or after a remount does it go to the pane's draft. * fix(native-chat): never evict a pane's attachments while a composer shows them or one is on its way The pane attachment cache is now where chips live, so its 128-scope bound passes over a scope a mounted composer subscribes to or that still holds a pending attachment, and evicts the oldest scope nobody uses instead. Protected scopes are bounded by mounted composers and attachments in flight. * fix(native-chat): never evict the scope just written when every older one is in use * fix(native-chat): keep the attachment store out of the RPC method table's imports The method table is imported far and wide (the SSH relay's CLI included), and the attachment RPCs pulled in the store, whose preview read loads modules that read fs constants at load: any test that mocks fs/promises without them failed to import. The installed store now lives in a small registry module; only the host wiring loads the store itself. * refactor(native-chat): the submission hook gets its own module Merging main added the operation receipt to the queued-message insert, which put journal-queued-messages.ts over the 300-line limit. The submission hook only uses the queued-messages object's public methods, so it moves out. * refactor(native-chat): fit the attachment sweep and paste upload into main's line budgets After merging main, the record store and the clipboard handlers each ran a few lines over the 300-line budget. The sweep now asks the record store whether a chat is recorded through the two lookups it already has (readable or unreadable) instead of a new id-listing method, and the paired-server paste upload lives beside the other attachment uploads. No behavior changes. * fix(native-chat): pending chips sit beside main's saved draft, which keeps server-stored images Main now saves a composer's draft (text and settled images) in one store that survives a reload or quit. This branch kept every chip, settled or not, in its own pane cache. After the merge the saved draft owns settled images and the pane cache holds only chips still on their way: they come back pending in a composer a prompt card remounted, hold Send, settle into the saved draft whichever composer is showing, and are never saved themselves, so a restored draft cannot bring back an upload as if it were attached. An image a paired server stored for the chat is saved with the draft as the real image (a pasted one is no longer turned into a "Not kept" placeholder), and the restore check leaves it to that server, whose claim at Send refuses one it no longer holds, instead of asking this machine's disk for a path that only exists on the server. Also: previews and the pending subscription move into their own small hooks to fit the line budget, the "local attachments" notice moves to the composer-target module so the attachments hook no longer pulls in the upload module's imports, and tests follow main's new mocks. * test(native-chat): give the claims test main's provider handle and message source * fix(native-chat): a paste still uploading dies with its tab or workspace A pasted image still uploading to a paired server sits in the pending chip cache, which by design outlives the composer and its tab. Removing the workspace deleted its saved drafts but left those chips, so the upload finishing afterwards recreated and saved the deleted draft. With the workspace's tabs gone it had no owner either, so no later workspace cleanup could ever remove it. A pending chip now records the owner its draft would have, resolved the way the draft store resolves one when the chip is added. Workspace removal and a user's tab close drop pending chips by the same owner, conversation and tab matches they delete drafts by, and a chip that settles after its chat's tab closed writes its draft under the owner it was begun in. * fix(native-chat): the claim type uses the journal's own SQLite type * refactor(native-chat): read the pending-image handlers off the composer's attachments Main's multi-file picker (#23956) and this branch together took NativeChatComposer over the 400-line limit; reading the two pending-image handlers the way the neighbouring ones already are keeps it at the limit. * test(native-chat): pass the launch-args resolver main now requires in the attachment grant test * fix(native-chat): grant the attachment store beside folders the saved Claude Arguments add Main (#25721) now builds additionalDirectories from the user's --add-dir arguments; this branch's attachment-store grant replaced that list instead of adding to it. Both are granted now. Moves the Claude session-id derivation to its own module to keep the resolver under the line limit. * test(native-chat): run the attachment store's SQLite tests in the Node runtime project * test(native-chat): follow main's attach ownership recheck (#25749) in the attachment tests * test(native-chat): a named upload chip still shows when the agent takes no images * fix(native-chat): a paired-server image paste failure keeps its error apart, as main's notice card does * fix(native-chat): the attachment sweep reads the chat list directly now that records have no import still owed |
||
|
|
d68f3bb5b2 |
fix(ssh): bind reattached SSH panes into the target's own session partition (#26088)
* fix(ssh): bind reattached SSH panes into the target's own session partition A relay reattach persisted the pane binding into the local partition, minting a minimal copy of every reattached SSH tab there. When nothing saved over it (a reconnect with no window open), the next startup kept that copy and skipped the SSH partition's rows for the workspace: its open editor tabs were missing for a launch and its agent-resume records were dropped (STA-9544). Bind into ssh:<target>, where the spawn bound the same pane. * test(e2e): run the reattach home-partition spec only in the Docker SSH lane * test(e2e): keep Electron running after its last window closes on Linux in the reattach spec * test(e2e): sync the SSH reattach spec on state, not sleeps Waits for the SSH partition to persist the open file and for main's SSH state to report the reattached connection, instead of fixed 3s/30s sleeps that could pass vacuously on a slow reconnect. |
||
|
|
40d35fb7bf | test: run real SSH command contracts under Node (#26226) | ||
|
|
2584ed906e |
feat: play video and music files in mobile previews (#26148)
Play workspace video and music files on mobile using bounded authenticated downloads and native media controls. Adversarial review fixes cover rewritten files and Android policy tests. Includes iOS simulator screenshots and a playback recording in PR #26148. Co-authored-by: ChangJun Park <40492343+ckdwns9121@users.noreply.github.com> Co-authored-by: Lirone Levy <lirone88@outlook.fr> Co-authored-by: dupi <david.li.du@gmail.com> Co-authored-by: John Cusack <5961784+John-Cusack@users.noreply.github.com> |
||
|
|
5cafefe726 |
Phase 3: every SSH host runs a managed Orca server (orcad), replacing the relay (#24863)
* Revert "revert: take the 26 Phase 3 (#16741 port) PRs back out of main (#24559)" This reverts commit |
||
|
|
4077fb4a8d |
fix(ci): actually exclude Electron probes from the headless node-server lanes (#26093)
vitest 5 does not apply a CLI --exclude to inline projects, so the glibc floor container ran profile-state-writer-stall.electron.test.ts and failed with 'spawn xvfb-run ENOENT'. Resolve the selectors to files and filter them before vitest sees them. |
||
|
|
67dc092f18 |
fix(windows): keep generated skills and snapshots LF (#25972)
Pin generated skill JSON and Vitest snapshots to LF so Windows autocrlf checkouts pass byte-for-byte verification. Cover the checkout behavior with a real Git regression test. Co-authored-by: Shuhei Konno <shuhei.konno@gmail.com> |
||
|
|
0acf039b5d |
Keep native contracts on Node and wait for Git upgrade completion (#26034)
* Keep native watcher and supervision contracts on Node * Route real permission and new ledger contracts through Node * Synchronize handshake cleanup with forced termination |
||
|
|
acea59c6d8 |
refactor(native-chat): remove the records-file and per-chat journal imports (#26038)
* refactor(native-chat): remove the records-file and per-chat journal imports Every native chat user is on a build that already moved chat records and history into the app-wide database, so the one-time copies are dead code: the agent-sessions.json import and its owed-copy flag, the per-chat journal.db import and the write-queue hold behind it, and the pre-SQLite log.jsonl notice. * refactor(native-chat): remove what only the deleted imports used - JournalHostDatabase.unsyncedTransaction and the synchronous-pragma reset that undid it; JOURNAL_SYNCHRONOUS is now module-local. The stranded rollback test drives transaction() instead. - deleteUnpublishedJournalRows and its SQL, plus its test. - boundJournalStatusText. - Stale comments naming the removed first-use copy (queued messages, send refusal example, JOURNAL_SYNCHRONOUS doc). - Write-queue and readInOrder comments: a write also lags when issued behind one still waiting in line. - Resolved-append ordering test binds the sink from inside a running read, so a resolver that read at handover now fails it. * test(native-chat): run resolved-append in the SQLite runtime project |
||
|
|
d0b0f13b74 |
fix(native-chat): hold a queued message until the turn ahead opens (follow-up to the Stop-event plan, fixes STA-9348) (#25217)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused
Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.
* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered
A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.
* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget
* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it
The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.
* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card
* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows
Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.
One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.
The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.
Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.
* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones
A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.
* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction
The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.
* fix(native-chat): stop creating the unused queue pause table
The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.
* fix(native-chat): a Stop's pause never hides the restart pause
A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.
Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.
* test(native-chat): pin the Stop's no-resend, lift and held-card rules
- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
again" at one instant, before a queue ignoring the pause re-sends. They
now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
whether or not a person's turn lifts it; it now reads the Stop's pause
before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
queued before a rewind.
* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller
The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.
* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event
* test(native-chat): pin that Stop and Resume rows never reach apps or count as history
* test(native-chat): only a person's Stop event pauses the queue
* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop
Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.
* test(native-chat): a card held at a starting agent is checked before the Stop's timing
Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.
* test(native-chat): a released build keeps and folds a journal holding Stop events
Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.
* style(native-chat): format the Stop event changes
* test(native-chat): type the released build's exports through one checked helper
* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only
* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade
The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.
Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.
* fix(native-chat): a Stop that stops nothing new writes no Stop event
A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.
It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.
* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop
* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled
* fix(native-chat): any later Stop event ends a person's Stop pause
A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.
An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.
* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed
A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.
A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.
Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.
* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes
A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.
The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.
* fix(native-chat): a Stop still reads as yours after Orca restarts before the turn ends
Every stop that ends work now writes the Stop's event before it ends the child: a
person's close of the chat, an eviction (worktree teardown, orchestration stop, tab
cleanup) and the idle sweep's stop of a start that never landed. A stop that ends
nothing writes nothing, and quit writes none: its resume marker records why.
The turn-end write reads the latest Stop event where every turn row is built, so the
adapter's settle, the host's fallback and the relaunch's settle all agree: a turn a
person's Stop or close named, ending with no verdict of its own after that Stop, ends
as their cancellation. A relaunch's probe-bounded end is no earlier than a Stop that
found the turn running. When the provider refuses the interrupt and the turn runs on,
a refusal row answers the Stop, so a later crash still reads Failed; pressing Stop
again after a refusal is a new Stop.
* refactor(native-chat): a stop no longer carries its cause; the turn's end reads the Stop event
The cause of a stop was threaded in memory from each entry through the host's stop
step, the adapter router and each adapter's close onto the `ended` it settled with,
and Claude kept a per-turn copy of a Stop it sent. All of that is gone: adapters
settle a turn they cut as interrupted with no verdict, the host's fallback does the
same, and the one rule where a turn row is built (`turnEndAfterStop`) reads the
journal's latest Stop event to say whether it was a person's.
- `closeSession` / `disposeSession` take no cause; `ended` has no `stopCause`.
- Claude reads an error result after a person's Stop as their cancellation from the
journal's Stop event (through the event sink), not from a per-turn slot, and a
refused interrupt is the host's refusal row, not `withdrawTurnStop`.
- An owed wind-down keeps no cause: its retry's fallback reads the Stop event.
- The mutation context's Stop passes no cause: its step already wrote the event, and
the delivery loop's child-end reason is read back from it.
- A Stop pressed before its turn showed applies to the turn that opens under it,
unless a send a person made since was accepted.
* test(native-chat): a turn a later send opened is no Stop's that named no turn
* test(native-chat): the restart test's death proof carries its detail
* refactor(native-chat): a refused Stop leaves no record; a Stop only ever ends the turn it names
The stop-refused mark is gone: its tombstone kind, its fold, the clock-keyed match that tied it to
a Stop, and the exception that let a second press after a refusal write a new Stop. A Stop that
stops nothing writes nothing. A Codex refusal names a turn that is no longer its active one, and
the Stop names that turn, so the turn running instead never reads as the person's by its id alone.
* fix(native-chat): a Stop pressed before any turn showed stops only the turn opened next
A Stop that named no turn read as the person's cancellation for every later turn that opened
after it, until a send a person made was accepted. The queue's drain, orchestration mail and a
restart continuation send as the host, so a turn they opened long after, cut by a crash, read
"Interrupted" as if the person had stopped it. The Stop now applies only to the first turn
opened after it.
* fix(native-chat): an older Claude's error end after a Stop pressed before its echo reads Interrupted
Claude CLIs before 2.1.91 end an interrupted turn with an error result that names no reason. The
translator judged whether a person's Stop explained it by its own copy of the Stop rule, which
ignored a Stop that named no turn, so a Stop pressed before Claude echoed the send read "Failed".
The translator now writes such an end as interrupted with no verdict and no error row whenever a
person's Stop may name the turn, and the journal's one rule decides as it writes the end.
* fix(native-chat): a person's Stop and /clear each name why they end the agent
The host's mutation path ended the agent with one "recorded" ending for every caller, which read
back the reason of whatever Stop event the journal held last, however old. /clear writes no Stop
event, so its end took an unrelated earlier reason. Each caller now names its own: the chat's Stop
`user-stop`, whose event its own step wrote, and /clear `user-close`, the user replacing this chat.
* fix(native-chat): a host stop judges whether it ends work after the provider's rows land
A close, eviction or host stop decided whether it ended a running turn from the journal as it
stood, while the provider's own rows (the turn its echo opened) could still be in the session's
event sink. A close landing in that gap wrote no Stop event, so the turn it cut read as news. It
now reads after the sink drains, as a person's Stop does, through the same check; a drain that
fails or takes over a second reads working.
* fix(native-chat): a Claude Stop naming a turn that just ended still marks the follow-up it cuts
A phone names the turn it last saw. When that turn had ended and a follow-up was still unechoed,
Claude's Stop interrupted the follow-up and ended the child, but the Stop's event named the ended
turn, so the follow-up's turn the child's end cut read "Failed" under "Cancellation requested.".
A Stop that ends the provider's session ends whatever is in flight, so its event now names the
live turn or none, and a Stop that names none binds the turn opened next. Codex keeps naming only
the turn the Stop names.
The Claude Stop turn-end tests move to their own file, since the session-ending Stop suite is at
its line budget.
* fix(native-chat): the idle sweep reads working by the same rule as a stop's event
The sweep judged a chat resting while a send whose reply was lost was still unanswered, but the
stop's event writer counts that send as work. So the sweep evicted it and wrote an evict event,
which ends a person's Stop pause and let the cards behind it drain on their own. The sweep's owed
work now reads the main agent working the way every session list and the event writer do.
* test(native-chat): an aborted eviction's injected drain failure lands on the eviction's own drain
A host stop now drains the session's sink once to judge whether it ends work, so the tests that
fail the eviction's drain-published step skip that first drain.
* fix(native-chat): the idle sweep's rest writes no Stop event; it evicts a send that never echoes
The previous commit made the sweep count an unanswered send as owed work, which pins a chat whose
admitted send Codex never echoes forever, and the sweep exists to retire exactly that. That rule
returns. The sweep stops only an agent it judged resting, so its eviction now writes no Stop
event, whatever send it retires: a person's Stop pause holds through it.
* fix(native-chat): stopping a start that carries no send writes no Stop event
A host stop, eviction or close of a starting child wrote a Stop event whatever the start carried.
A start with a send already reads working, so the clause only mattered for a start with none,
which ends no turn and no send: its event only lifted a person's Stop pause and bumped the idle
clock, which is why the idle sweep had been changed to close the conversation in the same pass.
The clause goes and the sweep is #24072's again. The child's end still reads host-stop, as before.
* test(native-chat): a Stop's pause across a restart is tested with a restart that writes no event
The rig's restart closes the chat with an eviction, which now writes a Stop event when work runs
and so ends a person's Stop pause. "A Stop never hides a restart's pause" then passed with no Stop
pause left to hide anything. Those tests, and the pause-lift test whose dropped assertion returns,
restart as a process that dies with no close, which like a quit writes no Stop event, and assert
that both the Stop's and the restart's pauses are in force first.
* fix(native-chat): a host stop of a turn a person's Stop is still ending keeps that Stop's reason
An eviction or host stop that landed while a person's Stop or close was already ending the same
turn wrote a newer Stop event, and the turn's end reads only the latest, so the person's Stop of
that turn read as news. A host reason now writes nothing while a person's Stop still decides what
runs: the live turn it names or bound, or, with none, the turn a send opens next. The person's
own close still writes. The E2 tests now open and end the stopped send's own turn, as Codex does,
so the mail turn after it is not the turnless Stop's.
* fix(native-chat): an older Claude's error on a later turn keeps its error text after a Stop
The translator left an error result that names no reason to the journal's Stop rule whenever a
person's Stop named the turn or none, but the rule binds a Stop naming no turn only to the turn
opened next. So a real error on a later turn read "Failed" with its error text dropped. The
translator now asks the journal's rule itself (`personStopDecidesTurn`, the one core
`turnEndAfterStop` and a host stop's in-force check share), so the two cannot disagree.
* fix(native-chat): a Stop of a start that never landed binds no later turn, whatever sent it
A person's Stop pressed while the agent starts names no turn, and the send it stopped is
cancelled before it opens one. The Stop then bound the next turn anything opened (orchestration
mail, a restart continuation, the queue's drain, all of which send as the host), so a host
eviction of that turn wrote nothing and its crash or close read as the person's cancellation. A
Stop that named no turn now binds only a turn no send journaled after it opened: any send since,
of any origin and not refused, opens its own. The E2 test's mail send is accepted as Codex
accepts it, instead of opening the stopped send's own turn first.
* test(native-chat): a rewind's restated turnless Stop binds no turn opened after the rewind
A Codex rewind restates a person's Stop still in force after the turns it keeps, at a new
sequence, so by sequence alone it would bind the next turn opened after the rewind. A send
journaled after the restated row voids that binding (the previous commit), which this pins.
* fix(native-chat): a relaunch settles a person's stopped turn with no "stopped while in progress" row
After a restart, a turn a person's Stop ended reads "Interrupted after N" with the muted mark, but
the relaunch still added the error row saying the provider stopped mid-response, which a live Stop
never writes. The settle now skips that row when every turn it interrupts is the person's Stop's
by the journal's one rule; a crash nobody stopped keeps it.
* test(native-chat): the unexpected-exit settle's journal fake answers whether a person's Stop decides a turn
* fix(native-chat): a host stop whose sink drain fails reads the journal as it stands
A host stop drains the session's sink before judging whether it ends work, and a failed or slow
drain read as working. So an eviction of an agent at rest wrote a Stop event that ended nothing,
which lifts a person's Stop pause, and a close wrote a person's event naming no turn. The drain is
now best effort: the stop goes ahead either way and only its record is at stake, so a failed or
slow drain leaves the journal's read as it stands. A person's Stop keeps its own rule.
* fix(native-chat): a Stop that named no turn applies only to a turn a send it stopped opened
A person's Stop pressed before any turn showed names no turn. It bound the first turn opened
after it, then (
|
||
|
|
64bb9373da |
Claude account profiles: dormant WSL guest setup (Step 3 of 4) (#24384)
* feat(claude): add dormant profile setup and history sharing
* fix(claude): make profile setup one gated, typed, fail-safe entry
Review round 1 of the dormant profile setup found that the pieces could
be called without their safety checks, that one failed write or an
unreadable bookkeeping file could silently stop sharing for good, and
that Windows prompt history could bring back history the user cleared.
- One entry, provisionClaudeAccountProfile: the profile gate (namespace,
no linked components, outside ~/.claude and ~/.config/claude, and an
ownership marker beside the home naming the account and target) runs
first and refuses before creating anything; then history sharing,
config provisioning, and the hook install after the settings merge.
Results come back per surface with closed warning codes instead of
message text.
- The sharing ledger is keyed by surface name, records a value only
after its write succeeded, and an unreadable ledger starts empty and
is rewritten instead of blocking every surface.
- The profile state file goes through the same locked writer as folder
trust (Claude's <file>.lock plus the in-process queue), generalized as
updateClaudeGlobalConfig. Onboarding and trust are still applied when
the personal state file is unreadable.
- WSL descriptors build guest POSIX paths; the state-file path style
follows the injected platform.
- Orca's managed statusLine has one owner in a profile: the settings
merge never shares it, a user's own statusLine is shared over it, and
the profile installer follows the default home's slot so a default
opt-out reaches every profile. remove() takes the same destination;
the remote installer cannot accept one.
- Prompt history compares file identity (bigint dev+ino) on every
platform, never drains the shared file into itself, drains retained
copies in generation order, never reuses a stale cursor, and on Windows
keeps a replaced default's old copy aside instead of replaying it.
Directory merges keep going past a failed entry.
* fix(claude): share the user's own hooks and keep merged history whole
A user's own Claude hooks in ~/.claude (notifications, formatters) did
not run under a managed account, because the whole hooks key stayed
private. They are now shared like any other settings key: Orca's own
hook entries and its managed statusLine are stripped from both the
personal value and the profile's current value before the per-key
ledger comparison, so they never travel through the merge and never make
the key look user-owned. Orca entries already in the profile are kept on
write, and the profile hook installer adds them on top as before.
Prompt history: merged bytes that lack a final newline are terminated,
so Claude's next record no longer fuses onto the last merged line. When
a CLI rewrote the profile's history file (old records plus new), only
the lines past the part it shares with the default history are added,
instead of the whole file again.
* fix(claude): close review round 2 gaps in profile setup
Hooks and statusLine sharing:
- When ~/.claude holds only Orca's hook entries, the user's shared hooks
now read as an empty value instead of a missing key. Removing the
user's last own hook in ~/.claude therefore reaches profiles that
never edited it, and deleting the only shared hook inside a profile
stays deleted.
- A custom statusLine Orca shared, and the profile never edited, goes
away when the default home drops it. When a shared custom line
replaced Orca's line in a profile, the profile's statusline marker is
dropped so Orca's line comes back once the default returns to it; a
profile that opted out stays opted out. No other key gains deletion.
- install/remove/getStatus with a profile directory refuse when it is
the default home, or its settings.json resolves to the default one,
instead of editing System Default's hooks and opt-out state.
- The profile statusline rule reads the default settings under the
userHome passed to the setup entry, not os.homedir().
Profile state and ownership:
- A malformed `projects` value skips only folder trust (new warning
code trust-refused); onboarding and shared keys still apply.
- The ownership marker stores only host-local facts (account, runtime,
distro). The execution host id is the caller's view of the host, so
it stays in the in-memory descriptor and is not compared.
Prompt history interruption paths:
- With no cursor yet, a retained copy starts past the bytes it shares
with the default history, so an interrupted share no longer replays
the whole history.
- A retained name for the shared file itself is removed with its cursor
instead of lingering until a later scrub makes it look new.
- The Windows link record is read three-state: unreadable stops the
share instead of reading as "no link". If the record cannot be
written after linking, the fresh link is undone.
- An unreadable retained copy is reported and no longer blocks linking.
* build(cli): list the new Claude hook modules in the CLI project
hook-service.ts and hook-settings.ts are compiled into the packaged CLI
project, which lists every file explicitly. The statusline policy and
profile destination modules they now import were missing, so the CLI
typecheck failed with TS6307. The CLI still loads hook-service through
the existing managed-agent-hook-controls build entry, which bundles
both modules; neither imports electron.
* fix(claude): close review round 3 regressions in profile setup
- A profile whose hooks hold only Orca's entries and that sharing never
recorded is no longer treated as a user edit, so the user's first own
hook in ~/.claude reaches it (for example when the profile was set up
before ~/.claude had any hooks).
- A retained prompt-history file is removed as a second name for the
shared file only when the default history does not itself link to it;
otherwise it holds the only copy and is kept.
- Default-home checks compare file identity: the profile hook
destination check uses device and inode, and the profile/default
separation check resolves on-disk case, so a case-only alias of
~/.claude is refused on case-insensitive filesystems.
- A test pins that an unreadable leftover session tree no longer blocks
linking.
* fix(claude): let shared keys leave a profile when ~/.claude drops them
QA found that removing a setting from ~/.claude never reached a managed
account: deleting the whole `hooks` block left the user's hook running
there. Only statusLine followed the default away.
Every shared key now follows the same rule through the existing per-key
ledger: when a key disappears from ~/.claude/settings.json (or
mcpServers/theme from the personal state file), it is removed from the
profile if the profile still holds exactly what Orca last shared. A
value changed inside the account is kept. Keys Orca never shared,
including denylisted ones, are never touched. Deleting the whole hooks
block removes the user's shared hooks and keeps Orca's own entries. A
missing source counts as empty; an unreadable source removes nothing.
* feat(claude): add dormant profile routing and account consumers
* fix(claude): drop the dormant profile selection RPC; clients negotiate by capability
Restores the inline mobile allowlist so its source-scan guard sees every
accounts.* method again, and the generated params catalog to generator order.
* fix(claude): guard the claude shell function and honour a hand-exported config dir
The function is defined only in a routed pane where claude is a real
executable (the codex function's guard), re-reads the pointer only while
CLAUDE_CONFIG_DIR is unset or still Orca's injected twin, accepts Git Bash
drive paths, and starts on its own line after the fish/PowerShell codex text.
* fix(claude): spawn-time profile env, total account listing, setup at lifecycle triggers
Round-1 review fixes for the dormant profile routing:
- Panes get the selected profile's CLAUDE_CONFIG_DIR plus an Orca twin at
spawn, so nested shells and scripts inherit the account; System Default
injects nothing and its home is the inherited CLAUDE_CONFIG_DIR.
- An absent routing owner is System Default, never a throw; AI Vault and
session-search scans receive profile roots from their parent, and the
capability is advertised only where an owner is installed.
- Account listing never throws: per-account readiness, a stale pointer is
republished in the background and reported on the snapshot.
- Profiles are set up at select and startup; a launch only sets up one that
never was, and a worker fault on a prepared profile is a warning. The
Claude version probe is cached per binary identity.
- Pre-trust goes through the existing deadline- and realpath-guarded writer
against the launch env's profile config.
- Skill discovery keeps a caller's Claude root and a broken Claude selection
no longer fails other providers.
- The durable record carries a provider-neutral launchAccountHome, read
through one helper by the launch fallback and the model catalog.
* test(claude): pin the version-probe cache, launch-account record and temp-home readers
* test(claude): pin dormant bash rc text alongside fish and PowerShell
* test(claude): read the fish launch init without a nullable index
* fix(claude): withdraw the profile pointer when a selection cannot be published
A pointer left naming the previous account would launch it silently; a
missing pointer makes the claude function refuse visibly. A newer selection
that raced the failed one keeps its pointer.
* fix(claude): read the fish profile pointer with read -z for fish older than 3.4
Shell tests skip system config and abort unless claude resolves to the fake.
* fix(claude): only the newest publish withdraws the pointer; total config dir lookup
- An overtaken publish that fails leaves the newer selection's pointer.
- The runtime config dir falls back to the legacy home for an unresolvable
account or a WSL target, so skill roots never fail for other providers.
- WSL guest reader roots merge verbatim, never realpathed on this thread.
- History readers include ~/.claude, where step-1 setup pools profile history.
- System Default ignores a config dir an outer Orca injected (twin-marked).
* fix(claude): System Default launches and probes use the structured create resolver
A Claude agent-env CLAUDE_CONFIG_DIR the create path stored is now the home
the launch pins and the model probe accepts.
* test(claude): type the System Default launch record as an agent-session record
* fix(claude): install profile hook scripts under the setup job's home
A worker thread's os.homedir() ignores its own env, so the hook and
statusline scripts now go under the home the job names. The worker test pins
the process HOME to a sentinel, refuses to run unless the worker sees it, and
asserts nothing lands there.
* test(claude): skip shell cases whose shell the runner lacks
* fix(claude): remove env vars in the PowerShell claude function instead of setting null
On .NET 9+ (pwsh 7.5+) SetEnvironmentVariable with $null creates an empty
variable, so stripped auth vars reached claude as empty strings and the
restore left CLAUDE_CONFIG_DIR empty in the user's session.
* feat(claude): add dormant WSL guest profile setup
* fix(claude): open WSL panes without guest calls and coalesce same-profile publishes
A WSL pane now gets the same non-throwing, guest-free spawn env as a host
pane; only select, startup and Claude launches publish into the guest.
Overlapping publishes of one target share the newest publish while the
selection still names the same profile, instead of failing as superseded.
Publish issues name their WSL distro and drop out when the target is no
longer routed. A late inspect from an older selection no longer replaces
the newer one's verification, a failed guest request evicts the cached
guest, and readiness is derived per account from the guest's owned homes.
* fix(claude): roll back only the target whose selection failed
With profiles, a failed select or remove republishes just its own target
instead of running startup over every WSL distro, and a rollback failure is
logged instead of replacing the error that caused the rollback.
* fix(claude): scan WSL profile history only in running distros
Vault and usage scans pass Claude profile roots through the same
running-distro filter as every other WSL root, so a stopped distro's UNC
paths are never walked.
* fix(wsl): ship the Claude profile helper only in the WSL bundle dir
The helper only ever runs inside WSL from the desktop, so it moves out of
the SSH relay artifacts (no upload, no relay version change) into
out/relay/wsl beside the other WSL-only guest bundles. The three WSL bundle
resolvers share one candidate list.
* fix(wsl): refuse old glibc before downloading, and keep the shared download per caller
The pinned Node runtime needs glibc 2.28, so a distro below the floor is
refused before any download with a message naming both versions, as SSH
hosts are. The shared download again owns its own deadline and each caller
waits on its own signal, and the OpenCode reader keeps its architecture
error text.
* fix(claude): bound each WSL guest operation and run the helper through the WSL runner
A cached guest no longer carries its 180 s preparation deadline into later
requests. The helper runs through runWslProcess (stdin payload, WSL_UTF8),
the distro is confirmed running once per preparation and once per request,
a failed `claude --version` probe continues with an unknown version like
native setup, the helper resolves from the WSL bundle dir, and the guest
entry decodes stdin once so split UTF-8 survives.
* test(claude): cover WSL profile pre-trust routing and its deadline
* refactor(claude): drop WSL refresh cleanup that the failed publish's withdraw already does
* fix(claude): catch rollback failures only when profiles route the selection
With the gate off, select and remove surface the rollback error exactly as
before; only profile routing logs it and keeps the original error.
* fix(claude): give every WSL pane a guest-relative Claude profile pointer
WSL panes now always carry `~/.local/share/orca/claude-profiles/selected-wsl`,
which the bash/zsh and fish claude functions expand against the guest $HOME
at each invocation, so a pane opened before Orca has met the distro still
follows the selected account instead of falling back to ~/.claude. Absolute
pointers are untouched, PowerShell is unchanged, and a missing pointer file or
profile still refuses visibly. CLAUDE_CONFIG_DIR is set at spawn only when the
selection resolves without a guest call.
* test(claude): assert a missing guest-relative pointer refuses with a visible message
* test(claude): type the WSL runner mock in the transport test
* fix(claude): route only WSL distros that hold an Orca account, and re-derive their publish
A WSL distro is routed only while host settings hold an Orca Claude account
for it, decided from settings with no guest call. An unrouted distro behaves
as before profiles: its panes get no pointer or profile env, and a Claude
launch is System Default with no guest prepare. A distro that loses its last
account has its pointer withdrawn best-effort so older panes stop launching
the removed account.
A routed distro without a current publish (for example stopped at startup)
gets one non-blocking background publish from its next pane spawn, coalesced
per target; its failure stays that distro's issue and a later success clears
it. A late setup result from an older publish no longer replaces the newer
selection's verification. The owner contract moves to its own module so the
routing service stays under the size limit.
* fix(claude): read WSL profile history in native chat and adoption only in running distros
Native chat resolves Claude transcripts from host roots first and reads WSL
profile roots only after a miss, filtered to running distros like Codex's WSL
homes. Structured adoption candidates go through the same filter.
* fix(claude): target registration rollbacks and keep their errors in profile mode
A failed add or re-authentication rolls back only the account's own target.
With profiles, a failed re-authentication rollback is logged instead of
replacing the original error; with the gate off both behave as before.
* fix(claude): spell the guest pointer location once and keep set -u safe
The guest helper, the withdraw script and the pane pointer all derive from
one home-relative constant, and the posix claude function reads ${HOME:-}
so `set -u` with HOME unset refuses cleanly instead of aborting.
* test(claude): cover the IPC preflight and daemon WSLENV paths for WSL profile env
The renderer preflight is tested for wsl.exe and Windows shells with a \\wsl$
cwd (which always launch wsl.exe) and with the gate off, the daemon launch
plan imports the pointer and profile home without a WSLENV flag, and the
Windows launch test uses the guest-relative pointer production sends.
* fix(wsl): report why the guest runtime failed, with download context and trimmed stderr
The install's promote output is classified with the SSH classifier, so a
self-test failure shows the exit code and the loader's words (for example a
missing libstdc++ on Alpine) and a security-software change is named. A failed
runtime download says it was Orca's Node runtime for WSL, while a checksum
mismatch keeps its own text. Guest stderr is trimmed before it reaches a
refusal message.
* test(claude): pin that pointer retirement never runs for host targets or with the gate off
* test(claude): give the routed WSL preflight fixture its required authMethod
* fix(claude): let the pane-triggered WSL publish repair a distro stopped at startup
"Distro not running" is now a typed refusal: it never withdraws the pointer
(the distro's last pointer cannot be stale, and a withdraw racing the boot
could delete a valid one) and never records a distro issue. The background
publish a pane fires now waits a few seconds for the pane's own spawn to boot
the distro, probing three times, and is dropped silently and re-armed if the
distro stays down. It joins any publish already in flight for that target
instead of preparing the guest a second time. Per-target generations and
pointer-write ordering move to ClaudeProfilePointerQueue so the routing
service stays under the size limit.
* fix(claude): remove the last selected WSL account without a guest publish
With profiles, removal writes the account list and the selection in one
update, so a distro losing its last account is already unrouted when it syncs
and its pointer is retired best-effort. Removal no longer needs the distro to
be running or able to run Orca's runtime. The gate-off order is unchanged.
* fix(claude): keep native chat's legacy Claude roots first and unfiltered
Only roots added by WSL profiles are read after a miss and filtered to
running distros; a host CLAUDE_CONFIG_DIR on a \\wsl$ share is searched first
and unfiltered, as before profiles.
* test(claude): cover stopped-at-startup repair, launch join and last-account removal end to end
* test(claude): assert no running probe before the pane has had a turn to boot the distro
* fix(claude): let user-initiated profile work boot an idle-stopped WSL distro
WSL distros idle-stop on their own, and the legacy path boots them with its
spawn or \\wsl$ write. With profiles on, a Claude launch, a select, a remove,
a failed-change rollback and the retire after removing a distro's last
account now skip the running pre-check and let their first bounded guest
command (`wsl -d <distro> --exec ...` through runWslProcess) boot the
distro. They refuse only if that command fails, with wsl.exe's own reason,
for example a distro that does not exist. Startup, the pane-triggered repair
and the history readers keep the running pre-check and its typed refusal, so
background work never boots a distro. With the gate off nothing changes.
* fix(claude): let startup join a launch or select already publishing a WSL distro
Startup no longer overtakes a user's in-flight publish for the same target,
so a launch that is booting an idle-stopped distro is not handed startup's
"not running" refusal.
* fix(claude): remove accounts of a WSL distro that no longer exists, and name the helper once
wsl.exe's own failures (exit 0xFFFFFFFF, empty stderr, the diagnostic and its
WSL_E_* code on stdout) are now read by one shared reader used by the git
runner and the WSL profile transport, so profile refusals show wsl.exe's
message. WSL_E_DISTRO_NOT_FOUND becomes ClaudeProfileHostMissingError: with
profiles, removing an account from a distro that no longer exists keeps the
removal and logs a warning, while select and launch still refuse visibly.
The helper's file name is defined once in shared/relay-artifacts.ts and used
by the relay build and the transport.
* fix(claude): give plain fish tabs the claude function through the codex hand-off
Main now gives a plain fish tab Orca's codex function through a vendor_conf.d
snippet instead of a -C init. The claude function only rode the -C path, so a
plain fish tab would not re-read the account selection per invocation once
profiles are on. Define it at the first prompt beside codex; it stays empty
while the profile gate is off.
* fix(claude): share personal rules, themes, workflows and keybindings into account profiles
A managed account launches Claude with its own config folder, so user-level
rules/, custom themes/ (which a shared `custom:<slug>` theme points at),
personal workflows/ and keybindings.json silently stopped applying. Link the
three directories like skills and commands, and copy keybindings.json with the
same edit-preserving ledger as CLAUDE.md. routines/ stays unshared: routines
belong to the claude.ai account and the folder holds per-run state.
* test(claude): wait for the running child to read its account before switching
The test switched the selection after a fixed 20 ms, so under load the backgrounded claude
had not yet read the pointer and picked up the new account. The stand-in now marks when it has
started, and the test waits for that mark (bounded) before switching.
* fix(claude): accept WSL setup warnings for every shared Claude file
The guest reply schema listed CLAUDE.md by name, so a warning about the newly shared
keybindings.json would have rejected the whole reply. It now takes the shared-file list
from provisioning, like the shared folders.
* fix(claude): import the personal CLAUDE.md into account profiles instead of copying it
Claude also loads ~/.claude/CLAUDE.md as a parent folder's memory for any project under home,
so a copied account CLAUDE.md made every such session read the user's instructions twice
(checked live with Claude 2.1.288). An @~/.claude/CLAUDE.md import resolves to the same real
file, which Claude loads once from home, from projects under home and from folders outside it.
* refactor(claude): simplify account profile setup toward the prior art
- Windows keeps each account's history private; drop the hardlink, link
record and conflict-copy machinery that only Windows reached.
- Share hooks and statusLine as ordinary settings keys: Orca writes the
same entries into every folder, so the installer finds them present.
Drops the Orca-entry carve-out, the per-profile statusline follow
logic and its marker.
- Unreadable ledger is just an empty ledger.
- Share from the user's own CLAUDE_CONFIG_DIR when they set one (marked
so Orca's injected value is never mistaken for it), and refuse a
profile at or around it.
- Pin the one canonical profile path spelling in a test.
* refactor(claude): route launches through one account router, superset-shaped
Replace the routing service, owner interface, setup worker thread, reader-root
merging, persisted launch account and capability string with one
ClaudeProfileRouter: the pointer is written first and setup runs best-effort
after it (superset's order); a missing pointer means System default.
The claude shell function re-reads the pointer on every launch, refuses only
a selected account whose folder is missing, and prints a note when the user's
own CLAUDE_CONFIG_DIR overrides the selected account in that terminal.
Still dormant: claudeProfileRoutingEnabled() is false.
* test(claude): type router test settings instead of casting
* fix(claude): run account setup on a worker thread, never Electron main
publish() writes the pointer and starts setup in the background, so neither
startup nor an account switch blocks on a history merge. Each setup runs in a
one-shot worker (the profile-state backup worker's pattern); one setup per
account at a time, reused by later requests. A launch waits only for a folder
that was never set up, and refuses with a clear message if that setup fails.
* fix(claude): do not await the synchronous pointer publish
* refactor(claude): route WSL distros through a small guest router on the Step 2 shape
Replaces the WSL owner/transport/guest-inspect stack with ClaudeWslProfileRouter:
publish writes the guest pointer with one sh command and kicks Step 1's setup
best-effort; prepareLaunch checks the folder over the distro share and waits only
for a never-set-up folder; preparation returns main's WSL shape, so trust, rate
limits and readers need no new code. Setup runs as Linux in the guest on Orca's
pinned Node via a bundled helper (argv in, exit code out), without hooks.
Restores OpenCode's WSL runtime prep, git's wsl-host-failure, wsl-runner,
workspace trust, readers and account selection/registration to Step 2.
Names the guest pointer per Orca build so dev and packaged never share it.
* test(claude): give the routing launch test the merged resolver deps and handle shape
* test(claude): type the WSL routing mock's original() without an inline import()
* fix(claude-accounts): dedupe merged prompt history, drop drained copies, link setup folders by path
- Prompt-history drain appends only lines the shared file lacks, so a purge never re-adds lines.
- A set-aside history copy whose saved offset reaches its end is deleted on the next run.
- Setup folders link to the default home's own entry, not its resolved target.
- The profile gate and folder creation run once, in provisionClaudeAccountProfile.
- installHooks receives only configDir; drop a duplicate test key that fails CI.
* fix(claude-accounts): refuse a routed resume whose transcript is in another account; zsh claude function; setup timeout
- With account routing, a chat resume checks its transcript is in the launch folder; a missing one
with a stored leaf refuses with historyInOtherAccount instead of starting fresh.
- The launch folder of a selected account comes from prepareLaunch(); the resolver stays for System default.
- zsh panes get the claude function like bash, fish and PowerShell (empty while routing is off).
- The setup worker is terminated after 60 s so a later launch can retry.
- Document that the setup marker means setup started, not finished.
* fix(claude-accounts): write the WSL account pointer before a launch returns; one relay bundle candidate list
- prepareLaunch awaits writePointer, so a missing or stale guest pointer cannot run another account.
- Startup's WSL republish runs inside serializeMutation, like rollback.
- relayBundleCandidates takes 'wsl'; the hook relay, browser relay and Claude helper use it, and
wsl-relay-bundle-dirs.ts is gone.
- One setup-marker path helper for host and WSL; the guest pointer path is home-relative and only
the pane value carries '~/'; drop a no-op esbuild external.
* fix(claude-accounts): refuse a routed resume only when the transcript is found in another folder
A transcript found in no known folder keeps the old stored-leaf resume.
* fix(claude-accounts): a WSL launch writes the pointer for the selection current at write time; bound the pointer read
A selection made while a launch waited on setup was overwritten by the launch's stale account.
A hung \\wsl.localhost read no longer stalls startup's serialized publish.
* fix(claude-accounts): record installed hooks as Orca-shared; skip symlink tests on Windows
After Orca installs its hooks into an account, record the account's hooks in
the settings ledger so a later run can still bring the user's own hooks in.
Tests that create real symlinks now skip on Windows.
* fix(claude-accounts): trim the which-account file in the PowerShell claude function
Co-Authored-By: Claude <noreply@anthropic.com>
* test(claude-accounts): spell the user's own config folder as an absolute path on every platform
Co-Authored-By: Claude <noreply@anthropic.com>
* test(claude): skip the POSIX-only WSL profile test on Windows
A WSL profile's data root is a POSIX path, so building one from a Windows
temp dir fails the absolute-path check there.
---------
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
||
|
|
2f49377425 |
feat(native-chat): Grok as a structured chat over the Agent Client Protocol (#25225)
* Leave a stopped turn's running tools to the agent's own end
When another writer settles the open turn (a person's Stop), the assembler
now only stops that turn's text and cancels its pending prompts. Running tool
calls stay the agent's: a progress update or completion it reports after the
Stop lands as reported, and whatever is still running settles at the agent's
turn end for that turn, the next turn's open, or the session's end.
An agent's end for an earlier turn while a newer one is open no longer clears
the open turn's activity line or ends its anonymous reply. An unnamed end right
after a Stop ends the stopped turn instead of being dropped. The test rig's
restart no longer writes the dead assembler's window text, matching dispose.
* Pin that a stopped turn's running tools hold budget until the agent's end
* Type the stopped turn's tool progress update as a tool body
* List every event the assembler hands to the decision step
The type-aware lint requires an exhaustive switch with no default case.
Also retitle a Stop test to say what it asserts.
* End a running call as its turn's journal row ends
A call still running when its turn ends takes the state of that turn's
row: a row another writer settled first (a person's Stop) stands, so its
calls read interrupted whatever the provider's later end reports. The
no-ending path that settled calls from the Stop row is gone, since a Stop
now leaves running calls to the provider. Adds the two Spanish strings.
* Say why a Grok turn failed, and keep task rows in Grok's own words
A failed Grok turn ended with no reason on screen: the translator dropped
every copy of Grok's message. The failed turn now gets one status row in
Orca's existing "provider did not accept this message" words with Grok's
reason, read from whichever copy arrives first (the given-up retry, the
turn's end, the prompt's completion notice, or the prompt's error answer);
later copies only fill a reason the row still lacks.
A running background command no longer reads "Background task <id>
started": a task's summary is mapped only once it has settled. A monitor
stays a monitor when the agent reads its output: a frame that names no
kind keeps the known one, and a "[monitor" command is a monitor.
A prompt's turn is marked started, so a late frame for an ended prompt
neither reopens it nor becomes the active turn. A tool's turn is held in
one place at a time.
* Read a monitor from Grok's exact output prefix
* Word a failed Grok turn in Grok's own text, not as a refused message
A turn that started and then failed was told "The provider did not accept
this message", Orca's sentence for a message refused before its turn. The
row now reads as a Codex turn-ending error does: an error status row with the
provider's own words. With no words, the dialect names the failure ("Grok
ended this turn with an error." / "Grok usage limit reached."), else the
agent's display name does.
* Settle a stopped turn's running call as its turn row ended after a restart too
The restart sweep ended every running call by the death evidence alone, so after
a person's Stop with no proof the child died the call read failed under a turn
that read interrupted. The sweep and the live dead-generation settlement now ask
the same rule the assembler does: a call in a turn already settled ends as that
row ended; only a turn still running leaves its calls to the evidence.
* Keep the dead-generation settlement under the line cap
* Register the ACP schema verify step in the PR preflight phase test
* refactor(agent-session): one required agent registry; declarations admit what they claim
A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.
/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.
Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).
* Read ACP permissions, session events and prompt errors through the protocol client's own types
The translator now reads a permission request with the client's lenient reader, a session update
with its session-event reader, and takes only the agent's own error answer as a failed prompt's
reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values
newer than this build.
* refactor(agent-session): the router applies the declared rewind itself
The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.
* test(agent-session): register the agents the merged-in tests now need
The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.
* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record
* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop
The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.
The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.
One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.
* fix(agent-session): a changed agent definition never hides that agent's chats
A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.
Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.
* refactor(agent-session): each agent's registration says where it runs and which account it pins
createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.
Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.
* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state
A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.
* fix(agent-session): a scoped dismiss-all persists no per-session fence
The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.
* fix(agent-session): refuse an attach whose agent is not the session's own
The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.
* fix(agent-session): offer to start a chat only when the start would accept it
The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.
* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it
A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.
* refactor(agent-session): the record store admits agent ids; comments say where transport is checked
The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.
* docs(agent-session): the record store admits the registered agents' ids
* refactor(native-chat): Grok's registration declares where it runs; ACP no longer borrows Codex's location rule
The rule a self-supervised agent child runs under (this machine, no WSL, Windows only with process
start-time proof) is its own module that Codex and the ACP adapter both use. Grok's registration
takes the full account-home resolver signature, and D3's tests build hosts with the agent registry.
* fix(native-chat): Grok follows the ACP runtime's request contract and the managed process's close
A request the agent or a Stop cancels is answered with the agent's own cancelled reply by the code that
owns it (the runtime no longer answers a silent handler), so a Stop needs no separate decline pass. A
permission answer still being saved when the agent stopped waiting is reported unconfirmed, since the
protocol already answered it cancelled. Cancelling the agent's own turn is the plain cancel. Request
rows are matched under their generation-scoped ids. A refusal's reason comes from the dialect's wording
path. The child drops its own stderr tail and close policy for the managed process's, and a close
whose process tree was not proven gone is reported as the adapter contract asks.
* fix(native-chat): a Grok chat Orca already holds resumes without writing what Grok replays
A chat with a saved Grok session reattaches with session/resume where the agent offers it, else
session/load. Either way the call runs inside the translator's load window, so what Grok sends while
it reattaches (its saved exchange, a task the dead process left running, ended by the restart) opens
no turn and writes no row; only context usage reads on. A reply an Orca or Grok crash cut short is no
longer completed from Grok's saved history: it reads like a Claude or Codex chat's, with the existing
notice. The attach window also closes after a failed attach, and a created session that session/resume
reports missing is replaced like one session/load reports missing.
The replay reconciliation is removed: the lane no longer reads the journal, and D3's replayed-input
grammar test and completed-turn check in the assembler go with it.
* refactor(native-chat): a failed Grok reattach needs no window close of its own; its lane is replaced
* test(native-chat): D3's merged tests use the shipped declarations and the launch options main requires
* fix(native-chat): typecheck fallout of the base merges; any agent's empty chat is reusable
Main's idle-empty-chat lookup and launch join now take any registered agent, as the rest of the
launch path does. The refusal check moved into the prompt turns and the prompt-block conversion beside
the turns that send it, keeping both files in their line limit.
* fix(native-chat): a Grok Stop ends the process once Grok settles its turn; the next send resumes
Grok's session/cancel ends only the running turn: work it already moved to the background keeps
running and can begin a turn of its own after the person pressed Stop. Stop is now a session
boundary, as it is for Claude: the cancel answers open requests and lets Grok end the turn, the host
waits a bounded grace for that, then ends the process; the next send relaunches and resumes.
The adapter's own bounded close of a turn Grok began is gone. Its named-turn check stays: the host
ends the session unless the provider declines a Stop naming a turn that has since ended.
* test(native-chat): a Grok Stop ends the process only after Grok answered the cancel
* fix(native-chat): Steer on a Grok card cancels the running prompt, then sends it
A send that reached Grok while a prompt ran was held in the adapter until that turn ended: Steer
on a queued card took the card out of the host's editable queue and meant 'send after this turn'.
It now cancels the running prompt (session/cancel; the session stays) and sends as the next prompt
once Grok answers the cancel, as the common pattern does; a steer behind another cancels it in
turn, so the last one runs. The adapter holds a send only while that cancel lands, so its general
held-send queue and its holdsDispatch report are gone (every send it holds has its turn open in
the journal). An older client's mid-turn send takes the same path. capabilities.steering is
unchanged and still unread.
* refactor(native-chat): a close or Stop cancels a start through the acquire's own abort signal
The host owns the acquire it runs, so it now owns its cancellation: each attach's acquire gets an
AbortSignal, aborted from outside the session's queue by a close and by a Stop admitted now (the
same admission rule as before). The optional abandonStart adapter hook, the router's fan-out to
every adapter and the ACP adapter's session-keyed start map are gone; the ACP adapter keeps an
unkeyed set of starts only so quit can prove their children gone, and keeps a failed start's
unproven child until its exit is proven.
The hook also let a later close ask that child again. The host now does that from state it holds:
a close of a chat with no live child whose record still names an owner process with no death
evidence asks the adapter to release it. The answer is not recorded as proof (the lease probe
does that), so an owner pid an earlier Orca left is never killed or marked gone. Claude and Codex
ignore the signal and hold no such child; their release is a no-op (tested).
* fix(native-chat): a Grok crash that closes stdout before its exit still ends with Grok's last words
On macOS and Linux the agent's stdout ends before its exit is observed, with or without the
supervisor's EOF forwarding, so the connection's loss closed the journal first and its error text
became the session's ended reason, dropping Grok's stderr. The reason is now read at the proven
exit: the agent's last words when it left any, else why the connection closed. The failure already
carried them. Comments that assumed the exit comes first, that early frames past the cap refuse the
start, and that dispatch re-checks image support are corrected.
* fix(native-chat): nothing Grok sends while a held chat reattaches is written, marked as replay or not
The reattach window relied on the dialect's replay verdict, and Grok's frames read as live unless
they carry isReplay, so an unmarked chat frame during session/resume opened a turn that never
ended. D3 now marks every frame inside the window as replay before the translator reads it, so the
translator keeps only context usage whatever the agent marked; options and commands are still
adopted. The translator's load semantics are unchanged.
* test(native-chat): a Stop after a resume finds no turn an unmarked old reply opened
* test(native-chat): a resumed Grok chat keeps its last context reading; the resume refreshes only the window
* test(native-chat): a Grok background task a Stop ended reads as stopped reporting
* refactor(native-chat): quit's stop of each start answers through one promise kind
* fix(native-chat): quit aborts every start the host has in flight before draining attaches
A Grok that never answered its handshake held quit until the start's own 60 s bound, past the
20 s quit deadline. The host's teardown now aborts each in-flight acquire (and any the drain
still begins), so the adapter's own quit controller and its map of starts are gone: a start
has one canceller, the host's signal.
* fix(native-chat): a Grok start's abort stops reaching its child once the start has returned
The listener stayed on the host's signal until the attach finished committing, so a Close in that
window killed the now-live child behind the host's back and it read as Grok crashing. The start
now detaches it when it ends; a later Close goes through the session's own stop.
* fix(native-chat): a close or Stop during any attach phase stops the start before it launches
The attach began its abort controller only after reconciling leases, resolving recovery and
probing the previous owner, so a close or admitted Stop in those phases reached nothing and Grok
launched anyway. The controller now begins first, and the acquisition checks it before asking the
adapter to start.
* test(native-chat): a close during the attach's owner probe asks no adapter to start
Also renames the close test after the hook it no longer exercises.
* test(native-chat): a close's re-ask closes a Claude or Codex child a failed cleanup left
The re-ask is not a no-op for them: when the adapter still holds the child its cleanup could not
prove gone, the close stops it again as a requested close, and Claude persists the handle of the
conversation it ran so the next send resumes it. Corrects the tests' and comment's wording; the
close awaits the re-ask, bounded by each adapter's kill ladder.
* fix(native-chat): Steer during a turn Grok began itself cancels it and sends once it ends
A send while Grok ran a turn of its own (a background task waking it) went straight to Grok, which
queued it behind that turn where Orca could no longer withdraw it, while Stop treated the same turn
as the running reply. The send now waits as a steer, the turn is cancelled once, and the message
goes when the turn ends; a Stop withdraws it and an exit rejects it as never sent.
* test(native-chat): a steer whose cancel Grok never answers ends Grok and is rejected as never sent
Pins the bounded steer cancel kept from the runtime: past the bound the connection closes, the
running reply reads unverifiable, Grok's end reads as its exit, and the waiting steer is rejected
as never sent.
* fix(native-chat): a Grok crash stays a crash when a stop lands before its exit is proven
After the connection broke and the close could not prove Grok's exit, any later stop Orca asked
for (the next start, a Stop, a Close) marked the child as closed by Orca, so the crash read as a
requested close and Grok's last words were dropped; a send meanwhile was recorded unconfirmed.
The connection loss now decides the cause, and a send on that session is rejected as never sent.
* test(native-chat): fixtures this PR's registered Grok and desktop capability made stale
CI's unit shards failed on tests outside the PR's own lists. Each encodes something this PR changes
on purpose: Grok is now a registered agent (the seam test's unregistered agent is now Cursor); the
desktop now advertises registered agents (the restart-offer tests' older client drops that
capability explicitly); the attach context carries the start's abort controllers (the forget-status
double gains them); and the ACP real-host test rig sends to the host directly (listed beside the
other real-host rig in the send ratchet).
* fix(native-chat): a start quit stops is not the queued message's start failure
With quit now aborting a start it would have waited for, the delivery step recorded the aborted
start as the message's failure ("couldn't restart"). After quit has stopped delivery, the step
leaves the message to quit, which settles it as a close does ("The chat closed before this message
was sent."). The test that pinned quit waiting for that start and stopping its child now pins that
nothing is launched behind quit.
* fix(native-chat): a message sent after a Stop or close aborted a start gets its own start
A start the host aborts (an admitted Stop, a close, or quit) returned its refusal to the delivery
loop, which then rejected whatever was queued at that moment with "couldn't restart", including a
message the user sent after the Stop. The attach now reports that the host aborted it, and the loop
re-derives from the journal instead: what the Stop or close withdrew is already settled, a message
accepted since gets a start of its own, and quit's next step stops the loop. This replaces the
quit-only carve-out with the same rule for every abort and every agent.
* test(native-chat): the message sent after an aborted start is answered, so no settlement outlives the test
* fix(native-chat): a Grok model pick Grok never answers no longer holds Stop or Close
The pick runs on the session's queue. It now registers in the host's out-of-queue
abort registry beside a start, so a close, an admitted Stop or quit abandons it, and
the ACP adapter bounds it at 30 s like Claude and Codex. A late answer is still adopted.
* fix(agent-launch): a phone's launch opens a terminal for an agent whose chat it cannot show
agent.launch now reads the caller's capabilities by the rule tabs and restart offers
use (clientRendersStructuredAgent). A phone without registered-agents.v1 gets Grok as
a terminal again, as on main; the host's own callers and desktop clients are unchanged.
* fix(acp): strip every agent hook variable from the ACP child, from the shared list
ACP_CHILD_ENV_TO_DELETE was a second copy of the hook runtime keys that missed
ORCA_AGENT_HOOK_TRANSPORT; it now spreads AGENT_HOOK_RUNTIME_ENV_KEYS beside the pane
identity keys.
* refactor(native-chat): the mutation context carries the provider-wait registry itself
Keeps the host file within its line limit; one field instead of two closures over it.
* fix(agent-launch): agent.launch.v2 still vouches for Claude and Codex chats
The caller rule from the previous commit also turned Claude and Codex into terminals
for a client advertising only agent.launch.v2, whose contract says it opens a chat
(mobile retry-authority tests). Only an agent beyond those two now needs the client to
read it (clientRendersStructuredAgent); the test fixtures go back to what they were.
* refactor(native-chat): drop saved-history adoption from the timeline assembler
The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.
* refactor(acp): drop session/load history adoption from the translator
The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.
* refactor(native-chat): a pending input is only Orca's send now
Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.
* test(acp): keep the task-result status table on live frames
Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.
* test(acp): a frame helper for a shell command Grok is running
* fix(acp): a Grok crash settles through the host's provider-exit batch, scoped to the turn it ended
A Grok crash ended the journal unverifiable before the adapter reported the exit, so the host's
provider-exit settlement found no running turn and wrote nothing: the adapter's failure (with
Grok's last words) never reached the journal, and a later stale-session pass wrote a bare,
thread-scoped cut-short row, so the partial reply was not folded as Claude's and Codex's are.
At a proven exit the ACP lane now ends its running turn interrupted at the exit instant, as the
host's exit contract expects of a child's own translator (Codex's does the same). When Grok's
stdout closed first (every POSIX crash), the turn is unverifiable only until the exit is proven:
the host's provider-exit settlement now takes the exit as proof naming the child's fence and
revises what that child left unverifiable in the same batch, with the turn-scoped row and the
adapter's failure. Claude and Codex write no unverifiable turn of a live child except a command
whose hand-off is in doubt; that turn is now revised at the exit instead of at the next open.
* test(acp): a crash seen first leaves the host no Grok turn to revise
* refactor(native-chat): what a gone generation left unfinished gets its own module
The settlement file passed 300 lines with the exit-proof revision. The unfinished-work reads
(capture, interrupted-by-the-exit, in-progress) are their own concept and move out unchanged,
apart from the exit proof they now take.
* refactor(native-chat): a watched exit revises what its child left unverifiable without reading Stop marks
An exit's own instant is the turn's end, so the revision needs only each row's fence: the
settlement's journal type gains itemFence alone, and the host test fakes say so.
* test(native-chat): drop the duplicate itemFence on the fake that already had one
* test(claude, codex): an exit whose stdout ended first still reports as it always did
The provider supervisor now ends Orca's stdout when the agent's ends, so on every crash EOF
arrives before the exit is seen. Claude's and Codex's connections report nothing at EOF and
report the exit, with its usual reason, once it is seen.
* fix(acp): reopen a chat with session/load, as the common pattern does
An agent that offers both now reloads its session instead of resuming it; the
reattach window still discards what it replays except context usage.
* fix(acp): drop the 60 s handshake bound; an abort fails the start's waits at once
Neither common design bounds an ACP handshake: Close, Stop and quit end a start
that never answers. The abort now also closes the connection, as a kill there
does, so the start settles even before the child's exit is proven. The
host-stopped start refusal only this bound produced goes with it; the idle
sweep keeps its words.
* fix(acp): a Stop naming an ended turn follows Claude's rule
It still stops nothing while another turn is live, but in the gap before a
follow-up's turn opens, which no client can name, it now stops what is in
flight and the session ends, as a Claude Stop does.
* fix(native-chat): a close no longer re-asks a failed start's unproven child
Neither common design retries that stop at Close, and Orca's Claude contract
re-asks only at the next start and at quit. The ACP adapter keeps the child
until its exit is proven and asks it again there, as Claude does.
* fix(acp): a message sent during a turn the agent began itself goes at once
Both common designs send it straight to the agent with no cancel; only Orca's
own running prompt is steered (cancelled, then re-prompted).
* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer
Uses an audience production sends (one that cannot show every agent), per review.
* fix(acp): launch Grok as `grok agent stdio`, without the update and leader flags
The common pattern passes neither --no-auto-update, --no-leader nor
GROK_DISABLE_AUTOUPDATER; full access still adds --always-approve.
* fix(acp): an agent that ends its stdout, or answers unreadably, is not a lost connection
As in the common pattern, only a broken stdin (or Orca's own close) ends the
agent; one that closed its output but can still be written to stays until a
Stop, a close or its exit. The provider supervisor goes back to its base
content, so Claude and Codex no longer get the forwarded stdout end either.
* fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts
On main every close of the chat writes its own Stop and settle. Here a later close joins the
first and writes no row, and the first's settle closed when its kill failed, so a turn that opened
in between and was cut by the next close read as failed. A person's close joining a person's close
whose Stop opened a settle now reopens that settle until its attempt is done.
Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and
the next reads as the person's cancellation (each fails without its half of the fix).
* test(native-chat): Grok opens as a chat only behind the structured-chat setting
agent.launch and orchestration worker-start read the same setting as the
renderer route; pin both states for Grok on each. The setting's description no
longer names only Codex and Claude, in every catalog.
* docs(acp): generic ACP comments say what holds for every agent, not Grok
Stop ends the session for every ACP agent, as in the common pattern; the
adoption hook comment goes (adoption is not planned); a failed start's child is
retried at the next start or quit.
* test(claude, codex): type the EOF-before-exit test's streams; the supervisor no longer forwards EOF
The Claude test wrote to the child's stdout and stderr through their Readable
type, which the node typecheck rejects; it now holds its own PassThrough
streams. The comments no longer credit the reverted supervisor change.
* feat(acp): a steer's cancel asks once and never ends the agent
The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.
* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge
* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns
Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.
* fix(native-chat): drop the stopDelivery the A3 merge doubled
* fix(acp): a steer's cancel asks Grok once and never ends it
A steer now uses D1's notify-only cancel. Two messages sent during a reply Grok began itself
cut that reply, as the common pattern does, and then both run; before, the queued first
message could not answer the bounded cancel and Orca ended Grok although Grok answered.
A Stop keeps the bounded cancel and its 4 s grace.
* fix(acp): a permission Grok asks with no prompt of Orca's running is declined
During a turn Grok began itself nobody asked it to act, so the request is answered
cancelled at once instead of opening a card that waits, as the common pattern does.
* fix(acp): a Grok that dies while starting is reported with its own last words
A dying process's stdout ends before its exit is seen, so the start failed as a closed
connection and Grok's stderr was lost. A start whose connection closed now waits, bounded by
the Stop grace (or a Close/Stop), for the exit before it is told.
* test: a Stop after a steer sends its own cancel; drop the import the A3 merge doubled
* test(native-chat): main's Stop-note test builds its turn context with the agent registry
* test(claude): say why the close test's fake child cast is safe
* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way
The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.
* test(native-chat): build the Stop-opened-turn test's identity with the opaque handle
The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle.
* test(ratchet): require src/main/provider-process now that it has landed
* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test
Main's #25706 and this branch both added the import at different lines; the merge kept both.
* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test
Main's #25706 made the same fix as this branch at a different line; the merge kept both imports.
* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore
A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.
Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.
* test(native-chat): build this stack's journal identities with main's opaque provider handle
Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test
files from this stack still wrote the old shape. Same lines the downstream ACP branch uses.
* docs(acp): every reattach drops the agent's replay, not only for a chat the journal holds
* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row
Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.
* refactor(native-chat): read hosts' structured agents from the app-shell services
Main grew the startup hydration hook to its line limit; the host agents sync is an app-lifetime subscription like the structured session tabs sync beside it, so it moves there.
* Use current provider handles in transition tests
* Use current provider handles in timeline fixtures
* test(native-chat): prove replacement rows survive downgrade and re-upgrade
* Require the ACP directory in the runtime import check
* test(ratchet): require src/main/acp now that this PR lands it
* feat(acp): a saved session the agent cannot reopen continues in a new one, with one warning row
When session/load (or session/resume) of a saved ACP session fails, the chat starts a new session and records it as a creation that replaces the lost one (#25747's 'replaces' link), and writes one warning row that the agent no longer remembers the earlier messages. A created session the agent reports missing is still superseded silently; a signed-out agent or a start that is over (Close, Stop, a lost agent) still fails the start.
* chore(acp): rewrap the acquire header comment
* test(acp): a start closed while the agent reopens fails without opening or announcing a new session
* Let ACP connections own their supervised agent process
* Preserve ACP cleanup evidence and isolate exit observers
* Expose ACP cleanup observations and type the permission fixture
* refactor(native-chat): the registered-agents capability lives in its own module
Main's growth put protocol-version.ts one counted line over its 300-line limit once the capability
was added; like main's other per-feature capabilities, it now has its own module, and importers read
it from there.
* refactor(acp): one connection owns the Grok process and its protocol
D3 now opens each ACP agent through createAcpAgentConnection (ACP-ALIGN #25810): one object spawns the
process on the execution host, owns its stdio and protocol, and reports its proven exit. It is built and
tracked before the handshake, so a start's abort (Close, Stop, quit) still reaches it, and a failed start
keeps that same connection for the next close to retry rather than spawning another process.
Deleted: the spawnAcpStructuredChild wrapper and its test, the raw-stream runtime assembly, the caller's
exit -> runtime.close wiring, the stdout-EOF heuristic (the connection no longer treats stdout EOF as
exit), and the 10 s steer/Stop cancel bound with requestSteerCancel. Reader control maps to
pauseReading/resumeReading; a close is connection.close after the host's existing 4 s Stop grace.
The adapter owns what the protocol no longer does: one session/cancel per running prompt however many
steers arrive (cleared with that send's settlement, retried after a failed write), and a Stop or steer
answers every open agent request the person has not already answered with the agent's own cancelled
reply. An answer already being saved when the Stop lands is sent.
Tests: blocked cancel write never holds Stop's grace, two quick steers send one cancel, a failed cancel
write is retried, a real process exiting while a child holds its stdout ends the session, and the
existing start-abort, retention, crash, connection-loss and reload-failure suites on the new rig.
* fix(acp): Grok signs in on its own machine with its API key or cached sign-in
When Grok reports that it needs authentication, Orca now names a sign-in method on the machine Grok runs
on, read from the same environment Grok was launched with: xai.api_key when XAI_API_KEY is set there and
Grok offers that method, else cached_token when Grok offers it, else none and the chat keeps the existing
not-signed-in refusal. The rule lives in Grok's launch spec; the adapter applies any agent's rule for new
and reopened sessions through the protocol client's caller-named method (authenticate, then retry once).
No new sign-in UI; interactive methods are never chosen.
* fix(acp): the adapter decides which of Grok's requests reach the person
The turn owner now admits every agent request, permission or question, from its own turn state: a
request reaches the person only while Orca's prompt runs and no steer or Stop is cutting it short (a
question may also come from a turn Grok began itself, until a Stop). Anything else gets the agent's
own cancelled reply and opens no card, so a question arriving after Stop or during a steer never
appears. A steer, like a Stop, withdraws the requests already open; an answer already being saved is
still sent. The protocol client's abort-on-cancel path is no longer used: after the connection
change its request signal aborts only when the connection closes.
* fix(acp): a plan Grok proposes shows as a plan, with no approval card
When Grok leaves plan mode it asks the client to approve its plan (x.ai/exit_plan_mode). Orca showed a
blocking 'Approve plan / Request changes' card for it; the common pattern has no such gate. Now the
plan goes into the chat's existing Plan row (the plan-document status row Codex and ACP plan updates
already use) and the request is answered at once with 'abandoned' plus feedback telling Grok to stop and
wait for the person's feedback or a request to implement it in a later turn, so nothing is approved on
the person's behalf. Dialects gain settleRequest for requests answered without asking anyone.
* fix(orchestration): worker-start opens a Grok worker in a terminal, as before
With the structured chat setting on, worker-start decided 'structured' for Grok and then the structured
worker factory (Claude and Codex only) refused it, so the start failed; main opened a terminal Grok
worker. Worker-start now decides with no registered agents beyond Claude and Codex, so Grok gets a
terminal worker as before. agent.launch and the app's own launches still open Grok as a structured
chat. Temporary until structured workers take registered agents.
* fix(acp): a prompt answer Orca can't read ends the turn instead of hanging it
A session/prompt rejection that was not the agent's own error answer (an answer that fails Orca's
schema, or one too large to read) left the turn running: the next message became a steer with nothing
to cancel and was never sent or settled, and Stop waited its full grace. As in the common pattern, any
prompt failure now ends the turn as failed (a failed-turn row without words, since none are the
agent's) and settles the send, so the next message goes. Only a closed connection keeps the send
running, for the connection-loss path to settle.
* fix(acp): send Grok's prompt-identity extension only to agents that echo it
session/prompt carried _meta {promptId, requestId} for every ACP agent, though only Grok's dialect
echoes it (injectedPromptIdentity). Now only an agent whose dialect declares it gets the extension;
other ACP agents get a plain prompt.
* refactor(native-chat): the registered-agents capability lives in protocol-version again, as on main
This reverts
|
||
|
|
76d480f808 |
Fix unsafe test fixtures and the Bun version pin (#26051)
* Keep test interruption signals within owned processes * Pin Bun and add optional unit runner shutdown diagnostics * Unblock CI lint without changing session host runtime * Avoid duplicating runtime import-check dependency bundles * Leave unit runner diagnostics disabled by default * test: keep runner incident follow-up focused on durable guards * test: apply transcript replacements as authoritative snapshots |
||
|
|
d3e1494674 | test: bound memory used by the runtime Electron audit (#26049) | ||
|
|
0bdcaf36ed |
fix(claude): start a Claude chat with its saved options and send the first message at once (#25152)
* fix(claude): end a Claude start that never answers initialize after 120 s * Read the Claude startup deadline inside startup; fix a stale test comment * fix(claude): start a Claude chat on its initialize answer, not on a frame only a SessionStart hook sends Startup waited for system/init or a SessionStart hook frame as well as the initialize answer. Before the first turn only a SessionStart hook sends one, and Orca adds that hook only through its optional status hooks, so with them off the first message was held forever. Startup now lands on the initialize answer; a start frame already seen is still checked, and one naming another session ends a started session. The deadline drops to 90 s so it fires inside the host's 120 s start wait. * fix(claude): time the Claude start by silence, and fail it at once on another session's frame Claude answers initialize only after its SessionStart hooks finish, so a total-time deadline would fail every start behind a slow hook. Each start frame now restarts the clock. A frame naming another session fails a start still waiting on initialize at once, as before. * test(claude): a real Claude chat starts and answers with every hook disabled * Say what the start-frame re-arm covers, and check only start frames in the hook-less real test * fix(native-chat): a Claude chat starts with its saved options and takes its first message at once Saved model, effort, Fast and permission mode are passed as launch options, checked against the account's cached model catalog, instead of restored by control requests after initialize. With nothing left to restore, the host no longer holds a message until the CLI answers initialize, and the 90 s startup deadline is gone. A Stop on a start that never answers ends the child and settles what it was handed as stopped. A failed result that repeats the turn's own API error reply writes no second row. * fix(native-chat): a host stop of a Claude start fails the message it was handed, with one row With no start-hold the delivery loop no longer sees a host stop of a start it waited on. The child's end now rejects what it handed over with the host-stopped words and writes the one row, as an exit of its own would; an idle start the host stops still goes quietly. * test(native-chat): a Claude chat's first message is written before initialize answers Rewrites the tests that encoded the start-hold, the startup deadline and the option restore to the new contract, and adds: saved options at launch (catalog checks, bypass, fresh-session Fast), a message written before initialize answers (adapter and runtime), Stop on a start that never answers (stopped, child closed, nothing working), and an API error said once. * revert(native-chat): keep a failed Claude turn's error row The shared turn fold already shows a failed turn's error once after it settles, as an error; dropping the row left the CLI's synthetic reply looking like something Claude said. * fix(native-chat): pass saved Claude options unchecked, heal a retired model on the CLI's word, and never leave an unrun message in doubt - Saved model, effort and Fast are launched as picked; only values no Claude can parse are left out. The pre-spawn cache check is gone. - A saved Fast on for a new conversation is applied once the settings readback shows no per-session opt-in (dropped when there is one, or when the model is listed without Fast), with nothing waiting on it; the record keeps the pick. - Under an Agent Permissions bypass, a saved narrower mode launches with the allow flag so bypass stays reachable. - A turn whose reply is the CLI's model_not_found for the launched model drops that model from the record; the launch's own row for it is kept out of the account model cache. - A child that ends before it answered initialize, for any reason, settles every message it was handed as not sent (cancelled for a Stop). - A launched effort the CLI reports only as `applied.effort` is confirmed from there. - The untimed-initialize comment is back to main's text. - A real-CLI test for a message written before initialize answers, under saved options. * fix(native-chat): type the close's ended event and the start-exit test fixtures The close's ended event is typed as the adapter event so its optional startupUnanswered spread fits exactOptionalPropertyTypes; two tests guard the fixture's optional generation, and the hung-start fixture records initialize on the fake connection it holds. * fix(native-chat): a Claude model heal keeps a later pick, a refused Fast is dropped, no allow flag - `options-skipped` carries the retired value; the record drops it only while it still holds it. - `started` carries the values a heal retired, and the record does not take them back from the CLI's report of the same value. - A saved Fast on a new conversation is applied before `started`: a refusal drops it and records it skipped, as main's refused restore did; silence keeps it wanted and unconfirmed. - A saved narrower mode under an Agent Permissions bypass launches without any bypass flag again: the allow flag is one older CLIs reject at start. Kept as a known limit. - The real-CLI test asserts the message was written before initialize answered. - The fake reports a launch effort only under `applied`, and a misplaced doc comment moves back. * fix(native-chat): a new Claude chat reports started before its saved Fast is applied The Fast apply on a new conversation now runs after `started`, so a Stop interrupts a running first turn and an option write is not refused while the round trip is out. A refusal drops the pick through `options-skipped`, in order after `started`; silence keeps it unconfirmed. The launch's unreachable skipped-model branch is gone. * fix(native-chat): a healed Claude chat goes back to the default model live; comments match the no-hold design When the CLI says the launched model does not exist, the live child is also put back on the CLI's own default (set_model with no model, fire-and-forget), so later messages in the same chat run; a user's pick sent after it wins, and a refused or unanswered reset only logs. Comments that still described the start-hold or the option restore now describe the launch options and the handed-over, never-echoed rule. * fix(native-chat): a message handed to a Claude start that never answered is kept as main keeps an unsent one A child that ended before it answered initialize ran nothing it was handed, the same fact as a send accepted and never handed over. Its end now settles those sends exactly as the chat settles a queued send for that end: a quit keeps a person's message as a held card (restart words), a close keeps it as a held card (closed words), a person's Stop withdraws it as cancelled, and a host stop fails the start with one row. * fix(native-chat): a quit during a Claude start that never answered offers no resume for the message it keeps as a card The restart snapshot now reads the same never-answered fact the exit does, so a message handed to such a start counts as queued work, not as work to resume. The retired-model reset comment names the default it really applies. * refactor(native-chat): the saved permission-mode launch helpers live with the spawn options that use them Keeps claude-structured-launch-resolution.ts within max-lines once merged with main, and names the hung-start test envelope's field type. * fix(native-chat): a Claude chat's saved options take precedence over the agent Arguments' own flags Main now passes the saved agent Arguments to the Claude child, and the SDK writes them after its own options. An Arguments --model or --effort therefore reached the CLI as a second flag after the chat's saved pick (a commander CLI keeps the last), and a saved Fast's launch settings replaced an Arguments --settings file outright. The saved model and effort now stand in for the Arguments' flags, and a saved Fast beside an Arguments --settings is applied by the start instead of at launch. * test(native-chat): the hand-built Claude session in the options test carries fastModeAtStart * test(native-chat): the queued rig's start spy carries a named Mock type An unannotated vi.fn() inferred @vitest/spy's internal Procedure, which CI's typecheck cannot name in the factories' inferred return types (TS2883). * fix(ci): run the PR's SQLite-backed tests in the Node runtime project Lists the hung-start Stop test, renames the send-during-startup entry from its old name, and carries main's own two entries from #26010 so the boundary test passes before the next merge. |
||
|
|
b3b6c5dc13 |
Give native chat names one source for tabs, sidebar and AI Vault (list and search) (#25986)
* Give native chat names one renderer source and drop Vault's name repair copies The host's saved conversation name now rides the structured session status feed, which already exists per host, is keyed by the durable session id, and keeps a closed chat's summary. Tab strip, sidebar rows and AI Vault (list and search) read it through one hook and one display order (tab alias, saved name, host label). Vault no longer copies names into its cached results, so the projection, recovery and pending-title modules and their tab-snapshot lanes are removed. Indexed search hits now carry the native owner and saved name from the host that indexed them. * Type the sidebar name test fixture without an assertion * Keep Vault search working when the chat host will not install Naming and owning search hits is bookkeeping: if the native chat host fails to install, return the plain hits instead of failing the search. The runtime RPC only installs the host for clients that will receive the owners. * Publish chat names to the feed independently of the tab retitle A failed feed publication no longer skips retitling the open tab. The publish now lives in the naming deps, where a test covers it. * Note why the status feed must keep closed chats' summaries * Bound names and owner ids that come from a paired host Drop a published chat name the record store would refuse, and cap a search hit's owner workspace id at the same length the list row and record use. * Let native chat search hits from a paired host open their chat A paired host's search hits carry no resume command, so the row disabled every open action even for a native chat it can open through its owner, as its list row does. * Ignore workspace ids that name object members in tab lookups A paired host's row or search hit could carry a workspace id such as "constructor", which read an Object.prototype member as a tab list and broke the render. The shared tab index now reads only own workspace entries. |
||
|
|
825d7bd5a9 |
test(vitest): run agent-launch-instant-tab in the SQLite runtime project (#26028)
#25430 added a test that opens a real agent-session record store, but not to the SQLite runtime list, so vitest-sqlite-runtime-boundary fails on main. |
||
|
|
3fb72d135d | Run Node event-loop measurement after ordinary test suites (#26015) | ||
|
|
37ff3873a0 | Run combined localization catalog verification on Bun (#25999) | ||
|
|
f0520851ab | Keep mobile restore SQLite fixture in the Node test runtime | ||
|
|
1040f673b1 | Keep orchestration SQLite fixtures in the Node test runtime | ||
|
|
3308ff8b26 |
Bound YAML merge conversion and close SQLite routing review gaps (#25998)
* Close test runtime and YAML merge review gaps * Bound YAML conversion inside explicitly tagged pairs * Make completion notification fixture cadence deterministic |
||
|
|
faed899cd3 |
fix(ci): run three new SQLite-backed tests in the Node runtime project (#26010)
#25888 and #25766 added tests that import the orchestration database or the structured session runtime without registering them in the Node runtime list, so vitest-sqlite-runtime-boundary fails on main and every PR. |
||
|
|
d8c871a1f0 |
Speed up unit tests with cross-runtime duration scheduling (#25967)
* Schedule unit tests across runtimes by measured duration * Keep new SQLite fixture suites on Node after updating main * Inject scheduling timings instead of mocking the module * Keep agent-session database lifecycle contracts on Node |
||
|
|
72b84118e0 |
test(e2e): terminal layout parity check against main (#25681)
* test(e2e): add terminal layout parity check for topology refactor PRs Runs fixed terminal-layout journeys in the real app on two builds, captures the renderer topology and the saved workspace session, normalizes volatile values and fails on any difference not declared for a named bug. Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): record quit exit status and report paths main does not reproduce Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): invoke pnpm correctly under corepack and silence the typeless-module warning Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): accept pnpm's forwarded -- in the layout parity runner Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): close parity panes through the user's chord and treat absent maps as empty Driving PaneManager.closePane directly left main to learn of the close from the PTY exit, which raced quit; the keyboard path commits the close in main first. Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): build parity checkouts before running, forward -g, and add a drag-out scenario Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
6aa12c30ef |
Name native Claude and Codex chats after their first message (#25724)
* feat: generate chat names through configured text agents * Project structured chat names across stored tabs and session rows * Name native chats from their first live message * Restore the journal test provider handle import * Fix first-message naming and live Vault title updates * Preserve Unicode characters in bounded chat naming prompts * Read chat naming settings only when a turn needs them * fix(chat): store only generated conversation names * fix(chat): keep ordinary labels across unnamed chat surfaces * fix(chat): preserve naming after fast first turns * Preserve native chat names across command-first sends and Vault lifecycle * Route native chat SQLite contracts through the existing Node test pool * feat(settings): add chat naming controls to Chat page * Preserve chat naming drafts and configure Custom commands through host settings * Scope synthetic command output to its journal thread in naming test * Keep naming test fixtures within their typed project boundaries * Resolve the real Vault hook directly in its integration test * Publish chat naming save refs after render commits * fix: keep Japanese chat naming copy stable during repair * Update test contracts for the naming integration * Keep journal fixture reads on the host clock * Give Chat names a separate settings section |
||
|
|
f7c542c7a3 |
Keep profile saving alive after a stalled main loop (#25318)
* Keep profile saving alive after a stalled main loop After a long main-loop stall (overnight sleep, dark wakes), the profile writer's overdue 30s timeout could run before an acknowledgement that was already queued, permanently retiring the writer until restart. Terminal creation then failed because pane bindings could not be saved. - Writer deadlines measure lateness on the monotonic clock and grant a bounded fresh window when the callback is overdue or the system reports suspended; resume re-arms without spending grace. Applies to initialization, every command, and the close/exit wait. - The "Saving stopped" alert is parented to a visible main window (never a parentless synchronous macOS alert), deferred until shown, deduplicated, and says whether the latest change is unconfirmed. - Timeouts, grace, writer faults, and alert presentation leave sanitized durable breadcrumbs. * Fix profile writer timeout and shutdown races * Run profile writer stall regression on Linux and Windows * Keep Electron probes out of headless runtime qualification |
||
|
|
83cf7cf5e2 |
perf(ci): compile release JavaScript once for all packaging hosts (#25828)
* perf(ci): share release JavaScript across packaging hosts * fix(ci): verify the projected web entry in release archives * fix(ci): use the Windows system archive tool for release bundles * test(ci): retain stylesheet evidence in release build comparisons * test(ci): verify release parity across native color rounding * test(ci): normalize manifest asset references without changing import order * test(ci): compare portable outputs across Windows text and color formatting * test(ci): preserve module identity across dependent asset hashes * fix(ci): keep SVG build inputs identical across release hosts * fix(ci): stabilize compiler inputs and projected web bindings * fix(ci): retain vendor minification in projected web output * test(ci): normalize platform-specific pnpm manifest source paths * fix(packaging): exclude shared build staging from application files * test: align thinking-state fixtures with the current source shape * test(mobile): reuse message fixtures within the line limit |
||
|
|
ac8ea9f958 |
fix(codex): write Orca's hook into ~/.codex only when something changed (#25743)
* fix(codex): write Orca's hook in ~/.codex only when something changed One reconcile replaces the per-launch writer and its background approval session. It runs at app start after PATH hydration, when the setting turns on, once per native pane spawn, and (with a bounded 3 s wait) on Codex launches and resumes. Orca's entry lives alone in a matcherless group, appended last unless already in place; other copies are removed and the shifted user approvals move verbatim under both key spellings. The approval, with Codex's own hash, goes in first. The routing gate closes only for a hooks.json with a bad shape. * fix(codex): report hook status for the home the next pane uses Status reads ~/.codex (both key spellings) when launches use it, or the CLI has no selection, and the selected managed home otherwise. It explains: update Codex, Codex not found, not asked yet, approved by Orca but not yet confirmed by Codex, a hooks.json shape that moves panes to Orca's own home, and inline config.toml approvals Orca cannot add to. * test(codex): pin the ~/.codex approval against a real Codex, and run it on those files The contract now also writes Orca's entry into a throwaway ~/.codex and checks that Codex lists it trusted and enabled, turns it back on over a /hooks switch-off, lists it for review after a user inserts a hook ahead until the next check, keeps an inline config.toml loadable, and shows no review in a real TUI start. The real-binary CI job now runs when the ~/.codex writer changes. * test(ci): keep the scope test under its line cap; assert the ~/.codex paths beside the contract job * test(codex): type the hoisted test holders instead of asserting * test(codex): cover the matcher rule, an older build's event, and the spawn trigger Adds tests that failed against mutants which survived the first pass: an entry alone in a matcher group, a current copy's approval in an event left to an older build, one run per spawn, a spawn riding a running reconcile, and the native-only spawn trigger. The opt-out's Codex-hash cleanup now runs once, after the sweep, instead of twice. * docs(codex): name the ~/.codex reconcile in the legacy sweep's lane comment * fix(codex): approve ~/.codex with Orca's own hash while Codex's answer is pending The reconcile waits at most 0.5 s for Codex's hash. If the lookup is still running, an entry already in place keeps its approval and nothing is written; a missing entry or approval gets Orca's computed hash, approval first, inside a launch's 3 s wait, as main's did. When the lookup lands, the reconcile runs again and Codex's hash replaces it. * refactor(codex): one start for the hook lookup and the ~/.codex reconcile startCodexHooks replaces the two start functions and takes the PATH wait that the reconcile request's after option carried. A spawn reconciles unless one is running, without a scheduled flag, launches call one reconcileCodexHooksForLaunch on the shared withTimeout, and the reconcile alone reads the hooks setting. The startup test now runs the ready phase instead of matching its source text. * refactor(codex): plan ~/.codex with the shared Orca-hook pruner and approval reader The planner prunes with removeManagedCommands, as the opt-out does (so a hook that runs Orca's script through its args goes too), and returns a prune or settle union. While Codex has not answered, the kept approval comes from the shared per-event reader. The reconcile result is its outcome alone, a concurrent save spends a pass of the one bound, the opt-out removes Orca's approvals once, the approval-first writer takes the hooks path, and getRealHomeConfigTomlPath gives way to getSystemCodexConfigTomlPath. * fix(codex): ~/.codex keeps only an approval holding a hash Orca's entry may carry * refactor(codex): name the ~/.codex hooks-file check for what it reports * test(codex): the ~/.codex entry tests know Codex's earlier hashes, as the app's lookup does * refactor(codex): the ~/.codex reconcile uses the lookup's answer names * test(codex): ~/.codex status tests keep a Codex on PATH unless one is missing * refactor(codex): one stopgap for both homes; the ~/.codex pass finds its own home, hash and user data Also passes every approval to the approval-first writer, which already skips the ones in place. * refactor(codex): the reconcile start owns the warm-up; no rerun-on-answer flag The lookup start only records the hydrated PATH; one catch, one request type, and the app config is kept as given. * refactor(codex): the ~/.codex pass returns nothing; tests read the files Also makes the ~/.codex opt-out take Codex's hashes, as every production caller passes them. * refactor(codex): the ~/.codex approval cleanup takes Codex's hashes * test(startup): check the Codex hook start waits for PATH by behavior, not identity * fix(codex): a status-hook problem never moves the system default off ~/.codex * chore(ci): run the real-Codex contract when the ~/.codex reconcile changes * docs(codex): the ~/.codex reconcile's refused branch covers every definitive refusal * test(codex): start the failure-memo lookup with the reconcile's PATH-only start * fix(codex): ~/.codex gets nothing while Codex is missing, a pending answer uses the saved one, a conversion waits for a run that writes, and status and the reconcile read one home and one Orca-hash rule * test(codex): ~/.codex and managed-home status tests name their home; ~/.codex rewrites the backslash key Codex on Windows reads * test(codex): the Windows upgrade test reads status for the managed home it installs * test(codex): the real-TUI contract trusts its workdir by its real path, and its userData exists * fix(codex): keep Orca's approval at every slot its entry holds in ~/.codex |
||
|
|
ab6389045e |
fix(native-chat): start Windows chats without reading process creation times (#25718)
* fix(native-chat): start Windows chats without reading process creation times
Native chat on Windows refused to start ("Orca can't run this agent in a chat
here") whenever the process-table addon could not report process creation
times. Chat never needed them; only the bookkeeping around stopping the agent
did.
Windows now follows the common pattern: Stop ends the agent's tree with
`taskkill /T /F` on the child Orca still holds, and reports the tree gone only
when taskkill exits 0. A saved pid is never signalled after a restart.
- Remove the Claude and Codex location gate and the Codex launch refusal.
- A start time that cannot be read is recorded as unknown instead of refusing
the session; recovery already releases such an owner without signalling it.
- Delete Claude's Windows creation-time descendant snapshot and its verifier.
- Codex's Windows teardown reports the real taskkill outcome.
- The renderer no longer waits on the capability flag; the host keeps
publishing it for older clients (temporary).
* fix(native-chat): treat a Windows Claude exit after stdin end as a proven close
On Windows an idle Claude leaves on its own once its stdin ends, so every Stop
and close read as unproven: live background work settled as unknown and the
trace logged a close that "did not finish cleanly". As with the Codex close,
that exit is now the close and Orca makes no claim about processes Claude
started; a forced close still rests on taskkill's own report.
Also drop the Settings clause about Windows needing process start times, and
fix comments and the tracked process-enumeration doc that still described the
removed start-time gate and descendant snapshot.
* fix(native-chat): renew a held child's lease without a PID probe
An owner recorded without a process start time could never renew: the renewer
re-proves every live owner by PID identity, and with no start time and no
spawn-token echo that probe is indeterminate. The lease's last renewal then
stayed at the spawn, so a turn cut by an Orca crash was dated to its own start,
and every tick logged a failed renewal and split the batch into one store
transaction per chat.
The runtime now renews a lease for a child it still holds at the record's
fence and whose exit it has not received, recorded as a `held-child` match.
Receipt of the exit ends that proof before the exit is settled, and a restart
holds no child, so a dead owner's lease still expires and restart adjudication
is unchanged. Records this runtime does not hold keep the PID probe.
Also correct the identity probe's comment about shipped addons and note that
recovery's stop ladder is POSIX-only.
* fix(native-chat): keep a failed Windows taskkill unproven across a retried close
A Windows close counts Claude leaving on its own after its stdin ends as the
close. That shortcut also caught a retried close whose first attempt forced a
taskkill that failed: once the root exited, the retry returned true and the
failure read as a proven close. The tree reaper now records whether a reap
ever reached the live root, here, on an earlier close or from a transport
failure, and the shortcut applies only when none did; otherwise taskkill's
verdict stands and no new taskkill runs against the exited root.
Pin the platform on the existing tests that assume the POSIX close, and say
what `exit-proven` means on Windows in the acquisition-failure docs.
* fix(native-chat): derive a held child's liveness from the adapter's own handle
Lease renewal trusted a held child until the host settled its exit, and the
host hears of an exit late: Claude runs its close ladder and a store write
first, unexpected exits wait on one delivery chain shared by every chat, and a
Claude close that cannot prove its tree publishes nothing at all. A dead root
could keep renewing through that window, so a crash in it dated the cut turn
late. Renewal now asks the adapter, which owns the process handle and sees the
exit first: a child is held only while it is on record at the lease's fence
and its adapter still runs that exact acquisition with no root exit seen. An
adapter that cannot answer falls back to the PID probe. The stored
exit-received mark is gone.
The held-child read is now required by the runtime state and the renewer, and
a host-level test drives renewal through the real host wiring.
|
||
|
|
8d049b594d |
fix(codex): approve Orca's hook in managed Codex homes with Codex's own hash (#25742)
* feat(codex): ask Codex for its hash of Orca's hook in a throwaway home, cross-checked by position and path * feat(codex): cache Codex's hook hashes per binary and version, asked one at a time and only by the app * feat(codex): write a hook approval before its entry, and take back only its own on failure * feat(codex): approve Orca's hook in managed Codex homes with Codex's own hash, written first Managed homes (the shared mirror and per-account homes) no longer run a background approval session. Status reads the home's files against Codex's answer, and turning hooks off recognizes every saved version's hashes. * feat(codex): managed homes approve Orca's hook with Codex's hash; drop their background approval The previous commit carried only the managed resume's wait; this one holds the managed install it relies on. Managed homes (the shared mirror and per-account homes) write Codex's hash before the entry, fall back to their own approvals when the answer is late, and strip Orca's entry only when Codex itself answered with nothing to approve. Status reads the home's files against Codex's answer, and turning hooks off recognizes every saved version's hashes. * feat(codex): only an Orca-launched Codex waits up to 3 s for the hook hash; warm it after PATH hydration * feat(cli): name the file each agent hook status reports on * test(codex): real-Codex contract for the derived hook hash in managed homes, on both pins and latest * test(codex): type the hook-hash test fixtures and drop a duplicate import * fix(codex): give a Codex launch its own install run instead of joining a plain terminal's * test(codex): a user hook's approval stays put in an event Codex does not list * test(codex): cover late answers, first-install mirroring, stale approvals and opt-out re-asking * test(codex): a long managed home gets the daemon guard on its first install * fix(codex): until Codex answers, approve a managed home's hook with Orca's own hash, as main did A late, temporary or missing answer with no earlier approval in the home now writes main's self-computed approval instead of leaving the hook out. Codex's answer replaces it at the next install, a definitive answer (no hooks/list, a refused cross-check, 0.128) never uses it, and status says the approval is Orca's until Codex confirms it. * test(codex): status flags an unapproved entry while Codex has not answered * fix(codex): managed stopgap fills each missing event Until Codex answers, a managed home kept only the events it had already approved and dropped Orca's entry from the rest. Each event now keeps the home's approval, else gets Orca's own hash, as main wrote every event. One reader of the approval at Orca's entry serves the stopgap and status. * refactor(codex): one Codex answer type, one in-process answer map, a disk-only memo - One answer type with a kind (hashes, refused, pending) replaces two types and the three-field decoding at each caller. - The lookup keeps one in-process answer per binary path, replacing the process memo, the global latest answer and the transient-failure map; status now reads the answer for the codex on PATH, not the last one asked. - The memo file keeps Codex's refusals per version, like its hashes. - Derivation takes the version it is given; one 30 s version-probe timeout. - The launch wait reuses withTimeout, and launch prep passes launchesCodex down instead of a wait in milliseconds. - Turning hooks off no longer forgets Codex's answer. - Tests mock the derivation instead of a test-only resolver in production. * chore(codex): list the approval reader for the CLI build; fold two identical scope checks * fix(codex): count an approval at Orca's key only when it holds a hash Orca's entry may carry * refactor(codex): one append for hook trust tables * refactor(codex): the lookup keeps no entry for a missing Codex, and status checks the binary's fingerprint Also names the lookup functions for the answer they return. * refactor(codex): one stopgap reader for the managed home; the refused branch reads its own status * refactor(codex): drop defaults and exports only tests relied on * fix(codex): a failed ask of Codex stays pending instead of refusing its version * chore(ci): run the real-Codex contract when the approval reader changes * fix(codex): only a scratch home Codex loaded can refuse; the memo takes any hash and writes only on change * fix(codex): hooks turned off during a launch's wait win, Off re-keys mirrored user approvals, and one rule says which hashes are Orca's * fix(codex): an approval counts only under every key spelling Orca writes, as Codex on Windows reads only the backslash one |
||
|
|
468e4e1167 |
fix(native-chat): a prompt card owns the chat input until its answer lands (terminal-backed chat, desktop and phone) (#25761)
* fix(native-chat): an answerable prompt card owns the chat input until its answer lands * fix(mobile): a terminal chat's composer waits while its prompt card is up * test(native-chat): type the prompt card fixtures without casts * fix(native-chat): scope replies to acknowledged prompt occurrences * fix(native-chat): preserve answer ordering and verified delivery * test(native-chat): keep mock RPC client inside test boundary * test(native-chat): place mock fixtures in the test-only scope * Keep runtime comments within the module size limit * test: preserve prompt delivery coverage in desktop CI * Treat an older host's accepted write as delivered A newer desktop or phone talking to a host that predates the write settlement field read every accepted reply as "unconfirmed". Prompt cards never dismissed, the phone showed "Response unconfirmed" on every tap and ordinary chat messages were held as "Delivery unconfirmed". The reader now uses writeSettlement when present and otherwise keeps the host's whole-write accepted/refused verdict, exactly as before this branch. Only prompt answers ask for provider settlement; ordinary callers (follow-up delivery, paste drafts, option commands, composer sends) are back on the original contract, so the legacy-handoff error class, its flag, the sequence-only send helper and the mobile handoff hook are gone. * Keep terminal-pane Escape on the plain accepted write Every pane's Escape/Ctrl+C goes through pty:writeAccepted. This branch had switched that IPC to wait for provider settlement, which dropped the "remount this pane" signal for a daemon session awaiting recovery and could stall later keystrokes behind a slow daemon acknowledgment. pty:writeAccepted is back to its original local-only, synchronous write. Prompt answers opt into settlement with requireWriteSettlement on the same channel, and a settled refusal while the daemon recovers now sends the same remount signal. Ordinary verified sends regain their original fallback write. * Report a partly accepted local paste as unconfirmed A settled local write split into chunks returned plain false when a later chunk was refused after earlier ones were accepted. Callers read false as "nothing was written", so chat showed "Message not sent" with a prefix already in the agent's input. It now reports the write as unconfirmed, the same verdict the paired host gives for a partial write. * Hide the chat composer under a prompt card instead of unmounting it When an approval or question card took the input region, the composer unmounted. A message still waiting for its Enter was cancelled and its bubble deleted after the draft had already been cleared, so the message vanished without a notice; composer history was also wiped each time. The composer now stays mounted but hidden while a card owns input, so its state survives. A send that has not submitted yet is still stopped (its Enter would answer the card), but its bubble stays with "Message not sent" so the text is not lost. The composer ref is detached while hidden, so root typing, paste and reveal focus never reach it. * Keep an answered prompt hidden after the chat view remounts The "answered" dismissal lived in component state. Toggling chat to terminal and back, a PTY reconnect, or leaving the phone session and coming back while the approved tool was still running brought the answered approval back, and it then took over the input again. Desktop now keeps the answered occurrence per pane outside the view; phone keeps it per chat tab outside the controller. Both still retire it when the pane observes the prompt clear or change, desktop also when the tab retires, and both maps are size-bounded. * Update the prompt-reply reliability gate for the review fixes Older hosts' accepted answers now dismiss like acknowledged ones, the composer stays mounted under a card, and answered prompts survive a view remount. The gate's invariant, oracle, assertion list, new test files and the two latest evidence runs now describe that contract. * Let users hide a prompt card, keep Escape from denying, and gate only Send on the phone The chat input could stay locked behind a card the host never closes (for example after a Deny typed in the terminal), and Escape on a focused approval card denied the tool even when the user meant to close a picker. - A Hide control (chevron) on terminal approval and question cards, desktop and phone, hides that prompt occurrence and gives the input back. It writes nothing to the agent and uses the same per-occurrence dismissal as an acknowledged answer, so a new occurrence shows the card again. - On desktop, Escape on a card now does the same Hide instead of Deny, and a card that appears while the user is typing no longer takes focus. - On the phone, a card blocks only Send: typing, dictation and image attach keep working on the draft. The placeholder is back to the normal one. * Fix two comments that still called older-host replies unconfirmed Since an older host's accepted write now counts as delivered, the requireWriteSettlement comment and the reliability gate's oracle said the opposite of the code. Both now describe the current rule. * Collapse prompt cards to a strip instead of hiding them, and close the round-2 gaps Hide removed a card completely, so nothing on screen said a prompt was still waiting, and several edges let the chat type into a live prompt. - Collapse (the header chevron, or Escape on desktop) folds the card to a one-line strip above the composer; the strip's chevron expands it back. Collapsing writes nothing, frees the composer, and is disabled while an answer is still being written. Each pane or tab keeps the occurrence as answered or collapsed, so a remount restores the same view. - Questions now carry the host wait's start like approvals, so an identical question in a new wait shows again (desktop and phone). A transcript-only prompt, which has no wait start, is dropped when the view stops observing it, and a transcript still loading no longer clears a dismissal. - Desktop: while a card owns the input, the hidden composer cannot send or interrupt even if it still has keyboard focus, and the card takes focus in the same commit. A send the card retires no longer types Ctrl+U under it. - Phone: an Ask hides the heuristic card read from the same waiting status, and the dismissal store is scoped by host, worktree and tab. * Keep a collapsed card's partial answer, and scope its focus to its own pane Collapsing a question card unmounted it, so expanding it again lost the chosen step, selections and typed "Other" text; Escape typed in that text field collapsed the card. A card arriving while the user typed in another surface (sidebar, notes, a browser URL bar) also took the keyboard. - The collapsed card now stays mounted but hidden (and inert on desktop) under its strip, on desktop and phone, so a partial answer survives collapse and expand. Escape inside the card's text field no longer collapses it. The question card shows the same focus ring as the approval card. - A card takes focus only from inside its own pane (its hidden composer) or from the page body, never from a text field elsewhere. - Desktop and phone share one dismissal store in src/shared, bounded by the existing scope-cache helper, which moves to src/shared with it. - The card send imports the verified helper from its own module, and the phone files are split so each name matches its contents (header action, strip, lane selector). * Return focus to the composer after a prompt card collapses Since a collapsed card stays mounted, Escape or the chevron left keyboard focus inside the now hidden, inert card. The composer's reveal-focus took that as focus already in the pane and stood down, then the browser dropped focus to the page body, so typed keys went nowhere. Reveal-focus now treats focus inside a hidden or inert subtree as not in the pane and focuses the composer. On the phone, collapsing a card also dismisses the keyboard so a hidden reply field does not keep it. * Keep the question card's collapse chevron beside its Cancel button The question card header spread its three items with justify-between, which put the new chevron in the middle of the header. The title now takes the free space, as in the approval card, so the chevron sits next to Cancel at the right edge. * Run the prompt tests on the merged main Main now runs Vitest under Bun, which resolves a long data: URL import as a package name, so the SSH delivery test loads its bundled mobile module from a temp file instead. The phone prompt harnesses mock the live line that main's view now renders, and add Platform, which main's text-selection helper reads, the same way main's own view tests do. |
||
|
|
3ec38b8c6d |
Run Vitest on Bun with Node runtime contracts (#25840)
* Run Vitest on Bun while preserving Node runtime contracts * Preserve runtime timing provenance and keep the Bun pin in config * Scope builtin compatibility mocks to test-only lint exceptions * Give capture retention fixtures distinct filesystem timestamps * Await the copy button success state in the React fixture * Bound Node test worker shutdown and tighten migration fixtures |
||
|
|
13ea35973c |
Stop expensive checks when an unmerged PR closes (#25829)
* Cancel active checks when an unmerged PR closes * Register owned-branch cancellation qualification * Keep temporary cancellation qualification outside the review diff |
||
|
|
f6f96db6be | Build SSH hostile-host Linux slots independently (#25821) |