mirror of
https://github.com/stablyai/orca.git
synced 2026-10-09 08:02:35 +00:00
64bb9373dafdef3bb24d2e409899802704a372ec
1055
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
64bb9373da |
Claude account profiles: dormant WSL guest setup (Step 3 of 4) (#24384)
* feat(claude): add dormant profile setup and history sharing
* fix(claude): make profile setup one gated, typed, fail-safe entry
Review round 1 of the dormant profile setup found that the pieces could
be called without their safety checks, that one failed write or an
unreadable bookkeeping file could silently stop sharing for good, and
that Windows prompt history could bring back history the user cleared.
- One entry, provisionClaudeAccountProfile: the profile gate (namespace,
no linked components, outside ~/.claude and ~/.config/claude, and an
ownership marker beside the home naming the account and target) runs
first and refuses before creating anything; then history sharing,
config provisioning, and the hook install after the settings merge.
Results come back per surface with closed warning codes instead of
message text.
- The sharing ledger is keyed by surface name, records a value only
after its write succeeded, and an unreadable ledger starts empty and
is rewritten instead of blocking every surface.
- The profile state file goes through the same locked writer as folder
trust (Claude's <file>.lock plus the in-process queue), generalized as
updateClaudeGlobalConfig. Onboarding and trust are still applied when
the personal state file is unreadable.
- WSL descriptors build guest POSIX paths; the state-file path style
follows the injected platform.
- Orca's managed statusLine has one owner in a profile: the settings
merge never shares it, a user's own statusLine is shared over it, and
the profile installer follows the default home's slot so a default
opt-out reaches every profile. remove() takes the same destination;
the remote installer cannot accept one.
- Prompt history compares file identity (bigint dev+ino) on every
platform, never drains the shared file into itself, drains retained
copies in generation order, never reuses a stale cursor, and on Windows
keeps a replaced default's old copy aside instead of replaying it.
Directory merges keep going past a failed entry.
* fix(claude): share the user's own hooks and keep merged history whole
A user's own Claude hooks in ~/.claude (notifications, formatters) did
not run under a managed account, because the whole hooks key stayed
private. They are now shared like any other settings key: Orca's own
hook entries and its managed statusLine are stripped from both the
personal value and the profile's current value before the per-key
ledger comparison, so they never travel through the merge and never make
the key look user-owned. Orca entries already in the profile are kept on
write, and the profile hook installer adds them on top as before.
Prompt history: merged bytes that lack a final newline are terminated,
so Claude's next record no longer fuses onto the last merged line. When
a CLI rewrote the profile's history file (old records plus new), only
the lines past the part it shares with the default history are added,
instead of the whole file again.
* fix(claude): close review round 2 gaps in profile setup
Hooks and statusLine sharing:
- When ~/.claude holds only Orca's hook entries, the user's shared hooks
now read as an empty value instead of a missing key. Removing the
user's last own hook in ~/.claude therefore reaches profiles that
never edited it, and deleting the only shared hook inside a profile
stays deleted.
- A custom statusLine Orca shared, and the profile never edited, goes
away when the default home drops it. When a shared custom line
replaced Orca's line in a profile, the profile's statusline marker is
dropped so Orca's line comes back once the default returns to it; a
profile that opted out stays opted out. No other key gains deletion.
- install/remove/getStatus with a profile directory refuse when it is
the default home, or its settings.json resolves to the default one,
instead of editing System Default's hooks and opt-out state.
- The profile statusline rule reads the default settings under the
userHome passed to the setup entry, not os.homedir().
Profile state and ownership:
- A malformed `projects` value skips only folder trust (new warning
code trust-refused); onboarding and shared keys still apply.
- The ownership marker stores only host-local facts (account, runtime,
distro). The execution host id is the caller's view of the host, so
it stays in the in-memory descriptor and is not compared.
Prompt history interruption paths:
- With no cursor yet, a retained copy starts past the bytes it shares
with the default history, so an interrupted share no longer replays
the whole history.
- A retained name for the shared file itself is removed with its cursor
instead of lingering until a later scrub makes it look new.
- The Windows link record is read three-state: unreadable stops the
share instead of reading as "no link". If the record cannot be
written after linking, the fresh link is undone.
- An unreadable retained copy is reported and no longer blocks linking.
* build(cli): list the new Claude hook modules in the CLI project
hook-service.ts and hook-settings.ts are compiled into the packaged CLI
project, which lists every file explicitly. The statusline policy and
profile destination modules they now import were missing, so the CLI
typecheck failed with TS6307. The CLI still loads hook-service through
the existing managed-agent-hook-controls build entry, which bundles
both modules; neither imports electron.
* fix(claude): close review round 3 regressions in profile setup
- A profile whose hooks hold only Orca's entries and that sharing never
recorded is no longer treated as a user edit, so the user's first own
hook in ~/.claude reaches it (for example when the profile was set up
before ~/.claude had any hooks).
- A retained prompt-history file is removed as a second name for the
shared file only when the default history does not itself link to it;
otherwise it holds the only copy and is kept.
- Default-home checks compare file identity: the profile hook
destination check uses device and inode, and the profile/default
separation check resolves on-disk case, so a case-only alias of
~/.claude is refused on case-insensitive filesystems.
- A test pins that an unreadable leftover session tree no longer blocks
linking.
* fix(claude): let shared keys leave a profile when ~/.claude drops them
QA found that removing a setting from ~/.claude never reached a managed
account: deleting the whole `hooks` block left the user's hook running
there. Only statusLine followed the default away.
Every shared key now follows the same rule through the existing per-key
ledger: when a key disappears from ~/.claude/settings.json (or
mcpServers/theme from the personal state file), it is removed from the
profile if the profile still holds exactly what Orca last shared. A
value changed inside the account is kept. Keys Orca never shared,
including denylisted ones, are never touched. Deleting the whole hooks
block removes the user's shared hooks and keeps Orca's own entries. A
missing source counts as empty; an unreadable source removes nothing.
* feat(claude): add dormant profile routing and account consumers
* fix(claude): drop the dormant profile selection RPC; clients negotiate by capability
Restores the inline mobile allowlist so its source-scan guard sees every
accounts.* method again, and the generated params catalog to generator order.
* fix(claude): guard the claude shell function and honour a hand-exported config dir
The function is defined only in a routed pane where claude is a real
executable (the codex function's guard), re-reads the pointer only while
CLAUDE_CONFIG_DIR is unset or still Orca's injected twin, accepts Git Bash
drive paths, and starts on its own line after the fish/PowerShell codex text.
* fix(claude): spawn-time profile env, total account listing, setup at lifecycle triggers
Round-1 review fixes for the dormant profile routing:
- Panes get the selected profile's CLAUDE_CONFIG_DIR plus an Orca twin at
spawn, so nested shells and scripts inherit the account; System Default
injects nothing and its home is the inherited CLAUDE_CONFIG_DIR.
- An absent routing owner is System Default, never a throw; AI Vault and
session-search scans receive profile roots from their parent, and the
capability is advertised only where an owner is installed.
- Account listing never throws: per-account readiness, a stale pointer is
republished in the background and reported on the snapshot.
- Profiles are set up at select and startup; a launch only sets up one that
never was, and a worker fault on a prepared profile is a warning. The
Claude version probe is cached per binary identity.
- Pre-trust goes through the existing deadline- and realpath-guarded writer
against the launch env's profile config.
- Skill discovery keeps a caller's Claude root and a broken Claude selection
no longer fails other providers.
- The durable record carries a provider-neutral launchAccountHome, read
through one helper by the launch fallback and the model catalog.
* test(claude): pin the version-probe cache, launch-account record and temp-home readers
* test(claude): pin dormant bash rc text alongside fish and PowerShell
* test(claude): read the fish launch init without a nullable index
* fix(claude): withdraw the profile pointer when a selection cannot be published
A pointer left naming the previous account would launch it silently; a
missing pointer makes the claude function refuse visibly. A newer selection
that raced the failed one keeps its pointer.
* fix(claude): read the fish profile pointer with read -z for fish older than 3.4
Shell tests skip system config and abort unless claude resolves to the fake.
* fix(claude): only the newest publish withdraws the pointer; total config dir lookup
- An overtaken publish that fails leaves the newer selection's pointer.
- The runtime config dir falls back to the legacy home for an unresolvable
account or a WSL target, so skill roots never fail for other providers.
- WSL guest reader roots merge verbatim, never realpathed on this thread.
- History readers include ~/.claude, where step-1 setup pools profile history.
- System Default ignores a config dir an outer Orca injected (twin-marked).
* fix(claude): System Default launches and probes use the structured create resolver
A Claude agent-env CLAUDE_CONFIG_DIR the create path stored is now the home
the launch pins and the model probe accepts.
* test(claude): type the System Default launch record as an agent-session record
* fix(claude): install profile hook scripts under the setup job's home
A worker thread's os.homedir() ignores its own env, so the hook and
statusline scripts now go under the home the job names. The worker test pins
the process HOME to a sentinel, refuses to run unless the worker sees it, and
asserts nothing lands there.
* test(claude): skip shell cases whose shell the runner lacks
* fix(claude): remove env vars in the PowerShell claude function instead of setting null
On .NET 9+ (pwsh 7.5+) SetEnvironmentVariable with $null creates an empty
variable, so stripped auth vars reached claude as empty strings and the
restore left CLAUDE_CONFIG_DIR empty in the user's session.
* feat(claude): add dormant WSL guest profile setup
* fix(claude): open WSL panes without guest calls and coalesce same-profile publishes
A WSL pane now gets the same non-throwing, guest-free spawn env as a host
pane; only select, startup and Claude launches publish into the guest.
Overlapping publishes of one target share the newest publish while the
selection still names the same profile, instead of failing as superseded.
Publish issues name their WSL distro and drop out when the target is no
longer routed. A late inspect from an older selection no longer replaces
the newer one's verification, a failed guest request evicts the cached
guest, and readiness is derived per account from the guest's owned homes.
* fix(claude): roll back only the target whose selection failed
With profiles, a failed select or remove republishes just its own target
instead of running startup over every WSL distro, and a rollback failure is
logged instead of replacing the error that caused the rollback.
* fix(claude): scan WSL profile history only in running distros
Vault and usage scans pass Claude profile roots through the same
running-distro filter as every other WSL root, so a stopped distro's UNC
paths are never walked.
* fix(wsl): ship the Claude profile helper only in the WSL bundle dir
The helper only ever runs inside WSL from the desktop, so it moves out of
the SSH relay artifacts (no upload, no relay version change) into
out/relay/wsl beside the other WSL-only guest bundles. The three WSL bundle
resolvers share one candidate list.
* fix(wsl): refuse old glibc before downloading, and keep the shared download per caller
The pinned Node runtime needs glibc 2.28, so a distro below the floor is
refused before any download with a message naming both versions, as SSH
hosts are. The shared download again owns its own deadline and each caller
waits on its own signal, and the OpenCode reader keeps its architecture
error text.
* fix(claude): bound each WSL guest operation and run the helper through the WSL runner
A cached guest no longer carries its 180 s preparation deadline into later
requests. The helper runs through runWslProcess (stdin payload, WSL_UTF8),
the distro is confirmed running once per preparation and once per request,
a failed `claude --version` probe continues with an unknown version like
native setup, the helper resolves from the WSL bundle dir, and the guest
entry decodes stdin once so split UTF-8 survives.
* test(claude): cover WSL profile pre-trust routing and its deadline
* refactor(claude): drop WSL refresh cleanup that the failed publish's withdraw already does
* fix(claude): catch rollback failures only when profiles route the selection
With the gate off, select and remove surface the rollback error exactly as
before; only profile routing logs it and keeps the original error.
* fix(claude): give every WSL pane a guest-relative Claude profile pointer
WSL panes now always carry `~/.local/share/orca/claude-profiles/selected-wsl`,
which the bash/zsh and fish claude functions expand against the guest $HOME
at each invocation, so a pane opened before Orca has met the distro still
follows the selected account instead of falling back to ~/.claude. Absolute
pointers are untouched, PowerShell is unchanged, and a missing pointer file or
profile still refuses visibly. CLAUDE_CONFIG_DIR is set at spawn only when the
selection resolves without a guest call.
* test(claude): assert a missing guest-relative pointer refuses with a visible message
* test(claude): type the WSL runner mock in the transport test
* fix(claude): route only WSL distros that hold an Orca account, and re-derive their publish
A WSL distro is routed only while host settings hold an Orca Claude account
for it, decided from settings with no guest call. An unrouted distro behaves
as before profiles: its panes get no pointer or profile env, and a Claude
launch is System Default with no guest prepare. A distro that loses its last
account has its pointer withdrawn best-effort so older panes stop launching
the removed account.
A routed distro without a current publish (for example stopped at startup)
gets one non-blocking background publish from its next pane spawn, coalesced
per target; its failure stays that distro's issue and a later success clears
it. A late setup result from an older publish no longer replaces the newer
selection's verification. The owner contract moves to its own module so the
routing service stays under the size limit.
* fix(claude): read WSL profile history in native chat and adoption only in running distros
Native chat resolves Claude transcripts from host roots first and reads WSL
profile roots only after a miss, filtered to running distros like Codex's WSL
homes. Structured adoption candidates go through the same filter.
* fix(claude): target registration rollbacks and keep their errors in profile mode
A failed add or re-authentication rolls back only the account's own target.
With profiles, a failed re-authentication rollback is logged instead of
replacing the original error; with the gate off both behave as before.
* fix(claude): spell the guest pointer location once and keep set -u safe
The guest helper, the withdraw script and the pane pointer all derive from
one home-relative constant, and the posix claude function reads ${HOME:-}
so `set -u` with HOME unset refuses cleanly instead of aborting.
* test(claude): cover the IPC preflight and daemon WSLENV paths for WSL profile env
The renderer preflight is tested for wsl.exe and Windows shells with a \\wsl$
cwd (which always launch wsl.exe) and with the gate off, the daemon launch
plan imports the pointer and profile home without a WSLENV flag, and the
Windows launch test uses the guest-relative pointer production sends.
* fix(wsl): report why the guest runtime failed, with download context and trimmed stderr
The install's promote output is classified with the SSH classifier, so a
self-test failure shows the exit code and the loader's words (for example a
missing libstdc++ on Alpine) and a security-software change is named. A failed
runtime download says it was Orca's Node runtime for WSL, while a checksum
mismatch keeps its own text. Guest stderr is trimmed before it reaches a
refusal message.
* test(claude): pin that pointer retirement never runs for host targets or with the gate off
* test(claude): give the routed WSL preflight fixture its required authMethod
* fix(claude): let the pane-triggered WSL publish repair a distro stopped at startup
"Distro not running" is now a typed refusal: it never withdraws the pointer
(the distro's last pointer cannot be stale, and a withdraw racing the boot
could delete a valid one) and never records a distro issue. The background
publish a pane fires now waits a few seconds for the pane's own spawn to boot
the distro, probing three times, and is dropped silently and re-armed if the
distro stays down. It joins any publish already in flight for that target
instead of preparing the guest a second time. Per-target generations and
pointer-write ordering move to ClaudeProfilePointerQueue so the routing
service stays under the size limit.
* fix(claude): remove the last selected WSL account without a guest publish
With profiles, removal writes the account list and the selection in one
update, so a distro losing its last account is already unrouted when it syncs
and its pointer is retired best-effort. Removal no longer needs the distro to
be running or able to run Orca's runtime. The gate-off order is unchanged.
* fix(claude): keep native chat's legacy Claude roots first and unfiltered
Only roots added by WSL profiles are read after a miss and filtered to
running distros; a host CLAUDE_CONFIG_DIR on a \\wsl$ share is searched first
and unfiltered, as before profiles.
* test(claude): cover stopped-at-startup repair, launch join and last-account removal end to end
* test(claude): assert no running probe before the pane has had a turn to boot the distro
* fix(claude): let user-initiated profile work boot an idle-stopped WSL distro
WSL distros idle-stop on their own, and the legacy path boots them with its
spawn or \\wsl$ write. With profiles on, a Claude launch, a select, a remove,
a failed-change rollback and the retire after removing a distro's last
account now skip the running pre-check and let their first bounded guest
command (`wsl -d <distro> --exec ...` through runWslProcess) boot the
distro. They refuse only if that command fails, with wsl.exe's own reason,
for example a distro that does not exist. Startup, the pane-triggered repair
and the history readers keep the running pre-check and its typed refusal, so
background work never boots a distro. With the gate off nothing changes.
* fix(claude): let startup join a launch or select already publishing a WSL distro
Startup no longer overtakes a user's in-flight publish for the same target,
so a launch that is booting an idle-stopped distro is not handed startup's
"not running" refusal.
* fix(claude): remove accounts of a WSL distro that no longer exists, and name the helper once
wsl.exe's own failures (exit 0xFFFFFFFF, empty stderr, the diagnostic and its
WSL_E_* code on stdout) are now read by one shared reader used by the git
runner and the WSL profile transport, so profile refusals show wsl.exe's
message. WSL_E_DISTRO_NOT_FOUND becomes ClaudeProfileHostMissingError: with
profiles, removing an account from a distro that no longer exists keeps the
removal and logs a warning, while select and launch still refuse visibly.
The helper's file name is defined once in shared/relay-artifacts.ts and used
by the relay build and the transport.
* fix(claude): give plain fish tabs the claude function through the codex hand-off
Main now gives a plain fish tab Orca's codex function through a vendor_conf.d
snippet instead of a -C init. The claude function only rode the -C path, so a
plain fish tab would not re-read the account selection per invocation once
profiles are on. Define it at the first prompt beside codex; it stays empty
while the profile gate is off.
* fix(claude): share personal rules, themes, workflows and keybindings into account profiles
A managed account launches Claude with its own config folder, so user-level
rules/, custom themes/ (which a shared `custom:<slug>` theme points at),
personal workflows/ and keybindings.json silently stopped applying. Link the
three directories like skills and commands, and copy keybindings.json with the
same edit-preserving ledger as CLAUDE.md. routines/ stays unshared: routines
belong to the claude.ai account and the folder holds per-run state.
* test(claude): wait for the running child to read its account before switching
The test switched the selection after a fixed 20 ms, so under load the backgrounded claude
had not yet read the pointer and picked up the new account. The stand-in now marks when it has
started, and the test waits for that mark (bounded) before switching.
* fix(claude): accept WSL setup warnings for every shared Claude file
The guest reply schema listed CLAUDE.md by name, so a warning about the newly shared
keybindings.json would have rejected the whole reply. It now takes the shared-file list
from provisioning, like the shared folders.
* fix(claude): import the personal CLAUDE.md into account profiles instead of copying it
Claude also loads ~/.claude/CLAUDE.md as a parent folder's memory for any project under home,
so a copied account CLAUDE.md made every such session read the user's instructions twice
(checked live with Claude 2.1.288). An @~/.claude/CLAUDE.md import resolves to the same real
file, which Claude loads once from home, from projects under home and from folders outside it.
* refactor(claude): simplify account profile setup toward the prior art
- Windows keeps each account's history private; drop the hardlink, link
record and conflict-copy machinery that only Windows reached.
- Share hooks and statusLine as ordinary settings keys: Orca writes the
same entries into every folder, so the installer finds them present.
Drops the Orca-entry carve-out, the per-profile statusline follow
logic and its marker.
- Unreadable ledger is just an empty ledger.
- Share from the user's own CLAUDE_CONFIG_DIR when they set one (marked
so Orca's injected value is never mistaken for it), and refuse a
profile at or around it.
- Pin the one canonical profile path spelling in a test.
* refactor(claude): route launches through one account router, superset-shaped
Replace the routing service, owner interface, setup worker thread, reader-root
merging, persisted launch account and capability string with one
ClaudeProfileRouter: the pointer is written first and setup runs best-effort
after it (superset's order); a missing pointer means System default.
The claude shell function re-reads the pointer on every launch, refuses only
a selected account whose folder is missing, and prints a note when the user's
own CLAUDE_CONFIG_DIR overrides the selected account in that terminal.
Still dormant: claudeProfileRoutingEnabled() is false.
* test(claude): type router test settings instead of casting
* fix(claude): run account setup on a worker thread, never Electron main
publish() writes the pointer and starts setup in the background, so neither
startup nor an account switch blocks on a history merge. Each setup runs in a
one-shot worker (the profile-state backup worker's pattern); one setup per
account at a time, reused by later requests. A launch waits only for a folder
that was never set up, and refuses with a clear message if that setup fails.
* fix(claude): do not await the synchronous pointer publish
* refactor(claude): route WSL distros through a small guest router on the Step 2 shape
Replaces the WSL owner/transport/guest-inspect stack with ClaudeWslProfileRouter:
publish writes the guest pointer with one sh command and kicks Step 1's setup
best-effort; prepareLaunch checks the folder over the distro share and waits only
for a never-set-up folder; preparation returns main's WSL shape, so trust, rate
limits and readers need no new code. Setup runs as Linux in the guest on Orca's
pinned Node via a bundled helper (argv in, exit code out), without hooks.
Restores OpenCode's WSL runtime prep, git's wsl-host-failure, wsl-runner,
workspace trust, readers and account selection/registration to Step 2.
Names the guest pointer per Orca build so dev and packaged never share it.
* test(claude): give the routing launch test the merged resolver deps and handle shape
* test(claude): type the WSL routing mock's original() without an inline import()
* fix(claude-accounts): dedupe merged prompt history, drop drained copies, link setup folders by path
- Prompt-history drain appends only lines the shared file lacks, so a purge never re-adds lines.
- A set-aside history copy whose saved offset reaches its end is deleted on the next run.
- Setup folders link to the default home's own entry, not its resolved target.
- The profile gate and folder creation run once, in provisionClaudeAccountProfile.
- installHooks receives only configDir; drop a duplicate test key that fails CI.
* fix(claude-accounts): refuse a routed resume whose transcript is in another account; zsh claude function; setup timeout
- With account routing, a chat resume checks its transcript is in the launch folder; a missing one
with a stored leaf refuses with historyInOtherAccount instead of starting fresh.
- The launch folder of a selected account comes from prepareLaunch(); the resolver stays for System default.
- zsh panes get the claude function like bash, fish and PowerShell (empty while routing is off).
- The setup worker is terminated after 60 s so a later launch can retry.
- Document that the setup marker means setup started, not finished.
* fix(claude-accounts): write the WSL account pointer before a launch returns; one relay bundle candidate list
- prepareLaunch awaits writePointer, so a missing or stale guest pointer cannot run another account.
- Startup's WSL republish runs inside serializeMutation, like rollback.
- relayBundleCandidates takes 'wsl'; the hook relay, browser relay and Claude helper use it, and
wsl-relay-bundle-dirs.ts is gone.
- One setup-marker path helper for host and WSL; the guest pointer path is home-relative and only
the pane value carries '~/'; drop a no-op esbuild external.
* fix(claude-accounts): refuse a routed resume only when the transcript is found in another folder
A transcript found in no known folder keeps the old stored-leaf resume.
* fix(claude-accounts): a WSL launch writes the pointer for the selection current at write time; bound the pointer read
A selection made while a launch waited on setup was overwritten by the launch's stale account.
A hung \\wsl.localhost read no longer stalls startup's serialized publish.
* fix(claude-accounts): record installed hooks as Orca-shared; skip symlink tests on Windows
After Orca installs its hooks into an account, record the account's hooks in
the settings ledger so a later run can still bring the user's own hooks in.
Tests that create real symlinks now skip on Windows.
* fix(claude-accounts): trim the which-account file in the PowerShell claude function
Co-Authored-By: Claude <noreply@anthropic.com>
* test(claude-accounts): spell the user's own config folder as an absolute path on every platform
Co-Authored-By: Claude <noreply@anthropic.com>
* test(claude): skip the POSIX-only WSL profile test on Windows
A WSL profile's data root is a POSIX path, so building one from a Windows
temp dir fails the absolute-path check there.
---------
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
||
|
|
2f49377425 |
feat(native-chat): Grok as a structured chat over the Agent Client Protocol (#25225)
* Leave a stopped turn's running tools to the agent's own end
When another writer settles the open turn (a person's Stop), the assembler
now only stops that turn's text and cancels its pending prompts. Running tool
calls stay the agent's: a progress update or completion it reports after the
Stop lands as reported, and whatever is still running settles at the agent's
turn end for that turn, the next turn's open, or the session's end.
An agent's end for an earlier turn while a newer one is open no longer clears
the open turn's activity line or ends its anonymous reply. An unnamed end right
after a Stop ends the stopped turn instead of being dropped. The test rig's
restart no longer writes the dead assembler's window text, matching dispose.
* Pin that a stopped turn's running tools hold budget until the agent's end
* Type the stopped turn's tool progress update as a tool body
* List every event the assembler hands to the decision step
The type-aware lint requires an exhaustive switch with no default case.
Also retitle a Stop test to say what it asserts.
* End a running call as its turn's journal row ends
A call still running when its turn ends takes the state of that turn's
row: a row another writer settled first (a person's Stop) stands, so its
calls read interrupted whatever the provider's later end reports. The
no-ending path that settled calls from the Stop row is gone, since a Stop
now leaves running calls to the provider. Adds the two Spanish strings.
* Say why a Grok turn failed, and keep task rows in Grok's own words
A failed Grok turn ended with no reason on screen: the translator dropped
every copy of Grok's message. The failed turn now gets one status row in
Orca's existing "provider did not accept this message" words with Grok's
reason, read from whichever copy arrives first (the given-up retry, the
turn's end, the prompt's completion notice, or the prompt's error answer);
later copies only fill a reason the row still lacks.
A running background command no longer reads "Background task <id>
started": a task's summary is mapped only once it has settled. A monitor
stays a monitor when the agent reads its output: a frame that names no
kind keeps the known one, and a "[monitor" command is a monitor.
A prompt's turn is marked started, so a late frame for an ended prompt
neither reopens it nor becomes the active turn. A tool's turn is held in
one place at a time.
* Read a monitor from Grok's exact output prefix
* Word a failed Grok turn in Grok's own text, not as a refused message
A turn that started and then failed was told "The provider did not accept
this message", Orca's sentence for a message refused before its turn. The
row now reads as a Codex turn-ending error does: an error status row with the
provider's own words. With no words, the dialect names the failure ("Grok
ended this turn with an error." / "Grok usage limit reached."), else the
agent's display name does.
* Settle a stopped turn's running call as its turn row ended after a restart too
The restart sweep ended every running call by the death evidence alone, so after
a person's Stop with no proof the child died the call read failed under a turn
that read interrupted. The sweep and the live dead-generation settlement now ask
the same rule the assembler does: a call in a turn already settled ends as that
row ended; only a turn still running leaves its calls to the evidence.
* Keep the dead-generation settlement under the line cap
* Register the ACP schema verify step in the PR preflight phase test
* refactor(agent-session): one required agent registry; declarations admit what they claim
A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.
/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.
Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).
* Read ACP permissions, session events and prompt errors through the protocol client's own types
The translator now reads a permission request with the client's lenient reader, a session update
with its session-event reader, and takes only the agent's own error answer as a failed prompt's
reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values
newer than this build.
* refactor(agent-session): the router applies the declared rewind itself
The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.
* test(agent-session): register the agents the merged-in tests now need
The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.
* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record
* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop
The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.
The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.
One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.
* fix(agent-session): a changed agent definition never hides that agent's chats
A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.
Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.
* refactor(agent-session): each agent's registration says where it runs and which account it pins
createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.
Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.
* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state
A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.
* fix(agent-session): a scoped dismiss-all persists no per-session fence
The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.
* fix(agent-session): refuse an attach whose agent is not the session's own
The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.
* fix(agent-session): offer to start a chat only when the start would accept it
The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.
* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it
A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.
* refactor(agent-session): the record store admits agent ids; comments say where transport is checked
The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.
* docs(agent-session): the record store admits the registered agents' ids
* refactor(native-chat): Grok's registration declares where it runs; ACP no longer borrows Codex's location rule
The rule a self-supervised agent child runs under (this machine, no WSL, Windows only with process
start-time proof) is its own module that Codex and the ACP adapter both use. Grok's registration
takes the full account-home resolver signature, and D3's tests build hosts with the agent registry.
* fix(native-chat): Grok follows the ACP runtime's request contract and the managed process's close
A request the agent or a Stop cancels is answered with the agent's own cancelled reply by the code that
owns it (the runtime no longer answers a silent handler), so a Stop needs no separate decline pass. A
permission answer still being saved when the agent stopped waiting is reported unconfirmed, since the
protocol already answered it cancelled. Cancelling the agent's own turn is the plain cancel. Request
rows are matched under their generation-scoped ids. A refusal's reason comes from the dialect's wording
path. The child drops its own stderr tail and close policy for the managed process's, and a close
whose process tree was not proven gone is reported as the adapter contract asks.
* fix(native-chat): a Grok chat Orca already holds resumes without writing what Grok replays
A chat with a saved Grok session reattaches with session/resume where the agent offers it, else
session/load. Either way the call runs inside the translator's load window, so what Grok sends while
it reattaches (its saved exchange, a task the dead process left running, ended by the restart) opens
no turn and writes no row; only context usage reads on. A reply an Orca or Grok crash cut short is no
longer completed from Grok's saved history: it reads like a Claude or Codex chat's, with the existing
notice. The attach window also closes after a failed attach, and a created session that session/resume
reports missing is replaced like one session/load reports missing.
The replay reconciliation is removed: the lane no longer reads the journal, and D3's replayed-input
grammar test and completed-turn check in the assembler go with it.
* refactor(native-chat): a failed Grok reattach needs no window close of its own; its lane is replaced
* test(native-chat): D3's merged tests use the shipped declarations and the launch options main requires
* fix(native-chat): typecheck fallout of the base merges; any agent's empty chat is reusable
Main's idle-empty-chat lookup and launch join now take any registered agent, as the rest of the
launch path does. The refusal check moved into the prompt turns and the prompt-block conversion beside
the turns that send it, keeping both files in their line limit.
* fix(native-chat): a Grok Stop ends the process once Grok settles its turn; the next send resumes
Grok's session/cancel ends only the running turn: work it already moved to the background keeps
running and can begin a turn of its own after the person pressed Stop. Stop is now a session
boundary, as it is for Claude: the cancel answers open requests and lets Grok end the turn, the host
waits a bounded grace for that, then ends the process; the next send relaunches and resumes.
The adapter's own bounded close of a turn Grok began is gone. Its named-turn check stays: the host
ends the session unless the provider declines a Stop naming a turn that has since ended.
* test(native-chat): a Grok Stop ends the process only after Grok answered the cancel
* fix(native-chat): Steer on a Grok card cancels the running prompt, then sends it
A send that reached Grok while a prompt ran was held in the adapter until that turn ended: Steer
on a queued card took the card out of the host's editable queue and meant 'send after this turn'.
It now cancels the running prompt (session/cancel; the session stays) and sends as the next prompt
once Grok answers the cancel, as the common pattern does; a steer behind another cancels it in
turn, so the last one runs. The adapter holds a send only while that cancel lands, so its general
held-send queue and its holdsDispatch report are gone (every send it holds has its turn open in
the journal). An older client's mid-turn send takes the same path. capabilities.steering is
unchanged and still unread.
* refactor(native-chat): a close or Stop cancels a start through the acquire's own abort signal
The host owns the acquire it runs, so it now owns its cancellation: each attach's acquire gets an
AbortSignal, aborted from outside the session's queue by a close and by a Stop admitted now (the
same admission rule as before). The optional abandonStart adapter hook, the router's fan-out to
every adapter and the ACP adapter's session-keyed start map are gone; the ACP adapter keeps an
unkeyed set of starts only so quit can prove their children gone, and keeps a failed start's
unproven child until its exit is proven.
The hook also let a later close ask that child again. The host now does that from state it holds:
a close of a chat with no live child whose record still names an owner process with no death
evidence asks the adapter to release it. The answer is not recorded as proof (the lease probe
does that), so an owner pid an earlier Orca left is never killed or marked gone. Claude and Codex
ignore the signal and hold no such child; their release is a no-op (tested).
* fix(native-chat): a Grok crash that closes stdout before its exit still ends with Grok's last words
On macOS and Linux the agent's stdout ends before its exit is observed, with or without the
supervisor's EOF forwarding, so the connection's loss closed the journal first and its error text
became the session's ended reason, dropping Grok's stderr. The reason is now read at the proven
exit: the agent's last words when it left any, else why the connection closed. The failure already
carried them. Comments that assumed the exit comes first, that early frames past the cap refuse the
start, and that dispatch re-checks image support are corrected.
* fix(native-chat): nothing Grok sends while a held chat reattaches is written, marked as replay or not
The reattach window relied on the dialect's replay verdict, and Grok's frames read as live unless
they carry isReplay, so an unmarked chat frame during session/resume opened a turn that never
ended. D3 now marks every frame inside the window as replay before the translator reads it, so the
translator keeps only context usage whatever the agent marked; options and commands are still
adopted. The translator's load semantics are unchanged.
* test(native-chat): a Stop after a resume finds no turn an unmarked old reply opened
* test(native-chat): a resumed Grok chat keeps its last context reading; the resume refreshes only the window
* test(native-chat): a Grok background task a Stop ended reads as stopped reporting
* refactor(native-chat): quit's stop of each start answers through one promise kind
* fix(native-chat): quit aborts every start the host has in flight before draining attaches
A Grok that never answered its handshake held quit until the start's own 60 s bound, past the
20 s quit deadline. The host's teardown now aborts each in-flight acquire (and any the drain
still begins), so the adapter's own quit controller and its map of starts are gone: a start
has one canceller, the host's signal.
* fix(native-chat): a Grok start's abort stops reaching its child once the start has returned
The listener stayed on the host's signal until the attach finished committing, so a Close in that
window killed the now-live child behind the host's back and it read as Grok crashing. The start
now detaches it when it ends; a later Close goes through the session's own stop.
* fix(native-chat): a close or Stop during any attach phase stops the start before it launches
The attach began its abort controller only after reconciling leases, resolving recovery and
probing the previous owner, so a close or admitted Stop in those phases reached nothing and Grok
launched anyway. The controller now begins first, and the acquisition checks it before asking the
adapter to start.
* test(native-chat): a close during the attach's owner probe asks no adapter to start
Also renames the close test after the hook it no longer exercises.
* test(native-chat): a close's re-ask closes a Claude or Codex child a failed cleanup left
The re-ask is not a no-op for them: when the adapter still holds the child its cleanup could not
prove gone, the close stops it again as a requested close, and Claude persists the handle of the
conversation it ran so the next send resumes it. Corrects the tests' and comment's wording; the
close awaits the re-ask, bounded by each adapter's kill ladder.
* fix(native-chat): Steer during a turn Grok began itself cancels it and sends once it ends
A send while Grok ran a turn of its own (a background task waking it) went straight to Grok, which
queued it behind that turn where Orca could no longer withdraw it, while Stop treated the same turn
as the running reply. The send now waits as a steer, the turn is cancelled once, and the message
goes when the turn ends; a Stop withdraws it and an exit rejects it as never sent.
* test(native-chat): a steer whose cancel Grok never answers ends Grok and is rejected as never sent
Pins the bounded steer cancel kept from the runtime: past the bound the connection closes, the
running reply reads unverifiable, Grok's end reads as its exit, and the waiting steer is rejected
as never sent.
* fix(native-chat): a Grok crash stays a crash when a stop lands before its exit is proven
After the connection broke and the close could not prove Grok's exit, any later stop Orca asked
for (the next start, a Stop, a Close) marked the child as closed by Orca, so the crash read as a
requested close and Grok's last words were dropped; a send meanwhile was recorded unconfirmed.
The connection loss now decides the cause, and a send on that session is rejected as never sent.
* test(native-chat): fixtures this PR's registered Grok and desktop capability made stale
CI's unit shards failed on tests outside the PR's own lists. Each encodes something this PR changes
on purpose: Grok is now a registered agent (the seam test's unregistered agent is now Cursor); the
desktop now advertises registered agents (the restart-offer tests' older client drops that
capability explicitly); the attach context carries the start's abort controllers (the forget-status
double gains them); and the ACP real-host test rig sends to the host directly (listed beside the
other real-host rig in the send ratchet).
* fix(native-chat): a start quit stops is not the queued message's start failure
With quit now aborting a start it would have waited for, the delivery step recorded the aborted
start as the message's failure ("couldn't restart"). After quit has stopped delivery, the step
leaves the message to quit, which settles it as a close does ("The chat closed before this message
was sent."). The test that pinned quit waiting for that start and stopping its child now pins that
nothing is launched behind quit.
* fix(native-chat): a message sent after a Stop or close aborted a start gets its own start
A start the host aborts (an admitted Stop, a close, or quit) returned its refusal to the delivery
loop, which then rejected whatever was queued at that moment with "couldn't restart", including a
message the user sent after the Stop. The attach now reports that the host aborted it, and the loop
re-derives from the journal instead: what the Stop or close withdrew is already settled, a message
accepted since gets a start of its own, and quit's next step stops the loop. This replaces the
quit-only carve-out with the same rule for every abort and every agent.
* test(native-chat): the message sent after an aborted start is answered, so no settlement outlives the test
* fix(native-chat): a Grok model pick Grok never answers no longer holds Stop or Close
The pick runs on the session's queue. It now registers in the host's out-of-queue
abort registry beside a start, so a close, an admitted Stop or quit abandons it, and
the ACP adapter bounds it at 30 s like Claude and Codex. A late answer is still adopted.
* fix(agent-launch): a phone's launch opens a terminal for an agent whose chat it cannot show
agent.launch now reads the caller's capabilities by the rule tabs and restart offers
use (clientRendersStructuredAgent). A phone without registered-agents.v1 gets Grok as
a terminal again, as on main; the host's own callers and desktop clients are unchanged.
* fix(acp): strip every agent hook variable from the ACP child, from the shared list
ACP_CHILD_ENV_TO_DELETE was a second copy of the hook runtime keys that missed
ORCA_AGENT_HOOK_TRANSPORT; it now spreads AGENT_HOOK_RUNTIME_ENV_KEYS beside the pane
identity keys.
* refactor(native-chat): the mutation context carries the provider-wait registry itself
Keeps the host file within its line limit; one field instead of two closures over it.
* fix(agent-launch): agent.launch.v2 still vouches for Claude and Codex chats
The caller rule from the previous commit also turned Claude and Codex into terminals
for a client advertising only agent.launch.v2, whose contract says it opens a chat
(mobile retry-authority tests). Only an agent beyond those two now needs the client to
read it (clientRendersStructuredAgent); the test fixtures go back to what they were.
* refactor(native-chat): drop saved-history adoption from the timeline assembler
The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.
* refactor(acp): drop session/load history adoption from the translator
The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.
* refactor(native-chat): a pending input is only Orca's send now
Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.
* test(acp): keep the task-result status table on live frames
Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.
* test(acp): a frame helper for a shell command Grok is running
* fix(acp): a Grok crash settles through the host's provider-exit batch, scoped to the turn it ended
A Grok crash ended the journal unverifiable before the adapter reported the exit, so the host's
provider-exit settlement found no running turn and wrote nothing: the adapter's failure (with
Grok's last words) never reached the journal, and a later stale-session pass wrote a bare,
thread-scoped cut-short row, so the partial reply was not folded as Claude's and Codex's are.
At a proven exit the ACP lane now ends its running turn interrupted at the exit instant, as the
host's exit contract expects of a child's own translator (Codex's does the same). When Grok's
stdout closed first (every POSIX crash), the turn is unverifiable only until the exit is proven:
the host's provider-exit settlement now takes the exit as proof naming the child's fence and
revises what that child left unverifiable in the same batch, with the turn-scoped row and the
adapter's failure. Claude and Codex write no unverifiable turn of a live child except a command
whose hand-off is in doubt; that turn is now revised at the exit instead of at the next open.
* test(acp): a crash seen first leaves the host no Grok turn to revise
* refactor(native-chat): what a gone generation left unfinished gets its own module
The settlement file passed 300 lines with the exit-proof revision. The unfinished-work reads
(capture, interrupted-by-the-exit, in-progress) are their own concept and move out unchanged,
apart from the exit proof they now take.
* refactor(native-chat): a watched exit revises what its child left unverifiable without reading Stop marks
An exit's own instant is the turn's end, so the revision needs only each row's fence: the
settlement's journal type gains itemFence alone, and the host test fakes say so.
* test(native-chat): drop the duplicate itemFence on the fake that already had one
* test(claude, codex): an exit whose stdout ended first still reports as it always did
The provider supervisor now ends Orca's stdout when the agent's ends, so on every crash EOF
arrives before the exit is seen. Claude's and Codex's connections report nothing at EOF and
report the exit, with its usual reason, once it is seen.
* fix(acp): reopen a chat with session/load, as the common pattern does
An agent that offers both now reloads its session instead of resuming it; the
reattach window still discards what it replays except context usage.
* fix(acp): drop the 60 s handshake bound; an abort fails the start's waits at once
Neither common design bounds an ACP handshake: Close, Stop and quit end a start
that never answers. The abort now also closes the connection, as a kill there
does, so the start settles even before the child's exit is proven. The
host-stopped start refusal only this bound produced goes with it; the idle
sweep keeps its words.
* fix(acp): a Stop naming an ended turn follows Claude's rule
It still stops nothing while another turn is live, but in the gap before a
follow-up's turn opens, which no client can name, it now stops what is in
flight and the session ends, as a Claude Stop does.
* fix(native-chat): a close no longer re-asks a failed start's unproven child
Neither common design retries that stop at Close, and Orca's Claude contract
re-asks only at the next start and at quit. The ACP adapter keeps the child
until its exit is proven and asks it again there, as Claude does.
* fix(acp): a message sent during a turn the agent began itself goes at once
Both common designs send it straight to the agent with no cancel; only Orca's
own running prompt is steered (cancelled, then re-prompted).
* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer
Uses an audience production sends (one that cannot show every agent), per review.
* fix(acp): launch Grok as `grok agent stdio`, without the update and leader flags
The common pattern passes neither --no-auto-update, --no-leader nor
GROK_DISABLE_AUTOUPDATER; full access still adds --always-approve.
* fix(acp): an agent that ends its stdout, or answers unreadably, is not a lost connection
As in the common pattern, only a broken stdin (or Orca's own close) ends the
agent; one that closed its output but can still be written to stays until a
Stop, a close or its exit. The provider supervisor goes back to its base
content, so Claude and Codex no longer get the forwarded stdout end either.
* fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts
On main every close of the chat writes its own Stop and settle. Here a later close joins the
first and writes no row, and the first's settle closed when its kill failed, so a turn that opened
in between and was cut by the next close read as failed. A person's close joining a person's close
whose Stop opened a settle now reopens that settle until its attempt is done.
Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and
the next reads as the person's cancellation (each fails without its half of the fix).
* test(native-chat): Grok opens as a chat only behind the structured-chat setting
agent.launch and orchestration worker-start read the same setting as the
renderer route; pin both states for Grok on each. The setting's description no
longer names only Codex and Claude, in every catalog.
* docs(acp): generic ACP comments say what holds for every agent, not Grok
Stop ends the session for every ACP agent, as in the common pattern; the
adoption hook comment goes (adoption is not planned); a failed start's child is
retried at the next start or quit.
* test(claude, codex): type the EOF-before-exit test's streams; the supervisor no longer forwards EOF
The Claude test wrote to the child's stdout and stderr through their Readable
type, which the node typecheck rejects; it now holds its own PassThrough
streams. The comments no longer credit the reverted supervisor change.
* feat(acp): a steer's cancel asks once and never ends the agent
The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.
* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge
* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns
Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.
* fix(native-chat): drop the stopDelivery the A3 merge doubled
* fix(acp): a steer's cancel asks Grok once and never ends it
A steer now uses D1's notify-only cancel. Two messages sent during a reply Grok began itself
cut that reply, as the common pattern does, and then both run; before, the queued first
message could not answer the bounded cancel and Orca ended Grok although Grok answered.
A Stop keeps the bounded cancel and its 4 s grace.
* fix(acp): a permission Grok asks with no prompt of Orca's running is declined
During a turn Grok began itself nobody asked it to act, so the request is answered
cancelled at once instead of opening a card that waits, as the common pattern does.
* fix(acp): a Grok that dies while starting is reported with its own last words
A dying process's stdout ends before its exit is seen, so the start failed as a closed
connection and Grok's stderr was lost. A start whose connection closed now waits, bounded by
the Stop grace (or a Close/Stop), for the exit before it is told.
* test: a Stop after a steer sends its own cancel; drop the import the A3 merge doubled
* test(native-chat): main's Stop-note test builds its turn context with the agent registry
* test(claude): say why the close test's fake child cast is safe
* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way
The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.
* test(native-chat): build the Stop-opened-turn test's identity with the opaque handle
The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle.
* test(ratchet): require src/main/provider-process now that it has landed
* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test
Main's #25706 and this branch both added the import at different lines; the merge kept both.
* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test
Main's #25706 made the same fix as this branch at a different line; the merge kept both imports.
* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore
A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.
Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.
* test(native-chat): build this stack's journal identities with main's opaque provider handle
Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test
files from this stack still wrote the old shape. Same lines the downstream ACP branch uses.
* docs(acp): every reattach drops the agent's replay, not only for a chat the journal holds
* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row
Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.
* refactor(native-chat): read hosts' structured agents from the app-shell services
Main grew the startup hydration hook to its line limit; the host agents sync is an app-lifetime subscription like the structured session tabs sync beside it, so it moves there.
* Use current provider handles in transition tests
* Use current provider handles in timeline fixtures
* test(native-chat): prove replacement rows survive downgrade and re-upgrade
* Require the ACP directory in the runtime import check
* test(ratchet): require src/main/acp now that this PR lands it
* feat(acp): a saved session the agent cannot reopen continues in a new one, with one warning row
When session/load (or session/resume) of a saved ACP session fails, the chat starts a new session and records it as a creation that replaces the lost one (#25747's 'replaces' link), and writes one warning row that the agent no longer remembers the earlier messages. A created session the agent reports missing is still superseded silently; a signed-out agent or a start that is over (Close, Stop, a lost agent) still fails the start.
* chore(acp): rewrap the acquire header comment
* test(acp): a start closed while the agent reopens fails without opening or announcing a new session
* Let ACP connections own their supervised agent process
* Preserve ACP cleanup evidence and isolate exit observers
* Expose ACP cleanup observations and type the permission fixture
* refactor(native-chat): the registered-agents capability lives in its own module
Main's growth put protocol-version.ts one counted line over its 300-line limit once the capability
was added; like main's other per-feature capabilities, it now has its own module, and importers read
it from there.
* refactor(acp): one connection owns the Grok process and its protocol
D3 now opens each ACP agent through createAcpAgentConnection (ACP-ALIGN #25810): one object spawns the
process on the execution host, owns its stdio and protocol, and reports its proven exit. It is built and
tracked before the handshake, so a start's abort (Close, Stop, quit) still reaches it, and a failed start
keeps that same connection for the next close to retry rather than spawning another process.
Deleted: the spawnAcpStructuredChild wrapper and its test, the raw-stream runtime assembly, the caller's
exit -> runtime.close wiring, the stdout-EOF heuristic (the connection no longer treats stdout EOF as
exit), and the 10 s steer/Stop cancel bound with requestSteerCancel. Reader control maps to
pauseReading/resumeReading; a close is connection.close after the host's existing 4 s Stop grace.
The adapter owns what the protocol no longer does: one session/cancel per running prompt however many
steers arrive (cleared with that send's settlement, retried after a failed write), and a Stop or steer
answers every open agent request the person has not already answered with the agent's own cancelled
reply. An answer already being saved when the Stop lands is sent.
Tests: blocked cancel write never holds Stop's grace, two quick steers send one cancel, a failed cancel
write is retried, a real process exiting while a child holds its stdout ends the session, and the
existing start-abort, retention, crash, connection-loss and reload-failure suites on the new rig.
* fix(acp): Grok signs in on its own machine with its API key or cached sign-in
When Grok reports that it needs authentication, Orca now names a sign-in method on the machine Grok runs
on, read from the same environment Grok was launched with: xai.api_key when XAI_API_KEY is set there and
Grok offers that method, else cached_token when Grok offers it, else none and the chat keeps the existing
not-signed-in refusal. The rule lives in Grok's launch spec; the adapter applies any agent's rule for new
and reopened sessions through the protocol client's caller-named method (authenticate, then retry once).
No new sign-in UI; interactive methods are never chosen.
* fix(acp): the adapter decides which of Grok's requests reach the person
The turn owner now admits every agent request, permission or question, from its own turn state: a
request reaches the person only while Orca's prompt runs and no steer or Stop is cutting it short (a
question may also come from a turn Grok began itself, until a Stop). Anything else gets the agent's
own cancelled reply and opens no card, so a question arriving after Stop or during a steer never
appears. A steer, like a Stop, withdraws the requests already open; an answer already being saved is
still sent. The protocol client's abort-on-cancel path is no longer used: after the connection
change its request signal aborts only when the connection closes.
* fix(acp): a plan Grok proposes shows as a plan, with no approval card
When Grok leaves plan mode it asks the client to approve its plan (x.ai/exit_plan_mode). Orca showed a
blocking 'Approve plan / Request changes' card for it; the common pattern has no such gate. Now the
plan goes into the chat's existing Plan row (the plan-document status row Codex and ACP plan updates
already use) and the request is answered at once with 'abandoned' plus feedback telling Grok to stop and
wait for the person's feedback or a request to implement it in a later turn, so nothing is approved on
the person's behalf. Dialects gain settleRequest for requests answered without asking anyone.
* fix(orchestration): worker-start opens a Grok worker in a terminal, as before
With the structured chat setting on, worker-start decided 'structured' for Grok and then the structured
worker factory (Claude and Codex only) refused it, so the start failed; main opened a terminal Grok
worker. Worker-start now decides with no registered agents beyond Claude and Codex, so Grok gets a
terminal worker as before. agent.launch and the app's own launches still open Grok as a structured
chat. Temporary until structured workers take registered agents.
* fix(acp): a prompt answer Orca can't read ends the turn instead of hanging it
A session/prompt rejection that was not the agent's own error answer (an answer that fails Orca's
schema, or one too large to read) left the turn running: the next message became a steer with nothing
to cancel and was never sent or settled, and Stop waited its full grace. As in the common pattern, any
prompt failure now ends the turn as failed (a failed-turn row without words, since none are the
agent's) and settles the send, so the next message goes. Only a closed connection keeps the send
running, for the connection-loss path to settle.
* fix(acp): send Grok's prompt-identity extension only to agents that echo it
session/prompt carried _meta {promptId, requestId} for every ACP agent, though only Grok's dialect
echoes it (injectedPromptIdentity). Now only an agent whose dialect declares it gets the extension;
other ACP agents get a plain prompt.
* refactor(native-chat): the registered-agents capability lives in protocol-version again, as on main
This reverts
|
||
|
|
76d480f808 |
Fix unsafe test fixtures and the Bun version pin (#26051)
* Keep test interruption signals within owned processes * Pin Bun and add optional unit runner shutdown diagnostics * Unblock CI lint without changing session host runtime * Avoid duplicating runtime import-check dependency bundles * Leave unit runner diagnostics disabled by default * test: keep runner incident follow-up focused on durable guards * test: apply transcript replacements as authoritative snapshots |
||
|
|
d3e1494674 | test: bound memory used by the runtime Electron audit (#26049) | ||
|
|
0bdcaf36ed |
fix(claude): start a Claude chat with its saved options and send the first message at once (#25152)
* fix(claude): end a Claude start that never answers initialize after 120 s * Read the Claude startup deadline inside startup; fix a stale test comment * fix(claude): start a Claude chat on its initialize answer, not on a frame only a SessionStart hook sends Startup waited for system/init or a SessionStart hook frame as well as the initialize answer. Before the first turn only a SessionStart hook sends one, and Orca adds that hook only through its optional status hooks, so with them off the first message was held forever. Startup now lands on the initialize answer; a start frame already seen is still checked, and one naming another session ends a started session. The deadline drops to 90 s so it fires inside the host's 120 s start wait. * fix(claude): time the Claude start by silence, and fail it at once on another session's frame Claude answers initialize only after its SessionStart hooks finish, so a total-time deadline would fail every start behind a slow hook. Each start frame now restarts the clock. A frame naming another session fails a start still waiting on initialize at once, as before. * test(claude): a real Claude chat starts and answers with every hook disabled * Say what the start-frame re-arm covers, and check only start frames in the hook-less real test * fix(native-chat): a Claude chat starts with its saved options and takes its first message at once Saved model, effort, Fast and permission mode are passed as launch options, checked against the account's cached model catalog, instead of restored by control requests after initialize. With nothing left to restore, the host no longer holds a message until the CLI answers initialize, and the 90 s startup deadline is gone. A Stop on a start that never answers ends the child and settles what it was handed as stopped. A failed result that repeats the turn's own API error reply writes no second row. * fix(native-chat): a host stop of a Claude start fails the message it was handed, with one row With no start-hold the delivery loop no longer sees a host stop of a start it waited on. The child's end now rejects what it handed over with the host-stopped words and writes the one row, as an exit of its own would; an idle start the host stops still goes quietly. * test(native-chat): a Claude chat's first message is written before initialize answers Rewrites the tests that encoded the start-hold, the startup deadline and the option restore to the new contract, and adds: saved options at launch (catalog checks, bypass, fresh-session Fast), a message written before initialize answers (adapter and runtime), Stop on a start that never answers (stopped, child closed, nothing working), and an API error said once. * revert(native-chat): keep a failed Claude turn's error row The shared turn fold already shows a failed turn's error once after it settles, as an error; dropping the row left the CLI's synthetic reply looking like something Claude said. * fix(native-chat): pass saved Claude options unchecked, heal a retired model on the CLI's word, and never leave an unrun message in doubt - Saved model, effort and Fast are launched as picked; only values no Claude can parse are left out. The pre-spawn cache check is gone. - A saved Fast on for a new conversation is applied once the settings readback shows no per-session opt-in (dropped when there is one, or when the model is listed without Fast), with nothing waiting on it; the record keeps the pick. - Under an Agent Permissions bypass, a saved narrower mode launches with the allow flag so bypass stays reachable. - A turn whose reply is the CLI's model_not_found for the launched model drops that model from the record; the launch's own row for it is kept out of the account model cache. - A child that ends before it answered initialize, for any reason, settles every message it was handed as not sent (cancelled for a Stop). - A launched effort the CLI reports only as `applied.effort` is confirmed from there. - The untimed-initialize comment is back to main's text. - A real-CLI test for a message written before initialize answers, under saved options. * fix(native-chat): type the close's ended event and the start-exit test fixtures The close's ended event is typed as the adapter event so its optional startupUnanswered spread fits exactOptionalPropertyTypes; two tests guard the fixture's optional generation, and the hung-start fixture records initialize on the fake connection it holds. * fix(native-chat): a Claude model heal keeps a later pick, a refused Fast is dropped, no allow flag - `options-skipped` carries the retired value; the record drops it only while it still holds it. - `started` carries the values a heal retired, and the record does not take them back from the CLI's report of the same value. - A saved Fast on a new conversation is applied before `started`: a refusal drops it and records it skipped, as main's refused restore did; silence keeps it wanted and unconfirmed. - A saved narrower mode under an Agent Permissions bypass launches without any bypass flag again: the allow flag is one older CLIs reject at start. Kept as a known limit. - The real-CLI test asserts the message was written before initialize answered. - The fake reports a launch effort only under `applied`, and a misplaced doc comment moves back. * fix(native-chat): a new Claude chat reports started before its saved Fast is applied The Fast apply on a new conversation now runs after `started`, so a Stop interrupts a running first turn and an option write is not refused while the round trip is out. A refusal drops the pick through `options-skipped`, in order after `started`; silence keeps it unconfirmed. The launch's unreachable skipped-model branch is gone. * fix(native-chat): a healed Claude chat goes back to the default model live; comments match the no-hold design When the CLI says the launched model does not exist, the live child is also put back on the CLI's own default (set_model with no model, fire-and-forget), so later messages in the same chat run; a user's pick sent after it wins, and a refused or unanswered reset only logs. Comments that still described the start-hold or the option restore now describe the launch options and the handed-over, never-echoed rule. * fix(native-chat): a message handed to a Claude start that never answered is kept as main keeps an unsent one A child that ended before it answered initialize ran nothing it was handed, the same fact as a send accepted and never handed over. Its end now settles those sends exactly as the chat settles a queued send for that end: a quit keeps a person's message as a held card (restart words), a close keeps it as a held card (closed words), a person's Stop withdraws it as cancelled, and a host stop fails the start with one row. * fix(native-chat): a quit during a Claude start that never answered offers no resume for the message it keeps as a card The restart snapshot now reads the same never-answered fact the exit does, so a message handed to such a start counts as queued work, not as work to resume. The retired-model reset comment names the default it really applies. * refactor(native-chat): the saved permission-mode launch helpers live with the spawn options that use them Keeps claude-structured-launch-resolution.ts within max-lines once merged with main, and names the hung-start test envelope's field type. * fix(native-chat): a Claude chat's saved options take precedence over the agent Arguments' own flags Main now passes the saved agent Arguments to the Claude child, and the SDK writes them after its own options. An Arguments --model or --effort therefore reached the CLI as a second flag after the chat's saved pick (a commander CLI keeps the last), and a saved Fast's launch settings replaced an Arguments --settings file outright. The saved model and effort now stand in for the Arguments' flags, and a saved Fast beside an Arguments --settings is applied by the start instead of at launch. * test(native-chat): the hand-built Claude session in the options test carries fastModeAtStart * test(native-chat): the queued rig's start spy carries a named Mock type An unannotated vi.fn() inferred @vitest/spy's internal Procedure, which CI's typecheck cannot name in the factories' inferred return types (TS2883). * fix(ci): run the PR's SQLite-backed tests in the Node runtime project Lists the hung-start Stop test, renames the send-during-startup entry from its old name, and carries main's own two entries from #26010 so the boundary test passes before the next merge. |
||
|
|
b3b6c5dc13 |
Give native chat names one source for tabs, sidebar and AI Vault (list and search) (#25986)
* Give native chat names one renderer source and drop Vault's name repair copies The host's saved conversation name now rides the structured session status feed, which already exists per host, is keyed by the durable session id, and keeps a closed chat's summary. Tab strip, sidebar rows and AI Vault (list and search) read it through one hook and one display order (tab alias, saved name, host label). Vault no longer copies names into its cached results, so the projection, recovery and pending-title modules and their tab-snapshot lanes are removed. Indexed search hits now carry the native owner and saved name from the host that indexed them. * Type the sidebar name test fixture without an assertion * Keep Vault search working when the chat host will not install Naming and owning search hits is bookkeeping: if the native chat host fails to install, return the plain hits instead of failing the search. The runtime RPC only installs the host for clients that will receive the owners. * Publish chat names to the feed independently of the tab retitle A failed feed publication no longer skips retitling the open tab. The publish now lives in the naming deps, where a test covers it. * Note why the status feed must keep closed chats' summaries * Bound names and owner ids that come from a paired host Drop a published chat name the record store would refuse, and cap a search hit's owner workspace id at the same length the list row and record use. * Let native chat search hits from a paired host open their chat A paired host's search hits carry no resume command, so the row disabled every open action even for a native chat it can open through its owner, as its list row does. * Ignore workspace ids that name object members in tab lookups A paired host's row or search hit could carry a workspace id such as "constructor", which read an Object.prototype member as a tab list and broke the render. The shared tab index now reads only own workspace entries. |
||
|
|
825d7bd5a9 |
test(vitest): run agent-launch-instant-tab in the SQLite runtime project (#26028)
#25430 added a test that opens a real agent-session record store, but not to the SQLite runtime list, so vitest-sqlite-runtime-boundary fails on main. |
||
|
|
3fb72d135d | Run Node event-loop measurement after ordinary test suites (#26015) | ||
|
|
37ff3873a0 | Run combined localization catalog verification on Bun (#25999) | ||
|
|
f0520851ab | Keep mobile restore SQLite fixture in the Node test runtime | ||
|
|
1040f673b1 | Keep orchestration SQLite fixtures in the Node test runtime | ||
|
|
3308ff8b26 |
Bound YAML merge conversion and close SQLite routing review gaps (#25998)
* Close test runtime and YAML merge review gaps * Bound YAML conversion inside explicitly tagged pairs * Make completion notification fixture cadence deterministic |
||
|
|
faed899cd3 |
fix(ci): run three new SQLite-backed tests in the Node runtime project (#26010)
#25888 and #25766 added tests that import the orchestration database or the structured session runtime without registering them in the Node runtime list, so vitest-sqlite-runtime-boundary fails on main and every PR. |
||
|
|
d8c871a1f0 |
Speed up unit tests with cross-runtime duration scheduling (#25967)
* Schedule unit tests across runtimes by measured duration * Keep new SQLite fixture suites on Node after updating main * Inject scheduling timings instead of mocking the module * Keep agent-session database lifecycle contracts on Node |
||
|
|
72b84118e0 |
test(e2e): terminal layout parity check against main (#25681)
* test(e2e): add terminal layout parity check for topology refactor PRs Runs fixed terminal-layout journeys in the real app on two builds, captures the renderer topology and the saved workspace session, normalizes volatile values and fails on any difference not declared for a named bug. Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): record quit exit status and report paths main does not reproduce Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): invoke pnpm correctly under corepack and silence the typeless-module warning Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): accept pnpm's forwarded -- in the layout parity runner Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): close parity panes through the user's chord and treat absent maps as empty Driving PaneManager.closePane directly left main to learn of the close from the PTY exit, which raced quit; the keyboard path commits the close in main first. Co-Authored-By: Claude <noreply@anthropic.com> * test(e2e): build parity checkouts before running, forward -g, and add a drag-out scenario Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
6aa12c30ef |
Name native Claude and Codex chats after their first message (#25724)
* feat: generate chat names through configured text agents * Project structured chat names across stored tabs and session rows * Name native chats from their first live message * Restore the journal test provider handle import * Fix first-message naming and live Vault title updates * Preserve Unicode characters in bounded chat naming prompts * Read chat naming settings only when a turn needs them * fix(chat): store only generated conversation names * fix(chat): keep ordinary labels across unnamed chat surfaces * fix(chat): preserve naming after fast first turns * Preserve native chat names across command-first sends and Vault lifecycle * Route native chat SQLite contracts through the existing Node test pool * feat(settings): add chat naming controls to Chat page * Preserve chat naming drafts and configure Custom commands through host settings * Scope synthetic command output to its journal thread in naming test * Keep naming test fixtures within their typed project boundaries * Resolve the real Vault hook directly in its integration test * Publish chat naming save refs after render commits * fix: keep Japanese chat naming copy stable during repair * Update test contracts for the naming integration * Keep journal fixture reads on the host clock * Give Chat names a separate settings section |
||
|
|
f7c542c7a3 |
Keep profile saving alive after a stalled main loop (#25318)
* Keep profile saving alive after a stalled main loop After a long main-loop stall (overnight sleep, dark wakes), the profile writer's overdue 30s timeout could run before an acknowledgement that was already queued, permanently retiring the writer until restart. Terminal creation then failed because pane bindings could not be saved. - Writer deadlines measure lateness on the monotonic clock and grant a bounded fresh window when the callback is overdue or the system reports suspended; resume re-arms without spending grace. Applies to initialization, every command, and the close/exit wait. - The "Saving stopped" alert is parented to a visible main window (never a parentless synchronous macOS alert), deferred until shown, deduplicated, and says whether the latest change is unconfirmed. - Timeouts, grace, writer faults, and alert presentation leave sanitized durable breadcrumbs. * Fix profile writer timeout and shutdown races * Run profile writer stall regression on Linux and Windows * Keep Electron probes out of headless runtime qualification |
||
|
|
83cf7cf5e2 |
perf(ci): compile release JavaScript once for all packaging hosts (#25828)
* perf(ci): share release JavaScript across packaging hosts * fix(ci): verify the projected web entry in release archives * fix(ci): use the Windows system archive tool for release bundles * test(ci): retain stylesheet evidence in release build comparisons * test(ci): verify release parity across native color rounding * test(ci): normalize manifest asset references without changing import order * test(ci): compare portable outputs across Windows text and color formatting * test(ci): preserve module identity across dependent asset hashes * fix(ci): keep SVG build inputs identical across release hosts * fix(ci): stabilize compiler inputs and projected web bindings * fix(ci): retain vendor minification in projected web output * test(ci): normalize platform-specific pnpm manifest source paths * fix(packaging): exclude shared build staging from application files * test: align thinking-state fixtures with the current source shape * test(mobile): reuse message fixtures within the line limit |
||
|
|
ac8ea9f958 |
fix(codex): write Orca's hook into ~/.codex only when something changed (#25743)
* fix(codex): write Orca's hook in ~/.codex only when something changed One reconcile replaces the per-launch writer and its background approval session. It runs at app start after PATH hydration, when the setting turns on, once per native pane spawn, and (with a bounded 3 s wait) on Codex launches and resumes. Orca's entry lives alone in a matcherless group, appended last unless already in place; other copies are removed and the shifted user approvals move verbatim under both key spellings. The approval, with Codex's own hash, goes in first. The routing gate closes only for a hooks.json with a bad shape. * fix(codex): report hook status for the home the next pane uses Status reads ~/.codex (both key spellings) when launches use it, or the CLI has no selection, and the selected managed home otherwise. It explains: update Codex, Codex not found, not asked yet, approved by Orca but not yet confirmed by Codex, a hooks.json shape that moves panes to Orca's own home, and inline config.toml approvals Orca cannot add to. * test(codex): pin the ~/.codex approval against a real Codex, and run it on those files The contract now also writes Orca's entry into a throwaway ~/.codex and checks that Codex lists it trusted and enabled, turns it back on over a /hooks switch-off, lists it for review after a user inserts a hook ahead until the next check, keeps an inline config.toml loadable, and shows no review in a real TUI start. The real-binary CI job now runs when the ~/.codex writer changes. * test(ci): keep the scope test under its line cap; assert the ~/.codex paths beside the contract job * test(codex): type the hoisted test holders instead of asserting * test(codex): cover the matcher rule, an older build's event, and the spawn trigger Adds tests that failed against mutants which survived the first pass: an entry alone in a matcher group, a current copy's approval in an event left to an older build, one run per spawn, a spawn riding a running reconcile, and the native-only spawn trigger. The opt-out's Codex-hash cleanup now runs once, after the sweep, instead of twice. * docs(codex): name the ~/.codex reconcile in the legacy sweep's lane comment * fix(codex): approve ~/.codex with Orca's own hash while Codex's answer is pending The reconcile waits at most 0.5 s for Codex's hash. If the lookup is still running, an entry already in place keeps its approval and nothing is written; a missing entry or approval gets Orca's computed hash, approval first, inside a launch's 3 s wait, as main's did. When the lookup lands, the reconcile runs again and Codex's hash replaces it. * refactor(codex): one start for the hook lookup and the ~/.codex reconcile startCodexHooks replaces the two start functions and takes the PATH wait that the reconcile request's after option carried. A spawn reconciles unless one is running, without a scheduled flag, launches call one reconcileCodexHooksForLaunch on the shared withTimeout, and the reconcile alone reads the hooks setting. The startup test now runs the ready phase instead of matching its source text. * refactor(codex): plan ~/.codex with the shared Orca-hook pruner and approval reader The planner prunes with removeManagedCommands, as the opt-out does (so a hook that runs Orca's script through its args goes too), and returns a prune or settle union. While Codex has not answered, the kept approval comes from the shared per-event reader. The reconcile result is its outcome alone, a concurrent save spends a pass of the one bound, the opt-out removes Orca's approvals once, the approval-first writer takes the hooks path, and getRealHomeConfigTomlPath gives way to getSystemCodexConfigTomlPath. * fix(codex): ~/.codex keeps only an approval holding a hash Orca's entry may carry * refactor(codex): name the ~/.codex hooks-file check for what it reports * test(codex): the ~/.codex entry tests know Codex's earlier hashes, as the app's lookup does * refactor(codex): the ~/.codex reconcile uses the lookup's answer names * test(codex): ~/.codex status tests keep a Codex on PATH unless one is missing * refactor(codex): one stopgap for both homes; the ~/.codex pass finds its own home, hash and user data Also passes every approval to the approval-first writer, which already skips the ones in place. * refactor(codex): the reconcile start owns the warm-up; no rerun-on-answer flag The lookup start only records the hydrated PATH; one catch, one request type, and the app config is kept as given. * refactor(codex): the ~/.codex pass returns nothing; tests read the files Also makes the ~/.codex opt-out take Codex's hashes, as every production caller passes them. * refactor(codex): the ~/.codex approval cleanup takes Codex's hashes * test(startup): check the Codex hook start waits for PATH by behavior, not identity * fix(codex): a status-hook problem never moves the system default off ~/.codex * chore(ci): run the real-Codex contract when the ~/.codex reconcile changes * docs(codex): the ~/.codex reconcile's refused branch covers every definitive refusal * test(codex): start the failure-memo lookup with the reconcile's PATH-only start * fix(codex): ~/.codex gets nothing while Codex is missing, a pending answer uses the saved one, a conversion waits for a run that writes, and status and the reconcile read one home and one Orca-hash rule * test(codex): ~/.codex and managed-home status tests name their home; ~/.codex rewrites the backslash key Codex on Windows reads * test(codex): the Windows upgrade test reads status for the managed home it installs * test(codex): the real-TUI contract trusts its workdir by its real path, and its userData exists * fix(codex): keep Orca's approval at every slot its entry holds in ~/.codex |
||
|
|
ab6389045e |
fix(native-chat): start Windows chats without reading process creation times (#25718)
* fix(native-chat): start Windows chats without reading process creation times
Native chat on Windows refused to start ("Orca can't run this agent in a chat
here") whenever the process-table addon could not report process creation
times. Chat never needed them; only the bookkeeping around stopping the agent
did.
Windows now follows the common pattern: Stop ends the agent's tree with
`taskkill /T /F` on the child Orca still holds, and reports the tree gone only
when taskkill exits 0. A saved pid is never signalled after a restart.
- Remove the Claude and Codex location gate and the Codex launch refusal.
- A start time that cannot be read is recorded as unknown instead of refusing
the session; recovery already releases such an owner without signalling it.
- Delete Claude's Windows creation-time descendant snapshot and its verifier.
- Codex's Windows teardown reports the real taskkill outcome.
- The renderer no longer waits on the capability flag; the host keeps
publishing it for older clients (temporary).
* fix(native-chat): treat a Windows Claude exit after stdin end as a proven close
On Windows an idle Claude leaves on its own once its stdin ends, so every Stop
and close read as unproven: live background work settled as unknown and the
trace logged a close that "did not finish cleanly". As with the Codex close,
that exit is now the close and Orca makes no claim about processes Claude
started; a forced close still rests on taskkill's own report.
Also drop the Settings clause about Windows needing process start times, and
fix comments and the tracked process-enumeration doc that still described the
removed start-time gate and descendant snapshot.
* fix(native-chat): renew a held child's lease without a PID probe
An owner recorded without a process start time could never renew: the renewer
re-proves every live owner by PID identity, and with no start time and no
spawn-token echo that probe is indeterminate. The lease's last renewal then
stayed at the spawn, so a turn cut by an Orca crash was dated to its own start,
and every tick logged a failed renewal and split the batch into one store
transaction per chat.
The runtime now renews a lease for a child it still holds at the record's
fence and whose exit it has not received, recorded as a `held-child` match.
Receipt of the exit ends that proof before the exit is settled, and a restart
holds no child, so a dead owner's lease still expires and restart adjudication
is unchanged. Records this runtime does not hold keep the PID probe.
Also correct the identity probe's comment about shipped addons and note that
recovery's stop ladder is POSIX-only.
* fix(native-chat): keep a failed Windows taskkill unproven across a retried close
A Windows close counts Claude leaving on its own after its stdin ends as the
close. That shortcut also caught a retried close whose first attempt forced a
taskkill that failed: once the root exited, the retry returned true and the
failure read as a proven close. The tree reaper now records whether a reap
ever reached the live root, here, on an earlier close or from a transport
failure, and the shortcut applies only when none did; otherwise taskkill's
verdict stands and no new taskkill runs against the exited root.
Pin the platform on the existing tests that assume the POSIX close, and say
what `exit-proven` means on Windows in the acquisition-failure docs.
* fix(native-chat): derive a held child's liveness from the adapter's own handle
Lease renewal trusted a held child until the host settled its exit, and the
host hears of an exit late: Claude runs its close ladder and a store write
first, unexpected exits wait on one delivery chain shared by every chat, and a
Claude close that cannot prove its tree publishes nothing at all. A dead root
could keep renewing through that window, so a crash in it dated the cut turn
late. Renewal now asks the adapter, which owns the process handle and sees the
exit first: a child is held only while it is on record at the lease's fence
and its adapter still runs that exact acquisition with no root exit seen. An
adapter that cannot answer falls back to the PID probe. The stored
exit-received mark is gone.
The held-child read is now required by the runtime state and the renewer, and
a host-level test drives renewal through the real host wiring.
|
||
|
|
8d049b594d |
fix(codex): approve Orca's hook in managed Codex homes with Codex's own hash (#25742)
* feat(codex): ask Codex for its hash of Orca's hook in a throwaway home, cross-checked by position and path * feat(codex): cache Codex's hook hashes per binary and version, asked one at a time and only by the app * feat(codex): write a hook approval before its entry, and take back only its own on failure * feat(codex): approve Orca's hook in managed Codex homes with Codex's own hash, written first Managed homes (the shared mirror and per-account homes) no longer run a background approval session. Status reads the home's files against Codex's answer, and turning hooks off recognizes every saved version's hashes. * feat(codex): managed homes approve Orca's hook with Codex's hash; drop their background approval The previous commit carried only the managed resume's wait; this one holds the managed install it relies on. Managed homes (the shared mirror and per-account homes) write Codex's hash before the entry, fall back to their own approvals when the answer is late, and strip Orca's entry only when Codex itself answered with nothing to approve. Status reads the home's files against Codex's answer, and turning hooks off recognizes every saved version's hashes. * feat(codex): only an Orca-launched Codex waits up to 3 s for the hook hash; warm it after PATH hydration * feat(cli): name the file each agent hook status reports on * test(codex): real-Codex contract for the derived hook hash in managed homes, on both pins and latest * test(codex): type the hook-hash test fixtures and drop a duplicate import * fix(codex): give a Codex launch its own install run instead of joining a plain terminal's * test(codex): a user hook's approval stays put in an event Codex does not list * test(codex): cover late answers, first-install mirroring, stale approvals and opt-out re-asking * test(codex): a long managed home gets the daemon guard on its first install * fix(codex): until Codex answers, approve a managed home's hook with Orca's own hash, as main did A late, temporary or missing answer with no earlier approval in the home now writes main's self-computed approval instead of leaving the hook out. Codex's answer replaces it at the next install, a definitive answer (no hooks/list, a refused cross-check, 0.128) never uses it, and status says the approval is Orca's until Codex confirms it. * test(codex): status flags an unapproved entry while Codex has not answered * fix(codex): managed stopgap fills each missing event Until Codex answers, a managed home kept only the events it had already approved and dropped Orca's entry from the rest. Each event now keeps the home's approval, else gets Orca's own hash, as main wrote every event. One reader of the approval at Orca's entry serves the stopgap and status. * refactor(codex): one Codex answer type, one in-process answer map, a disk-only memo - One answer type with a kind (hashes, refused, pending) replaces two types and the three-field decoding at each caller. - The lookup keeps one in-process answer per binary path, replacing the process memo, the global latest answer and the transient-failure map; status now reads the answer for the codex on PATH, not the last one asked. - The memo file keeps Codex's refusals per version, like its hashes. - Derivation takes the version it is given; one 30 s version-probe timeout. - The launch wait reuses withTimeout, and launch prep passes launchesCodex down instead of a wait in milliseconds. - Turning hooks off no longer forgets Codex's answer. - Tests mock the derivation instead of a test-only resolver in production. * chore(codex): list the approval reader for the CLI build; fold two identical scope checks * fix(codex): count an approval at Orca's key only when it holds a hash Orca's entry may carry * refactor(codex): one append for hook trust tables * refactor(codex): the lookup keeps no entry for a missing Codex, and status checks the binary's fingerprint Also names the lookup functions for the answer they return. * refactor(codex): one stopgap reader for the managed home; the refused branch reads its own status * refactor(codex): drop defaults and exports only tests relied on * fix(codex): a failed ask of Codex stays pending instead of refusing its version * chore(ci): run the real-Codex contract when the approval reader changes * fix(codex): only a scratch home Codex loaded can refuse; the memo takes any hash and writes only on change * fix(codex): hooks turned off during a launch's wait win, Off re-keys mirrored user approvals, and one rule says which hashes are Orca's * fix(codex): an approval counts only under every key spelling Orca writes, as Codex on Windows reads only the backslash one |
||
|
|
468e4e1167 |
fix(native-chat): a prompt card owns the chat input until its answer lands (terminal-backed chat, desktop and phone) (#25761)
* fix(native-chat): an answerable prompt card owns the chat input until its answer lands * fix(mobile): a terminal chat's composer waits while its prompt card is up * test(native-chat): type the prompt card fixtures without casts * fix(native-chat): scope replies to acknowledged prompt occurrences * fix(native-chat): preserve answer ordering and verified delivery * test(native-chat): keep mock RPC client inside test boundary * test(native-chat): place mock fixtures in the test-only scope * Keep runtime comments within the module size limit * test: preserve prompt delivery coverage in desktop CI * Treat an older host's accepted write as delivered A newer desktop or phone talking to a host that predates the write settlement field read every accepted reply as "unconfirmed". Prompt cards never dismissed, the phone showed "Response unconfirmed" on every tap and ordinary chat messages were held as "Delivery unconfirmed". The reader now uses writeSettlement when present and otherwise keeps the host's whole-write accepted/refused verdict, exactly as before this branch. Only prompt answers ask for provider settlement; ordinary callers (follow-up delivery, paste drafts, option commands, composer sends) are back on the original contract, so the legacy-handoff error class, its flag, the sequence-only send helper and the mobile handoff hook are gone. * Keep terminal-pane Escape on the plain accepted write Every pane's Escape/Ctrl+C goes through pty:writeAccepted. This branch had switched that IPC to wait for provider settlement, which dropped the "remount this pane" signal for a daemon session awaiting recovery and could stall later keystrokes behind a slow daemon acknowledgment. pty:writeAccepted is back to its original local-only, synchronous write. Prompt answers opt into settlement with requireWriteSettlement on the same channel, and a settled refusal while the daemon recovers now sends the same remount signal. Ordinary verified sends regain their original fallback write. * Report a partly accepted local paste as unconfirmed A settled local write split into chunks returned plain false when a later chunk was refused after earlier ones were accepted. Callers read false as "nothing was written", so chat showed "Message not sent" with a prefix already in the agent's input. It now reports the write as unconfirmed, the same verdict the paired host gives for a partial write. * Hide the chat composer under a prompt card instead of unmounting it When an approval or question card took the input region, the composer unmounted. A message still waiting for its Enter was cancelled and its bubble deleted after the draft had already been cleared, so the message vanished without a notice; composer history was also wiped each time. The composer now stays mounted but hidden while a card owns input, so its state survives. A send that has not submitted yet is still stopped (its Enter would answer the card), but its bubble stays with "Message not sent" so the text is not lost. The composer ref is detached while hidden, so root typing, paste and reveal focus never reach it. * Keep an answered prompt hidden after the chat view remounts The "answered" dismissal lived in component state. Toggling chat to terminal and back, a PTY reconnect, or leaving the phone session and coming back while the approved tool was still running brought the answered approval back, and it then took over the input again. Desktop now keeps the answered occurrence per pane outside the view; phone keeps it per chat tab outside the controller. Both still retire it when the pane observes the prompt clear or change, desktop also when the tab retires, and both maps are size-bounded. * Update the prompt-reply reliability gate for the review fixes Older hosts' accepted answers now dismiss like acknowledged ones, the composer stays mounted under a card, and answered prompts survive a view remount. The gate's invariant, oracle, assertion list, new test files and the two latest evidence runs now describe that contract. * Let users hide a prompt card, keep Escape from denying, and gate only Send on the phone The chat input could stay locked behind a card the host never closes (for example after a Deny typed in the terminal), and Escape on a focused approval card denied the tool even when the user meant to close a picker. - A Hide control (chevron) on terminal approval and question cards, desktop and phone, hides that prompt occurrence and gives the input back. It writes nothing to the agent and uses the same per-occurrence dismissal as an acknowledged answer, so a new occurrence shows the card again. - On desktop, Escape on a card now does the same Hide instead of Deny, and a card that appears while the user is typing no longer takes focus. - On the phone, a card blocks only Send: typing, dictation and image attach keep working on the draft. The placeholder is back to the normal one. * Fix two comments that still called older-host replies unconfirmed Since an older host's accepted write now counts as delivered, the requireWriteSettlement comment and the reliability gate's oracle said the opposite of the code. Both now describe the current rule. * Collapse prompt cards to a strip instead of hiding them, and close the round-2 gaps Hide removed a card completely, so nothing on screen said a prompt was still waiting, and several edges let the chat type into a live prompt. - Collapse (the header chevron, or Escape on desktop) folds the card to a one-line strip above the composer; the strip's chevron expands it back. Collapsing writes nothing, frees the composer, and is disabled while an answer is still being written. Each pane or tab keeps the occurrence as answered or collapsed, so a remount restores the same view. - Questions now carry the host wait's start like approvals, so an identical question in a new wait shows again (desktop and phone). A transcript-only prompt, which has no wait start, is dropped when the view stops observing it, and a transcript still loading no longer clears a dismissal. - Desktop: while a card owns the input, the hidden composer cannot send or interrupt even if it still has keyboard focus, and the card takes focus in the same commit. A send the card retires no longer types Ctrl+U under it. - Phone: an Ask hides the heuristic card read from the same waiting status, and the dismissal store is scoped by host, worktree and tab. * Keep a collapsed card's partial answer, and scope its focus to its own pane Collapsing a question card unmounted it, so expanding it again lost the chosen step, selections and typed "Other" text; Escape typed in that text field collapsed the card. A card arriving while the user typed in another surface (sidebar, notes, a browser URL bar) also took the keyboard. - The collapsed card now stays mounted but hidden (and inert on desktop) under its strip, on desktop and phone, so a partial answer survives collapse and expand. Escape inside the card's text field no longer collapses it. The question card shows the same focus ring as the approval card. - A card takes focus only from inside its own pane (its hidden composer) or from the page body, never from a text field elsewhere. - Desktop and phone share one dismissal store in src/shared, bounded by the existing scope-cache helper, which moves to src/shared with it. - The card send imports the verified helper from its own module, and the phone files are split so each name matches its contents (header action, strip, lane selector). * Return focus to the composer after a prompt card collapses Since a collapsed card stays mounted, Escape or the chevron left keyboard focus inside the now hidden, inert card. The composer's reveal-focus took that as focus already in the pane and stood down, then the browser dropped focus to the page body, so typed keys went nowhere. Reveal-focus now treats focus inside a hidden or inert subtree as not in the pane and focuses the composer. On the phone, collapsing a card also dismisses the keyboard so a hidden reply field does not keep it. * Keep the question card's collapse chevron beside its Cancel button The question card header spread its three items with justify-between, which put the new chevron in the middle of the header. The title now takes the free space, as in the approval card, so the chevron sits next to Cancel at the right edge. * Run the prompt tests on the merged main Main now runs Vitest under Bun, which resolves a long data: URL import as a package name, so the SSH delivery test loads its bundled mobile module from a temp file instead. The phone prompt harnesses mock the live line that main's view now renders, and add Platform, which main's text-selection helper reads, the same way main's own view tests do. |
||
|
|
3ec38b8c6d |
Run Vitest on Bun with Node runtime contracts (#25840)
* Run Vitest on Bun while preserving Node runtime contracts * Preserve runtime timing provenance and keep the Bun pin in config * Scope builtin compatibility mocks to test-only lint exceptions * Give capture retention fixtures distinct filesystem timestamps * Await the copy button success state in the React fixture * Bound Node test worker shutdown and tighten migration fixtures |
||
|
|
13ea35973c |
Stop expensive checks when an unmerged PR closes (#25829)
* Cancel active checks when an unmerged PR closes * Register owned-branch cancellation qualification * Keep temporary cancellation qualification outside the review diff |
||
|
|
f6f96db6be | Build SSH hostile-host Linux slots independently (#25821) | ||
|
|
eaaae0196f |
feat(native-chat): record fresh sessions after failed restoration (#25747)
* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore A chat whose saved conversation the agent cannot reopen can now continue in a fresh one: the handle chain records the new conversation as a creation that replaces the lost one (which, why, and when), keeping every earlier link. Rows keep a shape older builds read: the stored chain starts at the latest replacement and carries the earlier links inside it. * refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row Older builds only read Claude and Codex records, so the nested stored form protected rows no replacement can reach while adding a cap mismatch after a downgrade. Store the chain as held, refuse a replacement in a Claude or Codex chain until one has a stored shape older builds read, and refuse a supersession key on a replacement that names no creation in the chain. * test(native-chat): prove replacement rows survive downgrade and re-upgrade |
||
|
|
f9c8cd4fc3 | Reduce redundant test coverage and unnecessary CI waits (#25806) | ||
|
|
020cebeff6 |
Add standalone Agent Client Protocol client layer (#24990)
* Add standalone ACP protocol client and session runtime * Protect ACP transport teardown from late stream errors * Retire incoming ACP request ids before publishing responses * Narrow ACP configuration requests and transport message types * Remove redundant ACP request handler return unions * Keep ACP waits caller-owned and preserve protocol extensions * Preserve open ACP decisions through prompt completion * Generate open ACP enums and check the generated schema offline A newer or vendor enum value (tool kind, tool status, option kind, stop reason) no longer fails the whole message: generated enums accept the known literals plus any other string, typed so callers can still narrow on the known ones. The generated header now records the pinned input digests, the generator digest and a body hash, so `verify:acp-protocol` catches a stale or hand-edited file without network access; it runs in lint and the PR workflow. * Land the ACP runtime contract the agent adapters use - Deliver notifications other than session/update through onExtensionNotification, in arrival order with session updates. - Accept _meta on prompt, setMode, setModel, setConfigOption and cancel. - cancel() always sends session/cancel once the session runs, since the agent can be in a turn it began itself; only a successful send is shared, so a failed write is retried. - Cancel aborts each open agent request's signal and lets its handler send its own answer; -32800 only when the handler rejects. - Permission requests validate only the session, tool call id and options; unreadable fields are dropped with a diagnostic, and any answer Orca cannot send is `cancelled` instead of a JSON-RPC error. Agent-started turns may ask; whether to show it is the caller's decision. - AcpAgentError marks the agent's own errors; AcpInvalidResponseError keeps the raw answer and validation issues for answers Orca could not read. - Lines over the size limit are classified by prefix (shared with the Codex reader): the owed request fails, an oversized agent request is answered with an error, and an unattributable response closes the connection. * Answer every agent request after an ACP cancel A cancel that lands before a permission handler starts now still runs the permission path, so the agent gets the `cancelled` outcome rather than a request-cancelled error. A handler that ignores the abort no longer leaves the agent waiting: once the abort has run through, any request still unanswered gets request-cancelled. Handlers that answer on abort keep their own reply. Also renames a lint-rejected helper parameter, replaces a Reflect.apply in a test, and stops the permission diagnostic from firing with an empty list. * Let each ACP request handler own its answer after a cancel Removes the next-event-loop-turn fallback that answered request-cancelled for any handler still silent after a cancel. It raced answers that were still being saved (an approval mid-journal-write reached the agent as an error) and made the outcome depend on event-loop timing. The handler that owns an agent request now always sends its answer, or throws for request-cancelled; a request it never answers ends when the connection closes. A permission whose handler had not started still answers `cancelled`. * Register the ACP schema verify step in the PR preflight phase test * feat(acp): a steer's cancel asks once and never ends the agent The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle, then close the connection, which ends the agent. A steer used it too, so a slow agent lost its process just because the person added a message. requestSteerCancel() now sends session/cancel once per prompt, cancels the agent's open requests and answers later permissions cancelled, and never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel() stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel paths move into acp-prompt-cancel.ts over one cancel channel. * fix(acp): a repeated steer shares the cancel in flight; say what the caller owns Per review: a second steer before the first write lands returns that write instead of resolving early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that fails instead must not take the steer until the caller rebuilds the session; the Stop's says a prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives the runtime a handler that would allow: the open permission's signal aborts and the late one never reaches it. * test(ratchet): require src/main/acp now that this PR lands it |
||
|
|
ca4e239861 | Remove low-value test inventories and duplicate fuzz oracles (#25791) | ||
|
|
6c693edf40 |
Show native chat tool calls as plain sentences (#25654)
* feat(native-chat): add chat-scoped color tokens * feat(native-chat): soften transcript and composer appearance * fix(native-chat): refine code spacing and faint text styling * fix(native-chat): wrap prose links at word boundaries * test(native-chat): refresh background task strip snapshots * feat(native-chat): show tool calls as plain sentences * fix(native-chat): make tool sentences reflect call state * fix(native-chat): clarify failed commands and subagent sentences * test(native-chat): exercise command disclosure with real results * fix(native-chat): preserve command inputs and localize failure rows * fix(native-chat): retain complete padded command input * fix(native-chat): read named tool input fields safely * Refresh native chat tool rows when the UI language changes |
||
|
|
059e81a106 |
chore(i18n): use 智能体 for Chinese Agent copy (#25767)
Simplified Chinese rendered the Agent concept as 代理, which collides with 代理 = proxy. Standardize on 智能体 for Agent (and 子智能体 for subagent), while keeping 代理 for proxy senses: HTTP/network proxy, SSH Proxy Command, reverse proxy, and browser user agent. - Converted 131 zh catalog values (incl. 子代理 -> 子智能体); 26 proxy / user-agent values left as 代理. - locale-phrase-fixes.mjs: 客服人员/代理商/座席 -> 智能体, 代理 -> 智能体 when the English names an agent (guard excludes "user agent"); removed the old 智能体 -> 代理 rule so the pipeline no longer reverts it. - Updated value/key/search/macos-tcc overrides to 智能体; proxy keyword and proxy override entries unchanged. - Updated the two policy tests that pinned the old 代理 output. Gates: catalog verify, coverage --check, extraction, runtime-catalog, and the locale vitest suites (296 tests) pass. |
||
|
|
d73efccc7d | Reduce localization audit and relay setup work in CI (#25665) | ||
|
|
b01df814e6 | test(ratchet): require src/main/provider-process now that it has landed (#25710) | ||
|
|
1168e0f8c8 |
fix(ssh): let a placed worktree seed while its host is in conflict (#23213)
One host tab on a folder this client lacks marked every worktree on the host unverifiable, so reopening an emptied one never got a terminal. Fixes #22015 Co-authored-by: mmarabel <166927047+mmarabel@users.noreply.github.com> |
||
|
|
d791421568 |
Extract provider process supervision and stream reading from Codex (#24989)
* Move provider process supervision and stream reading out of Codex * Preserve teardown behavior with checked mock types after move * Apply provider launch environment and caller teardown labels * Give provider child env one owner and gate Codex contract on the shared reader resolveProviderChildEnv is now the only place that overlays and strips a provider's environment; the spawn spec and the request-scoped Codex session both call it. supervisedPosixLaunch only accepts a launch without env fields, so an override can no longer be silently ignored there. Edits to the shared stream reader or the env rule now run the real-binary Codex contract job. |
||
|
|
7208be6902 |
test(native-chat): cover paired runtime launch compatibility (#25072)
* test(native-chat): cover paired runtime launch compatibility * test(native-chat): use production prompt response fingerprints * test(native-chat): pass journal items to the paired-runtime outbox hook Main made journalItems a required outbox input; mirror production by passing the read state's items. * test(native-chat): pin paired-runtime routing and capability questions - Move the paired-runtime chat test into its own file and render the real session hook against a paired server that advertises today's capabilities while this machine advertises none, so Stop, prompt cancel, repeated Stop and question answers fail if their capability question goes to the wrong runtime. Add a local and a paired session sharing one id. - Cross-version: state that the file pins the capability strings a v1.4.219 server really advertises, use the desktop's real handshake list, and keep the current server only as the control for the refusal. - agent-launch-routing: give the released-server case a capable client so it is refused by the server gate, not the client one. * test(native-chat): route the paired launch suite and cover the outline read - Start the cross-version job when the launch route or the client list the desktop sends a paired host changes, since the paired launch suite runs them. - Check the rail's conversation outline is read from the paired server, and narrow the hook test's header to what it covers. |
||
|
|
508419f11e |
test: check structured-chat code for Electron imports even before the runtime loads it (#24988)
* test: keep structured chat free of Electron imports
* test: widen runtime Electron ratchet to structured chat
* test: check whole structured-chat directories for Electron imports
Cover src/main/{native-chat,claude,codex}, src/shared and every structured-*
or agent-session-* file under src/main/runtime, so new files are checked by
default. A lane that exists today now fails loudly if it goes missing; only
the not-yet-landed acp/ and provider-process/ may be absent.
Move the test-only file rule into one classifier shared with the
localization audit, so -test-<thing> helpers and test doubles no longer
enter the production gate.
* test: run the Electron-import CLI path in tests and retire may-be-absent lanes
Export main() so the real-tree tests exercise the entry list CI uses,
check every required lane for the missing-directory error, and fail once
acp/ or provider-process/ exists while still allowed to be absent.
|
||
|
|
e817b0e237 |
refactor(native-chat): keep the provider resume handle opaque to shared code (#24991)
* refactor(native-chat): keep the provider resume handle opaque to shared code
Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).
Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.
The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).
No user-visible change.
* fix(native-chat): derive journal-row provider handles from the journal identity
The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.
* fix(native-chat): refuse a stored provider handle written in both forms
A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.
* test(native-chat): use opaque handle in queued rejection fixture
* test(native-chat): share one Codex journal identity in the integration suite
Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.
* refactor(agent-session): name the handle's adapter state resumeCursor
Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.
State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.
|
||
|
|
761d63a4e5 |
feat(agent-launch): keep long prompts off the launch line and paste them after readiness (step 1 of 7) (#24257)
* feat(agent-launch): host-side prompt delivery for agent.launch The host's agent.launch typed any launch prompt into the shell as part of the launch command. A long or multi-line prompt then ran line by line in the shell, and an agent that never showed readiness or crashed at startup had nothing guarding where its text went. agent.launch now carries a prompt on the typed line only when the line stays one line, control-free and at most 512 bytes; otherwise the agent starts clean and the host pastes the prompt once the agent's own ready signal fires (bracketed paste plus its composer marker or a quiet render, read only after the shell's last hand-off, never while the pane's own shell is proven in front), with main's draft-paste bytes and an Enter 50 ms later. Orchestration worker starts wait on tui-idle as before. A replay-safe launch admits and claims its ledger row in one write, Qwen Code gets a second Enter, the desktop and phone share one launch-refusal classifier, and hosts advertise agent.launch.prompt-carry.v1. Split out of #23748, which moves the desktop source-control buttons onto this path. * fix(agent-launch): keep a short-lined multi-line prompt on a local zsh launch line, as main did #24257 moved every multi-line or over-512-byte prompt off the typed launch line and pasted it after readiness. The phone's AI buttons and review notes, whose multi-line prompts main typed whole into zsh, then reached Claude 0.5-3 s later and their RPC reply waited for the paste. The host now names the shell a local macOS or Linux line is typed into, the way the spawn picks it, and a multi-line prompt rides a zsh line when every line is at most 512 bytes and the whole line at most 8 KB. A real-zsh test types such a line through Orca's own ready barrier and startup write, including when a slow user config makes the write land early. Elsewhere the measured unsafe cases keep the paste: bash 3.2 runs multi-line lines piecemeal, fish drops an early multi-line write, and any shell loses a line over 1 KB written early. * test(agent-launch): keep the real-zsh launch-line test out of the Windows lane's gate scan The Windows lane registration check read `const ZSH_PATH = process.platform === 'win32'` (the head of a multi-line ternary) as a Windows-true flag, so `describe.skipIf(!ZSH_PATH)` looked like a Windows-only suite. The file is POSIX-only; the zsh lookup is now a function. * refactor(protocol): move the agent.launch capabilities into their own module Main's protocol-version.ts sits at the 300-line cap, so the prompt-carry capability pushed it over. The four agent.launch capabilities and their doc move to agent-launch-runtime-capability.ts, re-exported by name and spread into RUNTIME_CAPABILITIES at the same position; the advertised lists and every export are unchanged. * refactor(protocol): import the agent.launch capabilities from their own module `export *` from protocol-version left the four names undefined under the mobile recording loader, which resolves a relative import through a Proxy with no own keys, so 37 phone recordings lost agent.launchReplay. Importers now name agent-launch-runtime-capability directly; protocol-version only spreads its list. * refactor(agent-launch): drop the unshipped viewMode field and trusted local caller id Both were inert in step 1 and existed only for step 2. agent.launch will become a public plugin API, so every wire field is permanent once shipped; a top-level viewMode reads as "choose terminal vs chat", which the host decides. Step 2 introduces placement and view intent under a placement object instead. * fix(agent-launch): read Codex's provisional startup from the rule files' hold anchor Main (#24375) moved Codex's provisional-header check into codex.json's provisional_startup hold anchor and deleted codex-terminal-readiness.ts, so the launch readiness hold now asks showsHoldAnchor, as main's own settled check does. * fix(agent-launch): hold rule-file name titles to quiet for a launch, and census the zsh fixture Main (#24375) answers a name-only title from each agent's rule file ahead of the sustained-title lane, so gemini.json's name_title settled a launch readiness wait on the shell's auto-title while Gemini was still booting. A launch now asks quiet of every weak idle verdict, as that lane did. Main's readiness census requires a recorder for every runtime fixture; the zsh prompt recording is a non-agent control. Gemini's synthetic baseline is regenerated for this PR's stated change: a bare gemini title is no longer its rest mark, so name-only rows settle weak, and a fresh working or blocked status is no longer overridden. * fix(agent-launch): paste a launch prompt only when the launched agent is proven in front A launch pasted its prompt unless a shell was proven in the terminal's foreground, so any read that could not prove one let the prompt through. After an agent exited at startup, its shell turned bracketed paste on at the next prompt, readiness fired on it, and the prompt was typed into the shell: - macOS: a pane runs its shell under login, so the process-group fence's root was never the shell's group and never proved it; the cached foreground name could also still name the exited process. - Windows Git Bash and WSL: the shell-alone-in-its-job check never answers. Now one fresh read of the terminal's foreground decides: agent, shell or unknown. Only 'agent' lets a write through (paste, Enter, second Enter, reused panes too); 'shell' still drops a ready signal. A Windows host never proves the agent, so there the launch line carries the prompt at any size, as on main. * test(agent-launch): cover the Windows QA stub, a grok override that exits at once * fix(agent-launch): keep the local socket alive while a prompted launch waits for its agent A launch with a prompt now waits up to 60 s for the terminal agent to be ready before it writes the prompt, and reports not-delivered when the agent never is. The local runtime socket closes a connection idle for 30 s unless the request is a long poll, so a launch whose agent exited at startup lost its reply and the caller saw 'runtime closed the connection' instead of not-delivered. Classify a prompted agent.launch and agent.launchReplay as a long poll, as orchestration.workerStart already is for the same wait. * refactor(agent-launch): narrow the launch params by 'in' instead of a cast * fix(agent-launch): find a launched agent behind a wrapper that leads its process group A tcsh or nu launch line runs the agent from /bin/sh '<script>', and a wrapper script that does not exec its agent does the same: the wrapper leads the terminal's foreground process group and the agent is a member of it. The fresh foreground read names the group's leader, sh, so a prompted launch was refused or pasted late (M4Air tcsh: 2 of 4 not delivered, 2 pasted ~9 s late). Before that read, take the host's process-group observation as positive proof when it names the launched agent among the foreground group's members and is younger than a ready signal's quiet window. It never proves a shell. * fix(agent-launch): judge the foreground-group proof by when its capture began, not how long ps took The age the host stamps on a process-group observation runs from the start of its whole-machine ps, so on a loaded Mac a capture begun after the read was asked for still read as older than 1 s and the proof was dropped. Count an observation whose capture began after the read was asked for, less the window a shared capture is reused across. * test(agent-launch): keep the crash-guard live test out of the Windows lane's gate scan The Windows-lane registration scan read the const assigned from a platform check as a Windows-only gate, though the suite runs everywhere but Windows; find zsh in a function instead, as the real-zsh typed-line test does. Under load the fresh foreground scan can fail to answer, which lets the shell's prompt settle readiness (2 of 4 paired runs). The guard still refuses that write, so assert the refused write, the property that must always hold. * perf(agent-launch): read a local pane's foreground from its own terminal, not the whole process table The foreground read that gates every launch paste ran the daemon's inspectProcess capture and then a fresh scan, each a whole-machine ps; the fresh one also waits for any capture already running before it starts its own. Measured here at load 5: 1.2 s a read (M4Air QA: 3.4-5.0 s, and worker starts 17.6-32 s against main's 9-12 s at load 25-84). On a local macOS or Linux host, take the pane's root pid from the provider's session inventory and run one ps limited to that pane's terminal. Its foreground process group decides: the launched agent or any non-shell member is the agent (a wrapper that did not exec its agent leads the group), a group of shells alone is the shell. Same pane, same verdict: 2.7 ms a read. SSH hosts keep the relay's observation and name. * test(mobile): re-measure the web app's script sweep after agent.launch's capabilities moved out of protocol-version The mobile web bundle check failed at 124 assets against a ceiling of 123. Main already sat exactly on that ceiling: its sweep table read 69 scripts at 16 routes while the tree builds 73, the whole margin of 4. This branch imports the agent.launch capabilities from their own module, so protocol-version is no longer pulled into the root layout and four other routes. That moves which routes share which modules, and the Qoder capability module, imported by protocol-version and the AI-vault resume path, no longer shares an importer set with anything, so it gets a chunk of its own: 74 scripts. The fence says to re-derive the bound rather than raise it, so the sweep is re-measured on this head (every prefix of the sorted route list). The worst route now adds 10 scripts (session), not 9, which moves the pinned shell crossing from 32 to 30 routes; main re-measured on its own lands on the same crossing. * fix(agent-launch): a worker's brief needs its agent found in front, and Grok's start answers on its composer A paired-server worker start whose agent exited at startup typed its brief into the server's shell, which ran it: the idle wait can settle on a shell back at its prompt, and the brief was written with no foreground read. Both worker-start paths now check before each brief write, as a launch prompt is checked: on a host that can find the agent in front it must be there; on one that cannot (Windows) a shell proven in front still refuses, and anything else writes as before. A Grok worker start waited ~10 s more than main: its only rest signal is its bare name, which a launch holds to quiet output, and Grok animates its logo for ten seconds after its composer glyph. A worker start for an agent whose rest signal is its bare name and whose composer draws a marker (Grok, DSH, mimo-code) now also answers on that marker, whichever comes first. |
||
|
|
a68ee67cf9 |
fix(codex): stop prompting a Codex restart when only its home changed (#25421)
* fix(codex): stop prompting a Codex restart when only its home changed A Codex terminal that outlived the move to ~/.codex was blocked behind "This Codex session is using an outdated configuration" until the user restarted it, which starts a fresh codex and ends the conversation. That terminal keeps working on Orca's old Codex home, which is refreshed from ~/.codex, so the change did not need the user's answer; the prompt also never appeared for Codex in Git Bash terminals, whose process tree Orca cannot see through. The restart prompt now covers only an account switch, which the user caused. The pane record keeps its home route, which the retained-home refresh and the shared-server check still read; only the prompt and the custom-home downgrade that existed to keep it from guessing are gone. * refactor(codex): drop what the home-route prompt left behind - Record a pane's own CODEX_HOME override as-is. The filter that kept only overrides Orca could re-derive existed for the removed route comparison; its one remaining reader, the shared-server check, returns null for the shared-home route those panes record either way. - Treat custom-home like shared-home in resolveCodexPaneHome: the removed downgrade only wrote it without an override, so it never named a home. - Inline the restart notice key; its route/account prefix was its only job. - Delete the two pane-local override tests, including a POSIX-only one that still expected custom-home and would have failed on Linux and macOS CI. - Cover the account recheck branches the deleted route-recheck file was the last to exercise: main reporting a pane current, and main not answering. * refactor(codex): retire custom-home and the last re-check helpers - Read an older build's custom-home record as shared-home, which it always was, and drop custom-home from the route type and every check. - Delete shellStartupCodexHomeOverrideMatches and its comparison helper; nothing re-checks a recorded override any more. - Build the pane launch record in one return: the route is always set, and only a resumed launch differs, in how it picks the account. - Drop the restart dialog's notice key; its focus effect now depends on the pane and both account labels directly. - Give the account recheck test a typed window stub, so CI's type-assertion gate passes, and remove timing entries for deleted test files. * fix(codex): drop the removed dialog strings main added to es.json * refactor(codex): stop recording a pane's custom CODEX_HOME The pane record kept a pane's CODEX_HOME override so the removed route prompt could re-check it later. Its one other reader, the shared-server check, only looked at it for a real-home pane, and a custom CODEX_HOME always routes a pane to Orca's mirror, so it was never used. Drop both record fields, their validators, equality checks and spawn plumbing. getCustomCodexHomeOverrideForLaunch folds into the existing hasCustomCodexHomeOverrideForLaunch, and real-home resolves to ~/.codex. Records from older builds still parse; the parser keeps only known keys. The fish test's decoy no longer sets CODEX_HOME, so the boolean check still fails if the launch env XDG_CONFIG_HOME is ignored. * test(codex): drop setup the yes/no CODEX_HOME checks no longer need --------- Co-authored-by: Orca Worker <orca-worker@localhost> |
||
|
|
67fc708b3b | Group CSV editor modules and tests in their own folder (#25476) | ||
|
|
1bec53ceb2 |
Reduce repeated CI setup and overlap mobile typechecks (#25359)
* Measure remaining CI import, diagnostic and checkout savings * Qualify remaining CI candidates on hosted runners * Qualify independent mobile typecheck overlap on Actions * Keep explicit RPC test registries from loading unused methods * Qualify complete RPC registry cohort and mobile cancellation * Promote measured CI setup and typecheck savings * Recognize the shared RPC test guard in lint policy * Align the mobile barrier contract with independent typechecks |
||
|
|
822fc5bed4 |
Add repository OpenCode permission defaults (#25326)
* test(config): reproduce rejected repository OpenCode config * Add repository OpenCode permissions and allow its reviewed root config * fix: preserve sensitive OpenCode confirmation prompts --------- Co-authored-by: Orca campaign recovery <campaign-recovery@example.invalid> Co-authored-by: Orca OpenCode Campaign <opencode-campaign@local.invalid> |
||
|
|
cecb62158a |
fix(ui): restore IME Enter protection in workspace details (#24099)
Restore IME Enter protection in workspace details by reusing the existing composition tracker. Reset Notes ownership at textarea detachment and preserve sizing behavior. Repair isolated native test-window delivery without changing the production foreground policy or original native input assertions.
Fixes #24097
Related contributor history: #10711, #11067, #13128, #13282.
Original implementation and macOS recordings: @setodeve, commit
|
||
|
|
1978469fd2 |
fix(mobile): keep the working rings turning on the OTA page (#25299)
* fix(mobile): keep the working rings turning on the OTA page Animated.loop starts a native loop whenever the timing asks for the native driver; the web has none, so the JS fallback ran one turn and froze at 360deg. Ask for the native driver only off the web. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): pin the native driver on native spinners Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): share the working ring rotation between both rings Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
0c761a7610 | Admit short required auxiliary checks after PR preflight (#25317) | ||
|
|
baa56fd10d |
Simplify the phone-control and phone-size terminal dialogs (#25307)
* Redesign the phone-control and phone-size terminal dialogs Drop the eyebrow label and circled icon, shorten the copy so it no longer restates the buttons, and give each state one primary action with a quieter "all" action. Collapse moves out of the button row into a Minimize icon in the corner. Behavior is unchanged. * Point the phone settings copy at the renamed Restore button; drop dead ko overrides The phone app and desktop update independently, so name only "Restore", which matches both the old and new desktop banner labels. |
||
|
|
8e8efb1947 | Reduce avoidable work in PR checks and SSH test setup (#25309) | ||
|
|
b32462f246 |
Replace patched JSON parser with stream-json (#25202)
* Replace patched JSON parser with stream-json * Isolate dependencies for historical server compatibility builds |
||
|
|
d77c57022e |
Verify shared preflight selection and record full unit timings (#25239)
* Strengthen shared preflight contracts and record unit timing results * Record rejected shard-weight holdouts |