mirror of
https://github.com/stablyai/orca.git
synced 2026-10-03 08:02:12 +00:00
3e5c8d9f8e1dfe10490bd52cd111a9fa22bbfff9
12257
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3e5c8d9f8e |
feat(relay): declare Asia cell c31 at the c30 shape (#24310)
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
d2dfc79764 |
ci(daemon): runtime-launcher protocol ratchet and Node slot marker (#24108)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working tree must attach the newest release tag's daemon. Rollback crossing is reported only. Runs in the cross-version-wire job, which already has full tags; tag selection moves to config/scripts/stable-release-tags.mjs so both use one rule. * feat(persistence): run profile backups in the worker whenever its entry is bundled * refactor(orcad): make profile and native preflight runtime-neutral The profile preflight parser now takes the expected runtime identity from the caller (shipped callers pass the pinned Bun identity), and the native preflight is renamed to orcad-runtime-native-preflight with neutral wording. * feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS, NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball), generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no network, that the pin tracks the locked Electron, matches engines.node's major, and covers exactly SERVER_TARGETS; it runs in the static analysis job. ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target list; orcad's Bun runtime and build output are unchanged. * test(persistence): skip plain-Node backup selection tests in the Bun profile suite * fix(runtime): reject a pinned archive that belongs to another target * ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change, so one PR must not do both. The launcher file list lives in the check script; the allow-runtime-launcher-protocol-bump label overrides it. * feat(orcad): select pinned-Node slots by a .runtime-node marker D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through .runtime-node instead of .build-target, so Bun-era clients read it as a legacy slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet. * fix(runtime): load the Node pin without the typeless-module warning check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from its own module, so it no longer loads the update script's build graph. * fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
d4ae6a904b | Update README downloads badge | ||
|
|
4bd55c07ea |
fix(claude): a Stop ends Claude's process, and the next message resumes the conversation (#24235)
* fix(claude): a Stop ends Claude's process, and the next message resumes the conversation
Claude's Stop now ends its child after the interrupt, whatever Claude answered, once the stopped turn ends or a 3 s grace runs out. The grace lets Claude's own result move the resume point past that turn, so the next send resumes after it. Background work ends with the child. Codex keeps its child.
* test(native-chat): the router answers stopEndsSession for the session's live owner
* test(native-chat): type the Stop test's close mock as the adapter's optional method
* fix(native-chat): a Claude Stop answers on its interrupt, and the next step on the chat's lane ends the child
The Stop now answers as soon as the interrupt step finishes. Ending Claude's
process runs as a second task on the session's serialized lane, queued in the
same tick as the Stop, so a message sent meanwhile reaches only the resumed
child. That step waits for the stopped turn to end (its result, the CLI's idle,
the child's exit or its close), then ends the child; a failure there is
reported and never fails the Stop, and the wind-down it leaves owed is retried.
The interrupt's answer and the turn's end share one 3 s grace counted from when
the interrupt goes out, so an interrupt Claude never answers ends the child in
about 3 s instead of after the 30 s control deadline. The wait is derived per
call from Claude's open turn: the armed latch and the close's user-stop branch
are gone, and the close keeps its whole deadline for proving the exit.
A Stop naming a turn that has since ended (the phone names the turn it last
saw) now ends the child too when Claude took it, as it does when Claude
interrupts a handed-over follow-up whose turn has not opened.
* test(native-chat): pin a Claude Stop's race, failure and timing cases on the shipping adapter
- Nothing queued while the Stop's first step waits on the interrupt runs before
the child ends.
- A message sent with the pre-Stop fence during the Stop is accepted with no
notice, and reaches only the resumed child, after the old child's close.
- An interrupt Claude never answers ends the child within the grace.
- A close that cannot prove the exit leaves the Stop answered with its row; the
failure goes to the host's error hook.
- A Stop naming the turn that just ended ends the child when Claude interrupts
the handed-over follow-up.
- A second Stop pressed while the first ends the child stays quiet.
* fix(native-chat): a Stop naming an ended turn ends Claude's child when its interrupt fails or goes unanswered
The phone names the turn it last saw. When Claude interrupted a handed-over
follow-up for that Stop and the answer failed or never came, the child stayed.
Only a provider that answers it did not take the Stop keeps the child now.
Also pins a queue-if-active send made while the Stop ends the child.
* fix(native-chat): a Claude card's Cancel denies an approval and stops on a question, never a bare interrupt
A permission card's Cancel interrupted the turn and kept Claude running, the
old Stop on a second control: a refused interrupt let the turn run on and
background work survived. Now the host asks the provider how a card's Cancel
is answered. For Claude, an approval's Cancel sends the same reply as its Deny
option, so the turn goes on without that tool; a question's Cancel runs the
chat's Stop, which ends the child, and the next send resumes the conversation.
A provider that gives no answer keeps today's path, so Codex is unchanged.
The host decides, so desktops and phones of any version keep sending the same
cancel and get the new behaviour; nothing on the wire changes.
With no Claude caller left, the prompt-cancel interrupt is gone: the bound
claim, the cancellation observation, the admission of the prompt's
cancellation, the prompt's turn binding, and withdrawing a refused stop (every
Claude interrupt now precedes the child's end, so the recorded stop stands).
* fix(native-chat): a Claude approval card's Stop option is the chat's Stop
The approval card's "Stop" option answered Claude with a deny that interrupts
the turn and kept Claude running, the last control on a Claude chat that
interrupted without ending the child. The host now routes it, like the card's
own Cancel, through the provider: for Claude that option is the chat's Stop.
It runs the same Stop body as the Stop button, in the same order (withdraw,
the Stop takes effect, interrupt), and the step that ends the child is queued
in the same synchronous call. The body now lives in one place,
structured-agent-session-chat-stop.ts, shared by the Stop button, a question
card's Cancel and this option.
The host decides, so an old card or client that sends option `cancel` gets the
real Stop; nothing on the wire changes. The interrupting deny reply is gone.
* test(native-chat): the option-route test's answer carries a whole journal resolution
* fix(native-chat): Claude cards drop their Stop option, and a plan card's Cancel asks Claude to wait
- Approval and plan cards no longer offer "Stop". Neither reference offers a
Stop while an approval is up: the user denies, then stops. An older card's or
client's `cancel` answer is a plain deny, and the respond path is main's again.
- A plan card's X / Esc no longer answers "Keep planning", which sent Claude
straight back to revise and re-propose. It dismisses the plan: a deny that
tells Claude to end its turn and wait for the user. No interrupt; the child
stays. A tool approval's Cancel stays the Deny reply, and a question's Cancel
stays the chat's Stop.
- The provider's routing takes the pending card itself
(`routePromptCancel({ sessionId, prompt })`), so a plan is told from a tool
approval, instead of a kind plus an optional option id that meant Cancel when
absent. A routed answer may be one the card does not offer, like the
dismissal.
- The chat's Stop is reachable only through `mutateWithChatStop`, which owns the
mutation call and queues the step that ends the child in the same call, so
no caller can run the Stop and forget the child's end.
* refactor(native-chat): cancelPlan no longer takes a stopChild nothing passes
The Stop's own body ends the child; the plan's default run serves only a card's
interrupt and the background-task stop, which never end it.
* fix(native-chat): a question card's Cancel settles the card in the Stop's first step, and answers Claude when nothing stops
A question card's Cancel runs the chat's Stop. The card stayed pending until
the child ended (up to about 3 s), so it could still be pressed and the second
press was refused once the chat rested. When the Stop found nothing in flight
(Claude asking after its main turn ended), the card stayed pending for good and
Claude's request went unanswered; the phone froze the card with every button
disabled.
Now the card is recorded as cancelled by the user in the Stop's own step, with
the receipt "Cancelled on <device>". A Stop that ends the session leaves
Claude's request to end with the child, so no reply races the interrupt; one
that ends nothing declines the request itself. The provider forgets the card
once the host records it, so neither Claude's own cancel nor the child's end
writes over the user's cancel. The card's Stop stops whatever the chat has in
flight, like the Stop button, instead of a turn the card names that may have
ended.
The same dismissal, a new provider member beside the cancel route, now answers
a plan card's Cancel: recorded as cancelled by the user, and Claude told to end
its turn and wait. That replaces resolving it as an option the card does not
offer.
* test(native-chat): the question-card Stop test has Claude cancel its held request, as it does once interrupted
* fix(native-chat): a Claude Stop before the echo reads Interrupted, and a question card's Stop stays in its own turn
- A Stop pressed after Send but before Claude echoed it read as a normal finish
("Worked for 0s", a green done check): with no turn row to wait on, the child
was ended at once, and the echo that would have opened the stopped turn
landed on a retired send. The wait before the child ends is now keyed on what
Claude has in flight (an open turn or a send it has not answered), not on a
journal turn id, within the same 3 s from the interrupt and woken by the same
settles; a settle that leaves something in flight waits on. The echo opens
the turn, the aborted result ends it Interrupted. A Claude that says nothing
is ended when the grace runs out, the unanswered send doubt as before.
- A question card's Cancel takes the chat's Stop only when the card belongs to
the turn running now, judged after draining the provider's lifecycle. A card
a finished turn raised, such as a background agent's, is dismissed instead:
Claude's request is declined and nothing stops.
- Claude forgets a dismissed card before the host records it, so Claude's own
cancel landing during that write cannot replace the user's "Cancelled on
<device>".
* test(native-chat): Stop tests read background work from the host's child records, and guard the optional dismissal
Main's child records replaced the adapter's background-task callback: the
background-work test now reads the host's child record (live before the Stop,
settled after the child ends) and the session's agent status (done, Interrupted,
nothing left for Monitoring).
* fix(native-chat): a tool card's Cancel reads Cancelled, and a Stop's background work reads stopped
- A tool approval card's X / Esc already sent Claude the plain deny, but the
card read "Deny · Answered on <device>", the same as pressing Deny. It is
now a dismissal like a plan card's: recorded as cancelled by the user, with
the same deny reply, no interrupt and the child kept. Every card's Cancel
now routes to a dismissal or the chat's Stop, so the option route is gone.
- When a Stop ended Claude's child, a background agent or shell it was running
settled as an unknown ending (a neutral dot), while the same task's own stop
reads Interrupted. A close Orca asks for (a Stop, a rest, a quit) now ends
what still runs as stopped; only a session that dies on its own leaves the
ending unknown. The close reads no stop cause. A rest never meets live child
work: the idle sweep keeps a chat with any.
- If the host fails to record a dismissed card, Claude gets the card back, and
a withdrawal Claude made meanwhile closes it, as before the host took it.
* refactor(native-chat): the session mutation path gets its own module, breaking the chat-stop import cycle
chat-stop.ts imported mutateStructuredAgentSession from host-mutations.ts,
which imports mutateWithChatStop back. The mutation context and the one
admit-then-serialize path now live in structured-agent-session-mutation-context.ts,
which both import; host-mutations re-exports the context type for its other
readers.
* fix(claude): a close that saw a descendant survive leaves background work's ending unknown
A Stop, rest or quit ended still-live background tasks as stopped even when the
close found Claude's root gone but a descendant still running. Only a close
that proved the whole tree gone stops them now; otherwise they settle unknown,
as for an exit of the session's own. The stopped ending is marked as Orca's,
so a frame of the child's own still replaces it.
* docs(native-chat): comments describe the Stop as it now works
What re-drives a failed child end, when the child-work decoder stops what is
live, what a Stop's wait re-reads on a result, and why a card's Stop is
unnamed.
|
||
|
|
b245b3e019 | feat(sidebar): let the host pill be turned off per card (#24299) | ||
|
|
a37e5026fc |
test(native-chat): cancel the journal import test's tick before teardown (#24071)
The test re-arms a setImmediate tick while the copy runs and stopped it with a flag. A tick already queued when the import resolved still ran, and when the afterEach teardown closed the database first it threw journal_closed as an unhandled error, failing a CI shard with every test green. Clear the queued tick instead. |
||
|
|
2a83c9536f |
ci(daemon): gate PRs on daemon protocol crossing from the newest release (#24089)
Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working tree must attach the newest release tag's daemon. Rollback crossing is reported only. Runs in the cross-version-wire job, which already has full tags; tag selection moves to config/scripts/stable-release-tags.mjs so both use one rule. Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
49a83deaef |
refactor(orcad): make profile backup and preflight runtime-neutral (#24088)
* feat(persistence): run profile backups in the worker whenever its entry is bundled * refactor(orcad): make profile and native preflight runtime-neutral The profile preflight parser now takes the expected runtime identity from the caller (shipped callers pass the pinned Bun identity), and the native preflight is renamed to orcad-runtime-native-preflight with neutral wording. * test(persistence): skip plain-Node backup selection tests in the Bun profile suite --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
a8a434db98 |
fix(native-chat): a resting chat lists its model the way the running agent does, so no stray effort picker appears (#24267)
* fix(native-chat): a resting chat lists its current model as a live child does, with no phantom effort
A chat at rest (no live child) answered its options from the host's model
catalog alone. When the chat's model is one the catalog does not list, the
answer left it out, the client filled the gap from its static seed, and the
composer showed an effort control ("Medium") that the live child never offers;
the control vanished again once the child was back.
The live Claude and Codex answers and the resting answer now build their model
list through one function: the catalog's rows, plus the current model when the
catalog does not list it, with no effort levels. Live and rest can no longer
disagree about it.
* fix(native-chat): a resting chat with no catalog yet lists what a live child would, so a listed model keeps its effort
When the host has no model catalog for the account yet (the first read before
its probe returns, a failing probe, an account-home lookup that failed, or a
WSL-pinned chat no live child has written one for), the resting answer listed
only the saved model, as unlisted, and a listed model such as opus lost its
effort control. It now rests on the list a running child falls back to:
Claude's built-in models, then the shared function. Codex lists nothing
without its catalog, so the resting answer leaves the list to the client's own
defaults, as a failed live read does.
* fix(native-chat): a resting chat with no pick and no catalog names no model, as before
With no catalog, the built-in list's default model is a guess, not the
account's: naming it moved a chat the running agent reported on Opus to
Sonnet. Only a real listing names the default now; otherwise the answer names
none and the client keeps the model the agent last reported. The comment on
Codex's no-catalog answer now says what it does: the client fills the current
model from its own defaults, unchanged from before.
|
||
|
|
ebc479b9c0 |
fix(native-chat): /clear starts nothing; the new chat's first message starts its agent (#23935)
* fix(native-chat): every journal append reaches the chats that are open
A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.
A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.
* test(native-chat): an epoch replacement reaches the open chat
* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map
* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite
The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.
* test(worktree-activation): restore the OMP surfaced-agent resume test
The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.
* perf(native-chat): a publish behind a delivered commit reads nothing
Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.
* test(native-chat): state why the teardown test's fake journal is safe to cast
* docs(native-chat): say mutation admission checks only the writer lease
* docs(native-chat): drop the send rebase from comments that still described it
* fix(native-chat): a message is accepted, then delivered
A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".
A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.
Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.
A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.
Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.
* fix(native-chat): settle queued messages only for the child that ended
A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.
A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.
The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.
* fix(native-chat): an adoption that fails to import keeps the conversation open
The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.
* perf(native-chat): the recovering open reads the journal once
Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.
* fix(native-chat): an attach that fails after indexing its child leaves no child behind
A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.
* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer
The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.
A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.
* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down
The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.
* fix(native-chat): a message rejected while its chat was closed reads as not sent
A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.
* test(orchestration): name why the readiness settlement fakes are cast
* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent
* docs(native-chat): drop the fence from the admission the send effects run behind
* docs(native-chat): give the fence move on release the reason that still holds
* docs(native-chat): stop citing a write fence check in launch and mailbox comments
Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.
* refactor(native-chat): the provider child is its own record
A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.
- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.
* fix(native-chat): the delivery loop alone settles a message its start or child failed
A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.
- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
reads how it ended: a Stop continues; anything else writes one failure row and rejects every
queued message with the same words, then stops. A child still starting whose start the adapter
says did not land fails the same way. The exit, eviction and the settlement retry only settle
the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
closed, with or without a child, and a start the loop already has in flight is waited for so the
child it produces is stopped rather than left behind.
* refactor(native-chat): a stopped child ends on the one reading of its stop
The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.
* feat(native-chat): the host says it accepts a send before any agent has it
The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.
* refactor(native-chat): an attach never opens a journal of its own
The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.
* fix(native-chat): a moved fence resends nothing on a host that accepts first
The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.
The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.
* refactor(native-chat): a child's end says whether the user or the host stopped it
The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.
* fix(native-chat): a chat whose only work is a queued message is not offered for resume
A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.
* fix(native-chat): the conversation outlives its agent
Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.
- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.
* test(native-chat): type the queued-message fixtures in the resume-offer tests
* fix(native-chat): a start that dies while a message waits on it is that message's failed start
Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.
* fix(native-chat): a request that failed reads as failed
A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.
The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".
* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now
* test(native-chat): a verdict change republishes the mobile status projection
* refactor(native-chat): the store's retention trigger keeps its flag compare
A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.
* test(native-chat): a user message the provider journaled keeps its session listed
* test(native-chat): pin what a failed start settles, and what a resume offer names
A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.
* test(native-chat): the failed-start pins fail on what the message became, not on a timeout
* fix(native-chat): a restart offer ends when the chat's agent starts again
The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).
A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.
Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.
* fix(runtime): end a transcript stream when its client unsubscribes
Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.
Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.
* fix(native-chat): a late provider-session update keeps a failed recovery record failed
A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.
* test(orchestration): the preamble's host stub is typed, not cast
The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.
* test(native-chat): the terminal-bell check asserts the renamed verdict field
The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.
* fix(native-chat): a failed turn ranks like a completion for attention
Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.
The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.
* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart
The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.
The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.
Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.
* test(native-chat): an older build reads the restart offer this build records
The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.
* fix(native-chat): read a restart offer against where the journal stood when it was taken
"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.
The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.
* test(native-chat): wait for the listing's retire write before reading the recovery file
* refactor(native-chat): every journal row states which turn it belongs to
Rows gain a turn scope stated by the write that creates them: the open root
turn, or the conversation. A queued message takes its scope from its handover.
Rows stored before scopes existed are placed on replay by the root turn open
when they were created, so no persisted state is needed for them. Rewind keeps
each retained row's scope and producer, so a subagent's row stays its own.
* fix(native-chat): keep the terminal-backed chat's read error over its local echoes
Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.
* fix(native-chat): a start retries the exit settlement a failed journal write left owed
An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.
* fix(native-chat): a failed main agent reads failed while its subagents still work
The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.
Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.
worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.
* perf(native-chat): answer the owner check without opening the chat
Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.
* fix(native-chat): a read waiting on the session lock opens nothing once quit began
The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.
* test(native-chat): pin stated turn scopes, the upcast of unscoped rows, and rewind attribution
* fix(native-chat): /compact is a message the chat sends, run as a turn of its own
The conversation command RPC now accepts /compact into the queue like any
send and answers once it is handed over. The delivery loop opens the command's
own turn, starts the provider on it, and waits for the provider's end off the
session's queue, so messages typed meanwhile are held and delivered after it,
even when it fails. It settles by re-reading the journal: a child that died
meanwhile already wrote the verdict. Stop ends the command at once. The 180 s
completion window, the unconfirmed row and the recovery of an older build's
compaction record are gone; that record no longer gates anything. On Codex the
provider turn the command opens is claimed into the command's turn.
* fix(native-chat): read a failed resume's chat before calling it retryable
Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.
* test(native-chat): type the provider event sink the settlement test reaches for
* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it
The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.
* fix(native-chat): say the structured read keeps trying only where it does
The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.
* fix(native-chat): rows group under the turn their record names, not the one above them
Each row's turn is the turn its stated scope names, anchored on the entry
that opened it, or on the turn itself when the provider opened it unasked.
So /compact groups its own rows and the previous turn is untouched, a message
typed into a running turn joins it, and a provider-resumed turn folds under
its own Worked-for. A row reporting how a turn ended, an error or the
compaction separator, never folds. Desktop and mobile read the same keys; a
host that states no scope keeps today's positional grouping.
* test(native-chat): await the send's settlement instead of polling for the start
The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.
* docs(native-chat): the status-store listing rule names provider-journaled user messages
* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations
The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.
Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.
* fix(native-chat): a /compact is not a request the sidebar, notifications or restart resume report
The sidebar's prompt, preview, verdict and instant, the turn-completion feed,
and the restart-resume marker read past a conversation command and its turn to
the last real request, so a /compact neither notifies nor re-dates the row,
and a command in flight is never offered as work to resume. An older client
shown a command's turn in the legacy form names the session's own agent.
* fix(orchestration): route no mail to a structured worker its orchestration released
A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.
* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation
A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.
* test(native-chat): pin what a conversation command's admission refuses at rest and at handover
* test(native-chat): tests merged from the base state which turn their rows belong to
* fix(native-chat): a refused send notifies failed through the completion feed
The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.
* fix(orchestration): read the released row optionally, as the authority does
worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.
* chore(native-chat): one import per module and no unexplained casts in the turn-scope changes
* test(claude): pin which turn a Claude row joins, including a subagent's after the turn ends
* fix(native-chat): the status bar drops a restart offer the chat moved on from
The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.
* fix(native-chat): a refused steer is read from the turn its handover named
The latest-request reader decided whether a refused send had joined a running turn by comparing
host clocks: its handover time against the previous turn's end. The handover row now states the
turn it delivered into, so the reader reads that instead and the clock comparison goes. A journal
written before handover rows stated a turn is scoped on replay from the turn open when each row
was written, which can differ from the clock reading only when a send and a turn's end share a
millisecond.
* fix(mobile): the native-chat controller contract carries the turn journal
The controller and overlay already pass nativeChatTurnJournal, but the
contract type never declared it, so mobile failed to typecheck.
* fix(native-chat): the live turn is the running turn, not the newest user row
A turn the provider opened on its own (a background wake, a resumed turn)
anchors on its own record, but the list still treated the newest user row
as the live turn. While such a turn ran, the settled user turn before it
lost its duration and the running turn's own rows were drawn as settled,
so its tool calls lost their live state.
nativeChatTurnMembership now answers both questions from the turn record:
each row's turn, and the live turn (the running root turn's anchor, else
the newest user row, which is also all an unscoped host has). Desktop and
mobile key liveness, the timing clock and the live status's row on it.
* test(native-chat): a turn the provider opened keeps its own clock
Pins that the local turn clock follows the live turn, so a wake after a
settled turn does not restart that turn's clock when no host durations
are recorded.
* fix(native-chat): a running turn no message opened draws its status on no row
Its live status belongs to the transcript-tail indicator alone. Once it
settles, its duration draws at its first row as before; a running turn a
message opened still draws on that message.
* fix(native-chat): every copy of a row carries the main agent's own status
History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.
- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
rebuilding one; the sync key and history equality compare it.
* test(native-chat): pin the worktree ps verdict across host and phone versions
Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.
* test(mobile): name the parity table's row for its role
* test(native-chat): a roster of idle or finished children does not keep an agent awake
The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.
* fix(native-chat): a request that settles while the user is asked something notifies once
The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.
The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.
* fix(orchestration): a task dispatched into a resting structured worker keeps it running
The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.
* fix(native-chat): a command's wait ends when its child does
The delivery loop waited for a /compact only on the adapter's compaction
tracker, which learns of the child's end only on some exit paths: a Codex
exit or close, and a Claude close, never reach it. The wait then never
ended, so nothing queued behind the command was delivered again, Stop had
no child to answer through, and the tracker's leftover entry refused the
next /compact.
Every way a child ends passes endProviderChild, so the host now offers a
per-child end signal there. The loop races the tracker against it (the
dead-generation settlement has already written the command's verdict),
and on that end asks every adapter to release the command, so a later
command runs and no later provider turn is claimed into the dead one.
The adapters' own exit-time releases were unreachable (Codex) or covered
one path of several (Claude), and are removed.
The Codex RPC test harness moves to its own module so the exit can be
driven through the real adapter's connection callback.
* fix(native-chat): keep refusing sends during a command on an older host
An older host's controller still refuses a send while a conversation
command runs, so dropping the client's block turned every message typed
during /compact into a 'not sent' row with Retry there. The block stays
for hosts that do not run the command as a send-path turn, and goes only
for those that do.
The signal is one the client already holds: a host that runs /compact on
the send path states a turn scope on every journal row it writes, the
same fact turn membership uses to tell it from an older host. Both now
read it from one predicate. On an empty conversation, or one whose rows
all predate the upgrade, the signal is absent until the command's own
entry streams in, so that brief window keeps the old local refusal; no
capability or wire field is added.
* docs(native-chat): comments stop describing the hold this PR removed
Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.
* fix(native-chat): the completion says when the user is being asked
A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.
The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.
* fix(native-chat): a restart offer keeps the start its own continuation made
Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.
* fix(native-chat): a rewound turn still names the message that opened it
A Codex rewind rebuilds the epoch without submissions, so each sent message survives only under
its provider key. The kept turn records still named the submission key, so each turn anchored on
itself and its rows grouped apart from the message that opened it. The rewind now renames the
turn's opener along with the message.
* fix(native-chat): Stop ends only the command it names
Stop on a command turn abandoned whatever compaction the session had pending, so a late Stop for
an earlier /compact cancelled the one running now. The tracker now ends a command only when the
Stop names its turn, and the cancel reply reports whether it did.
* fix(native-chat): an agent gets a full idle window after its owed work ends
The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.
* test(claude): the options-read fixture runs a live child
The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.
* test(native-chat): host tests reach its collaborators through a typed seam
The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.
* fix(worktree-status): a departed agent's failure yields to live work on the worktree card
A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.
* refactor(orchestration): one owner answers a structured worker's custody
Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.
* refactor(orchestration): owed work is an open dispatch on the worker's incarnation
A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.
* docs(agent-status): a departed agent's failure ranks below live work on the worktree card
* fix(native-chat): a restart offer knows its continuations by a tag in their id
The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.
* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait
The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.
Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.
* fix(native-chat): a message held behind /compact is drawn where it was handed over
A message typed while /compact runs was drawn above the compaction's result, between
itself and its own answer. The reducer kept every item at the sequence and timestamp of
the row that created it, and a queued message is created at acceptance, long before the
command it waits behind writes its result. The phone orders by that sequence and the
desktop by that timestamp, so both put the message first.
A queued message now takes its position from its handover row, the same row that already
states its turn scope. Everything the agent did before the handover, a command it waited
behind included, draws above it. This holds for every held message, not only /compact's,
and needs no client change: every client, older builds included, reads the position the
host publishes. A live batch already carries the item when its dispatch row lands, and
history pages cut the reduced timeline by sequence, so paging stays contiguous.
* fix(native-chat): a phone's send during /compact answers without waiting out the compaction
A client that predates accepted-send replies, which is every phone build, has its send
reply held until the host hands the message over. A message sent during /compact is not
handed over until the compaction ends, so the phone's 15 s request timeout fired first
and showed the message as unconfirmed.
That wait now also ends once the message is queued behind a running command. This is
read from the journal's running turn and needs no new state. Every other wait still
ends at the handover: behind a starting child or an ordinary turn, and for restart
resume, the command front door and orchestration, which keep the plain handover point.
* perf(native-chat): a rewind places provider items with one pass over the merged rows
A Codex rewind gives each provider item the old epoch never held the turn record for its
provider turn. It found that record by scanning every merged row, restoring each row's
body, once per provider item. That is quadratic, and it runs on the host's main thread
up to the journal's 10,000-row cap, twice per rewind. A rewind record written before
rows carried their scope holds no scope for any provider item, so it paid the full cost.
The merge now indexes turn records by provider turn id once, keeping the first match as
the scan did, and each provider item looks its record up.
* fix(native-chat): a view never restarts a chat whose last start failed
A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.
* fix(native-chat): a message waiting behind /compact is drawn after it until it is sent
A message sent while /compact runs is placed where it was handed over. It was still
drawn where it was accepted until then. /compact writes its result one step before the
handover, so for that step the waiting message sat above the compaction's separator.
A message the host accepted but has not handed over is not part of the conversation
yet, so both clients now draw it after everything the agent has done. The shared
projection moves it to the end, which is the order the phone draws. The desktop ranks
it with the other not-yet-sent rows, after the streaming preview. At handover it takes
its place from its handover row, which is also after the separator, so it never
appears above the compaction it waited for.
* fix(native-chat): the idle sweep reads owed work every tick
Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.
* fix(native-chat): a continuation handed to the agent stays sent
The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.
* test(native-chat): start the child the loop waits on with an attach, not a second view
A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.
* fix(native-chat): settle a gone generation's turn wherever a conversation opens
A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.
* test(native-chat): prove the next child's start settles the turn an earlier child left
The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.
* test(native-chat): count a failed start's rows by row, not by text
Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.
* test(cross-version): load the phone row readers without mobile's toolchain
Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.
The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.
* test(cross-version): keep the checkout path-guard message and justify the copy import's cast
* fix(native-chat): a command ends only by its own provider answer or its child's end
Stop no longer settles a conversation command. It interrupts it like any turn,
and when the provider cannot take that (Codex has not opened the command's turn
yet, or Claude refuses the interrupt) it stops the child, whose dead-generation
settlement writes the verdict.
The pending command now lives on the provider child's own session instead of an
adapter-wide map keyed by session, so it dies with the child and nothing has to
release it. Claude's /compact is sent under a uuid the slot records, and only a
root result naming that input (or naming none) ends it; its outcome is read with
the ordinary result reading, so a stopped /compact is a cancellation.
* fix(native-chat): a command's settle answers its message before ending its turn
The two writes are not one batch. Writing the message's answer first means a
crash between them leaves a running command turn, which the stale-turn sweep
already settles, instead of an ended turn whose message reads as in flight
forever. The settle now writes only while the command turn is still running.
* fix(native-chat): "Worked for" counts from the handover, not the send
A message held behind /compact, or behind a cold start, used to count the wait
as the agent's work, although its row is drawn at the handover. Every handed-over
submission's turn, the command's own included, now starts at the handover row's
instant, falling back to the send time for a host that recorded none.
* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget
* test(native-chat): the interrupted create's own retry continues again
The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.
* docs(native-chat): three comments that still had views starting agents
A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.
* test(native-chat): pin the open's and the send's start and row counts, however the view binds
Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.
* fix(native-chat): a second Stop on a command ends its child; one compaction verdict for every provider
A Stop's note now names itself in its key, so a later Stop on a command still
running reads, from the journal, that the provider was already asked and never
answered, and stops the child instead of interrupting again. Nothing is held in
memory for it.
Adds the rule both translators will read a compaction's end by: only a
compaction the provider reported is a success; none after Orca's interrupt is a
cancellation; anything else is a failure. A real Claude capture, pinned as a
fixture, is why: a stopped /compact ends in the same success result as a
finished one.
* test(native-chat): a reader's open settles the turn a failed exit settlement left running
An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.
* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner
A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.
* fix(native-chat): settle a gone generation's turn at every open but an acquisition's
The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.
* test(native-chat): hold the create's start open until the views bind
The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.
* refactor(native-chat): the provider's translator ends a command's turn; the loop holds no command state
A conversation command is now a turn of the provider child's own journal
pipeline. The adapter-wide tracker, its promise and the loop's settle step are
gone.
- Codex: the translator claims the provider turn that carries the command, scopes
its rows to the command's turn, and writes the command's end in the same batch
that settles that turn. Codex's own compaction marker is the success row.
- Claude: the command's turn is the translator's open turn until the result that
answers the /compact input ends it. The command's own frames, such as the
continuation summary, its echo and "Compaction canceled.", draw nothing.
- Both read the end with the one compaction rule: success needs the provider's
report of the compaction; none after Orca's interrupt is a cancellation.
- The message resolves at the provider's receipt, as any send does: the Codex
ack, or the Claude slash-command waiter on its result. The host writes a
command's end only when the provider never took it.
- The delivery loop stops while a command's turn runs, and every journal commit
re-wakes it through the session's serialize, so an end that lands while a step
decides to stop is never lost. A child that ends first is settled with it.
* test(native-chat): pin a command's end to real /compact frames and to each path it threads
The captured /compact frames drive the Claude translator's command turn: a
finished compaction ends as a success with only the separator drawn; a stopped
one ends as a cancellation with no failure row, and the next send answers in its
own turn; a result naming another input ends nothing. The command's end is
checked at each point the ordinary result path threads through: the reopen latch
after a failure, the settling of a child still working, the context facts the
result reports, and the provider's own error row.
On the host: a message held behind a command is handed over when the command
ends just as the loop stops for it, a refused command settles as a failure and
the loop moves on, and a Claude child that exits mid-command settles the command
and hands what waited to a fresh child.
* test(native-chat): tests merged from the base state which turn their rows belong to
* refactor(native-chat): drop the child-end waiter nothing waits on
A command no longer waits for its child here: its turn ends from the provider's frames or from
that child's settlement, and the delivery loop is woken by the commit. The waiter and its test
were left from the earlier shape.
* fix(native-chat): a command holds the queue only while its child runs it
The delivery loop stopped whenever the journal showed a command's turn running. When the
command's child ended and its settlement could not be written, that turn stayed running with
no child to end it, and the loop's gate kept it from ever starting the next child, which is
what settles a gone generation's leftovers. Every later send was held for good, and Stop had
no child to end.
The gate now holds only while the conversation has a child: with none, the command belongs to
a gone generation, and the loop's start settles it like any turn a dead child left running.
* fix(native-chat): a Claude /compact succeeds only on its compaction boundary
The command's evidence counted Claude's `compact_result: 'success'` status as the compaction
done. That status comes before the boundary that replaces the history, so a Stop landing
between the two read as a finished compaction even though no boundary was ever written. Only
the boundary now counts, as the rule for both providers states; the capture's finished
compaction carries one, so it still reads as a success.
* fix(native-chat): a Claude child's exit says why the turn it ended stopped
When a Claude child exited mid-/compact, the command showed "Worked for 0s" and no reason. The
child's translator ends its open turn the moment the exit is reported, stamped with the exit's
instant, so by the time the exit settlement ran nothing was running. The settlement recognises a
turn the exit already ended by that same instant, but the Claude lifecycle event dropped it on the
way to the host, which then used its own clock, matched nothing, and wrote no row. When the clocks
did agree, the row was scoped to the running turn, of which there was none, so it landed outside
the turn it explained.
The exit's instant now reaches the host, and the exit row belongs to the turn the exit ended:
still running, or ended by the translator at that instant.
* fix(native-chat): a message waiting behind /compact draws below its live activity
A message sent while /compact runs waits on the host until the command ends. Both clients moved
it to the end of the transcript rows, but the running turn's live activity line ("Compacting the
conversation") draws after every row, so the waiting message sat between the command and its own
live status.
A row that is queued, and not what the live turn is for, now draws after that live activity: on
desktop outside the transcript window, below the activity line; on the phone in the list footer,
below the live status. A message whose own start is pending still draws above the activity that
start reports.
* fix(native-chat): only a running command holds a message below its live activity
A message is accepted, then handed over a moment later, and in between it reads as waiting. Every
message waiting behind a live turn drew below that turn's activity line, so an ordinary message
sent while the agent was working crossed below "Thinking" and jumped back up once it was handed
over, on desktop and phone. Only a conversation command's turn holds the queue on the host.
A message now waits below the live activity only while the running turn is one a command opened,
read from the entry that opened it. The phone test also typechecks, which the mobile test ratchet
requires.
* test(codex): the claim test names its notification params as a record
* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn
On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.
* refactor(native-chat): drop the composer's second error formatter
After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.
* test(native-chat): pin the reason on a message rejected while its chat was closed
The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.
* docs(native-chat): drop the removed dispatch hold from six comments
A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.
* test(native-chat): rest the owner-status chat through the idle sweep, not a hold
The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.
* fix(native-chat): show the structured pane's retrying line when a read fails
The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.
* fix(native-chat): a send the provider never received after a restart has no verdict
Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.
* fix(native-chat): a failed Codex compaction's late completion writes no turn of its own
Codex ends a failed turn with an error and then still completes it as failed.
The error settled the compaction and released its claim on the provider turn,
so the completion read that turn as an ordinary one and wrote a stray record.
The claim now lasts until the completion, which adds nothing to a command the
error already ended.
* test(native-chat): the mid-command exit case resumes its next child as a real one does
The case's fake started every child as a newly created thread with the same generation. The
store refuses a created link once the conversation has a thread, so the next child's start
failed and wrote its own error row, which landed before or after the case read the journal.
The next child now resumes the thread under its own generation, and the case reads the
journal once the waiting message is delivered, which also proves the loop moved on.
* fix(native-chat): a /clear that never committed no longer locks the chat
A /clear wrote a durable "prepared, outcome unknown" record before starting
the replacement conversation. When that start was refused without a definite
answer (or Orca died), the record stayed forever, and while it did the chat
refused every send, /compact, a new /clear and rewind. Its only exit was a
rerun under the same operation id, which only the renderer held.
The record guarded nothing the process does not already know: a clear in
flight holds the session's serialize for its whole run and the command
controller refuses sends meanwhile, and the replacement's id and start
operation are pure functions of the clear's operation id. So the clear now
writes nothing durable before its commit, the gates refuse only a committed
clear (an older build's prepared record is inert), and a clear with no
committed answer reruns: a same-op retry re-attaches the same replacement,
a new op id runs a fresh clear.
A crash between the replacement's start and the commit leaves a replacement
record nothing points at. Verified: it has no tab, is not in the
replacement list, and a restart opens and starts nothing for it (restore
reads only the visible tab index); restart reconciliation releases its lease
like any dead owner's. In a live process its agent is stopped by the idle
sweep like any quiet agent. Session History lists provider transcripts and
only annotates them with an owner, so it can list this only if the provider
wrote a transcript for a thread that never got a message. Its record stays
on disk, as every closed chat's does; the store deletes none.
* fix(native-chat): a Codex rewind the provider did not keep no longer fails every attach
When Codex acknowledged a revert and Orca stopped before proving it, the
rewind stayed prepared with providerApplied set. On the next attach,
recovery read the provider's history, found the target turn still there
(provider-refused), and threw, because that settlement was limited to
reverts never sent. The throw ran inside the attach, so every attach, and
every send that needs one, failed for good.
The journal is replaced only once the provider proves the revert, so both
the provider and the journal still hold the target turn: settling the
rewind refused is consistent whether or not the provider acknowledged it.
* test(native-chat): a clear retried after a crash starts no second replacement
The replacement's id is the only thing that keeps a retried clear from leaving a second one, and no test held it across a restart.
* chore(native-chat): the clear rerun comment claims only the stable replacement id
* test(native-chat): wait for a send's background start before the refusal oracle removes its store
An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals.
* fix(native-chat): a start a message waited on gets one failure row, the delivery loop's
When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice.
The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit.
* fix(native-chat): a /compact whose start failed says to run /compact again
The failure-words context named only /clear as a command to retry, so a
/compact whose agent failed to start read "Send your message to try again."
on its row, its rejected message and the command reply. The context now
carries any conversation command; the host derives it from the oldest
message still waiting on the provider, which is the one a failed start
fails first, and the /compact reply names it directly.
* fix(native-chat): a Codex /compact ends only on its turn's completion, below Codex's own error row
Since only turn/completed ends a Codex turn, Codex's turn-ending `error` is a row
inside the still-open command turn, and the failed completion that follows it is
the command's end: completed, outcome failure, at the completion's receipt time.
The command's own "Compaction failed" row was written on that completion too, so a
failed /compact read its reason twice.
The command turn now notes when Codex's turn-ending error for the turn it carries
was written as a row, and its end then adds no second row. A retried stream error
ends nothing and is not counted. The flag that let the error end the command and
kept the claim until the completion is gone with the error-driven end.
A test replays the captured failed compaction from the real app-server through a
claimed command turn.
* test(native-chat): main's crash-turn test states its row's turn, and a dead /compact settles on its recorded exit
Two tests the main merge brought together:
- The crash-turn test from #23456 writes a turn record through the event sink
without options; every row here states its turn scope, and a turn record's is
the thread.
- The /compact whose exit settlement could not be written no longer stays running
until the next start: main now settles an open chat from the exit it recorded, so
the command reads interrupted before the next message, which is then delivered.
* test(native-chat): main's new journal tests state each row's turn
The crash-turn, stale-turn and sink-queue tests main added wrote rows without a
turn scope, which every item write now states. Rows written inside a running
turn name that turn; the sink-queue batch and a send handed over with no live
turn name the thread.
* fix(native-chat): draw a queued turn's message after the earlier turn's rows
A message sent while A runs is written to the journal when it is sent.
When the provider queues it (Claude answers it after A), A's remaining
rows - its last tool run and its answer - are written after that
message, and the message's own turn opens only after them. Grouping put
those rows in A's turn, but the transcript still drew them in journal
order, below B's bubble and bar, where A's answer read as B's reply. This
is the residual #23671 left open.
A message that opened a turn now draws after the earlier turns' rows the
journal wrote after it, just before its own turn's rows
(nativeChatTurnDrawOrder, returned by nativeChatTurnMembership as
drawOrder). Desktop and mobile both draw in that order. A steer, and a
message that has opened no turn yet, stay where they were written. It
applies on hosts that state turn scopes and, through journal order, on
older ones.
* test(native-chat): run #23026's Stop tests against #23059's command turns
Two of #23026's tests call APIs #23059 changed, and failed after the
merge:
- codex-structured-conversation-stop: a compaction now goes through
adapter.compact with the command run the host wrote (#23059), not a
bare turn id, and answers with the provider's receipt. With the command
claimed, a Stop that names no turn while the compaction's provider turn
has not opened still interrupts nothing.
- main-agent-working-agreement: a provider row states its turn scope
(#23059's appendItem contract); the retry and subagent rows are
conversation-scoped.
* fix(native-chat): typecheck main's Stop and restore-grouping code against #23059
A Stop's compaction interrupt reads the narrowed requested turn, and the
restore-grouping test states whether each row reports its turn's outcome.
* fix(native-chat): say a /clear cut off by a restart left the chat unchanged
A /clear retried under the same operation after Orca restarted could not reuse the new conversation its first try started, and its row said "Codex couldn't start. Run /clear again." The agent did not fail to start: the earlier try was cut off. The row now reads "This /clear didn't finish, so the chat is unchanged. Run /clear again to start fresh.", from a new clearUnfinished failure fact written through agentSessionFailureWords.
The clearUnconfirmed and conversationCommandUnconfirmed reasons stay, with their words, for older hosts that still send them.
* fix(native-chat): a retried /clear finishes onto the conversation its earlier try started
When an earlier try of the same /clear started its replacement conversation and a restart or the
idle sweep has since stopped it, the retry could not replay that settled start and reported the
chat unchanged. That replacement is a fresh conversation at rest, so the retry now commits onto it
and its first message starts its agent. A replacement whose start definitely failed still reads
that failure, and one Orca can't prove stopped still commits nothing. The clearUnfinished failure
kind this made unnecessary is removed.
* refactor(native-chat): stop recording that Codex acknowledged a rewind
A refused rewind recovery now settles as refused whether or not Codex acknowledged the revert,
so nothing reads providerApplied any more. Stop writing it and drop the hook that wrote it.
Records that still carry the field load as before; the schema ignores the extra key.
* fix(native-chat): a /clear retried under a new operation id finishes the same replacement
A /clear's replacement id came from the client's operation id, so a retry the client sent
under a fresh id started a second replacement and orphaned the first. The host now derives
it from this caller's oldest /clear since its last commit whose replacement start reached
the operation ledger, so any retry from that caller finishes the same replacement, including
after a restart. A /clear after a committed one starts a new replacement. Another caller's
/clear is refused only while such a replacement is running or not proven stopped. An older
client that resends the same operation id still lands on the same replacement.
* fix(native-chat): a /clear retry never repeats a failed start or waits on an unproven stop
A retry under a new operation id could pick an earlier try whose replacement start had already
failed, replay that failure and commit it again, so a user who had since signed in was told
they were still signed out. Such a try is now skipped, and the retry starts afresh.
Another window's /clear was refused while the first window's leftover replacement was merely
not proven stopped. Nothing but the first window's own retry would settle that, so the refusal
could last until its ledger row expired a day later. It now waits only on a replacement whose
agent is running.
* fix(native-chat): a /clear retry finishes only a replacement that started
A retry picked an earlier try whose replacement start never answered, because a crash left
that start unsettled. Replaying it could only repeat "couldn't start" or, with the old agent
unproven, refuse every /clear from that window. Only a start that succeeded left a
conversation to finish; any other try is skipped and the retry starts afresh.
* refactor(native-chat): a record's identity fields are built in one place
A created record and a founded one (a conversation no agent has run yet, at
rest) share who and where the agent is and how it launches. The founding
builder is used by the /clear commit that follows.
* feat(native-chat): the store commits a /clear and its new conversation in one write
commitConversationClear founds the at-rest replacement from the cleared
record's identity and writes the committed marker and tab move in the same
transaction, so neither can land without the other. It refuses to overwrite
an existing record under the replacement id.
* fix(native-chat): /clear starts nothing; the new chat's first message starts its agent
/clear used to start the new conversation's agent before it committed, so it
could fail on that start ("Run /clear again"), and a crash between the start
and the commit left a running conversation nothing pointed at. #23524 then
needed a ledger scan to find an earlier try's replacement, a nonce half of the
derived ids, a refusal of another window's /clear while a leftover agent ran,
and a check for a start that had already finished.
Now /clear opens the chat for writing (it no longer starts an at-rest chat's
agent either) and makes one store write: the at-rest replacement under a
random id, the committed marker, and the tab move. The first message in the
new chat starts its agent through the existing send and delivery path, fresh
because its handle chain is empty. A failed start shows on that message with
the typed failure and a Retry, and a conversation no agent ever ran now reads
"couldn't start" rather than "couldn't restart".
Deletes clearTryToFinish, otherCallersClearIsLive, the attach block and the
committed start-failure branch, and the tests of that retry machinery.
* test(native-chat): drop the /clear retry wording test; no start runs for a /clear now
* test(native-chat): another window and a phone read a /clear's replacement from the host
Both list the replacement the committed marker names, under the chat's tab,
and each one's session list shows it with nothing unread until its first
message runs. A reader that recomputed the id from the operation turns this
red.
* test(native-chat): a never-started replacement closes as settled
A worktree delete closes every chat in it and asks the user to force any it
cannot prove stopped. A replacement no agent has run is released, so its
close settles like any at-rest chat's.
* fix(native-chat): /clear settles an interrupted Codex rewind the way a send does
/clear moved from starting the chat's agent to only opening the
conversation. A Codex rewind cut off mid-way on a chat at rest can only be
settled by its agent, so /clear was refused as "rewind unconfirmed" every
time until the user happened to send a message. It now prepares like a send
or /compact: the agent starts only when such a rewind is in doubt.
* fix(native-chat): a chat whose agent is not running keeps its `/` commands
Claude reports its skills and project commands only from a running process,
and the host served the `/` menu only from the running agent. Now that
/clear starts nothing, the new chat's menu lost those entries until its
first message; a chat stopped by the idle sweep already did.
The host now keeps, in memory, the list a running agent last reported for
its launch (provider, host, workspace, account and launch arguments) and
serves it to a chat of the same launch whose agent is not running. A new
report replaces it; nothing is stored on disk, so a relaunch still shows
the short menu until the agent reports again, and no list is ever served
across accounts, workspaces or hosts.
* test(native-chat): queued drafts around /clear follow what a /clear now is
Three queue tests from #23726 are red on main
|
||
|
|
3135fbbf49 |
feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check (#24087)
* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS, NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball), generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no network, that the pin tracks the locked Electron, matches engines.node's major, and covers exactly SERVER_TARGETS; it runs in the static analysis job. ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target list; orcad's Bun runtime and build output are unchanged. * fix(runtime): reject a pinned archive that belongs to another target --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
7e5950c1c3 |
fix(ssh): keep keystrokes typed while a restored SSH terminal reattaches (#24166)
A restored SSH pane that remounts takes keyboard focus as soon as its new xterm mounts, but its transport only binds once the relay answers the reattach. Keys typed in that window were refused by the transport and lost, so the start of a command vanished while the rest ran against the live shell. Buffer input on SSH panes that reattach to an existing PTY and flush it in order once the reattach binds. If the reattach fails or the pane falls back to a fresh shell, drop the buffered keys with a console warning instead of typing them into a replacement shell. Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
3fe4b18dae |
fix(ai-vault): require node:sqlite backup support in remote SQLite probes (#24086)
* fix(ai-vault): require the full SyncDatabase node:sqlite surface in host SQLite probes The SSH and WSL OpenCode probes admitted any Node with DatabaseSync, so Node 22.13-22.15 hosts (no backup export) skipped the pinned-runtime fallback. Share one admission predicate with isSqliteAvailable() and embed its source in both probe scripts. * build(cli): list the node:sqlite admission predicate in the CLI project --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
07e9fdfd13 |
fix(native-chat): a message the chat said was not sent is never sent later on its own (#24232)
* fix(native-chat): keep a message the chat said was not sent held until its Retry A native-chat send the host refused (for example "Chats were saved by a newer Orca. Your message was not sent.") or that never reached the host showed "not sent" with a Retry button, but the hold that stopped it lived only in the outbox hook's memory. The message itself was saved in the outbox, so the next launch lifted the hold and sent it with no Retry; a copy the user retyped in the meantime was held behind it and went out as well. The hold is now read from the failure the message already saves: a queued entry that carries its last failure waits for the user's Retry, on this launch and every later one, and an entry saved by an earlier build in that shape is held too. The drain passes over a held message instead of stopping behind it, so what the user sends next goes out as they send it. Retry clears the saved failure. On a host from before accepted-send, a new agent owner observed while the chat is open still sends the refused message again, as before; a relaunch does not. A held message keeps its operation id, and a host refuses an id older than a day as expired for good, so its Retry could never go through; a new id could deliver a message an earlier attempt already delivered. Such a message now goes back to the composer with a notice to check the chat before sending it again, and leaves the outbox. * fix(native-chat): keep refused messages as rows until Retry, and only release them for an older host's new owner - A message the host refuses as expired under an id it kept stays a saved row reading "Orca couldn't confirm what happened. Check the chat.", and its Retry sends it under a new id. It no longer moves into the message box, where an automatic resend after a relaunch could put text the user never asked for, held only in memory. - An owner change releases a refused message only on a host known to predate accepted sends, and only for the refusals such a host gives while it restarts the chat's agent. Those rows say Orca will send it again when the agent restarts, beside their Retry. A host whose capability check has not answered, or failed, no longer releases anything. - Every failed message ahead of the one the queue stopped on keeps its Retry, since that Retry sends it at once. - A journal row saying the host cannot tell whether a message landed, and the unconfirmed probe's resend, replace an earlier attempt's saved failure, so the message is probed rather than held. - The drain stages from the hook's own outbox, so a hold kept only in memory after a failed save survives the next send; the hold is written once more after that failed save. * fix(native-chat): a refused message waits for its Retry on every host, and a send is staged from the latest outbox A message the chat showed as not sent no longer goes out on its own when an older host's chat gets a new agent owner. Resending it on the owner change sent it after messages typed later, still delivered a retyped copy twice, and its "Orca will send it again when the agent restarts" row promised a resend that often never came. It now waits for the user's Retry, as it does on every current host. A send still in flight when the owner changes is still sent again under its id; it was never shown as failed. The drain admitted and staged the next send from the render's outbox. An owner change requeues the send it interrupted in an effect earlier in the same commit, and staging from the render's list wrote the old list back, leaving that send stuck as sending. The drain now reads the latest list, which still carries a hold kept only in memory after a failed save. Tests pass the view's target as one stable object, as the view does: a new object each render re-ran the owner-change requeue, which hid the drain bug. * fix(native-chat): keep a not-sent message out of newer turns, and word its saved cause only when seen A message shown as not sent stays in the outbox and draws below every newer turn. It was an ordinary user row there, so it counted as the newest user row: while a new send waited for its turn to open, that turn's "Working for" clock drew under the old message, and once the turn ended an empty "Worked for" divider was left under it. The projection now marks such a bubble (held for its Retry, or rejected) as unsent; turn membership gives it no turn and never makes it the live one, on hosts that state turn scopes and on those that do not; and the transcript draws it after the live activity, as it draws a message waiting behind /compact. Mobile has no outbox, so its rows never carry the mark and its grouping is unchanged. A held message read back from storage repeated the cause it was saved with, which may no longer hold: "Update Orca to keep using them" after the user updated Orca. The outbox hook now remembers, in memory only, which messages failed while the chat was open; only those word their cause. Any other held message reads "Your message was not sent." with its Retry, and a Retry the cause still stops brings the full words back. An expired id keeps its words, since that cause cannot clear. * fix(native-chat): follow the bottom and light a tick for a chat whose only rows are not sent A message shown as not sent draws after the windowed transcript. When it was the only row, the windowed list was empty, and following the bottom or "Jump to latest" asked the virtualizer for an end it computes from its own rows: the top. It now scrolls to the container's own bottom when no row is windowed. A rejected send the journal recorded keeps its place but opens no turn, so the rail lit no tick when it was the row being read. A user row in no turn now lights its own tick. Also pins that a refusal seen while the chat is open reaches the rendered notice in full, and reads only "not sent" after the chat is reopened until a Retry is refused again. * test(native-chat): name the relaunch test parameter for how the refusal arrives The low-evidence lint rejects "shape" as a symbol name. |
||
|
|
be575800c2 |
fix(codex): steer a mid-turn send into the running turn by name (#21062)
* fix(native-chat): settle in-flight sends when their turn ends A send the provider admits gets no dispatch row, by design: the provider's later acknowledgement is what settles it. If the turn carrying that send ends first, the acknowledgement can never arrive and the submission stays pending for the life of the session, so the chat reports work forever with no running turn. It also blocks /clear and /compact and holds the session open. Turn settlement now settles the sends that were in flight inside it. One routine owns the behaviour and both journal write paths use it, because the turn record reaches its terminal state through a plain item append on one provider and through a lifecycle batch on the other. Ownership is derived from journal order rather than stored: the pending set is captured inside the serialized row build, so exactly the sends preceding the terminal row are settled and later ones are untouched. Settlement records doubt rather than a rejection, since an unacknowledged send is never proof of non-delivery. A failed settlement is reported and never blocks the turn from settling or the next send. No schema or wire change: settlement writes ordinary dispatch rows that every client already decodes. * fix(native-chat): retire settled dispatch ownership * fix(native-chat): correlate terminal dispatch ownership * fix(native-chat): make dispatch ownership provider-authoritative * fix(native-chat): complete durable late settlement recovery * Resolve mainline conflicts in settlement plumbing * fix(codex): steer a mid-turn send into the running turn by name A message sent while a Codex turn runs, a queued card's Send-now included, went out as turn/start. Codex 0.148 and later steer that into the running turn and answer with its id, so the turn's end settles the send. Before 0.148, turn/start answers with its own submission id, which never opens or ends as a turn: the send was bound to a turn that never exists, and a Stop left it pending, so the chat read as working and /clear stayed blocked until the app exited. The send now goes in as turn/steer with expectedTurnId set to the turn Codex last reported started, and is bound to the turn the answer names. A steer Codex refuses took no input (the turn ended or changed, it cannot be steered, or this Codex has no turn/steer), so the send falls back to turn/start. A steer that times out is never re-sent and stays armed for its echo. Per-turn options ride on the next turn/start. * fix(codex): steer a send made before Codex opens the previous send's turn On a Codex before 0.148, a second send made after Codex answered the first but before it reported that turn started went out as turn/start. Codex folded it into the first turn but answered with an id that never opens or ends, so a Stop left it pending: the chat kept reading Working. A send now resolves its target the way Stop already does: the running turn, or else the turn Codex answered an earlier send into, once it opens (same bounded wait). The resolver moves into the turn-open-wait module and Stop and send share it. When Codex refuses a steer and a different turn is now running, the send steers that turn once before falling back to turn/start. The lifecycle fake's legacy mode now mints a false id for a start made while a turn is picked but unopened, and refuses a steer until that turn starts. * fix(codex): wait at most once for an answered turn Codex never opens A Codex before 0.148 can answer a send with a turn it then fails before starting, reporting only an `error` and no turn end. That turn stayed the answered-but-unopened turn for the rest of the session, so every later send made while the chat was idle, and every Stop naming no turn, waited the full open-wait first. When a wait ends without the turn opening, the dispatch correlation now records it, and later lookups skip it. A turn that opens later is still found through turn/started. The runtime test for a send made before the first turn opens now waits on a signal that the send is inside the open-wait instead of a fixed sleep, and pins that the send was steered. * test(codex): prove a legacy mid-turn send settles on Stop, and stop faking a steer Adds the user-visible outcome on a Codex before 0.148: a send made while a turn runs is withdrawn when a Stop ends that turn, the chat no longer owes work, and /compact is admitted after. The shared fake Codex servers no longer answer an unrouted turn/steer as a success; they refuse it as a Codex without that method would. Adds the case where Codex refuses both the steer and the fallback start, so the send is rejected in Codex's words and disarmed. * test(codex): type the fake Codex connection and wait recorder instead of casting |
||
|
|
e10b8357be |
fix(editor): honor the Markdown Review Notes setting (#24057)
Turning off Settings > Editor > Markdown Review Notes now hides markdown review notes everywhere: the controls in preview, rich, and source modes; existing notes in the rich editor; and markdown notes in the Source Control notes list and the diff note menus (count, copy, send). Clearing notes there keeps hidden markdown notes. Fixes #23966. Co-authored-by: yi111 <153097222+Yi-111-a@users.noreply.github.com> Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> |
||
|
|
1762a138f7 |
feat(mobile): slide the page's host stack on push and Back (#24268)
expo-router's Stack on web renders native-stack's web view, which flips display and ignores animation. The page's host stack now keeps expo-router's StackRouter under its public Navigator and draws the slide with the Web Animations API; a popped screen stays mounted until it has slid out. Native is a pure move. |
||
|
|
069eaf5668 |
fix(mobile): iOS shell keeps WebKit text interaction on so page fields take text (#24270)
With isTextInteractionEnabled = false (#21589) WebKit delivered keydown to a focused page field but never inserted text. The page's user-select rule (#24277) now keeps WebKit's selection off long presses. Needs an iOS shell release. |
||
|
|
e9ec63168f |
fix(mobile): hold-to-dictate, repeat keys and the browser long-press survive the page's long-press (#24277)
On the OTA page a held press died ~500 ms in: the WebView's long-press selected nearby text and that selection's selectionchange/touchcancel ended the press. Page text is now unselectable unless it opts in (as native), hold surfaces declare onLongPress, the browser pane refuses contextmenu termination, and the chat mic's swapped icons no longer steal the touch target. |
||
|
|
c0ac2b1fcc |
feat(browser): add a rebindable shortcut for Annotate page element (#23879)
Fixes #23470 Co-authored-by: fruit <200041037+guozi-lab@users.noreply.github.com> |
||
|
|
cb11a9ff1a |
Improve scroll indicator visibility with larger sizes and opacity (#24276)
* Improve scroll indicator visibility with larger sizes and opacity - Idle height: 2px → 3px, expanded: 3px → 4px - Increased background and thumb color opacity for better contrast * add test |
||
|
|
daf63e659c |
fix(runtime): read Antigravity, Cline and Prime Agent readiness from the live screen (#24222)
* fix(runtime): decide Antigravity readiness from the live screen agy paints its composer with cursor addressing, so the line-folded wait text misses the 1.2.14 accept-edits and plan composers and an ended turn, while the grid keeps the bare `>` caret painted mid-turn and behind the /model picker. Read the screen's bottom rows instead: rule, caret, rule, `? for shortcuts`. A clocked pane is held to quiescence (tier 1b) because the submit repaint reads ready for a moment; a clockless restored pane settles from the screen alone. When a trustworthy screen exists it decides, so the name-only title lane no longer settles an open picker. Retires the visible-read probe's Antigravity branch: the probe now runs the shared screen rule for any screen-ruled agent without an output clock, and keeps its generic empty-pane read for everyone else. Adds twelve agy 1.2.14 recordings and a replay suite shared by screen-ruled agents. STA-8741. * fix(runtime): decide Cline readiness from the live screen Cline paints its composer box with cursor addressing on the alternate screen, so no text rule saw it and worker-start timed out at agent_readiness (#23268). Read the box off the grid: rule, an empty composer with one of the captured placeholders, rule, the Plan/Act row and the auto-approve row, with no braille spinner above it. A streaming reply repaints the same box once its spinner has scrolled away, so Cline is tier 1b only: a clocked pane waits for quiet and a clockless one never settles from the screen. The screen now decides for a Cline pane, which shuts the quiet-process lane that would have settled its unworded tool-approval prompt and the Cline Desktop promo. readLiveTerminalScreenLines now returns raw rows: the read projection blanks a composer it takes for a draft, and it takes Cline's placeholder for one, so a typed draft and an empty composer looked the same. Adds nine cline 3.0.66 recordings (macOS) and the 3.0.65 Windows capture from #23269. STA-8741. * fix(runtime): decide Prime Agent readiness from the live screen Prime redraws its composer on the alternate screen, so the text tail never showed a settled prompt and tui-idle timed out (#22153). Read the grid instead: a bare `>` directly over the `<- manage` footer, with no braille status row (`Writing - 6s`) above it. The footer and caret alone stay painted for a whole turn. Replayed chunk by chunk, Prime erases that status row before redrawing it, and on first launch paints the idle composer just before the trace-sharing question covers it. Both keep repainting, so a clocked pane is held to quiescence (tier 1b); a clockless restored pane settles from the screen alone. Adds nine prime-agent 0.9.8 recordings (isolated HOME, OpenRouter) and the two 0.9.5 captures from #22154. STA-8741. * refactor(runtime): drop Cline-only readiness branches Cline now follows the same pattern as Antigravity and Prime: a screen rule plus table entries. - Drop MID_TURN_COMPOSER_AGENTS. onPtyData stamps lastOutputAt on every chunk, so a re-attached streaming pane has an output clock from its first byte; the exception only guarded a pane that printed nothing since attach. A clockless Cline pane now settles from its screen like the other two. - Drop the 'ready-body' rest-signal entries for all three agents. The rest signal is read only by quietForegroundLane, and a readable screen already shuts that lane and the title lane (isReadinessDecidedByScreen), so the entries only removed the quiet-process fallback for a pane with no trustworthy grid. The census now checks that screen-shut instead. - Drop the Cline rule's auto-approve row check; no recorded verdict depends on it. Kept: raw rows from readLiveTerminalScreenLines. Every frame of every codex-* and qoder-* capture at 120x40, 80x24 and 100x32 gives the same isKnownReadyPromptBody (with and without a clock) and isQuietReadyScreenBody verdict through both readers. Serializer known-failures for the new captures are pre-existing serializer behaviour, not this branch: row-0 cells restore with a true-colour background where the source has the default (the DSH class), and Prime's cursor restores at column 119 instead of the pending wrap at 120 (the qoder class). STA-8741. * fix(runtime): trust a screen rule only on the PTY's own grid Review findings on the screen-ruled readiness (STA-8741): - A grid out of step with the PTY garbles cursor-addressed chrome, and a model resize does not make the TUI repaint. readLiveTerminalScreenLines now returns null unless the emulator's grid matches the PTY's reported size and was never reflowed without a repaint (a re-attach that learned the real size late), so the pre-existing lanes decide there instead of timing out. - The visible-read probe reads the draft-blanking projection, which turns Cline's `❯ Ask anything...` into a bare `❯`. It now restores the blanked composer row before the rule reads it; `terminal read --screen` output is unchanged. - The quiet lane no longer ORs the text rules over a trustworthy screen that refused; without one, tier 1 already ran them. No recorded verdict changes. Tests: ready recordings on a mismatched and on a reflowed grid settle through the old lanes; the restored-pane probe runs every ready recording through the real projection; the rest-signal census checks the lane verdict with and without a screen. * test(runtime): trim STA-8741 recordings to the screens they prove * refactor(runtime): one screen verdict for every screen-ruled lane readScreenRuledReady, readScreenRuledQuietReady and isReadinessDecidedByScreen each re-derived the same thing: the agent's rule applied to a trustworthy live screen. They collapse into readScreenRuledVerdict (true / false / null), which tier 1, tier 1b and the lane gate read. This also makes a refusal final in tier 1: a clockless pane whose trustworthy screen refused fell through to the text rules, so retained ready text could settle over an open picker (Greptile review). The quiet tier already refused there; now both do. The tier-1b agent set derives the screen-ruled agents from the rule table instead of listing them again, and the lane test that repeated the census case is dropped. * refactor(runtime): let the visible-read probe read its own output clock The probe's clock was captured at start and threaded through the wait dependencies as a one-off parameter. The probe now reads it from the live record when its screen read returns, which is also the fresher answer. * fix(runtime): trust a reflowed grid again once a PTY resize repaints it The reattach-reflow flag was never cleared, so a pane stayed on the old lanes for the rest of its life even after a real resize made the TUI repaint (Greptile review). The record now keeps the reflowed grid, and a PTY resize off that grid clears it; an echo of the same size sends no SIGWINCH and keeps it. Tests: the reflow case in every screen-ruled suite now includes a same-size echo, and an Antigravity recording only the screen reads ready settles after a resize and repaint. * refactor(runtime): keep screen-rule trust and raw rows to screen-ruled agents Two shared changes reached agents this PR does not target: the live screen reader returned raw rows, and it refused a grid that did not match the PTY. Both now live in readScreenRuledLines, which only the screen-ruled agents read (screenReader picks it from the rule table); readLiveTerminalScreenLines is main's again. The probe keeps main's Antigravity-banner trigger, so a Codex or unknown pane is probed exactly as before. Proof: the non-screen-ruled suites give identical pass sets on this branch and its base (1,781 tests), and replaying every other recording frame by frame through the readiness and blocked verdicts, for its agent and for an unknown pane, gives identical results (93 pairs). A new test keeps a Codex pane reading its screen when the PTY reports another grid; it fails if the trust check moves back into the shared reader. |
||
|
|
a4606ccae3 |
fix(cli): orca file open no longer moves your view unless you pass --focus (#24244)
* docs(cli): file open/diff/open-changed say they switch the user's view and are for user requests only Refs #9944 * fix(cli): file open/diff/open-changed leave the user's view alone unless --focus `orca file open`, `file diff` and `file open-changed` always switched the desktop to the target worktree, selected the tab and revealed it in the sidebar. An agent skill that opens its answer pulled the user out of whatever they were typing in (#9944), and a phone opening a file moved the desktop too. The commands now add the tab in its worktree without changing anything on screen, including when that worktree is the one being viewed: the new tab is added to the tab bar but the active tab, tab type and focus stay put. In a worktree the user is not viewing, the tab becomes that worktree's selection so it is in front when they go there. `--focus` keeps today's behavior. files.open / files.openDiff take an optional `navigation` target (the existing RUNTIME_NAVIGATION_TARGETS vocabulary); the CLI sends 'all' for --focus, like `worktree create --activate`, and nothing otherwise. The renderer moves the host view only when the target reaches the host; a missing field (phones, older CLIs) leaves it still. Editor opens for a worktree other than the on-screen one no longer write the global activeFileId/activeTabType. Refs #9944 * test(cli): justify the window and runtime stubs in the file-open notification test * fix(cli): keep phone file opens switching the desktop; the CLI asks for 'caller' Phone opens send no `navigation` field, and the phone's diff-review "Open in session" relies on the desktop selecting the diff it opened. A missing field now keeps the original switch exactly; the CLI says what it wants instead: 'caller' (no host move) by default and 'all' for --focus. Older CLIs, which send nothing, keep switching as they always have. Refs #9944 * fix(cli): background file opens select the tab without counting as a visit A CLI open into a worktree the user is not viewing selected the new tab with the same activation a user click uses, which stamps lastFocusedAt and the group's recency list. The worktree jump palette sorts recent tabs by that time, so every agent `orca file open` into another worktree jumped to the top of the user's recent tabs. Editor opens now take a selection mode: 'focus' (default, unchanged), 'background' (select within its worktree without recording focus or recency) and 'none' (add only). createUnifiedTab and activateTab gain recordFocus:false for the background case. Also: tests for reopening an already-open file or diff without --focus, a comment that file opens move only the host window ('all' acts as 'host'), root help lines back under 100 columns, and an accurate remote test title. Refs #9944 * fix(tabs): a background-selected tab still joins its group's tab history recordFocus:false skipped both the focus-time stamp and the group's recentTabIds append while still making the tab the group's active tab. Ctrl+Tab looks the active tab up in that history, so after a background CLI open it did nothing (or went to the wrong tab) once the user switched to that worktree, and hydrate kept the broken history across a restart. Only the focus-time stamp is skipped now; the jump palette's recent rows sort by that alone, so the palette fix stands. Refs #9944 * fix(cli): file open/diff/open-changed --focus help says it brings the user to the file The three commands borrowed the shared --focus line written for terminal create ("Reveal the created terminal session in Orca"). They now use the per-command flag help table; terminal create's line is unchanged. Refs #9944 |
||
|
|
78daf71268 |
fix(status-bar): show Antigravity's model-group pools instead of an empty segment (#24074)
The verbose bucket allowlist was written for Gemini's experimental models, and its fallback window was session-or-monthly. Antigravity reports one pool per model group and some tiers meter weekly only, so a signed-in account with an exhausted pool rendered an icon and no number. Antigravity bypasses the allowlist by provider, because its group names come from the account's tier and cannot be enumerated. Cursor stays name-matched so an unrecognised pool still falls back to the plan total. Refs #22511 Refs #16704 |
||
|
|
a76a0bcddc |
fix(rate-limits): read real Antigravity quota from the agy CLI, not the Gemini mirror (#24073)
* fix(rate-limits): read real Antigravity quota from the agy CLI Orca published a successful Gemini `retrieveUserQuota` read under the antigravity provider id. That reported Gemini CLI per-model buckets on a 60-minute window, so Antigravity's real pools were never shown, the weekly window was always null, and the segment depended on an installed @google/gemini-cli for token refresh that an Antigravity user has no reason to have. Read the quota from `agy -p "/usage" --output-format json` instead, which is the only caller that can authenticate it — agy keeps its credential in the OS keyring and mints its own token against daily-cloudcode-pa. Fixes #9122 Fixes #22511 * test(rate-limits): stub the Antigravity CLI fetch in every service suite Without the stub, each RateLimitService suite spawned the developer's real `agy` and resolved a login shell, which turned service-window-activation from 214 ms into 14 s and broke its fake-timer fetch counts. * fix(rate-limits): never pass --disable-slash-commands to the agy quota read The flag stops agy expanding `/usage` as a command, so the text goes to the model as an ordinary prompt: the call starts a conversation, spends quota, and on an account near its limit answers RESOURCE_EXHAUSTED (429) instead of a reading. Adds an opt-in real-CLI suite that catches exactly this. * fix(rate-limits): stop polling agy once it answers /usage as a prompt In print mode an unrecognised slash command is not an error — agy sends the text to the model. On a build that does not know `/usage`, polling would start a conversation and spend the user's quota every cycle while Orca reported no quota. The envelope distinguishes the two: a command reply has an empty conversation_id and num_turns 0. A successful parse is checked first, so a real reading can never trip the latch. |
||
|
|
7c119465b0 |
fix(relay): define restart-safe by the cell runtime, and refuse waves without headroom (#24259)
* fix(relay): let a same-cap drain finish when only unplaceable hosts remain The c28 canary on 2026-10-01 drained the cell to zero live connections, but four hosts with no free slot anywhere kept redialling and held director leases on it, so the restart-safe wait timed out and left the cell isolated and empty. The drain wait now also passes once the runtime has carried nothing for a sustained quiet window while a small, capped number of leases remain, and logs the escape. Apply modes also refuse a cell whose hosts exceed 80% of the free slots on the other general cells, so a wave cannot strand hosts in the first place. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): define restart-safe by the cell runtime, not director leases Replaces the opt-in stranded-host escape with a corrected definition. A restart is safe when the cell runtime carries nothing live and no migration is open, sustained for the drain pace window. Director activity leases lag hosts that already left or cannot be placed, so they are reported in a progress line and the verified result instead of blocking the restart. The same-cap drain passes its existing pace window. The headroom script is added to the trusted evidence code paths with the other production scripts. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): print stranded director counts on every restart-safe sample Each restart-safe poll now prints its sample count and the director's restart-blocking leases, request units, reserved remainder, and migrations under `stranded`; the verified line carries the same object. Open migrations still block because each is pinned to the cell incarnation a restart replaces. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): require the pace window for every live restart-safe wait Pre-auth and total connections no longer reset the restart-safe window: on drained c28 they flickered with unauthenticated redials in a third of samples, which a restart does not lose. They stay in the progress output. Every live restart-safe call must now pass --pace-window-ms. The capacity job and staging proof drain unpaced, so they pass the production 300000 ms window, and the calls that relied on the 180000 ms default get 480000 ms. Headroom free slots now follow the director's placement rule: the admission pause minus the larger of observed and enforced units, minus outstanding control reservations. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
1b0ee969bc |
fix(browser): stop the first browser command on a new Windows tab hanging until it fails (#24237)
* fix(browser): stop slow browser commands failing with "runtime closed the connection" On Windows, the first helper-backed browser command on a fresh tab (snapshot, screenshot, wait, get, capture start, ...) produced nothing for 30 s and then failed with runtime_unavailable and "Restart Orca", which never helped. Two defects combined: - The local RPC transport destroyed any connection that had written nothing for 30 s, including one still waiting on our reply. Only long-poll methods sent keepalives, so every slower unary RPC (browser helper commands allow 90 s, emulator 180 s, skill share 10 min) was cut off and the CLI reported the close as a dead runtime. The idle timer is now suspended while a dispatch is in flight and re-armed when none remain; the client's own deadline bounds the wait. No wire change. - On Windows, agent-browser's daemon inherits the helper's stdio pipes (vercel-labs/agent-browser#1407), so EOF never arrives and execFile waited for the 90 s timeout even though the helper had printed its result and exited in ~300 ms. The bridge now releases the pipes shortly after the helper exits, so the call settles with its full output. Every new tab and every daemon restart after the 10-minute idle timeout hit this. Fixes #23911 Fixes #18093 * test(browser): remove the raw-process test's temp dir; fix stale keepalive comment * revert(runtime): move the RPC idle-timer change to its own branch The browser fix alone resolves the ticket: with the pipe release, a fresh-tab snapshot takes ~1.4 s, well under the 30 s idle window. Whether the server should stop cutting off slow in-flight RPCs is a separate trade-off against fail-fast on genuine hangs, now on sta-8934-rpc-idle-timeout. |
||
|
|
12b8ef8c0b |
fix(worktree): update local main safely, once per branch, alongside the checkout (#23698)
* fix(worktree): retry local main refresh through git lock contention and skip false alarms * fix(worktree): overlap the local main refresh with the checkout and run one refresh per repo at a time * fix(worktree): skip the local base refresh when the create makes that branch itself Creating a workspace named feature-x from origin/feature-x runs `worktree add -b feature-x`, which now overlaps the refresh. The refresh's drift probe could see refs/heads/feature-x missing and its presence probe then see it (the add just wrote it), which reported "not fast-forward" and showed a sticky "Local feature-x was not refreshed" warning. `-b` refuses an existing branch, so there is nothing to refresh in that case: skip it on the local, prepared-checkout and SSH create paths. The SSH overlap tests move to their own file so the existing suite stays under the line limit. * fix(worktree): say plainly what happens after the local base refresh queue wait expires * test(worktree): prove SSH local base refreshes of one repo run one at a time * test(worktree): drop type assertions from the SSH refresh overlap test mocks * fix(worktree): fast-forward local main with one host-owned merge --ff-only per branch Moves the whole local base refresh into one shared routine that runs on the execution host (main process for local and WSL repos, the relay for SSH), so the app no longer keeps a second copy of the checks, queue and retry. A checked-out branch now moves with merge --ff-only (hooks, auto-gc and autostash off) instead of status-then-reset --hard, which silently overwrote an untracked file the new commit adds and could discard an edit or a commit made after the check. A free branch moves with a compare-and-swap update-ref that writes a reflog message. Status reads no longer take index.lock. Creates of one branch share one run plus at most one trailing run; a create waits at most 30 s and never starts a competing mutation. The failure toast is keyed by repo and branch because every create that joined a run reports the same fact. * fix(worktree): fast-forward local main even when the repo requires signed merges With merge.verifySignatures=true, the owner-checkout fast-forward refused an unsigned origin/main tip, so every create warned "Local main was not refreshed" where the old reset moved main. The new workspace is already created from that same unsigned commit, and a branch that is not checked out moves without a signature check, so the refusal protected nothing. Turn the setting off for this one merge, like the hooks, gc and autostash overrides. * fix(worktree): clear git read caches when a shared local main update lands late The update of local main can finish after a create stopped waiting for it, so the shared run now invalidates git read caches itself. The index.lock real-git test also no longer reads the developer's global git config. * fix(worktree): never overwrite an ignored file when fast-forwarding local main A plain `git merge --ff-only` silently replaces an ignored file (for example a local `.env`) at a path the new commit starts tracking. Pass `--no-overwrite-ignore` so git refuses instead and the create reports the checkout as having local changes. Supported on the fast-forward path since well before Git 2.25. Also make the relay test for one-refresh-per-branch hold the first merge until the second request has reached the relay, so it fails without the coalescing. * fix(worktree): keep the local main update a plain fast-forward whatever the user's merge settings say A per-branch mergeOptions such as '-s ours' or '--squash', or pull.twohead=ours, made the update create a merge commit that dropped upstream, or stage upstream without moving main, while reporting success. The command now clears the branch's mergeOptions and passes the strategy and signature choice on the command line, which beats any config. After the move Orca confirms local main is exactly the target before reporting it updated. The exact command also runs in the Git 2.25 compatibility suite. * fix(worktree): make the Git 2.25 fast-forward contract pass in CI and rerun on every change to it The new real-Git contract for the local main fast-forward wrote a post-merge hook into .git/hooks, which does not exist when the repo is created by the uninstalled Git 2.25.5 build CI uses (no templates), so the Git compatibility check failed. Create the directory first. The Git compatibility check also did not run when only the fast-forward module changed, so a later edit to its merge arguments (for example a flag Git 2.25 lacks) would skip the one check that tests them. Add the module to the check's paths. * fix(worktree): answer every create from a local main update toward its own base Creates from different remotes' main (origin/main and upstream/main) shared one queued update per repo and branch, which ran only the latest caller's target: a create could get no result for its own base, or a false "not refreshed" warning computed for another remote's main. The per-branch runner now queues one run per distinct target, still one at a time per branch, and only callers toward the same target share a queued run. Applied in the app and the relay. * test(worktree): record the third create's result in the mixed-remote burst tests and update the toast id rationale * chore(worktree): correct the toast id rationale |
||
|
|
0b79720c2e |
feat(native-chat): the chat strip and the sidebar read the host's child records (#22614)
* feat(native-chat): publish the host's child records to the status summary and the chat strip
The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.
* feat(native-chat): the chat strip reads the host's child records with its parent's verdict
The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.
Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.
* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open
- The view decoder ignores unknown keys, degrades unknown kinds, states,
outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
roster of finished children and never the views themselves; a stop-only
reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.
* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered
* test(native-chat): type the switch tests' mocks instead of asserting them
* test: remote clients advertise reading child views
* docs(agent-status): the structured row folds the store's child records
* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary
The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.
* refactor(native-chat): the status summary's broadcast equality gets its own module
The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.
* fix(native-chat): command admission reads the strip's child records
A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.
Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.
* refactor(native-chat): command admission takes only what it reads of a turn
* fix(native-chat): the session list drops a session's children when the store does
A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.
The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.
* test(native-chat): write the Codex frame script's parent row out step by step
Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.
* fix(native-chat): the idle sweep and the restart snapshot read the host's child records
The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.
The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.
* test(native-chat): the child-record tests follow the merged command lifecycle
A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.
Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.
* refactor(native-chat): the status feed's journal projection cache gets its own module
The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.
* test(native-chat): the admission test's compaction resolves with a real outcome
Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.
* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished
The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.
This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.
* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source
`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.
A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.
* test(native-chat): the switch test passes the startup child key main's status bar takes
* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own
Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.
Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.
* fix(native-chat): a background Stop reaches the tasks the child records show
The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.
The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.
* fix(native-chat): one rule for a finished child that still owns live work, at any depth
The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.
* fix(native-chat): an older client sees a Codex child's shell as it did before views
Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.
* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives
The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.
* fix(native-chat): the strip channel forgets a closed conversation's roster
It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.
* docs(native-chat): rewrap the retention comment
* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays
Two lifecycle gaps from the round-1 fixes.
A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.
A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.
Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.
* fix(native-chat): the strip keeps one empty list for a roster that omits one
A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.
* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent
The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.
* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once
A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.
The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.
Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.
* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader
CI on
|
||
|
|
24540300f0 |
fix(agent-status): preserve hook presence when process checks cannot answer (step 1 of 3) (#23947)
* fix(agent-status): admit hook process presence on the execution host * fix(agent-status): restrict process checks to real hook ingress * fix(agent-status): keep presence checks from causing false exits or losing real ones - Pin the macOS process start time to UTC on both the hook and the host so a shell TZ or a time-zone change cannot turn a live Claude into an exit. - An unanswered process check falls back to the foreground confirmation, so Codex, SSH and Windows panes still leave the agent state on a real exit and the Codex late-completion recovery still runs. - A nested agent that inherits the pane key cannot take over the pane's presence; a retired session is replaced by the next session even when its SessionStart was lost. - Drop the unused Windows process read (no Windows hook captures an identity yet) so shared code no longer imports main-process modules. - Skip the capture outside Orca panes, gate relay re-checks to real title changes, and list the capture module in the CLI project. * refactor(agent-status): own pane presence by the agent's process, not its session Presence now exists only when a hook carries the agent's process identity, and only that process's evidence changes it: its SessionEnd ends the pane, its /clear and /resume keep it, and hooks from any other process (a nested agent inheriting the pane key) update status without taking ownership. Hooks without an identity (Windows, sessions started before the capture) behave exactly as before, so a nested agent can no longer end a pane it does not own. An ended owner stops answering 'exited', and the host probes the owner only when another process reports in the pane. * fix(agent-status): close round-3 review gaps in presence handling - Relay retries and transcript polls schedule against the row the relay cached, and identity-less events pass through the transition unchanged, so SSH Grok replies and Codex transcript polls deliver again. - A live process check restores the runtime's agent status and releases queued orchestration mail, like a foreground read that finds the agent. - An answered foreground read naming a non-agent (wsl.exe, tmux) is still an exit; only silence is not. - A suspended (Ctrl-Z) agent is unverifiable, not live, so the foreground read decides as before. - An agent of another type started mid-turn cannot own or end the pane. - Replayed spool hooks check each pane once, and SessionEnd ends presence only for reasons that end the process. * refactor(agent-status): record the pane's owning agent even without a process id Ownership is now decided only by comparing the recorded owner with the sender, never from the row's agent type or turn state (which identity resolution rewrites). The first agent hook in an empty pane claims it; a live owner keeps it against any other agent (a nested claude -p, a Codex started inside Claude or the reverse); only the owner's own proven process can end it. An owner no hook identified never ends from a hook and is not probed, so those panes behave as before. * test(agent-status): read the optional process id in the relay presence test * fix(agent-status): other agents' SessionEnd hooks settle their status again Only an admitted exit (Claude's process-ending SessionEnd, or a host-proved exit) is marked ended on the event, and ownership keys on that marker, not the hook name. Devin, Qoder, CodeBuddy and Copilot SessionEnd hooks are ordinary status updates again, locally and through the relay. * test(agent-status): cover the owner's own SessionEnd through the relay * fix(agent-status): rows without a process identity keep today's command-finished cleanup The renderer's command-finished cleanup kept every row when its shell check could not answer, which left Codex, hookless-agent and old-relay rows over SSH showing done after a real exit. It now asks the host whether the pane's agent process can be checked: only a pane with an identified, running owner keeps its row on an unanswered check; every other pane drops it exactly as before. The drop stays armed while the host answers, so a new command still cancels it, and a missing or failing answer (web client, older host) keeps today's behaviour. * fix(agent-status): panes without an identified owner keep today's exit confirmation The title-driven exit confirmation applied 'silence is never an exit' to every pane. It now applies only when the pane has an identified owner whose process cannot be checked right now; a pane with no process identity (Codex, hookless agents, Windows, old relays, sessions started before the update, or an owner that already ended) confirms exits exactly as before. |
||
|
|
2ab179ccf7 |
fix(emulator): take serve-sim 0.1.47 so the iOS simulator works on Xcode 27 (#24228)
* fix(emulator): take serve-sim 0.1.47 so the iOS simulator works on Xcode 27 serve-sim 0.1.40's helper binary hard-linked SimulatorKit at its pre-Xcode 27 path, so every iOS emulator start failed with a dyld error on Xcode 27. 0.1.47 replaces that binary with a napi addon that finds SimulatorKit in either location, and runs the stream helper as a node process. Match the new helper process (serve-sim.js ... --exit-on-simulator-shutdown) while still matching legacy serve-sim-bin helpers left by an older Orca, and mark the package's new standalone helpers executable. * refactor(emulator): drop redundant serve-sim chmod; the package already ships helpers executable serve-sim publishes its helpers as 755 and pnpm, fs.cpSync and electron-builder all preserve the mode, so the executables list and every chmod over it are dead. * fix(emulator): give the materialized serve-sim runtime its node dependencies serve-sim 0.1.47's entry imports `ws`. Packaged macOS builds run serve-sim from a copy under userData, and that copy had no node_modules, so every iOS emulator attach failed with ERR_MODULE_NOT_FOUND. Dev builds were unaffected because they run serve-sim straight from pnpm's node_modules. Lay the copy out as <version>/node_modules/serve-sim and link each dependency the package declares to the bundle's installed sibling, so transitive deps keep resolving from the bundle. A runtime in the old flat layout, or one whose links dangle because the app moved, is rebuilt instead of reused. |
||
|
|
6729f1b8e2 |
fix(notes): send AI notes to open structured chat sessions (#24221)
* fix(notes): send AI notes to open structured chat sessions The notes "Send notes to" menu, browser annotations and the sidebar send targets only listed terminal agents, so an open structured chat was never offered. Send targets now include structured chats, and one sendMessageToAgent delivers to either kind. A chat receives the message through its own outbox, the same enqueue its composer uses, so the note shows in the chat and a failed send keeps Retry. To make that possible the desktop outbox state moves out of the chat view's React hook into a per-session store the view subscribes to; the launch-prompt settlement, the composer and outside senders all write that one copy. Also settles the notes loading toast in place, so an instant send can no longer leave "Sending notes..." stuck on screen. * fix(notes): settle failed sends, keep unsaved launch prompts queued, limit sidebar targets - A failed notes send settled the "Sending notes..." toast with toast.message, which keeps sonner's loading type: the spinner never cleared and the toast could not be closed. Settle it with toast.info on the same id. - A launch prompt whose staging save failed stayed "dispatching" in the open chat, so it was never sent. Staging is now a required save; on failure the entry stays queued and the chat sends it. - A chat before its first turn has no sidebar row, but counted as a sidebar send target, so opening the notes menu revealed and highlighted an empty card. Only the notes menu lists such chats now. * fix(notes): show an agent's failed or stopped verdict in the notes menu The notes menu picked a row's dot and state word from the plain row state, while the sidebar row uses agentRowDisplayDotState, which shows the main agent's verdict first. A chat whose turn failed read "Done" with a green check in the menu and "Failed" with a red dot in the sidebar. The menu now uses the sidebar's function; terminal rows with a verdict get the same correction. |
||
|
|
3ab3c9239f |
fix(native-chat): the working line shows only what the agent is doing now (#24218)
* fix(native-chat): the working line shows only what the agent is doing now A chat's live "Working…" line could show an old notice, such as "Claude hit a temporary problem and is retrying.", long after the agent had moved on and was running new commands. When the host had no live activity for the turn, the line fell back to the newest status row in the turn, and any status row qualified: retry warnings, "Context compacted", "Cancellation requested.", and notification summaries. Those rows record the past and already appear in the transcript. The line now reads only the host's live, per-turn activity, which is never saved and is cleared at turn boundaries. Without it, the line says Thinking or Working…. Desktop and mobile share the selector, so both change. * test(codex): guard that a subagent's compaction never becomes the parent's live activity |
||
|
|
6f2a7d05c9 |
fix(worktrees): let git delete removed checkouts so chat sends never wait behind them (#23837)
* fix(worktrees): delete removed checkouts in git, not in Orca's file pool Local worktree removal renamed the checkout into a sibling trash root and deleted it in the background with a recursive fs.rm in the main process. That queued one request per entry on libuv's shared 4-thread file pool, so for minutes every other async fs call in the main process (the agent-session store behind chat sends, file explorer reads) waited behind the delete. `git worktree remove` now deletes the checkout inline in git's own process again, so the card stays in its Deleting state for the length of the delete while Orca's file pool stays free. No timeout applies to the call, so a large delete is never killed halfway. If git reports success but the path still exists (Git for Windows leaves junctions and their parent directories in place), the leftover is deleted with the existing removeHostTree; WSL checkouts stay with the distro. Nothing creates trash any more: the scheduling queue, rename/restore helpers and the trash_rename span are gone. The startup sweep stays to drain entries older releases left behind, and now removes each emptied trash root so the obligation ends. * fix(worktrees): let Git delete Windows checkouts with long paths enabled Removal now always runs Git's own recursive delete, and worktree creation checks out with core.longpaths on Windows, so a deep checkout Orca created could fail to delete with "Filename too long" (#6433). The Windows recovery then finishes the delete but keeps the branch. Pass the same command-scoped core.longpaths option to `git worktree remove` so Git can delete what it created. Also point the CI shard timing entry at the renamed real-git removal suite. * fix(worktrees): keep an inherited GIT_ASK_YESNO out of the worktree delete Git for Windows asks $GIT_ASK_YESNO whether to retry when a file stays locked during a recursive delete. Orca's git env inherits the user's environment, so an inherited value would run an arbitrary prompt program in the middle of a removal. Drop it for the removal call only. * perf(worktrees): run worktree deletes under their own limit, outside git admission `git worktree remove` now deletes the whole checkout in Git's own process, which takes 20-35 s on a large tree. It took a general git admission slot at status tier for that whole time, and that cap is as small as two slots on a machine with six or fewer cores, so two deletes blocked every status read. Deletes now skip general admission and queue under their own limit of two per host instead: two concurrent deletes already saturate one disk, and more only slow each other down. Leftover cleanup runs inside the same slot. * fix(worktrees): delete removed checkouts in the background and mark them removing Since the checkout is deleted by `git worktree remove` in Git's own process, a large delete takes 20-35 s. Answering the request only after that made web and mobile (30 s), paired desktop (60/180 s) and the CLI (60 s) report a failure for a delete that was still going, and mobile silently re-showed the row. The request now does everything that can refuse (lock, cleanliness, archive hook, watcher/terminal gate, terminal stop, shared-link unlink), records the removal in an in-memory table on the host and answers `removing: true`. The delete, branch cleanup and metadata purge run after it in the same order as before, and the watcher/terminal gate stays held until they finish. - Listings mark rows in the table `removing` for clients that advertise `worktree.background-removal.v1` (the desktop renderer, paired desktop and web), and leave them out for everyone else (older clients, mobile, the CLI), which already dropped the row when the request answered. - The outcome (removed, with any preserved branch, or the error) rides the existing worktrees-changed event as an optional field, sent after the row has left the table. - A repeat delete while Git runs joins it. A create at the same path or with the same branch is refused with "Cleanup is pending; try again shortly"; create's name search skips the path, so generated names move on. - Nothing is persisted: after a quit or crash Git still lists the checkout and it can be deleted again. WSL checkouts still delete inline. - `orca worktree rm` says the checkout is still being deleted. * fix(worktrees): keep the existing Deleting card until the host's Git finishes The host now answers a local worktree delete on acceptance and deletes in the background. The renderer keeps the existing delete state set until the host publishes how it ended: - The delete that asked waits for the outcome on the worktrees-changed event (local IPC or the paired runtime's client event), then runs the same teardown, preserved-branch toast and card error an inline delete did. If that event is lost to a dropped connection, a listing that shows the row gone after it was marked removing finishes the wait, and one that shows it back without the marker fails it. - Any other renderer (a reload, a paired desktop, web) sets the same delete state from the host's `removing` marker and clears it when the marker goes. A failure the host publishes lands on that card's existing error. - Web advertises `worktree.background-removal.v1` so the host sends it the marker; paired desktop does through the Electron capability list. No new component, style or state: the card reads the delete state it always did. A host that predates this answers when done without `removing`, and the renderer takes that as finished, as before. * test(worktrees): type the removal harness and projection for the node typecheck * fix(worktrees): don't fail a delete retry with an earlier attempt's buffered failure A background removal's outcome that reached this renderer with no waiter (another client's delete, a host-marked card, or one already settled from listings) was buffered for 60 s and consumed by the next delete of the same workspace, so retrying a failed delete failed at once with the old error while the host was deleting. Drop the buffered outcome before sending the request; only an outcome that arrives after it can belong to it. * fix(worktrees): let only a gap in host events settle a background delete from listings Git unlists the checkout before the host deletes the branch, cleans the push target and purges metadata, and the worktree-directory watcher refetches within 250 ms. The renderer read the missing row as a finished delete, so the waiter resolved without the preserved branch (no toast) and a failure in those last steps showed as success; the real outcome was then dropped. The listing fallback exists only for a lost outcome event, so it now applies only after this host's event stream had a gap: a new subscription or a replay after reconnect. * perf(worktrees): let a bulk delete start each same-repo checkout delete once the host accepts the last A bulk delete ran one worktree at a time per repo (#2259, for packed-refs and ref-lock races in branch cleanup). With Git now deleting each checkout for 20-35 s before the request settles, N worktrees in one repo took N times that. The renderer now queues same-repo deletes only until the host accepts each one; a parent still waits for its nested children to finish. The host serializes the branch cleanup step per repo itself, which also covers removals started by different clients. * test(worktrees): pin the host platform in the mocked removal suites so they pass on Windows Removal now passes -c core.longpaths=true on Windows, so the exact-argv assertions and command-keyed mocks never matched there (17 failures on a Windows host). Pin darwin as the add-worktree suites already do, and drive the one Windows-specific case through the same spy. * test(worktrees): type the blocked git remove result instead of a broad object The anti-slop static-analysis gate rejects `object` parameters. * test(worktrees): clear the changed-code quality gate in the removal suites Merge the duplicate node:fs import, build the mock child without a cast, read worktrees:list rows through one typed helper, and give the remaining casts a SAFETY line. * fix(worktrees): record each background delete durably and finish it after a quit or crash A quit mid-delete left git to finish the checkout on its own while the branch delete and metadata purge never ran; a crash left a normal-looking row. Each accepted local removal now writes a record beside the profile state before git starts, clears it on success or failure, and the host runs the same delete again for any record left at startup, re-deriving what remains from git and disk. An orderly quit stops the checkout delete without waiting for it. * test(worktrees): type the interrupted-removal assertions for the node typecheck * fix(worktrees): finish an interrupted delete that already removed the checkout's .git file Quit stops git worktree remove mid-delete, and Git deletes the checkout's .git file wherever it falls in directory order. Git then refuses the checkout ("validation failed ... .git does not exist") on every retry, so the startup finish failed and the row could never be deleted from Orca. A registered checkout this record owns that has lost its .git file now finishes like an unregistered one: leftover files, prune, then the branch. * fix(worktrees): let Git finish an interrupted delete, and never take a different checkout A quit or crash that stops `git worktree remove` after it deleted the checkout's .git file left a registered checkout Git refuses to remove. The previous fix deleted that leftover inside Orca's process, which is the bulk delete this change exists to avoid (and on Windows the leftover can be most of the checkout). The startup finish now rewrites the missing .git file from Git's own admin entry for that path and lets `git worktree remove --force` delete it. `git worktree repair` is not used: it also re-points every other registered path, including a checkout another repository now owns there. Orca deletes the leftover itself only when no admin entry claims the path. The startup finish forces, so it now leaves the path alone when the checkout there is not the one recorded: a registered worktree on a different branch or head, or a `.git` at a path Git already unregistered. The record is dropped and the card shows why. The record write before Git starts is now bounded (2 s, logged when exceeded) so a stalled disk cannot hold the delete, and the outcome is published before the record's clear reaches disk. * test(worktrees): compare worktree paths by value and tear down with Windows lock retries Git prints forward slashes in `git worktree list` on Windows, so the real-Git removal suites never found a joined path there: positive checks failed and negative ones passed without proving anything. They now compare Git's parsed rows by value. Teardown uses the shared retrying removeTree, since Windows can hold the deleted checkout busy for a moment after Git exits. Adds a relative-path worktree case for the .git restore (skipped before Git 2.48). * fix(worktrees): reply to a worktree delete when it has finished, not on a broadcast event A current client's delete request now waits for the host's background delete and gets its real result (removed, a preserved branch, or the error) as the reply, the way it did before the delete moved off the request. A request that arrives while the delete runs joins it and gets the same result. Every other view keeps reading the host's `removing` marker: the row leaving means the delete finished, and the row listed again without the marker shows "The delete did not finish. Try again." on a card that view had marked Deleting. A request whose reply is lost (a timeout or a dropped connection) settles the same way from a fresh listing instead of reporting a failure. Clients without the background-removal capability (mobile, the CLI, older desktops) are still answered on acceptance and have rows under removal left out of their listings. This removes the outcome on worktreesChanged and everything it needed: the renderer's outcome waiters, early-outcome buffer and TTL, per-host event-gap generations, the request pre-registration, and the accept callback bulk delete used. Bulk delete runs same-repo deletes in parallel only on this machine, whose host serializes branch cleanup per repo; SSH and paired hosts stay serialized. * test(worktrees): type the pending-removal host id in the background-removal suite * fix(worktrees): answer a delete request even when a concurrent removal of the same worktree replaced its record The desktop app's removal and the runtime removal (CLI, paired clients) coalesce separately, so both can be accepted for one worktree. The second replaced the first's record, and the first delete then finished without resolving the request waiting on it, leaving the desktop card on Deleting indefinitely. Each delete now settles the request it was started for. * fix(worktrees): run same-repo removal archive hooks and teardown one at a time on the host Local bulk delete now sends same-repo removals in parallel, so their archive hooks, terminal teardown and preflight ran at once; a hook that writes refs can race the repo's ref locks (#2259). The host now serializes each local removal up to acceptance per repo, for every client; Git's checkout delete still runs in parallel under the delete limit. * fix(runtime): keep waiting worktree deletes out of a host's foreground call slots worktree.rm now replies only after Git deletes the checkout (up to minutes), so on paired desktop and web each waiting delete held one of the host's 8 foreground call slots, and a bulk delete queued listing refreshes and every other foreground call behind it. Deletes now run in their own lane with the same bound; the 2-slot background lane stays for status polls. * fix(worktrees): join a same-worktree delete accepted while a removal waited its repo turn The desktop app and the runtime (CLI, paired clients, web) check for a running delete before they queue for the repo's acceptance turn. A delete of the same worktree from the other path, accepted while this one queued, was missed: this request re-ran the archive hook, stopped the terminals again and started a second `git worktree remove` on the directory Git was deleting. The queued acceptance now re-checks and joins the running delete. * fix(worktrees): fence a resumed delete's checkout from startup, and drop rows a listing read before the delete finished A delete a quit or crash interrupted took its terminal and file-watcher gate only when the resume job ran, after the first window was shown; session restore could open a shell or watcher inside the half-deleted checkout first, and on Windows that handle can fail the resumed git delete. Loading the records now fences each recorded path, and the resumed job takes the fence over in the same tick it takes its own gate. A listing that read git's registration before a delete finished, and replied after the removal record cleared, returned the row unmarked, so other views briefly showed "The delete did not finish". Listings now capture the pending removals before reading git and leave out a row whose delete finished successfully since; a row whose delete failed stays listed as before. * test(worktrees): keep git's auto-maintenance out of the real-git removal suite CI's Git 2.55 failed the file-pool test in teardown with ENOTEMPTY on the scratch repo's objects/pack after the test body passed: the 3,000-file commit's detached auto-maintenance was still writing a pack. The scratch repo now disables auto-maintenance and auto-gc. * fix(worktrees): one archive-hook approval covers a same-repo bulk delete again Local same-repo deletes now start together, so each queued its trust prompt with a state snapshot taken before the first prompt was answered; approving the first still showed the same prompt once per remaining worktree. The queued check now reads the store when its turn comes. |
||
|
|
b594898533 |
test: stop main failing on three tests the store and close-cause changes crossed (#24231)
* test: two tests catch up with the journal-database store and the close cause #24006 moved the session store into the chat journal database, and #23467 made close take a cause; #23684's and an older tab-table test still used the old calls, so main's typecheck fails. * test: the startup-reconcile close test passes the close cause too |
||
|
|
3727100cc9 |
fix(opencode): report OpenCode 2 status from each pane's own TUI (#23722)
* fix(opencode): report OpenCode 2 status from each pane's TUI OpenCode 2 serves every pane from one shared server whose env names only the pane that started it, so every same-folder pane's work showed on that pane, and the plugin's single aggregate swallowed the Idle of a pane whose turn ended while another pane was busy. Install the status plugin a second time as an OpenCode 2 TUI plugin (plugins/<name>-tui/tui.js, written only when its bytes differ, locally, in overlays and in the SSH/WSL relay installs). In a TUI process setup() runs the same engine, fed through the same event translation as the server, but only for root sessions this pane owns: the route's session from when it starts (or when the route reaches it while running, hydrated from the TUI's session status) until it settles, is deleted, or the TUI exits. The server plugin stands down in any serve process when the TUI copy is installed beside it; a relay that predates the TUI copy keeps the old behavior. OpenCode 1 loads no plugin directories and is unchanged. Delete the session-to-pane binder, registry, client sweep and ingest reattribution: with every post stamped by its own pane nothing is left for them to correct. Known gap: OpenCode 2 `opencode run` in a pane has no TUI, so it reports no pane status. * fix(opencode): settle missed OpenCode 2 turn ends and order TUI installs - The TUI copy now settles an owned root whose run the TUI's session data reports ended, with no pending permission or form, when the engine still holds it busy. An execution end missed across a service restart or reconnect no longer leaves the pane Working until the TUI exits. - The TUI copy stays idle without ORCA_PANE_KEY. post() cannot report without it, so OpenCode 2 TUIs outside Orca no longer run the route poll and engine. - Installers write the TUI copy before the server plugin file. The server decides at load whether to stand down, so a reload between the two writes now finds the TUI copy. * fix(opencode): reconcile OpenCode 2 TUI status only once queued events drain The TUI's session data applies each event before this plugin's queued handling reaches it, so settling against it mid-backlog published a false Done before a fast turn's later steps (Working, Done, Working, Done). Reconcile only when no event is queued; a mismatch then is a start or end missed across a reconnect, and both directions are now re-derived (a missed start left the pane on Done). Also install the TUI copy beside the server plugin in the retired shared hooks dir, so a TUI or service still loading that dir reports per pane instead of leaving the service reporting under its starter pane. * fix(opencode): keep OpenCode 2 TUI panes silent on plugin dispose A TUI plugin hot reload disposes the plugin while the pane's turn keeps running. The server path already passes sessionsOutliveDispose so dispose publishes nothing; the TUI adapter now does the same, or every TUI reload would still show a false Done. Also refreshes the generated-bytes digest after rebasing onto the write-if-changed installers. * fix(opencode): refresh installed OpenCode status plugins at app start After an Orca upgrade, an OpenCode 2 service that was already running kept the previous plugin, and with it the old wrong-pane status, until any new terminal pane rewrote the file. OpenCode 2 reloads a plugin whose file changes, so Orca now refreshes its existing installs once after the first window shows: the global config dir, source overlays and the retired shared dir, TUI copy first. It reuses the per-pane writers, which skip unchanged files, never creates an install the user did not have, and honours the status-hook and per-agent switches. An SSH relay does the same for its canonical install when Orca connects and ships the plugin sources. The plugin source assembly moves to its own module (re-exported unchanged) to keep hook-service.ts under the line limit. * fix(opencode): derive each OpenCode 2 pane's status from its TUI session data The TUI copy of the status plugin translated OpenCode events into the server-side engine and then patched the engine's latches back toward the TUI's own session data: a settle on every tick, a queued-event counter, a re-assert, synthetic Busy and Idle. Each review found another place where the two copies disagreed: a missed end left the pane Working, a backlog of slow posts flickered Done, a missed start showed Done mid-turn, a turn held only by a subagent never settled, and a request answered while disconnected pinned Needs input. The TUI reporter now reads the pane's level straight from the session data OpenCode keeps current (and re-hydrates on reconnect) on every event and on a 100 ms tick: for each root session this pane owns, Needs input (an open permission, else form, anywhere in its family while it runs) outranks Working (any family member running) outranks Done. It posts one status per level change through the plugin's existing delivery functions (retry, dedupe, message-part throttle, ordering), so the wire and Orca's ingest are unchanged. Ownership is kept in OpenCode's storage.memory, which survives a plugin hot reload, so a turn that ends during a reload still shows Done; dispose publishes nothing, and a level the old generation could not deliver is re-posted by the next. On reconnect it re-syncs blockers for the owned sessions OpenCode would not re-sync itself, ignoring requests already answered. Events are handled synchronously, so a slow post can no longer hold up event processing. The server path, OpenCode 1 and mimo are unchanged. * fix(opencode): keep an OpenCode 2 pane on Needs input while its root streams text Orca treats every OpenCode MessagePart as Working. While a background subagent waits on a permission or form, the root session can keep streaming reply text, and the TUI reporter forwarded that text as a MessagePart. The pane then flipped from Needs input to Working, and nothing restored it until the level changed, so the user could miss the open request. Skip reply text while the pane's level is Needs input, as the reporter already does for queued prompts. * fix(opencode): leave OpenCode 2 step events to the server path's own change The shared-server translation re-derived Working from session.step.started. The pane reporter no longer uses it, so it only changed the server path and duplicated a separate open change. Drop it; server behaviour matches main. * fix(opencode): reset an OpenCode 2 pane to idle when its TUI starts Before this change, a pane could keep a status an earlier process left on it. The case that matters is an upgrade mid-turn: the old shared-service plugin posted pane B's turn under pane A's key, then stood down, and pane A's TUI loaded the new reporter owning nothing, so it never posted and A showed a wrong Working until its own next turn. A freshly started reporter that owns no running turn now posts the host's existing session-start boundary once, which the host shows as connected idle: no completion, no notification, and an unseen Done it lands on stays unread. A plugin hot reload keeps its memory and skips it, and where OpenCode keeps no plugin memory it is never sent. * fix(opencode): keep OpenCode 1 serve + attach sessions on the attaching pane OpenCode 1 `opencode serve` in one pane plus `opencode attach` in others reports every session from the serve process, whose environment names only the serve pane. Before this branch the session-to-pane binder moved those posts to the attaching pane; deleting it for OpenCode 2 put them back on the serve pane. Restore the binder, registry, correlation and client sweep for OpenCode 1 (and mimo-code, which it also served), gated so OpenCode 2 never uses it: - The status plugin now sends `opencodeMajor: 2` on every post from a process whose loader called setup(), which only OpenCode 2 does. The host skips the binder (no rewrite, no kicked round) for any post that carries it. OpenCode 1 and older plugins send nothing and keep the previous behaviour. - The binder reads only OpenCode 1's `session` table. OpenCode 2 writes `session_v2`, so its sessions never bind; on a database both versions wrote, OpenCode 1 sessions are no longer hidden behind the v2 table. OpenCode-1-only; it goes when OpenCode 1 support is removed. * fix(opencode): show OpenCode 2 `opencode run` as Working, then Done, on its own pane OpenCode 2's `opencode run` loads no plugin, and the shared service that runs its session cannot tell which pane it belongs to, so a pane running it showed nothing. The runtime now reports it from the pane's own process lifetime. On the pane's OSC 133 command start (main already parses 133 for every local PTY; the start callback was never wired), after the pane tracker's 350 ms settle it reads the foreground process name through the existing foreground reader, and only when that name is OpenCode, the foreground command line from one fresh shell-foreground process-table capture. An `opencode run` posts Working through the existing terminal-status path into the hook server's store; the pane's 133;D (or the daemon's background fact) posts Done, marked interrupted on Ctrl-C. No polling. Exclusive with plugin reporting: each write carries the command's start time, and the store drops it once a hook has written the pane since then, so an OpenCode 1 `run` (in-process plugin) or a server plugin without the TUI copy owns its command alone and there is never a second Done. Local macOS and Linux panes only; SSH, WSL and Windows panes stay silent because their foreground cannot be read on this host. * fixup! fix(opencode): show OpenCode 2 `opencode run` as Working, then Done, on its own pane Type the selected descendants as process-table rows; ReturnType of the generic collector widened them to bare identity rows. * fix(opencode): keep an `opencode run` pane's Done after the command exits The run's Done was published inside the chunk that carried its OSC 133;D, before the chunk's command-finished fact. The renderer drops an exited agent's row on command-finished when the row has not changed since that fact arrived, so it took the Done as the stale row and dropped it, in the renderer and in main. Publish the run's Done after the chunk's side-effect facts are emitted (and after the daemon's background fact). The renderer then sees command-finished while the row is still Working, and the Done that follows counts as a change, so it stays, the same way a hook Done that lands after exit already does. * fix(opencode): let an `opencode run` revive a pane Orca retired When an agent Orca launched exits, command completion retires the pane, and only a hook new-turn event revived it. A later `opencode run` in that pane posts no hook, so its process-lifetime Working was refused and the pane stayed silent. A process-lifetime Working is posted only after a fresh OSC 133;C and the pane's own foreground argv prove a new OpenCode run, which is at least as strong as a hook new turn. The store now revives the retired pane on it the same way: it clears the retirement, drops the launch-token fence, and rebinds the observation. OSC status and a lone process Done still cannot write a retired pane, and a closed tab stays closed. * fix(opencode): report `opencode run` on local Windows panes too The run producer skipped every Windows pane, although Orca already resolves a Windows pane's foreground agent from the native process table. On main an OpenCode 2 `run` showed there (on the service's pane); on this branch it was silent. The argv read now has a Windows branch: one fresh native table read (windows-process-table, no interpreter spawn), the same foreground identity the Windows resolver already computes, and the command line of the process that identity names. It runs only after the name read says OpenCode. Local Windows panes whose shell prints OSC 133 C/D (PowerShell with PSReadLine, Git Bash with Orca's wrapper) now go Working, then Done; cmd.exe prints no markers and stays silent. SSH and WSL panes stay silent. * fix(opencode): re-read an `opencode run` foreground on the pane tracker's ladder The producer read the pane's foreground once, 350 ms after the command started. A wrapper, a shim or `sleep 1; opencode run` execs OpenCode later, so those runs stayed silent. The renderer's pane tracker already re-reads a command's foreground at 350, then 1200, then 6000 ms for exactly this. Move those delays to one shared module and use it from both. The producer re-reads only while the foreground is a non-shell process that is not OpenCode, and at most on those three rungs; no new polling. * fix(opencode): end an armed `opencode run` when the next command starts A new OSC 133;C in a pane whose `opencode run` was still armed dropped the armed state without a Done, so a run whose 133;D never arrived left the pane on Working. A new command start proves the previous command ended, so it now posts that run's Done (not marked interrupted: no exit code is known) before the new command is inspected. * fix(opencode): skip the session-start row inside an OpenCode 1 `run` process OpenCode 1 `run` loads the status plugin in its own process, and the plugin posts SessionStart when the run's session is created. The host lands that as an idle session boundary, so a pane the run producer had already shown as Working blinked idle before the plugin's Busy. The plugin now skips SessionStart when its own process is a `run`, read from its argv the same way isOpenCodeRunCommand reads a pane's foreground (first positional after global options, past the compiled binary's entry path). A `run` session goes Busy at once, and its first prompt part still revives a retired pane and resets the turn caches. The TUI and `serve` still post it, and OpenCode 2's `run` loads no plugin. * chore(opencode): say where the plugin's opencodeMajor comes from The field is set when OpenCode's loader calls setup(), which only OpenCode 2 does; the plugin never reads OpenCode's version. Say so at the field. Comment only; the generated plugin is unchanged. * test(opencode): type the late-Done pane fixture without bare casts The new renderer test passed its fixtures with `as never`; use the checked fixture-tuple cast with a SAFETY note, like the sibling pty-connection tests. * fix(opencode): keep `opencode run` silent when OpenCode status is turned off The run producer ignored the status-hooks switch, so a user who turned status off globally or for OpenCode (#23667) still got a run's Working and Done. It now checks isAgentStatusHooksEnabledForAgent for the run's agent (opencode or opencode2), the same predicate the plugin install honours, before reading argv or posting anything. * test(opencode): read the late-Done row without an untyped property access The mock store types agentStatusByPaneKey values as unknown, so reading `.state` failed tc:web. Assert the row with toMatchObject instead. * perf(opencode): read a command's foreground only while it could become `opencode run` Two costs the run producer paid on every command in every local pane: - The retry ladder re-read the foreground (a process-table capture on the local provider) for any non-shell program still running, including other agents, editors and dev servers, up to three times. It now re-reads only while the foreground is still the shell (the command has not exec'd yet) or an unrecognised launcher that may still exec OpenCode (node, bun, bunx, npx, npm, pnpm, pnpx, yarn). Any other program is read once. - With OpenCode status turned off for both opencode and opencode2, it still set the timer and read the foreground only to discard the result. It now returns before any timer or read. The per-agent check after the name read stays for the case where only one of them is off. * fix(opencode): let a new hook turn end a pane's process-exit completion A confirmed agent process exit records a pane-wide completion identity that names only the agent. Hook Dones are matched against it by agent alone, and only a working title cleared it, so in a pane whose agent paints no working title (OpenCode) every later hook-reported Done was treated as already notified: no notification and no unread mark. An unreported exit (a status-off `opencode run`, or quitting an idle client) was enough to set it. A fresh hook Working now clears a process-exit identity, since a new turn cannot be a duplicate of an earlier exit. Hook identities stay, so same-turn and replay dedupe is unchanged. * refactor(opencode): share the foreground read schedule as one value The pane tracker's import of the two shared foreground-read delays took four lines where its old local constants took three, which put the file one line over max-lines once merged with main. Export the settle delay and retry ladder as one object so each consumer imports a single name. No behaviour change: same 350 ms settle and 1200/6000 ms retries. |
||
|
|
8dab53fb5a |
fix(worktrees): a finished create no longer pulls you off the workspace you switched to (#23974)
* fix(worktrees): a finished create no longer pulls you off the workspace you switched to Show a ready toast with an Open action instead (#9944). * fix(worktrees): a local agent create no longer opens its workspace from the host For a composer create on a local repo that starts a terminal agent, main spawned the agent terminal with `activate: true`. The renderer turned that into a workspace switch before `createWorktree` resolved, which cleared the creation panel pointer: a user who had moved to another workspace was still pulled onto the new one, and every such create also showed a redundant "ready" toast and lost its completion focus. The host now adopts the startup and setup terminals silently (`surfaceOwner: false`, no activation), the same way runtime-managed creates do when not asked to activate. The submitting renderer's completion rule is the only thing that decides whether to open the new workspace. Also: the ready toast's Open is a deliberate user open (`navigationIntent: 'user-open'`), and the renderer tests now drive the composer's live entry and reach the post-create cancel check. * test(worktrees): pin split-mode setup silence and agent-tab focus on completion - Pin that split-mode setup is split into the startup terminal without surfacing the new workspace. - The watching-user completion test is now a real agent create: activation returns no primary tab, as it does in the app. - New store-backed test drives the host's silent reveals through the real bridge, then the completion's activation and focus queue, and checks focus lands on the host-adopted agent tab rather than the setup tab. * fix(worktrees): start a backgrounded create's renderer-owned startup without opening it When the host does not start the agent (SSH repos, local repos with project default tabs, VM creates) and the user has moved on, completion seeds the agent/setup/issue terminals without activating them. Nothing mounted those panes, so the agent, draft prompt and setup script waited until the user opened the workspace while the toast already said it was ready. The non-activating branch now asks for a background mount of just the created tabs that still owe startup work, the same request the terminal bridges and sleeping-agent wake already use. Idle tabs and host-started terminals are left unmounted. * test(worktrees): pin background mount for setup-split and issue-split only tabs * fix(worktrees): name the ready toast by kind and skip it for a user already there - The toast now reads "Worktree <name> is ready" ("Workspace <name> is ready" for folder repos, which are not git worktrees) with a "Go to worktree" action led by the external-link icon. Old readyToast/openReady keys are replaced in all six catalogs. - The created row is listed before completion, so a user can open the new workspace while it is still being created. Completion now treats that user like one still watching the creation panel (open + focus the agent, no toast), and the toast is also skipped if they opened it while a native chat was starting. * fix(worktrees): drop the unreachable toast-time "already on it" re-check Nothing between the completion decision and the toast yields to user input (the structured-chat launch awaits nothing), so the re-check could never fire, and a future await would have let a user who reopened the panel get neither activation nor a toast. The decision-time check stays. Also clarify why a folder workspace's title says Workspace while the button keeps one label. * fix(worktrees): folder workspace ready toast says Go to workspace |
||
|
|
b4b708c2c4 |
fix(native-chat): a Codex stream retry is one warning row that updates in place (#23684)
* fix(native-chat): a provider's own retry progress is quoted in its retry row A retry row whose fact carries a detail the provider wrote for a person now quotes it, the same way a rejected message or failed compaction does, so the row says how the retry is going. A log detail still stays out of the sentence. * fix(native-chat): a Codex stream retry is one warning row that updates in place An error Codex says it will retry used to fall through to the generic frame row: red, and a new row for every attempt. It now writes one providerRetrying row per retry run, warning-toned, revised by each attempt with Codex's own progress sentence. A run is the retry frames of one turn with nothing else the thread journals between them; every attempt still publishes, so the idle sweep keeps seeing activity. Errors Codex will not retry are unchanged. * fix(native-chat): a Codex retry row says it is retrying and keeps the frame behind Details The quoted retry sentence leads with "is retrying", which holds for any provider's progress text. The Codex retry row also keeps the whole bounded frame behind the row's Details, as the generic row did, so Codex's additionalDetails stays available to diagnose a retry. * test(native-chat): a Codex retry re-handled after backpressure keeps its one row Pins the run being opened before the write: a first attempt whose publish is refused and is handed back must revise the row it already wrote, not open a second run. Also stops the fixture claiming Codex sends an idle thread status beside each retry, which the app server does not do. * perf(native-chat): a Codex frame with no retry run open is not classified Ending a retry run classified every non-retry frame, and classifying walks the whole payload: every streaming delta, and every large item/completed, paid a walk about as costly as parsing the frame. Only a thread with a run open needs the answer, so the classification now runs only there. * fix(native-chat): each Codex retry attempt is its own warning row, with what failed on its second line A stream error Codex says it will retry is written as its own warning row with a providerRetrying fact, under the same per-frame identity every Codex frame row gets. The host no longer tracks retry runs or rewrites one row in place, so there is no run state to open, end or clear, and no frame has to be classified to end a run. Every attempt publishes, which keeps renewing the idle clock while Codex retries. Codex's additionalDetails, which its own UI shows under the progress message, is kept on the fact as the retry's cause and printed on the row's second line. Errors Codex will not retry are unchanged. * fix(native-chat): a transcript draws only the latest row of a provider retry run The shared structured message projection, which both the desktop and the mobile transcript read, collapses a run of retry rows into its latest row. A run is retry rows from the same agent with no other drawn row between them; a row that draws nothing, or a queued send drawn after the conversation, does not split it. The earlier attempts stay in the journal. * test(native-chat): the retry-run render test uses the message list's current props * fix(native-chat): a Codex frame row is named for its connection, so a later one never revises it Frame rows were named provider-frame:codex:<n> from a counter that starts over with every connection, so the first frame row after a reconnect in the same session revised an earlier connection's row in place, at its old spot. Each connection's frame rows now carry the acquisition generation, minted once before the translator is built: provider-frame:codex:<generation>:<n>. Rows already written keep their identities. * fix(native-chat): agents retrying at once each keep one row, read from the row's own agent * test(native-chat): a reconnect's Codex rows are named for the acquisition that received them * fix(native-chat): a retry run is one agent's, so another agent's row never splits it Each agent's rows are drawn apart: the session's own rows are the conversation, and a subagent's rows open in that subagent's section. Splitting a run on any other row in the flat list left two adjacent retry rows on screen whenever another agent wrote between two attempts: a subagent finishing a command while the session reconnected, or the session working while a subagent reconnected. A run is now per agent: an agent's retry rows with none of its own other rows between them, drawn as its latest. * test(native-chat): a subagent's retry run is checked in its own section, and the run rule's words say same agent * test(mobile): the retry-rows test typechecks, so the test ratchet keeps checking it |
||
|
|
24edf0f64b |
fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm (#23467)
* refactor(native-chat): remove the unused terminal handoff
No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.
* fix(native-chat): never let the pre-stop snapshot hold a chat's stop
Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(native-chat): drop helpers only the terminal handoff called
`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.
Co-Authored-By: Claude <noreply@anthropic.com>
* docs(native-chat): stop citing the removed handoff in lifecycle comments
Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): type the stalled snapshot drain without a cast
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): pin that a start dead before proving owes no settlement
The removed restart handoff test pinned this branch; nothing else did.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(native-chat): keep the owner-status read behind an in-flight attach
The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(terminal): remove the agent-session PTY write gate
The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(native-chat): drop the transcript helpers only the handoff called
appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(native-chat): stop calling a starting chat "mid-handoff"
A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): type the stand-in roster decoder without a cast
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(codex): name the pinned rollout lookup for what it does
With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.
* refactor(native-chat): type the owner-status reply as the host sends it
The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.
* refactor(native-chat): normalize terminal-handoff lease values once at decode
Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.
The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:
- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
`conflicted`, the claim every build probes but never stops. A plain native
owner would be stopped by restart recovery, here and in older builds.
Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.
The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.
* refactor(native-chat): stop threading the owner kind through a reservation
A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.
* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else
Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.
* fix(native-chat): name a chat write by its target, not the owner generation
A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.
Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.
Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.
* fix(native-chat): every journal append reaches the chats that are open
A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.
A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.
* test(native-chat): an epoch replacement reaches the open chat
* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map
* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite
The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.
* test(worktree-activation): restore the OMP surfaced-agent resume test
The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.
* perf(native-chat): a publish behind a delivered commit reads nothing
Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.
* test(native-chat): state why the teardown test's fake journal is safe to cast
* docs(native-chat): say mutation admission checks only the writer lease
* docs(native-chat): drop the send rebase from comments that still described it
* fix(native-chat): a message is accepted, then delivered
A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".
A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.
Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.
A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.
Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.
* fix(native-chat): settle queued messages only for the child that ended
A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.
A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.
The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.
* fix(native-chat): an adoption that fails to import keeps the conversation open
The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.
* perf(native-chat): the recovering open reads the journal once
Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.
* fix(native-chat): an attach that fails after indexing its child leaves no child behind
A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.
* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer
The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.
A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.
* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down
The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.
* fix(native-chat): a message rejected while its chat was closed reads as not sent
A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.
* test(orchestration): name why the readiness settlement fakes are cast
* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent
* docs(native-chat): drop the fence from the admission the send effects run behind
* docs(native-chat): give the fence move on release the reason that still holds
* docs(native-chat): stop citing a write fence check in launch and mailbox comments
Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.
* refactor(native-chat): the provider child is its own record
A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.
- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.
* fix(native-chat): the delivery loop alone settles a message its start or child failed
A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.
- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
reads how it ended: a Stop continues; anything else writes one failure row and rejects every
queued message with the same words, then stops. A child still starting whose start the adapter
says did not land fails the same way. The exit, eviction and the settlement retry only settle
the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
closed, with or without a child, and a start the loop already has in flight is waited for so the
child it produces is stopped rather than left behind.
* refactor(native-chat): a stopped child ends on the one reading of its stop
The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.
* feat(native-chat): the host says it accepts a send before any agent has it
The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.
* refactor(native-chat): an attach never opens a journal of its own
The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.
* fix(native-chat): a moved fence resends nothing on a host that accepts first
The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.
The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.
* refactor(native-chat): a child's end says whether the user or the host stopped it
The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.
* fix(native-chat): a chat whose only work is a queued message is not offered for resume
A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.
* test(native-chat): type the queued-message fixtures in the resume-offer tests
* fix(native-chat): a start that dies while a message waits on it is that message's failed start
Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.
* fix(native-chat): a request that failed reads as failed
A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.
The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".
* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now
* test(native-chat): a verdict change republishes the mobile status projection
* refactor(native-chat): the store's retention trigger keeps its flag compare
A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.
* test(native-chat): a user message the provider journaled keeps its session listed
* test(native-chat): pin what a failed start settles, and what a resume offer names
A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.
* test(native-chat): the failed-start pins fail on what the message became, not on a timeout
* fix(native-chat): a late provider-session update keeps a failed recovery record failed
A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.
* test(orchestration): the preamble's host stub is typed, not cast
The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.
* test(native-chat): the terminal-bell check asserts the renamed verdict field
The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.
* fix(native-chat): a failed turn ranks like a completion for attention
Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.
The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.
* fix(native-chat): a failed main agent reads failed while its subagents still work
The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.
Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.
worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.
* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it
The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.
* docs(native-chat): the status-store listing rule names provider-journaled user messages
* fix(native-chat): a refused send notifies failed through the completion feed
The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.
* fix(native-chat): every copy of a row carries the main agent's own status
History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.
- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
rebuilding one; the sync key and history equality compare it.
* test(native-chat): pin the worktree ps verdict across host and phone versions
Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.
* test(mobile): name the parity table's row for its role
* fix(native-chat): a request that settles while the user is asked something notifies once
The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.
The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.
* fix(native-chat): the completion says when the user is being asked
A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.
The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.
* fix(worktree-status): a departed agent's failure yields to live work on the worktree card
A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.
* docs(agent-status): a departed agent's failure ranks below live work on the worktree card
* fix(native-chat): a view never restarts a chat whose last start failed
A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.
* test(native-chat): start the child the loop waits on with an attach, not a second view
A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.
* fix(native-chat): settle a gone generation's turn wherever a conversation opens
A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.
* test(native-chat): prove the next child's start settles the turn an earlier child left
The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.
* test(native-chat): count a failed start's rows by row, not by text
Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.
* test(cross-version): load the phone row readers without mobile's toolchain
Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.
The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.
* test(cross-version): keep the checkout path-guard message and justify the copy import's cast
* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget
* test(native-chat): pin the open's and the send's start and row counts, however the view binds
Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.
* fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm
When the provider gave no verdict, the structured status projection now derives
one from the newest turn's lifecycle: interrupted -> interruption, unverifiable ->
unconfirmed. Nothing new is journaled, the completion feed stays provider-only, and
the legacy interrupted flag stays a user stop only. Every verdict reader handles
both arms explicitly.
* fix(native-chat): settle a gone generation's turn at every open but an acquisition's
The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.
* fix(native-chat): a folded turn a crash cut off reads Interrupted after N
The settled-turn timing now carries the turn's verdict, derived by the same
agentTurnVerdict the status row uses. A turn that ended interrupted with no
provider verdict heads its fold 'Interrupted after N'; a user's stop keeps
'Worked for N'.
* test(native-chat): hold the create's start open until the views bind
The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.
* fix(native-chat): the user's close of a chat records the turn it cuts short as their cancellation
The expected-close settle writes outcome cancellation when the user aimed the
stop at this chat: a Stop while the agent starts, the chat's tab closed (the
agentSession.close RPC, or session.tabs.close with a user reason), or /clear.
A quit, an idle eviction, a worktree teardown or an orchestration stop leaves the
turn with no verdict, so it still reads Interrupted.
* test(native-chat): a Claude turn a newer send superseded reads Interrupted
The supersede fires for any send Orca dispatched, the user's or another agent's,
and nothing at that site records the sender, so the turn keeps no verdict and
folds as Interrupted after N.
* fix(native-chat): the user's close records cancellation on the turn the provider settled on its way out
The Codex and Claude adapters settle their open turn as interrupted, with no
verdict, while the host stops them. The expected-close settle then found no
running turn, so a user's close of a mid-turn chat read Interrupted. The stop
now reads the running turns before it reaches the provider and records the
user's cancellation on each one it cut short, unless the provider gave a
verdict of its own. The close-verdict test's adapter now settles its turn on
close the way the real adapters do.
* test(native-chat): update the close and settled-turn expectations for the host-observed verdict
agentSession.close now passes the user's word to the host, and a settled
interrupted turn with no provider verdict carries `interruption`. Also merge a
duplicate import the code-quality gate rejects.
* refactor(native-chat): drop the composer's second error formatter
After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.
* test(native-chat): pin the reason on a message rejected while its chat was closed
The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.
* fix(native-chat): the user's close cancels a turn whose start landed as the provider stopped
The close read which turns were running before the stop. A turn whose start was
still in flight (a send echo not yet journaled) was absent from that read, so the
provider's verdict-less settle on the way out left it Interrupted. The close now
reads which turns were already over instead, and records the user's cancellation
on every other turn the stop left running or interrupted with no verdict. A turn
cut off earlier, or finished during the stop, keeps its end.
* fix(native-chat): a send the provider never received after a restart has no verdict
Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.
* refactor(native-chat): the adapter settles the turn a stop cuts with the stop's typed cause
The host hands its stop's cause to the adapter's close. Each adapter settles its own
open turn on 'ended' through one mapping, turnVerdictForChildEnd, and the host's
dead-generation fallback uses the same mapping for any turn no adapter settled. A
user's close or stop of this chat is their cancellation; a quit, eviction, teardown
or an exit the adapter saw first is news.
Deletes the snapshot-and-diff reconstruction (endedTurnItemIds,
userStoppedTurnRevisions, settleUserStoppedTurns) and the requestedByUser flag.
host.close now takes a required cause.
* test(native-chat): a user's close drops the chat's status row like an eviction
* test(native-chat): expect the eviction cause on the host closes of idle release, worker stop and worker discard
The stop's typed cause now travels into host.close and the adapter's close, so these
three non-user closes assert the 'evict' they pass.
* refactor(native-chat): every stop names its cause, so none defaults to the user's cancellation
stopStructuredAgentSessionAgentUnderSerialize defaulted its ending to 'user-stop', which now
settles the cut turn as the user's cancellation. Every caller already passes a cause; the
parameter is now required, and a type-level test fails to compile if the default returns.
* test(native-chat): pin who a chat's session.tabs.close speaks for, older clients' reasonless close included
The mapping lived inline in a type-unchecked file, and only the explicit user reason had a test: an older client's reasonless close, or a lifecycle echo read as the user's, stayed green. It is now one exhaustive, type-checked function with a case per reason.
* style(mobile): draw the unconfirmed dot in the theme's status amber, not an inline hex
* docs(agent-status): the main agent's outcome also carries the host-observed end, interruption or unconfirmed
* fix(native-chat): a chat the user closed while its agent started is not a failed start
A still-starting child the user's close cut counted as a failed start, since only 'user-stop' was excluded: a start-failure row, and queued messages rejected as a provider failure. Whether an ending fails its start is now one exhaustive switch, shared by the failed-start read and the delivery loop's handover: a user's stop or close never does; an exit, a failed attach and the host's own stops still do.
* fix(native-chat): a chat the user closed closes its queued messages, and starts no agent for them
After
|
||
|
|
84f58a1fc7 |
fix(relay): refuse a redial at once while the host's own release holds its row (#24225)
* fix(relay): refuse a redial at once while the host's own release holds its row During an Asia drain the host whose socket closes is the one that redials. Its release on the draining cell locks its assignment row first, then waits on the cell's busy row for up to the lock timeout. The director's sticky and placement paths waited on that assignment row inside the single sticky slot, and the sticky path then locked the busy cell row itself before it checked isolation. The slot backed up and dials timed out fleet-wide. Both paths now take the host's assignment row NOWAIT and throw RelayAssignmentRowBusyError when it is held. /v1/assign answers that with 503, Retry-After 1 and error assignment_row_busy, logged with its own reason. The sticky path decides isolation before it touches the pinned cell row. An isolated retry keeps its own tier as its retry scope, so a busy lock inside it no longer falls to the all-rows path. The local drain arm drops its zero-release-failures bar, which the Asia arm never had. The drain harness gains a departing-host arm: each host releases its own lease, then redials after 150, 400 or 1000 ms on the desktop client's 5-5.5 s pacing. At 400 ms, main rejected 83 of 180 first dials by sticky wait timeout, placed 11.6/s with 3.1 director backends lock-waiting, and took 11.2 s at p95 from release to placed. Now: 16 fast refusals, 18/s, no lock waits, 5.7 s at p95. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): wait briefly for a calm host's row and keep the dead-cell sweep going The dead-cell sweep treated RelayAssignmentRowBusyError as fatal, so one busy host ended the sweep for every later host each tick. It now skips that host and carries on. The sticky path refused a busy row at once for every host. A calm host redialling after its own clean close often meets its own short release, and a refusal costs it the client's 5 s assign gate. When the pinned cell is general and live, the sticky path now waits up to 1 s for the row before refusing. A roll-isolated, parked or dead cell still gets the immediate refusal. Placement keeps NOWAIT, because it holds cell rows while it would wait. A resume refused for a busy row now carries Retry-After 1 as well. The departing-host harness arm now bounds the busy refusals at 20% of hosts and the p95 at 8 s, and counts unexpected errors apart from retryable refusals. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): never wait on a host's row while the sticky retry holds its cell row The inventory-first sticky retry takes the pinned cell row before the assignment row. With the calm-host bounded wait it could then wait up to 1 s on the assignment row while holding the cell row, the reverse of the ranked lock order, against this host's own release, which holds its row and wants the cell's. The bounded wait now applies only when no cell row is held; the retry stays NOWAIT. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
5cda0f4508 |
refactor(native-chat): keep agent-session records in the chat journal database (#24006)
* fix(native-chat): report a failed startup chat reconcile instead of failing app startup At startup the chat host re-checks every saved chat's lease and writes the result to agent-sessions.json. If that write failed (the file lock gave up, the file could not be written, or the file was written by a newer Orca and is read-only here), reconcileRestartLeases rejected, the startup IPC call rejected, and the renderer fell into its degraded "Session restore failed. Changes won't be saved until restart" mode. The reconcile is bookkeeping: a lease left unreconciled grants no writer, and every attach, send and read of a chat reconciles its own lease again. So the startup reconcile now reports its failure through a new optional host dependency, onStartupReconcileFailure, and resolves. The runtime routes it to its onError sink under the scope structured-agent-session-startup-reconcile, or logs it when no sink is installed (the desktop installs none). * fix(native-chat): read restored chats without waiting on lease bookkeeping With native chat on and a chat tab open at quit, the renderer's startup also awaits the chat tab restore (session.tabs.listAll). That restore re-ran the lease reconcile before reading each chat and rethrew its store failure, then recorded each restored tab as visible through a store transaction that throws on a held lock or a read-only store. Either one failed the restore, so startup still fell into "Session restore failed". Reading a chat grants no writer, so the reconcile startup and the restore run is now a reader's: createReaderReconcile never throws, answers whether every lease is settled (recovery is resolved only then; the journal opens either way), and reports each distinct failure once until a reconcile settles. Attach and agent start keep the strict reconcile. The restore's tab republish logs a failed visibility write and still publishes the tab, since a client drops every unpublished chat tab; user-driven publishes still refuse. The host dependency is renamed onLeaseReconcileFailure (scope structured-agent-session-lease-reconcile), since it now also reports for reads. * fix(native-chat): keep every record-store write off the startup chat read path Round-2 review found two more writes on the startup chat restore that could still fail it and put the app into "Session restore failed": republishing a /clear replacement recorded its tab visibility strictly, and resolving a chat's recovery rethrew its store error. The restore also paid one lock wait per tab and per batch of chats while the lock stayed held. The restore now derives tabs from state it already holds: - publishStructuredAgentSessionTab splits into the strict write and projectStructuredAgentSessionTab, which only updates the runtime's snapshot. The restore and /clear replacements only project: a saved tab index already lists every restored chat, and a /clear moves the tab in the same write that commits it. visibilityWriteMayFail is gone. - Chats a legacy profile restores that the index does not list are recorded in one best-effort transaction (store.showSessionTabs), so a failure leaves the index absent to seed again rather than partial. - The read restore's recovery resolution is caught and reported through onLeaseReconcileFailure, deduplicated with the reconcile's reports. - Once lease bookkeeping fails in a restore pass, the rest of that pass skips it, so a held lock costs one wait for the startup reconcile and one for the restore, however many chats are open. User actions (create, reveal, attach, send, the /clear commit) keep their strict writes. * test: open, seed and read the agent-session record store through one harness Tests that open the durable agent-session record store, seed it, or read back what it persisted now go through agent-session-record-store-test-harness.ts instead of calling AgentSessionRecordStore.open or touching agent-sessions.json themselves. A later change that moves the store into the chat database then changes the harness instead of every test. No production code changes. Tests whose subject is the JSON file itself (its .bak recovery, salvage, schema versions, permissions, and what older builds read back) keep reading and writing the file directly; the storage move rewrites or deletes them. * fix(native-chat): start each restore pass from one lease check and stop its bookkeeping at the first failure The restore now runs one reader lease check for the pass and lets each chat re-check and resolve recovery only while the pass is still settled. The first refusal or failed write clears it for the rest of the pass, and every chat is still opened for reading. With another process holding the lock, startup waits on it once in prepare and once in the restore, however many chats are open; a legacy profile waits once more for its tab-index seed. * docs(native-chat): correct restore comments and a test name to match the final design * test: address the record-store harness by the host's state directory The harness took the store's own folder, so each caller picked one (join(root, 'store'), or 'agent-sessions' where a test read the store the runtime owns). A later change that moves the store into the state directory's journal database could not tell those apart, and would have had to edit every caller again. Every harness function now takes the state directory, the one the test's journal database and recovery capsule already live in, and keeps the store in the same subfolder the runtime uses. Callers pass that directory; store-only tests pass their temp directory unchanged. Format tests that share a directory with harness calls take the file path from testAgentSessionStoreFilePath. The folder name moves from a private constant in the runtime to AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness shares it without importing the runtime. Its value and every path built from it are unchanged. * refactor(native-chat): keep agent-session records in the chat journal database The record store's records, operation ledger, retired claim keys and chat tab index become tables in agent-session-journal.db (user_version 4). The version-4 migration copies agent-sessions.json in its own transaction and never writes, renames or deletes that file or its .bak. Each store write is one journal transaction over exactly the rows it changed, checked with the load rules; the file lock, the external-change refresh and its hash, the .bak rotation, salvage and the hot-path recovery fence are gone from the store. * wip: importer tests * test(native-chat): cover the records migration, the import, row writes and read-only records * docs(native-chat): retire comments that describe the records file as the live store * test(native-chat): drop the record-store harness's leftover file path and type the import fixture * test(native-chat): let the host harness cleanup wait out a recovery-offer read's lock * fix(native-chat): let Stop reach the agent when its ledger row cannot be written Stop's operation-ledger row now shares the database with the chat history, so damage, a full disk or a stranded transaction on that write refused the Stop before the interrupt. A cancel plan now takes its decision from the committed ledger in memory, runs without settling, and warns that the row was skipped. Other mutations answer proven damage with the typed "Unable to load this chat." refusal instead of the raw SQLite error. * fix(native-chat): answer whether a profile holds chats from the database's rows Every host install creates agent-session-journal.db, chats or not, and the version probe created it too, so its mere existence made every profile that ever installed the host wait on host install and reconcile at startup. The check now opens the database read-only and looks for a record or tab row, lets the records file answer while its import is still owed, and counts an unreadable database as present. The version probe no longer creates the file. * fix(native-chat): open a chat from history when its tab index cannot be written Over records a newer Orca wrote, every write is refused, so opening a closed chat from Agent Session History failed on the tab-visibility write and the chat read as unreachable. Like closing a tab, opening one now reports a failed restore-index write and still publishes the tab. * fix(native-chat): keep the records import owed when the backup read fails transiently A torn records file whose .bak could not be read (EACCES, EIO) was reported as unusable, so the migration completed with nothing copied and never retried. A non-ENOENT read failure of either copy now carries its cause, which the importer classifies as a read that can clear. * test(native-chat): pin that an unreadable records file never falls back to its backup * fix(native-chat): restore imported chats' tabs when the records file had no tab index A chat created while the import was owed recorded a tab index holding only itself. When the file it later imported had no index, that index still read as recorded, so the imported chats' tabs never came back. The import now clears the recorded marker in that case, and restore falls back to the profile's tabs. * refactor(native-chat): drop the unused in-transaction store write Nothing called it, and it bypassed the write queue and the read-only refusal. * docs(native-chat): say that an unusable records file is left untouched but never re-imported * refactor(native-chat): keep the provider handle chain check as main has it The chain-validation refactor has no measured need in this change. * docs(native-chat): retire lease-renewer comments that describe the records file as the live store * fix(native-chat): keep a throwing failure sink from failing the startup chat read The lease bookkeeping failure reporter called the host's failure sink directly, so a sink that threw turned a reported, recoverable store failure back into a rejected startup reconcile or read restore. The reporter now catches a sink throw and logs both the original failure and the sink error with console.warn. * test(native-chat): wait for a replaced host's restart-offer writes before cleanup A restart test replaces the host without tearing the old one down, so the old host's fire-and-forget restart-offer withdrawal could still hold the recovery capsule's lock directory when cleanup removed the test directory (ENOTEMPTY). The harness now hands hosts a capsule that tracks running operations and waits for them before removing the directory, replacing the rm retries. * docs(native-chat): retire the abandon helper's note that the store re-creates its directory * fix(native-chat): restore a chat opened while the import was owed beside the profile's chats When the imported records file had no tab index, restore fell back to the profile's saved tabs, which never list a Claude chat, and the seed then rewrote the tab table without the chat opened while the import was owed. The tab rows that chat left are now loaded as unrecorded, restore takes them together with the profile's chats, and the seed keeps their tab ids. * test(native-chat): pin that a create whose tab index write fails still opens the chat * docs(native-chat): say why restore puts chats opened while the import was owed first * test(native-chat): replace a ledger row rather than change it in place in the Send-now rerun test The record store freezes published rows in tests, so setting a row's outcome in place threw; the test now swaps in a changed copy, as its sibling cases do. |
||
|
|
cfa43e7eab |
fix(codex): opening a terminal no longer strips Codex hooks from the real ~/.codex (#23552)
* fix(codex): a real-home restore leaves a file alone once someone else changed it Orca writes ~/.codex/hooks.json (and a trust rebase writes config.toml), then runs a Codex trust session for up to 10 s, then restores the original bytes if the session fails. The restore wrote unconditionally, so a save that landed during the session, from the user or another Orca, was silently reverted. Each restore now compares first: it writes the original back only while the file still holds the generation Orca's mutation left, and otherwise logs and leaves it alone. This covers the real-home install and opt-out sweep (restoreRealHomeHooksJson), the legacy sweep's hooks restore, and config.toml rollback (restoreCodexTrustConfig). For hooks.json the generation is the exact bytes Orca wrote. For a config.toml that a trust rebase changed it is the file as the rebase left it. When Codex itself wrote config.toml inside the session that just failed, Orca never knew those bytes, so that rollback compares against the file as the session settled. The next commit keeps other Orca instances out of that window; a user edit made during such a session can still be rolled back. * fix(codex): serialize real-home Codex writes across Orca instances Every Orca on one HOME (a dev and a packaged app, or an offline CLI) writes the same ~/.codex/hooks.json, config.toml and ~/.orca/agent-hooks/codex-hook.sh. The per-file lane that orders capture, mutate and restore was in-process only, so another instance could write inside this one's restore window, or undo it. The lane for the user's real config.toml now also holds the existing crash-safe managed-hook install lock (~/.orca/managed-hook-install.lock, the one relay installers take for the same home). It is taken only by the outermost acquire, because the lock file is not reentrant and grants and trust rebases nest inside an install. Managed-home installs, the real-home install and opt-out sweep, and the legacy sweep all enter through it. Compare-and-swap on restore stays as the backstop. A lock that cannot be taken within its 10 s wait fails that install, which is already best effort: launch prep logs it, and the real-home lane falls back to the managed lane until its retry. * fix(codex): opening a terminal no longer strips the shared Codex entry from ~/.codex Every Orca instance on one HOME writes the same status-hook entry into the user's ~/.codex/hooks.json, with its trust in config.toml. Launch prep runs on every pane spawn, and under a managed Codex account it ran the legacy system sweep. That sweep matched Orca entries by script file name, so it removed the current shared entry and the trust blocks the grant ledger recorded. On a live laptop hooks.json went 4139 -> 18 bytes about 150 ms before a new pane opened. With hooks off, the real-home lane's launch prep swept the same way. Now nothing automatic removes the current entry or its trust: - The legacy sweep removes only an enumerated list of retired command forms that no build writes any more (#1019's double-quoted form, #1536's exec-guarded form, and Windows' per-userData bare path), plus their trust. - ensureRealHomeCodexHookState with hooks off writes nothing; that covers launch prep, session resume and startup. - Only the user's explicit opt-out (codexHookService.remove()) strips the entry and its ledger-recorded trust from the real home. - The sweep-suppression gate existed only to stop the sweep from deleting the current entry, so it is deleted with its main-process wiring. Startup with hooks off already skipped the real-home install; with this change the first pane's launch prep with hooks off also leaves ~/.codex untouched. * fix(codex): a pane's prepare-codex only repairs a home its own HOME's app installed On macOS a pane starts through login(1), so it gets the user's real HOME even when its Orca app runs with another one. The pane's `codex()` preflight installed hooks in the CLI process with that real HOME: it rewrote ~/.orca/agent-hooks/codex-hook.sh, promoted trust into the real config.toml, and wrote the real HOME's script path into the app's managed home. The preflight now acts only when the managed home's hooks already run this process's own shared script, which proves the app that installed them shares its HOME. Otherwise it writes nothing; the app installed the home at spawn. Why not a no-op: the preflight was added (#14326) because trust can go stale between opening a pane and typing `codex`, for example in a pane that survives an app update, and Codex then stops in hook review. For a same-HOME pane it still repairs that. Why keep promotion: the install drops runtime trust the system config does not back, so skipping promotion would delete approvals the user gave inside Orca-launched Codex. * test(agent-hooks): await every installer in the refresher coverage test The test fired each managed installer without awaiting it and read ~/.orca/agent-hooks straight after. Codex's install now takes the cross-process real-home lock before it writes its script, so the script landed after the read. Await the installers, and stub Codex's trust sessions so the awaited install cannot start a real `codex app-server`. * fix(codex): retire the two real-home command forms the list missed The real-home lane wrote two Codex hook forms into ~/.codex that no build writes any more and that the enumerated retired list did not name: - POSIX, #9501 until #10885: the file-guarded form draining with a bare `cat`. - Windows, #9501 until #10221 took Windows off the real-home lane: the encoded PowerShell launcher for a non-cmd-safe script path. The file-name sweep removed both before; the enumerated sweep left them in place, trusted, still passing the script's exit status to Codex. Both now match as frozen literals. Also corrects the startup ordering comment: the real-home install runs first so its in-slot upgrade lands before the managed install's sweep retires the prior command; nothing re-arms a legacy sweep any more. * fix(codex): take the real-home lock only when a write is needed The previous commit made every entry to the real-home config lane take the cross-process lock. That lane runs on every pane spawn and every typed `codex` preflight, so the steady state paid an owner probe (a `ps` spawn on macOS) and could wait up to 10 s behind another instance's trust session, even though it wrote nothing. Each real-home writer now compares the desired state with the files on disk first, without the lock. Only when a write is needed does it take the lock, re-read and recheck, then write: - real-home install: the planned hooks.json, the shared script and the ledger-recorded grant are compared; the locked path re-plans from disk. - legacy sweep: locks only when a retired entry is present; the sweep re-reads. - approval promotion: locks only when there is something to promote; the promotions are recomputed under the lock. - the shared ~/.orca/agent-hooks script: locks only when its bytes differ. The explicit opt-out always takes the lock. The lock is reentrant through async context, since grants and rebases nest inside an install, so the config-lane option the previous commit added is removed. * fix(codex): a shared script without its exec bit is not the steady state The compare-first check matched the shared ~/.orca/agent-hooks script on bytes alone. writeManagedScript also restores 0755 on every call, and the POSIX hook guard skips a script that is not executable, so a script whose mode was lost (a dotfiles restore, a plain copy) now stayed that way: every Codex hook drained stdin and reported nothing until an app restart refreshed the script. The check now also requires the mode the writer sets, so that case takes the lock and the write path repairs it. * test(codex): the retired encoded launcher never matches today's shared one The shared encoded Windows launcher is still current for other agents, so the comment claiming today's launcher is never encoded was wrong. What keeps the retired matcher off it is the exact payload: since #14825 the shared launcher prefixes its payload and drops -ExecutionPolicy Bypass. Pin that with a case. * fix(codex): the pane step recognises its own script under a home path with an apostrophe The same-HOME check looked for the script path wrapped in bare single quotes, but both hook writers escape an apostrophe inside the quotes. A home such as C:\Users\O'Brien never matched, so the pane-step repair never ran there. * fix(codex): the trust-RPC escape hatch still keeps the real home off its lane The no-write check reported a recorded grant as current, so with ORCA_DISABLE_CODEX_TRUST_RPC set the real-home lane stayed in use. The grant itself refuses before reading its ledger; the check now does the same. * fix(codex): the shared script write no longer waits on the real-home lock The write is atomic and skips identical bytes; waiting behind another instance's trust session could only fail a pane's managed-home install. * fix(codex): an in-Orca approval survives a launch that cannot get the real-home lock The install drops runtime trust the system config does not back, so a promotion skipped for want of the lock lost the approval for good. It now writes unlocked, as it did before the lock existed. * refactor(codex): take the cross-process real-home lock back out The lock fixed no observed failure. The three that were observed each have their own fix in this series: the legacy sweep matches only frozen retired command forms, hooks-off launch prep writes nothing, and a pane's prepare-codex repairs only a home its own HOME's app installed. The lock instead brought its own defects: a steady-state spawn waiting behind another instance's trust session, a compare-first split to avoid that, a script write and an approval promotion that could fail for want of the lock. Removed, with their tests: the real-home write lock and its async-context reentrancy, the plan/compare split that kept it off steady-state spawns, the compare-first legacy sweep, the locked approval promotion and its unlocked fallback, the compare-first shared script write (writeManagedScript already skips identical bytes and restores the exec bit), and the CLI tsconfig entries the lock pulled in. Kept: the retired-forms matcher, the hooks-off no-op, removal only on an explicit opt-out, the pane own-script check, and the compare-and-swap rollbacks. Every instance now writes identical bytes idempotently. * fix(codex): an opt-out that cannot read hooks.json keeps Orca's trust and ledger The opt-out swept the real-home entry, then dropped Orca's ledger-proven trust whenever a ledger existed, even when the sweep could not read hooks.json. The entry could still be there, now untrusted, and the ledger that proves ownership was gone for the retry. Drop that trust only after a sweep that read the file. * refactor(agent-hooks): one predicate for whether an agent's status hooks are on "Global switch on and this agent not turned off" was spelled out separately in the startup controls, the settings reconcile, the retained-home reconcile, the WSL preflight RPC, the CLI preflight and the OpenCode plugin selection. They now share one function, in a module light enough for the CLI's per-launch Codex preflight to load. The PTY spawn env derives the Codex flag from the switch and opt-out list it already carries, the same way it does for OpenCode and Pi, instead of receiving a second copy. * fix(codex): launch and resume prep honour Codex's per-agent hook opt-out Turning Codex off in the per-agent hook settings removes Orca's Codex hook entry, but launch prep and session resume read only the global hooks switch, so the next Codex launch or resume wrote the entry straight back into the real ~/.codex or the account's home. Both now read the per-agent predicate, which the PTY spawn env and startup already honoured. * fix(codex): turning Codex off per agent clears the real ~/.codex entry While the real-home lane owns ~/.codex/hooks.json, the legacy system-home sweep stands down. That gate read only the global switch, so turning Codex off per agent ran remove() with the sweep still suppressed and left Orca's entry in the real ~/.codex. The gate now reads the per-agent predicate, the same as turning every hook off. * test(codex): cover the system ~/.codex sweep gate for Codex turned off The gate that lets the legacy system-home sweep run was an inline closure in startup, so reverting it to the global switch left CI green. It is now a pure function beside the gate it feeds, with a table test and a remove() test on a seeded ~/.codex: turning Codex off strips Orca's entry and keeps user hooks; with Codex on the entry stays. * fix(cli): keep the agent-status hooks predicate loadable by the packaged CLI The CLI's prepare-codex handler imported the predicate from src/main, but the Electron build rebuilds out/main from its declared entries only, so the packaged `orca agent hooks` commands could not load it (package jobs and the CLI bundle-parity test were red). The predicate reads only settings, so it now lives in src/shared, which the CLI compiles itself. * feat(codex): every Orca build writes one frozen Codex hook command The Codex hook command was built from this build's wrapper, so two builds on one HOME disagreed about the bytes of the shared ~/.codex entry and kept rewriting it, with a Codex trust session each time. The command is now fixed per form and carries its form number: - POSIX: one command with no path in it. It runs the shared script only in an Orca pane with hooks on (pane key and hook port set), drains stdin everywhere else, and always exits 0. A branch for a per-build script root is written now and stays dormant until Orca sets ORCA_AGENT_HOOK_ROOT, so that change will not move these bytes. - Windows: the bare forward-slash path to the shared .cmd, which runs under PowerShell 7 and 5.1, Codex's hook hosts. A profile path that is not one PowerShell token gets a plain PowerShell form with the same branches. The literals live in the form module, so a change to the shared hook constants cannot move them; goldens pin the bytes. Every form keeps `agent-hooks/codex-hook.*` in plain text, so older builds still recognize it. * fix(codex): one main-process owner adds the real-home entry; nothing restores files Each Orca writer of ~/.codex decided what Orca's entry must be from its own build and instance, then removed or reverted whatever differed: launch prep rewrote any Orca-shaped entry to this build's command and stripped Orca entries from events this build does not use, and a failed trust session restored hooks.json and config.toml from snapshots. With several instances and builds on one HOME, every disagreement became a deletion or a revert. The main process is now the one writer, and its writes are add-only: - A launch or resume adds Orca's frozen entry to an event that has none and leaves every Orca entry it finds, so a running older build is never fought. - App start also converts an older Orca form to the frozen command, once, in its own slot: one hooks.json write (one .bak) and one trust grant per home. - A newer form is never rewritten or appended beside, and Orca entries in events this build does not use are kept. - After a failed trust grant, only an entry this call wrote that is still untrusted is withdrawn, putting back the handler it replaced. Both files are re-read, so a concurrent edit, or the identical entry another Orca trusted meanwhile, survives. Deleted: the compare-and-swap hooks.json restore, the config.toml snapshot restore after a grant session and after a user-trust re-key, and the rollback module. A grant session writes trust only at Orca's own keys, and every caller settles those keys itself. A failed re-key of moved user hooks now keeps the write and reports it; Codex lists those hooks for review. * fix(codex): the pane CLI asks the app to prepare its Codex home `orca agent hooks prepare-codex` ran Codex's install inside the pane. That process can have the real HOME (login(1)) and runs outside the app's in-process queues, so it was a second writer of ~/.codex and ~/.orca beside the app. A check that the home ran "its own script" guarded it. The pane step now only asks the app, over the same kind of local RPC the WSL pane step already uses (agentHooks.prepareCodexForPane). The app checks that the pane's CODEX_HOME is one its own userData owns, reads its own hooks setting, and installs on its own queue. An app that is not running, or is too old to know the method, makes the step a no-op, as it is on WSL. The own-script check and the CLI's settings read are gone, and the preflight module leaves the CLI bundle. * fix(codex): delete the pane step on native hosts The previous commit had `orca agent hooks prepare-codex` ask the app to prepare the pane's Codex home. The case it existed for (#14326, a pane that survives an app update with stale hook trust) did not reproduce, and no other desktop agent host writes agent config from a terminal or launch wrapper. - Deleted: the agentHooks.prepareCodexForPane RPC method, its params and catalog entry, and prepareManagedCodexHomeBeforeShellLaunch with its module, tests and CLI build entry. - `agent hooks prepare-codex` is a no-op on native hosts. It stays for one release so shell wrappers from older builds, which still call it, exit 0. - WSL panes are unchanged: they still ask the app over agentHooks.prepareCodexForWslPane. The shell wrappers and ORCA_CODEX_LAUNCH_PREFLIGHT stay, because WSL panes use the same wrappers and variable (forwarded through WSLENV). A native pane still starts the CLI once per `codex` it runs; skipping that is a follow-up. * test(codex): a failed trust session keeps concurrent edits to both files QA case 9 at host level, on a real file system in a temp HOME: Codex's trust session fails after another writer saved hooks.json and config.toml. - Both saves survive, and no Orca entry is left that Codex would list for review: this call's entry is withdrawn. - A failed one-time conversion puts the older Orca entry back in its slot and keeps both saves. Both tests fail on the previous head, which restored config.toml from a snapshot and left the untrusted entries in hooks.json. Removing the withdrawal turns both red. * feat(codex): read whether an Orca entry's stored trust is still current A Codex release that changes how it hashes a hook leaves Orca's stored trust stale: the entry is present, but Codex lists it as modified. Checking only whether the entry is missing cannot see that. readOrcaEntryTrust sorts a present entry into four states: - trusted: the stored hash is the current one; - untrusted: there is no stored hash; - stale: the stored hash is not the current one; - disabled: the user turned the entry off. The caller can pass Codex's current hash, for example one a grant recorded. The failed-grant withdrawal now uses it, and also keeps an entry the user turned off. Nothing re-grants on 'stale' yet. * fix(codex): a slow Codex start retries on the next launch, never for minutes On a loaded Mac a cold `codex app-server` took over 10 s (QA case 4). The grant timed out, the entry was withdrawn, and a 5-minute cooldown in both the grant and the real-home install then refused every retry. - The native session deadline is 30 s, the same as WSL's. - A timeout starts no cooldown in the grant or in the real-home install. The next launch retries. Other failures keep their cooldown. - Launches that queue behind a slow session share one follow-up run, so a launch waits for at most two sessions, not one per earlier launch. Tests: a 15 s cold start still grants and keeps the entry; after a timeout, the next launch runs a session at once; four queued launches run two sessions. Each is red on the previous head, and each mechanism was removed in turn to confirm its test turns red. * fix(codex): Orca's automatic writes never move a user hook Codex keys a hook's trust by its position in hooks.json. App start's collapse of Orca duplicates removed every Orca entry and appended one at the end. That moved any user hook that followed a removed entry, so the write waited on a session to re-key the moved hook's trust. App start now: - converts the first Orca entry that sits in a plain slot to the frozen command, in place; - drops any other Orca entry only when that moves no user hook; - keeps a duplicate that a user hook follows, and trusts every frozen copy, so none is listed for review; - appends only when no frozen entry is left. Tests check user positions and user trust blocks byte-for-byte for each automatic write: add-missing (append), the one-time conversion (in place), a trailing duplicate, a duplicate before a user hook, and older duplicates normalized to one entry. The three collapse cases fail on the previous head. Removing the position check, or the in-place conversion, turns its tests red. Only the explicit opt-out still removes an entry that user hooks follow. * fix(codex): removing an Orca entry never waits on a Codex session Removing an Orca entry from ~/.codex/hooks.json moves every user hook behind it up a slot, and Codex keys trust by slot. The retired-form sweep, the opt-out and a failed-grant withdrawal all asked a `codex app-server` session to list the old trust before writing, and to re-key it afterwards. A timeout there threw before the write and latched a 5-minute cooldown, so a slow cold start blocked the retired-form sweep at boot (QA case 4). Each moved hook's [hooks.state] block now moves to its new key, body bytes unchanged, straight after the hooks.json write. Codex hashes a hook's content, not its position or its file path, so the moved block stays exactly as valid as it was: a trusted hook stays trusted, an untrusted one stays untrusted, and one the user turned off stays off. No removal waits on or depends on a session. A failed config.toml write keeps the hooks write and logs. Deleted: the inspect and repair sessions, their client, and their cooldown. The generation guards on the hooks.json writes stay, for other processes. Tests: the retired sweep removes the retired entry and carries the trust of the user hook behind it while every Codex session times out (red on the previous head); the opt-out carries an appended user hook's trust; the move carries trusted, disabled and untrusted states byte for byte. Removing the move turns all of them red. * fix(codex): a Codex launch never waits on Codex's approval of Orca's entry A launch on the real-home lane awaited Codex's trust grant for the entry it had just added. A cold `codex app-server` on a loaded Mac took over 10 s, so the launch could wait that long, and a failure then latched a 5-minute cooldown. - Codex's approval runs in the background, with a 30 s cold-start budget. - A launch uses the real home only when the ledger shows trust is already current. Otherwise it goes to the managed home at once, and the next launch picks up the finished grant. - A launch that arrives while a grant runs does no work and does not queue behind it. - A resume into the real home has no managed home to fall back to. It waits for the grant, but no longer than the 10 s a launch always could. - A background grant that times out starts no cooldown; the next launch retries. Any other failure backs off for 10 s instead of 5 minutes. Success is what the ledger remembers. - A failed grant still withdraws only what that install added and is still unapproved. The log now says how many entries it took back and when the next try comes. Managed-home grants keep their 10 s deadline and stay on launch prep, as before; they fall back to Orca-computed trust. Tests: - A 15 s start: the launch returns in under a second on the managed home, a second launch starts no session, the grant lands in the background, and the next launch uses the real home. - A timeout sets no cooldown, withdraws its adds and logs it. - Another failure retries after 10 s, not before. - A resume waits only as long as allowed. - Case 9 checks the log line and the retry. Making the launch await the grant, a 10 s budget, either timeout cooldown, and a 5-minute backoff were each tried, and each turns its test red. * fix(codex): move a hook's trust only when every stored key has the known shape Orca now edits Codex's trust store directly when a removal moves a user hook. Three safeguards keep that honest: - Fail safe. If any [hooks.state] key in config.toml does not have the shape `<path>:<event>:<group>:<handler>`, nothing moves and Codex asks the user to review. That shape was checked unchanged from Codex 0.141 to 0.158. - Targeted. The file is read immediately before the atomic rename, and only the moved keys' blocks change. Every other byte stays, and no snapshot is restored. - Verbatim. Each block's body moves as Codex wrote it, including fields Orca does not know. No hash is ever computed, and a hook with no block gets none. Tests: - An unknown key shape stops every move. - Everything except the moved block survives byte for byte, and the moved body keeps an unknown field. - In case 9, a hook the user approved during the failed session keeps its approval when the withdrawal moves it, beside the concurrent project edit. Removing the shape check, or writing a computed block instead of the stored body, turns these tests red. * refactor(codex): keep only the trust read the failed-grant withdrawal uses A capture across Codex 0.141, 0.150 and 0.158, switching in all six directions, showed Orca's entry keeps the same hash and stays trusted. A Codex upgrade does not make its trust stale, so nothing needs to re-grant on staleness. readOrcaEntryTrust keeps the four states the withdrawal needs, but loses the parameter that let a caller pass a different current hash, and the test for a Codex that hashes differently. * fix(codex): native panes no longer start the Orca CLI before each codex The pane step is a no-op on native hosts, but native panes still carried ORCA_CODEX_LAUNCH_PREFLIGHT, so every `codex` typed in a pane started the Orca CLI for nothing. Only a packaged Windows build's WSL pane now gets the variable; the app prepares every native Codex home itself. The resolver loses the dev-launcher path and its userDataPath option, which only native panes used. Tests: a native macOS, Linux and Windows pane gets no preflight, packaged or not, even with the bundled CLI present; a WSL pane still gets the verified absolute launcher. Letting native panes through again turns them red. * chore(cli): say when the native prepare-codex no-op can go Native pane wrappers from builds up to v1.4.216 still call it. It can be deleted once no supported build's wrapper does. * test(codex): check the WSL launcher path instead of asserting it * fix(codex): a launch no longer waits behind the background real-home approval The background grant ran its whole codex app-server session inside the shared ~/.codex/config.toml lane, and on a cold host its session was also the shared capability probe. A launch sent to the managed home then waited on both: the managed install and the project-trust write queue on that lane, and the managed install's own grant waited for the probe. On a cold app-server that was up to 30 s per launch. The lane was held across the session only to protect the retired capture-and-restore. Codex writes its own records, so the lane is now taken only around Orca's own pre-grant write. The background grant runs its session without publishing it as the shared probe, and the whole grant is bounded by its deadline, so a hang outside the session cannot leave the lane 'granting'. * fix(codex): a failed re-grant no longer strips Codex's own approval of Orca's entries Before each trust session, the grant deleted every Orca record whose hash matched the one Orca computes. That exists because a managed home's fallback writes Orca-computed trust under both Windows path-separator spellings, and Codex rewrites only its own spelling, so the other copy would linger. On failure the managed and WSL fallbacks write that trust back, and before this fold a snapshot restore covered it. The real ~/.codex has neither: Orca never writes computed trust there (the real-home lane does not run on Windows at all), so a matching record there is Codex's own approval. After a ledger miss (another Orca profile, a Codex update, a lost ledger) and a failed session, nothing put it back, and every Orca entry showed "Hooks need review". The clear now runs only for homes whose fallback writes that trust. * fix(codex): a real-home resume spawns only once Orca's entry is approved or withdrawn A resume that must run in ~/.codex waited at most 10 s for the background approval, then spawned anyway. On a cold app-server that left Codex beside an unapproved Orca entry, so the resumed pane showed hook review. The resume now waits for the grant to settle. Settled means Codex approved the entry, or the grant failed and withdrew its own unapproved write; the grant's deadline bounds the wait (30 s, the cold-start budget), and a failed approval never fails the resume. Why this over the alternatives: - Spawning at 10 s keeps the review prompt this fold exists to remove. - Withdrawing at 10 s from the resume races the still-running session: Codex can write the frozen entry's hash after the withdrawal, and for a converted entry that marks the older command Orca put back as modified. - A resume cannot use the managed home: the session lives in ~/.codex. So the only states that cannot race Codex are the grant's own settle. The cost is a longer worst case on a cold app-server (up to the 30 s deadline, plus any managed-home install that holds the config.toml lane); a warm approval takes seconds, and an approved entry costs no wait. * fix(codex): keep the 5-minute trust cooldown for launch-path grants The fold shortened the host's trust-grant cooldown from 5 minutes to 10 seconds for every grant. That was meant for the background ~/.codex approval, which blocks no launch. The managed-home and WSL grants run inline on the launch path, so with a hung app-server every launch more than 10 s after the last failure paid the full inline timeout again (10 s native, 30 s WSL). Cooldowns are now kept per lane: inline grants keep 5 minutes, the background grant retries after 10 s, and neither lane's failure cools the other down. A success, or a proven-missing surface, still clears both. The real-home install's own retries (an unreadable hooks.json, unknown keys) are back on the 5-minute interval they had before the fold. The cooldown moves to its own module so the grant stays within the file limit. * fix(codex): a failed grant withdraws the exact copy it wrote The withdrawal re-found "this call's" entry by command, taking the first frozen handler in the event. When app start converted a later slot while an earlier frozen copy sat in a matcher group (which conversion skips), a failed grant acted on that earlier copy: it put the older command into it, or skipped it, and left the converted, unapproved copy in place. Each write now records where its handler landed, after any duplicate drops, and the withdrawal acts only on that slot. A copy that has since moved is left alone; the next launch's grant retries it. * fix(codex): the failed-grant withdrawal checks hooks.json is unchanged before writing The install and the retired-form sweep both refuse to replace ~/.codex/hooks.json if it changed since they read it. The withdrawal did not: a save landing between its read and its atomic replace was lost. The window is small, since the withdrawal is synchronous, but it now carries the same guard. * refactor(codex): drop rationale left over from the snapshot restore; name the trust-move module for what it does Comments on the config.toml lanes still justified them by a grant's capture-and-restore window, which the fold deleted, and the trust-write deadline still counted a grant session holding the lane. They now give the reason that remains: Orca's own multi-step reads and writes, and managed-home installs that hold the lane across their inline grant. codex-user-hook-trust-rebase no longer rebases through Codex; it moves stored trust records, so it is now codex-user-hook-trust-moves. The grant test that pinned two sessions on one config.toml to run one at a time is removed: its reason was an interleaved capture and restore. Callers that write config.toml around a grant hold their own lane, which the nested installer test still covers. * build(cli): list the trust-grant cooldown module in the CLI program The CLI's agent-hooks handler loads the hook controls, which reach the Codex trust grant; the CLI project is composite, so every module in that graph must be listed. * docs(codex): say which Windows hosts each hook command form runs under Codex runs a hook under the turn's shell (PowerShell 7 or 5.1 in every captured session) and, with no single local turn shell, under %COMSPEC% /C. The bare forward-slash path ran under all three in the Windows host census. The PowerShell form used for a profile path with a space does not parse under cmd.exe; no form valid in all three hosts has been run for such a path, so the form stays and the gap is stated here and in the PR. * test(codex): type the withdrawal seam without an assertion * fix(codex): a real-home resume starts at once, trusting Orca's entries for that process A resume that must run in ~/.codex waited for Codex's background approval of Orca's newly written hook entry: up to 30-40 s on a cold app-server. That made the user's resume wait on bookkeeping, and the alternatives (start at 10 s with Codex's hook review showing, or withdraw the entry and race Codex's own write) were worse. Codex reads hook trust from its session-flag config layer as well as the user's config.toml, merged per key, and has since hook trust shipped. So the resume no longer waits. When Orca's own frozen entries in ~/.codex are untrusted (or hold a stale hash), the resume command carries `-c hooks.state={'<key>'={trusted_hash='<hash>'},...}` for exactly those entries: the key under both the logical and the real path of ~/.codex (Codex keys an explicit CODEX_HOME by its real path), and the hash of that entry's content, so it can trust nothing else at that slot. The user's hooks are never included, nothing is written, and the background approval still runs for later plain `codex` launches. An approved entry adds nothing; a Codex known to lack hook trust gets nothing. One inline table, because Codex splits a `-c` key on every `.` and the key holds `.codex/hooks.json`. TOML literal strings keep `"` out of Windows native-argument quoting. The flag goes before `resume <id>`, quoted for the pane's shell (portable Unix, PowerShell or cmd), in the launch command and in the setup-sequenced copy of it; a cmd line whose path cmd would expand, or a key with an apostrophe, is left unchanged. SSH and WSL resumes get no preparation, so no local path reaches them. * Revert "fix(codex): a real-home resume starts at once, trusting Orca's entries for that process" This reverts commit |
||
|
|
2be67ce891 | fix(source-control): show git history commit times to the second (#23954) | ||
|
|
afa81dc3ad |
fix(native-chat): chat failure messages appear in the app's language (#23674)
* fix(native-chat): a read whose history will not open is refused with its reason
History, subscribe, snapshot and options reads reach a chat through one accessor, whose open had no
catch: a journal that would not open reached every client as a runtime error carrying the storage's
own text (a path, "file is not a database"). The accessor, and the options read's own open, now throw
the classified journal refusal: journalCorrupt when SQLite reports damage, journalUnavailable
otherwise. The storage text goes to the log only.
The wire code stays runtime_error and the message becomes the bare code, as for every thrown
refusal; the reason rides in the error's data.
* refactor(native-chat): the idle sweep's stop of a hung start carries no hand-written reason
The sweep passed an English sentence as the stop's reason. It lived only in memory and nothing read
it: the delivery loop words the error row and the rejection from the hostStopped fact. Dropped, with
the display-name lookup that built it.
* fix(native-chat): a refusal the host throws is worded from its data, never its message
A thrown agent-session refusal reaches the client as runtime_error with the bare code as its message
and the typed refusal in error.data. Stop, answers, options and goals, a launch's held option pick,
the option picker's failure toast, and the Retry line of a chat that could not start now word it
from that refusal through the shared notice table. What each caller decides about the outcome is
unchanged: only the words move.
The Retry line of a failed start no longer prints the host's message or a thrown error's text; it
keeps the refusal as a fact and says the cause and step its reason names, or only that the chat
could not be started. The option toast keeps a local option surface's own sentence.
* fix(native-chat): an unreadable history is worded from its refusal, and damage stops the retry
The structured chat's read failure showed the host's text on the status line, and the pane always
said Orca keeps trying. The read transport now takes the refusal from the error's data (a stream
payload or a thrown RPC error), the reducer keeps it beside the failure text, and the pane and the
status line word it through the notice table, once: on the pane when the failure took it, else
beside the transcript that stays.
A damaged journal says "Unable to load this chat." and the read stops reconnecting for that run;
reopening the chat reads again. An open that can clear names its cause without "Try again", since
the pane retries on its own. A failure that names no reason keeps today's generic line. Finality
comes from the refusal's reason, never its message, which is the bare code for both.
* fix(native-chat): a rejected message is worded from its stored fact
A message the host recorded and then rejected keeps the host's typed fact beside its reason, but
the Retry words re-read the reason alone. Now the fact decides: a hand-over failure says Orca
couldn't reach the agent, a kind whose reason may be a legacy marker gets its fact's own sentence
(a full queue now says so instead of only "not sent"), and any other kind shows the sentence the
host wrote for it, which carries the agent's name and any words the provider wrote for a person. A
row with no fact reads as before.
* fix(native-chat): each message that did not go through says why on its own row
The structured chat showed one Retry strip under the transcript for whichever single entry it
picked, so a second failed message had no reason and no Retry of its own. The terminal-backed
chat's existing per-row delivery marker now carries a notice and an optional Retry, and the
structured pane derives one per message from the outbox on each render: every rejected message,
and the one the queue stopped on (read through the drain's own rule, so a Retry never names a
message waiting behind it). Each is worded from that message's stored failure. The single strip is
deleted. Nothing new is stored, and the shared message projection is untouched.
* fix(native-chat): say each chat failure's words where its own control already acts
Three wording rules for the desktop chat:
- A chat that could not start shows Retry beside its reason, so the reason stops at its cause
where the Retry is the step: a reason whose action is to retry, and a start failure's "send
your message again". Any other step stays (quit the terminal agent, start a new chat). The
start-failure sentences take a retryControl context for this; what the host writes is unchanged.
- A history that couldn't open right now still reconnects, so the pane keeps "Orca keeps trying
to load it" under its cause. Only a damaged history, which no retry reads past, drops it.
- A read failure that names no reason while the transcript is shown is only the pane
reconnecting: the status line says "Reconnecting to this chat…" in muted text, not an error.
New key components.native-chat.state.reconnecting, hand-translated for es/fr/ja/ko/zh.
* fix(native-chat): a rejected message offers Retry only once the queue is moving
Each rejected message's row offered its own Retry even while the queue was stopped on another
message. Any Retry clears the stopped queue, so pressing a rejected message's Retry also sent the
message the queue was holding, which the user had not retried; behind a message whose delivery is
unconfirmed, the retried one instead went back into the queue with no notice and waited there.
While the queue is stopped, only the message it stopped on offers Retry, as the single Retry strip
this replaced did. A rejected message keeps its words on its row and gets its Retry back once the
queue moves.
* fix(native-chat): a chat whose history will not open logs once, not on every reconnect
A reader reconnects every 750 ms while a journal open can clear, and each attempt logged the
failure with its full stack. The read door now logs a session's failure once until that session
opens, closes, or fails differently; every attempt is still refused with its reason.
* fix(native-chat): a message's own Retry is its resend step, so its notice stops at the cause
A rejected message offering Retry read "Claude stopped before it finished starting. Send your
message to try again." beside that button. Its row now takes the rule the launch strip already
follows: beside its own Retry the words leave out sending or trying again, worded from the stored
fact with the chat's agent name. The stored fact keeps less than the host wrote from (a refusal,
the provider's words), so a reason it cannot rebuild exactly is kept as written. A rejected
message without a Retry, while the queue is held, keeps the step. What the host writes and the
phone's notice are unchanged.
* fix(native-chat): a message's Retry sends only that message, never the one the queue is held on
Retry released the queue's refusal hold whichever message it was pressed on. While a queued message waited ahead of a held one, every rejected message offered Retry, and pressing it also sent the held message the person had not retried. Retrying an unconfirmed message ahead of a held one did the same. Retry now releases the hold only for its own message.
* fix(native-chat): a not-signed-in failure beside Retry still says to sign in first
Beside a Retry the notice dropped the whole next step, so a chat that could not start because the agent was not signed in read only the cause. Pressing Retry without signing in fails the same way again. The words now keep the sign-in step and leave out only the resend, which the Retry button is.
* test(native-chat): a rejected message's hidden Retry only avoids waiting unseen
* fix(native-chat): every Retry beside a notice leaves out the retry step the same way
A message the queue stopped on worded its refusal with no agent name and with its retry step, beside its own Retry, while a rejected message next to it named the agent and left the step to the button. The launch strip and the history pane each had their own copy of the same rule. One wording context now goes through the one notice table for every surface: a Retry beside the words, or a pane that reconnects on its own, is the step for a reason whose action is to retry, and every other step stays. What the phone and the host write is unchanged.
* test(native-chat): read the sent message id without a type assertion
* fix(native-chat): a rejected message is worded from the journal's own fact, never by comparing sentences
A message the host recorded and then rejected kept only the rejection's kind on the message, so its notice was reworded from that smaller copy only when it rebuilt the host's sentence word for word. A different agent name, an older host's wording, or anything the copy dropped (why a start failed, the provider's own words) left the host's sentence in place, beside a Retry that repeated its resend step. The notice now reads the journal's own rejection for that message, found by id, with the pane's agent name and Retry, and shows the provider's words only when they were written for a person. The message's smaller copy words it only when that journal row is not loaded. Nothing new is stored.
* fix(native-chat): a message rejected before a restart retries under a new id the first time
Whether a Retry needed a new message id was remembered in memory for one message, or read from the journal row when it was loaded. After a restart, or for an older message whose row was not loaded, the first Retry resent under the old id, the host answered with the same settled rejection, and nothing visibly happened. The message now says so itself: one the host recorded and rejected always retries under a new id, including after a restart. A refusal that already gave the message a fresh id, and a message whose delivery is unconfirmed or in flight, keep their id as before.
* fix(native-chat): a rejected message older than the loaded history keeps the provider's words
When the journal row that rejected a message is not loaded, the message's own copy of the
fact has no provider detail or start refusal. For the kinds worded from those, the row now
shows the sentence the host wrote for the person instead of a thinner rebuilt one.
* test(native-chat): the chat pane words a rejected message from its loaded journal row
Nothing covered the pane handing the journal's rows to the per-message notices, so a pane that stopped passing them would quietly fall back to the message's smaller copy of the rejection and show the host's sentence, resend step and all. The new case renders the pane with a rejected message whose journal row is loaded and checks that it reads that row's refusal in the chat's own agent name.
* fix(native-chat): a chat whose history won't load says why in one line
A read the host refused for a named reason put its sentence under the generic
"Could not load conversation" title, so a damaged history read as two lines
saying the same thing. The pane's own sentence now takes the title's place; a
history that can come back keeps its line saying Orca keeps trying. A failure
that names nothing keeps the generic title.
* fix(native-chat): a message a failed start rejected says only that it was not sent
When an agent stopped before it finished starting, the chat showed the start's
red row ("Claude stopped before it finished starting. Send your message to try
again.") and then repeated that cause under every message the start rejected.
Each of those messages now reads "Your message was not sent." beside its Retry.
The match is made on typed facts, not on the words: the host writes the start's
row and the rejection of its queued messages from the same failure fact, and the
row is keyed by the start. The pane finds the loaded start-failure rows by that
key and shortens a message's notice only when its loaded journal submission was
rejected with the same fact. Any other rejection, or one whose submission or row
is not loaded, keeps its full notice. The row key moves to a shared module so the
host that writes it and the pane that reads it use one definition.
* fix(native-chat): a chat whose history keeps failing to open retries less often
A read the host kept refusing (its history store could not be opened right now)
reopened every 750 ms for as long as the chat stayed open, about 40 opens every
30 seconds. Each reconnect now waits twice as long as the last, from 750 ms up
to 30 seconds, and never gives up; the first read that delivers anything starts
the wait over at 750 ms. A damaged history still stops reconnecting at once.
Reset happens on a delivered read, not on connect: a local subscribe resolves
before the host's open refuses, so resetting there would keep the 750 ms loop.
* test(native-chat): the pane harness types its journal rows without a cast
* test(native-chat): the admission test passes no start-failure rows to the notices
* test(native-chat): import the journal types once
* fix(native-chat): a remote chat reads again as soon as its host is back
The read retry doubles its wait up to 30 s during an outage, and nothing
reset it when the remote runtime reconnected, so the transcript could
lag the reconnect by up to 30 s. The read now watches the runtime
status store's contact-regained edges (hostContactEpoch for a
same-runtime return, connectionGeneration for a new runtime session)
and, when one lands, runs a waiting retry immediately with the wait
reset to its base.
* refactor(native-chat): build each failure sentence from whole pieces
Every sentence agentSessionFailureWords writes is now assembled from a
table of whole English pieces, so a reader can supply its own words for
each piece. The host still fills them in English, byte for byte as
before.
* fix(native-chat): translate the failure sentences desktop notices show
A refused start and a rejected message now carry their failure fact to
the notice instead of its English sentence, and desktop words that fact
through translate keys whose English defaults are the host's own
pieces. The host keeps writing English into rows and reasons, and a
host sentence with no fact beside it still shows as written.
* fix(native-chat): the history pane says only that Orca keeps trying
When the pane's title already says Orca couldn't open this chat's
history right now, the line under it no longer repeats that the
transcript could not be read; it says only that Orca keeps trying to
load it. The pane with no named reason keeps its two-part line.
* test(native-chat): type the failure pieces a refusal notice shares
* test(native-chat): the Chinese failure words use no Japanese-only characters
* fix(native-chat): every history pane that says it didn't load says only that Orca keeps trying
A pane whose title is a code's own words ("This chat's history couldn't
be loaded.") now gets the short retrying line too. Only the pane with no
named reason keeps the two-part line.
* test(native-chat): a provider's words with nesting and markup stay as written in a translated notice
* fix(native-chat): French and Spanish say a withdrawn message was withdrawn before the agent began working on it
* fix(native-chat): Japanese and Chinese notices run their sentences on without a space
A notice joined its sentences with a space in every language, so Japanese and
Chinese read "Claude 无法启动。 请重新发送消息。" with a stray gap after the full
stop. The failure-sentence builder now takes the joiner alongside its words, and
desktop joins in the UI language: no space in Japanese and Chinese (including a
plugin pack that declares either), one space elsewhere. The host and the phone
keep English, joined with a space as before.
* fix(native-chat): a failed /clear or /compact says why in the app's language
The line under the composer printed the host's English sentence although the
result carries the typed failure beside it. It now words that failure the way
the host does (the chat's agent and /clear for a failed /clear, nothing for
/compact), in the app's language; an older host that sends no failure keeps
its sentence.
* fix(native-chat): Spanish says a rate limit, and Korean says a withdrawn message was never processed
The Spanish retry notice said the agent hit a usage limit, a different thing
from the rate limit the English names. The Korean withdrawn-message notice
said the agent had not started, which reads as the agent not launching; it now
says the agent had not begun processing the message, as the other languages do.
* refactor(native-chat): one rule says which words already say the history didn't load
* fix(native-chat): an image size limit says its unit the way the reader's language does
* test(native-chat): the option picker's i18n stand-in knows the reader's locale, which a refusal notice now reads
* fix(native-chat): a sentence a language pack left in English keeps its space
Sentences were joined by the UI language: no space in Japanese and Chinese,
one elsewhere. A plugin pack for a Chinese or Japanese variant that predates
the failure words falls back to English for them, so a notice read
"您的訊息未傳送。Claude couldn't start.Send your message to try again."
Each gap now follows the sentence before it: none after a full-width 。!?,
one space after anything else. The built-in Japanese and Chinese catalogs end
every sentence in 。, so they read as before, and English is unchanged. Since
the rule no longer needs the language, one joiner serves every surface and
the failure-sentence builder no longer takes one alongside its words.
* fix(native-chat): a failed /clear this build only partly understands shows the host's own sentence
The line under the composer words a failed /clear or /compact from the fact
the host sends beside its sentence. The reader drops any part this build
cannot place, such as a refusal code a newer host added, and the rest of the
fact can then give different advice: "Run /clear again." where the host said
"Start a new chat to continue."
When any part the host sent did not survive the read, the line now shows the
host's sentence as written, the same as for a host that sends no fact. A fact
this build reads whole is still worded in the app's language.
* fix(native-chat): a failure fact this build reads only in part shows the host's sentence everywhere
The previous check compared only a fact's top-level parts, so a known refusal
code carrying a reason a newer host added still counted as read: the reader
dropped the reason and the notice re-worded what was left, which can advise
differently from the host ("Run /clear again." against "Start a new chat to
continue.").
One shared reader now answers whether this build read the whole fact: it reads
the fact and keeps it only when the read equals what arrived, at every depth.
Every place that chooses between wording a fact and showing the host's text
uses it: the line under the composer after /clear or /compact, a rejected
message's notice from the journal's fact, and the smaller copy a rejected
message keeps for when its journal row is not loaded. Matching a rejected
message to the start row that already says why still uses what this build can
read, since that is identity, not wording. Facts this build reads whole are
worded as before.
* fix(native-chat): a failed /compact names /compact as its next step on desktop too
The host names the command a failed start was waiting on, for /clear and /compact alike.
The line under the composer re-worded only /clear with it, so a /compact whose start failed
read "The agent couldn't restart. Send your message to try again." in the reader's
language. It now words every command the host answers with the agent and the command the
host used, so desktop English matches the host and French or Japanese keep /compact.
* fix(native-chat): a /compact whose start failed no longer says the operation was not confirmed
A /compact on a chat whose agent is not running starts it first, and that start takes a new
lease, so the chat's fence moves before the command's reply arrives. The write settles as one
for a fence this pane no longer shows, and the composer line read that as "Conversation
operation was not confirmed." Such a write now says nothing there, as every other write
already does: the chat's own start-failure row says why, and the command's message is
rejected in the journal. The sentence it printed is gone from the catalogs.
* fix(native-chat): keep the message outbox within its line limit after main's growth
* fix(native-chat): a command's own reply is kept when its start moved the fence
A /clear or /compact on a chat whose agent is at rest starts the agent first, and that start
moves the chat's fence before the command's reply arrives. Every reply from an earlier fence was
discarded, so a /clear whose new chat failed to start said nothing at all, and a /compact that
started left "/compact" in the composer.
A conversation command's reply is now kept while the pane still shows the chat it was sent for;
a closed pane or another chat still drops it, and every other write keeps the fence rule. The
line under the composer says nothing only when the failure is the chat's own start and that
start's loaded row already says why, the rule a message that start rejected already follows.
* fix(native-chat): a returned queued message shows the host's sentence for a fact read in part
The caption under a returned queued message re-worded its failure from whatever this build could
read of the fact. A newer host's fact with a refusal code or reason this build drops read as a
shorter sentence with different advice. It now re-words only a fact read whole, and otherwise
shows the host's own sentence, as every other surface that re-words a fact does.
* test(native-chat): a command reply after any fence move, for /clear and /compact
The fence-move tests now state the rule as the code has it: a conversation command's reply is kept
whenever the pane still shows its chat, whatever moved the fence. They cover a /clear that
completed, and a failed start for each command in French with the fact that command really meets
(a /clear's new chat fails to start; a /compact's chat fails to restart).
* fix(native-chat): only the reply a pane still waits on outlives a fence move
A conversation command's reply was kept across a fence move whenever the pane still showed the
same chat. A reply the pane had stopped waiting on, because it left the chat and came back or a
newer command replaced it, was applied as if it answered the current one.
The pane now remembers the one command request it waits on; only that request's reply is kept
after the fence moves, and closing the pane or showing another chat forgets it. Tests now drive the
fence move during the request itself rather than through /clear starting an agent, and cover a
/clear that stops a running agent.
* fix(native-chat): a newer-Orca history error keeps a whole retry line, and the phone shows the host's words for a fact it reads in part
When a chat's history was saved by a newer Orca, the pane's title reads "Chats were saved by a newer
Orca. Update Orca to keep using them." and the line under it said only "Orca keeps trying to load
it.", with nothing for "it" to mean (in French and Spanish the pronoun also disagreed with "chats").
That title names every chat rather than this one, so the pane keeps the full line: "The transcript
could not be read. Orca keeps trying to load it."
On the phone, a returned queued message whose failure fact this build reads only in part was
re-worded from what it could read, dropping advice the host gave; it now shows the host's own
sentence, as the desktop card already does.
The comment on the reply the pane waits on now says what forgets it: a newer command, or disabling
the pane.
* fix(native-chat): a command reply this build can't place shows the host's own words
A newer host can answer a conversation command with a command name this build doesn't know. This
build re-worded that reply from its failure fact as if it were a command it knew, naming the
command in its own words, or said nothing when a loaded start row matched the fact. It now shows
the host's sentence as written, as it already does for a fact it can read only in part.
|
||
|
|
7cae036ebf |
fix(native-chat): a failed Codex turn shows its error, not a raw thread-status row (#23704)
* fix(native-chat): a failed Codex turn shows its error, not a raw thread-status row
Codex reports `thread/status/changed` {systemError} just before the `error`
frame of a failed turn. The translator read it for session state, then let it
fall through to the generic-frame fallback, whose payload check reads
`systemError` as a failure and printed "codex · notification:thread/status/changed"
in red above the real error row.
The translator now owns the notification: it reads the stopped-running verdict
exactly as before, then journals nothing for any arm (idle, active, notLoaded,
systemError). The error row that follows still carries Codex's sentence and
still fails the turn.
* refactor(native-chat): the Codex thread status is owned through the typed-translator registry
The translator already reads every thread status for session state. Listing
the kind beside Claude's background-task frames makes that ownership visible
where the fallback consults it, instead of a second method check in the
translator.
* chore(native-chat): the typed-translator registry key is checked against the provider table
A mistyped provider key compiled and silently brought the row back. Also
retire the renamed constant from a Claude test comment and say which path
the classifier-only fallback test covers.
* fix(native-chat): the Codex translator vouches only for the thread status it read
The covered flag means this exact frame was handled; passing it for every
notification would silently hide any Codex kind later added to the registry.
Also say why a state-only kind may be listed, and retitle the cross-provider
coverage test now that Codex has entries.
|
||
|
|
007b7c0d32 |
fix(claude): end a message Claude started but never confirmed, and keep Claude running while it holds one (#23898)
* fix(claude): settle a queued send the CLI withdrew from its own cancelled frame Claude reports each uuid-stamped command's lifecycle (queued, started, completed, cancelled). A send it withdraws from its queue gets `cancelled` before the interrupt or cancel_async_message answer, so a lost or failed answer no longer leaves that send pending: it settles as withdrawn, with the same reason and words as the receipt path. A command the CLI already started also ends `cancelled` when its turn is interrupted or fails, so `cancelled` after `started` is not a withdrawal; an echoed send has left the waiter lists and is never reached. Tests replay real 2.1.280 captures, scrubbed. * fix(claude): release a doubted send when the CLI reports its session idle A Claude send whose write ended in doubt is recorded `unknown`, and a live `unknown` reads as work still owed, so the chat showed Working until the child exited. Claude sends `session_state_changed idle` only once its whole queue has drained, so it can no longer be holding that send. The runtime now routes that report to the host's existing release, the same one Codex's thread-stopped report uses; it retires `unknown` only, never `pending`. * fix(claude): keep a command's started mark when a redelivery re-emits queued; fixtures name msg_lifecycle_v1 * fix(claude): settle every terminal lifecycle state of a send the CLI never echoed A send the CLI started, then cancelled before any echo, stayed pending: it may already be in the conversation, so it is released as doubt (unknown, recovered), never withdrawn and never re-sent. A late echo still accepts it. The 2.1.280 schema has two more terminal states. `discarded` (the CLI ended its session with the send still queued) settles as not delivered; `refused` (declined before it queued) settles as not accepted by the provider. After `started`, either one is doubt, as `cancelled` is. The late-settlement path gains an `unknown` outcome, which the host records as released doubt. * fix(claude): release a send the CLI took but left unanswered when it goes idle `session_state_changed idle` comes only once the CLI's queue has drained, so a send it took that is still unanswered there got no echo and never will: a turn that throws can leave `started` with no terminal state. Idle releases it as doubt. What proves the CLI took a send is its lifecycle frame. On a CLI that reports no lifecycle, it is the send's place on stdin: one whose write finished before an interrupt went out was read before the interrupt was, so the first idle after that interrupt releases it too. A send armed ahead of the interrupt but written after it is left alone, since the CLI may still run it. * fix(native-chat): keep the idle sweep off a Claude child that holds a send A Claude retrying a rate-limited request has taken the send but echoes nothing, so no turn row exists yet and the sweep rested the child after the idle window, turning the send into doubt. The adapter now reports whether the CLI holds a send (lifecycle `queued` or `started`, not yet echoed or ended), derived from the live waiters, and owed work counts it. Nothing is stored: every held send leaves the live set on its echo, its terminal lifecycle state, the CLI's idle, or the child's exit, so the hold ends with the send. * fix(claude): count only a started send at idle and as a held send 2.1.280's end-of-turn cleanup can report idle before it re-reads its queue, so a send read in that window goes queued, idle, started. Releasing every taken send at idle doubted that live send and dropped Working. Only a `started` send is released at idle or keeps the child from the idle sweep; a `queued` one ends by starting and echoing, by a terminal lifecycle frame, or with the child. The stdin-order path for CLIs without lifecycle frames is removed: a doubted send retired there disables content matching on CLIs that mint their own echo ids, and no Orca failure called for it. Those CLIs keep the earlier behaviour. Comments that said only a failed write or child exit ends a waiter, or that idle comes only once the queue has drained, now say what ends one. * docs(claude): say only what the CLI's lifecycle frames and idle actually prove * fix(claude): hold the idle sweep while Claude has a send queued, not only started The sweep rested a child whose CLI had queued a follow-up behind a turn, dropping the send it had already taken. The hold now spans the CLI reporting it took the send until its echo, a terminal lifecycle state, or the child's exit. The idle release still covers only started sends: 2.1.280 can report idle before it re-reads its queue. * refactor(native-chat): give provider-proven late dispatch settlement its own module * test(claude): pin a steer a Stop interrupts after it started as doubt, not withdrawn |
||
|
|
ed462b2caa |
fix(relay): re-place hosts off a draining cell without locking its row (#24216)
* test(relay): reproduce drain-release contention against director placement Adds a Postgres harness that drives releases from an isolated cell over a 171 ms per-statement pool while five directors re-place reconnecting hosts through the sticky lane. At 18 releases/s placements fall from 20/s to about 5/s and every active director backend is blocked on relay_cells. Moves the per-statement delay pool into a shared test fixture so the rehome target-row test and this harness use one implementation. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): re-place hosts off a draining cell without locking its row Sticky re-placement of a host whose cell is isolated for a roll locked every relay_cells row. During an Asia drain the source row is held by the cell's own releases for a round trip each, so the placement waited on it, the single sticky slot backed up, and /v1/assign returned 503 fleet-wide. The isolation decision now comes from an unlocked read, and the placement locks only the same-region general rows it can move to. It no longer writes the source row: the host's source leases stay, and each one's own release or expiry takes its units back off the source. The assignment keeps its activity counters and adds one control instead of resetting them. With no same-region headroom the path falls back to the all-rows lock, as before. Dormant hosts hold no units, so their placement also locks only the general rows and skips the zero write to their old cell. The drain harness now asserts the after picture: 20 placements/s at Asia latency with no director lock waits, against 9.2/s and 4.8/s before. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): take a moved host's units off the cell that holds them After a narrowed re-placement a host keeps leases on its old cell while its assignment row names the new one. Two paths charged the row's whole counted total to the row's cell: aggregate expiry, and the lease deletion in dead-cell and stranded re-placement. Both over-charged the new cell and left the old cell's units stranded. Aggregate expiry now skips hosts that still hold any lease; the lease sweep takes each lease's units off its own cell. Placement frees each deleted lease's units on that lease's cell, charges the old cell only for units no lease backs, and sets the counters from the leases it keeps plus the new control. The narrowed path runs only when the counters already match the leases, so it never needs to write the old cell's row. The all-rows re-placement off an isolated cell follows the same rule, so it no longer decrements the source at placement either. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): try other regions before the all-rows lock when re-placing off a roll With every same-region neighbour at its connection cap, the narrowed path found no target and fell back to the all-rows lock behind the busy source row, which is the drain brownout again. It now tries a second tier, general cells in every other region, in its own transaction over one ordered lockCellRows, still never the source row. Only when no general cell in any region has room does it fall back to the all-rows path, which keeps the pin. This changes the policy from #21911, which refused to re-place an isolated host across a region. The unit tests that encoded that rule now assert the tier order instead. The drain harness gains a US cell and an arm with every Asia neighbour capped: 200 of 200 dials placed cross-region at 20/s with no lock waits, against 0 placed and 144 sticky rejections on the previous head. Its pass bars are now the rejection share and the lock-waiting share, not the placement rate a slow runner's pacing can move. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): take the narrowed path for hosts whose counters sit below their leases Main's old placement reset a moved host's counters while keeping its source leases, and those leases' releases floored the counters at zero. Such hosts hold fewer counted units than lease units, and requiring equality sent them down the all-rows path behind the busy source row. The narrowed path now requires only that the host holds no units no lease backs, the one case that needs a write to the old row. Its placement already rebuilds the counters from the kept leases plus the new control, so a drifted host is healed by its next re-placement. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * test(relay): judge the local drain arm on completion and lock waits, not rate The local-latency arm asserted at least 18 placements/s at a 20/s dial rate, which a slow runner's pacing alone can miss. It now asserts what the Asia arm does: no dial failures, sticky rejections under 10% of dials, every other dial placed, and director lock-waiting under half a backend. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
95e8725b40 |
fix(claude): stop Orca making a WSL user's ~/.claude.json world-readable (#23973)
* fix(claude): keep a WSL guest's ~/.claude.json mode when Orca writes folder trust * test(claude): assert a skipped trust write leaves no replacement file behind The replacement file is now created before Claude's lock is taken, so the locked path must remove it. |
||
|
|
cfe4c633eb |
fix(mobile): publish the Android APK's size and checksum with the release (#24037)
An APK that fails to install with a missing certificate or a package-parse error is usually a download that died near the end: the signature block sits in the last ~100 KB of a 133 MB file, so a truncated APK looks complete and carries no signature at all. The release published neither a size nor a digest, so there was no way to tell that apart from a bad build without deriving both from the asset by hand. The release now uploads app-release.apk.sha256 next to the APK in `sha256sum -c` format (binary marker, so Git Bash cannot translate line endings while hashing) and puts the exact byte size and digest in the release body, naming `shasum -a 256 -c` for readers on macOS. The upload path rewrites the body too: --clobber replaces the APK, so a digest left over from the previous build would describe a file nobody can download, and a reader comparing against it would reject a good APK. Both paths reserve the section's own length out of the release-body cap before truncating, so the section always survives and the body always fits; MAX_RELEASE_BODY_LENGTH is exported from the desktop release script rather than restated. Refs #24011, #12248, #11444. |
||
|
|
2caa79e097 |
fix(source-control): drop reasoning model think blocks from generated messages (#24005)
* fix(source-control): drop reasoning model think blocks from generated messages Custom commit-message commands that run a reasoning model print the reasoning before the answer, either as a <think>...</think> block or, when the chat template prefills <think>, as text ending in a lone </think>. The cleaner kept all of it, so the first reasoning line became the commit subject. Strip everything through the first </think> when the output starts with <think> or has no <think> before the close tag. Output that quotes both tags is left unchanged. Fixes #24004 Signed-off-by: FenjuFu <fufenjupku@gmail.com> * fix(source-control): scope lone </think> stripping to custom commands and add Kimi-VL tags A lone closing tag is only stripped for custom commands, so a built-in agent's message that mentions </think> is kept. PR fields parse the raw JSON first and strip reasoning only when that fails, so a body quoting the tag still parses. Adds Kimi-VL-Thinking's ◁think▷ tags. --------- Signed-off-by: FenjuFu <fufenjupku@gmail.com> Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> |