Commit Graph
12 Commits
Author SHA1 Message Date
OrcaWinandm4air 8c2cd7d331 feat(ai-vault): read remote OpenCode history with the pinned Node; remove Bun (#24128)
* feat(ai-vault): read OpenCode history with the pinned Node instead of Bun

SSH hosts whose Node lacks node:sqlite (or its backup(), which 22.13-22.15
omit) now get the pinned Node in the shared ~/.orca-remote/runtimes/node-<sha>
store orcad uses: POSIX hosts receive the official archive and extract and
hash-verify it on the host; Windows hosts receive the verified node.exe the
client extracted, promoted by host Node with the same hash check. WSL distros
use the same layout and checks under ~/.cache/orca/runtimes/.

The Bun release pin table and its materializer are deleted. Old relays keep
reading their vault-sqlite/<sha>/bun references; nothing deletes those files.
An unconfirmed runtime upload now keeps its stage instead of removing it.

* refactor(sqlite): drop the Bun SQLite adapter; node:sqlite is the only backend

Nothing outside Electron runs on Bun any more (design D4), so SyncDatabase
loses its Bun branch, and bun-sqlite-database, bun-sqlite-statement and
bun-readonly-wal go, with the relay's bun:sqlite external. The profile-state
backup worker admits Electron or an entry that exists, and startup errors
name the pinned Node. The D7 cross-runtime gate still runs Bun 1.4.2, now
reaching Bun's SQLite through its node:sqlite.

* test(native-chat): drop the Bun SQLite driver case now that node:sqlite is the only backend

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 01:10:31 -07:00
OrcaWinandm4air 3fe4b18dae fix(ai-vault): require node:sqlite backup support in remote SQLite probes (#24086)
* fix(ai-vault): require the full SyncDatabase node:sqlite surface in host SQLite probes

The SSH and WSL OpenCode probes admitted any Node with DatabaseSync, so
Node 22.13-22.15 hosts (no backup export) skipped the pinned-runtime
fallback. Share one admission predicate with isSqliteAvailable() and
embed its source in both probe scripts.

* build(cli): list the node:sqlite admission predicate in the CLI project

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 22:57:01 -07:00
Brennan BensonandClaude a68d62911e fix(native-chat): the host writes chat failures for a person, with a typed fact beside them (#23116)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* feat(native-chat): a typed failure fact beside every failure sentence

Adds the shared vocabulary the host writes a failure with: a closed failure kind, a
provider diagnostic that says who it is for (a person, or a log), and a refusal cause
beside the refusal code. Status rows gain an optional failure fact and rejected
submissions an optional rejection fact; the dispatch row carries it, the reducer reads
it field by field, and the projection forwards it. Older rows and older readers are
untouched: every field is optional and the schemas stay open.

* fix(native-chat): durable failure rows and rejection reasons are written for a person

Every host writer that records a failure now writes a sentence for a person beside a
typed fact, instead of embedding a refusal's message, an exception or a composed exit
string. A provider's own words travel as a separate diagnostic from the places Orca
composes them - the Claude and Codex exit stderr (a log), Codex's JSON-RPC message,
Claude's compact_error and Codex's turn error (for a person) - and are never inferred
from a string afterwards. Not signed in and oversized history are typed at the
adapter that detects them.

Covers start and restart failures, the delivery loop, dispatch rejections (content,
queue-full, write failures, provider refusals), cancel and answer confirmation rows,
compaction, the rewind fallback, and not_delivered, which released clients printed
as it was. Two leaks close on the way: a settlement retry no longer writes Orca's
probe evidence into the exit row, and an attach or journal-sink failure is recorded
as Orca's fault rather than as the provider stopping. The legacy rejection markers
and the reasons on sends in doubt stay byte-identical.

* feat(native-chat): refusals name their cause, and a failed start is worded in one place

A refusal now carries an optional cause beside its code: one closed enum of the situations
a chat write can meet, set at every emitter a structured-chat write reaches. Returned
refusals build it with refuse(code, cause, message). Store and host paths that raised a
bare Error(code) now throw AgentSessionRefusalError, whose message is still the code and
which has no code property; the RPC error mapper handles it before any other passthrough,
keeps today's wire code and message byte-identical, and adds { refusal: { code, cause } }
to the error's data. The hold throws it, and restart-resume files the cause beside the
unchanged reason. The operation ledger stores the cause beside the code, so a replay names
the same situation as the first answer. The store fallback copy picks its words and cause
by situation, so a stale replay or a moved lease no longer reads as a latched owner.

Every failed start is worded by structuredAgentSessionStartFailure(cause, context), which
returns the row sentence and the typed fact together; the delivery loop, the exit
settlement and the dispatch that met a starting child all call it. Provider diagnostics
are capped at the lease record's 512 characters wherever a fact is built.

* refactor(native-chat): one reader of why a submission was rejected

classifyDispatchRejection(submission) returns { category, verdict, kind? }. It reads the
typed rejection fact when the row carries one this build can place, and the legacy
markers otherwise - all six, including not_delivered, which released clients printed as
it was. The verdict is null only for a withdrawal, a host restart and a closed chat;
write failures, a full queue and not-delivered stay failures. It replaces
dispatchRejectionReasonIsInternal and every string comparison against the markers: the
outbox reconcile, the rejection notice, and the send disposition, where a replayed
Stop-withdrawn send no longer surfaces as a failed send.

The journal reducer's echo-aliasing guard reads the narrow isWriteFailureSubmission,
which matches the typed kind or the legacy prefix in any dispatch state, so its behaviour
on legacy unknown rows is unchanged.

* test(native-chat): pin provider diagnostics where they are composed

The Claude exit status and stderr, the Codex stderr tail and Codex's JSON-RPC message are
each checked at the place Orca composes its own error around them, so the typed detail
is proven to come from the provider's value and never from Orca's wording.

* fix(native-chat): an attachment Orca refuses says which limit it broke

The content check's refusals (20 images, 5 MB per image, 20 MB in total, supported types) are
written for a person, but the rejection writer replaced them all with one generic sentence.
Each refusal now carries its own sentence, in MB rather than bytes, and the writer records it
beside kind attachmentInvalid. Only an attachment that could not be read keeps the generic
sentence.

* fix(native-chat): the chat tab table refuses with a typed cause

Showing a chat tab refused with bare Error('agent_session_conflict') and
Error('agent_session_identity_required'), the only chat-reachable refusals still thrown without a
cause (opening a chat from history can reach the second when the chat is removed mid-open). Both
now throw the typed refusal; wire code and message are unchanged.

* fix(native-chat): a compaction Codex refuses up front keeps Codex's words

When Codex refused thread/compact/start, the adapter passed on only Orca's wrapped error text and
dropped Codex's own message, so the chat's row read just "Compaction failed." The refusal now
carries Codex's message as the failure detail, as a compaction that fails later already did.

* fix(native-chat): an unreadable chat record no longer promises an update fixes it

The recordUnreadable copy said a newer Orca saved the chat, but the store marks a record
unreadable for damage and key mismatches too, where updating does nothing. The sentence now says
Orca can't read it and gives both next steps.

* fix(native-chat): a restart or a close leaves released clients a sentence, not a marker

host_restarted_before_delivery and provider_closed_before_delivery are not in the markers released
desktop and mobile builds hide, so they printed raw on every rejected message a restart or a chat
close left. Neither marker has shipped. New rows carry a sentence plus kind hostRestarted or
chatClosed, as not_delivered already did; the classifier still reads both markers, and the
verdict for both stays no-failure.

* fix(native-chat): log why an attachment could not be read

The rejected message now says only that the attachment couldn't be read, so the error that said
why (a missing file, a permission, or an unexpected throw) went nowhere. It is logged instead. The
content rejection moves beside the content check that owns its errors.

* refactor(native-chat): drop an unused thrown-refusal cause reader

It had no callers, and its comment claimed it looked through wrappers, which it did not.

* fix(native-chat): a compaction the provider never confirmed is recorded as unconfirmed, not failed

* fix(native-chat): an empty or non-user message is not recorded as a bad attachment

* fix(native-chat): an undelivered preamble's error ends in one period

* refactor(native-chat): a compaction ends as compacted, failed, or unconfirmed, never an unlabelled error

* refactor(native-chat): a failure's sentence is written only from its fact

A writer could choose its sentence and its kind separately, so five writers
put hand-written words beside a fact that said something else. One shared
constructor, agentSessionFailureWords(fact, { surface, agentName }), now
makes both, and the journal types refuse anything else: a status row, a
rejected message or a conversation command that carries a fact must carry
the sentence that constructor branded. Persisted rows keep their shape.

- The sentence table and the restart table move to src/shared, as the English
  default a client copy table can reuse.
- A rejection's legacy markers come from the same constructor. A write
  failure is now the bare `provider_write_failed` marker, which released
  clients already hide; its error goes to the log.
- An image Orca refuses carries which check it failed (and the limit) in the
  fact instead of a sentence; an empty message gets its own kind, and a
  non-user message is Orca's fault.
- The exit row, the interrupted compaction, the /clear failure and the rewind
  placeholder no longer carry their own words beside a fact: the exit row
  says what its fact says, and the placeholder carries no fact.

* fix(native-chat): a start that failed without an observed exit no longer blames the provider

Every untyped start error was recorded as "The provider stopped before it
finished starting.", so an Orca fault, a failed spawn or a close that ended
a start was blamed on the provider. Only an exit the adapter observed says
so now: the Claude adapter marks the error it saw the child exit with, and
anything else is a new `startFailed` kind, "<Agent> couldn't start.", keeping
the provider's diagnostic when the error carried one. A child gone with no
end observed, and a /clear whose new conversation was refused, read the
same way.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): a chat a terminal agent still holds says to quit that agent

The merged base words a restart a terminal agent's claim refused with the
refusal's own message, which names the process. That message is Orca's text
and never reaches a durable row here, so the row read only "<Agent>
couldn't restart." and lost the one step that frees the chat. The restart
sentence now derives it from the refusal's cause: a claimConflicted refusal
adds "This chat is still open in a terminal agent. Quit that agent to
continue the chat here." The live refusal still names the process.

The restart-resume ledger test now sets up a claim the base still refuses:
a terminal owner that is proven running.

* refactor(native-chat): keep the changed files inside their lint limits

The legacy-marker lookup is a table, not a non-exhaustive switch; the
preamble tests read the error without a cast; and the Codex history refusal
goes through a named constructor so its file stays under the line limit.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* refactor(native-chat): refusals carry details keyed by their code

A refusal named its situation with one flat cause list shared by every
code, so nothing stopped a site pairing a code with a situation that code
never means, and the loose fields released clients read (fence, revision,
resolution, verdict, rewind reason) were written by hand at each emitter.

A refusal is now one variant per code with optional details: a reason that
code lists plus that code's own facts. refuse(code, details, message)
rejects a reason the code does not list at compile time, and it is the one
place the loose top-level fields are copied from details, so released
clients read exactly what they read before. A site that cannot name its
situation uses refuseUnclassified, which carries facts but no reason, the
same as an older host; there is no catch-all reason.

Thrown refusals put { refusal: { code, details } } in the RPC error's data
(wire code and message unchanged). The operation ledger, the restart-resume
record and a restart failure's embedded refusal keep details beside the
code and read them back against it; a row an unreleased build wrote with a
cause parses and reads as naming none.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* test(native-chat): a restart-resume failure keeps its refusal details

The recovery capsule reads a failure's details back against the refusal
code in its reason: facts the code does not list and a reason another code
owns are dropped, and a record an unreleased build wrote with a cause
still parses, naming nothing.

* test(native-chat): import the failure words once in the provider child test

The merge left two imports of the same modules.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* fix(native-chat): a start the provider refused or Orca broke no longer says the provider stopped

A failed start whose cleanup proved the child gone was typed as
`providerStartFailed` whatever failed it: Codex refusing to resume a thread,
a timeout, or Orca's own store fault. The chat then read "The provider
stopped before it finished starting.", which was untrue, and the provider's
own words were dropped. A refused restart also inferred the same from an
`exited` verdict, which only says nothing runs now.

Only an exit the adapter observed names that situation now; anything else
is refused with no reason, keeping its verdict, and the chat reads
"<Agent> couldn't restart.". The provider's words travel host-side from
where the acquisition failed into the start-failure fact's detail, never
onto the refusal, and the sentence does not quote them. What failed is
logged once where the start failed.

* fix(native-chat): a refused /clear start keeps the situation it named

The replacement start that /clear makes built its own start-failure fact,
so a refusal that named its situation, such as not being signed in or a
history too large to restore, was recorded as a bare "couldn't start". It
now takes its fact from the same start-failure function as every other
start, as a new session that failed to start.

* fix(native-chat): the journal schema and comments describe a refusal's details, not its cause

The persisted failure fact's schema still described `refusal.cause`, which
this branch replaced with `details`. It now describes `details` as an
optional open object; a row an earlier build wrote with a `cause` still
parses. A comment and three test descriptions that still named the cause
now name the details.

* fix(native-chat): a Claude child that exits while being acquired still reads as the provider stopping

Now that only an exit the adapter observed says the provider stopped, the
exit the Claude adapter saw during acquisition has to be marked where it is
seen, as the exit after acquisition already is. Without the mark, a Claude
CLI that exited at spawn read "Claude couldn't restart." instead of "The
provider stopped before it finished starting."

* fix(native-chat): word a failed chat start's refusal as its start failure

A chat whose agent failed to start answered the create with the raw error: the launch strip read "Chat could not be started. claude stream-json exited (code 1): claude: not signed in", and the ledger replayed the same text. The first answer and the replay now carry the sentence the chat's start-failure row reads as ("The provider stopped before it finished starting.", "Claude couldn't start.", the not-signed-in and history-too-large sentences), and the raw error goes to the log. A store refusal's code and the unproven-exit marker are unchanged.

* fix(native-chat): show Claude's API retries as one sentence row

While Claude retried a refused request (a 429, say), the chat gained one red row per attempt reading "rate_limit", with the raw retry frame behind Details. Each retry run now writes one warning row that later attempts revise in place: "Claude is rate-limited and retrying." for a rate limit (error `rate_limit` or status 429), and "Claude hit a temporary problem and is retrying." otherwise. The row carries a `providerRetrying` fact with the provider's error type and status, and the frame as a log detail capped at 512 characters.

* fix(native-chat): tell the user to run /clear again when its new conversation can't start

When /clear's replacement conversation failed to start, the result told the user to "send your message again", which would go into the old conversation. The failure words now take the command the start was for, so a failed /clear reads "Codex is not signed in for the selected account. Sign in, then run /clear again.", "Codex couldn't start. Run /clear again." or "The provider stopped before it finished starting. Run /clear again." A message send keeps its wording.

* test(native-chat): expect a failed Claude create to be refused in a sentence

The runtime suites asserted the CLI's stderr reached the create refusal; it now goes to the log and the refusal reads as the chat's start failure.

* fix(native-chat): name the agent that stopped starting and say how to retry a failed start

A start the provider ended now reads "Claude stopped before it finished starting." (or "The agent ..." when the chat's agent is unknown) instead of naming "the provider". A start or restart that failed with a chat left to retry now ends in "Send your message to try again.", or "Run /clear again." for /clear; released clients print only this sentence. A message rejected at dispatch because its child died while starting names the same agent as the start's row.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): answer a create refused before spawn in the words its replay reads

A create that failed before any process started threw Orca's own error text as its first answer,
while its replay from the operation ledger read the generic start sentence. The two refusals a
person can act on, a launch whose Anthropic sign-in variables override the managed Claude account
and a Claude account switch in progress, are now typed where they are thrown and worded by the
shared failure constructor, so the first answer and the replay say the same thing. Every other
pre-spawn failure reads the generic start sentence, with its own text in the log. The first answer
keeps its wire code; a message that is itself a code is unchanged.

* fix(native-chat): say how to retry after an agent stopped before it finished starting

"<Agent> stopped before it finished starting." gave no next step outside /clear, unlike every other
failed start. It now ends "Send your message to try again.", the same step a start or restart that
could not run gives; after /clear it still says "Run /clear again."

* fix(native-chat): say how to reach a Claude chat when a WSL Claude account blocks it

A Claude chat that restarts while a Claude account is added in WSL and no Windows Claude account is selected was refused before spawn with Orca's own text as its first answer, and the generic start sentence on replay. The refusal is now typed where the account gate throws, and both answers read the same sentence: choose or add a Windows Claude account in Claude Accounts settings, then send the message again. Account settings that cannot be read name no situation and keep the generic sentence.

* fix(native-chat): say the reason a host names for a refused chat write

A refused Stop, answer, setting, goal, command or queued message now reads the refusal's reason as
well as its code. A code stands for several situations, so the code alone could only say what did
not happen; with the reason, the notice says why and, where the person has a step to take, what it
is: "The agent is still responding. The command didn't run. Wait for the agent to finish
responding, or stop it." The phone uses the same words.

The notice table keys on code, then reason, then the kind of write. Every reason of every code has
an entry, so a reason the host adds does not compile until it has words; a reason whose honest
words are its code's keeps the code's row. A refusal with no reason, or one this build does not
know, reads exactly as before, which is what an older host gets. A start that failed reuses the
failure row's own sentence rather than a second one.

A queued message keeps the reason, the rewind reason and the owner's verdict with its saved
failure, and a rejected one keeps the host's typed fact without its provider detail. Nothing that
moves with the owner or comes from the provider is saved; entries saved before this load as they
were. The saved failure moves to its own module beside the words chosen from it.

* fix(native-chat): retry a refused send under a new id only once its agent is proven gone

A send refused because Orca could not tell who owns the chat keeps its operation id, since the
first attempt may still land. When the refusal also says the agent process has exited, nothing can
run that attempt, so the message moves to its Retry row under a new id instead of holding the queue.

The owner's verdict is read as a floor: a saved `exited` is final, and any other saved verdict never
changes the id on its own. Only a verdict re-derived from the current lease can raise it to
`exited`, and nothing lowers a saved `exited`. Today no host sends a verdict on a send refusal, so
nothing a person sees changes; the rule is in place for the saved verdict a reload reads back.

* fix(native-chat): say a /clear that never finished did not finish, instead of that it cleared the chat

A send into a chat whose /clear started but never committed was refused as if the conversation had been cleared: "This conversation has been cleared. Your message was not sent. Open the current conversation to continue." The clear never finished, so its new conversation may not exist and there is nothing to open. That refusal now has its own reason and reads "The last /clear didn't finish. Your message was not sent. Start a new chat to continue.", which is the only way on today. Its message for released clients says the same: "The last /clear didn't finish. Start a new chat to continue." Only a committed /clear still says the conversation was cleared.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send the provider's history
proves it never received. Nobody failed that send, but the verdict table
treated it as a failure. Each rejection kind now has its verdict in one
exhaustive table, so a new kind does not compile until its verdict is chosen;
no verdict for a withdrawal, a host restart, a chat close, or this lost send.

A rejection whose kind this build cannot place, such as one a newer host
added, now reads as undelivered with no verdict instead of falling back to the
reason beside it: all it proves is that the message did not happen. The host
keeps such a fact's kind when it reads the row back, rather than dropping it
and letting the reason decide.

Only kinds that can be why a message was not sent may reject one, by type:
compaction, cancel/answer confirmation and provider-retry kinds stay on status
rows. The one dispatch-row builder takes its input from the type that makes a
rejected row carry its fact.

* fix(native-chat): a send refused after its agent exited keeps its place in the queue

When Orca cannot tell who owns a chat but the refusal says the agent process
has exited, the next attempt may use a new id, since nothing can run the old
one. It no longer marks the message as rejected: nothing recorded it, so it
still holds the head of the queue, and later messages wait behind it instead
of being sent ahead of it.

* fix(native-chat): stop reading a provider's words from an error that contains itself

A cleanup that aggregates errors restarted the depth count for each one, so an
aggregate error that contains itself recursed until the host ran out of stack.
One depth bound now covers both the cause chain and the aggregated errors.

* fix(native-chat): say a chat whose history can't be read can't continue, and to start a new one

A read of a chat's history is refused with `agent_session_journal_unreadable` only when the chat's journal file is corrupt or not a database, which no retry can change. The notice table had no words for a read at all, so a pane had nothing to show but the raw code. Reading a chat's history is now its own request kind, `read-history`, and that refusal on it reads "This chat's history couldn't be read, so it can't continue here. Start a new chat to continue.", whether the host names the reason or raises the bare code. A write refused under the same code keeps its words, because its cause is any failed open, which can clear. Any other read refusal says only "This chat's history couldn't be loaded." `agentSessionReadHistoryRefusalParts(code, details)` gives a pane those words from a read error.

The notice sentences move to their own module so the table stays within its size limit.

* fix(native-chat): decide what every way a child ends means for queued messages in one table

What a child's end means for the messages queued behind it was an if-chain: a user's Stop was checked in one place, a host stop in another, and every other cause, including one added later, fell through to "the provider exited". It is now one table over every end cause, so a new cause does not compile until someone says whether it fails what is queued and how. A user's Stop still fails nothing, a host stop is still Orca's fault, and an exit, a failed attach or an eviction still carry the failure the end recorded. Nothing a person sees changes.

* test(native-chat): a close that stops the child and then fails rejects what is queued by how the child ended

When a chat closes, stops its agent, and then fails a later step, the chat stays open with its messages still queued. The delivery loop then rejects them by the way the child ended: an eviction during startup reads as a failed start, otherwise as the provider having stopped, and either counts as a failure. The cases are rows over the end cause, so another way a chat closes is one more row.

* test(native-chat): type the refusal a persisted-schema test admits

* fix(native-chat): say whether a chat's history is damaged or just couldn't open

A write refused because the chat's journal would not open named one reason, `journalUnreadable`, for every failed open, and a read of the history took that same reason as final. So the words depended on what was asked, not on what happened: a busy or permission-denied open could tell a person to start a new chat, and a damaged one could read as something that clears.

The host now decides at the refusal which it was. `journalCorrupt` is set only when SQLite itself reports the journal damaged or not a database (SQLITE_CORRUPT or SQLITE_NOTADB, extended codes included, read from the driver's result code and never from message text, through any `cause` chain). Every other failed open is `journalUnavailable`. A corrupt history reads "This chat's history couldn't be read, so it can't continue here. Start a new chat to continue.", with "Your message was not sent." before the step on a send. One that couldn't open reads "Orca couldn't open this chat's history right now. Try again.", likewise on a send. A host that names no reason gets "Orca couldn't read this chat's saved history.", which promises neither, because damage can't be proven from the code alone.

The refusal's message, which released clients print for a send, is now that person sentence instead of the open error's own text; the error is logged instead. `journalUnreadable` is replaced outright: no released build wrote it, and a stored one reads as a refusal with no reason.

* docs(native-chat): say why a history that couldn't open names its retry step

* fix(native-chat): a failed start's row keeps the words its rejected messages carry

When the delivery loop settles a failed start before the child's exit is published, it writes the
start's error row and rejects every queued message with the adapter's startup answer. The exit
settlement then rewrote the same row from the exit event, so the row could say one thing while the
rejected messages, which are terminal, said another. The exit now leaves a start's row it finds
already written.

* fix(native-chat): a Claude start Orca itself failed no longer says Claude stopped

Every error that ended a Claude session was marked as an exit the adapter
observed, including a start Orca failed while the CLI was still running: a
saved option whose restore lost its answer, an init frame naming another
session, or a journal write fault. Those read "Claude stopped before it
finished starting." although Claude never stopped on its own. The mark now
stays where the child's exit is seen (the connection's exit callback), so
those starts read "Claude couldn't start." with any diagnostic beside it,
and a real exit before the start lands still says Claude stopped.

* fix(native-chat): name the agent that stopped, and blame Orca for its own closes

A chat that lost its agent mid-response said "The provider stopped…", and a
started Claude session that Orca itself closed after a journal fault said the
same, as if Claude had exited on its own.

The exit row and the rejected-message reason now name the chat's agent ("Claude
stopped while this response was in progress…", "Codex stopped before this
message was sent."), or "The agent" when the name is unknown; the stale-state
settlement now passes the agent name too. After a Claude start has landed, the
ended event reports providerExited only when the child's own exit was observed;
any other close is Orca's fault and reads as one.

* test(native-chat): pin the sidebar verdict to the rejection classifier for every kind and legacy marker

* fix(native-chat): blame Orca, not Codex, when Orca closes the Codex child

A Codex chat that Orca itself closed (a journal sink that could not take a
frame, or a forced close) said "Codex stopped while this response was in
progress", as if Codex had exited on its own.

Orca's own close path now reports hostFault. providerExited is left to the
app-server connection's exit callback, which the connection withholds while Orca
is closing the child, so it only ever reports the child's own exit.

* fix(native-chat): a chat whose history is damaged reads "Unable to load this chat."

* fix(native-chat): route the conversation-outlives-agent writers through the typed refusals

Three writers that arrived with the merge wrote refusals the old way:

- An operation that starts the agent itself, such as a goal change, turned any error the start
  threw into a refusal whose message was Orca's own error text, which released clients print.
  It now logs the error and says only that the agent couldn't restart.
- An option picked while the chat is at rest, for a key the provider would not accept, is refused
  with the rejected-option reason, like the same pick on a running agent.
- A restart continuation whose agent was refused a start filed the rejected message's sentence as
  the failure's reason. It files the refusal's code with its details again, which is what the
  restart-failure guidance keys on.

* fix(native-chat): a start Orca stopped because it never finished reads as that

The idle sweep now stops an agent whose start never finished and rejects the messages waiting on
it. The chat read "Orca ran into a problem, so this didn't go through. Try again." for that,
because every host stop was worded as Orca's own fault. It now reads "Codex never finished
starting, so Orca stopped it." in the chat's row and on each rejected message, carried as its own
failure kind so newer clients can tell it apart. The message counts as failed, like any start that
did not land.

* test(native-chat): pin the merged close and host-stop rows to their typed facts

The merge left two expectations on the old words: the close tests looked for the marker a close
used to write, and the host-stop test for the host-fault sentence. A close now writes "The chat
closed before this message was sent." with its fact, and a host stop the hostStopped sentence the
constructor gives, whatever reason the stop carried. Also folds the conversation command's two
imports from send preparation into one.

* fix(native-chat): a read of a chat this host cannot open says why

Reads now reach a chat through one accessor, which refused a missing record and a provider this
host does not run as a bare code with nothing beside it. Revealing the same chat already names
those reasons, so a client could tell "this chat no longer exists" and "update Orca" apart there
but not on the history or subscribe read that follows. The accessor now throws the same typed
refusals. The wire code and message are unchanged; the reason rides only in the error's data,
which released clients ignore.

* test(native-chat): a Claude retrying past the idle window keeps its conversation open

Every api_retry frame publishes the journal, and that publish is the activity the idle sweep reads, so a retry run revised into one row still renews the clock on each attempt.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-28 01:53:11 -07:00
OrcaWin 6fc3cdcad6 Bundle Bun for headless Orca and profile persistence (#22635)
Bundle a pinned, verified Bun runtime for headless Orca so existing Node launch commands can hand off before opening a profile. Keep desktop execution on Electron.

Add the Bun SQLite adapter and terminal backend, bounded shutdown, process inspection and cross-platform artifact qualification. Keep future managed SSH deployment separate from current production launch paths.
2026-09-25 22:49:06 -07:00
OrcaWin 82412dab8b Persist profile state in SQLite with background writes (#22612)
Migrate profile state to SQLite and move writes and backups into a background worker. Acknowledge terminal, SSH and automation changes only after durable saves. Preserve JSON import, recovery, rollback and compatibility exports.

Validate migration, worker failures, maintenance, cross-profile moves and terminal lifetime races with unit, integration and end-to-end coverage.
2026-09-25 22:47:33 -07:00
Jinwoo Hong 06a607a1d7 feat(orchestration): make multi-agent workflows durable (#16904)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->

| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 225 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$​21666 | $\color{#cf222e}{\Huge{\mathbf{−}}}$​2820 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$​18846 |
| Prod | 348 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$​17107 | $\color{#cf222e}{\Huge{\mathbf{−}}}$​4706 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$​12401 |

<!-- /orca-pr-loc -->

## ELI5

Orca now treats orchestration like a durable control plane instead of inferring success from terminal keystrokes. Agents can tell whether a prompt was accepted or a turn started, replay an ambiguous request without sending twice, and recover coordinator mail after a crash. Completed workers can be inspected, released, or retained, and their panes no longer auto-resume as if the work were still running.

## What changed

- **Run receipts** from `run-create/use/current/show/list` are the row without routing plumbing (`home_database`, `coordinator_pane_key`) and without the duplicate `binding` object.
- **`terminal send` receipts are honest and idempotent.** `input_accepted` and `turn_started` are the only stages; `--wait-submit` observes without resending; `--retry-request <uuid>` replays the exact request against the same process incarnation. A transport timeout keeps the retry ID; only a different runtime answering strips it. Value-less or non-UUID `--retry-request` is rejected on the CLI and the SSH shim.
- **Mailbox delivery is committed before wakeup.** Pointer writes are staged in the DB before any PTY byte, replayed once after restart, and never emit a naked Enter. The watermark that parks concurrent deliveries is released with the DB reservation. Restart rescans pointer-pending and `dispatch:` mailboxes.
- **Lifecycle is a guarded transition graph** (`lifecycle-transition.ts`) with a table-driven test over every caller edge. Task reopen/overturn stays in the public contract. A PTY exit during `worker-stop` is the stop succeeding, not a failure.
- **Worker lifecycle CLI:** `worker-start` (`--spec` creates Task + attempt in one call), `worker-show`, `worker-read` (provider transcript first, bounded terminal fallback with a typed reason, local/WSL/SSH), `worker-stop`, `worker-abandon`, `worker-release`, `worker-retain`, `worker-list` (rowid-fenced pagination, fleet liveness, `attention`, literal `nextAction`).
- **Release is an explicit ownership table** (`decideWorkerTerminalRelease`): only an `owned` resource can be settled, the archive is mandatory where reachable, and an owner whose process is proven exited can always get out of `retained` via `archive_status: unavailable`. User-taken-over, external, and transferred panes stay retained.
- **Settled-worker resume fence** (folds in #17651): a settled dispatch whose pane is still open is fenced at settlement, on stop/abandon/exit, and at startup; lifted on release, retain, takeover, and pane reuse.
- **Liveness is `live` / `unverifiable` / `exited` only**, from execution-host evidence. Fleet projection reads the evidence clock, not the relay delivery clock. A host-certified exit outranks the worker's settled state. `unverifiable` never authorizes stop, abandon, retry, or release, in code or in the guide.
- **Federation:** structured reads negotiate by `method_not_found` so every shipped host keeps transcript-first output; exited remote workers are closed before being reported closed; epoch fencing holds across peer restart, downgrade, and pairing rotation; no per-second forced capability probe.
- **Schema v35:** repairs databases stamped v34 by the pre-fix branch (mailbox_handle default, index predicates), drops the write-only `lifecycle_transition_receipts` ledger and five never-read v31 identity columns.
- **Schema v36:** `dispatch:<id>` mailboxes get a real consumer generation on `dispatch_contexts` and `remote_dispatch_attachments`, bumped and fenced in the same transaction on every re-attach (manual inject, worker-start, federated attach). A stale worker whose Dispatch moved to another process now gets `consumer_fenced` instead of silently acking the new worker's Delivery. Run mailboxes already worked this way.
- **Schema v37:** `dispatch_contexts` records its creator (`creator_handle`, `creator_pane_key`), so a coordinator's context-only self-dispatch is bookkeeping rather than a nesting parent; before this, one self-dispatch made every later `worker-start` from that coordinator fail the depth cap. Pre-v37 rows keep counting (fails closed).
- **Dispatch-mailbox ownership is checked, not inferred.** A `check` from a process whose pane no longer holds the Dispatch, or whose last Attempt was abandoned/failed and moved to another terminal, gets `consumer_fenced` instead of an empty inbox that reads as "no mail yet". `--peek`/`--all` stay readable. A paneless caller still gets `stable_pane_required` with the rebind recovery.
- **Liveness certification is stricter:** a `process_exited` stage whose termination reason is `unknown` (a stop that was issued but never observed) projects `unverifiable`, not `exited`. Federated `worker-show` carries the execution host's verdict and host kind instead of a local guess. A live, ready worker with nothing pending has `nextAction: none` rather than pointing at the `worker-show` that produced it.
- **Wire:** `workerShow` keeps `dispatch.task_id` next to `taskId` for shipped CLIs. `ask --json` uses the standard `{ok, result}` envelope like every sibling verb.
- **Migration start-version detection** treats the two v32 recovery columns as versioned. Before this, every shipped database stamped below 32 resolved to the v6 floor and replayed the whole chain (the v23 backfill synthesized 68 phantom retained workers on a real v30 profile). Verified on a copy of a real 62 MB v30 profile: starts at 30, no row delta, integrity ok, 11 ms.
- **Skill guide** rewritten as a ≤200-line kernel plus seven references, to the outcome-first standard (Result / Done / Safe failure first, conditions not case lists, one done bar, references loaded at the point of use). The canonical loop uses `worker-start --spec`, names `worker-list` for completion accounting, documents `--retry-request` / `request-show` / `--wait-submit`, and requires positive evidence before any stall action. The other seven guides get the same treatment in #18724, split out so this PR stays orchestration-only.
- **`rpc/methods/orchestration-*`** (126 flat files) regrouped into `orchestration/{worker,federation,messaging,runs,gates}/`.

## Why

User reports showed the same boundary failures: false `agent_prompt_stalled` causing duplicate sends (#15180), coordinators unable to trust screen scrapes, cold-parked terminals receiving a pointer without the submit, settled workers accumulating as live tabs and auto-resuming after restart, and no way to tell a stalled worker from a working one.

## Linked issues

Fixes #15180. Fixes #17935 (orchestration skill description is 866 characters; a guard now caps every bundled skill at 1,024). Supersedes #17651 (fence folded in). Advances #16660, #16522, #14907, #13047.

## Review record

This PR was reviewed adversarially after revival: eight independent lenses (lifecycle, mailbox, send, worker, federation, transcript, complexity, live ergonomics), each required to prove findings with a failing test. That produced 16 proven blockers, all fixed with red-then-green regression tests, followed by two re-review rounds and a third fix wave that caught 3 regressions introduced by the fixes and 7 fixes that missed their target; all closed. A final pass (five lenses incl. a live built-runtime smoke, then a re-review of the fix wave) found and fixed seven more, chiefly the stale-worker mailbox steal, the self-dispatch depth wedge, and the unproven-exit certification. Three independent Codex (gpt-6-astra) passes followed: the first found nothing new, the second found and fixed 3 defects (task-status reachability, WSL-local host classification, peer-capability epoch), the third found and fixed 6 (production PTY controller never installed settled writes, ambiguous in-flight pointer failures allowed duplicate replay, SSH/relay deadlines cut off a valid `--wait-submit`, stop-vs-exit race during inspection, and two release-recovery paths for vanished or exited terminals). The full record (findings, proof tests, triage, declines with reasons) is archived outside the repo.

**Rework after the live smoke.** A first live cross-host run on the shipped adhoc build (this Mac, a paired Windows host on the same build, a paired Mac on 1.4.195, and an SSH host) found a P1: a running local worker read `unverifiable`/`missing_status` because the fleet snapshot rows lacked the terminal handle the matcher keyed on. A 59-row failure table over every bug fixed during review showed the same two classes recurring: a fact dropped in transit through optional fields, and two authorities for one fact. Two blind designs (Opus, Codex) converged on the same mechanisms, and the scoped tranches landed here with red-then-green seam tests from the real producer to the real consumer, faults injected only at the transport or hook-ingest boundary:

- **Settlement (data-loss class):** one three-valued `WriteSettlement` (`accepted | refused{reason} | unverifiable{reason, bytesHandedToTransport}`) from the SSH multiplexer through daemon client, providers, controller, to pointer staging. No boolean, no rejection-as-third-state. The two silent degrades that fabricated a handoff are deleted; a provider that cannot settle refuses before any effect. Pointer text and Enter share the contract; a partial flush is `unverifiable`, never `refused`.
- **Evidence identity (false-liveness class):** fleet agent-status evidence is a tagged union (`binding: worker | pane | unresolved{reason}`, `clock: observed | delivery`) minted once at ingest, so a hook row captured on one process incarnation can never bind to a later dispatch on the same pane. The matcher's `!worker.paneKey ||` defaults are gone. One host-scope parser replaces two.
- **Small pre-merge items:** `capability_unsupported` from an old peer is no longer relabelled `host_unavailable`; a producer census test asserts every agent-status consumer path projects a pane-only hook row as `live`.

Two ergonomics defects the second live run surfaced on a real database are fixed here too: a pre-v3 dispatch already marked `completed` projected as `outcome_unknown` / `requiresAction: true` forever (three copies of the outcome ladder disagreed on legacy rows; now one resolver, legacy `completed` reads `succeeded` with nothing to act on, legacy `failed` stays actionable on the failure), and an unscoped `worker-list` enumerated the entire database (now defaults to the Run bound to the calling terminal, `--run` overrides, and the receipt's additive `scope` field says which).

A third live round on the shipped adhoc build of `b082443e1f` (same four hosts) plus an unscripted run in the user's own prompt style (a plain Claude Code shell, `/orchestration`, three workers, zero errors, bound-Run default confirmed) found two more branch defects, fixed with red-then-green tests: a worker freshly started on a paired server projected `unverifiable`/`host_indeterminate` with `requiresAction` for ~3 minutes, including after its own `worker_done`, because the host's federation observation returned `missing_liveness_verdict` for any PTY the liveness register had not yet swept (the host now reads a connected pane it owns locally as `live`; disconnected or SSH-scoped panes stay `unverifiable`); and six pre-v3 completed rows still carried an `input` category because settling through the task-status path or `failDispatch` never closed the Dispatch's pending question threads (both paths close them now, and schema v38 closes threads already pending on settled rows). The guide's `worker-start` examples now show `--model sonnet`, since an omitted model inherits the launcher's default.

A Codex adversarial pass on the tranche diff found one real design hole (identity minted at read time instead of ingest, now closed) and two daemon settlement paths that threw instead of settling (fixed). Two `@ts-nocheck` runtime mixins on these paths were extracted into checked modules; the repo-wide `@ts-nocheck` count is unchanged at 171.

Deletions during review: ~1,900 lines (write-only ledger, unread columns, dead v1 archive path, test harnesses shipped in prod, duplicated liveness and state-machine copies, self-capability checks that were compile-time true).

## Testing

- `pnpm typecheck:tsc:node|cli|web` clean
- `pnpm run check:code-quality:changed` 0 findings; `check:react-doctor:changed` 0
- `pnpm verify:bundled-skill-guides`, `verify:skill-bundle-manifest`
- full `pnpm test` on the integrated head: 72,332 pass / 292 skipped; the only failures were three non-PR files (two zsh live-shell suites hit a node-pty spawn-helper ENOENT while a concurrent native rebuild ran, 44/44 in isolation; `release-checkout.unit.test.ts` is a known 30 s load timeout that passes in isolation on `origin/main` too).
- CI on 70b4811267 (rerun, pre-Codex): the only reds are five SSH e2e specs plus `terminal-send-agent-prompt-submit:198`, each shown failing identically on main (main's E2E workflow is red on its last 40 runs). The terminal-send spec is root-caused and fixed separately in #18707. The Windows hook-service flake (#17721) and the federation load flake did not recur.
- Skills: `pnpm exec vitest run` over the skill gate files plus `src/cli`, `config/scripts`, `src/main/skills` pass; live smoke on the built CLI of `skills get orchestration` and `--full` (7 references).
- live headless runtime (`orca-dev serve`, isolated profile): canonical loop, stop, release, archive read, retry rejection, stale-handle check, SIGKILL-and-replay all verified with receipts
- Live cross-host smoke on the shipped adhoc build of `0d465e7931` (this Mac and a paired Windows host on the build, a paired Mac left on 1.4.195, an SSH host): local, paired-new, paired-old and SSH loops all settle; running workers read `live` on every host and `exited` after release; the old peer reads `capability_unsupported` and refuses release honestly. Injected 10 s relay stall with a send in flight: delivered exactly once after recovery, zero duplicates. Every liveness field across 104 receipts is only `live` / `unverifiable` / `exited`.
- Final live cross-host smoke on the shipped adhoc build of `b082443e1f` (same hosts): every loop settles; 942 of 948 legacy completed rows read settled with `requiresAction: false` before the question-thread fix and all of them after; `worker-list` scope reads `bound` / `flag` / `all` correctly; 122 JSON receipts carry only `live` / `unverifiable` / `exited`. Unscripted prompt-style run: clean.
- Confirmation smoke on the shipped adhoc build of `2da076d4e9` (this Mac and the paired Windows host, both updated): a freshly started Windows worker reads `live` on the first fleet poll and on all 20 that follow, with no `host_indeterminate` at any point, and `exited` after release; all 948 legacy completed rows read `requiresAction: false` with `nextAction: none` after schema v38; every verdict across 60 receipts is `live` / `unverifiable` / `exited`.
- Not physically exercised: WSL hosts, the renderer notification bell (headless has no renderer), same-session fence via a real pane close (renderer-only state), restart mid-delivery on a real app (covered by e2e only).

## Notes

- Remote-wire additions are optional fields or `method_not_found`-negotiated methods; one new Electron-only IPC channel (`agentStatus:legacyWorkerTerminalResumeFence`) never crosses the wire.
- SSH contact loss remains `unverifiable`; the execution host stays authoritative.
- Intentional wire projection change: an SSH host scope with an empty `targetId` now projects host id `ssh` instead of an empty string (remote-wire-compatibility rule 3, old clients decode the same field). A fleet pane key without a terminal handle is now `unidentifiable` rather than matched by pane key alone.
- Found live but pre-existing on main, filed separately: a relay daemon-start collision during transport loss rewrites the endpoint credential and wedges the surviving relay (host needs a manual kill); `terminal create` on a reconnecting SSH host reports an opaque `No PTY provider for connection`; `terminal list` reports `orphaned:false` and `terminal close` reports `ptyKilled:true` for a pane whose relay is gone (orchestration's own projection reads `unverifiable` correctly at the same moment).
- Downgrade after this PR is not a supported path: main opens a v37 database and early-returns (its inserts still work against the v36/v37 defaulted columns), but its one-outstanding-Delivery-per-Run index is a no-op against the branch's mailbox-scoped index of the same name.
- Known follow-ups (not blockers): `worker-list` materializes every dispatch row per call; a positive "agent absent" signal distinct from PTY liveness is a product decision left open (a headless fake agent never reaches `live`, so its `nextAction` stays `inspect`); a context-only self-dispatch still lists as `role: worker` in `worker-list`; `dispatch` task-not-found / task-not-ready / inject-rejected still surface as `runtime_error`; task and inbox receipts still expose raw row columns. Deferred skill product decisions live on #18724.
2026-09-06 14:34:03 -04:00
Neil 1924c8f5b1 feat(perf): lint repeated sort setup and schedule regression contracts (#18822)
* feat(perf): audit comparator setup and schedule performance contracts

* test(sqlite): close readers after expected busy failures

* ci(perf): trigger contract workflow on the contract files themselves

Without these paths a contract rename lands green on PR CI and only
breaks the next nightly, where nobody owns the failure. Also run the
OS-independent source audit once instead of on all three runners.
2026-09-05 13:56:06 -07:00
Brennan BensonandMerge Sim f4c2821167 refactor(agent-session-journal): move the session journal onto SQLite (#18652)
* refactor(agent-session-journal): move the session journal onto SQLite

The agent-session journal kept its state in three hand-rolled file formats: an
append-only `log.jsonl` with torn-tail repair, a `snapshot.json` holding folded
state plus a retained tail, and byte-quarantine files for anything unreadable.
This replaces all of it with one SQLite database per session — `journal.db`
beside the existing `blobs/` store — using the in-house adapter and the
open/pragma/migrate/harden pattern the orchestration database already follows.

Two tables: `journal_rows` (the append-only log, keyed by
`(session_id, epoch, seq)`) and `journal_sessions` (the derived projection,
upserted in the SAME transaction as every insert). Rows stay JSON in one
column, so the row schema, the version upcast chain, and the reducer survive
byte for byte — `journal-reducer.test.ts` and four other suites pass unchanged
and are the regression proof.

Deleted: `journal-log-file.ts`, `journal-compaction.ts`,
`journal-corruption-quarantine.ts`, and the public `compact()` /
`compactionBoundary` / `autoCompact` members, none of which had a non-test
caller.

Existing `log.jsonl` / `snapshot.json` journals are deliberately abandoned. No
importer: a session created on the old path stops working, which is acceptable
because the feature is off by default.

## The physical quota is repriced, because SQLite does not charge like a file

The 256 MiB per-session bound is unchanged, but the arithmetic under it could
not survive: SQLite grows the database in pages and the WAL in frames, and the
checkpoint that copies the WAL forward holds the same pages in both files at
once, so a transaction's peak is about twice its content. Admission now charges
the candidate transaction's own measured page cost, validated against a sweep
that runs as a regression test (`journal-database-space.test.ts`) rather than
derived from reasoning about the allocator.

Four things are load-bearing rather than tuning, each measured:

- `auto_vacuum = INCREMENTAL` must be set BEFORE `journal_mode = WAL`. Set it
  after and it is ignored with no error, reclamation silently becomes a no-op,
  and the file never shrinks again. Both halves are asserted.
- `wal_autocheckpoint = 0` plus an explicit `wal_checkpoint(TRUNCATE)` at the
  end of every write path, so the one moment the same pages live in two files is
  a moment the charge accounts for.
- Reclamation runs in bounded chunks. A single unbounded `incremental_vacuum`
  took a 252 MB directory to 504 MB — the reclamation added to defend the bound
  would have breached it. `PRAGMA incremental_vacuum(N)` also frees exactly one
  page unless it is stepped to completion, which no size assertion catches, so
  the freed page count is asserted directly.
- A blocked checkpoint leaves the WAL on disk together with the database growth
  it already copied, so admission charges that deferred copy explicitly. The
  term is zero whenever the last checkpoint succeeded, so the uncontended path
  admits and refuses an identical set.

The epoch discard is `DELETE FROM journal_rows` with no WHERE clause, which
takes SQLite's truncate optimization: measured at ~0.26% of the database in WAL
bytes where the `WHERE session_id = ?` form rewrote every emptied leaf at up to
99%. One database per session is what makes the unqualified form correct.

An open, empty journal costs 57,344 bytes before a single row exists, so a
configured quota below `JOURNAL_MIN_SESSION_BYTES` now fails loudly at open with
the existing `journal_bound_exceeded` instead of as a run of identical append
failures. No production caller configures one; the affected surface is test
fixtures, rescaled to the smallest value that restores what each case proves.

## One deliberate behaviour change

Compaction was the only mechanism that shed bytes inside an epoch, and the write
path called it precisely so an append at the bound was not refused. The
SQLite-shaped replacement — a bounded prefix delete — cannot be used: with the
snapshot gone the surviving rows ARE the state, so dropping the oldest of them
loses the oldest transcript silently at the next reopen. So no row is ever shed
inside an epoch, and a session whose row bytes alone reach the bound now refuses
every append where it previously compacted and continued. A loud typed refusal
beats silent data loss.

What still sheds is unreferenced BLOB bytes — the dominant and unbounded byte
source — on the same write-path hook. The escape from the hard stop is the fold
that already exists, `replaceEpochItems`, which now actually returns bytes to
the filesystem instead of leaving them on the freelist.

The prune's protected set is a union of live reducer digests AND the candidate
row's own digests, including those cited only by a nested lifecycle-batch
mutation. Content addressing never rewrites a digest already on disk, so
protecting live state alone deletes the blob the append is about to cite — a
dangling reference that surfaces one reopen later as an empty expansion on an
item the user can see. `journal-store-blob-budget.test.ts` pins it, and it goes
red when the set is narrowed back.

## Handle ownership

A file handle used to be opened and closed per append; a SQLite handle is held
for the session's lifetime. Every path that can open a connection now has one
owner: the open function owns its raw connection until it returns, the store
owns its retained one and releases it in a new `close()`, and every other
connection is closed by the call that opened it. The attach, recovery,
eviction, map-overwrite and host-teardown paths close what they drop, and host
teardown is failure-complete — the sink-barrier flush throws by design, so a
trailing close statement would be skipped on exactly the path that leaks.

`close()` has a stated contract: admission at enqueue and permanent, the close
step on the same queue past that gate, one shared in-flight attempt, fulfilment
terminal, and the release last and deliberately unguarded so a retry re-enters
it. Guarding the release would skip it on retry, guaranteeing a permanent leak
in exactly the case where it did not release.

`journal_closed` joins the error union for a write after `close()`; no file
outside the directory references any of these codes.

* fix(agent-session-journal): make a COMMIT final, stop repairs deleting valid rows, and keep rejected closes retryable

Six review findings on the SQLite journal migration.

1. A successful COMMIT is now the point of no return. The ordinary append,
   the epoch roll and the epoch replacement each adopt the committed row or
   epoch BEFORE any post-commit filesystem work; checkpoint, reclaim, blob
   prune and directory measurement run through `runJournalPostCommit`, which
   is best-effort by design and falls back to the transaction's own charge as
   a conservative footprint. Previously a post-COMMIT scan failure rejected a
   durable append and the next one reused its sequence, and a failed epoch
   housekeeping step left the store writing into a prefix already deleted.

2. Corruption repair preserves instead of destroying. A rejected suffix is
   copied into a new `journal_quarantine` table and removed from the live
   epoch in ONE transaction per chunk, charged against the session bound
   before a byte is written; a journal that cannot afford the copy refuses to
   open rather than falling back to deletion. The repair state is exposed as
   `journal.repair` and the rows are readable through
   `recoverQuarantinedRows()`, so Orca-owned submission, receipt and
   lifecycle identity survives a gap or a malformed row.

3. The physical charge covers the B-tree key payload. `session_id` and
   `epoch` are stored in both tables and both primary-key indexes and appear
   nowhere in `row_json`, so the journal boundary now bounds them and
   `journalTxnPhysicalCost` charges those bounds plus the projection upsert.
   The charge sweep runs the exact production transaction at maximum admitted
   key sizes.

4. A rejected `close()` no longer orphans its handle. Callers hand the
   journal to `agentSessionJournalCloseRetries` instead of swallowing the
   rejection, the attach map replacement is ABORTED when the previous
   journal will not close, host teardown retries what the registry holds, and
   a failed runtime teardown is retained so the next stop is a real retry.

5. `journalWalBytes()` returns zero only for ENOENT and propagates every
   other stat error, so admission and reclamation fail closed.

6. The WAL contention test closes the writer before removing its temp root
   and asserts the directory is removable once handles close.

Regression coverage: post-commit divergence (4), corruption repair (5),
key bounds (5), WAL stat (8), close retry (5), plus a runtime stop-retry
case. Each fix was ablated on this head and the matching tests go red.

* fix(agent-session-journal): anchor replay at sequence 1, make quarantine append-only, and charge it in bytes

Three ways the corruption quarantine still lost rows it was written to keep.

Replay validated contiguity from the first row that HAPPENED to remain, so an
epoch missing only its sequence-1 row declared the leftovers contiguous and set
no `truncateFrom`. The load was still corrupt, so recovery imported provider
history and `replaceEpochItems` deleted every live row — including Orca-minted
submission, receipt and lifecycle identity that no transcript can reconstruct,
and that nothing had quarantined. Replay now anchors at sequence 1, so a missing
epoch row rejects the whole surviving range before any replacement runs.

`journal_quarantine` was keyed on `(session_id, epoch, seq)` and copied with
`INSERT OR REPLACE`. A repair frees the sequences it removed and the live epoch
reuses them, so a second repair in the same epoch silently deleted what the
first preserved. The table is now keyed on a surrogate `quarantine_id`, the copy
is a plain append, and `(epoch, seq)` is metadata; existing v1 databases are
rekeyed in the migration that already bumps `user_version`.

The admission charge read `length(row_json)`, which counts CHARACTERS for a TEXT
value where `journalTxnPhysicalCost` expects physical UTF-8 bytes. A multibyte
suffix was charged at up to a third of what it writes, which defeats the
pre-write physical bound — over a megabyte on a maximum-size lifecycle batch.

* fix(agent-session-journal): keep a repaired epoch anchored and stop the v1 quarantine migration doubling the file

Replay validated numeric contiguity from sequence 1 but never that sequence 1
IS the epoch row. When the anchor was missing the repair set aside every
surviving row, and if provider-history import then failed — a transcript that
is temporarily gone is enough — the journal reopened as a clean, row-less
epoch: an ordinary append took sequence 1, replay accepted it, read-restore
published it as history, and automatic recovery never ran again while the
user's real messages sat in quarantine.

Replay now rejects an unanchored prefix, the open publishes an
`unreconcilable_prefix` anchor for an epoch its repair emptied, and that anchor
keeps reporting corrupt — so provider history is retried on every attach —
until the timeline is rebuilt or the session writes content of its own. A
repair also discloses rows it set aside when no line was unreadable at all,
which is the case that removes the most.

The v1 quarantine rekey copied every legacy row into the new table inside one
transaction and dropped the old one. A quarantine holds whole rejected rows: a
single 8 MiB row nearly doubled the database past the physical bound the open
had already checked, the dropped pages only reached the freelist, and the next
open refused the session it had just migrated. The v1 table is renamed and
frozen instead, and reads take both generations. Table creation also moves
inside the migration transaction, so a crash can no longer leave a v2-shaped
database still reporting version 0 for an older build to write into.

* fix(agent-session-journal): stop an empty provider transcript retiring the repair marker

A transcript that exists but decodes to zero messages was imported as a
success: the import published an empty `legacy_import` replacement that
deleted the `unreconcilable_prefix` anchor and its disclosure, so the next
probe read the session as clean and every later attach skipped provider
recovery while the user's rows sat in quarantine for good.

The import now leaves the epoch untouched when nothing decodes, reporting
`replaced: false`, and recovery treats that like a transcript it could not
read — the marker stands and a later attach with real history rebuilds the
timeline.

* style(agent-session-journal): merge the duplicate journal-database-space import

* refactor(agent-session-journal): drop quarantine, byte bound, blob spill and rate limit

Match what comparable implementations do: the journal is an unbounded
append-only SQLite log with no side tables and no admission control.

Corruption: the rejected suffix is DELETED rather than copied into a
quarantine table. The load still reports `corrupt` and recovery still
rebuilds the epoch from provider history, so the observable outcome is
unchanged — only the preservation half is gone. The schema is back to one
version with two tables; no v1 database exists outside unmerged commits of
this branch, so the rekey migration and the two-generation read path go with
it. Sequence-1 epoch anchoring and the empty-provider-transcript retry are
kept: both are about the corrupt signal being correct.

Size: no `maxSessionBytes`, so no page-cost arithmetic, reclaim band,
incremental vacuum, lifecycle byte reservations or `journal_bound_exceeded`.
`auto_vacuum` and `wal_autocheckpoint = 0` existed only to make a
transaction's physical cost predictable for that charge; with the charge gone
SQLite's default checkpointing is what the journal wants, and the explicit
pre-close checkpoint is redundant with the one `db.close()` performs. WAL,
`synchronous = FULL` and `busy_timeout` stay.

Payloads: an oversized body is truncated at the existing inline cap with the
existing marker and the remainder is discarded, bounded at the translation
layer that already calls these helpers. The truncation point and message do
not change; the content-addressed blob directory and all digest tracking do.

Rate: no `maxAppendsPerWindow` and no `journal_rate_exceeded`.

`JournalPayloadLimits` is now just the inline cap.

* fix(agent-session-journal): mark a partial repair pending and bound multi-block tool input

A repair that keeps its prefix had nothing durable to show for the suffix it
deleted: a sequence gap costs no malformed row, so no disclosure is appended,
and the surviving rows keep their epoch anchor. The next probe read a
contiguous anchored prefix, called it clean, and the deleted stretch of
timeline was never asked for again — silent loss, with the deletion already
committed. The deletion now writes a `journal_repairs` marker in the SAME
transaction, and replay keeps reporting corrupt while it stands. It retires
under exactly the rule the emptied-epoch anchor takes: a fresh epoch carries
the rebuild, or the session writes content of its own past the sequence the
repair left free. The repair's own disclosure is not that content.

Legacy import bounded a tool call's input only when it was the message's sole
block; the multi-block path returned `tool-call` unchanged, so a mixed message
from Claude, Grok or an omp execution cell persisted the whole input despite
`inlineHeadBytes`. `boundBlock` now routes it through `boundToolInput`.

Also drops canonical comments describing quarantine, snapshot files, blob
storage and blob compaction — none of which exist any more.

* fix(agent-session-wire): stop awaiting the synchronous journal probe

loadJournal runs on a sync-database connection and returns JournalLoad | null, so both wire call sites were awaiting a non-Promise. The type-aware code-quality gate flags it; the native gate does not.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-04 15:09:23 -07:00
Neil 9062494f9b fix(ai-vault): stop a whole opencode.db failure reading as one skipped transcript (#16587)
* fix(ai-vault): stop a whole opencode.db failure reading as one skipped transcript

#15036 reported "1 transcript skipped / database is locked" with both Agent
Session History scopes empty. Two separate defects.

The panel counts every unkinded scan issue as a skipped transcript, so a
failure that lost an entire *source* was reported as one lost *file*. The
whole-database failure is now kinded `scope`, and an unknown `kind` from a
newer host degrades to `scope` instead of failing validation and coming back
unkinded — a mixed-version remote host previously turned a source-level
failure into a phantom skipped transcript.

The read also inherited sqlite3's 0 ms busy timeout, so a genuinely contended
open failed in ~1 ms. It now opens once with a bounded timeout. No retry loop:
sqlite's own busy handler already blocks and retries internally for the whole
timeout, and WAL readers do not block on a writer at all (measured: 547/547
cross-process reads at timeout=0 while a writer held open transactions).

Measured against a real Ubuntu-24.04 distro, Windows cannot take SQLite's file
locks over \\wsl.localhost at all: an idle, never-WAL, nothing-attached
database still answers SQLITE_BUSY, a 5 s busy timeout does not change it, and
the identical bytes open fine once copied to local disk. So a lock-family error
on that share never means "a writer holds it" and no timeout can help. The copy
says so rather than sending the user after a write-ahead log that is not the
problem. Restoring those sessions needs an in-distro read; that is a follow-up,
and this PR no longer pretends a timeout will do it.

immutable=1 is deliberately not used as a workaround: over the same share it
opens and returns 100 of 150 rows, silently dropping everything still in the
uncheckpointed -wal — in a history panel, exactly the newest sessions.

* skip the provably futile busy wait on \\wsl.localhost paths
2026-08-26 23:37:16 -07:00
NeilandOrca 0fe04e2c91 perf(sqlite): cache prepared statements in SyncDatabase (#13769)
prepare() recompiled every statement, so orchestration reads re-parsed the same SQL
on the main thread. Adds a bounded LRU keyed by SQL, cleared on close() and before
schema-changing exec().

Wildcard selects are excluded: node:sqlite builds the first post-schema-change row
from stale column names, so a reused SELECT * can silently drop a freshly added
column. PRAGMAs stay uncached.

Co-authored-by: Orca <help@stably.ai>
2026-08-11 01:48:14 -07:00
Jinwoo HongandJinwoo-H cf16eac7f6 fix(agent-hooks): keep Node 18 relay companion loadable (#13135)
Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
2026-08-07 22:54:42 -07:00
Neil 1009ac9083 chore: update Electron to 42 (#3919) 2026-05-30 13:26:48 -07:00