mirror of
https://github.com/stablyai/orca.git
synced 2026-10-03 16:02:11 +00:00
* fix(native-chat): report a failed startup chat reconcile instead of failing app startup At startup the chat host re-checks every saved chat's lease and writes the result to agent-sessions.json. If that write failed (the file lock gave up, the file could not be written, or the file was written by a newer Orca and is read-only here), reconcileRestartLeases rejected, the startup IPC call rejected, and the renderer fell into its degraded "Session restore failed. Changes won't be saved until restart" mode. The reconcile is bookkeeping: a lease left unreconciled grants no writer, and every attach, send and read of a chat reconciles its own lease again. So the startup reconcile now reports its failure through a new optional host dependency, onStartupReconcileFailure, and resolves. The runtime routes it to its onError sink under the scope structured-agent-session-startup-reconcile, or logs it when no sink is installed (the desktop installs none). * fix(native-chat): read restored chats without waiting on lease bookkeeping With native chat on and a chat tab open at quit, the renderer's startup also awaits the chat tab restore (session.tabs.listAll). That restore re-ran the lease reconcile before reading each chat and rethrew its store failure, then recorded each restored tab as visible through a store transaction that throws on a held lock or a read-only store. Either one failed the restore, so startup still fell into "Session restore failed". Reading a chat grants no writer, so the reconcile startup and the restore run is now a reader's: createReaderReconcile never throws, answers whether every lease is settled (recovery is resolved only then; the journal opens either way), and reports each distinct failure once until a reconcile settles. Attach and agent start keep the strict reconcile. The restore's tab republish logs a failed visibility write and still publishes the tab, since a client drops every unpublished chat tab; user-driven publishes still refuse. The host dependency is renamed onLeaseReconcileFailure (scope structured-agent-session-lease-reconcile), since it now also reports for reads. * fix(native-chat): keep every record-store write off the startup chat read path Round-2 review found two more writes on the startup chat restore that could still fail it and put the app into "Session restore failed": republishing a /clear replacement recorded its tab visibility strictly, and resolving a chat's recovery rethrew its store error. The restore also paid one lock wait per tab and per batch of chats while the lock stayed held. The restore now derives tabs from state it already holds: - publishStructuredAgentSessionTab splits into the strict write and projectStructuredAgentSessionTab, which only updates the runtime's snapshot. The restore and /clear replacements only project: a saved tab index already lists every restored chat, and a /clear moves the tab in the same write that commits it. visibilityWriteMayFail is gone. - Chats a legacy profile restores that the index does not list are recorded in one best-effort transaction (store.showSessionTabs), so a failure leaves the index absent to seed again rather than partial. - The read restore's recovery resolution is caught and reported through onLeaseReconcileFailure, deduplicated with the reconcile's reports. - Once lease bookkeeping fails in a restore pass, the rest of that pass skips it, so a held lock costs one wait for the startup reconcile and one for the restore, however many chats are open. User actions (create, reveal, attach, send, the /clear commit) keep their strict writes. * test: open, seed and read the agent-session record store through one harness Tests that open the durable agent-session record store, seed it, or read back what it persisted now go through agent-session-record-store-test-harness.ts instead of calling AgentSessionRecordStore.open or touching agent-sessions.json themselves. A later change that moves the store into the chat database then changes the harness instead of every test. No production code changes. Tests whose subject is the JSON file itself (its .bak recovery, salvage, schema versions, permissions, and what older builds read back) keep reading and writing the file directly; the storage move rewrites or deletes them. * fix(native-chat): start each restore pass from one lease check and stop its bookkeeping at the first failure The restore now runs one reader lease check for the pass and lets each chat re-check and resolve recovery only while the pass is still settled. The first refusal or failed write clears it for the rest of the pass, and every chat is still opened for reading. With another process holding the lock, startup waits on it once in prepare and once in the restore, however many chats are open; a legacy profile waits once more for its tab-index seed. * docs(native-chat): correct restore comments and a test name to match the final design * test: address the record-store harness by the host's state directory The harness took the store's own folder, so each caller picked one (join(root, 'store'), or 'agent-sessions' where a test read the store the runtime owns). A later change that moves the store into the state directory's journal database could not tell those apart, and would have had to edit every caller again. Every harness function now takes the state directory, the one the test's journal database and recovery capsule already live in, and keeps the store in the same subfolder the runtime uses. Callers pass that directory; store-only tests pass their temp directory unchanged. Format tests that share a directory with harness calls take the file path from testAgentSessionStoreFilePath. The folder name moves from a private constant in the runtime to AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness shares it without importing the runtime. Its value and every path built from it are unchanged. * refactor(native-chat): keep agent-session records in the chat journal database The record store's records, operation ledger, retired claim keys and chat tab index become tables in agent-session-journal.db (user_version 4). The version-4 migration copies agent-sessions.json in its own transaction and never writes, renames or deletes that file or its .bak. Each store write is one journal transaction over exactly the rows it changed, checked with the load rules; the file lock, the external-change refresh and its hash, the .bak rotation, salvage and the hot-path recovery fence are gone from the store. * wip: importer tests * test(native-chat): cover the records migration, the import, row writes and read-only records * docs(native-chat): retire comments that describe the records file as the live store * test(native-chat): drop the record-store harness's leftover file path and type the import fixture * test(native-chat): let the host harness cleanup wait out a recovery-offer read's lock * fix(native-chat): let Stop reach the agent when its ledger row cannot be written Stop's operation-ledger row now shares the database with the chat history, so damage, a full disk or a stranded transaction on that write refused the Stop before the interrupt. A cancel plan now takes its decision from the committed ledger in memory, runs without settling, and warns that the row was skipped. Other mutations answer proven damage with the typed "Unable to load this chat." refusal instead of the raw SQLite error. * fix(native-chat): answer whether a profile holds chats from the database's rows Every host install creates agent-session-journal.db, chats or not, and the version probe created it too, so its mere existence made every profile that ever installed the host wait on host install and reconcile at startup. The check now opens the database read-only and looks for a record or tab row, lets the records file answer while its import is still owed, and counts an unreadable database as present. The version probe no longer creates the file. * fix(native-chat): open a chat from history when its tab index cannot be written Over records a newer Orca wrote, every write is refused, so opening a closed chat from Agent Session History failed on the tab-visibility write and the chat read as unreachable. Like closing a tab, opening one now reports a failed restore-index write and still publishes the tab. * fix(native-chat): keep the records import owed when the backup read fails transiently A torn records file whose .bak could not be read (EACCES, EIO) was reported as unusable, so the migration completed with nothing copied and never retried. A non-ENOENT read failure of either copy now carries its cause, which the importer classifies as a read that can clear. * test(native-chat): pin that an unreadable records file never falls back to its backup * fix(native-chat): restore imported chats' tabs when the records file had no tab index A chat created while the import was owed recorded a tab index holding only itself. When the file it later imported had no index, that index still read as recorded, so the imported chats' tabs never came back. The import now clears the recorded marker in that case, and restore falls back to the profile's tabs. * refactor(native-chat): drop the unused in-transaction store write Nothing called it, and it bypassed the write queue and the read-only refusal. * docs(native-chat): say that an unusable records file is left untouched but never re-imported * refactor(native-chat): keep the provider handle chain check as main has it The chain-validation refactor has no measured need in this change. * docs(native-chat): retire lease-renewer comments that describe the records file as the live store * fix(native-chat): keep a throwing failure sink from failing the startup chat read The lease bookkeeping failure reporter called the host's failure sink directly, so a sink that threw turned a reported, recoverable store failure back into a rejected startup reconcile or read restore. The reporter now catches a sink throw and logs both the original failure and the sink error with console.warn. * test(native-chat): wait for a replaced host's restart-offer writes before cleanup A restart test replaces the host without tearing the old one down, so the old host's fire-and-forget restart-offer withdrawal could still hold the recovery capsule's lock directory when cleanup removed the test directory (ENOTEMPTY). The harness now hands hosts a capsule that tracks running operations and waits for them before removing the directory, replacing the rm retries. * docs(native-chat): retire the abandon helper's note that the store re-creates its directory * fix(native-chat): restore a chat opened while the import was owed beside the profile's chats When the imported records file had no tab index, restore fell back to the profile's saved tabs, which never list a Claude chat, and the seed then rewrote the tab table without the chat opened while the import was owed. The tab rows that chat left are now loaded as unrecorded, restore takes them together with the profile's chats, and the seed keeps their tab ids. * test(native-chat): pin that a create whose tab index write fails still opens the chat * docs(native-chat): say why restore puts chats opened while the import was owed first * test(native-chat): replace a ledger row rather than change it in place in the Send-now rerun test The record store freezes published rows in tests, so setting a row's outcome in place threw; the test now swaps in a changed copy, as its sibling cases do.