mirror of
https://github.com/stablyai/orca.git
synced 2026-09-29 16:02:50 +00:00
orchestration-dispatch-error-codes
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
720c3299ba |
fix(ssh): require a host death certificate before recreating a pane, and unstick expired leases (#18013)
* fix(ssh): match an expired lease on where its leaf lives now, not its frozen tab A lease freezes tabId at write time, but detachTerminalPaneToTab moves a live pane, so the stored tab is the one the pane LEFT. getRecentExpiredSshLease required lease.tabId === tabId, which is wrong in both directions: a viewer on a stale mirror matched under the abandoned coordinates (and resolvePersistedStable PaneOwner then reads an empty layout for that tab, so adoptStablePane is skipped entirely and a fresh shell is spawned over a possibly-live one, binding the same leaf in two tabs), while a viewer using the pane's real coordinates matched nothing and got terminal_not_recoverable. Resolve the leaf's current tab the way restoreReattachedPtyRuntime already does and compare against that, falling back to the frozen tabId only when nothing can say where the leaf lives. Both workspace partitions are read because SSH spawns bind into ssh:<target> while reattach binds into local. * fix(ssh): let a proven reattach take an expired lease back to attached #17965 authorized reattach from `expired` but the state machine refused the transition back, so a lease that reattached and proved itself alive stayed `expired` forever. That silently exempted a demonstrably running remote shell from `ssh:reset` (skips `expired`), from the SSH_TERMINATE_RECONNECT_REQUIRED ownership fence in `ssh:terminateSessions` (marks it not-owned), and from the quit-time `detached` sweep, and made it permanently ineligible to win supersession so its own successors never retired their predecessors. Only the id-qualified caller carries per-pty proof: markSshRemotePtyLeases AttachedAsync is fed the relay's `attachedLeaseIds`, so an unqualified bulk mark over a whole target still cannot revive `expired`. `terminated` stays absorbing. Re-entering `attached` drops supersededBy/relayIdRecycled, since route retirement belongs to the shell that lost the pane and this one just proved it is not that shell — the same invariant upsertSshRemotePtyLease enforces. * fix(ssh): make the pane-recovery liveness gate refuse without positive evidence of life The gate refused only `live` and `unverifiable` and passed on `null` — but the register is an in-memory Map, so `null` is equally what a fresh app start, a never-asked host and a certified death look like. Absence of evidence was reading as authorization to spawn a shell over a possibly-live remote process: `!pty.connected` is cleared for every PTY a dropped relay owned, and `expired` only ever says the CLIENT lost its route. - `exited` is now RETAINED rather than deleted, so the register is three-valued in the map as well as in the type. Its one writer is a host-delivered exit frame — an exit with a real code, or an explicit `hostExitConfirmed` — which records the certificate instead of merely dropping the doubt. - `recoverTerminalPane` refuses on `live` and `unverifiable`, and deliberately does NOT demand a positive `exited`. The only answer that ever reaches this gate is a reachable relay reporting no such id, and that is a union: pty.attach throws not-found for an unknown id with no liveness check, and a relay restart makes every previously minted id unknown (ids carry a per-start `ptyIdMintEpoch`). No writer of `exited` co-occurs with a reattachable `expired` lease either — a host-delivered exit frame tombstones the lease `terminated` — so requiring one would close the gate permanently. - `handlePtyReattachFailure`'s not-found branch publishes `code: -1` to the renderer and does not call `runtime.onPtyExit`. The relay's not-found answer is not a death certificate, and #17963's ratchet on the same branch pins that. - The inventory's `observed === false` hunk keeps dropping doubt rather than asserting a death: `pty.listProcesses` returns the relay's CURRENT session map, so a restarted relay omits every previously minted id whether or not those shells died — the same union, one hop away. A live or unprovable pane refuses; a disowned one still recovers. No wire change. The gate's ratchets live in terminal-pane-recovery-liveness-gate.test.ts: config/vitest.config.ts — the config CI runs — matches only `*.test.ts`, so cases placed under orca-runtime-tests/*.spec.ts would never execute. * fix(ssh): gate paired-viewer pane recovery on the narrowed session-gone predicate isSshSessionGoneError landed on the IPC transport, which never calls terminal.recoverPane. The one caller that does — recoverExpiredHostPane in the paired-viewer transport — still triggered on a bare SSH_SESSION_EXPIRED substring, so the identity-mismatch reply (the relay found a LIVE PTY under that id owned by another pane, which is evidence of presence) still asked the HUB to replace the pane, putting a second agent on one transcript. Main already refuses the respawn on that same reply; this makes the two agree. A pane whose shell genuinely died is unaffected: plain SSH_SESSION_EXPIRED still matches. The mismatch reply now surfaces as an error instead of a respawn. * test(persistence): update the reattach ratchet for expired-lease reclaim markSshRemotePtyLeasesAttachedAsync is id-qualified, so a named pty that proved itself alive now returns to attached instead of staying expired. |
||
|
|
03fcfdfb92 |
feat(orcad): boot the Orca runtime on plain Node (#15968)
* refactor(host): resolve the app root through the port in fork-reachable modules
`parcel-watcher-entry-path.ts` and `session-scanner-service-entry-path.ts` read the
app root via `require('electron').app` inside a try/catch that already returns null
when Electron is absent. They were therefore correct under plain Node at runtime and
only failed the *static* text check — which is real, not pedantic: the comment in
`ports/port-scan-command-client.ts:19` records that the plain-node-entry-guard fails
on that literal text, try/catch or not.
`hasAppEnvironment() ? getAppEnvironment() : null` gives the identical "no app root
here" answer without the text. That restores `hasAppEnvironment`, which an earlier
commit in this stack deleted as unused — it now has the caller it was waiting for.
Ratchet baseline 27 → 25.
Verified: 74 files / 458 tests; `pnpm typecheck` clean; `oxlint` clean.
* feat(orcad): boot the Orca runtime on plain Node
Closes the last two Electron couplings and makes `orcad` a working artifact:
a 4.43 MB Node bundle that boots, pairs, registers a repo, creates a real git
worktree and round-trips a PTY — with zero `require("electron")`.
Ratchet 2 -> 0, so `config/runtime-electron-baseline.txt` is now empty and its
test asserts exactly that: any reachable electron import is a regression.
- speech: inject the service factories, so importing ModelManager for its type
no longer drags Electron's streaming net.request into the graph
- filesystem-watcher: add a WorktreeWatcherRemoval port. Every entry in those
maps arrives through an ipcMain handler carrying a renderer sender, so a host
with no renderer has nothing to close, restore or forget — the inert default
is what the desktop code does against empty maps, not a stub hiding work
- user-data-path / profile-storage-paths: resolve userData through
AppEnvironment. These surfaced only once orcad pulled the store in
Both host ports now anchor to a realm-global symbol. `vi.resetModules()` gives
the re-imported graph a fresh module copy, so a binding installed before the
reset silently read back as uninstalled.
The acceptance smoke drives both hosts through one code path (`--target
orcad|electron`) and seeds its own git repo, so it is hermetic and asserts the
same contract of each. Wired into PR CI.
* test(smoke): remove the seeded workspace container, not just the worktree
* test(smoke): surface the server's stderr when it dies before ready
* fix(smoke): build node-pty for Node before booting orcad in CI
* fix(smoke): drive the CLI built from this checkout, not one on PATH
* docs(ratchet): say the baseline must stay empty, not merely shrink
* build(orcad): externalize only the native modules actually in the graph
|
||
|
|
ed4d6979b1 |
fix(app): await durable checkpoints before restart actions (#12433)
* fix(app): await durable checkpoints before restart actions * fix(app): clear restart latch after refused reload * fix(persistence): invalidate hash after stale rename |
||
|
|
25fefa4072 |
fix(P1-A): async SSH consumer-recovery persistence and detach on failed connect (#12026)
* fix(P1-A): persist SSH consumer recovery without a sync store flush rememberPtyConsumerRecovery ran on the live establish/reconnect path and called flushOrThrow -> writeToDiskSync, parking the Electron main thread on the profile-directory write. On a stalled or slow profile mount that freezes the whole app during SSH recovery and reconnect. Add Store.flushAsync(): same debounce-cancel and write serialization as flushOrThrow, but awaits writeToDiskAsync instead of blocking. The consumer recovery upsert/remove pair is now async and awaits it, and the SSH callers await through to establish()/reconnect() so ownership is still durable before relay setup continues. In-memory state still mutates synchronously (before the first await), so no caller can observe a torn record and dispose() stays synchronous. * fix(P1-A): detach the SSH session when a connect attempt fails Both failure exits in doConnect dropped the session from activeSessions without calling detach(). claimSshPtyConsumerRecovery only reuses an existing in-memory entry when detached === true, so the next connect attempt fell through to minting a fresh clientInstanceId, discarding the remembered owner lease and its resume identity. Route both exits through abandonFailedSshSession(), which detaches (keeping PTY ownership, unlike dispose()) before removing the session, and tolerates a teardown throw so it can't mask the connect error being rethrown. * fix(P1-A): await async lease persistence in SSH relay teardown Failed connect attempts now wait for 'detached' leases to persist before throwing, preventing reconnects from claiming them before cleanup completes. Detach and dispose operations are now async and await store durability. * fix(ssh): make session detach lease writes retryable on failure Separate in-memory detach (identity recovery, provider cleanup) from lease write persistence so rejected writes can be re-issued without re-running provider teardown or re-minting the session identity. Introduce flushDurableStateOrThrowAsync to flush only SSH-recovery state on the live establish/reconnect path, avoiding snapshot writes of sidecars that belong to quit/startup. Use Promise.allSettled in test reset to prevent one rejected disposal from leaking state into the next test. * fix(ssh): dispose mux on failed establish and propagate sync errors - Dispose mux when session is disposed during establish to prevent resource leak - Propagate synchronous errors in teardown via the completion promise instead of leaving completion undefined - Add test coverage for terminated PTYs that exit mid-reattach and must stay dead |
||
|
|
8ab85c9bfc |
fix(quit): stop durable state writes from parking the main thread on quit (#11931)
* fix(quit): stop durable state writes from parking the main thread will-quit ran stats.flush() and store.flush() synchronously, before preventDefault(). Both fsync and rename a multi-MB file on the profile directory. When that directory sits on a stalled network mount the syscall enters an uninterruptible wait: the app stops repainting and stops responding to Force Quit, because a process blocked in the kernel ignores SIGTERM and SIGKILL alike. The existing 20s teardown deadline could not bound this. Its timer runs on the very thread the syscall parked, so it never fires. The fix is to make the quit path awaitable rather than to try to bound it — a quit that is slow but responsive stays killable by the OS. - preventDefault() now runs first, so every teardown step is free to await - stats and state gain flushAsync() twins that use node:fs/promises - both join the existing teardown barrier, which can now actually bound them - the pass-2 will-quit re-entry returns early instead of re-running teardown - quitFlushStarted makes the quit flush the last write, so a teardown step touching the store cannot arm a debounce that races process exit Making the swap async cost the atomicity of check-generation-then-rename: a writer parked on await rename has already cleared the guard, so a later synchronous flush could be clobbered by stale state. Both async writers now claim their temp path, and the sync writers delete it, turning that swap into a swallowed ENOENT. Atomic temp+rename is unchanged, so a write cut short by the deadline leaves the previous file whole — bounded loss, never corruption. * fix(quit): harden async persistence finalization * fix(persistence): bound best-effort flushes |
||
|
|
76a2317bc0 |
fix(persistence): stop blocking backup rotation (#11916)
Use async profile probes and serialize the current-time due decision plus mutations under one async owner so sync flushes and detached writers cannot double-shift or miss the recovery interval. Add syscall, timer-liveness, parity, and held-I/O interleaving coverage. |