mirror of
https://github.com/stablyai/orca.git
synced 2026-09-22 08:02:28 +00:00
f6d0bde6fb5badbf1949bc8d00a4dea3230ed617
202
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7da9368b78 |
fix(terminal): fence detached daemon endpoint ownership (#12709)
* fix(terminal): fence daemon endpoint ownership * fix(terminal): clean failed daemon PID claims * fix(terminal): close daemon ownership review gaps * test(daemon): release startup IPC in boot smoke * test(daemon): mirror production stdio in boot smoke * fix(daemon): exit after rpc shutdown cleanup * fix(terminal): make the socket name the daemon endpoint authority The reported failure was a live daemon hosting PTYs that nothing could reach: terminals acknowledged input and never ran it, listings diverged from reality, and restarting the app never helped because the detached helper survived. The ownership fence added for it could not fire in the sequence that produces the split brain. libuv unlinks the pathname a server bound to when that server closes, with no ownership check. A daemon that lost its endpoint name therefore deleted whichever socket then sat at that path — including a live replacement's — stranding a daemon that still hosted every session. Bind a private same-directory name and hard-link it into place instead: libuv can only ever unlink our own bind name, the exclusive link is a kernel-enforced endpoint claim, and the canonical name is removed only under an inode ownership check. The bind name replaces the basename rather than extending it, so it cannot overflow sun_path. killStaleDaemon removed the PID record unconditionally immediately before every fork, so the exclusive PID claim was always uncontested at bind time. It also unlinked a live daemon's endpoint whenever a connect probe merely timed out, and treated a `ps` timeout as proof of PID recycling. Now only positive evidence of a dead endpoint authorizes reclaiming it, SIGKILL is confirmed rather than assumed, and a daemon that cannot be proven stopped keeps its record and endpoint while the launcher refuses to fork beside it. A daemon whose endpoint was taken over now retires itself, draining rather than killing, so an unreachable orphan stops being permanent. A repaired PID record re-derives entryPath, appVersion and the Linux incarnation markers from the authenticated owner instead of dropping them; without appVersion a healthy daemon read as a permanently stale bundle and, on Windows, went unpinned against daemon-host pruning. Repair failure now fails open — abandoning a healthy daemon over a pid file write cost every persistent terminal on the machine. Also: treat only ENOENT as an unclaimed record so a Windows file lock is not reported as an ownership conflict; settle start() before close() so an accepted connection cannot defer it forever; sweep abandoned claim and bind names; and type the endpoint-identity seam so a rename cannot silently disable the fence. Adds a real-process handover smoke that reproduces the failure with two daemons racing one endpoint, and wires it into the native-smoke job. * fix(daemon): retire only on proven endpoint ownership loss The ownership watchdog read a null identity for any stat failure, so a transient EACCES or EIO on the runtime directory would retire a daemon that was still serving every terminal on the machine. Distinguish "the entry is gone" from "the probe failed" and act only on the former. Also require the loss to persist across two polls: a replacement publishes by unlink-then-link, and a single observation can land in that gap. * fix(daemon): source repaired ownership metadata from the authenticated hello Adversarial review found three defects in the previous two commits. Re-deriving entryPath from the owner's command line truncated it at the first space. A command line is a single space-joined string, so `C:\Program Files\Orca\...` and `/Applications/Orca 2.app/...` came back as `"C:\Program` and `/Applications/Orca`. getDaemonLaunchIdentity treats a present entryPath as authoritative, so a healthy daemon read as `different_app_path` and was killed and re-forked — worse than the missing-metadata case the derivation was added to fix. Carry entryPath and appVersion as optional fields on the daemon hello identity instead: the daemon already has both from its own argv, and per docs/reference/remote-wire-compatibility.md a new optional field is safe because every reader falls back when it is absent. This also removes a synchronous `ps` spawn from the Electron main thread during startup. `start()` rolled back the PID record even when it never published one. Losing the endpoint link now runs that path, and the ownership-checked unlink briefly renames the incumbent's record aside — enough to strand a live daemon's ownership. Roll back only what we actually wrote. publishDaemonSocketPath read its identity from the canonical name after linking, so a concurrent unlink returned null: no ownership watchdog and no endpoint cleanup on any shutdown path. Read it from the bound name before linking, which shares the inode. Refusing to fork beside an unconfirmed daemon left the user with no daemon at all and no in-app recovery, since restart re-entered the same fence. We have just proved something answers the endpoint, so adopt it in degraded mode: live sessions keep working, fresh terminals run locally. SIGTERM is also individually guarded now — an EPERM fell into the blanket catch and reported "nothing alive", authorizing the very duplicate this fence exists to prevent. Also reset the ownership-loss streak on an inconclusive probe so the confirmations are consecutive, and sweep scratch names before the launch so a failed launch still reclaims them. |
||
|
|
fde816e4ee | move folders (#12758) | ||
|
|
06780260c0 |
test(remote-runtime): run an old client and an old server against current code (#12682)
Mixed versions are the normal state of the remote-server feature: users update clients and servers independently. Until now nothing tested that. Every cross-version claim was made by code reading plus unit tests with hand-written old/new shapes — enough to catch design problems, not enough to catch a real skew regression. This runs the REAL protocol implementations from two builds against each other in one process: the actual host methods and RPC dispatcher on one side, the actual renderer multiplexer on the other, with a transport that reproduces the production asymmetry — each side decodes with its OWN codec and drops frames whose opcode it does not know. A frame survives only if the RECEIVING build understands it, which is what makes this level sufficient without launching two apps. The old side is a genuine checkout extracted from the release tag; the extracted client was confirmed to lack a symbol that exists only on main. Journey: subscribe, first snapshot, input reaching the process, live output, hide/reveal snapshot, transport drop, resubscribe, input landing again — across old->new, new->old, and a current/current control. Every step ends on an observed-state barrier; no sleeps. The oracle asserts the recorded step list, the exact 16-frame named sequence, negotiated capabilities, the exact input the host wrote to the PTY, rendered content, and zero decoder-rejected frames. A host method the stub lacks is recorded by name and asserted empty, so a harness gap cannot masquerade as a wire break. Detection is proven per violation shape, and it attributes each to the correct side: an unnegotiated opcode goes red only where a decoder would reject it, a removed published field goes red only where an old client consumes it, and a legal additive field stays green in all three pairings so the harness will not cry wolf on safe changes. It also documents the three compatibility rules in docs/reference/remote-wire-compatibility.md, linked from AGENTS.md, since they previously existed only as folklore — notably that "decoders reject unknown opcodes" is true for the desktop decoder but NOT for mobile, which silently drops them. Deliberately scoped: terminal stream only. The session-tab sync channel is not covered, nor agent-session publications, file/Git RPCs, mobile E2EE framing, or the relay transport. Two version points, so a regression introduced and reverted between them is invisible. CI selection was verified rather than assumed — `vitest list` confirms 0 matches under the shard's exclude and 4 under the dedicated job — because a lane silently running zero tests is precisely how a host-side defect escaped CI earlier in this series. Closes STA-3469. |
||
|
|
fb27702100 |
feat(updater): restart hourly build numbers per version, restyle the timestamp (#12587)
The number answers "which build of 1.4.163 is this", so carrying it across versions made it meaningless — 1.4.164 opened at 38 for no reason a reader could see. It now counts titles matching the base version being built, so a version bump restarts the series at 01. Deriving it moves from workflow jq into the script, because the number depends on the base version and only the script knows which base the published tags resolved to. Timestamps go from `07-31 13:54` to `Jul 31, 1:54PM`, still Pacific. Co-authored-by: Orca <help@stably.ai> |
||
|
|
c4ae923baf |
fix(ci): vet the adhoc build ref before running it with release secrets (#12161)
The adhoc workflow checked out any requested ref and ran its scripts and electron-builder config with MAC_CERTS, the notary password, and the adhoc publisher token in reach — including refs/pull/* fork code a maintainer could dispatch in one innocuous-looking click. Vet the ref before checkout: PR refs are refused, branches/tags resolve in a bare tree:0 scratch fetch, raw SHAs must be reachable from a repo branch or tag (a partial clone lazily serves PR-only commits by SHA, so name resolution alone is not a trust test), and checkout pins the vetted SHA so a race push cannot swap the commit. Also reference an adhoc-mac-build environment so the secrets can later be fenced off from stale workflow copies via repo settings. |
||
|
|
8ab7d8a110 |
fix(updater): base dev builds on published tags, not main's package.json (#12376)
main's version only moves on `release:` commits, and stable patches are cut from release branches that never merge back. On 2026-08-03 main read 1.4.165-rc.0 for twenty hours while 1.4.165, 1.4.166 and 1.4.167 all shipped, so every hourly built in that window was stamped 1.4.165-hourly.* while carrying code newer than 1.4.167 — and sorted below the stable its user was already running. Resolve the base from the main repo's published tags instead, taking the patch above the highest shipped stable. package.json stays a floor for the case where main leads the tags. Co-authored-by: Orca <help@stably.ai> |
||
|
|
484273844a |
feat(updater): add an adhoc release channel for branch builds (#12051)
* feat(updater): add an adhoc release channel for branch builds Hourly covers main. This covers everything that is not main yet: a dispatchable macOS build of an unlanded branch, published to stablyai/orca-adhoc, so the team can run an experimental feature for a few days instead of reasoning about it from a diff. Adhoc sits at the bottom of the version order — 'adhoc' < 'hourly' < 'rc' < stable — so no routine check can walk anyone onto somebody's branch; only an explicit pinned jump reaches one. It gets its own repo rather than sharing orca-hourly's, because a branch build must not appear in the list a developer riding main is looking at. Signed and notarized exactly like hourly, for the same reason: macOS anchors a notarized app's TCC grants on identifier + team, so an unnotarized build reads as a new client and silently loses file access under Documents/Desktop/Downloads. Tags stamp to the second rather than the minute. Hourly runs under a concurrency group and cannot overlap itself; adhoc builds are dispatched on demand, so two people cutting from different branches inside one minute is ordinary — and a minute-resolution tag would collide and fail the second build after its whole pack-and-notarize run. Channel-specific behaviour now derives from one DEDICATED_REPO_CHANNELS list: repo mapping, macOS-only support, and UpdateSource. The RPC schema that validates releaseChannelOverride was a hand-copied enum missing the new channel, which would have rejected the override on its way to the main process; it reads the predicate now. * fix(updater): merge the duplicated shared/types import Co-authored-by: Orca <help@stably.ai> * fix(ci): default the adhoc build ref to the dispatch branch The Actions UI puts its own "Use workflow from" branch picker directly above the ref field, and picking a branch there is what most people read as "build this". Making the field optional means the obvious action is also the correct one; naming a branch explicitly still wins, so main's copy of the workflow runs rather than a stale one on an old branch. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
2b44e9ed9e |
fix(updater): notarize hourly macOS builds so TCC grants survive updates (#12007)
macOS anchors a notarized Developer ID app's TCC grants on identifier + team, which is cdhash-independent and so survives an in-place update. Without a notarization ticket there is no such stable identity, so every hourly reads as a different client: the grant row stays but stops matching, and file access under Documents/Desktop/Downloads fails with EPERM and no re-prompt. `tccutil reset` fixes it until the next build — and orca-hourly has shipped as many as 14 builds in a day. Skipping notarization was chosen because Squirrel.Mac validates the replacement bundle's signature, not its notarization. That is true, but it is the wrong requirement; the in-place swap was never the problem. Budgets grow to absorb the notary round trip (publish 2x45, job 150), and the App token is re-minted after the build so its one-hour life starts at the first call that uses it rather than during `pnpm install`. |
||
|
|
edb5607e28 |
ci: block new root-level entries (#11903)
* ci: guard repository root additions * fix: clear existing type-aware lint warnings |
||
|
|
ad1e58d966 |
chore: declutter top-level repo layout (#11890)
Remove one-off incident docs and committed test-results noise, move dev/repro/bench tools under tests/tools, and relocate i18next config into config/ so the GitHub root scrolls to the description faster. |
||
|
|
676964b099 | ci: run only changed e2e specs on pull requests (#11834) | ||
|
|
cd2b62ed14 |
feat(updater): name hourly releases by version, build number, time, and sha (#11817)
* feat(updater): name hourly releases by version, build number, time, and sha Hourly releases were titled with their raw tag (`v1.4.163-hourly.202607312054`), which reads as one opaque digit run and does not say which commit it came from. Title them `1.4.163 • 01 • 07-31 13:54 • e698241` instead, and show that same string in the in-app build picker by having the picker render the release's stored name rather than deriving its own label. Composing it in one place means the two surfaces cannot drift. The build number is monotonic across the channel. It is read as the highest number already in use rather than as a count of releases: the prune step trims to 72, so a count would roll backwards after three days and reissue numbers. Drafts count toward it — unlike in the freshness check, which asks whether a commit shipped, this asks whether a number is free, and a stranded draft still holds one. Times are Pacific while the tag's stamp stays UTC. The stamp is a sort key and a local one would repeat an hour at every DST fall-back, making two distinct builds compare equal; the title is only ever read. * fix(updater): fail the hourly build when the release name is missing The workflow checks out `ref: main`, but a workflow_dispatch runs the workflow file from whatever branch was dispatched. A branch that edits this step while main still carries the old script produces an empty name and an untitled release — silent, and only visible once someone opens the releases page. Verified by hitting exactly that on run 30665586904. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
79251d7a98 |
[P2] fix(release,settings): restore signing preflight portability, bootstrap diagnostics, and skill re-check (#11692)
* fix(release): restore the SignPath composite action when cutting from an older ref Co-authored-by: Orca <help@stably.ai> * fix(startup): record a durable diagnostic before the bootstrap fatal-exit guard exits Co-authored-by: Orca <help@stably.ai> * fix(settings): make agent-skill Re-check rescan skill freshness Co-authored-by: Orca <help@stably.ai> * fix(startup): keep the bootstrap fatal diagnostic when the log override is unwritable Create the parent directory an overridden ORCA_BOOTSTRAP_FATAL_LOG names and fall back to the default location when that path still cannot be opened, so a missing parent no longer costs the only account of the failure. Also pins the Re-check freshness rescan to the completed install scan rather than the click. Co-authored-by: Orca <help@stably.ai> * refactor(settings): move the post-recheck surface sync out of the panel Co-authored-by: Orca <help@stably.ai> * fix(startup): retain diagnostics without node fs * fix(skills): keep freshness scoped to the local runtime * fix(settings): register freshness status translations * fix(settings): scope and sequence skill freshness refreshes * fix(settings): refresh freshness across runtime transitions --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
e467b3ff7b |
fix(remote): stabilize shared control and terminal parking (#11656)
* fix(remote): stabilize shared control and terminal parking * fix(remote): harden parking review edge cases * fix(terminal): restore parked local floating buffer * fix(ci): drop superseded paired parking evidence * fix(terminal): preserve floating park watchers * fix(ci): include web client in paired e2e artifact * fix(ci): reuse renderer build for paired e2e |
||
|
|
f998f7ec62 |
feat(updater): add hourly dev channel and build switching (#11250)
* feat(updater): add hourly dev channel and build switching Adds an hourly macOS build channel plus a dev-only surface for switching update channels and jumping to any published build, including older ones. Hourly builds publish to a separate stablyai/orca-hourly repo. The routine update path resolves tags from the main repo's releases atom feed, which exposes only its 10 newest entries — 24 hourly tags a day would evict every stable/RC entry there and leave real users with nothing to update to. Hourly artifacts carry the release bundle id and Developer ID signature so Squirrel.Mac can swap them in place; only notarization is skipped, which in-place updates never check. Version tails are stripped to the base (1.4.160-hourly.<stamp>, not 1.4.160-rc.3-hourly.<stamp>) so hourlies sort below both rc.N and stable and are reachable only by an explicit pinned jump, never by an ordinary check. The picker is revealed by Option-clicking the Updates header, matching the Help menu's existing hidden admin affordance. Pinned jumps set allowDowngrade and release the feed on every settle path so a jump can never leave background checks permanently deferred. * chore(hourly): create orca-hourly and add token provisioning script Adds setup-hourly-release-token.sh, which provisions HOURLY_RELEASE_TOKEN without the value ever reaching stdout, argv, or shell history: it is read with `read -rs`, passed to gh through GH_TOKEN in the environment rather than as an argument (argv is world-readable via ps), piped into `gh secret set` on stdin, and scrubbed by an EXIT trap. Verification creates and deletes a draft release in orca-hourly to prove Contents:write for real rather than trusting the permission checkbox. Drafts are absent from the releases atom feed, so the probe cannot disturb users. Refuses to run without a controlling terminal instead of falling through having set nothing, and refuses to run under xtrace, which would echo the token on every expansion. * fix(updater): address review feedback on the hourly channel Renderer: - Guard listBuilds against out-of-order responses. activeChannel flips once getVersion resolves, and rapid channel clicks stack requests, so a slower earlier load could land last and fill the list with builds from a channel the picker was no longer showing. - Selecting the running build's own channel now clears the override instead of pinning it. There was previously no way back to "follow this build's channel", so merely opening the panel left background checks pinned. - Validate releaseChannelOverride on hydration, matching every other enum-like field in that function. Main: - Exclude pinned jumps from recordCompletedUpdateCheck() in update-available. A dev browsing the picker was persisting lastUpdateCheckAt and suppressing the next real background check for a full day. - parseHourlyVersionStamp now anchors on the whole version and round-trips the parsed fields. It accepted garbage prefixes, and Date.UTC rolled impossible dates forward, so ...hourly.202602300000 rendered as March 2. Workflow: - Publish into a draft and flip it live only after the manifest check. The window between creating the release and verifying its assets previously exposed a tag the picker would offer and the download would 404 on; a draft is invisible to listReleaseBuilds, so a job that dies in that window — including a hard kill by the job timeout, which runs no cleanup step — leaves nothing user-visible behind. - Add a failure handler that discards the draft, gated on the publish step not having succeeded so a later prune failure cannot delete a live release. - Align retry budgets with the job timeout (was 60min against a worst case of ~185min, so a mid-retry kill skipped the cleanup that step exists for). - Exclude drafts from the freshness and retention queries. - persist-credentials: false; the job only reads this repo and never pushes. * refactor(hourly): authenticate with a GitHub App instead of a PAT A fine-grained PAT expires, and the hourly build would then fail silently on a schedule nobody watches. A GitHub App's private key has no expiry, so this is set up once. It is also owned by the org rather than by the person who created it, so the credential survives that person leaving. The workflow mints a short-lived installation token via actions/create-github-app-token and passes it as GH_TOKEN. Installation tokens live one hour, which is ample: this job runs no tests, no notarization, and no Windows signing, so it is pack + upload. The retry budgets and job timeout are re-sized to that reality rather than copied from the release pipeline, whose 3x45 publish budget exists for notarization and SignPath. setup-hourly-release-token.sh now provisions HOURLY_RELEASE_APP_ID and HOURLY_RELEASE_APP_PRIVATE_KEY. The key is redirected from a file straight into `gh secret set` on stdin, so its contents never enter a shell variable, argv, or the terminal. * fix(hourly): make the xtrace guard fire and cover cancelled runs The xtrace guard disabled tracing before testing for it, so `[[ -o xtrace ]]` read the state the previous line had just cleared and never fired. `bash -x` ran straight through, tracing exactly the key handling the guard exists to prevent. Test first, then disable. The draft cleanup only ran on failure(), but a run stopped from the Actions UI is cancelled(), not failed — a manual cancel mid-publish stranded the draft. Cover both. |
||
|
|
cc078a5021 |
perf(main): move hang watchdog into a worker thread (#11488)
* perf(main): add watchdog boundary memory benchmark Add a repeatable Electron 43 RSS harness that measures the production-built watchdog entry across the child-process and worker-thread boundaries. Record per-trial samples, the median, revision, runtime, and settling procedure for reproducible PR evidence. * perf(main): move hang watchdog into a worker thread Keep main-thread hang detection independent of the blocked Electron event loop without paying for a second ELECTRON_RUN_AS_NODE process. Preserve the marker and telemetry contract while moving timing configuration and heartbeats onto a bundled worker entry. * test(main): smoke packaged hang watchdog worker * fix(main): make packaged watchdog smoke able to fail The smoke reported failure only through process.exitCode, but its finally block quit Electron gracefully, and Electron takes its status from the browser exit code. Every failure mode — entry missing from app.asar, worker error, marker timeout, non-zero worker exit — exited 0 with the diagnostic discarded on stderr, so the required PR check could never go red. Propagate a real status via app.exit, assert the success line in stdout, and surface stderr. Verified against a packaged tree with the entry removed: exit 0 before, exit 1 after. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
14de3fa14d |
fix(computer): reap mac helper after client loss (#11493)
* perf(computer): add mac helper owner-loss benchmark Measure the release helper's resident memory before and after its owner-session deadline. Record exact revisions, per-trial RSS, retained state, and clean-exit latency so lifecycle reclamation is reproducible. * fix(computer): reap mac helper after client loss Bind the detached macOS helper lifetime to authenticated socket ownership. Reap the helper after its final authenticated client disconnects, and add a startup deadline for sessions that never authenticate. * test(computer): harden owner benchmark cleanup * test(computer): make owner benchmark cleanup failure-safe * test(computer): close remaining owner cleanup races |
||
|
|
8f7692aa12 |
Fix packaged skills CLI runtime ownership (#11627)
* fix(cli): make packaged skills runtime self-contained * fix(cli): address packaged skills review feedback * ci(cli): smoke packaged skills on Windows --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
a906f98baf |
fix(release): survive PSGallery outages in the Windows signing preflight
The Windows release job hard-failed in run 30125672117: every SignPath module install attempt got 403 Forbidden from the gallery's OData API, which is behind Azure Front Door and was also serving 502/504 at the time. That step was the only hard-fail in an otherwise fail-open signing chain, so a gallery incident blocked the whole release. The gallery CDN that serves the nupkg is a separate origin and stayed healthy throughout, so fall back to a pinned version fetched from it after the normal install path is exhausted. The fallback verifies a SHA-256 pin, since that route skips the gallery's own package validation. Extracted to a composite action so the release job and the signing rehearsal cannot drift apart. |
||
|
|
b370dc0900 | ci(release-cut): include source ref/commit and cutter in SignPath Slack (#10524) | ||
|
|
d0f341ad69 |
fix(computer-use): make modifier clicks interruption-safe (#11451)
* fix(computer-use): make modifier clicks interruption-safe * fix(computer-use): pace modified Windows multiclicks * fix(computer-use): address modifier safety review |
||
|
|
5e00a30e4e |
Decouple feature copy from locale parity (#8512)
* Decouple feature copy from locale parity * Fix undeclared dynamic localization key check * Fix localization code owner |
||
|
|
b339fe0346 |
Fix Node 26 test gate and happy-dom storage (#11434)
* ci: test PR shards on Node 26 * test: isolate happy-dom storage from Node globals |
||
|
|
fe6f929c6e |
fix(terminal): reconcile cross-platform IME composition lifecycle (#11293)
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: JeongUk Park <jeongph.dev@gmail.com> |
||
|
|
dde72f85de |
fix(windows): separate updater from orchestration migration (#11405)
* fix(windows): separate updater from orchestration migration * fix(terminal): attest adopted reveal identity --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
363e478909 |
fix(orchestration): preserve active workers across updates (#11271)
* fix(orchestration): preserve active workers across updates * test(ssh): model absent legacy adoption * test(orchestration): align compatibility contracts * fix(windows): escape updater PowerShell booleans * fix(windows): restore stock uninstall process check * fix(orchestration): keep recovery off renderer startup barrier * fix(orchestration): harden legacy recovery migration * fix(orchestration): close recovery review gaps * fix(orchestration): complete legacy worker cutover recovery * fix(orchestration): preserve legacy workers across updates --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
a7c8b8e071 |
fix(terminal): bound SSH & remote hidden-worktree terminal retention (C1) (#10625)
* fix(terminal): park SSH worktrees like local ones (C1 retention, slice A) SSH ptys were blanket-excluded from hidden-view parking, so a hidden SSH worktree retained every pane forever (C1: renderer heap climbs to the V8 ceiling). SSH bytes transit local main — fact-mode watchers already cover them, and main keeps a headless model served over pty:getMainBufferSnapshot that the SSH reattach path never consulted. - isParkRestorableTerminalPty: snapshot-backed OR (SSH + policy); threaded through both park verdicts, both selectors, watcher coverage, and the watcher start guard. Remote-runtime/fail-open/foreign/null unchanged. - Parked-SSH reveal paints from main's headless model (dimension-matched, ~5k rows) and degrades to the relay 100KiB replay unless the snapshot is a non-empty source==='headless' payload — never a blank/stale paint. - Kill switch: settings.terminalSshViewParking (default on). DESIGN.md records the approved plan and the H1 magnitude non-claim. Co-authored-by: Orca <help@stably.ai> * fix(terminal): bound hidden-worktree retention with a force-park budget (C1, slice B) Un-parkable worktrees (remote-runtime ptys, uncoverable tabs, SSH with the slice-A switch off) had unlimited retention: the parking cap/TTL only ever saw eligibility-passing worktrees, so one bad tab pinned a whole worktree's panes forever. Retention is now memory-bounded, not eligibility-bounded. - terminal-hidden-worktree-retention.ts: retention budget (12 hidden / 45min TTL, sized from the measured 2.5-19MB per-pane V8 cost, DESIGN.md §2) over hidden worktrees ordinary parking can never evict; reuses the hot-retain ranking so last-active exemption, deterministic ties, and deadline-driven rechecks hold. Fail-open/foreign-pty tabs are eviction-exempt (a remount would fresh-spawn and orphan the live shell). - Terminal.tsx: force-parked ids join the parked set AFTER the coverage veto (darkness for uncoverable tabs is the accepted cost); buffers captured via the sleep-flow registry before the unmount render; retention TTL added to the recheck deadlines for budget candidates only. - Verdict stays out of its own effect deps; policy test asserts idempotence and time-monotone membership (flip-loop dwell regression). - Kill switch: settings.terminalHiddenWorktreeRetentionBudget (default on). Co-authored-by: Orca <help@stably.ai> * fix(terminal): demote hidden scrollback for eviction-exempt worktrees (C1, slice C) The retention budget (slice B) must exempt worktrees holding fail-open or foreign-worktree ptys — a remount would fresh-spawn and orphan the live shell — which would leave that class unbounded again. Instead, past the same 45min retention TTL their hidden panes drop to the minimum scrollback tier (measured: ~19MB -> ~1.3MB V8 heap per 50k-row pane; trimmed history is gone by design, reveal restores the configured cap for future output). - terminal-hidden-scrollback-demotion.ts: module-state verdict registry (parked-watcher pattern) with content-equality notify damping; applied in the existing scrollback-rows effect in use-terminal-pane-lifecycle. - selectScrollbackDemotedTerminalWorktrees: pure, TTL-gated, time-monotone. - Retention TTL wakeups now also cover exempt worktrees so demotion fires. - Kill switch: settings.terminalHiddenScrollbackDemotion (default on). Co-authored-by: Orca <help@stably.ai> * fix(terminal): paint the SSH model snapshot inline, not via nested coordinator (C1 slice A fix) applyMainBufferSnapshot runs its own structuralReplayCoordinator.run; calling it from applyReattachPayload (already inside the coordinator when a relay replay exists) deadlocks on the coordinator's tail chain. The model paint now mirrors the daemon-snapshot branch inline (folded scrollback + rehydrate + screen, dimension-matched, escape tail last) and arms the restored-snapshot seq baseline so deferred/live chunks the snapshot covers dedupe instead of double-painting. Also falls through (no early return) so reattachPayloadApplied still latches. Adds the folder-workspace id parity unit case. Co-authored-by: Orca <help@stably.ai> * test(terminal): SSH park+reveal e2e round-trip + as-built design notes (C1) Docker-gated (ORCA_E2E_SSH_DOCKER=1) spec: SSH tab parks behind a decoy and reveal restores marker content at multi-viewport scrollback depth. DESIGN.md records the as-built deltas (inline paint, force-park shape, last-active floor) and the residuals so follow-ups aren't lost. Co-authored-by: Orca <help@stably.ai> * fix(terminal): paint SSH reveal from main's model even when the relay replay is empty (C1 review #1) A relay restart empties the replay buffer; the reveal previously painted nothing even when main's headless model held the session. The reattach now prefetches the model snapshot when no structural replay exists (SSH-shaped ptys only) and paints it inside the coordinator; emptiness is judged on the composed payload (scrollbackAnsi + data + pendingEscapeTailAnsi) so an alt-screen snapshot with an empty screen frame still paints. Co-authored-by: Orca <help@stably.ai> * fix(terminal): decouple scrollback demotion (slice C) from the retention-budget switch (C1 review #2) Per the approved contract each slice reverts behind its own switch: slice C now requires only the master terminalHiddenViewParking plus its own terminalHiddenScrollbackDemotion flag. The TTL wakeup timer fires for demotion candidates even with the budget switch off. No DEFAULT_SETTINGS entries exist for sibling flags (defaults are the '!== false' optional pattern), so no explicit defaults are added. Co-authored-by: Orca <help@stably.ai> * fix(terminal): scope eviction exemption to the tab, not the worktree (C1 review #3) One eviction-exempt tab (fail-open/foreign pty) previously vetoed force-park for its whole worktree, pinning co-located remote-runtime tabs forever. The worktree now force-parks while exempt tabs keep their mounted panes via a per-tab exclusion mirroring the Activity-portal pattern (legacy watcher sync, legacy render, and the overlay cold-parking hook). Ordinary parking is untouched — a worktree with an exempt tab still cannot ordinary-park. Slice C now also demotes exempt tabs' panes as soon as their worktree force-parks under the count budget (they are the only panes left mounted). Co-authored-by: Orca <help@stably.ai> * fix(terminal): demote un-parkable worktrees the force-park lever spared (C1 review #4) The last-active exemption means a single hidden un-parkable worktree never force-parks — and slice C previously only targeted exempt-tab worktrees, so its panes held full scrollback forever. Demotion now also covers un-parkable non-exempt worktrees past the retention TTL that are absent from the force-parked set (last-active spared, or slice B switched off). Membership stays time-monotone for fixed inputs; covered by new idempotence/monotone selector tests. Co-authored-by: Orca <help@stably.ai> * fix(terminal): keep the hidden clock running through transient background-measure windows (C1 review #5) Whole-worktree background mounts (browser-automation bootstrap lease, mobile mounts, agent wakes) open a ~3s self-clearing measure window that previously deleted hiddenSince — every remount restarted the 30s hysteresis and the 45min retention TTL, so a periodically re-mounted force-parked worktree never re-parked. The measure window still pauses parking/eviction verdicts (all selectors skip measuring candidates); only the clock survives, so the prior verdict resumes as soon as the window closes. Visible and portal-holding worktrees still reset the clock. Co-authored-by: Orca <help@stably.ai> * test(terminal): make the SSH park+reveal depth assertion prove the model paint (C1 review #6a) Pad the session with ~180KB of output after the numbered markers so the earliest marker falls outside the relay's 100KiB rolling replay buffer while staying inside main's ~5k-row headless model; asserting marker_1 after reveal now proves the headless-model paint rather than passing under the relay fallback. Co-authored-by: Orca <help@stably.ai> * docs(terminal): rewrite DESIGN.md as the single as-built C1 contract (review #7) One contract matching the code: status IMPLEMENTED around force-park (not the unmount proposal), real kill-switch names with coupling + revert matrices, the true retention-floor formula with measured per-pane and demotion numbers, an explicit when-OOM-is-still-possible paragraph naming the H2 pendingSideEffects residual, the applyMainBufferSnapshot deadlock constraint inside the slice-A section, stable-signal phrasing instead of a capability latch, fail-open AND foreign-worktree exemption class, verified cites, and a planned/landed/follow-up test matrix. Co-authored-by: Orca <help@stably.ai> * fix(terminal): resolve the eviction exemption per pane, not per tab (C1 review #8) isEvictionExemptTerminalTab read only tab.ptyId — the FIRST leaf's pty — while the coverage veto that makes a worktree a retention candidate walks every pane. A split tab whose second leaf held an unrestorable pty therefore failed coverage (→ force-park target) yet looked exempt-free, so force-park unmounted it and orphaned the live shell. The exemption now resolves panes through the same resolveParkedTerminalPaneCandidates, keeping tab.ptyId in the union for the no-layout/no-capture case. Also from the same review round: - force-park's capture passes includeLocalBuffers:false like every other shutdownBufferCaptures caller; it was serializing up to 512KB/pane of scrollback into the store inside a fix meant to bound renderer heap. - Terminal.tsx unmount resets the scrollback-demotion registry — module state with no reset path, read by a pane effect that runs before the host effect that would clear it, so a stale verdict trimmed restore replays. - memoize watcher coverage per tab within the parking pass; the retention candidates re-asked it for every mounted worktree, not just the parked few. * docs(terminal): drop DESIGN.md — the as-built C1 contract moves to the PR body Co-authored-by: Orca <help@stably.ai> * fix(terminal): cap the deferred PTY side-effect queue (C1 residual H2) pendingSideEffects grew without bound under background timer throttling (~64 drained/s vs hundreds queued/s overnight). Cap at 512 entries with oldest-first eviction: titles drop (last-wins), a pending bell latches onto the next survivor, agent-status payloads collapse onto the survivor keeping the newest 16 (last-wins store state, KB-scale strings). Co-authored-by: Orca <help@stably.ai> * fix(terminal): carry command-lifecycle facts through parked watchers (C1 follow-up) Parked fact-mode watchers omitted onCommandFinished/onCommandCode*, so OSC 133;D and Command Code scrape signals went dark while parked. New parked-terminal-command-status.ts ports the store-level subset: git-UI nudge on every command finish, same-turn status-row drop for SSH PTYs (exact mounted-path parity — the foreground tracker refuses SSH ids), and the Command Code working seed / 1500ms done settle. Byte mode scans the same shared parsers for authority-off parity. Local-PTY status drops stay with the mounted pane: they need pty-connection's process-confirm ladder to tell a leaked nested-shell 133;D from a real agent exit. Co-authored-by: Orca <help@stably.ai> * test(terminal): retention-budget force-park e2e with a retentionLimit override (C1 6b) ORCA_E2E_TERMINAL_RETENTION_LIMIT flows preload → e2e-config → getTerminalParkingPolicyOverrides (exposeStore-gated, positive-integer only) so a spec can shrink the force-park budget to 1. The Docker-gated spec opens two remote worktrees on one relay target (second pre-seeded remote repo), disables terminalSshViewParking to make both un-parkable, hides both behind the local context, and proves the older one force-parks while the last-active exemption spares the newest; re-activating the evicted worktree restores the marker tail via relay replay. Co-authored-by: Orca <help@stably.ai> * test(terminal): retention-budget e2e via same-repo remote worktrees (passes docker lane) The first draft added a second remote repo mid-session, whose pane pty spawn misroutes to the local daemon with the remote cwd (pre-existing multi-repo issue, reproducible without any retention override — a seeded local repo plus one remote repo shows the same misroute). The spec now budgets across three worktrees of the ONE connected repo, created through the product createWorktree path (an external git-worktree-add only lands as a detected worktree needing adoption) and polled through the relay's transient post-connect reconnect window. Verified green on the local Docker lane in 20.8s. Co-authored-by: Orca <help@stably.ai> * fix(terminal): prevent remount thrashing during post-measure cool-down ( Implements the C1 retention contract: preserve worktree `hiddenSinceMs` through a background-measure window (so TTL/ranking stay honest), but re-park waits for a full `coldParkDelayMs` cool-down after the measure ends. Without the cool-down, every ~3s measure lease on a past-deadline worktree thrashes remount/reattach. Core changes: - Terminal.tsx: add measure clock (measuringTerminalWorktreeIdsRef) and post-measure cool-down tracking (terminalWorktreeParkCooldownUntilRef); gate parking candidates until cool-down expires. - Extract snapshot replay choreography to shared terminal-snapshot-replay-paint.ts (used by SSH reattach + daemon restore paths). - Add SSH model snapshot timeout (750ms) with fallback to relay replay. - Move cold-park recheck deadline logic to terminal-cold-park-recheck-deadlines.ts; add cool-down deadline to scheduling. - useTerminalTabColdParking: implement matching measure-clock contract with per-tab cool-down gate to keep tab deadlines synced with worktree retention clock. - Add resolveTerminalMountScrollbackRows() to demote new xterms under demoted worktrees (pane births during demotion must take the demoted tier at create). - Add kill switches: terminalSshViewParking, terminalHiddenWorktreeRetentionBudget, terminalHiddenScrollbackDemotion. * fix(terminal): detect Command Code completion in parked mid-turn panes Seed the byte watcher with in-flight turn state from agent status: the watcher is recreated per park cycle with no startup command to arm it, and the banner scrolled away before parking. Also memoize eviction-exempt checks and use SSH PTY ID builder in tests. * fix(terminal): flush pending command-code settles on reveal remount When a parked pane reveals mid-Command Code turn, the new detector cannot re-observe the already-passed idle composer. Cancelling the settle leaves the row stranded at 'working', so dispose now flushes the pending settle instead. Extract readInFlightCommandCodeTurn to shared space and seed detectors with in-flight turns so remounts complete mid-flight commands. Also memoize SSH model probes to prevent double timeouts on reattach. * fix(terminal): remove scrollback demotion (C1 slice C) The scrollback demotion feature for eviction-exempt hidden worktrees is no longer needed. Retention budget limits are now sufficient without this additional bound. Remove the terminal-hidden-scrollback-demotion module, the selectScrollbackDemotedTerminalWorktrees function, and related per-pane demotion logic. * test(terminal): assert bounded probe during stalled reveal Add assertion to verify that a stalled reveal operation makes exactly one `getMainBufferSnapshot` call, ensuring retry logic doesn't introduce redundant probes that would extend the timeout window before relay fallback. * fix(terminal): implement C1 retention budget for hidden parked worktrees Addresses OOM regressions in hidden parked terminals by force-evicting worktrees past a retention budget: at most 12 mounted while hidden, none past 45 minutes (absolute, not exempted by last-active). Eviction is least-recently-hidden-first. Exempt tabs (unrestorable local PTYs) keep their panes to avoid orphaning shells; worktrees are force-parked even if they contain exempts, and their buffers released elsewhere. SSH/remote worktrees serialize buffers pre-eviction for reveal; local worktrees keep daemon snapshots. Command Code's done-settle window is transferred across park/reveal boundaries so the row cannot strand at 'working'. Model probe on SSH reattach is scoped to park-reveal only, not ordinary reconnects. Includes new E2E suite proving the budget actually releases memory. * memoize eviction-exempt terminal tabs to avoid redundant store reads Each tab's exemption check re-reads the store and walks the layout tree. Introduce selectEvictionExemptTerminalTabIds() to resolve all exempt tabs for a worktree in a single pass, then memoize the result in Terminal.tsx and useTerminalTabColdParking. This prevents O(n) store reads when checking exemptions across multiple tabs and ensures the set remains stable across unrelated re-renders. * refactor: reformat hidden-worktree retention comments Reflow to 80-character lines and remove internal ticket references (C1, C1 slice C). * fix(lint): split overlay slot and eviction-exempt tabs under max-lines Static analysis failed because TerminalPaneOverlayLayer (401) and terminal-parked-tab-watchers (304) exceeded oxlint max-lines. Extract the slot component and eviction-exempt helpers into dedicated modules. * test(terminal): stabilize retention budget e2e control arm Stage un-parkable remote pty ids only after both worktrees are hidden, and keep re-staging during the control-arm poll so a late updateTabPtyId cannot flip the decoy back to park-restorable and ordinary-park it before budget engages. * test(terminal): pin retention e2e decoy to a mounted pane snapshot Use the active pane-identity snapshot for the decoy tab instead of all worktree tabs, and re-assert un-parkable ids after the control-arm hold so a deferred/empty tab id cannot fail the budget-off mounted-count check. * fix: memoize terminal eviction exemptions on layout leaf PTYs Splits add leaf panes to the layout store without changing the tabs array. A memo keyed only on tabs misses this change, leaving new panes unexempted for unmount. Include layout leaf PTYs in the exemption memo key so it recalculates when splits occur or PTYs are re-minted. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
a40183389b |
feat: bound direct SSH reconnect fan-out and recovery (#11003)
* docs: design for direct SSH reconnect fan-out Capture the implementation-ready plan for host-qualified, epoch-fenced SSH reconnect recovery after two rounds of multi-model LLM counsel review. * docs: reconcile SSH reconnect fan-out design * docs: close reconnect design consistency gaps * feat: implement bounded direct SSH reconnect recovery * fix: bound direct SSH retry settlement * fix: harden direct SSH reconnect authority * fix: preserve split SSH retry ownership * fix: preserve SSH split continuation authority * docs: record final SSH reconnect validation * fix: preserve SSH authority through retained and detached state * fix: retain SSH authority across delayed split mounts * fix: close SSH authority recovery gaps * fix: fence stale SSH transport replacement * fix: serialize SSH target teardown * fix: settle SSH teardown failures before reconnect * fix: retire failed SSH reset sessions * test: reconcile current main E2E contracts * fix: close direct SSH reconnect review gaps * fix: fence stale SSH reconnect side effects * fix: close final SSH reconnect lifecycle gaps * test: stabilize current-main reliability gates * test: prove plugin navigation containment * test: make plugin navigation oracle authoritative * test: make plugin navigation oracle deterministic * ci: allow sharded e2e suite to finish * test: wait for runtime pane publication * test: classify pane readiness by error code * test: select close persistence terminal by tab identity * docs: mark reconnect implementation validated --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
1fa9ffb5ea |
ci(pr): run E2E when a PR touches tests/e2e paths (advisory) (#11131)
* ci(pr): run E2E when a PR touches tests/e2e paths Regression specs under tests/e2e never ran on PR CI — only schedule and release called e2e.yml — so a red regression test could merge green. Path-filter and workflow_call the E2E suite when E2E-relevant files change. Use merge-base diffs so base-branch drift does not false-trigger E2E, fail the detector when git diff cannot compute the PR range, and pin least-privilege contents:read on both the detector and reusable E2E workflow. Closes #10518 Co-authored-by: Wooseong Kim <innocarpe@gmail.com> Co-authored-by: Orca <help@stably.ai> * ci(pr): make the E2E path gate actually block, and match the real config path Two fixes to the new path-filtered E2E job. The gate did not gate. pr.yml's `verify` job is the required check, and it enumerates its dependencies explicitly — `e2e` was in neither `needs` nor the result list, so a failing shard left `verify` green. That reproduces the exact hole this job exists to close: a red spec merges green, just with a red box further down the page. Add `e2e` to both. Because the job is path-filtered, `skipped` is the normal result on a PR that touches no E2E files and has to keep passing. That allowance is checked after the strict loop rather than inside it, so it can never leak to the six jobs that are always required. The `playwright.` pattern matched nothing. The config is tests/playwright.config.ts — beside tests/e2e/, not inside it — so no tracked file starts with `playwright.` and editing the runner config would silently skip E2E. Anchor it at `tests/playwright.`. Adds a contract test alongside the existing release-e2e one. Verified it fails when either fix is reverted, and simulated the gate across success/skipped/failure/cancelled plus the skip-must-not-mask-a-real-failure case. * test(ci): close two gaps in the E2E gate contract CodeRabbit was right on both counts — verified by reverting each and watching the contract stay green. The path filter was unasserted, so `e2e` could lose its `if:` and run on every PR — the cost the filter exists to avoid — without failing anything. The strict-loop check hardcoded four of the six required jobs, so dropping GIT_COMPATIBILITY or SHELL_CONTRACTS left them unenforced while the contract passed. Derive the list from verify.needs instead, so a newly added required job that misses the loop fails here rather than silently going unchecked. * ci(pr): land the E2E path gate advisory instead of blocking The E2E suite is currently failing every scheduled run on main — 22 of the last 22 — so making verify depend on it would block any PR touching tests/e2e/**, including the PRs that fix the suite. This PR's own run reproduced that: 3 of 12 shards failed on specs unrelated to it (agent-session resume, Jira linking, plugin containment, terminal artifacts). So the job runs and reports on E2E-path PRs but is left out of verify.needs for now. The detector, the tests/playwright. path fix, and the contract tests are unaffected — those stand on their own and were the substance of the review. Flipping to blocking is a three-line change once the suite is green; the exact wiring, including why the skipped allowance must sit outside the strict loop, is recorded on verify's Require-successful-checks step. The contract test pins the advisory choice so it reads as deliberate rather than as the unwired-gate bug it originally caught, and still fails if the path filter, the strict-loop coverage, or the config path regress. --------- Co-authored-by: Wooseong Kim <innocarpe@gmail.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
e551d3ec0d |
perf(lint): consolidate code-quality gates into Oxlint (#11117)
Consolidate standalone code-quality scanners into Oxlint, preserve focused native/type-aware enforcement, add custom plugin coverage, and harden deferred PTY test cleanup. |
||
|
|
badf91101b |
fix(quality): enforce performance-safe lint baseline (#11074)
* fix(quality): clear safe existing lint findings * fix(quality): keep lint cleanup allocation-free * fix(quality): enforce performance-safe baseline * test(terminal): drain deferred confirmation cleanup |
||
|
|
12ef12c55b |
chore(quality): ratchet Oxlint, React Doctor, and Zustand performance (#11034)
* chore(quality): ratchet lint and Zustand performance * fix(ci): stabilize React peer lock snapshot * fix(ci): isolate PR diff and React Doctor CLI |
||
|
|
c6076a507c |
ci(release): detach non-blocking E2E (#11031)
* ci(release): detach non-blocking E2E * test(release): pin E2E dispatch retries |
||
|
|
abcdc04f6b |
fix(ci): mirror missing lint steps in PR workflow (#10601) (#10623)
Reviewed with an independent reproduction. Added the allowlist entry that unblocked verify:localization-coverage on main, the 4th drifted step, and a parity gate that fails when pnpm lint's chain contains a script absent from pr.yml. |
||
|
|
39a200d900 |
fix(release): restore the Windows inner-binary signature gate (#6487) (#10719)
* fix(release): restore the Windows inner-binary signature gate
electron-builder 26.9+ dropped the bundled 7zip-bin package, so the gate's
hardcoded node_modules/7zip-bin path stopped resolving in
|
||
|
|
0f91af821d |
ci: parallelize PR checks and accelerate Vite builds (#10989)
* ci: parallelize and accelerate PR checks * fix(ci): make accelerated checks runtime-safe * fix(ci): address review findings * fix(ci): retry transient Electron downloads * test(ci): cover Electron download retry limits |
||
|
|
97e4776dfe |
feat(plugins): Orca plugin system — kernel, content packs, panels, workers, marketplace v0 (experimental) (#8549)
* feat(plugins): Orca plugin system — kernel, content packs, panels, workers, marketplace v0 (experimental) Adds Orca's experimental plugin system behind a settings flag: a supervised kernel, declarative content packs (VM recipes, commands and keybindings, language packs), sandboxed iframe panels, forked worker hosts, and a Git-backed marketplace v0 with consent, provenance and kill-list enforcement. Theme, icon-theme and terminal-theme contributions are deferred to a follow-up pass. * fix(plugins): make unsupported marketplace listings unreachable by key findPlugin() backs preview/install/previewInstalledUpdate via requireListing(), so filtering only listPlugins() hid the catalog card while leaving the dead install path reachable one click later. * fix(plugins): fan Pi session-only status out to plugin subscribers The providerSessionOnly early-return in applyNormalizedStatus emitted to onAgentStatus (main-window fanout) but skipped enrichedStatusListeners, so plugins subscribed to agent.status.changed silently missed every Pi session_start event. Route both emit sites through one helper so a future early return cannot drop the plugin tap again. Co-authored-by: Orca <help@stably.ai> * plugins: drop dead code and hoist duplicated trust-boundary patterns Cleanup pass over the P1 diff, no behavior change: - Delete `readPluginTreeSnapshot`/`readSnapshotFile` and their types, plus the now-vestigial `directories`/`signal` plumbing in `collectFiles`. - Delete `resolveContainedPluginDirectory` (no callers). - Delete `plugin-content-load-pool.ts`; it reimplemented the existing `mapWithConcurrency`, whose index arg also removes the pairing wrapper in `buildPluginList`. - Hoist `PLUGIN_CONTENT_HASH_PATTERN` and `PLUGIN_COMMIT_PATTERN` into the install-lockfile module; 11 sites hand-rolled these identically. - Point the new reliability gate at the PR instead of gitignored docs paths, matching every other gate's link form. * fix(plugins): retry plugin state renames on Windows AV/EPERM locks Six plugin write paths (lockfile, provenance, current pointer, kill list, marketplace cache, staged install dir) did a plain rename, so an antivirus or indexer holding the target open surfaced as a failed install. The repo already retries this hazard for issue #1507, but only through a sync helper; these paths are all async. Adds one bounded async retry + atomic write used by all six, and trims a consent-provenance header that restated its own JSX. * test(plugins): cover the Windows rename retry path The retry loop shipped untested: both existing cases hit the non-retry path, and the temp-cleanup test passed identically with the `finally` removed. Mock `rename` to queue errno codes so CI can exercise locks it cannot provoke. Co-authored-by: Orca <help@stably.ai> * fix(plugins): pin bundled plugin resources to LF Windows CI checks out with autocrlf, so the byte-hashed launch tree arrived as CRLF and verify-packaged-plugin-resources rejected it — the packaged build could never pass on Windows. Reproduced locally: CRLF yields the exact CI error, LF verifies clean. Files are already LF, so nothing renormalizes. Co-authored-by: Orca <help@stably.ai> * test: guard the bundled-plugin LF pin against a CRLF checkout The byte-hash mismatch only surfaced in Windows packaging CI. Assert the .gitattributes pin and that a CRLF tree is rejected, so a regression fails on any platform instead of waiting for a packaged Windows build. Co-authored-by: Orca <help@stably.ai> * ci: trigger packaged-build check on bundled plugin resource changes The launch tree is byte-hashed during packaging, but no trigger path covered it — so the CRLF fix for that check would not have re-run the check. Add the resources, verifier and .gitattributes paths that can break packaging. Co-authored-by: Orca <help@stably.ai> * perf(plugins): rebuild the panel frame only when its baked theme values change The revision keys the panel iframe, so every bump destroys the sandboxed frame and its in-panel state. It counted root attribute mutations, but --workspace-sidebar-live-width is written every rAF of a sidebar drag, so dragging with a panel open blanked it ~60x/sec. Compare the two values the shell actually bakes in instead. Co-authored-by: Orca <help@stably.ai> * test: stop pinning a plugin name in the CRLF guard The CRLF case rewrites every launch file, so the reported mismatch is whichever plugin sorts first. P2 adds theme plugins that sort ahead of orca-navigation-shortcuts, which broke the assertion there. Co-authored-by: Orca <help@stably.ai> * style: drop stray blank lines left by the rebase resolutions Both sides of the agent-hooks and orca-runtime conflicts contributed a trailing blank, which oxfmt rejects. Whitespace only. Co-authored-by: Orca <help@stably.ai> * test(plugins): stop the startup budget failing on machine load P95 runs 16-34ms idle but exceeds the 50ms bound under full-suite parallelism, so the gate flaked. Widen it to catch an order-of-magnitude regression instead; the no-worker/no-plugin-code assertions are the real guarantee. Verified a 400ms regression still fails. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
fc513233cb |
fix(release-cut): gate an explicit RC against its own series (#10525)
* fix(release-cut): gate an explicit RC against its own series
semver_gt compares through strip_pre(), so the explicit-version override
only ever checked the stable line: 1.4.156-rc.0 read as 1.4.156, cleared
a 1.4.155 stable, and republished an RC below what clients already run.
Anchor a prerelease request on highest_rc_for_base -- the same rc history
the kind path uses -- so the override can only advance the series.
Two sibling gaps in the same block:
- version_suffix was silently dropped when version was set, because the
append lives in the kind branch the override skips.
- the shape regex rejected X.Y.Z-rc.N.suffix, so a suffixed RC the rc
path can produce could never be re-cut explicitly.
* fix(release-cut): close both ends of the rc-number range the gate compares
The new explicit-rc gate compares with `[[ -le ]]`, i.e. bash machine-width
integers, and the author closed only the low end. Past INTMAX bash saturates,
so `version=1.4.156-rc.99999999999999999999` reads as "above the published
rc.3" and the gate falls open — then the tag it cuts pins
highest_rc_for_base at 1e20 for that base forever, and every later cut wraps
to a lower rc the fleet never updates to. Bound the rc number to nine digits.
Also reject leading zeros on an all-digit prerelease identifier. `npm version`
renormalizes rc.4.01 to rc.4.1 while the tag step keeps the literal input, so
the shipped package.json version and its own release tag name different
releases. The explicit path's embedded identifier now goes through the same
validator the kind path uses instead of only the shape regex.
* fix(release-cut): stop the refusal pointing minor/major RCs at the wrong series
kind=rc derives its base from bump(latest_stable, patch), so the remedy the
refusal suggested only works when the requested base *is* that next patch. A
1.5.0-rc.N series exists only because this override created it, so an operator
resuming a stuck 1.5.0-rc.2 was told to dispatch kind=rc, which would have cut
an unrelated 1.4.156-rc.4. Spell the condition out and give the fallback that
does work for a non-patch base.
Also correct the mechanism in the comment I added in
|
||
|
|
f009500677 | ci(release-cut): always show resolved commit, branch, and tag in summary (#10482) | ||
|
|
7a01910f20 |
fix(skills): advance the release ledger at the cut so shipped revisions freeze (#10483)
* fix(skills): advance the release ledger at the cut so shipped revisions freeze #10340 made the released-skill registry a function of the committed ledger instead of a git tag walk, and #10460 reverted the cut step that advances that ledger because it violated the #9119 contract (a version-only cut must not regenerate or stage the content-addressed skill artifacts). Both were right; the result is a ledger that never advances. generate-skill-bundle-manifest.mjs:390 derives releasedCount solely from release-mapping.json and :461 assigns a changed skill releaseRevision = releasedCount + 1, while :518 protects only committedReleasedCounts[name] — so index releasedCount is unprotected. A tag ships that tail revision, nothing records it, and the next skill change rebuilds the same revision number over different bytes. Installs carrying the shipped digest then match no snapshot and degrade to unrecognized, which cannot be updated. Restore the advance in a form the #9119 contract can keep enforcing: --release now verifies that current-manifest.json and snapshot-registry.json already match the ref being tagged, appends the mapping row, and writes only release-mapping.json. The cut stages just that file, so it still cannot move a content-addressed artifact — the failure #9119 guarded against — and now fails loudly instead of recording a revision the tag does not ship. The contract test is narrowed to match: it asserts the cut runs --release (never --write) and stages exactly package.json and release-mapping.json. * test(release-cut): close the staging bypasses the narrowed gate left open The narrowed contract test anchored its `git add` scan to line start and only inspected staged paths, so three ways to reintroduce #9119 stayed green: a `git add` chained after `&&`, a write that never calls `git add` at all, and `pnpm run generate:skill-bundle-manifest` — the package.json alias for `--write`, which the hyphenated ban never matched. That last one also passed the pre-#10460 assertions, so it was never covered. Drop the anchor, require every `resources/skills` mention in the step to be exactly what is staged, and ban the alias and `commit -a`. Comments are stripped first so prose cannot trip a ban. Verified each bypass fails and the real workflow passes. * fix(release-cut): make the new provenance failure actionable to an operator Verifying the content-addressed artifacts is the only new way the cut can block, and it fails inside a step named "Bump package.json and tag" with a lint-shaped message. That names the files and the command but not the two things the operator needs: the regeneration has to land on main, and the cut is safe to re-run afterwards. Say so. Also pin down why assertReleasedHistoryPreserved takes the pre-append mapping. It pairs with artifacts.releasedSnapshotCounts, which seeding fixed before the row existed; handing it the post-append mapping makes every cut throw "Released snapshot history is incomplete", which points at tag fetching rather than the real cause. Nothing enforces the pairing. * test(release-cut): gate the whole cut job, not just the bump step Round-2 review defeated the previous gate twice, both proved by running the full contract file green with #9119 reintroduced. Every step in the cut job shares one workspace and one index, but the contract test only inspected `Bump package.json and tag`. A step inserted earlier could run --write and `git add resources/skills`, and the bump step's own commit swept it into the version commit and the tag. Assert job-wide instead: only the bump step may name the directory, and no step may regenerate under either the flag or its package.json alias. That lives in the generator suite because the contract file is at its max-lines cap. Two regexes were also evadable. The mention scan required a trailing slash, so a path held in a variable was invisible; it now matches the directory itself. The `commit -a` ban matched nothing at all — `commit\s` ate the only separator, so `-a`, `-am`, and `--all` all survived while only a trailing `-a` was caught. `--allow-empty` stays allowed. * fix(release-cut): assert the index, not the workflow text, before committing Round-3 review defeated the job-wide grep three ways, each proved by running both test files green with #9119 reintroduced into the tagged commit: an `env:` block holding `--write` and `resources/skills`, a composite action whose steps the workflow never spells out, and plain shell concatenation (`root=resources; leaf=skills`). Grepping shell source for path literals is inherently evadable, and the previous fix only relocated round-2's variable-indirection hole one step over. Move the invariant to where it cannot be dodged: immediately before committing, the cut diffs its own index and refuses anything that is not package.json or the release-mapping row. That does not care which step staged what, or how the path was spelled. The workflow grep stays as a cheap tripwire for literal spellings, now paired with a positive assertion that the index guard exists and precedes the commit — indirection cannot hide a missing guard. Mention matching dedupes and trims quotes, since the guard names the row a second time. * fix(release-cut): match the staged-path allowlist literally `grep -vx` treats its patterns as regexes, so the `.` in `package.json` matched any character: a staged `packageXjson` or a `resources/skills/release-mappingXjson` was silently accepted by the index guard. Verified both slip through `-vx` and are caught by `-vxF`. Exercised the guard against a legitimate cut, an empty index, a staged content-addressed artifact, paths containing a space and a non-ASCII character (git quotes the latter, so it fails closed), and a staged deletion. Only the two allowed paths pass. * test(release-cut): assert the index guard aborts, not just that it exists The positive assertion pinned the guard's shape and its position before the commit, but not its effect: replacing `exit 1` with `:` left both test files green while the cut logged the error and shipped the artifact anyway. That is the same failure this whole gate keeps having — asserting the shape of a defense rather than what it does. Pin the abort too. Verified the neutered guard now fails the suite. * test(release-cut): scope the abort check and catch clustered commit flags Two holes in the guards this PR added, both in the same shape-not-effect class the previous commit was meant to close. The abort assertion's lazy match was not scoped to the guard's own block, so it could borrow an `exit 1` from any later `if ... fi` in the step. Degrading the guard to a warning while adding a plausible HEAD precondition left every test green. Stop the match at the guard's `fi`. The `commit -a` ban only matched when `a` led the flag cluster, so `-vam`, `-va`, `-qam` and `-sam` all survived. That matters more than it looks: `commit -a` stages at commit time, after the index guard has already inspected a clean index, so it is the one way to defeat that guard. Match `a` anywhere in a short-flag cluster; `--allow-empty` and `--amend` stay allowed. Verified both mutants now fail. * fix(release-cut): validate the commit, not the index, before tagging The index guard asserted the wrong thing. `git commit` has a family of forms that commit the working tree rather than the index — `-a`, `-i`, `--only`, and a bare pathspec — so a rogue earlier step could leave regenerated artifacts unstaged and any of those forms would carry them into the tagged commit while the guard saw a clean index and passed. Reproduced end to end: `git commit -i resources` put current-manifest.json and snapshot-registry.json in the tag with all gates green, and `--only resources` additionally dropped package.json from the tag. Banning those flags one by one is the same enumeration game the earlier rounds kept losing. Assert the outcome instead: after committing and before tagging, diff-tree HEAD and refuse anything that is not package.json or the release-mapping row. That is indifferent to which step staged what and to how the commit was spelled. Verified the whole family is now blocked (-i, --only, -a, -am, -vam, pathspec, and an alias expanding to `commit -i`), that a stock commit and an --allow-empty re-cut still pass, and that deleting, neutering, un-anchoring, or relocating the guard each fails the suite. * fix(release-cut): make the commit guard fail closed on a merge commit Plain `git diff-tree` prints nothing for a merge commit, so the guard would have passed silently instead of failing closed — the one direction that matters on a release path. `-m --first-parent` reports the diff against the first parent; verified byte-identical output for an ordinary commit and still empty for the `--allow-empty` re-cut, so nothing else changes. Not reachable today (nothing in the cut job creates a merge, and npm version has no lifecycle hooks defined), but the failure mode is a guard that looks like it ran. Pin the flags in the assertion too, so neither dropping -m nor slipping in a `--diff-filter` can weaken it without failing the suite. |
||
|
|
2653794c82 |
fix(terminal): verify Windows PTY root identity before taskkill /T /F (#10484)
* fix(terminal): verify Windows PTY root identity before taskkill /T /F killWithDescendantSweep guarded its Windows tree kill with ownsRoot() alone, which is JS state only. node-pty's ConPTY exit watcher closes the last shell handle before it queues the JS exit callback, so Windows can recycle the PID while the session map still looks live — force-killing an unrelated process and its whole descendant tree. Walk the recycled PID's ancestry back to this process before taskkill: skip the sweep when the root is gone or resolves to a stranger, and keep the sweep when identity is unknown so #10004 orphan cleanup still runs. Also gate the local provider's ownsRoot on observed physical exit. * fix(terminal): dedupe the Windows root-identity scan, drop dead exit gate Review fixes on the PID-identity guard. The probe read the process table through a new uncached export, bypassing the reader that worktree teardown depends on: worktree-teardown.ts fans out 32-wide inside a 10s deadline, so a delete forked 32 powershell cold-starts (the churn windows-foreground-process-rows.ts:25-32 warns about, #6288/#6667). getFreshSnapshot() already guarantees a scan that starts after the request -- the exact property the bypass existed for -- and coalesces concurrent callers, so use it. Measured on the new test: 32 scans -> 1. The PhysicalExitTracker.hasExited gate could never fire. markExited() is only reached at local-pty-provider.ts:985/:1431, and both are followed synchronously by clearPtyState(), which deletes the ptyProcesses entry -- so ownsRoot's map check is already false whenever hasExited is true. Reverting it broke no test. Drop it and the shared getter it added; the identity probe already covers every ownsRoot caller from inside killWithDescendantSweep. Also point the Windows terminal-restart E2E job at the files that own this behavior, so a change to the new Windows-only module runs the one job that executes on a real Windows host. * docs(terminal): state what the Windows root probe actually proves The probe checks subtree membership, not root identity: a recycle that lands on another Orca descendant (another pane's shell, an agent CLI, a git.exe we spawned) still reads `own`, and that is not remote during teardown when Orca is itself allocating pids. It bounds the blast radius rather than closing the class. Say so at the type and at classifyWindowsTreeKillTarget, and name what a real close would need (a CreationDate baseline -- the analogue of the POSIX lstart check already used here -- or an inherited handle / Job Object). Also note why our own pid must classify `foreign`. * ci(windows): trigger the terminal-restart E2E on the shared snapshot reader The Windows root-identity probe now reads through getFreshSnapshot, so an edit to that module changes Windows teardown behavior without touching any path the job already watches. * test(terminal): guard the teardown probe against a reintroduced scan bypass The existing volume guard covers queryWindowsProcessRowsFresh directly, but the identity-probe cases all inject readRows, so nothing exercised the DEFAULT reader wiring -- a bypass reintroduced inside windows-pty-root-identity would have gone unnoticed. Drive verifyWindowsTreeKillTarget 32-wide through the real reader and assert one scan. Verified it fails at 32 when the bypass is put back. |
||
|
|
d5340fd191 |
ci(release-cut): restore the skill-independent version commit (#10460)
#10340 added a ledger-advance step to the release cut, which violates the contract test #9119 added: the cut must not run generate-skill-bundle-manifest or stage resources/skills. Both PRs were green on their own branches and only conflicted once merged, so nothing failed until main had both — main and every open PR have been red since. Revert the two workflow lines. The script's --release implementation stays: it is correct and harmless when unused, and the root fix in #10340 — verify no longer walking git tags — does not depend on the cut step. This leaves #10340 semantically incomplete and that must not be dropped. With the registry seeded from the committed ledger instead of a tag walk, nothing advances the ledger at cut, so each new skill change re-uses the same unreleased tail revision for different bytes; older installs then match no known snapshot and degrade to unrecognized, which reads in the UI as a skill that needs attention and cannot be updated. Follow-up is to reintroduce the advance narrowly — stage only resources/skills/release-mapping.json and narrow the assertion to forbid mutating the content-addressed artifacts while permitting the provenance row. |
||
|
|
8d61d76a59 |
fix(skills): decouple skill-manifest verify from local git tags (#10340)
* fix(skills): source released history from the committed ledger, not a tag walk verify:skill-bundle-manifest rebuilt the entire released-skill history by walking every local refs/tags/v* on each run and demanded byte-equality with the committed artifacts. Output was therefore a function of (skill bytes x local tag set x release timing), so any clone holding stray, deleted, or fork tags the committed artifacts predate rebuilt a divergent registry and failed lint. This was the 4th instance of one failure class (#8637 -> #9119 version bumps -> #9778 new tags -> local tag drift), each patched with a new tolerance rather than removing the tag coupling. Fix: the committed snapshot-registry + release-mapping ARE the released history; trust them instead of re-deriving from tags. - releasedHistoryFromCommitted() seeds generation from the committed ledger, dropping the floating unreleased tail (entries beyond what the mapping names). verify and --write are now pure functions of working-tree bytes with zero tag access. The tag walk survives only behind --rebuild-from-tags (disaster recovery), off the everyday path. - --release <version> + appendReleaseRow() perform the O(1) append of one mapping row at release cut (dedupes vs the last row, strips the v-prefix) -- the single authoritative point where working-tree bytes become an immutable released revision. - release-cut.yml runs generate --release "$VERSION" before the release commit (Node built-ins only, no install needed); pr.yml drops fetch-depth: 0 from the lint job since verify no longer needs tag history. Recognition is unaffected: the runtime uses knownSnapshots = registry.skills (all entries, incl. the tail committed at PR-merge time), so a missing mapping row only loses a version label, never recognition or the update nudge. Trade-off: lint no longer cross-checks committed historical snapshots against tags. A hand-edit to an old released entry is still caught by the runtime manifest<->registry consistency check when the current manifest points at it, and can be audited anytime with --rebuild-from-tags. Verified: verify passes committed-sourced; --write is zero-diff (byte parity); a planted stray v-tag no longer changes output; edit-stub -> --write -> --release appends the correct single row; double --release is idempotent; --rebuild-from-tags reproduces the committed artifacts. Generator tests 14 pass/ 1 skip; runtime skill-bundle-artifacts + freshness-inventory 14 pass; bundled skill guides verify passes. * fix(skills): keep one release-mapping row per version on a re-cut A cut that pushed the version bump to main but died before pushing the tag is re-cut at the same version. If skills changed in between, the second --release appended a duplicate row, and the stale one named revisions that tag never ships — which verify-skill-update-roundtrip then pairs with the tag's real bytes. Overwrite the trailing row instead (the tag is absent, so that version was never published). Refuse only when an earlier row claims the version, which the cut workflow already rejects upstream, so this cannot wedge a recovering cut. |
||
|
|
efe996a007 |
ci(release-cut): add explicit version override to the cut dispatch (#10329)
* ci(release-cut): add explicit version override to the cut dispatch Kind-based computation derives the next version from the latest *published* stable. When a shipped stable is deleted/rolled back, the release list regresses to the prior stable, so a `kind` cut recomputes a number at or below the deleted one — stranding every client that already installed it, since electron-updater only moves forward. The existing package.json floor only recovers this when the deleted version's bump commit is on the ref being cut, which a hotfix cut from an older RC ref does not carry. Add an optional `version` workflow_dispatch input that lets a human assert the exact target (e.g. leapfrog a deleted 1.4.154 to 1.4.155), bypassing kind-based computation. The updater-safety gate (must exceed the latest published stable) and the existing tag-collision recovery still apply. Empty by default, and forced empty for scheduled cuts, so normal automation is unchanged. * ci(release-cut): let explicit version override the package-floor recovery Per review: the package.json floor block can recover_unpublished_tag and exit 0 before the explicit-version branch runs, hijacking an explicit request to recover a floor tag instead — the exact rollback scenario the override targets. Skip floor-tag recovery when EXPLICIT_VERSION is set; latest_stable is still raised to the floor for the safety gate, and the requested tag's collision recovery runs later. |
||
|
|
aab112933e |
Revert "fix(memory): bound OOM-prone accumulators (#10179)" (#10255)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
801ff57e83 |
fix(mobile): unblock iOS releases (#10224)
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
8f40ddf328 | fix(memory): bound OOM-prone accumulators (#10179) | ||
|
|
56101422d2 |
fix(release): regenerate Windows blockmap via app-builder-lib JS (#10110)
electron-builder 26 dropped the app-builder-bin Go binary, so the signed-installer staging step failed with 'node_modules/app-builder-bin/ win/x64/app-builder.exe is not recognized'. Blockmap generation now lives in app-builder-lib's pure-JS buildBlockMap; call it through a small script in both the release-cut and signing-rehearsal workflows. Co-authored-by: Orca <help@stably.ai> |
||
|
|
5d0c29f729 |
ci(release-cut): show workflow_dispatch inputs in job summary (#9866)
Mirror the noqa deploy workflow pattern so Cut Release runs surface kind/ref/dry_run/version_suffix as a table under Workflow Input Parameters. |
||
|
|
0b71f3bfba |
test(e2e): prove the terminal daemon survives a main-process crash on Windows (#7742) (#9311)
* test(e2e): prove the terminal daemon survives a main-process crash on Windows (#7742) Add a win-crash-survival e2e harness (sibling to win-update-e2e) that force-kills ONLY the packaged app's real Electron main (resolved via app.evaluate -> process.pid, /F no /T) and asserts the detached orca-terminal-daemon.exe plus its ConPTY shell survive with no pwsh 0xE9 FailFast, then that a relaunch re-adopts the SAME daemon and the reattached UI binds to the SAME survivor shell (proved via a per-shell env sentinel read back through the restored terminal). This guards the #7742 fix (standalone relocated daemon that outlives main death) against regression. A directional `--expect orphaned` profile fails on a fixed build, keeping the survival assertions honest. Windows-only; reuses win-update-e2e app-driver/daemon-process modules. * test(e2e): harden Windows crash-survival proof * test(ci): keep crash survival gate durable * test(e2e): tolerate restart hydration navigation * test(e2e): prove exact shell input after crash * perf(ci): avoid crash harness installer rebuilds * test(ci): harden crash survival evidence and cost * test(e2e): fail closed on authoritative crash target * test(e2e): fail closed on crash liveness evidence --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |