mirror of
https://github.com/stablyai/orca.git
synced 2026-09-22 00:02:31 +00:00
6f7a30ac2e43db980d91265d5d83c09e05ce4dfd
322
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
339045b150 |
fix(runtime): coalesce concurrent host terminal focus (#11841)
Bound exclusive host navigation to a generation-aware latest-wins single-flight so bulk open and switch fan-out stay responsive on large remote fleets. Add freeze repro harnesses and navigated settlement. |
||
|
|
5887b36eff |
fix(updater): recover Linux .deb/.rpm installs that fail escalation (#12183)
* fix(updater): recover Linux .deb/.rpm installs that fail escalation A `.deb` install fails with `No authentication agent found` when the session has no polkit agent. Orca reported "Quit and reopen Orca, then try again" — wrong advice — and its only action was Retry Download, discarding a verified 160 MB package that was still in the updater cache. Keep the one-click install path, but make a failed root-package install recoverable without downloading again: - Retain the downloaded package and its expected SHA-512 from the `update-downloaded` event, mirroring electron-updater's cache-name rule. - Capture the child stderr that BaseUpdater logs but drops from the `error` event, redact it (ANSI, control bytes, `<home>`, `<package>`, `<user>`, 1 KiB cap), and classify the failure. Classification reads the original text — redaction can rewrite a matched phrase. - Send a structured `linux-package-install` recovery status and render a dedicated card: Copy Install Command / Try Automatic Install Again / Show Package. - Revalidate on every action: cache containment, lstat, streamed SHA-512, timingSafeEqual. Concurrent requests coalesce into one hash pass. - Build the command from fixed tokens plus one POSIX-single-quoted absolute path, resolving sudo and the package manager only from /usr/bin, /bin, /usr/sbin, /sbin. Orca never runs it. - Disable `autoInstallOnAppQuit` for .deb/.rpm so an ordinary quit cannot trigger the same failing escalation after the UI is gone. Extracts the error-card presentation into UpdateErrorCardContent so UpdateCard does not absorb another stateful surface. Lifecycle breadcrumbs carry package type, reason, exit code and version — never a path, command, username or raw child output. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Improve Linux package install recovery diagnostics - Distinguish invalid-package-path errors from missing package manager - Expand ANSI escape sequence stripping to handle OSC hyperlinks and DCS - Prevent generic error logs from overwriting specific diagnostic verdicts - Add error handling for shell.openUrl in update UI - Fix test isolation with proper afterEach hooks --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8c5371ebad |
fix(worktrees): respect Windows shell for setup runners (#6967)
* Honor configured shells during worktree setup
* Align setup launch paths with selected Windows shells
* Carry setup shell selection through deferred launches
* Prove Windows setup shell routing at its real adapters
* Ground remote PowerShell proof in the real writer
* Preserve Git Bash across deferred setup launches
* Harden Windows setup runner shell selection
- Resolve remote PowerShell binary without local pwsh probe: for SSH/remote
Windows worktrees, isPwshAvailable() reflects only the LOCAL host, so an
'auto' implementation could route the remote runner to a pwsh.exe the remote
lacks. Add resolveSetupRunnerShell(..., { probeLocalPwsh: false }) so remote
auto keeps the always-present powershell.exe; explicit pwsh.exe still honored.
- Preserve native exit codes in the PowerShell runner by checking
$LASTEXITCODE before $?, so a failing native command surfaces its real code
instead of a generic exit 1; $? still catches cmdlet soft-failures.
- Write the PowerShell runner with a UTF-8 BOM so Windows PowerShell 5.1 (the
new default powershell.exe) reads it as UTF-8 instead of ANSI, preventing
non-ASCII setup-script corruption.
- Add unit tests for the remote-probe behavior.
* Restore setup-shell scope narrowing over the rebase
The force-pushed rebase dropped five review-fix commits that were already
on this branch; this reapplies their combined effect on top of the new
base and the hardening commit:
- Keep SSH setup shell selection remote-owned (no local terminalWindowsShell
or pwsh routing for remote hosts; supersedes the probeLocalPwsh guard)
- Preserve cmd setup compatibility outside POSIX shells (no .ps1 runner
family, so the BOM/exit-code hardening is no longer applicable)
- Route WSL setup runners from the project runtime
- Avoid blocking PowerShell probes during setup creation
- Correct SSH and WSL background setup fixtures
* Satisfy the changed-code gates for the setup-shell runner
- createWorktreeRunnerScript took 7 positional parameters, tripping the
changed-code max-params gate; move it to a single options object.
- hooks-runner.test.ts deep-equals the createSetupRunnerScript result, so
assert the cmd shell now returned for native Windows worktrees.
* Carry the setup launch shell through observed and issue runners
- buildObservedSetupCommand takes the runner's launch shell so WSL-routed
Windows-drive setup replays use /mnt/c instead of Git Bash /c
- resolveSetupRunnerShell gates the posix runner on the same Git Bash
resolution the PTY uses, so a missing or non-MSYS bash keeps the cmd runner
- issue-command runners carry their launch shell, and the renderer passes it
when building the queued command
- treat a bare `bash` shell setting as POSIX like `bash.exe`
Co-authored-by: Orca <help@stably.ai>
* fix(worktrees): close counsel P1 gaps for Windows setup shells
Route windowless/headless creates through the shell-aware setup runner when a
PTY controller is available, existence-check explicit Git Bash paths before
committing to .sh runners, thread the resolved shell into issue-command
runners, and document the intentional Git Bash interpreter flip with a narrow
scope table.
* Convert setup env to MSYS form and harden the bare cmd runner launch
C3: a Git Bash setup runner now receives ORCA_*/CONDUCTOR_*/GHOSTX_* path
values in /c/... form, matching the runner path and the shell's own HOME/PWD.
C5: extension-less `bash` resolves to Git Bash everywhere, matching how
resolveWindowsShellStartupFamily already classifies it.
C7: runner paths carrying characters that cannot be quoted on a cmd command
line launch through a delayed-expansion PowerShell shim instead, and the batch
runner disables inherited delayed expansion so `!` in setup lines survives.
Co-authored-by: Orca <help@stably.ai>
* docs: note MSYS ORCA_* paths and bare bash Git Bash resolution
Keep the setup-shell release note aligned with C3 env conversion and C5 bare
bash resolution so the published claim matches runtime behavior.
* revert: drop windows-setup-shell doc allowlist and AGENTS link
Keep the counsel P1/P2 product fixes without expanding the docs allowlist
or AGENTS.md guidance surface.
* fix(plugins): contain Parcel unsubscribe rejections under Vitest
Dev plugin watchers fire-and-forget unsubscribe, and in-process Parcel
can reject when temp watch roots are already deleted. Catch those
rejections so they cannot fail the suite as unhandled errors.
* fix(plugins): keep in-process unsubscribe rejection surface
Swallowing Parcel unsubscribe errors broke mocked unsubscribe tests
that return non-Promises and expect rejections. Contain failures only
in PluginDevWatcher fire-and-forget paths.
---------
Co-authored-by: OrcaWin <alpha-eng@stably.ai>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Co-authored-by: Orca <help@stably.ai>
|
||
|
|
98ae8e4c8c |
Allow clearing all agents from AI Vault session history filter (#12128)
* Allow clearing all agents from AI Vault session history filter Add "Select all" / "Clear" buttons so users can quickly isolate one agent without unchecking each box individually. Previously, at least one agent had to remain enabled; now users can filter to zero agents and re-enable selectively. * Address PR #12128 review feedback - Make Select all / Clear real DropdownMenuItems so Radix roving focus reaches them by keyboard. - Rename the zero-agent empty state to a neutral "No agents selected" now that zero agents is a valid filter. - Use 모두 해제 for the Korean Clear label instead of 지우기 (erase). Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
484273844a |
feat(updater): add an adhoc release channel for branch builds (#12051)
* feat(updater): add an adhoc release channel for branch builds Hourly covers main. This covers everything that is not main yet: a dispatchable macOS build of an unlanded branch, published to stablyai/orca-adhoc, so the team can run an experimental feature for a few days instead of reasoning about it from a diff. Adhoc sits at the bottom of the version order — 'adhoc' < 'hourly' < 'rc' < stable — so no routine check can walk anyone onto somebody's branch; only an explicit pinned jump reaches one. It gets its own repo rather than sharing orca-hourly's, because a branch build must not appear in the list a developer riding main is looking at. Signed and notarized exactly like hourly, for the same reason: macOS anchors a notarized app's TCC grants on identifier + team, so an unnotarized build reads as a new client and silently loses file access under Documents/Desktop/Downloads. Tags stamp to the second rather than the minute. Hourly runs under a concurrency group and cannot overlap itself; adhoc builds are dispatched on demand, so two people cutting from different branches inside one minute is ordinary — and a minute-resolution tag would collide and fail the second build after its whole pack-and-notarize run. Channel-specific behaviour now derives from one DEDICATED_REPO_CHANNELS list: repo mapping, macOS-only support, and UpdateSource. The RPC schema that validates releaseChannelOverride was a hand-copied enum missing the new channel, which would have rejected the override on its way to the main process; it reads the predicate now. * fix(updater): merge the duplicated shared/types import Co-authored-by: Orca <help@stably.ai> * fix(ci): default the adhoc build ref to the dispatch branch The Actions UI puts its own "Use workflow from" branch picker directly above the ref field, and picking a branch there is what most people read as "build this". Making the field optional means the obvious action is also the correct one; naming a branch explicitly still wins, so main's copy of the workflow runs rather than a stale one on an old branch. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
5e9186f5ff |
chore(mobile): import the shared marine-creature corpus instead of mirroring it (#12090)
mobile/src/constants/marine-creatures.ts was a hand-maintained copy of src/shared/marine-creatures.ts, identical except for a comment header. The copy existed because Metro only watched mobile/ and could not resolve repo-root modules; mobile/metro.config.js:11 added src/shared to watchFolders five weeks later, and ~195 mobile files already import from src/shared. The renderer collapsed its copy to a re-export at the same time; mobile was the leftover. Point the one consumer at the shared corpus and delete the mirror, the bespoke regex-scraping parity test that policed it, and the now-stale max-lines baseline entry. No behavior change: same exported symbol, byte-identical name list. |
||
|
|
2b44e9ed9e |
fix(updater): notarize hourly macOS builds so TCC grants survive updates (#12007)
macOS anchors a notarized Developer ID app's TCC grants on identifier + team, which is cdhash-independent and so survives an in-place update. Without a notarization ticket there is no such stable identity, so every hourly reads as a different client: the grant row stays but stops matching, and file access under Documents/Desktop/Downloads fails with EPERM and no re-prompt. `tccutil reset` fixes it until the next build — and orca-hourly has shipped as many as 14 builds in a day. Skipping notarization was chosen because Squirrel.Mac validates the replacement bundle's signature, not its notarization. That is true, but it is the wrong requirement; the in-place swap was never the problem. Budgets grow to absorb the notary round trip (publish 2x45, job 150), and the App token is re-minted after the build so its one-hour life starts at the first call that uses it rather than during `pnpm install`. |
||
|
|
edb5607e28 |
ci: block new root-level entries (#11903)
* ci: guard repository root additions * fix: clear existing type-aware lint warnings |
||
|
|
ad1e58d966 |
chore: declutter top-level repo layout (#11890)
Remove one-off incident docs and committed test-results noise, move dev/repro/bench tools under tests/tools, and relocate i18next config into config/ so the GitHub root scrolls to the description faster. |
||
|
|
676964b099 | ci: run only changed e2e specs on pull requests (#11834) | ||
|
|
f56e6ade80 |
fix(ssh): recover orphaned relay install locks (#9828) (#10207)
* fix(ssh): recover orphaned relay install locks (#9828) * test(ssh): split staged upload relay specs (#9828) * fix(ssh): verify staged relay upload namespace * fix(ssh): bound stale relay stage cleanup * fix(ssh): complete bounded stage recovery * fix(ssh): generate valid PowerShell stage scripts * fix(ssh): make staged upload cancellation safe * fix(ssh): fence staged relay recovery * test(ssh): align deploy timeout oracle --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
4205d5f31f |
Refine macOS access prompt guidance (#11822)
* Refine macOS access prompt guidance * Capitalize MacOS in access prompt * Use official macOS styling in prompt |
||
|
|
cd2b62ed14 |
feat(updater): name hourly releases by version, build number, time, and sha (#11817)
* feat(updater): name hourly releases by version, build number, time, and sha Hourly releases were titled with their raw tag (`v1.4.163-hourly.202607312054`), which reads as one opaque digit run and does not say which commit it came from. Title them `1.4.163 • 01 • 07-31 13:54 • e698241` instead, and show that same string in the in-app build picker by having the picker render the release's stored name rather than deriving its own label. Composing it in one place means the two surfaces cannot drift. The build number is monotonic across the channel. It is read as the highest number already in use rather than as a count of releases: the prune step trims to 72, so a count would roll backwards after three days and reissue numbers. Drafts count toward it — unlike in the freshness check, which asks whether a commit shipped, this asks whether a number is free, and a stranded draft still holds one. Times are Pacific while the tag's stamp stays UTC. The stamp is a sort key and a local one would repeat an hour at every DST fall-back, making two distinct builds compare equal; the title is only ever read. * fix(updater): fail the hourly build when the release name is missing The workflow checks out `ref: main`, but a workflow_dispatch runs the workflow file from whatever branch was dispatched. A branch that edits this step while main still carries the old script produces an empty name and an untitled release — silent, and only visible once someone opens the releases page. Verified by hitting exactly that on run 30665586904. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
79251d7a98 |
[P2] fix(release,settings): restore signing preflight portability, bootstrap diagnostics, and skill re-check (#11692)
* fix(release): restore the SignPath composite action when cutting from an older ref Co-authored-by: Orca <help@stably.ai> * fix(startup): record a durable diagnostic before the bootstrap fatal-exit guard exits Co-authored-by: Orca <help@stably.ai> * fix(settings): make agent-skill Re-check rescan skill freshness Co-authored-by: Orca <help@stably.ai> * fix(startup): keep the bootstrap fatal diagnostic when the log override is unwritable Create the parent directory an overridden ORCA_BOOTSTRAP_FATAL_LOG names and fall back to the default location when that path still cannot be opened, so a missing parent no longer costs the only account of the failure. Also pins the Re-check freshness rescan to the completed install scan rather than the click. Co-authored-by: Orca <help@stably.ai> * refactor(settings): move the post-recheck surface sync out of the panel Co-authored-by: Orca <help@stably.ai> * fix(startup): retain diagnostics without node fs * fix(skills): keep freshness scoped to the local runtime * fix(settings): register freshness status translations * fix(settings): scope and sequence skill freshness refreshes * fix(settings): refresh freshness across runtime transitions --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
fdb58695e9 |
[P1] fix(checks): stop skipped and manual checks reporting as failures (#11700)
* fix(checks): stop skipped and manual checks reporting as failures Route every check-classification surface through one shared helper so desktop renderer, desktop main and mobile agree on the same verdict. - GitLab `manual` jobs and pipelines are neutral again, not action_required/failure - `skipped` counts as passed everywhere, including mobile - a neutral check no longer demotes a summary that has passing checks * fix(checks): move the check-classification parity test into the renderer project The parity table lived in src/shared but imported a renderer module, and both config/tsconfig.node.json and config/tsconfig.cli.json are composite projects that include src/shared without that renderer path, so `pnpm typecheck` failed with TS6307 on two of its three projects. Only the web project spans both trees. Co-authored-by: Orca <help@stably.ai> * fix(checks): stop the Tasks-grid pill contradicting its own verdict The checks pill's label, tone and icon all read one ProviderCheckSummary, but getChecksLabel short-circuited on the raw `neutral` counter while the tone and icon key off `state`. After the classification fix a PR with 19 success + 1 neutral renders an emerald CheckCircle2 pill that reads "1 unresolved", and mobile's own label (which keys off `state`) reads "19/20 passed" for the same summary. Move the label into src/shared/provider-check-summary.ts so desktop and mobile cannot fork it again, and key it off `state`. Also covers deriveWorkItemCheckSummary, the desktop-main producer of the summary that reaches the Tasks grid and the relay-paired mobile client. It was rewritten here with no test at all; the parity table stands in derivePRCheckStatusFromRollup, which is a different normalizer. The new main-process test drives getWorkItem with a real statusCheckRollup fixture, pinning the StatusContext `state` fallback that would otherwise be deletable with the whole suite still green. Co-authored-by: Orca <help@stably.ai> * fix(gitlab): route the pipeline job-array rollup through the shared check classifier The array path in derivePipelineStatus kept its own copy of the rollup rules, so manual-only read green and one unrecognized job status demoted a passing pipeline to neutral — both disagreeing with every other check surface. Also retry the packaged-CLI smoke temp cleanup on Windows: the copied Orca.exe can still be locked by AV/indexers after every assertion passed, failing the package job. Co-authored-by: Orca <help@stably.ai> * fix(gitlab): stop the skipped pipeline string diverging from the Checks tab - classifyPipelineString now counts a skipped pipeline as passing, matching the per-check classifier; canceled stays neutral and is pinned as an explicit, sign-off-pending divergence. - Pin the production string path (head_pipeline.status) in the parity table and note that the job-array branch has no production caller yet. - Count skipped checks in the Checks panel's passing header so it agrees with the checks pill. - Correct the packaged-CLI smoke retry comment: the EBUSY is the smoke's own just-exited Electron process, not AV/indexers. Co-authored-by: Orca <help@stably.ai> * fix(checks): finish cross-surface check parity and back out the skipped MR-card flip Review follow-ups on the check-classification PR. - PullRequestPage and GitHubItemDialog kept private copies of getCheckCounts / getChecksSummaryLabel that still counted only `success` as passing, so a 2-success/3-skipped PR read "2 passing · 3 skipped" there and "5 passing" in the sidebar. Both copies move to pr-check-counts.ts, which routes the passing bucket through classifyCheckOutcome; action_required keeps its own amber bucket. The summary icon now keys off passing count, so an all-neutral PR stops painting a green tick above "0 of N checks passing". - The sidebar checks header and triage strip still called `{status: completed, conclusion: null}` pending, contradicting the grey "Unresolved checks" pill. Both now read summarizeProviderChecks and render an unresolved chip/strip instead of an amber spinner that can never resolve. - classifyPipelineString('skipped') is reverted to neutral. That flip painted MR cards green for pipelines that never ran, on the only GitLab path with production callers, and contradicted the same function's deferral of `canceled`. Both tone changes stay deferred, pinned by one test. - classifyPipelineString('manual') resolves to pending rather than neutral: a blocked pipeline is outstanding, and neutral let the worktree card fall through to its emerald `open` default while GitLab still refuses the merge. - TaskPage's checks pill helpers move to task-page-checks-pill.ts so the "1 unresolved on a green pill" fix is actually pinned by a test. - smoke-packaged-cli no longer lets an EBUSY cleanup replace the real failure. * fix(checks): stop completed unknown checks from spinning --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
e467b3ff7b |
fix(remote): stabilize shared control and terminal parking (#11656)
* fix(remote): stabilize shared control and terminal parking * fix(remote): harden parking review edge cases * fix(terminal): restore parked local floating buffer * fix(ci): drop superseded paired parking evidence * fix(terminal): preserve floating park watchers * fix(ci): include web client in paired e2e artifact * fix(ci): reuse renderer build for paired e2e |
||
|
|
f998f7ec62 |
feat(updater): add hourly dev channel and build switching (#11250)
* feat(updater): add hourly dev channel and build switching Adds an hourly macOS build channel plus a dev-only surface for switching update channels and jumping to any published build, including older ones. Hourly builds publish to a separate stablyai/orca-hourly repo. The routine update path resolves tags from the main repo's releases atom feed, which exposes only its 10 newest entries — 24 hourly tags a day would evict every stable/RC entry there and leave real users with nothing to update to. Hourly artifacts carry the release bundle id and Developer ID signature so Squirrel.Mac can swap them in place; only notarization is skipped, which in-place updates never check. Version tails are stripped to the base (1.4.160-hourly.<stamp>, not 1.4.160-rc.3-hourly.<stamp>) so hourlies sort below both rc.N and stable and are reachable only by an explicit pinned jump, never by an ordinary check. The picker is revealed by Option-clicking the Updates header, matching the Help menu's existing hidden admin affordance. Pinned jumps set allowDowngrade and release the feed on every settle path so a jump can never leave background checks permanently deferred. * chore(hourly): create orca-hourly and add token provisioning script Adds setup-hourly-release-token.sh, which provisions HOURLY_RELEASE_TOKEN without the value ever reaching stdout, argv, or shell history: it is read with `read -rs`, passed to gh through GH_TOKEN in the environment rather than as an argument (argv is world-readable via ps), piped into `gh secret set` on stdin, and scrubbed by an EXIT trap. Verification creates and deletes a draft release in orca-hourly to prove Contents:write for real rather than trusting the permission checkbox. Drafts are absent from the releases atom feed, so the probe cannot disturb users. Refuses to run without a controlling terminal instead of falling through having set nothing, and refuses to run under xtrace, which would echo the token on every expansion. * fix(updater): address review feedback on the hourly channel Renderer: - Guard listBuilds against out-of-order responses. activeChannel flips once getVersion resolves, and rapid channel clicks stack requests, so a slower earlier load could land last and fill the list with builds from a channel the picker was no longer showing. - Selecting the running build's own channel now clears the override instead of pinning it. There was previously no way back to "follow this build's channel", so merely opening the panel left background checks pinned. - Validate releaseChannelOverride on hydration, matching every other enum-like field in that function. Main: - Exclude pinned jumps from recordCompletedUpdateCheck() in update-available. A dev browsing the picker was persisting lastUpdateCheckAt and suppressing the next real background check for a full day. - parseHourlyVersionStamp now anchors on the whole version and round-trips the parsed fields. It accepted garbage prefixes, and Date.UTC rolled impossible dates forward, so ...hourly.202602300000 rendered as March 2. Workflow: - Publish into a draft and flip it live only after the manifest check. The window between creating the release and verifying its assets previously exposed a tag the picker would offer and the download would 404 on; a draft is invisible to listReleaseBuilds, so a job that dies in that window — including a hard kill by the job timeout, which runs no cleanup step — leaves nothing user-visible behind. - Add a failure handler that discards the draft, gated on the publish step not having succeeded so a later prune failure cannot delete a live release. - Align retry budgets with the job timeout (was 60min against a worst case of ~185min, so a mid-retry kill skipped the cleanup that step exists for). - Exclude drafts from the freshness and retention queries. - persist-credentials: false; the job only reads this repo and never pushes. * refactor(hourly): authenticate with a GitHub App instead of a PAT A fine-grained PAT expires, and the hourly build would then fail silently on a schedule nobody watches. A GitHub App's private key has no expiry, so this is set up once. It is also owned by the org rather than by the person who created it, so the credential survives that person leaving. The workflow mints a short-lived installation token via actions/create-github-app-token and passes it as GH_TOKEN. Installation tokens live one hour, which is ample: this job runs no tests, no notarization, and no Windows signing, so it is pack + upload. The retry budgets and job timeout are re-sized to that reality rather than copied from the release pipeline, whose 3x45 publish budget exists for notarization and SignPath. setup-hourly-release-token.sh now provisions HOURLY_RELEASE_APP_ID and HOURLY_RELEASE_APP_PRIVATE_KEY. The key is redirected from a file straight into `gh secret set` on stdin, so its contents never enter a shell variable, argv, or the terminal. * fix(hourly): make the xtrace guard fire and cover cancelled runs The xtrace guard disabled tracing before testing for it, so `[[ -o xtrace ]]` read the state the previous line had just cleared and never fired. `bash -x` ran straight through, tracing exactly the key handling the guard exists to prevent. Test first, then disable. The draft cleanup only ran on failure(), but a run stopped from the Actions UI is cancelled(), not failed — a manual cancel mid-publish stranded the draft. Cover both. |
||
|
|
cc078a5021 |
perf(main): move hang watchdog into a worker thread (#11488)
* perf(main): add watchdog boundary memory benchmark Add a repeatable Electron 43 RSS harness that measures the production-built watchdog entry across the child-process and worker-thread boundaries. Record per-trial samples, the median, revision, runtime, and settling procedure for reproducible PR evidence. * perf(main): move hang watchdog into a worker thread Keep main-thread hang detection independent of the blocked Electron event loop without paying for a second ELECTRON_RUN_AS_NODE process. Preserve the marker and telemetry contract while moving timing configuration and heartbeats onto a bundled worker entry. * test(main): smoke packaged hang watchdog worker * fix(main): make packaged watchdog smoke able to fail The smoke reported failure only through process.exitCode, but its finally block quit Electron gracefully, and Electron takes its status from the browser exit code. Every failure mode — entry missing from app.asar, worker error, marker timeout, non-zero worker exit — exited 0 with the diagnostic discarded on stderr, so the required PR check could never go red. Propagate a real status via app.exit, assert the success line in stdout, and surface stderr. Verified against a packaged tree with the entry removed: exit 0 before, exit 1 after. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
14de3fa14d |
fix(computer): reap mac helper after client loss (#11493)
* perf(computer): add mac helper owner-loss benchmark Measure the release helper's resident memory before and after its owner-session deadline. Record exact revisions, per-trial RSS, retained state, and clean-exit latency so lifecycle reclamation is reproducible. * fix(computer): reap mac helper after client loss Bind the detached macOS helper lifetime to authenticated socket ownership. Reap the helper after its final authenticated client disconnects, and add a startup deadline for sessions that never authenticate. * test(computer): harden owner benchmark cleanup * test(computer): make owner benchmark cleanup failure-safe * test(computer): close remaining owner cleanup races |
||
|
|
8f7692aa12 |
Fix packaged skills CLI runtime ownership (#11627)
* fix(cli): make packaged skills runtime self-contained * fix(cli): address packaged skills review feedback * ci(cli): smoke packaged skills on Windows --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
a906f98baf |
fix(release): survive PSGallery outages in the Windows signing preflight
The Windows release job hard-failed in run 30125672117: every SignPath module install attempt got 403 Forbidden from the gallery's OData API, which is behind Azure Front Door and was also serving 502/504 at the time. That step was the only hard-fail in an otherwise fail-open signing chain, so a gallery incident blocked the whole release. The gallery CDN that serves the nupkg is a separate origin and stayed healthy throughout, so fall back to a pinned version fetched from it after the normal install path is exhausted. The fallback verifies a SHA-256 pin, since that route skips the gallery's own package validation. Extracted to a composite action so the release job and the signing rehearsal cannot drift apart. |
||
|
|
b370dc0900 | ci(release-cut): include source ref/commit and cutter in SignPath Slack (#10524) | ||
|
|
d73c61364f | fix(win): harden startup during partial updates (#11613) | ||
|
|
49cfbf014c |
fix(skills): stop OS sidecars marking an untouched skill as modified (#11471)
* fix(skills): stop OS sidecars marking an untouched skill as modified Package identity compared a live user directory against a tree read from a clean checkout, so anything the OS deposited counted as drift. One Finder visit writes .DS_Store, which sorts before SKILL.md and misaligns the index-aligned snapshot comparison — the copy became 'unrecognized', was reported as "may be modified... Remove it", and left out of the update. Running the update could not clear it either: the updater compares its lock to the source and never reads disk, so it correctly reports "up to date" and writes nothing. Ignore OS-authored names on both sides of the comparison. The generator half is not hypothetical: a stray sidecar in a working tree made the committed artifacts read as stale, failing lint for that developer. Scoped to OS-authored names only. Tolerating unexpected files in general would let an injected payload ride along beside a clean SKILL.md; these are safe because an official SKILL.md never references them, so no agent can be routed into one. Mode bits are deliberately untouched — that would weaken identity for real scripts. * fix(skills): keep guarding a directory or link wearing an OS metadata name The name-only skip dropped any entry matching an OS metadata name, so a directory named .DS_Store or ._scripts took its whole subtree out of identity and a symlink wearing one stopped tripping the link guard — a skill hiding either read as pristine. The OS writes these as plain files only, so the entry type decides, still ahead of the case-fold map. Also compares both walkers over the same fixture: an asymmetric skip is worse than none, since one side would bake in content the other can never observe. * chore: ignore the OS metadata names skill identity already skips Both skill-identity walkers ignore these names, but .gitignore covered only .DS_Store and Thumbs.db — so a stray ._SKILL.md showed as untracked and `git add -A` could commit it. That is the one way the two walkers can disagree: the disk walker skips such a file while the git-tree producer (collectGitPackageFiles, used by the unreferenced --rebuild-from-tags path) does not, so a committed sidecar would make released history and observation describe different content. Ignoring them keeps that asymmetry unreachable rather than adding a second skip to the released-history path, which is load-bearing and provably never sees one today: no committed sidecar exists on any ref. Nothing tracked matches the new patterns. * chore: correct the skill-identity ignore comment The previous wording claimed these names cannot be committed, which overstates what .gitignore provides: `git add -f` and `git apply --index` both bypass it, so a cherry-pick, rebase or fork branch already carrying a sidecar is unaffected. That clause was load-bearing — it was the stated reason for leaving the released-history producer unhardened — so it should not read as a structural guarantee. Also fixes the producer count (three, not two: two disk walkers plus the git-tree producer, which does not skip) and says plain file, since the skip is isFile()-gated so a directory or link wearing the name is still walked. |
||
|
|
650dd48ec9 |
feat(cli): add orca account add / account list for headless hosts (Claude + Codex) (#9177)
* feat(cli): add `orca account add` / `account list` for headless hosts The desktop "Add account" UI is disabled when the renderer drives a remote runtime (isRemoteAccountScope === kind:'environment'), so a headless server reached from a remote desktop/web client has no way to register managed Claude accounts. Add a host-local CLI path that reuses the existing capture logic: - ClaudeAccountService.addAccountFromConfigDir(): register a managed account by capturing credentials from an already-authenticated CLAUDE_CONFIG_DIR instead of spawning the interactive browser login (extracted persist/rollback helpers shared with the existing add flow) - RPC accounts.addClaudeFromConfigDir, bridged via OrcaRuntime; rejected for mobile device tokens (host-local only) - `orca account add` runs `claude login` in the user's own terminal into a temp CLAUDE_CONFIG_DIR, then registers it via the local runtime; `orca account list` lists managed accounts Switching (select) already works from a remote client; only adding was blocked. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(cli): support Codex in `orca account add` / `account list` Mirror the Claude headless-account CLI for Codex: - CodexAccountService.addAccountFromHome(): register a managed Codex account by importing auth.json from an already-authenticated CODEX_HOME, reusing a shared persist helper extracted from doAddAccount (no interactive login spawned here) - RPC accounts.addCodexFromHome + OrcaRuntime.addCodexAccountFromHome bridge, rejected for mobile device tokens (host-local only) - `orca account add --agent claude|codex` (default claude); `orca account list` now renders both Claude and Codex managed-account blocks Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: cover headless account-add capture paths (Claude + Codex) - ClaudeAccountService.addAccountFromConfigDir: registers a managed account by capturing an authenticated CLAUDE_CONFIG_DIR; rejects and rolls back when the dir has no .credentials.json - CodexAccountService.addAccountFromHome: imports auth.json from an authenticated CODEX_HOME into a managed account; rejects when auth.json is missing Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: address CodeRabbit review on headless account-add flows - CLI login spawn uses a shell on Windows so `.cmd` agent shims resolve without ENOENT (args are fixed literals, no injection risk) - Claude capture skips the `.credentials.json` precheck on macOS, where creds live in the Keychain and captureAuthFromConfigDir reads them - Claude add rollback is best-effort: a failed rematerialization no longer skips managed-auth cleanup or masks the original add error - Codex persist restores the prior account/selection if a post-write sync or rate-limit refresh fails, so a failure can't leave a dangling managed account - Codex sync passes the account's selection target (correct runtime for WSL) - Add JSDoc to the new public service methods and CLI functions Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(cli): harden headless account capture * fix(cli): correct account command flag surface and interrupt cleanup - `account` commands no longer accept or advertise the browser `--page` flag; `supportsBrowserPageFlag` allow-listed them by omission, so `orca account list --page x` was silently accepted and `--help` rendered a browser-only option - account specs declare GLOBAL_FLAGS, so `--help`/`--json` render in the Options block like every other command - `--agent` on `account add` documents the account provider instead of the terminal TUI-agent meaning inherited from the shared flag table - a SIGINT/SIGTERM during the interactive login now removes the temp login dir (and restores the macOS Keychain item) before exiting 130; Node terminates without unwinding `finally`, which stranded live OAuth credentials on disk * perf(cli): stop `account list` forcing a provider usage refresh `accounts.list` awaited refreshAccountsForMobile(), which runs fetchAll({ force: true }) — bypassing both the poll throttle and the per-provider Retry-After gate — then O(N) serial per-account round trips. `orca account list` renders only emails and the active ids, so all of that work was discarded. The RPC now takes `refreshUsage` (default true, so mobile and web keep the forced lane) and the CLI opts out. Older hosts declare `params: null` and ignore the field, so a newer CLI degrades to the previous behavior rather than failing. Also documents on `account list` that `--environment` does not retarget it, matching the host-local behavior of shouldIgnoreRemoteSelection. * fix(cli): survive repeated and hangup signals during account add withInterruptCleanup latched cleanup behind a boolean, so a second signal got an already-resolved promise and its process.exit fired while the first cleanup was still inside a Keychain call (3s each) — the temp dir's OAuth credentials and the swapped macOS Keychain item both survived. Memoize the cleanup promise so every signal awaits the same run, and register with `on` instead of `once` so a second Ctrl-C cannot fall through to Node's terminate-immediately default mid-cleanup. Handle SIGHUP too. This flow exists for headless/SSH hosts, where the most likely interrupt is the connection dropping, which hangs up the login's terminal and previously ran no cleanup at all. Warn when the interrupt lands after sign-in completed: the runtime finishes the add independently of this process, so exiting 130 silently would tell the user it was cancelled when the account may exist. Reject a valueless `--agent`; the parser turns it into boolean true, which silently ran a full OAuth login for Claude when the user asked for another provider. Also lock two behaviors the refactor changed but left uncovered: a WSL Codex add must sync the WSL runtime lane rather than the default host lane, and rename the account-spec help test to describe the Options block it actually asserts rather than the usage string it never reads. * fix(build): bundle the main modules the account CLI imports electron-vite cleans out/main and emits only its declared entries, and `build:desktop` runs it after `build:cli`, so the tsc-emitted copies of `claude-accounts/keychain`, `codex-cli/command` and `win32-utils` were deleted before packaging. Both `orca account add` and `orca account list` then died at require time with "Cannot find module '../../main/claude-accounts/keychain'" — reproduced against a real `--serve` host. `agent-hooks/managed-agent-hook-controls` already carried an entry for exactly this reason; these three were missing. Adds a parity test so any future CLI import of a `src/main` module fails in CI rather than at a user's shell after packaging. * test: cover the desktop add-path behavior this PR changes Both changes ride in the persist/rollback helpers the existing GUI add flow shares with the new headless path, and neither had coverage: - Claude: rollbackAddAccount now guards forceMaterializeCurrentSelection- ForRollback, so a rejecting rematerialization no longer replaces the real add error nor skips safeRemoveManagedAuth. Asserts the original error surfaces and the throwaway auth dir is gone. - Codex: the desktop add now passes the account's selection target to syncForCurrentSelection, matching reauthenticate and select. Asserts the host target alongside the existing WSL assertion. Both fail when the corresponding change is reverted. * fix(cli): close the remaining account-add interrupt and preflight gaps The round-1 interrupt fix detached the signal handlers before running the finally-path cleanup, so the very window it was meant to protect — the two serial 3s `security` calls plus rmSync on the success/error path — was still covered only by Node's terminate-immediately default. Both review lanes reproduced it independently. Await cleanup first, detach in a nested finally, and stop a cleanup failure from replacing the error that actually explains why the add failed. Do not burn the interactive login when the runtime is unreachable. The RuntimeClient is lazily constructed and the first call was the registration RPC itself, so "Requires the Orca runtime to be running" was discovered only after the user completed a full OAuth round trip. Preflight with the now-cheap `accounts.list { refreshUsage: false }`. Reject `--environment` / `--pairing-code` on `account add`. shouldIgnoreRemoteSelection pins account commands to the local runtime, so `orca account add --environment homelab` silently registered the account on the laptop instead of the headless host it names. Survive a daemon that cannot spawn `claude`. `allowFailure` is honored in onClose but not onError, and unlike the GUI flow nothing has run `claude` in the daemon before this point — so a launchd/systemd daemon with a minimal PATH hard-failed an add the user had already signed in for, even though identity resolves fine from the config dir's oauthAccount. Also align the `--agent` help description with the global flag column. * fix(cli): reject runtime selectors on `account list` too `orca account list --environment homelab` was accepted and silently listed the LOCAL machine's accounts, because shouldIgnoreRemoteSelection pins account commands to the local runtime. Documenting that in --help does not reach someone who already typed the flag, and answering with the wrong host's accounts is the specific wrong answer they would act on. `account add` already errors; this makes the new command group internally consistent. The other groups in shouldIgnoreRemoteSelection keep their existing silent-ignore behavior — changing those is not this PR's job. * test: harden account-add signal tests and cover cleanup failure - Identify the handler under test by set difference instead of `process.listeners(sig).at(-1)`. Vitest installs its own once-wrapped SIGINT teardown, so the positional lookup could grab the wrong listener; the helper also asserts exactly one new listener was added. - Mock rmSync while keeping the real implementation by default, so the temp-dir assertions elsewhere stay honest. - Cover that a cleanup failure in the `finally` does not replace the error explaining why the add failed. Fails when that guard is removed. Completes the review loop's final round; the loop died on an API error before it could commit this, and its `import()` type annotation would have failed oxlint. * fix(cli): harden interactive account add * test(cli): make account cancellation coverage portable * fix(cli): preserve merged skills runtime modules --------- Co-authored-by: Dominik <marketing@gavaplast.sk> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
8ad9448905 | revert: restore pre-worker process boundaries (#11481) | ||
|
|
1f2f809a11 | fix(computer): bind macOS helper to supervised peer pid (#11475) | ||
|
|
d0f341ad69 |
fix(computer-use): make modifier clicks interruption-safe (#11451)
* fix(computer-use): make modifier clicks interruption-safe * fix(computer-use): pace modified Windows multiclicks * fix(computer-use): address modifier safety review |
||
|
|
5e00a30e4e |
Decouple feature copy from locale parity (#8512)
* Decouple feature copy from locale parity * Fix undeclared dynamic localization key check * Fix localization code owner |
||
|
|
b339fe0346 |
Fix Node 26 test gate and happy-dom storage (#11434)
* ci: test PR shards on Node 26 * test: isolate happy-dom storage from Node globals |
||
|
|
fe6f929c6e |
fix(terminal): reconcile cross-platform IME composition lifecycle (#11293)
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: JeongUk Park <jeongph.dev@gmail.com> |
||
|
|
3f37e32e72 |
perf(main): move hang watchdog into a worker thread (#11344)
Keep main-thread hang detection independent of the blocked Electron event loop while reducing watchdog memory from 47.1 MiB to 11.5 MiB. Preserve marker, recovery, and telemetry behavior with a bundled worker-thread entry. |
||
|
|
363e478909 |
fix(orchestration): preserve active workers across updates (#11271)
* fix(orchestration): preserve active workers across updates * test(ssh): model absent legacy adoption * test(orchestration): align compatibility contracts * fix(windows): escape updater PowerShell booleans * fix(windows): restore stock uninstall process check * fix(orchestration): keep recovery off renderer startup barrier * fix(orchestration): harden legacy recovery migration * fix(orchestration): close recovery review gaps * fix(orchestration): complete legacy worker cutover recovery * fix(orchestration): preserve legacy workers across updates --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
a7c8b8e071 |
fix(terminal): bound SSH & remote hidden-worktree terminal retention (C1) (#10625)
* fix(terminal): park SSH worktrees like local ones (C1 retention, slice A) SSH ptys were blanket-excluded from hidden-view parking, so a hidden SSH worktree retained every pane forever (C1: renderer heap climbs to the V8 ceiling). SSH bytes transit local main — fact-mode watchers already cover them, and main keeps a headless model served over pty:getMainBufferSnapshot that the SSH reattach path never consulted. - isParkRestorableTerminalPty: snapshot-backed OR (SSH + policy); threaded through both park verdicts, both selectors, watcher coverage, and the watcher start guard. Remote-runtime/fail-open/foreign/null unchanged. - Parked-SSH reveal paints from main's headless model (dimension-matched, ~5k rows) and degrades to the relay 100KiB replay unless the snapshot is a non-empty source==='headless' payload — never a blank/stale paint. - Kill switch: settings.terminalSshViewParking (default on). DESIGN.md records the approved plan and the H1 magnitude non-claim. Co-authored-by: Orca <help@stably.ai> * fix(terminal): bound hidden-worktree retention with a force-park budget (C1, slice B) Un-parkable worktrees (remote-runtime ptys, uncoverable tabs, SSH with the slice-A switch off) had unlimited retention: the parking cap/TTL only ever saw eligibility-passing worktrees, so one bad tab pinned a whole worktree's panes forever. Retention is now memory-bounded, not eligibility-bounded. - terminal-hidden-worktree-retention.ts: retention budget (12 hidden / 45min TTL, sized from the measured 2.5-19MB per-pane V8 cost, DESIGN.md §2) over hidden worktrees ordinary parking can never evict; reuses the hot-retain ranking so last-active exemption, deterministic ties, and deadline-driven rechecks hold. Fail-open/foreign-pty tabs are eviction-exempt (a remount would fresh-spawn and orphan the live shell). - Terminal.tsx: force-parked ids join the parked set AFTER the coverage veto (darkness for uncoverable tabs is the accepted cost); buffers captured via the sleep-flow registry before the unmount render; retention TTL added to the recheck deadlines for budget candidates only. - Verdict stays out of its own effect deps; policy test asserts idempotence and time-monotone membership (flip-loop dwell regression). - Kill switch: settings.terminalHiddenWorktreeRetentionBudget (default on). Co-authored-by: Orca <help@stably.ai> * fix(terminal): demote hidden scrollback for eviction-exempt worktrees (C1, slice C) The retention budget (slice B) must exempt worktrees holding fail-open or foreign-worktree ptys — a remount would fresh-spawn and orphan the live shell — which would leave that class unbounded again. Instead, past the same 45min retention TTL their hidden panes drop to the minimum scrollback tier (measured: ~19MB -> ~1.3MB V8 heap per 50k-row pane; trimmed history is gone by design, reveal restores the configured cap for future output). - terminal-hidden-scrollback-demotion.ts: module-state verdict registry (parked-watcher pattern) with content-equality notify damping; applied in the existing scrollback-rows effect in use-terminal-pane-lifecycle. - selectScrollbackDemotedTerminalWorktrees: pure, TTL-gated, time-monotone. - Retention TTL wakeups now also cover exempt worktrees so demotion fires. - Kill switch: settings.terminalHiddenScrollbackDemotion (default on). Co-authored-by: Orca <help@stably.ai> * fix(terminal): paint the SSH model snapshot inline, not via nested coordinator (C1 slice A fix) applyMainBufferSnapshot runs its own structuralReplayCoordinator.run; calling it from applyReattachPayload (already inside the coordinator when a relay replay exists) deadlocks on the coordinator's tail chain. The model paint now mirrors the daemon-snapshot branch inline (folded scrollback + rehydrate + screen, dimension-matched, escape tail last) and arms the restored-snapshot seq baseline so deferred/live chunks the snapshot covers dedupe instead of double-painting. Also falls through (no early return) so reattachPayloadApplied still latches. Adds the folder-workspace id parity unit case. Co-authored-by: Orca <help@stably.ai> * test(terminal): SSH park+reveal e2e round-trip + as-built design notes (C1) Docker-gated (ORCA_E2E_SSH_DOCKER=1) spec: SSH tab parks behind a decoy and reveal restores marker content at multi-viewport scrollback depth. DESIGN.md records the as-built deltas (inline paint, force-park shape, last-active floor) and the residuals so follow-ups aren't lost. Co-authored-by: Orca <help@stably.ai> * fix(terminal): paint SSH reveal from main's model even when the relay replay is empty (C1 review #1) A relay restart empties the replay buffer; the reveal previously painted nothing even when main's headless model held the session. The reattach now prefetches the model snapshot when no structural replay exists (SSH-shaped ptys only) and paints it inside the coordinator; emptiness is judged on the composed payload (scrollbackAnsi + data + pendingEscapeTailAnsi) so an alt-screen snapshot with an empty screen frame still paints. Co-authored-by: Orca <help@stably.ai> * fix(terminal): decouple scrollback demotion (slice C) from the retention-budget switch (C1 review #2) Per the approved contract each slice reverts behind its own switch: slice C now requires only the master terminalHiddenViewParking plus its own terminalHiddenScrollbackDemotion flag. The TTL wakeup timer fires for demotion candidates even with the budget switch off. No DEFAULT_SETTINGS entries exist for sibling flags (defaults are the '!== false' optional pattern), so no explicit defaults are added. Co-authored-by: Orca <help@stably.ai> * fix(terminal): scope eviction exemption to the tab, not the worktree (C1 review #3) One eviction-exempt tab (fail-open/foreign pty) previously vetoed force-park for its whole worktree, pinning co-located remote-runtime tabs forever. The worktree now force-parks while exempt tabs keep their mounted panes via a per-tab exclusion mirroring the Activity-portal pattern (legacy watcher sync, legacy render, and the overlay cold-parking hook). Ordinary parking is untouched — a worktree with an exempt tab still cannot ordinary-park. Slice C now also demotes exempt tabs' panes as soon as their worktree force-parks under the count budget (they are the only panes left mounted). Co-authored-by: Orca <help@stably.ai> * fix(terminal): demote un-parkable worktrees the force-park lever spared (C1 review #4) The last-active exemption means a single hidden un-parkable worktree never force-parks — and slice C previously only targeted exempt-tab worktrees, so its panes held full scrollback forever. Demotion now also covers un-parkable non-exempt worktrees past the retention TTL that are absent from the force-parked set (last-active spared, or slice B switched off). Membership stays time-monotone for fixed inputs; covered by new idempotence/monotone selector tests. Co-authored-by: Orca <help@stably.ai> * fix(terminal): keep the hidden clock running through transient background-measure windows (C1 review #5) Whole-worktree background mounts (browser-automation bootstrap lease, mobile mounts, agent wakes) open a ~3s self-clearing measure window that previously deleted hiddenSince — every remount restarted the 30s hysteresis and the 45min retention TTL, so a periodically re-mounted force-parked worktree never re-parked. The measure window still pauses parking/eviction verdicts (all selectors skip measuring candidates); only the clock survives, so the prior verdict resumes as soon as the window closes. Visible and portal-holding worktrees still reset the clock. Co-authored-by: Orca <help@stably.ai> * test(terminal): make the SSH park+reveal depth assertion prove the model paint (C1 review #6a) Pad the session with ~180KB of output after the numbered markers so the earliest marker falls outside the relay's 100KiB rolling replay buffer while staying inside main's ~5k-row headless model; asserting marker_1 after reveal now proves the headless-model paint rather than passing under the relay fallback. Co-authored-by: Orca <help@stably.ai> * docs(terminal): rewrite DESIGN.md as the single as-built C1 contract (review #7) One contract matching the code: status IMPLEMENTED around force-park (not the unmount proposal), real kill-switch names with coupling + revert matrices, the true retention-floor formula with measured per-pane and demotion numbers, an explicit when-OOM-is-still-possible paragraph naming the H2 pendingSideEffects residual, the applyMainBufferSnapshot deadlock constraint inside the slice-A section, stable-signal phrasing instead of a capability latch, fail-open AND foreign-worktree exemption class, verified cites, and a planned/landed/follow-up test matrix. Co-authored-by: Orca <help@stably.ai> * fix(terminal): resolve the eviction exemption per pane, not per tab (C1 review #8) isEvictionExemptTerminalTab read only tab.ptyId — the FIRST leaf's pty — while the coverage veto that makes a worktree a retention candidate walks every pane. A split tab whose second leaf held an unrestorable pty therefore failed coverage (→ force-park target) yet looked exempt-free, so force-park unmounted it and orphaned the live shell. The exemption now resolves panes through the same resolveParkedTerminalPaneCandidates, keeping tab.ptyId in the union for the no-layout/no-capture case. Also from the same review round: - force-park's capture passes includeLocalBuffers:false like every other shutdownBufferCaptures caller; it was serializing up to 512KB/pane of scrollback into the store inside a fix meant to bound renderer heap. - Terminal.tsx unmount resets the scrollback-demotion registry — module state with no reset path, read by a pane effect that runs before the host effect that would clear it, so a stale verdict trimmed restore replays. - memoize watcher coverage per tab within the parking pass; the retention candidates re-asked it for every mounted worktree, not just the parked few. * docs(terminal): drop DESIGN.md — the as-built C1 contract moves to the PR body Co-authored-by: Orca <help@stably.ai> * fix(terminal): cap the deferred PTY side-effect queue (C1 residual H2) pendingSideEffects grew without bound under background timer throttling (~64 drained/s vs hundreds queued/s overnight). Cap at 512 entries with oldest-first eviction: titles drop (last-wins), a pending bell latches onto the next survivor, agent-status payloads collapse onto the survivor keeping the newest 16 (last-wins store state, KB-scale strings). Co-authored-by: Orca <help@stably.ai> * fix(terminal): carry command-lifecycle facts through parked watchers (C1 follow-up) Parked fact-mode watchers omitted onCommandFinished/onCommandCode*, so OSC 133;D and Command Code scrape signals went dark while parked. New parked-terminal-command-status.ts ports the store-level subset: git-UI nudge on every command finish, same-turn status-row drop for SSH PTYs (exact mounted-path parity — the foreground tracker refuses SSH ids), and the Command Code working seed / 1500ms done settle. Byte mode scans the same shared parsers for authority-off parity. Local-PTY status drops stay with the mounted pane: they need pty-connection's process-confirm ladder to tell a leaked nested-shell 133;D from a real agent exit. Co-authored-by: Orca <help@stably.ai> * test(terminal): retention-budget force-park e2e with a retentionLimit override (C1 6b) ORCA_E2E_TERMINAL_RETENTION_LIMIT flows preload → e2e-config → getTerminalParkingPolicyOverrides (exposeStore-gated, positive-integer only) so a spec can shrink the force-park budget to 1. The Docker-gated spec opens two remote worktrees on one relay target (second pre-seeded remote repo), disables terminalSshViewParking to make both un-parkable, hides both behind the local context, and proves the older one force-parks while the last-active exemption spares the newest; re-activating the evicted worktree restores the marker tail via relay replay. Co-authored-by: Orca <help@stably.ai> * test(terminal): retention-budget e2e via same-repo remote worktrees (passes docker lane) The first draft added a second remote repo mid-session, whose pane pty spawn misroutes to the local daemon with the remote cwd (pre-existing multi-repo issue, reproducible without any retention override — a seeded local repo plus one remote repo shows the same misroute). The spec now budgets across three worktrees of the ONE connected repo, created through the product createWorktree path (an external git-worktree-add only lands as a detected worktree needing adoption) and polled through the relay's transient post-connect reconnect window. Verified green on the local Docker lane in 20.8s. Co-authored-by: Orca <help@stably.ai> * fix(terminal): prevent remount thrashing during post-measure cool-down ( Implements the C1 retention contract: preserve worktree `hiddenSinceMs` through a background-measure window (so TTL/ranking stay honest), but re-park waits for a full `coldParkDelayMs` cool-down after the measure ends. Without the cool-down, every ~3s measure lease on a past-deadline worktree thrashes remount/reattach. Core changes: - Terminal.tsx: add measure clock (measuringTerminalWorktreeIdsRef) and post-measure cool-down tracking (terminalWorktreeParkCooldownUntilRef); gate parking candidates until cool-down expires. - Extract snapshot replay choreography to shared terminal-snapshot-replay-paint.ts (used by SSH reattach + daemon restore paths). - Add SSH model snapshot timeout (750ms) with fallback to relay replay. - Move cold-park recheck deadline logic to terminal-cold-park-recheck-deadlines.ts; add cool-down deadline to scheduling. - useTerminalTabColdParking: implement matching measure-clock contract with per-tab cool-down gate to keep tab deadlines synced with worktree retention clock. - Add resolveTerminalMountScrollbackRows() to demote new xterms under demoted worktrees (pane births during demotion must take the demoted tier at create). - Add kill switches: terminalSshViewParking, terminalHiddenWorktreeRetentionBudget, terminalHiddenScrollbackDemotion. * fix(terminal): detect Command Code completion in parked mid-turn panes Seed the byte watcher with in-flight turn state from agent status: the watcher is recreated per park cycle with no startup command to arm it, and the banner scrolled away before parking. Also memoize eviction-exempt checks and use SSH PTY ID builder in tests. * fix(terminal): flush pending command-code settles on reveal remount When a parked pane reveals mid-Command Code turn, the new detector cannot re-observe the already-passed idle composer. Cancelling the settle leaves the row stranded at 'working', so dispose now flushes the pending settle instead. Extract readInFlightCommandCodeTurn to shared space and seed detectors with in-flight turns so remounts complete mid-flight commands. Also memoize SSH model probes to prevent double timeouts on reattach. * fix(terminal): remove scrollback demotion (C1 slice C) The scrollback demotion feature for eviction-exempt hidden worktrees is no longer needed. Retention budget limits are now sufficient without this additional bound. Remove the terminal-hidden-scrollback-demotion module, the selectScrollbackDemotedTerminalWorktrees function, and related per-pane demotion logic. * test(terminal): assert bounded probe during stalled reveal Add assertion to verify that a stalled reveal operation makes exactly one `getMainBufferSnapshot` call, ensuring retry logic doesn't introduce redundant probes that would extend the timeout window before relay fallback. * fix(terminal): implement C1 retention budget for hidden parked worktrees Addresses OOM regressions in hidden parked terminals by force-evicting worktrees past a retention budget: at most 12 mounted while hidden, none past 45 minutes (absolute, not exempted by last-active). Eviction is least-recently-hidden-first. Exempt tabs (unrestorable local PTYs) keep their panes to avoid orphaning shells; worktrees are force-parked even if they contain exempts, and their buffers released elsewhere. SSH/remote worktrees serialize buffers pre-eviction for reveal; local worktrees keep daemon snapshots. Command Code's done-settle window is transferred across park/reveal boundaries so the row cannot strand at 'working'. Model probe on SSH reattach is scoped to park-reveal only, not ordinary reconnects. Includes new E2E suite proving the budget actually releases memory. * memoize eviction-exempt terminal tabs to avoid redundant store reads Each tab's exemption check re-reads the store and walks the layout tree. Introduce selectEvictionExemptTerminalTabIds() to resolve all exempt tabs for a worktree in a single pass, then memoize the result in Terminal.tsx and useTerminalTabColdParking. This prevents O(n) store reads when checking exemptions across multiple tabs and ensures the set remains stable across unrelated re-renders. * refactor: reformat hidden-worktree retention comments Reflow to 80-character lines and remove internal ticket references (C1, C1 slice C). * fix(lint): split overlay slot and eviction-exempt tabs under max-lines Static analysis failed because TerminalPaneOverlayLayer (401) and terminal-parked-tab-watchers (304) exceeded oxlint max-lines. Extract the slot component and eviction-exempt helpers into dedicated modules. * test(terminal): stabilize retention budget e2e control arm Stage un-parkable remote pty ids only after both worktrees are hidden, and keep re-staging during the control-arm poll so a late updateTabPtyId cannot flip the decoy back to park-restorable and ordinary-park it before budget engages. * test(terminal): pin retention e2e decoy to a mounted pane snapshot Use the active pane-identity snapshot for the decoy tab instead of all worktree tabs, and re-assert un-parkable ids after the control-arm hold so a deferred/empty tab id cannot fail the budget-off mounted-count check. * fix: memoize terminal eviction exemptions on layout leaf PTYs Splits add leaf panes to the layout store without changing the tabs array. A memo keyed only on tabs misses this change, leaving new panes unexempted for unmount. Include layout leaf PTYs in the exemption memo key so it recalculates when splits occur or PTYs are re-minted. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
747b241145 |
feat(main): record main-thread hangs so we can measure them (#10256)
A deadlocked main thread never crashes, so it leaves no crash report and no artifact — incidence has been unmeasurable (n=1 confirmed, macOS 26.5.1, FB24004458 / electron#52437). This forks a plain-Node watchdog sibling under ELECTRON_RUN_AS_NODE that survives the deadlock, listens for a 2s heartbeat, and after 45s of silence writes a marker to userData. The next launch consumes it, records a durable crash breadcrumb, and emits a main_thread_hang_detected telemetry event carrying unresponsive_ms and self_recovered. Observes only — it never kills or relaunches the parent. A true positive recovers nothing force-quitting wouldn't, while a false positive would SIGKILL a live main thread mid-write. self_recovered counts exactly the stalls such a killer would have gotten wrong, so recovery can be built on evidence if the field numbers justify it. macOS-only, packaged-only (ORCA_HANG_WATCHDOG_FORCE=1 to test), with sleep-gap suppression and idempotent shutdown on will-quit. |
||
|
|
3c0cd6069f |
fix(release): stop packaging plugin authoring examples into app.asar (#11087)
* fix(release): stop packaging plugin authoring examples into app.asar electron-builder's `files` is an all-negation list, so its default `**/*` packs anything without an explicit `!` entry. examples/ arrived with the plugin system in #8549 and never got one, so 1.4.160-rc.3 shipped examples/plugins/hostile-panel/panel.html — the adversarial fixture the panel containment tests point at, complete with its fetch-exfiltration probe — plus hello-orca, inside every user's app.asar. Verified against the installed 1.4.160-rc.3 artifact, not just the config. The two orchestration design docs landed at the repo root in the same span and shipped the same way; fold them into the existing root-doc negation. Neither has a runtime consumer: bundled plugins ship via extraResources from resources/plugins/launch/, which is already excluded from the asar for exactly this reason. * test(release): assert the examples exclusion through the real file matcher The added case mapped each negation to a bare top-level token, so it passed under '!examples/README.md' — a pattern that still ships the whole tree. Drive app-builder-lib's FileMatcher instead so the assertion matches the test name, and pin the root anchoring so the negation cannot grow into '!**/examples'. |
||
|
|
a8126a0a92 |
fix(macos): explain the TCC prompts, and surface Full Disk Access only to users macOS is prompting (#9756) (#9910)
* fix(macos): add a Full Disk Access nudge to reduce recurring TCC prompts (#9756) macOS shows the "Orca wants to access other apps' data" (kTCCServiceSystemPolicyAppData) prompt and it can keep reappearing. The reappearing loop is not a fixable app bug: it is TCC identity churn — an unsigned local rebuild mints a new code identity each build, so macOS treats each as a new app — and Orca's other-app reads are already gated behind opt-in settings or explicit user actions. The durable remedy for the population we can help (release users) is Full Disk Access, a superset macOS grant that stops these prompts for a stable identity. Surface it with an ambient, dismissable sidebar card that reuses the existing developer-permissions IPC. macOS-only; probes FDA status at most once per renderer session (the probe itself reads protected data, so it must not repeat on focus/remount); "Open System Settings" opens the Full Disk Access pane; permanent localStorage dismissal. * fix(macos): stop the FDA nudge promising macOS will stop asking The card said Full Disk Access makes "macOS stop asking", but the grant covers this app while terminals are spawned by the detached PTY daemon (daemon-init.ts forks execPath with ELECTRON_RUN_AS_NODE + detached:true, reparented to launchd), which macOS treats as its own TCC identity. A user who followed the card would grant FDA and still be prompted from terminals. Scope the claim to reducing prompts and name the terminal caveat. * fix(macos): drop stale focus refreshes in the FDA nudge refreshFullDiskAccessStatus() applied whichever getStatus() round-trip resolved last. Rapid blur/focus puts several in flight, so an earlier pre-grant 'unknown' landing after a newer 'granted' un-hid the card and also wrote 'unknown' into the module-level session cache, re-nagging a user who already has Full Disk Access for the rest of the session. The adjacent FullDiskAccessSetupPrompt already guards this with a refresh sequence; mirror it here. Also unmount React roots in afterEach: clearing document.body left them mounted, leaking each test's window focus listener into later tests. * test(macos): unmount the StrictMode FDA nudge root between tests The afterEach unmount added in |
||
|
|
1fa9ffb5ea |
ci(pr): run E2E when a PR touches tests/e2e paths (advisory) (#11131)
* ci(pr): run E2E when a PR touches tests/e2e paths Regression specs under tests/e2e never ran on PR CI — only schedule and release called e2e.yml — so a red regression test could merge green. Path-filter and workflow_call the E2E suite when E2E-relevant files change. Use merge-base diffs so base-branch drift does not false-trigger E2E, fail the detector when git diff cannot compute the PR range, and pin least-privilege contents:read on both the detector and reusable E2E workflow. Closes #10518 Co-authored-by: Wooseong Kim <innocarpe@gmail.com> Co-authored-by: Orca <help@stably.ai> * ci(pr): make the E2E path gate actually block, and match the real config path Two fixes to the new path-filtered E2E job. The gate did not gate. pr.yml's `verify` job is the required check, and it enumerates its dependencies explicitly — `e2e` was in neither `needs` nor the result list, so a failing shard left `verify` green. That reproduces the exact hole this job exists to close: a red spec merges green, just with a red box further down the page. Add `e2e` to both. Because the job is path-filtered, `skipped` is the normal result on a PR that touches no E2E files and has to keep passing. That allowance is checked after the strict loop rather than inside it, so it can never leak to the six jobs that are always required. The `playwright.` pattern matched nothing. The config is tests/playwright.config.ts — beside tests/e2e/, not inside it — so no tracked file starts with `playwright.` and editing the runner config would silently skip E2E. Anchor it at `tests/playwright.`. Adds a contract test alongside the existing release-e2e one. Verified it fails when either fix is reverted, and simulated the gate across success/skipped/failure/cancelled plus the skip-must-not-mask-a-real-failure case. * test(ci): close two gaps in the E2E gate contract CodeRabbit was right on both counts — verified by reverting each and watching the contract stay green. The path filter was unasserted, so `e2e` could lose its `if:` and run on every PR — the cost the filter exists to avoid — without failing anything. The strict-loop check hardcoded four of the six required jobs, so dropping GIT_COMPATIBILITY or SHELL_CONTRACTS left them unenforced while the contract passed. Derive the list from verify.needs instead, so a newly added required job that misses the loop fails here rather than silently going unchecked. * ci(pr): land the E2E path gate advisory instead of blocking The E2E suite is currently failing every scheduled run on main — 22 of the last 22 — so making verify depend on it would block any PR touching tests/e2e/**, including the PRs that fix the suite. This PR's own run reproduced that: 3 of 12 shards failed on specs unrelated to it (agent-session resume, Jira linking, plugin containment, terminal artifacts). So the job runs and reports on E2E-path PRs but is left out of verify.needs for now. The detector, the tests/playwright. path fix, and the contract tests are unaffected — those stand on their own and were the substance of the review. Flipping to blocking is a three-line change once the suite is green; the exact wiring, including why the skipped allowance must sit outside the strict loop, is recorded on verify's Require-successful-checks step. The contract test pins the advisory choice so it reads as deliberate rather than as the unwired-gate bug it originally caught, and still fails if the path filter, the strict-loop coverage, or the config path regress. --------- Co-authored-by: Wooseong Kim <innocarpe@gmail.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
89f32a121b | fix(lint): restore nested config discovery in changed-code gate (#11130) | ||
|
|
77d4c64f7a |
Improve orchestration migration safety for live legacy workers (#11107)
* fix(orchestration): clarify legacy migration safety * fix(cli): sanitize legacy formatted messages * test(runtime): allow near-cap fuzz under shard load --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
e551d3ec0d |
perf(lint): consolidate code-quality gates into Oxlint (#11117)
Consolidate standalone code-quality scanners into Oxlint, preserve focused native/type-aware enforcement, add custom plugin coverage, and harden deferred PTY test cleanup. |
||
|
|
c25d85cc4c |
perf(terminal): eliminate adverse control and frame-gate cases (#11045)
* perf(terminal): eliminate dense control and frame gate regressions * test(terminal): keep gate labels in valid expect shape * test(terminal): expose the surviving sub-threshold control-density case The only adverse strip fixture sat at 50% control density, which is exactly where the fallback fires and wins. A shape at 31 controls per 64-unit block evades the trigger and still loses to the per-character legacy (0.67x), so the benchmark structurally could not show it. Add that fixture, pin both density literals in the staleness guard so a retune fails loudly instead of silently measuring a boundary that moved, and export the probe constant the equivalence test was hardcoding. |
||
|
|
badf91101b |
fix(quality): enforce performance-safe lint baseline (#11074)
* fix(quality): clear safe existing lint findings * fix(quality): keep lint cleanup allocation-free * fix(quality): enforce performance-safe baseline * test(terminal): drain deferred confirmation cleanup |
||
|
|
12ef12c55b |
chore(quality): ratchet Oxlint, React Doctor, and Zustand performance (#11034)
* chore(quality): ratchet lint and Zustand performance * fix(ci): stabilize React peer lock snapshot * fix(ci): isolate PR diff and React Doctor CLI |
||
|
|
c6076a507c |
ci(release): detach non-blocking E2E (#11031)
* ci(release): detach non-blocking E2E * test(release): pin E2E dispatch retries |
||
|
|
3d98cda5b2 | fix(release): accept lowered telemetry declarations (#11019) | ||
|
|
0edc95fa35 |
perf(editor): cut per-keystroke work on two rich-markdown paths (#10862)
* perf(editor): cut per-keystroke work on two rich-markdown paths Doc links: both plugins walked every text node and ran matchAll on each — the auto-convert appendTransaction once per keystroke, the preview decorations once per keystroke and again per caret move. A link needs `[[`, so gate on a native substring check first. The two walks had duplicated their guard sequence; they now share one predicate. 3.1x-3.8x over the repo's own markdown. Annotations: resolving a comment's block re-serializes the whole document (every node, plus every adjacent pair), and both the highlight-range and comment-at-position paths did that once per comment — O(comments x document). Build the blocks once and pass them down. On a 12-node fixture with 8 comments that is 184 serializations down to 23. * test(editor): pin the one-build serialization baseline Review feedback, all four points: - The serialize-count assertions compared many-comments against one-comment, so they would have passed if BOTH built blocks twice. Pin the absolute count (23 = 12 nodes + 11 adjacent pairs) derived from the fixture size, so a regression to per-comment building fails instead of comparing equal. Verified by reverting the hoist: 2 tests fail. - Skip an empty benchmark corpus instead of evaluating `index % 0` and dereferencing undefined. - Build fixture paths with path.join. - Condense the benchmark header to purpose plus parity guarantee. Co-authored-by: Orca <help@stably.ai> * test(editor): harden doc-link performance evidence Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
a49d68f8c2 |
perf(git): overlap getBranchCompare's head-of-chain reads (#10895)
* perf(git): overlap getBranchCompare's head-of-chain reads
Four git spawns ran strictly in series before any compare work began:
branch --show-current, the base-ref probe, rev-parse HEAD, and rev-parse <base>.
Three are independent -- compareRef is display-only metadata and HEAD's oid does
not depend on the base ref -- so they now run concurrently. The fourth was
redundant outright: the probe already runs `rev-parse --verify --quiet
<ref>^{commit}` and discarded the oid it printed, which was then re-resolved by a
second spawn. resolveWorktreeBaseCommitOid returns that oid so it can be reused;
hasWorktreeBaseCommitRef now delegates to it, leaving its other 4 callers
untouched.
3.6-3.7x on a short remote base label (192ms -> 52ms), 1.44x on an
already-qualified refs/... base, which skips the probe by design.
Reuse is keyed by ref: resolveWorktreeAddBaseRef returns at its first successful
candidate, so only that ref's oid is ever read back. Peeling is safe because only
refs/heads and refs/remotes candidates reach the probe, where ^{commit} is a
no-op.
No new git features: this removes a spawn rather than adopting an option.
Co-authored-by: Orca <help@stably.ai>
* fix(git): preserve compare semantics across providers
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
|
||
|
|
6677b5f171 |
perf(cli): construct the runtime client only when a command needs it (#10919)
src/cli/index.ts was the only eager value-import of RuntimeClient, and five other eager modules imported just RuntimeClientError / RuntimeRpcFailureError from the runtime-client barrel -- dragging in client -> pairing -> zod -> ws -> e2ee on every invocation. Those error classes live in runtime/types.ts, which has zero children, so the five imports now point there and the client loads through the existing (already lazy by design) ctx.client getter. Eager modules 199 -> 46, with node_modules dropping 94 -> 0. `orca --help` 2.04x (59.6 -> 29.2 ms); the same for help, no-args, and both error paths, which return before constructing a client. Commands that DO construct one still gain 1.10-1.12x from not eagerly parsing the transport the local path never uses. Correction to an earlier note: websocket-transport alone is ~24 modules / ~8 ms, not the 107 / 28 ms once recorded -- that figure wrongly charged it for zod, which enters through shared/pairing on a different edge. Marginal cost, never isolated cost. Co-authored-by: Orca <help@stably.ai> |
||
|
|
077561f89a |
perf(terminal): measure stream byte length natively above a code-unit floor (#10916)
* perf(terminal): measure stream byte length natively above a code-unit floor The terminal RPC path counted UTF-8 bytes with a hand-rolled per-code-point scan that Buffer.byteLength does natively an order of magnitude faster. Routed through a small module rather than swapping the shared clipboard helper, which has ~50 renderer call sites and a partial-count contract on the over-limit path that must not change. 4.4x on an 8KiB batcher push, 4.2x on a 2MiB snapshot scan, 4.0x on the 48KiB chunk gate, and 1.26-1.33x on the adversarial early-trip shapes where the legacy scan bails after a third of the string. The floor is load-bearing, not defensive. Buffer.byteLength has a fixed ~14ns call cost against a scan iteration of ~1.5ns, so below the measured 8-12 code unit crossover the native call is a REGRESSION -- 4.2x slower at one code unit, which is keystroke echo, the most latency-sensitive PTY shape there is. Short inputs keep the scan verbatim; 16 leaves margin over the crossover so the worst sub-floor shape stays at parity. measureTerminalStreamByteLength takes the native count only when `length * 3 <= stopAfterBytes` proves the limit cannot trip, so the callers' truncated running total is never replaced by a full count. Co-authored-by: Orca <help@stably.ai> * test(terminal): benchmark production byte-length exports Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |