Fix native-chat tool-run summaries so Windows-style file paths display their basename instead of the full backslash path.
Also reuse the renderer basename helper so trailing Windows paths keep a useful summary hint.
* Apply #6826: reclaim xterm focus after window blur (verbatim)
Co-authored-by: Orca <help@stably.ai>
* Improve #6826 focus reclaim: pane-scoped helper + cross-platform focus-steal guard
- Reclaim the exact split helper that owned focus at blur, not the first
helper in the container (a single TerminalPane hosts all splits as siblings).
- Defer the reclaim refocus to the next frame on every platform and only
take focus if nothing newer grabbed it, matching the existing macOS guard,
so a click into the sidebar/dialog during reactivation isn't yanked back.
- Skip reclaim if the released helper was detached before refocus.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* Harden merged #6830 daemon degraded-mode handling
- shutdownFallbackSessions is now best-effort: a single un-killable local PTY
no longer throws and aborts the daemon restart (which would leave the user's
recovery path unusable, recreating the original lockup). Logs and continues.
- checkPtySpawnHealth retries once (timeout raised 2s->4s) so a transient stall
on a busy machine right after an upgrade does not mis-classify a healthy
daemon as unable to spawn PTYs and silently drop new terminals to the local
provider without daemon persistence.
- Surface degraded mode: DegradedDaemonPtyProvider exposes isDegraded, and
pty:management:listSessions returns { degraded } so the session UI can warn
instead of it being a silent console.warn. Clearer actionable warn message.
Refs #6814.
Co-authored-by: Orca <help@stably.ai>
* Add #6814 daemon failure-mode classification test
Drives the real DaemonServer + checkDaemonHealth over a real socket to lock in
the healthy / degraded(pty-spawn-unhealthy) / wedged(unreachable) / unreachable
classification. Documents the load-bearing boundary that the degraded-daemon
fallback rescues a degraded daemon but NOT a fully wedged one.
Co-authored-by: Orca <help@stably.ai>
* Add degraded flag to web preload listSessions stub (typecheck parity)
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* fix(win): fall back to software rendering after repeated GPU crashes
On old/flaky GPU drivers the GPU child process crashes
(STATUS_BREAKPOINT / ANGLE-D3D init failure) within seconds of launch,
repeatedly — Windows crash clusters F0BDNADU79Q (Win11) and
F0BDNRZ5MDG (Win10). GPU child deaths are intentionally suppressed as
recoverable churn, so Orca never reacted and the GPU kept crashing every
launch with no self-healing path.
Track GPU child crashes in a post-launch window (3 / 30s). On the
threshold, record a gpu_fallback_engaged breadcrumb, persist a marker
file in userData, and relaunch once. The marker is read before
app.whenReady() resolves (the Store isn't available that early) so
app.disableHardwareAcceleration() + --disable-gpu take effect, then it's
consumed/cleared so the fallback applies to exactly one launch and can't
relaunch-loop. Records gpu_fallback_applied on the software-render boot.
Pure logic lives in GpuCrashFallbackTracker and the marker read/write
module, both unit-tested; the index.ts wiring is thin. Desktop only —
headless serve already uses software rendering where needed.
* fix(win): harden GPU crash fallback policy
---------
Co-authored-by: Neil <neil@stably.ai>
- Update the PR state badge icon to dynamically match merged, closed, or draft states.
- Simplify PR merge direction layout using branch tags and a directional arrow.
- Add author avatars to the PR metadata row.
- Prevent scrolled task rows from bleeding through sticky elements by adjusting header z-indexes and applying opaque hover/surface backgrounds.
- Refine padding, text size, and spacing across the PR and task views.
The oxlint 1.71 upgrade (#6841) autofixed these imports to the node:
protocol, which Metro can't resolve in a React Native bundle, breaking
the Android release build ("Unable to resolve module node:buffer").
Revert to the npm 'buffer' polyfill and disable prefer-node-protocol for
the mobile package so the autofix can't reintroduce the regression.
Co-authored-by: Orca <help@stably.ai>
Patch @xterm/addon-webgl so clearTexture() and page merge/delete operations
request a model clear and bump a generation counter. Without this, the renderer
keeps drawing glyphs against a stale atlas after the texture is reset, leaving
garbled or blank cells until the next full repaint.
Co-authored-by: Orca <help@stably.ai>
The OOM reports (F0BDMD16LJ2 and the taifunk many-worktree sessions)
show memory creeping over long sessions with churning worktrees. One
contributor: in pr-refresh-coordinator, many local worktrees that track
the same linked PR coalesce into a single queue entry whose 'aliases'
map keeps one entry per worktree. Aliases were only pruned when a
candidate was re-enqueued as invalid — never when a worktree was simply
removed/closed — so the maps grew unbounded across a session.
Add pruneWorktreePRRefreshAliases(worktreeId) and call it from
removeWorktreeMetadataAndTransientState (the existing central
worktree-removal cleanup, alongside removeWorktreeMeta /
forgetWorktree / deleteWorktreeHistoryDir). It drops the removed
worktree's aliases, deletes the queue entry when none remain, and
rebinds the representative candidate if the removed worktree owned it.
Covered by 3 new coordinator tests (accumulate-then-prune, keep-entry-
on-remaining-aliases with candidate rebind, no-op for unknown worktree)
plus two test-only inspection helpers.
Co-authored-by: Neil <neil@stably.ai>
The user report F0BDXH978GJ surfaced 'Error invoking remote method
fs:readDir' — Electron's opaque message when the main-process handler
throws. On Windows this is typically a WSL \wsl$ / \wsl.localhostUNC path or network drive failing realpath/readdir after the distro or
share goes away, or a dropped SSH provider. Nothing recorded which throw
site fired or what kind of path was involved, so the crash report
carried no actionable cause.
Wrap the fs:readDir handler to record an 'fs_readdir_error' breadcrumb
on throw, tagging the throw site (ssh-provider | authorize | readdir),
the error name/code, and a REDACTED path shape (isUNC, isWsl,
driveLetter, hasConnectionId) — never the raw path. The error is
re-thrown unchanged, so renderer behavior is identical; only crash
reports gain structured context.
Pure classification lives in readdir-error-diagnostics.ts (8 unit tests
covering WSL UNC, \wsl$, network share, drive letter, SSH, and
no-path-leak); 4 handler tests assert the breadcrumb fires per throw
site and not on success.
Co-authored-by: Neil <neil@stably.ai>
A deterministic per-load renderer fault (bad GPU driver, corrupt chunk,
AV interference) crashes on every load. Orca auto-reloads recoverable
renderer deaths after 250ms, so without a limit it reloads every
~0.25-1.3s forever — the Windows clusters F0BDRAZN55L (launch-failed
exit 18, bootstrap loop ~1.3s) and F0BDPCL93UM (ACCESS_VIOLATION on
Win10 1909, ~0.8s loop). shouldRecoverRendererAfterProcessGone only
gated integrity-failure/launch-failed, never 'crashed'/'killed'.
Add a rolling-window circuit breaker (3 recoveries / 60s) consulted at
the reload point in createMainWindow. When it opens, Orca stops
auto-reloading, records a renderer_recovery_circuit_breaker_open
breadcrumb, and shows a native Reload/Quit prompt (the renderer is dead,
so a main-process dialog is the only available surface). The window is
NOT reset on did-finish-load: a crash loop renders on every cycle before
dying, so stale attempts must age out by time instead.
Covered by renderer-recovery-circuit-breaker.test.ts (pure helper) and a
createMainWindow.test.ts integration test that drives the loop and
asserts reloads stop after the limit.
Co-authored-by: Neil <neil@stably.ai>
A synchronous throw inside any terminal link provider's provideLinks
escapes to window.onerror and gets the renderer killed (reason=killed,
exit 1). The reported crash (F0BDKBHDAUE) is xterm web-links'
LinkComputer._getWindowedLineStrings raising 'RangeError: Invalid array
length' while scanning a pathological wrapped line during agent CLI
output (opencode).
Patch terminal.registerLinkProvider at construction so every provider
registered afterward — the web-links addon's internal provider plus
Orca's file-path and terminal-handle providers — has its provideLinks
wrapped in a try/catch that records a crash breadcrumb and degrades to
'no links this hover' instead of throwing.
Covered by terminal-link-provider-guard.test.ts (reproduces the
RangeError and asserts it no longer escapes).
Co-authored-by: Neil <neil@stably.ai>
* docs: add headless Linux server guide
* docs: add ldd/appimage-extract tip for diagnosing missing libraries
Salvaged from #6817 before closing it as a duplicate.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Jinwoo-H <jinwoo@stably.ai>
Co-authored-by: Orca <help@stably.ai>
Migrate fileURLToPath(import.meta.url) / dirname(...) boilerplate to the
native import.meta.dirname / import.meta.filename, then enable the rule
at error so new code stays on the native form.
The oxlint autofix rewrites the expression but leaves the now-unused
node:url / node:path imports behind (which the already-enabled
no-unused-vars=error would then flag), so this commit also removes those
34 orphaned imports — trimming the named import where other names are
still used, deleting the line where it was the sole import.
Scope is build scripts + Node-env tests only (config/scripts, tools/
benchmarks, *.test.{ts,mjs}, vitest configs); zero shipped runtime code.
The native properties are exact equivalents (Node >= 20.11; repo is on
24), so behavior is unchanged.
Verified: oxlint 0 errors tree-wide (root + mobile), oxfmt clean,
typecheck (node+cli+web) + mobile tsc pass, root vitest 22825 passed /
0 failed, mobile vitest 1018 passed. Exercised the rewritten scripts
directly: build:relay (6 targets), ensure-native-runtime,
verify-macos-entitlements all run correctly with import.meta.dirname.
mobile/app/h/[hostId]/session/[worktreeId].tsx grew past its 4988
counted-line ratchet (now 5015) in #6659 ("Show the git primary action
on mobile Source Control") without the override being bumped. That PR's
mobile CI passed because it branched off stale main where the file was
smaller; once merged, main went red for mobile lint — but mobile.yml
only runs on PRs touching mobile/, so no main-push build caught it.
Bump the per-file override 4988 -> 5015 to match the file's actual size
(what #6659 should have done), restoring green on every mobile-touching
PR. Per AGENTS.md the file should not carry a max-lines disable; this
keeps the tight per-file ratchet. The file is a candidate for a future
split, but that's a separate, owner-driven refactor.
Verified: mobile oxlint 0 errors, root oxlint 0 errors.
* chore(lint): upgrade oxlint to 1.71 and enable 7 new rules
Upgrade oxlint 1.67.0 -> 1.71.0 (1.72 was blocked by the repo's 3-day
minimum-release-age supply-chain guard; nothing here needs it). The
bump is a no-op on the existing config.
Enable 3 error rules (backlog autofixed to zero in this commit) and
4 warn rules (surface signal without gating CI):
error (autofixed, behavior-preserving):
- unicorn/prefer-node-protocol (~1531 sites: bare builtin -> node:)
- typescript/no-import-type-side-effects (~36: all-inline-type -> import type)
- unicorn/no-array-reverse (19: copy-then-reverse -> toReversed)
warn (real signal, current fires are test-only/correct):
- unicorn/no-array-fill-with-reference-type (aliasing footgun guard)
- typescript/no-unsafe-function-type (bans bare Function type)
- unicorn/prefer-array-flat-map (map().flat() -> flatMap())
- unicorn/prefer-regexp-test (.match() in bool ctx -> .test())
mobile/.oxlintrc.json extends root, so it inherits all 7; the autofix
ran from root and covered mobile/ too.
Verification (all green): oxlint 0 errors (root+mobile+aux configs),
oxfmt clean, typecheck (node+cli+web), vitest 22795 passed / 0 failed,
builds (electron-vite + web + cli) succeed. node: rewrites confirmed to
skip embedded SSH/CLI string payloads (AST-only); all toReversed sites
verified to operate on fresh copies or write-once locals.
* chore(lint): bump mobile oxlint to 1.71 so inherited rules parse
mobile/ is a standalone pnpm project pinning its own oxlint@1.67, which
lacks unicorn/no-array-fill-with-reference-type (needs >=1.70). Since
mobile/.oxlintrc.json extends the root config, mobile CI's 'cd mobile &&
oxlint' failed to parse the new rule. Bump mobile to match root (1.71).
Verified in mobile/: oxlint 0 errors, oxfmt --check clean, tsc --noEmit
pass, vitest 978 passed / 0 failed.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
A worktree created from a PR whose branch ships an .envrc with a failing
direnv command bounced the user straight back to the Landing screen. The
worktree's only terminal spawns in the worktree cwd, direnv runs on shell
startup, fails, and the login shell exits non-zero immediately. The
single-pane PTY-exit branch routed that through onPtyExitRef ->
handlePtyExit -> closeTerminalTab, which deactivated the just-created
worktree (setActiveWorktree(null)) and rendered Landing.
The existing newborn-death guard only protected freshly-split panes, never
the first/sole terminal. Track the freshly-spawned ptyId (onPtySpawn fires
only for fresh spawns, never reattach/coldRestore) and, in the sole-pane
branch, keep the dead pane mounted when a never-interacted fresh spawn
exits. A typed 'exit' or a reattached-dead session still tears down as
before. Not gated on hasReceivedPtyOutput on purpose: a failing direnv
prints its error first, so output is received.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): scope skill-install terminals to the floating runtime selector (#6789)
Inline setup/onboarding terminals (skill installers, feature tips) create a
PTY under a synthetic per-panel worktree id. On a remote runtime the remote
PTY transport sent that id as `id:<panel>` to `terminal.create`, which the
runtime cannot resolve, so installing a skill via Settings failed with
selector_not_found. Locally it worked because the IPC spawn uses cwd directly
and never resolves the selector.
These terminals are ephemeral floating terminals with no backing worktree, so
brand their id and resolve it to the floating-terminal selector
(global-floating-terminal), which every runtime already maps to the home dir.
Local tab isolation keeps the distinct per-panel id; only the runtime terminal
selector changes. The server is unchanged.
* docs(terminal): add JSDoc for functions touched by the skill-install terminal fix
Document the new ephemeral setup-terminal id helpers, the runtime terminal
selector, and the onboarding/remote-transport entry points so the changed
functions carry contract-level docstrings.
---------
Co-authored-by: vladmesh <vladmesh@gmail.com>
Adds daemon PTY-spawn health classification, POSIX daemon cwd repair before terminal spawn, and degraded-daemon fallback routing so fresh terminals keep opening without killing preserved sessions.
Fixes#5508.
Rejects mobile-scope phone-QR pairing in the full web client, keeps runtime browser access links on the instant startup path via advisory pairing scope metadata, and surfaces forbidden runtime scope errors instead of rendering empty workspaces or retry-looping setup checks.
Security note: pairing offer scope is UI metadata only. Runtime RPC authorization remains based on the server-side device token registry and mobile allowlist.
Restyle shared Sonner toasts (wrapping, bottom-chrome clearance with mobileOffset, custom delete-failure footer) and add a dev-only Dev Tools settings pane gated behind import.meta.env.DEV. Removes dead devToolsSearch locale keys, aligns toast description opacity/width with sibling toasts, and uses a distinct nav icon for Dev Tools.