Handle expected Monaco Delayer cancellation during renderer cleanup so ignored trigger promises do not surface as unhandled rejections. Adds regression coverage for the DisposableStore teardown path from the crash report.
Keep the synthetic floating workspace local while a remote runtime is active, including terminal/browser creation, activation, close, and remote snapshot handling.
Maintainer follow-ups:
- require worktreeId for runtime-session terminal create payloads
- add renderer-backed terminal create reply sender regression coverage
- merge current main and keep the WSL readDir breadcrumb test aligned with main's Windows-only handler coverage
Fix native Windows CJK terminal repaint and ConPTY wrap metadata handling.\n\nNormalize filesystem realpath results before authorization so WSL UNC roots compare consistently in CI.
Fix native-chat tool-run summaries so Windows-style file paths display their basename instead of the full backslash path.
Also reuse the renderer basename helper so trailing Windows paths keep a useful summary hint.
* Apply #6826: reclaim xterm focus after window blur (verbatim)
Co-authored-by: Orca <help@stably.ai>
* Improve #6826 focus reclaim: pane-scoped helper + cross-platform focus-steal guard
- Reclaim the exact split helper that owned focus at blur, not the first
helper in the container (a single TerminalPane hosts all splits as siblings).
- Defer the reclaim refocus to the next frame on every platform and only
take focus if nothing newer grabbed it, matching the existing macOS guard,
so a click into the sidebar/dialog during reactivation isn't yanked back.
- Skip reclaim if the released helper was detached before refocus.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* Harden merged #6830 daemon degraded-mode handling
- shutdownFallbackSessions is now best-effort: a single un-killable local PTY
no longer throws and aborts the daemon restart (which would leave the user's
recovery path unusable, recreating the original lockup). Logs and continues.
- checkPtySpawnHealth retries once (timeout raised 2s->4s) so a transient stall
on a busy machine right after an upgrade does not mis-classify a healthy
daemon as unable to spawn PTYs and silently drop new terminals to the local
provider without daemon persistence.
- Surface degraded mode: DegradedDaemonPtyProvider exposes isDegraded, and
pty:management:listSessions returns { degraded } so the session UI can warn
instead of it being a silent console.warn. Clearer actionable warn message.
Refs #6814.
Co-authored-by: Orca <help@stably.ai>
* Add #6814 daemon failure-mode classification test
Drives the real DaemonServer + checkDaemonHealth over a real socket to lock in
the healthy / degraded(pty-spawn-unhealthy) / wedged(unreachable) / unreachable
classification. Documents the load-bearing boundary that the degraded-daemon
fallback rescues a degraded daemon but NOT a fully wedged one.
Co-authored-by: Orca <help@stably.ai>
* Add degraded flag to web preload listSessions stub (typecheck parity)
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* fix(win): fall back to software rendering after repeated GPU crashes
On old/flaky GPU drivers the GPU child process crashes
(STATUS_BREAKPOINT / ANGLE-D3D init failure) within seconds of launch,
repeatedly — Windows crash clusters F0BDNADU79Q (Win11) and
F0BDNRZ5MDG (Win10). GPU child deaths are intentionally suppressed as
recoverable churn, so Orca never reacted and the GPU kept crashing every
launch with no self-healing path.
Track GPU child crashes in a post-launch window (3 / 30s). On the
threshold, record a gpu_fallback_engaged breadcrumb, persist a marker
file in userData, and relaunch once. The marker is read before
app.whenReady() resolves (the Store isn't available that early) so
app.disableHardwareAcceleration() + --disable-gpu take effect, then it's
consumed/cleared so the fallback applies to exactly one launch and can't
relaunch-loop. Records gpu_fallback_applied on the software-render boot.
Pure logic lives in GpuCrashFallbackTracker and the marker read/write
module, both unit-tested; the index.ts wiring is thin. Desktop only —
headless serve already uses software rendering where needed.
* fix(win): harden GPU crash fallback policy
---------
Co-authored-by: Neil <neil@stably.ai>
- Update the PR state badge icon to dynamically match merged, closed, or draft states.
- Simplify PR merge direction layout using branch tags and a directional arrow.
- Add author avatars to the PR metadata row.
- Prevent scrolled task rows from bleeding through sticky elements by adjusting header z-indexes and applying opaque hover/surface backgrounds.
- Refine padding, text size, and spacing across the PR and task views.
The oxlint 1.71 upgrade (#6841) autofixed these imports to the node:
protocol, which Metro can't resolve in a React Native bundle, breaking
the Android release build ("Unable to resolve module node:buffer").
Revert to the npm 'buffer' polyfill and disable prefer-node-protocol for
the mobile package so the autofix can't reintroduce the regression.
Co-authored-by: Orca <help@stably.ai>
Patch @xterm/addon-webgl so clearTexture() and page merge/delete operations
request a model clear and bump a generation counter. Without this, the renderer
keeps drawing glyphs against a stale atlas after the texture is reset, leaving
garbled or blank cells until the next full repaint.
Co-authored-by: Orca <help@stably.ai>
The OOM reports (F0BDMD16LJ2 and the taifunk many-worktree sessions)
show memory creeping over long sessions with churning worktrees. One
contributor: in pr-refresh-coordinator, many local worktrees that track
the same linked PR coalesce into a single queue entry whose 'aliases'
map keeps one entry per worktree. Aliases were only pruned when a
candidate was re-enqueued as invalid — never when a worktree was simply
removed/closed — so the maps grew unbounded across a session.
Add pruneWorktreePRRefreshAliases(worktreeId) and call it from
removeWorktreeMetadataAndTransientState (the existing central
worktree-removal cleanup, alongside removeWorktreeMeta /
forgetWorktree / deleteWorktreeHistoryDir). It drops the removed
worktree's aliases, deletes the queue entry when none remain, and
rebinds the representative candidate if the removed worktree owned it.
Covered by 3 new coordinator tests (accumulate-then-prune, keep-entry-
on-remaining-aliases with candidate rebind, no-op for unknown worktree)
plus two test-only inspection helpers.
Co-authored-by: Neil <neil@stably.ai>
The user report F0BDXH978GJ surfaced 'Error invoking remote method
fs:readDir' — Electron's opaque message when the main-process handler
throws. On Windows this is typically a WSL \wsl$ / \wsl.localhostUNC path or network drive failing realpath/readdir after the distro or
share goes away, or a dropped SSH provider. Nothing recorded which throw
site fired or what kind of path was involved, so the crash report
carried no actionable cause.
Wrap the fs:readDir handler to record an 'fs_readdir_error' breadcrumb
on throw, tagging the throw site (ssh-provider | authorize | readdir),
the error name/code, and a REDACTED path shape (isUNC, isWsl,
driveLetter, hasConnectionId) — never the raw path. The error is
re-thrown unchanged, so renderer behavior is identical; only crash
reports gain structured context.
Pure classification lives in readdir-error-diagnostics.ts (8 unit tests
covering WSL UNC, \wsl$, network share, drive letter, SSH, and
no-path-leak); 4 handler tests assert the breadcrumb fires per throw
site and not on success.
Co-authored-by: Neil <neil@stably.ai>
A deterministic per-load renderer fault (bad GPU driver, corrupt chunk,
AV interference) crashes on every load. Orca auto-reloads recoverable
renderer deaths after 250ms, so without a limit it reloads every
~0.25-1.3s forever — the Windows clusters F0BDRAZN55L (launch-failed
exit 18, bootstrap loop ~1.3s) and F0BDPCL93UM (ACCESS_VIOLATION on
Win10 1909, ~0.8s loop). shouldRecoverRendererAfterProcessGone only
gated integrity-failure/launch-failed, never 'crashed'/'killed'.
Add a rolling-window circuit breaker (3 recoveries / 60s) consulted at
the reload point in createMainWindow. When it opens, Orca stops
auto-reloading, records a renderer_recovery_circuit_breaker_open
breadcrumb, and shows a native Reload/Quit prompt (the renderer is dead,
so a main-process dialog is the only available surface). The window is
NOT reset on did-finish-load: a crash loop renders on every cycle before
dying, so stale attempts must age out by time instead.
Covered by renderer-recovery-circuit-breaker.test.ts (pure helper) and a
createMainWindow.test.ts integration test that drives the loop and
asserts reloads stop after the limit.
Co-authored-by: Neil <neil@stably.ai>
A synchronous throw inside any terminal link provider's provideLinks
escapes to window.onerror and gets the renderer killed (reason=killed,
exit 1). The reported crash (F0BDKBHDAUE) is xterm web-links'
LinkComputer._getWindowedLineStrings raising 'RangeError: Invalid array
length' while scanning a pathological wrapped line during agent CLI
output (opencode).
Patch terminal.registerLinkProvider at construction so every provider
registered afterward — the web-links addon's internal provider plus
Orca's file-path and terminal-handle providers — has its provideLinks
wrapped in a try/catch that records a crash breadcrumb and degrades to
'no links this hover' instead of throwing.
Covered by terminal-link-provider-guard.test.ts (reproduces the
RangeError and asserts it no longer escapes).
Co-authored-by: Neil <neil@stably.ai>