mirror of
https://github.com/stablyai/orca.git
synced 2026-09-22 08:02:28 +00:00
13c193a00a3ecccc4d4a589cd95363bfa92a53ec
184
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a40183389b |
feat: bound direct SSH reconnect fan-out and recovery (#11003)
* docs: design for direct SSH reconnect fan-out Capture the implementation-ready plan for host-qualified, epoch-fenced SSH reconnect recovery after two rounds of multi-model LLM counsel review. * docs: reconcile SSH reconnect fan-out design * docs: close reconnect design consistency gaps * feat: implement bounded direct SSH reconnect recovery * fix: bound direct SSH retry settlement * fix: harden direct SSH reconnect authority * fix: preserve split SSH retry ownership * fix: preserve SSH split continuation authority * docs: record final SSH reconnect validation * fix: preserve SSH authority through retained and detached state * fix: retain SSH authority across delayed split mounts * fix: close SSH authority recovery gaps * fix: fence stale SSH transport replacement * fix: serialize SSH target teardown * fix: settle SSH teardown failures before reconnect * fix: retire failed SSH reset sessions * test: reconcile current main E2E contracts * fix: close direct SSH reconnect review gaps * fix: fence stale SSH reconnect side effects * fix: close final SSH reconnect lifecycle gaps * test: stabilize current-main reliability gates * test: prove plugin navigation containment * test: make plugin navigation oracle authoritative * test: make plugin navigation oracle deterministic * ci: allow sharded e2e suite to finish * test: wait for runtime pane publication * test: classify pane readiness by error code * test: select close persistence terminal by tab identity * docs: mark reconnect implementation validated --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
1fa9ffb5ea |
ci(pr): run E2E when a PR touches tests/e2e paths (advisory) (#11131)
* ci(pr): run E2E when a PR touches tests/e2e paths Regression specs under tests/e2e never ran on PR CI — only schedule and release called e2e.yml — so a red regression test could merge green. Path-filter and workflow_call the E2E suite when E2E-relevant files change. Use merge-base diffs so base-branch drift does not false-trigger E2E, fail the detector when git diff cannot compute the PR range, and pin least-privilege contents:read on both the detector and reusable E2E workflow. Closes #10518 Co-authored-by: Wooseong Kim <innocarpe@gmail.com> Co-authored-by: Orca <help@stably.ai> * ci(pr): make the E2E path gate actually block, and match the real config path Two fixes to the new path-filtered E2E job. The gate did not gate. pr.yml's `verify` job is the required check, and it enumerates its dependencies explicitly — `e2e` was in neither `needs` nor the result list, so a failing shard left `verify` green. That reproduces the exact hole this job exists to close: a red spec merges green, just with a red box further down the page. Add `e2e` to both. Because the job is path-filtered, `skipped` is the normal result on a PR that touches no E2E files and has to keep passing. That allowance is checked after the strict loop rather than inside it, so it can never leak to the six jobs that are always required. The `playwright.` pattern matched nothing. The config is tests/playwright.config.ts — beside tests/e2e/, not inside it — so no tracked file starts with `playwright.` and editing the runner config would silently skip E2E. Anchor it at `tests/playwright.`. Adds a contract test alongside the existing release-e2e one. Verified it fails when either fix is reverted, and simulated the gate across success/skipped/failure/cancelled plus the skip-must-not-mask-a-real-failure case. * test(ci): close two gaps in the E2E gate contract CodeRabbit was right on both counts — verified by reverting each and watching the contract stay green. The path filter was unasserted, so `e2e` could lose its `if:` and run on every PR — the cost the filter exists to avoid — without failing anything. The strict-loop check hardcoded four of the six required jobs, so dropping GIT_COMPATIBILITY or SHELL_CONTRACTS left them unenforced while the contract passed. Derive the list from verify.needs instead, so a newly added required job that misses the loop fails here rather than silently going unchecked. * ci(pr): land the E2E path gate advisory instead of blocking The E2E suite is currently failing every scheduled run on main — 22 of the last 22 — so making verify depend on it would block any PR touching tests/e2e/**, including the PRs that fix the suite. This PR's own run reproduced that: 3 of 12 shards failed on specs unrelated to it (agent-session resume, Jira linking, plugin containment, terminal artifacts). So the job runs and reports on E2E-path PRs but is left out of verify.needs for now. The detector, the tests/playwright. path fix, and the contract tests are unaffected — those stand on their own and were the substance of the review. Flipping to blocking is a three-line change once the suite is green; the exact wiring, including why the skipped allowance must sit outside the strict loop, is recorded on verify's Require-successful-checks step. The contract test pins the advisory choice so it reads as deliberate rather than as the unwired-gate bug it originally caught, and still fails if the path filter, the strict-loop coverage, or the config path regress. --------- Co-authored-by: Wooseong Kim <innocarpe@gmail.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
e551d3ec0d |
perf(lint): consolidate code-quality gates into Oxlint (#11117)
Consolidate standalone code-quality scanners into Oxlint, preserve focused native/type-aware enforcement, add custom plugin coverage, and harden deferred PTY test cleanup. |
||
|
|
badf91101b |
fix(quality): enforce performance-safe lint baseline (#11074)
* fix(quality): clear safe existing lint findings * fix(quality): keep lint cleanup allocation-free * fix(quality): enforce performance-safe baseline * test(terminal): drain deferred confirmation cleanup |
||
|
|
12ef12c55b |
chore(quality): ratchet Oxlint, React Doctor, and Zustand performance (#11034)
* chore(quality): ratchet lint and Zustand performance * fix(ci): stabilize React peer lock snapshot * fix(ci): isolate PR diff and React Doctor CLI |
||
|
|
c6076a507c |
ci(release): detach non-blocking E2E (#11031)
* ci(release): detach non-blocking E2E * test(release): pin E2E dispatch retries |
||
|
|
abcdc04f6b |
fix(ci): mirror missing lint steps in PR workflow (#10601) (#10623)
Reviewed with an independent reproduction. Added the allowlist entry that unblocked verify:localization-coverage on main, the 4th drifted step, and a parity gate that fails when pnpm lint's chain contains a script absent from pr.yml. |
||
|
|
39a200d900 |
fix(release): restore the Windows inner-binary signature gate (#6487) (#10719)
* fix(release): restore the Windows inner-binary signature gate
electron-builder 26.9+ dropped the bundled 7zip-bin package, so the gate's
hardcoded node_modules/7zip-bin path stopped resolving in
|
||
|
|
0f91af821d |
ci: parallelize PR checks and accelerate Vite builds (#10989)
* ci: parallelize and accelerate PR checks * fix(ci): make accelerated checks runtime-safe * fix(ci): address review findings * fix(ci): retry transient Electron downloads * test(ci): cover Electron download retry limits |
||
|
|
97e4776dfe |
feat(plugins): Orca plugin system — kernel, content packs, panels, workers, marketplace v0 (experimental) (#8549)
* feat(plugins): Orca plugin system — kernel, content packs, panels, workers, marketplace v0 (experimental) Adds Orca's experimental plugin system behind a settings flag: a supervised kernel, declarative content packs (VM recipes, commands and keybindings, language packs), sandboxed iframe panels, forked worker hosts, and a Git-backed marketplace v0 with consent, provenance and kill-list enforcement. Theme, icon-theme and terminal-theme contributions are deferred to a follow-up pass. * fix(plugins): make unsupported marketplace listings unreachable by key findPlugin() backs preview/install/previewInstalledUpdate via requireListing(), so filtering only listPlugins() hid the catalog card while leaving the dead install path reachable one click later. * fix(plugins): fan Pi session-only status out to plugin subscribers The providerSessionOnly early-return in applyNormalizedStatus emitted to onAgentStatus (main-window fanout) but skipped enrichedStatusListeners, so plugins subscribed to agent.status.changed silently missed every Pi session_start event. Route both emit sites through one helper so a future early return cannot drop the plugin tap again. Co-authored-by: Orca <help@stably.ai> * plugins: drop dead code and hoist duplicated trust-boundary patterns Cleanup pass over the P1 diff, no behavior change: - Delete `readPluginTreeSnapshot`/`readSnapshotFile` and their types, plus the now-vestigial `directories`/`signal` plumbing in `collectFiles`. - Delete `resolveContainedPluginDirectory` (no callers). - Delete `plugin-content-load-pool.ts`; it reimplemented the existing `mapWithConcurrency`, whose index arg also removes the pairing wrapper in `buildPluginList`. - Hoist `PLUGIN_CONTENT_HASH_PATTERN` and `PLUGIN_COMMIT_PATTERN` into the install-lockfile module; 11 sites hand-rolled these identically. - Point the new reliability gate at the PR instead of gitignored docs paths, matching every other gate's link form. * fix(plugins): retry plugin state renames on Windows AV/EPERM locks Six plugin write paths (lockfile, provenance, current pointer, kill list, marketplace cache, staged install dir) did a plain rename, so an antivirus or indexer holding the target open surfaced as a failed install. The repo already retries this hazard for issue #1507, but only through a sync helper; these paths are all async. Adds one bounded async retry + atomic write used by all six, and trims a consent-provenance header that restated its own JSX. * test(plugins): cover the Windows rename retry path The retry loop shipped untested: both existing cases hit the non-retry path, and the temp-cleanup test passed identically with the `finally` removed. Mock `rename` to queue errno codes so CI can exercise locks it cannot provoke. Co-authored-by: Orca <help@stably.ai> * fix(plugins): pin bundled plugin resources to LF Windows CI checks out with autocrlf, so the byte-hashed launch tree arrived as CRLF and verify-packaged-plugin-resources rejected it — the packaged build could never pass on Windows. Reproduced locally: CRLF yields the exact CI error, LF verifies clean. Files are already LF, so nothing renormalizes. Co-authored-by: Orca <help@stably.ai> * test: guard the bundled-plugin LF pin against a CRLF checkout The byte-hash mismatch only surfaced in Windows packaging CI. Assert the .gitattributes pin and that a CRLF tree is rejected, so a regression fails on any platform instead of waiting for a packaged Windows build. Co-authored-by: Orca <help@stably.ai> * ci: trigger packaged-build check on bundled plugin resource changes The launch tree is byte-hashed during packaging, but no trigger path covered it — so the CRLF fix for that check would not have re-run the check. Add the resources, verifier and .gitattributes paths that can break packaging. Co-authored-by: Orca <help@stably.ai> * perf(plugins): rebuild the panel frame only when its baked theme values change The revision keys the panel iframe, so every bump destroys the sandboxed frame and its in-panel state. It counted root attribute mutations, but --workspace-sidebar-live-width is written every rAF of a sidebar drag, so dragging with a panel open blanked it ~60x/sec. Compare the two values the shell actually bakes in instead. Co-authored-by: Orca <help@stably.ai> * test: stop pinning a plugin name in the CRLF guard The CRLF case rewrites every launch file, so the reported mismatch is whichever plugin sorts first. P2 adds theme plugins that sort ahead of orca-navigation-shortcuts, which broke the assertion there. Co-authored-by: Orca <help@stably.ai> * style: drop stray blank lines left by the rebase resolutions Both sides of the agent-hooks and orca-runtime conflicts contributed a trailing blank, which oxfmt rejects. Whitespace only. Co-authored-by: Orca <help@stably.ai> * test(plugins): stop the startup budget failing on machine load P95 runs 16-34ms idle but exceeds the 50ms bound under full-suite parallelism, so the gate flaked. Widen it to catch an order-of-magnitude regression instead; the no-worker/no-plugin-code assertions are the real guarantee. Verified a 400ms regression still fails. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
fc513233cb |
fix(release-cut): gate an explicit RC against its own series (#10525)
* fix(release-cut): gate an explicit RC against its own series
semver_gt compares through strip_pre(), so the explicit-version override
only ever checked the stable line: 1.4.156-rc.0 read as 1.4.156, cleared
a 1.4.155 stable, and republished an RC below what clients already run.
Anchor a prerelease request on highest_rc_for_base -- the same rc history
the kind path uses -- so the override can only advance the series.
Two sibling gaps in the same block:
- version_suffix was silently dropped when version was set, because the
append lives in the kind branch the override skips.
- the shape regex rejected X.Y.Z-rc.N.suffix, so a suffixed RC the rc
path can produce could never be re-cut explicitly.
* fix(release-cut): close both ends of the rc-number range the gate compares
The new explicit-rc gate compares with `[[ -le ]]`, i.e. bash machine-width
integers, and the author closed only the low end. Past INTMAX bash saturates,
so `version=1.4.156-rc.99999999999999999999` reads as "above the published
rc.3" and the gate falls open — then the tag it cuts pins
highest_rc_for_base at 1e20 for that base forever, and every later cut wraps
to a lower rc the fleet never updates to. Bound the rc number to nine digits.
Also reject leading zeros on an all-digit prerelease identifier. `npm version`
renormalizes rc.4.01 to rc.4.1 while the tag step keeps the literal input, so
the shipped package.json version and its own release tag name different
releases. The explicit path's embedded identifier now goes through the same
validator the kind path uses instead of only the shape regex.
* fix(release-cut): stop the refusal pointing minor/major RCs at the wrong series
kind=rc derives its base from bump(latest_stable, patch), so the remedy the
refusal suggested only works when the requested base *is* that next patch. A
1.5.0-rc.N series exists only because this override created it, so an operator
resuming a stuck 1.5.0-rc.2 was told to dispatch kind=rc, which would have cut
an unrelated 1.4.156-rc.4. Spell the condition out and give the fallback that
does work for a non-patch base.
Also correct the mechanism in the comment I added in
|
||
|
|
f009500677 | ci(release-cut): always show resolved commit, branch, and tag in summary (#10482) | ||
|
|
7a01910f20 |
fix(skills): advance the release ledger at the cut so shipped revisions freeze (#10483)
* fix(skills): advance the release ledger at the cut so shipped revisions freeze #10340 made the released-skill registry a function of the committed ledger instead of a git tag walk, and #10460 reverted the cut step that advances that ledger because it violated the #9119 contract (a version-only cut must not regenerate or stage the content-addressed skill artifacts). Both were right; the result is a ledger that never advances. generate-skill-bundle-manifest.mjs:390 derives releasedCount solely from release-mapping.json and :461 assigns a changed skill releaseRevision = releasedCount + 1, while :518 protects only committedReleasedCounts[name] — so index releasedCount is unprotected. A tag ships that tail revision, nothing records it, and the next skill change rebuilds the same revision number over different bytes. Installs carrying the shipped digest then match no snapshot and degrade to unrecognized, which cannot be updated. Restore the advance in a form the #9119 contract can keep enforcing: --release now verifies that current-manifest.json and snapshot-registry.json already match the ref being tagged, appends the mapping row, and writes only release-mapping.json. The cut stages just that file, so it still cannot move a content-addressed artifact — the failure #9119 guarded against — and now fails loudly instead of recording a revision the tag does not ship. The contract test is narrowed to match: it asserts the cut runs --release (never --write) and stages exactly package.json and release-mapping.json. * test(release-cut): close the staging bypasses the narrowed gate left open The narrowed contract test anchored its `git add` scan to line start and only inspected staged paths, so three ways to reintroduce #9119 stayed green: a `git add` chained after `&&`, a write that never calls `git add` at all, and `pnpm run generate:skill-bundle-manifest` — the package.json alias for `--write`, which the hyphenated ban never matched. That last one also passed the pre-#10460 assertions, so it was never covered. Drop the anchor, require every `resources/skills` mention in the step to be exactly what is staged, and ban the alias and `commit -a`. Comments are stripped first so prose cannot trip a ban. Verified each bypass fails and the real workflow passes. * fix(release-cut): make the new provenance failure actionable to an operator Verifying the content-addressed artifacts is the only new way the cut can block, and it fails inside a step named "Bump package.json and tag" with a lint-shaped message. That names the files and the command but not the two things the operator needs: the regeneration has to land on main, and the cut is safe to re-run afterwards. Say so. Also pin down why assertReleasedHistoryPreserved takes the pre-append mapping. It pairs with artifacts.releasedSnapshotCounts, which seeding fixed before the row existed; handing it the post-append mapping makes every cut throw "Released snapshot history is incomplete", which points at tag fetching rather than the real cause. Nothing enforces the pairing. * test(release-cut): gate the whole cut job, not just the bump step Round-2 review defeated the previous gate twice, both proved by running the full contract file green with #9119 reintroduced. Every step in the cut job shares one workspace and one index, but the contract test only inspected `Bump package.json and tag`. A step inserted earlier could run --write and `git add resources/skills`, and the bump step's own commit swept it into the version commit and the tag. Assert job-wide instead: only the bump step may name the directory, and no step may regenerate under either the flag or its package.json alias. That lives in the generator suite because the contract file is at its max-lines cap. Two regexes were also evadable. The mention scan required a trailing slash, so a path held in a variable was invisible; it now matches the directory itself. The `commit -a` ban matched nothing at all — `commit\s` ate the only separator, so `-a`, `-am`, and `--all` all survived while only a trailing `-a` was caught. `--allow-empty` stays allowed. * fix(release-cut): assert the index, not the workflow text, before committing Round-3 review defeated the job-wide grep three ways, each proved by running both test files green with #9119 reintroduced into the tagged commit: an `env:` block holding `--write` and `resources/skills`, a composite action whose steps the workflow never spells out, and plain shell concatenation (`root=resources; leaf=skills`). Grepping shell source for path literals is inherently evadable, and the previous fix only relocated round-2's variable-indirection hole one step over. Move the invariant to where it cannot be dodged: immediately before committing, the cut diffs its own index and refuses anything that is not package.json or the release-mapping row. That does not care which step staged what, or how the path was spelled. The workflow grep stays as a cheap tripwire for literal spellings, now paired with a positive assertion that the index guard exists and precedes the commit — indirection cannot hide a missing guard. Mention matching dedupes and trims quotes, since the guard names the row a second time. * fix(release-cut): match the staged-path allowlist literally `grep -vx` treats its patterns as regexes, so the `.` in `package.json` matched any character: a staged `packageXjson` or a `resources/skills/release-mappingXjson` was silently accepted by the index guard. Verified both slip through `-vx` and are caught by `-vxF`. Exercised the guard against a legitimate cut, an empty index, a staged content-addressed artifact, paths containing a space and a non-ASCII character (git quotes the latter, so it fails closed), and a staged deletion. Only the two allowed paths pass. * test(release-cut): assert the index guard aborts, not just that it exists The positive assertion pinned the guard's shape and its position before the commit, but not its effect: replacing `exit 1` with `:` left both test files green while the cut logged the error and shipped the artifact anyway. That is the same failure this whole gate keeps having — asserting the shape of a defense rather than what it does. Pin the abort too. Verified the neutered guard now fails the suite. * test(release-cut): scope the abort check and catch clustered commit flags Two holes in the guards this PR added, both in the same shape-not-effect class the previous commit was meant to close. The abort assertion's lazy match was not scoped to the guard's own block, so it could borrow an `exit 1` from any later `if ... fi` in the step. Degrading the guard to a warning while adding a plausible HEAD precondition left every test green. Stop the match at the guard's `fi`. The `commit -a` ban only matched when `a` led the flag cluster, so `-vam`, `-va`, `-qam` and `-sam` all survived. That matters more than it looks: `commit -a` stages at commit time, after the index guard has already inspected a clean index, so it is the one way to defeat that guard. Match `a` anywhere in a short-flag cluster; `--allow-empty` and `--amend` stay allowed. Verified both mutants now fail. * fix(release-cut): validate the commit, not the index, before tagging The index guard asserted the wrong thing. `git commit` has a family of forms that commit the working tree rather than the index — `-a`, `-i`, `--only`, and a bare pathspec — so a rogue earlier step could leave regenerated artifacts unstaged and any of those forms would carry them into the tagged commit while the guard saw a clean index and passed. Reproduced end to end: `git commit -i resources` put current-manifest.json and snapshot-registry.json in the tag with all gates green, and `--only resources` additionally dropped package.json from the tag. Banning those flags one by one is the same enumeration game the earlier rounds kept losing. Assert the outcome instead: after committing and before tagging, diff-tree HEAD and refuse anything that is not package.json or the release-mapping row. That is indifferent to which step staged what and to how the commit was spelled. Verified the whole family is now blocked (-i, --only, -a, -am, -vam, pathspec, and an alias expanding to `commit -i`), that a stock commit and an --allow-empty re-cut still pass, and that deleting, neutering, un-anchoring, or relocating the guard each fails the suite. * fix(release-cut): make the commit guard fail closed on a merge commit Plain `git diff-tree` prints nothing for a merge commit, so the guard would have passed silently instead of failing closed — the one direction that matters on a release path. `-m --first-parent` reports the diff against the first parent; verified byte-identical output for an ordinary commit and still empty for the `--allow-empty` re-cut, so nothing else changes. Not reachable today (nothing in the cut job creates a merge, and npm version has no lifecycle hooks defined), but the failure mode is a guard that looks like it ran. Pin the flags in the assertion too, so neither dropping -m nor slipping in a `--diff-filter` can weaken it without failing the suite. |
||
|
|
2653794c82 |
fix(terminal): verify Windows PTY root identity before taskkill /T /F (#10484)
* fix(terminal): verify Windows PTY root identity before taskkill /T /F killWithDescendantSweep guarded its Windows tree kill with ownsRoot() alone, which is JS state only. node-pty's ConPTY exit watcher closes the last shell handle before it queues the JS exit callback, so Windows can recycle the PID while the session map still looks live — force-killing an unrelated process and its whole descendant tree. Walk the recycled PID's ancestry back to this process before taskkill: skip the sweep when the root is gone or resolves to a stranger, and keep the sweep when identity is unknown so #10004 orphan cleanup still runs. Also gate the local provider's ownsRoot on observed physical exit. * fix(terminal): dedupe the Windows root-identity scan, drop dead exit gate Review fixes on the PID-identity guard. The probe read the process table through a new uncached export, bypassing the reader that worktree teardown depends on: worktree-teardown.ts fans out 32-wide inside a 10s deadline, so a delete forked 32 powershell cold-starts (the churn windows-foreground-process-rows.ts:25-32 warns about, #6288/#6667). getFreshSnapshot() already guarantees a scan that starts after the request -- the exact property the bypass existed for -- and coalesces concurrent callers, so use it. Measured on the new test: 32 scans -> 1. The PhysicalExitTracker.hasExited gate could never fire. markExited() is only reached at local-pty-provider.ts:985/:1431, and both are followed synchronously by clearPtyState(), which deletes the ptyProcesses entry -- so ownsRoot's map check is already false whenever hasExited is true. Reverting it broke no test. Drop it and the shared getter it added; the identity probe already covers every ownsRoot caller from inside killWithDescendantSweep. Also point the Windows terminal-restart E2E job at the files that own this behavior, so a change to the new Windows-only module runs the one job that executes on a real Windows host. * docs(terminal): state what the Windows root probe actually proves The probe checks subtree membership, not root identity: a recycle that lands on another Orca descendant (another pane's shell, an agent CLI, a git.exe we spawned) still reads `own`, and that is not remote during teardown when Orca is itself allocating pids. It bounds the blast radius rather than closing the class. Say so at the type and at classifyWindowsTreeKillTarget, and name what a real close would need (a CreationDate baseline -- the analogue of the POSIX lstart check already used here -- or an inherited handle / Job Object). Also note why our own pid must classify `foreign`. * ci(windows): trigger the terminal-restart E2E on the shared snapshot reader The Windows root-identity probe now reads through getFreshSnapshot, so an edit to that module changes Windows teardown behavior without touching any path the job already watches. * test(terminal): guard the teardown probe against a reintroduced scan bypass The existing volume guard covers queryWindowsProcessRowsFresh directly, but the identity-probe cases all inject readRows, so nothing exercised the DEFAULT reader wiring -- a bypass reintroduced inside windows-pty-root-identity would have gone unnoticed. Drive verifyWindowsTreeKillTarget 32-wide through the real reader and assert one scan. Verified it fails at 32 when the bypass is put back. |
||
|
|
d5340fd191 |
ci(release-cut): restore the skill-independent version commit (#10460)
#10340 added a ledger-advance step to the release cut, which violates the contract test #9119 added: the cut must not run generate-skill-bundle-manifest or stage resources/skills. Both PRs were green on their own branches and only conflicted once merged, so nothing failed until main had both — main and every open PR have been red since. Revert the two workflow lines. The script's --release implementation stays: it is correct and harmless when unused, and the root fix in #10340 — verify no longer walking git tags — does not depend on the cut step. This leaves #10340 semantically incomplete and that must not be dropped. With the registry seeded from the committed ledger instead of a tag walk, nothing advances the ledger at cut, so each new skill change re-uses the same unreleased tail revision for different bytes; older installs then match no known snapshot and degrade to unrecognized, which reads in the UI as a skill that needs attention and cannot be updated. Follow-up is to reintroduce the advance narrowly — stage only resources/skills/release-mapping.json and narrow the assertion to forbid mutating the content-addressed artifacts while permitting the provenance row. |
||
|
|
8d61d76a59 |
fix(skills): decouple skill-manifest verify from local git tags (#10340)
* fix(skills): source released history from the committed ledger, not a tag walk verify:skill-bundle-manifest rebuilt the entire released-skill history by walking every local refs/tags/v* on each run and demanded byte-equality with the committed artifacts. Output was therefore a function of (skill bytes x local tag set x release timing), so any clone holding stray, deleted, or fork tags the committed artifacts predate rebuilt a divergent registry and failed lint. This was the 4th instance of one failure class (#8637 -> #9119 version bumps -> #9778 new tags -> local tag drift), each patched with a new tolerance rather than removing the tag coupling. Fix: the committed snapshot-registry + release-mapping ARE the released history; trust them instead of re-deriving from tags. - releasedHistoryFromCommitted() seeds generation from the committed ledger, dropping the floating unreleased tail (entries beyond what the mapping names). verify and --write are now pure functions of working-tree bytes with zero tag access. The tag walk survives only behind --rebuild-from-tags (disaster recovery), off the everyday path. - --release <version> + appendReleaseRow() perform the O(1) append of one mapping row at release cut (dedupes vs the last row, strips the v-prefix) -- the single authoritative point where working-tree bytes become an immutable released revision. - release-cut.yml runs generate --release "$VERSION" before the release commit (Node built-ins only, no install needed); pr.yml drops fetch-depth: 0 from the lint job since verify no longer needs tag history. Recognition is unaffected: the runtime uses knownSnapshots = registry.skills (all entries, incl. the tail committed at PR-merge time), so a missing mapping row only loses a version label, never recognition or the update nudge. Trade-off: lint no longer cross-checks committed historical snapshots against tags. A hand-edit to an old released entry is still caught by the runtime manifest<->registry consistency check when the current manifest points at it, and can be audited anytime with --rebuild-from-tags. Verified: verify passes committed-sourced; --write is zero-diff (byte parity); a planted stray v-tag no longer changes output; edit-stub -> --write -> --release appends the correct single row; double --release is idempotent; --rebuild-from-tags reproduces the committed artifacts. Generator tests 14 pass/ 1 skip; runtime skill-bundle-artifacts + freshness-inventory 14 pass; bundled skill guides verify passes. * fix(skills): keep one release-mapping row per version on a re-cut A cut that pushed the version bump to main but died before pushing the tag is re-cut at the same version. If skills changed in between, the second --release appended a duplicate row, and the stale one named revisions that tag never ships — which verify-skill-update-roundtrip then pairs with the tag's real bytes. Overwrite the trailing row instead (the tag is absent, so that version was never published). Refuse only when an earlier row claims the version, which the cut workflow already rejects upstream, so this cannot wedge a recovering cut. |
||
|
|
efe996a007 |
ci(release-cut): add explicit version override to the cut dispatch (#10329)
* ci(release-cut): add explicit version override to the cut dispatch Kind-based computation derives the next version from the latest *published* stable. When a shipped stable is deleted/rolled back, the release list regresses to the prior stable, so a `kind` cut recomputes a number at or below the deleted one — stranding every client that already installed it, since electron-updater only moves forward. The existing package.json floor only recovers this when the deleted version's bump commit is on the ref being cut, which a hotfix cut from an older RC ref does not carry. Add an optional `version` workflow_dispatch input that lets a human assert the exact target (e.g. leapfrog a deleted 1.4.154 to 1.4.155), bypassing kind-based computation. The updater-safety gate (must exceed the latest published stable) and the existing tag-collision recovery still apply. Empty by default, and forced empty for scheduled cuts, so normal automation is unchanged. * ci(release-cut): let explicit version override the package-floor recovery Per review: the package.json floor block can recover_unpublished_tag and exit 0 before the explicit-version branch runs, hijacking an explicit request to recover a floor tag instead — the exact rollback scenario the override targets. Skip floor-tag recovery when EXPLICIT_VERSION is set; latest_stable is still raised to the floor for the safety gate, and the requested tag's collision recovery runs later. |
||
|
|
aab112933e |
Revert "fix(memory): bound OOM-prone accumulators (#10179)" (#10255)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
801ff57e83 |
fix(mobile): unblock iOS releases (#10224)
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
8f40ddf328 | fix(memory): bound OOM-prone accumulators (#10179) | ||
|
|
56101422d2 |
fix(release): regenerate Windows blockmap via app-builder-lib JS (#10110)
electron-builder 26 dropped the app-builder-bin Go binary, so the signed-installer staging step failed with 'node_modules/app-builder-bin/ win/x64/app-builder.exe is not recognized'. Blockmap generation now lives in app-builder-lib's pure-JS buildBlockMap; call it through a small script in both the release-cut and signing-rehearsal workflows. Co-authored-by: Orca <help@stably.ai> |
||
|
|
5d0c29f729 |
ci(release-cut): show workflow_dispatch inputs in job summary (#9866)
Mirror the noqa deploy workflow pattern so Cut Release runs surface kind/ref/dry_run/version_suffix as a table under Workflow Input Parameters. |
||
|
|
0b71f3bfba |
test(e2e): prove the terminal daemon survives a main-process crash on Windows (#7742) (#9311)
* test(e2e): prove the terminal daemon survives a main-process crash on Windows (#7742) Add a win-crash-survival e2e harness (sibling to win-update-e2e) that force-kills ONLY the packaged app's real Electron main (resolved via app.evaluate -> process.pid, /F no /T) and asserts the detached orca-terminal-daemon.exe plus its ConPTY shell survive with no pwsh 0xE9 FailFast, then that a relaunch re-adopts the SAME daemon and the reattached UI binds to the SAME survivor shell (proved via a per-shell env sentinel read back through the restored terminal). This guards the #7742 fix (standalone relocated daemon that outlives main death) against regression. A directional `--expect orphaned` profile fails on a fixed build, keeping the survival assertions honest. Windows-only; reuses win-update-e2e app-driver/daemon-process modules. * test(e2e): harden Windows crash-survival proof * test(ci): keep crash survival gate durable * test(e2e): tolerate restart hydration navigation * test(e2e): prove exact shell input after crash * perf(ci): avoid crash harness installer rebuilds * test(ci): harden crash survival evidence and cost * test(e2e): fail closed on authoritative crash target * test(e2e): fail closed on crash liveness evidence --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
803a442868 |
test(e2e): harden terminal restart regressions on Windows (#9064)
* test(e2e): harden Windows terminal restart regressions * test(e2e): cover renderer replacement rejection * test(e2e): tolerate ESRCH when force-killing the daemon on POSIX * fix(ci): preserve Windows restart test selection --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
cc1ad064d7 |
fix(skills): decouple bundled skill artifacts from the release train (#9119)
The current manifest stamped package.json's version into itself (9 lines), so every RC/stable version bump made the committed artifact stale on every open branch: lint failed until authors committed content-free regeneration diffs, which also dragged the resources/skills-filtered update-roundtrip matrix onto unrelated PRs. Cutting a release tag whose skills tree changed had the same effect through release-mapping.json. - current-manifest.json is now schema 2 and content-only; the generator no longer reads package.json. Registry and mapping stay schema 1 so the append-only released-history guard keeps its schema gate. - The running build's version enters at the IPC boundary (skills:freshnessInventory passes app.getVersion()) and threads through the inventory to placement observation; current-revision placements are labeled with it while historical revisions keep resolving through the release mapping. The artifact loader and its cache stay content-only. - verify tolerates a committed release mapping that is a byte-exact prefix of the derived one when every missing trailing row's revisions equal the current manifest (a just-cut tag over unchanged-since bytes); such rows are provably redundant until the next real regeneration adds them. Artifacts now change only when skills/ content changes. |
||
|
|
f6f2561623 |
ci(release): regenerate skill manifest on version bump (#9117)
* ci(release): regenerate skill manifest on version bump The release-cut "Bump package.json and tag" step bumped package.json but never regenerated resources/skills/current-manifest.json, so its appVersion stayed at the prior release. That drift shipped in v1.4.144-rc.1, rc.2, and rc.3 (all carried an rc.1 manifest) and turns verify:skill-bundle-manifest red on every branch after a cut, since that check runs in `pnpm lint` and the PR `verify` job. Regenerate the manifest right after `npm version` and stage resources/skills into the release commit so the bundled manifest always matches the shipped version. The generator is dependency-free (node builtins + git), so it runs without a pnpm install, and the step's fetch-depth:0 checkout supplies the tag history it reads. * test(release): guard skill manifest regeneration * test(release): require full history for skill manifest |
||
|
|
68fca0b076 |
Add safe skill freshness detection and update rail (#8637)
* Add safe skill freshness detection * Accept observed copy-mode rail outcomes * chore(skills): regenerate snapshot artifacts for the merged guide content The rebase onto main picked up the reviewed guide fixes (#8624), so the current manifest hashes and a new appended snapshot generation must match those bytes; the registry keeps all prior snapshots so existing installs classify as outdated rather than unrecognized. * fix(skills): canonicalize snapshot file order and guard released history Historical snapshots kept git ls-tree byte-order while the working-tree walk and runtime observation use the sorted depth-first order, so any future multi-file skill would misclassify older installs as unrecognized and churn spurious registry revisions; all producers now share one canonical order (no digest changes for today's single-file packages). Also rejects executable files from shipped skills (Windows observation cannot see execute bits, which would misclassify pristine Windows installs) and adds an explicit append-only invariant for released snapshots so a generation-logic change cannot rewrite them silently. * fix(skills): throttle focus rescans and correct self-blocked placement copy Every window focus re-read and re-hashed all installed packages, and the nudge and panel each forced their own trailing rescan for one event; a 15s cooldown plus a shared invalidation latch keep one bounded scan per event while install-change events stay immediate. Bundle artifacts are now loaded once per run instead of re-parsed on every scan. A read-only or otherwise unsupported outdated placement now explains that it blocks itself instead of blaming a phantom sibling placement; the supported topology set moved to shared so eligibility and copy cannot drift. * feat(skills): move freshness surfacing to a lingering toast and update modal The Skills page has been unreachable since its toolbox menu entry was removed (#4535), so surfacing freshness there buried the feature behind its own nudge. The nudge now lingers until acted on (ignoring it records nothing; only the explicit close persists dismissal keys) and opens an update modal hosting the pre-filled editable terminal, an honest current/blocked summary, and the per-placement rows in a collapsed Details section. A compact 'Check for skill updates' row in CLI settings is the manual re-entry point. Skills page restored to main; design-doc surfacing section records the venue decision. * fix(skills): mount update dialog inside the link-routing provider and fold freshness into the setup rails The dialog hosts a live terminal pane that requires the link-routing preference context; mounted outside the provider it crashed the renderer the moment an eligible update existed (caught by live QA — unit tests mock the terminal). It now mounts inside the provider behind its own recoverable boundary. The separate 'Check for skill updates' settings button is gone: the setup rails' own pill now carries freshness (Update available / Up to date, falling back to Installed for blocked or unrecognized copies and for non-local runtimes the local-only scan cannot vouch for), and Re-check refreshes both installation detection and the freshness inventory. Wired for the CLI, Orchestration, Computer Use, and Per-Workspace Environments rails. * fix(skills): use the sleek scrollbar style in the update dialog * chore(skills): regenerate manifest for merged main (v1.4.142-rc.1) Main advanced to 1.4.142-rc.1 with a v1.4.141 release, so the embedded appVersion and release mapping were stale on the PR's merged tree. Only appVersion and the new release entry change; no snapshot digests move (released history preserved). * fix(skills): bound and batch freshness work * fix(skills): harden freshness integrity checks * fix(skills): accept observed copy topology outcomes * chore(skills): regenerate manifest for current main * fix(skills): preserve update terminal lifecycle * chore(skills): regenerate manifest for current main * fix(skills): fail closed on stale freshness scans * chore(skills): regenerate manifest for current main * fix(skills): preserve freshness safety under focus churn * feat(skills): group the update modal by skill with plain-language status The Update skills modal now lists only skills that will update or that can't (with why), grouped by skill with their install locations nested underneath — no more one row per placement. - Statuses collapse to "Update available" / "Can't update" at the skill level. - A location's problem is a chip (Duplicate, Unrecognized, Inaccessible, Read only, In a repo, External/Broken link, Plugin cache) with a hover tooltip that explains what it means for the user and what to do. - Up-to-date, unrecognized-only, and unreadable-only skills are hidden; a current/unrecognized/etc. location only appears when it explains a shown skill. - Copy is de-jargoned (drops "copy"/"placement"/"snapshot"/"official copy") and names the mechanism as the npx skills update command, not "Orca's update". - Rename the section to "Update details"; drop the unreachable newer-known state. Renderer-only: derivation is a pure module (groupSkillFreshness) with unit tests; no IPC or main-process change. Locales updated for all five languages. * chore(skills): regenerate manifest for current main (v1.4.143-rc.0) * feat(skills): don't let a duplicate block the update; clearer skipped copy - Eligibility: a clean standalone duplicate no longer poisons the whole name — the canonical copy still updates and the duplicate is flagged; a duplicate-only skill stays unoffered. - Update modal: "Can't update" -> "Skipped" with a reason-specific sentence (edited/read-only/in-a-repo/plugin/link); chips describe only the location state; footer "Check now" -> "Re-check". - Settings sidebar nav pills go amber "Update available" when a skill is updatable, matching the setup cards. - Localized new strings across en/es/ja/ko/zh. * chore(skills): regenerate manifest for merged main (v1.4.144-rc.1) |
||
|
|
d3fedb9ba9 |
ci(computer-e2e): run native-smoke vitest with the shared config (#8909)
The native-smoke job ran bare `pnpm vitest run <files>` with no --config. There is no root-level vitest/vite config, so vitest fell back to its built-in defaults: a 5000ms testTimeout and no Windows worker cap. The real csc.exe launcher-compile tests (build-windows-cli-launcher and ssh-remote-cli-launcher) cold-start in 1.4-7.4s on hosted Windows and tipped over the 5s default, failing the PR check. Pass --config config/vitest.config.ts like every other vitest invocation in the repo (pnpm test, pr.yml) so the job gets the shared 30s timeout and Windows maxWorkers:4 cap (less CPU contention -> lower compile variance). Positional file args still filter to just the listed files. |
||
|
|
59a7fffcd6 |
fix(terminal): keep WebGL glyph atlas pages within the shader sampler budget (#8672)
* fix(terminal): keep WebGL glyph atlas pages within the shader sampler budget The fragment shader has sampler slots for maxAtlasPages (16 on most Macs) and leaves outColor uninitialized for any higher page index, so glyphs rasterized onto pages past the budget render as garbled pixels. Long sessions grow past the budget via the merge fallback, and the previous wipe fix re-activated those unbindable pages, so every atlas wipe re-allocated glyphs onto them (post-wipe allocation prefers the last, highest-index active page) and garbled whole panes mid-stream. Fix, matching the direction xterm.js maintainers are pursuing upstream (xtermjs/xterm.js#6043): a shared _evictAllPages resets the atlas to one fresh page, called from clearTexture and from the two allocation paths that could otherwise push a page past the budget (merge fallback and oversized-glyph page creation), so the page count can never exceed the renderer's texture capacity. Defensive backstops: a one-time warn plus bind-loop clamp, and an else branch in the generated shader so an unexpected overflow renders blank instead of undefined pixels. * test(terminal): cover WebGL atlas sampler budget * fix(terminal): align WebGL atlas invalidation source |
||
|
|
77b154d5dd |
Add Orca Relay desktop and mobile transport (#8536)
* feat(mobile): define relay protocol groundwork Co-authored-by: Orca <help@stably.ai> * feat(mobile): implement replay-safe E2EE v2 sessions Co-authored-by: Orca <help@stably.ai> * test(auth): lock cloud refresh single-flight Co-authored-by: Orca <help@stably.ai> * test(mobile): complete E2EE v2 adversarial coverage Co-authored-by: Orca <help@stably.ai> * refactor(runtime): unify mobile socket wiring Co-authored-by: Orca <help@stably.ai> * feat(runtime): add relay control and data clients Co-authored-by: Orca <help@stably.ai> * feat(runtime): coordinate desktop relay sessions Co-authored-by: Orca <help@stably.ai> * fix(auth): fence stale cloud session mutations Co-authored-by: Orca <help@stably.ai> * feat(runtime): add relay pairing and durable revoke Co-authored-by: Orca <help@stably.ai> * feat(runtime): add relay credential pairing RPCs Co-authored-by: Orca <help@stably.ai> * feat(settings): show Orca Relay sign-in status Co-authored-by: Orca <help@stably.ai> * test(relay): prove desktop lifecycle and E2EE splice Co-authored-by: Orca <help@stably.ai> * feat(mobile): persist relay pairing state Co-authored-by: Orca <help@stably.ai> * feat(mobile): race direct and relay pairing Co-authored-by: Orca <help@stably.ai> * feat(mobile): recover pairing through relay director Co-authored-by: Orca <help@stably.ai> * fix(relay): preserve origin controls during drain Co-authored-by: Orca <help@stably.ai> * feat(mobile): recover interrupted relay pairing Co-authored-by: Orca <help@stably.ai> * feat(mobile): add stable relay RPC sessions Co-authored-by: Orca <help@stably.ai> * feat(mobile): supervise direct and relay endpoints Co-authored-by: Orca <help@stably.ai> * Cover mobile relay director fallback matrix Co-authored-by: Orca <help@stably.ai> * Fix relay settings component test isolation Co-authored-by: Orca <help@stably.ai> * Remove unrelated merge formatting drift Co-authored-by: Orca <help@stably.ai> * Update runtime connection count integration assertion Co-authored-by: Orca <help@stably.ai> * Run mobile typecheck through pnpm Co-authored-by: Orca <help@stably.ai> * feat(relay): gate desktop controls on mobile demand Co-authored-by: Orca <help@stably.ai> * test(mobile): cover served relay recovery Co-authored-by: Orca <help@stably.ai> * feat(mobile): upgrade direct pairings to relay Co-authored-by: Orca <help@stably.ai> * fix(relay): harden mobile reconnect and teardown Co-authored-by: Orca <help@stably.ai> * fix(auth): clarify account sign-in state Co-authored-by: Orca <help@stably.ai> * fix(auth): polish sign-in completion flow Co-authored-by: Orca <help@stably.ai> * fix(auth): clarify sign-out confirmation Co-authored-by: Orca <help@stably.ai> * fix(auth): simplify sign-in completion page Co-authored-by: Orca <help@stably.ai> * feat(mobile): add per-device pairing connection mode Co-authored-by: Orca <help@stably.ai> * fix(mobile): stabilize pairing option layout Co-authored-by: Orca <help@stably.ai> * fix(mobile): give pairing choices stable space Co-authored-by: Orca <help@stably.ai> * fix(mobile): stabilize pairing QR regeneration Co-authored-by: Orca <help@stably.ai> * Animate mobile pairing flow height Co-authored-by: Orca <help@stably.ai> * Configure auth in packaged builds Co-authored-by: Orca <help@stably.ai> * Make Orca Relay pairing an opt-in beta Co-authored-by: Orca <help@stably.ai> * Show Relay beta details on hover Co-authored-by: Orca <help@stably.ai> * Refine mobile relay pairing choice Co-authored-by: Orca <help@stably.ai> * Polish Orca Relay pairing controls Co-authored-by: Orca <help@stably.ai> * Keep mobile contract fallback test additive Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
31f643ca42 |
Add version-matched skill guides to the CLI (#8624)
* Add version-matched bundled skill guides * Clarify skill freshness rollout PRs * Add canonical skills show alias * fix(skills): address guide review feedback * fix(skills): make guide commands cross-platform * fix(skills): apply the ORCA convention to the emulator guides Review follow-up: the emulator guides still instructed literal `orca emulator ...` in sh fences with no Linux disambiguation, so on unmanaged Linux they could launch the GNOME screen reader — the exact failure the executable-selection preamble prevents. Both emulator guides now carry the preamble and ORCA placeholder across fences, tables, and prose, and the cross-platform safety test covers all four converted guides. Also replaces computer-use's "unless a block names a shell" carve-out, which contradicted its own POSIX example, with the unconditional placeholder rule. |
||
|
|
302b97029a |
P2 windows cli hardening (#8638)
* fix(cli): harden Windows launcher transports * Fix csc.exe compile failures on space-bearing Windows install paths - Legacy csc.exe mangles absolute paths containing spaces, so the compile step now cd's into the bin directory and passes bare file names for /out and the source file instead of full paths |
||
|
|
b69c6043e3 |
Ssh watcher isolation e2e (#8494)
* Add a Docker SSH watcher-isolation E2E test to verify remote relay-watch - Covers two scenarios: crashed watcher children are respawned under the same relay without dropping the terminal PTY or file-explorer view, and a missing deployed relay-watcher.js artifact is repaired on reconnect - Extracts shared connect/disconnect/reconnect logic out of the perf spec into docker-ssh-relay-connection.ts, and adds docker-ssh-relay-processes.ts for inspecting/signaling remote relay and watcher PIDs - Wires the new spec into a dedicated CI job and pnpm script * Fix Windows and Linux-only issues in Docker SSH watcher-isolation E2E ha - node-gyp override only applies on Linux runners now, since the CI job moved to ubuntu-latest but shares the workflow with non-Linux jobs - spawn the e2e runner scripts through a shell on win32 to satisfy Node's CVE-2024-27980 restriction on unshelled .cmd spawns - harden relay process row parsing against empty pid/ppid fields so a vanished /proc entry fails loudly instead of coercing to pid 0 - dedupe the reconnect helpers and export shellQuote for reuse across the docker-ssh-relay test helpers |
||
|
|
43e481b1c3 |
Revert "Decouple feature copy from translated locale catalogs (#8488)" (#8500)
This reverts commit
|
||
|
|
a5e9e139b1 |
Decouple feature copy from translated locale catalogs (#8488)
* Decouple feature copy from locale catalogs * Update PR workflow contract tests * Address localization review findings * Document localization cache context |
||
|
|
e3c47eff17 |
Fix ssh watcher isolation (#8463)
* fix(ssh): isolate relay filesystem watchers
* Fix relay watcher fault-harness pid file and in-process fallback isolati
- Use exclusive ('wx') creation for the fault-harness pid file so a leaked
ORCA_WATCHER_CHILD_PID_FILE env var can't clobber an existing file, and
have the harness remove the file after reading a replacement pid.
- Force useInProcessVitestFallback to false in the relay watcher pool so a
leaked VITEST env var can never load the native watcher addon in-process
on the relay; fail closed instead when the isolated child is missing.
- Thread an injectable RelayWatcherProcessPool into FsHandler/
RelayFilesystemWatchRegistry for tests, and add coverage for both fixes.
|
||
|
|
433df4be3c |
fix(runtime): prevent file watcher SIGSEGV from crashing orca serve (#8370)
* Fix crash-isolated file watcher process pool for orca-serve SIGSEGV afte Replace the worker-thread runtime file watcher with a forked, crash-isolated @parcel/watcher child process pool so a native FSEvents fault can no longer take down the main/serve process, and add bounded event batching, delivery backpressure, and quarantine-based recovery for faulty watch roots. * Fix crash-isolated file watcher teardown and shutdown leaks - Fault harness could throw before mkdtemp/realpath completed, skipping cleanup; now tracks each temp path independently and races an async watcher-callback error so it can't escape the try/finally unhandled. - In-process fallback swallowed unsubscribe failures via a bare rejection handler that could still throw; use .catch() instead. - Watcher process entry's cancel-subscribe handler now reuses the async unsubscribe path when a crawl already finished, releasing the native handle instead of leaking it (blocks worktree unlock on Windows). - Runtime watcher process pool exposed no real dispose(); shutdown now kills pooled children so they don't outlive the main process. * Fix disposeSlot double-iteration bug in file watcher pool teardown Remove the unnecessary array snapshot in dispose(): disposeSlot mutates allSlots by deleting the slot being visited, and deleting the in-progress element during Set iteration is well-defined, so the spread copy was dead weight left over from prior debugging. * Fix pending file watcher installs not aborting on unsubscribe - Local/WSL watcher installs and SSH fs.watch setup now honor the in-flight AbortSignal, so the last unwatch cancels a slow native subscribe or remote setup instead of waiting for it to finish. - Thread signal through IFilesystemProvider.watch and SSH-backed file explorer watches for the same early-cancel behavior. * Fix crash-resubscribe hangs and SSH watch teardown races in file watcher - Add a bounded deadline for post-crash resubscription crawls so one stuck root quarantines instead of pinning its whole shard forever. - Report FSEvents overflow as recoverable so delivery continues after a dropped-events error instead of surfacing as terminal. - Make WSL watcher abort errors real DOMException instances so AbortSignal-based cancellation checks recognize them. - Rework SSH watch registration so ownership of the shared setup request (not just the first caller) decides teardown, preventing one caller's abort from cancelling another's shared watch and guaranteeing exactly one fs.unwatch per registration. - Reformat reliability-gates.jsonc arrays and refresh WSL/SSH coverage entries and evidence runs to match the above. * Add CI gate to run the file-watcher SIGSEGV fault harness under Electron - The reliability gate and release workflows (mac, Linux) previously only exercised the crash-isolation harness under vanilla Node, which doesn't catch runtime differences in the actual Electron binary that ships to users. - Adds an `ELECTRON_RUN_AS_NODE=1 pnpm exec electron ...` run of the same harness alongside the existing Node run, so #8212's SIGSEGV-survival contract is proven against both runtimes before packaging. * Add CI gate blocking Linux/macOS release packaging on watcher fault reco Adds a contract test asserting release-cut.yml and release-mac-build.yml run the runtime-file-watcher-fault-harness after building and before publishing artifacts, so a regression in watcher process fault recovery fails release packaging instead of shipping silently. * Fix use-after-clear crash in failAllWatcherSubscriptions Snapshot the records map before iterating, since onTerminalError hooks can dispose the supervisor and clear `records` mid-loop, causing a crash. Also update the matching test to assert against the shared buildParcelWatcherIgnoreOptions helper instead of a loose arrayContaining match. * Fix use-after-clear crash in failAllWatcherSubscriptions Snapshot watcher records with Array.from instead of spread, since spread syntax over an iterator that's mutated mid-loop by onTerminalError hooks can produce inconsistent results. |
||
|
|
ee82d66a35 |
fix(cli): preserve multiline arguments on Windows (#8374)
* fix(cli): preserve multiline Windows arguments * test(cli): run Windows launcher regression in CI * fix(cli): support Windows Framework C# compiler |
||
|
|
4f9fbecb9e |
Assign GitHub issue types (Bug/Feature) via issue templates (#8346)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
533992bdda |
fix(git): cache unsupported capabilities per host (#8109)
* fix(git): cache unsupported capabilities per host Old Git worktree, ref-search, and merge-tree fallbacks retried unsupported flags on recurring operations, flooding subprocess traces. Centralize capability probing per native, WSL, and SSH execution host, coalesce concurrent probes, and retry periodically for in-place Git upgrades. * fix(git): recognize real old-Git merge-tree rejection * test(git): enforce real binary compatibility matrix * fix(ci): preserve Git compatibility test ownership * fix(git): retain supported capability state |
||
|
|
e84a8ddec9 |
Terminal performance initiative: pipeline fixes + term-speed-2 revival + PTY flow control (integration branch) (#7214)
* Skip legacy hidden skip grammar assertions
* Fix hidden TUI snapshot test setup
* Fix sleep wake history test contract
* Fix hidden delivery startup gate helper
* Fix hidden Latin skip branch predicate
* Fix hidden synchronized split-boundary replay
* Stabilize remote runtime mixed subscription test
* Keep hidden startup query parser active during window
* Stabilize raw emoji golden restore width
* Stabilize raw emoji golden fixture completion
* Keep terminals responsive under agent output load
* Add frozen-terminal repro harness and silent-drop regression tests
Investigation harness for the frozen-terminal reports (Discord
#performance, issue #2836): pane shows content, shell alive, daemon
output.log flat while typing.
- e2e: renderer crash -> auto-reload recovery and three restart/restore
shapes (live daemon, SIGSTOP-wedged daemon, daemon killed between
launches), each probing input at both drop layers. Post-crash phases
drive the renderer from the main process because a crashed target
severs Playwright's CDP session even though the app recovers.
- e2e helpers: layer-discriminating probes (direct pty.write vs
transport input, plus pty:listSessions ownership-rebuild revival).
- unit repro: vendored xterm 6.1.0-beta.287 WriteBuffer permanently
wedges when a sync throw escapes a write-completion callback or a
custom parser handler (xterm-write-buffer-stall.repro.test.ts).
- unit repros for both silent input-drop layers: main drops writes for
a live PTY once ptyOwnership loses the id (revived by listSessions),
and the renderer transport stays unbound after a failed connect.
- pty.test.ts: unregister every leaked SSH provider id in afterEach so
module-level provider state cannot leak across tests.
Co-authored-by: Orca <help@stably.ai>
* Harden xterm write pipeline against sync-throw wedge that freezes panes
A synchronous exception escaping xterm's WriteBuffer loop permanently
wedges that terminal: _innerWrite has no try/catch around the parse
action or the write-completion callback, the tail re-schedule never
runs, and write() only re-arms on an empty buffer. The pane stops
rendering and, if a replay was in flight, the replay guard latches and
pty-connection's onData silently eats every keystroke — matching the
field reports (Discord #performance, issue #2836: content visible,
shell alive, daemon output.log flat). Both vectors verified against
vendored xterm 6.1.0-beta.287 in xterm-write-buffer-stall.repro.test.ts.
Three layers of defense:
- Guard every write-completion callback Orca hands xterm at the two
choke points (writeForegroundTerminalChunk, writeBackgroundTerminalChunk),
with settle and onParsed guarded separately so a WebGL/renderer
failure during viewport settle cannot starve the replay-guard release.
- Guard all throwing-capable custom parser handlers (DA1, OSC 10/11,
CSI ?h/?l mode reports, OSC 52 clipboard, OSC 7 cwd), degrading a
throw to "not handled" — same escape class as
terminal-link-provider-guard.ts.
- Replay-guard watchdog: each engagement releases exactly once, from
xterm's completion or a 10s watchdog, so a lost completion (wedged
pipeline, disposed-terminal race) cannot latch the guard on a live
pane; replayIntoTerminalAsync resolves on either path so restore
chains cannot hang. Force-releases record a crash breadcrumb.
All guard trips record rate-capped crash breadcrumbs, so the next field
occurrence names the throwing stack instead of failing silently.
Co-authored-by: Orca <help@stably.ai>
* Cap unbounded terminal output buffers in main and the foreground queue
Field evidence (Discord #performance / #2836): renderer memory climbs to
~1.5 GB and terminals freeze; a force reload does not help until memory
recovers. Two unbounded buffers matched that shape:
- Main-process pendingData grew by string concatenation without bound
while the renderer could not receive (frozen, starved, mid-reload) —
main-heap bloat a renderer reload cannot clear. Now capped at 2 MB per
PTY: past the cap the buffered bytes are dropped and the entry stays
O(1) until the renderer ACKs again, then a droppedOutput sentinel is
delivered and the pane repaints from the authoritative main-owned
buffer snapshot (existing hidden-output restore path) instead of
continuing a stream with a silent gap.
- The renderer output scheduler capped only hidden-pane backlogs; the
foreground path could queue a visible pane's flood without bound when
the drain could not keep up. The 2 MB cap now applies to every
foreground enqueue branch too, with a foreground-specific skip notice.
Verified: new main-side cap test (starve → flood → sentinel → normal
flow resumes), renderer sentinel-to-snapshot-restore test, two
foreground scheduler cap tests; full pty/terminal-pane/pane-manager
suites (1981 tests) and typecheck pass.
Co-authored-by: Orca <help@stably.ai>
* Make replay-guard stall release probe-certified instead of time-based
The previous stall watchdog blindly released the input guard after 10s.
If a replay were genuinely still parsing on a starved machine, that
early release could leak xterm's auto-replies into the shell — and into
agent TUIs, where a leaked ESC reads as the user pressing Escape.
Replace the blind release with a probe: when a completion looks
overdue, enqueue an empty write behind the replay. xterm parses writes
in order, so every outcome is provably safe:
- probe parses after the replay completion ran: normal release already
happened; probe is a no-op.
- probe parses but the replay completion never ran: all replay bytes
have parsed, no further auto-replies can exist — the completion was
genuinely lost. Release + breadcrumb.
- probe never parses (bounded wait): the pipeline is wedged, and a dead
parser can never emit auto-replies, so releasing cannot leak input.
Release + breadcrumb naming the pane as needing recovery.
While the probe is pending — a slow-but-alive replay — the guard now
HOLDS instead of releasing early; that case is pinned by a regression
test.
Co-authored-by: Orca <help@stably.ai>
* Scale output backlog caps with the scrollback setting and breadcrumb drops
The 2 MB pending-output caps were flat, which risked dropping lines a
50k-row scrollback user would have retained. Both caps (main pendingData
and the renderer output queue) now derive from one shared policy:
max(2 MB, scrollbackRows x 120 chars) — 2 MB at the 5k default, 6 MB at
the 50k max. The main side reads the setting live via getSettings; the
renderer scheduler is configured where the terminal lifecycle already
reads the scrollback setting.
Every drop now records a rate-limited crash breadcrumb with dropped and
cap sizes (terminal_output_backlog_dropped in the renderer,
terminal_pending_output_dropped in main — no pty ids, session ids can
embed workspace paths). Field drop frequency and size decide whether the
cap constants need raising, replacing theory with data (#2836, #7017).
Backlog skip notices are now cap-agnostic since the limit varies.
Co-authored-by: Orca <help@stably.ai>
* Extract breadcrumb recording into a collection-safe leaf module
Playwright loads spec imports at collection time, and e2e specs import
terminal-module constants (e.g. terminal-attention.spec.ts pulls
POST_REPLAY_MODE_RESET from layout-serialization, whose chain reaches
replay-guard). The breadcrumb import added to the terminal modules made
that chain reach crash-diagnostics.ts, whose top-level import.meta.hot
and webview-registry import crash Playwright's transform
("ReferenceError: exports is not defined in ES module scope") — every
e2e shard failed at collection before running a single test.
Move recordRendererCrashBreadcrumb into crash-breadcrumb-recorder.ts
(type-only imports, no import.meta) and point the terminal modules and
their test mocks at it; crash-diagnostics re-exports for existing
callers. Full e2e suite collects again (262 tests / 94 files); unit
suites, typecheck, lint green. No runtime behavior change.
Co-authored-by: Orca <help@stably.ai>
* Add cross-terminal pipeline benchmark (DSR-fenced throughput + latency probe)
Run inside any terminal (Orca pane, iTerm2, Ghostty, Terminal.app, VS Code)
to measure its full byte path. DSR round-trip latency at idle and under a
paced agent-TUI load, plus fenced throughput over four deterministic
fixtures. The DSR fence forces 'all bytes parsed' before the clock stops so
xterm.js-class ingest queues can't flatter the result.
First piece of the terminal performance initiative's measurement rig.
Co-authored-by: Orca <help@stably.ai>
* Add terminal performance initiative plan
Working plan for the orca-performance branch: verified architecture
findings, workstreams (baselines, #7153 validation, term-speed-2 revival
with merge-scout numbers, stall fixes, flow control, rig extensions,
utilityProcess router, telemetry), benchmark protocol, sequencing, and
baseline-relative success criteria.
Co-authored-by: Orca <help@stably.ai>
* Add cross-terminal baseline results (Orca 1.4.91 prod vs Terminal.app vs Ghostty)
Headline: Orca DSR latency under 1MB/s agent-TUI load is p50 134ms / p99 292ms
vs 0.45ms (Terminal.app) and 0.21ms (Ghostty). Idle latency is fine (0.69ms
p50) — the problem is queueing under load, not the pipeline hop. agent-tui
fenced throughput: Orca 2.0 MB/s vs Terminal.app 37 MB/s, Ghostty 78 MB/s.
Co-authored-by: Orca <help@stably.ai>
* Add pipeline-loss decomposition benches (headless xterm + daemon ingest)
Both isolate layers of the 51x agent-tui gap found in baseline-jul02:
bare @xterm/headless parses agent-tui at 103 MB/s and daemon Session
ingest (emulator + pending-output recording + fanout) at 103 MB/s —
on the byte stream the full Orca pipeline delivers at 2.0 MB/s.
Parser and daemon are exonerated; the loss is in main per-chunk
processing, delivery/ACK pacing, or renderer layers above xterm.
Co-authored-by: Orca <help@stably.ai>
* Record baseline + decomposition findings in initiative plan
Co-authored-by: Orca <help@stably.ai>
* Add dev-build orca-performance bench result (confounded: dev mode, 282-col window, 3MB fixtures)
DSR under load p50 161ms — the #7139/#7150 branch does not move the
under-load latency class. Expected in hindsight: DSR replies are ordered
within the output stream, so the metric measures output-queue depth;
cooperative drain paces input responsiveness but cannot reorder the queue.
Shrinking the queue itself (producer flow control, task 6) and raising
agent-tui throughput (task 9) are the levers for this number.
Co-authored-by: Orca <help@stably.ai>
* Record dev-build #7153 check in findings log
Co-authored-by: Orca <help@stably.ai>
* Parse-clock high-priority terminal drains instead of fixed-nap dripping
Attribution (task #9): the drain loop wrote at most 2x16KB then slept
4/16ms regardless of parse speed — an isolation bench (new
pane-terminal-output-scheduler-throughput.bench.test.ts) measures that
drip at 1.9 MB/s background / 27 MB/s foreground against xterm's
~103 MB/s parse rate, matching the baseline-jul02 end-to-end numbers
(agent-tui 2.0 MB/s in prod 1.4.91).
Fix: high-priority (visible-pane) drains now re-arm on xterm's
parse-completion callback and carry 8 writes per tick; the isolation
ceiling rises 27 -> 117.6 MB/s (parse-limited). Background cadence is
deliberately unchanged (2 MB/s drip protects the focused pane; hidden
delivery is term-speed-2's job). DRAIN_TIME_BUDGET_MS still bounds
per-tick work, preserving #7139's cooperative-drain intent.
Validation: 621 scheduler/guard/pty tests green, typecheck clean.
Co-authored-by: Orca <help@stably.ai>
* Record task #9 attribution + parse-clock fix in findings log
Co-authored-by: Orca <help@stably.ai>
* Findings: 51x loss attributed to O(tail) retained-tail redraw path in main onPtyData
Co-authored-by: Orca <help@stably.ai>
* Window the retained-tail redraw path to the cursor's reach
Attribution (findings log 2026-07-03): main's onPtyData consumed ~93% of
the event loop under an agent-TUI flood, and the dominant term was
appendNormalizedToMultilineTailBuffer + finalizeRetainedTerminalRows
materializing ~2x tail-length row objects plus a per-row trailing-space
regex on every chunk — 0.888ms/chunk at the 2,000-line cap, on every
Claude-Code-shaped frame (cursor-up + erase-below).
The multiline algorithm now runs on a suffix window sized by the chunk's
maximum upward cursor excursion (plus the inherited redraw cursor and a
safety margin); the untouched prefix is shared by reference with a cheap
last-char trailing-space check to match the reference trim. Pathological
full-height cursor-ups fall back to the unwindowed implementation, which
is kept verbatim and exported as the reference for the 500-case
differential fuzz (retained-tail-redraw-window.equivalence.test.ts).
Micro-bench at a full 2,000-line tail: 0.888 -> 0.073 ms/chunk (12x).
1,415 runtime tests green, typecheck clean.
Co-authored-by: Orca <help@stably.ai>
* Add dev bench results: parse-clock and windowed-tail fixes
Co-authored-by: Orca <help@stably.ai>
* Record windowed-tail partial win + next-cycle recipe in findings log
Co-authored-by: Orca <help@stably.ai>
* Findings: remaining whale is the per-chunk blocked-reason check (~85% of onPtyData post-fix)
Co-authored-by: Orca <help@stably.ai>
* Throttle the terminal wait-blocked check off the PTY hot path
Post-windowed-tail attribution (findings log 2026-07-03): the blocked-
reason complex — two full-tail buildTerminalWaitText builds plus
toLowerCase and multi-pattern scans per chunk, existing only to stamp
waitBlockedAt — consumed ~85% of onPtyData's remaining cost (~700-790ms/s
under an agent-TUI flood).
The check now runs at a 50ms cadence over coalesced chunks (PTY chunk
boundaries are arbitrary, so coalescing preserves semantics), with a
trailing-edge timer so burst-final state is always evaluated, and an
immediate bypass when the incoming chunk (plus a 31-char split carry)
contains a prompt keyword — so actionable-prompt stamping stays
per-chunk-immediate while keyword-free flood frames skip the complex
entirely. Previous wait text is cached per pty instead of rebuilt, and
state is cleared at both pty teardown sites.
1,415 runtime tests green (including the cross-chunk prompt test, which
exercises the keyword bypass), typecheck and lint clean.
Co-authored-by: Orca <help@stably.ai>
* Findings + results: three stacked fixes unlock the pipeline (agent-tui 16x, DSR-load p50 161->18.8ms in dev)
Co-authored-by: Orca <help@stably.ai>
* Add producer flow-control design to initiative plan
Co-authored-by: Orca <help@stably.ai>
* Findings: revival branch green but perf-gated — daemon Session ingest regressed 103->40-48 MB/s (chain emulator restructure); merge blocked until blockedfix parity
Co-authored-by: Orca <help@stably.ai>
* Pre-filter daemon OSC/mouse scanners for introducer-free chunks
Skips the scan-tail copy and full-chunk walks when a chunk cannot contain
an OSC or private-mode sequence (single native includes() checks), with
split-sequence correctness preserved via explicit tail retention. Strictly
positive micro-optimization on the daemon per-chunk path; 641 daemon tests
green (1 pre-existing WSL failure unrelated).
Co-authored-by: Orca <help@stably.ai>
* Retract confounded daemon conviction; mandate load-controlled A/B protocol for the revival merge gate
Co-authored-by: Orca <help@stably.ai>
* Record A/B gate pass in findings log; add A/B result JSONs
Co-authored-by: Orca <help@stably.ai>
* Add producer-side PTY flow control (watermarks + protocol v19)
Main now pauses the actual PTY when a pane's renderer-pending backlog
crosses the 256KB high watermark and resumes once it drains below the
32KB low watermark (wide hysteresis band so a draining queue cannot flap
pause/resume per flush slice). node-pty pause() stops the pty fd read, so
the kernel/ConPTY buffer fills and a flooding shell blocks on write —
flood-induced buffered lag becomes shell blocking instead of unbounded
main-process buffering (terminal-performance-initiative §5).
Transport: new fire-and-forget pausePty/resumePty daemon notifications
(protocol v19; 18 added to PREVIOUS_DAEMON_PROTOCOL_VERSIONS), routed
DaemonServer -> TerminalHost -> Session -> subprocess pause()/resume().
LocalPtyProvider pauses node-pty directly. Router/degraded providers
forward; IPtyProvider gains optional pauseProducer/resumeProducer.
Safety invariants:
- Lost-resume failsafe: daemon Session auto-resumes 5s after a pause with
no matching resume; main re-asserts the pause at most once per 5s while
still above the high watermark, so a lost resume can never wedge a shell
and a sustained flood stays throttled.
- Resume on every teardown path: Session kill/exit/dispose/detach; main
releases on pty exit and on window-destroyed bookkeeping wipes; the
adapter owes paused sessions a resumePty on the next connect after a
socket drop.
- Providers without support (SSH relay, legacy protocol <= v18) no-op
silently, and the scrollback-scaled pending-output cap still bounds
main memory when pause is unavailable.
- Kill switch: PRODUCER_FLOW_CONTROL_ENABLED in ipc/pty.ts flips the
whole mechanism off in one line.
daemon-errors.ts is split out of types.ts to stay under the max-lines cap.
Tests: watermark transitions/hysteresis/re-assert (controller unit),
lost-resume failsafe + resume-on-kill/exit/dispose/detach (session),
notification routing + v18 gating + reconnect owed-resume (adapter),
direct pause/resume (local provider), and a flood test asserting pause
fires once, pending stays bounded at HIGH + one chunk, and resume fires
once after drain (ipc/pty).
Co-authored-by: Orca <help@stably.ai>
* Findings: flow control merged; definition-of-done accounting; prod verification re-scoped to packaged RC
Co-authored-by: Orca <help@stably.ai>
* Fix stray brace from revival merge in long-table-scroll-restore e2e spec (broke e2e transform in CI)
Co-authored-by: Orca <help@stably.ai>
* Prod verdict: v1.4.121-rc.0 bench — DSR-load p50 134->18.6ms (7.2x), agent-tui 2.0->11.2 MB/s, idle at Terminal.app parity; pipeline now cadence-bound
Co-authored-by: Orca <help@stably.ai>
* Recover terminal output delivery after system sleep
Root cause: main gates every pty:data send on a global + per-PTY
in-flight counter that only renderer ACKs decrement. If ACKs are lost
across a system suspend, the counters pin at the cap and every PTY —
old and newly created — is silently gated forever while output piles up
in pendingData. A focus-preserving display wake also fires no renderer
focus/visibilitychange events, so terminal wake recovery (and the WebGL
context-loss latch clear) never runs. Only a renderer reload recovered.
Three fixes:
- ACK-stall watchdog (src/main/ipc/pty.ts): if sends stay gate-blocked
for 10s with zero ACK progress while the renderer webContents is
alive, warn once, reset the in-flight delivery counters, and flush
held pendingData. Armed lazily on the first gate-blocked send and
disarmed by every ACK, so it can never fire under healthy heavy load.
- Renderer lifecycle reset now also zeroes the in-flight counters — a
reload/navigation destroys the renderer dispatcher, so outstanding
ACKs can never arrive and stale counters would gate the new renderer.
- System-resume wake IPC: main relays powerMonitor 'resume' as
system:resumed to live windows (plus forceRepaint); preload exposes
ui.onSystemResumed; the terminal wake-recovery hook runs the same
recovery path as window focus/visibilitychange.
Co-authored-by: Orca <help@stably.ai>
* VS Code head-to-head: Orca beats/ties 5 of 6 metrics (16x idle, 5x styles-stress, better p99); load p50 gap attributed to ACK window + timer-clamped drain cadence
Co-authored-by: Orca <help@stably.ai>
* Schedule zero-delay terminal drains via MessageChannel
Chromium clamps nested setTimeout(0) to ~4ms, stacking dead gaps onto
every parse-clocked drain tick; the explicit 4ms high-priority re-arm
interval added more. A posted message is still a macrotask — input and
paint are serviced between posts — so cooperative yielding survives
without the clamp. Generation-tokened cancellation; vitest keeps the
timer path (fake timers can't advance channel posts) plus a real-timer
smoke test for the channel path. Standing-queue target: VS Code's ~7ms
class (measured us 18.6ms, them 7.18ms, same rig).
Co-authored-by: Orca <help@stably.ai>
* Cut daemon and main PTY batch windows 8ms -> 2ms
At 9% pipeline utilization the DSR-under-load latency is fixed batching
windows, not queue depth (proved by the MessageChannel drain lever
moving nothing). Both hops charged an expected half-window per chunk;
2ms keeps burst coalescing at negligible IPC overhead (~500 msgs/s
worst case vs MB/s payloads).
Co-authored-by: Orca <help@stably.ai>
* Findings + tests: batch windows were the DSR-load gap (19->8.0ms dev); timing tests updated to 2ms windows
Co-authored-by: Orca <help@stably.ai>
* Fix PR CI and guard resume relay during shutdown
Co-authored-by: Orca <help@stably.ai>
* Chain e2e specs 6/6 green — gate x drain validation debt paid
Co-authored-by: Orca <help@stably.ai>
* Replace ack-stall watchdog with cumulative ACKs + solicited delivery resync
Design review: the 10s blind-reset watchdog decided correctness from a
wall-clock threshold. Rework piece 1 into a deterministic two-part design
(pieces 2 and 3 — lifecycle-reset counter zeroing and powerMonitor wake
IPC — are unchanged):
- Cumulative ACKs (TCP-style): the renderer dispatcher now tracks a
monotonic per-pty total of processed chars (terminal-pty-ack-gate) and
sends it on every ACK alongside the legacy per-chunk delta. Main keeps
per-pty sentChars/ackedChars and max-merges received totals — idempotent
and reorder-tolerant, so a lost ACK self-heals when any later ACK
arrives instead of becoming permanent in-flight debt. Provider
(SSH/daemon) backpressure is credited only the derived delta, clamped,
never negative. Main tolerates both payload shapes keyed by field
presence (dev hot-reload can mix renderer/main versions); totals reset
on pty exit and renderer lifecycle reset on both sides.
- Solicited resync (replaces the blind reset): when new pty data arrives
while that pty's delivery is fully gated and no probe is outstanding,
main sends pty:requestDeliveryResync; the renderer replies with its
cumulative totals and main reconciles via max-merge, then flushes held
pendingData. Event-triggered, verified-state recovery — no wall-clock
threshold decides correctness. The only timer is a 5s request/response
hygiene timeout that clears the outstanding flag and logs one
diagnostic warn per silent streak; it never mutates counters (a
renderer that cannot answer has dead IPC — reload is the only cure).
The 10s corrective watchdog is deleted.
Co-authored-by: Orca <help@stably.ai>
* Starting point: prior agent's garble differential fuzz harness
Three files recovered (were untracked) from a prior agent killed by API
outages, plus a trivial curly-brace lint fix in the op dispatcher so the
pre-commit hook passes:
- src/shared/agent-tui-ansi-fuzz-stream.ts (seeded agent-TUI byte-stream gen)
- src/shared/terminal-restore-parity-fixture.ts (renderer-parity fixture)
- src/main/daemon/headless-emulator-fidelity.fuzz.test.ts (suite 1: differential
HeadlessEmulator vs @xterm/headless reference on identical bytes)
Co-authored-by: Orca <help@stably.ai>
* Suite 1 findings: two new serialize round-trip bugs (B bold-loss, C cursor)
Scanned seeds 1..2000. Beyond the pre-documented serialize wrap-null-cell bug
(A, 27 seeds, tolerated), the fuzz surfaced two NEW real @xterm/addon-serialize
0.15.0-beta.287 round-trip defects, both of which garble a revealed hidden pane:
- Bug B (seeds 435, 770, 1321): serializing a dim cell followed by a bold-only
cell emits \x1b[1;22m; SGR 22 clears bold too, so restored bold is lost.
Minimal repro: '\x1b[2mA\x1b[22m\x1b[1mB' -> restored 'B' loses bold.
- Bug C (seeds 454, 1696): a final content row filled to the right margin leaves
xterm wrap-pending; the serializer's relative cursor restore lands one column
short. Minimal repro: '0123456789\x1b[3;5H' at cols=10 -> cursor x=3 not x=4.
Both isolated to pure serializer replay (no Orca preamble), confirming upstream.
Parity fixture verified faithful to the renderer pane's buffer options. Each is
pinned as a standalone it.skip repro; full evidence + classification in
notes/garble-fuzz-divergences.md. Seed 113 (handoff's DECSC/DECRC case) does not
diverge on the current harness. No production code changed.
Co-authored-by: Orca <help@stably.ai>
* Add perf prerelease update check modifier
Co-authored-by: Orca <help@stably.ai>
* Suite 2: hidden-reveal seq-reconciliation fuzz + two new snapshot bugs (D, E)
Property-tests the reveal seq-reconciliation byte-stitch (getChunkDataAfterSnapshot
/ reconcileChunkAgainstRestoredSnapshot in pty-connection.ts), mirrored exactly:
N=200 seeded hide/reveal scenarios with a rich agent-TUI hidden prefix snapshot
and an append-only racing tail, chunked with seq/rawLength meta, seq-domain
restarts, unmetered chunks and droppedOutput markers. Asserts snapshot-at-S +
reconciled tail == snapshot-of-everything (seq-neutral) and == always-visible
(end-to-end). Runtime ~5s at 200; FUZZ_ITERATIONS override documented.
Two NEW real snapshot-limitation garbles found while building it, both distinct
from suite 1's serialize bugs and pinned as standalone it.skip repros:
- Bug D: the DECSC saved-cursor register is not serialized. A hidden TUI that
saves the cursor (ESC 7 / CSI s) and restores it on reveal (ESC 8 / CSI u)
lands the restore at home. Repro: 'AB\x1b7\x1b[4;10HCD' + '\x1b8X' -> 'XB' vs 'ABX'.
- Bug E: a snapshot taken mid-escape-sequence (a PTY read split an escape) drops
the partial sequence (it's parser state, not buffer), so the tail's
continuation renders literal. Repro: 'AB\x1b[3' + 'mCD' -> 'ABmCD' vs 'ABCD'.
Fired on ~24% of the corpus (tolerated + counted via prefixEndsMidSequence).
The append-only-tail design isolates seq reconciliation from these and the Bug C
cursor cascade. Full evidence + fix directions in notes. No production changes.
Co-authored-by: Orca <help@stably.ai>
* Suite 3: 25-cycle park/reveal drift e2e test
Extends terminal-hidden-view-parking.spec.ts with a deterministic 25-cycle
park->reveal test on a static rich alt-screen TUI frame (box drawing, SGR
colors, wide CJK/emoji). Baselines against the frame after the first snapshot
restore (so both sides pass through identical machinery — the alt-screen restore
correctly drops normal-buffer scrollback, which is contract not garble), then
asserts every subsequent reveal reproduces it byte-for-byte with no accumulated
drift and no hidden-skip banner. Exercises the real renderer teardown +
HeadlessEmulator snapshot restore + PTY reattach path the fuzz suites model in
isolation. Passes in ~29s (electron-headless, workers=1).
Co-authored-by: Orca <help@stably.ai>
* Fix two serialize round-trip bugs garbling hidden-terminal snapshot restore
BUG B (addon patch): @xterm/addon-serialize's SGR diff emitted bold/dim set
params before the shared intensity reset 22, so "1;22" wiped a freshly set
bold and a bare "22" dropped a still-set bold/dim. Patched via pnpm
patchedDependencies (config/patches) to diff bold+dim as one intensity
group with the clearing 22 emitted first. Other flag pairs (4/24, 3/23,
7/27, ...) have dedicated resets and were verified unaffected.
BUG C (Orca-side hardening): the addon restores the cursor with relative
moves computed from where it assumes replay leaves the cursor; a final row
filled exactly to the right margin leaves replay wrap-pending and the
restore lands one column short. New shared
serializeWithAbsoluteCursor appends an absolute CUP from the source
terminal's authoritative cursor at every restore/replay serialize site
(daemon/runtime HeadlessEmulator.getSnapshot, renderer mobile snapshot
serializer, shutdown layout capture). It skips empty snapshots and
wrap-pending sources so it never changes already-correct behavior.
Round-trip repros + non-regression coverage in
src/main/daemon/terminal-snapshot-serialize-roundtrip.test.ts (verified
failing with the fixes stashed). buildRehydrateSequences extracted to its
own module to keep headless-emulator.ts under the max-lines budget.
Co-authored-by: Orca <help@stably.ai>
* Gates: tolerate+count Bugs B/C in deep mode; drop inverse from reconciliation tail
- Fidelity suite: add snapshotHasSelfCancellingBoldReset (Bug B) and
isMarginWrapPendingCursorOffByOne (Bug C) predicates so the corpus tolerates +
counts them like Bug A. FUZZ_ITERATIONS=2000 is now green (~113s) and fails
only on genuinely new divergences; each tolerance keeps its <50% degeneracy
guard. Default 300 unchanged (~17s).
- Reconciliation suite: drop SGR 7 (inverse) from the append-only tail. Inverse
marks trailing blanks with an inverse-fg the serializer round-trips slightly
differently by capture depth — a Bug-B-class serialize nuance, not seq
reconciliation. FUZZ_ITERATIONS=1000 is now green; default 200 unchanged.
- Notes updated: every bug class is both pinned (skipped repro) and tolerated in
its corpus; combined default runtime ~19s.
Regex uses String.fromCharCode(27) to stay oxlint no-control-regex clean.
Co-authored-by: Orca <help@stably.ai>
* Keep RC update checks off perf prereleases
Co-authored-by: Orca <help@stably.ai>
* Fix snapshot DECSC register loss (Bug D) and mid-escape boundary drop (Bug E)
Bug D: the serialized screen cannot carry the VT100 DECSC saved-cursor
register, so a hidden ESC 7 followed by a post-reveal ESC 8 restored to
home and clobbered live cells. The snapshot epilogue now re-saves at the
source's saved position before the final absolute CUP
(readSavedCursorRegister + serializeWithAbsoluteCursor; the active
buffer's own register, so alt screens carry theirs). Position-only by
design; never-saved terminals are left untouched.
Bug E: a PTY read ending mid-escape leaves the sequence in the emulator's
parser, so serialize dropped it and the racing tail's continuation bytes
rendered literally after reveal (~24% of the fuzz corpus). The emulator
now tracks the unparsed trailing partial at ingest
(terminal-partial-escape-tail.ts, committed post-parse like the mouse
mirror) and ships it as TerminalSnapshot.pendingEscapeTailAnsi.
applyMainBufferSnapshot writes it LAST, after POST_REPLAY reset — any
later ESC would abort the dangling sequence. Seq accounting is unchanged:
the tail is a suffix of bytes the snapshot seq already counts, so
reconcile slicing needs no adjustment.
Fuzz suites: unskip the Bug B/C repros (fixed on this branch) and the new
D/E repros; remove the B/C/E tolerance predicates so regressions fail
loudly. Only Bug A (upstream wrap null-cell) stays tolerated + counted.
Green at FUZZ_ITERATIONS=2000 (fidelity) and 1000 (reconciliation).
Co-authored-by: Orca <help@stably.ai>
* Count suffixed RC tags (rc.N.perf) in the shared rc counter — second suffixed cut collided with the first
Co-authored-by: Orca <help@stably.ai>
* Classify suffixed rc tags (rc.N.perf) as rc telemetry identity in release builds
The build-identity guard only knew vX.Y.Z and vX.Y.Z-rc.N, so suffixed
perf RCs cut fine but every platform build refused the tag and the
releases published empty.
Co-authored-by: Orca <help@stably.ai>
* Cut the hidden-restore flood feedback loop (A) + query carve-out on drops (B)
(A) Under a foreground flood, the hidden-output-restore loop re-fetched
snapshots endlessly: each synchronous applyMainBufferSnapshot starved ACK
processing, main pinned at the in-flight cap, dropped at the pending cap,
and every droppedOutput/modelRestoreNeeded marker re-armed another
restore until the flood ended (rc.7.perf DSR timeouts).
- Restore loop: a foreground live-chunk queue overflow now abandons the
restore immediately (the stream is outrunning snapshot fetch+replay),
with a 3-iteration hard cap + lifecycle warn as backstop.
- Re-arm gate: drop markers/sentinels and reconcile seq-gaps on a visible
pane during its own in-flight/just-abandoned restore no longer re-arm;
live bytes write through and ONE deferred repaint (2s after the last
backpressure signal) heals the gap. Hidden-pane gate semantics are
unchanged.
- Query salvage: discarding queued restore bytes (overflow/refetch) now
extracts DSR/CPR/DA/OSC-color queries and replays them to xterm so
replies still flow.
(B) Main-side: dropOversizedPendingPtyData carves reply-eliciting query
sequences out of the dropped buffer (and out of post-drop latched data,
bounded) and ships them on the droppedOutput sentinel, so DSR probes
survive bulk drops. Query scanning moved to
src/shared/terminal-reply-query-extraction.ts, shared verbatim with the
renderer's hidden-startup query extraction.
Co-authored-by: Orca <help@stably.ai>
* ACK terminal output at parse-drain, not dispatcher enqueue (C)
The renderer credited main's per-PTY in-flight window the moment a
pty:data chunk entered the dispatcher, so the 512KB window meant "bytes
received", never "bytes parsed". Under flood the renderer write queue
grew unbounded behind instant ACKs; main saw no backpressure, crossed
the pending cap, and bulk-dropped output (rc.7.perf DSR timeouts).
Crediting is now parse-deferred: each delivery carries a fire-once
credit (deliverPtyDataWithDeferredAck); the pane's first scheduler write
claims it (writeTerminalOutput.ackCredit) and the output scheduler fires
it when the bytes are consumed — after terminal.write in the
parse-clocked drain, or on ANY discard path (backlog cap replacement,
discardTerminalOutput, disposed-terminal drops, flush recovery).
Deliveries that never reach the scheduler (reconcile drops, restore
queueing, pre-mount eager buffer) settle at handler return, so the
invariant holds: every delivered chunk credits exactly once, parsed or
discarded. E2E ack-gate hold/release and delivery-resync semantics are
unchanged (all crediting still routes through ackPtyData).
Main-side equilibrium: with ACKs at parse cadence, in-flight becomes
true backpressure — pendingData stays near the 256KB producer-pause
watermark, far under the >=2MB drop cap, so bulk floods block the shell
(node-pty pause) instead of dropping.
Co-authored-by: Orca <help@stably.ai>
* Synthesize salvaged query replies directly instead of replaying into xterm
The 10MB dev bench proved the write-back salvage insufficient: a
pending-cap drop always triggers a snapshot restore, whose replay guard
swallows xterm auto-replies and whose discardTerminalOutput races away
still-queued query writes — the salvaged DSR died both ways and the
fence still timed out.
Salvage now answers directly on the input path (immune to both): CPR
(CSI 6n) from the live buffer via transport.sendInput, DA1 with the
renderer's canned response, OSC color probes via the existing direct
responder. Rare queries (DECRQM, DA2) keep the best-effort xterm
replay.
Co-authored-by: Orca <help@stably.ai>
* Untrack branch-added bench result JSONs (20 files); keep numbers in the findings log
Files stay on disk; main's 7 pre-existing results are untouched.
Co-authored-by: Orca <help@stably.ai>
* Branch guide: document merge-not-rebase sync strategy and conflict pattern
Co-authored-by: Orca <help@stably.ai>
* Merge origin/main (#7316 tab-strip click-vs-drag fix); adapt #7290 recovery-reload tests to this branch's dual did-finish-load listeners
The three tests grabbed the FIRST did-finish-load listener; on this branch
the renderer delivery-gate reset registers before the orphan sweep, so the
sweep tests exercised the wrong handler (one failing, two vacuously green).
They now fire all listeners like a real reload.
Co-authored-by: Orca <help@stably.ai>
* Fix branch CI lint: split pane-interaction functions out of artificial-opencode-terminal-load.spec (815>800 lines), modernize perf-html-report script
No max-lines disable per repo rules; extracted to
artificial-opencode-pane-interactions.ts. toReversed() and
import.meta.filename replace reverse()/fileURLToPath.
Co-authored-by: Orca <help@stably.ai>
* Fix Windows update-relaunch killing the live terminal daemon
On a Windows update relaunch the daemon can be wedged past every RPC
budget (final checkpoint flush + installer/AV disk pressure), so the 3s
health check AND the 5s session-list hello both time out while sessions
are still alive - and the launcher failed closed, killing the daemon and
every terminal session it owned.
- Adopt an unresponsive daemon whose pipe still accepts a raw
connection; a new rejected health state keeps replacing daemons that
answered and refused the handshake (never adoptable).
- Give Windows pid files a real startedAtMs (daemon self-reports it in
the ready IPC message) and verify it via CIM CreationDate piggybacked
on the existing command-line query, so the pid-recycling guard is no
longer inert on win32.
- Only delete legacy daemon pid/token files when the pid-file process is
provably dead; deleting a live daemon''s token made its sessions
permanently unadoptable after a protocol bump.
- Capture agent resume records every 60s in the renderer (skipping
unchanged records) so hard kills still leave a fresh resume record.
* Heal blank terminals when main→renderer push delivery dies (renderer-pull delivery watchdog)
Field evidence (v1.4.121-rc.0 debug snapshot, 2026-07-06): a wedged window
held 530,115 un-ACKed in-flight chars — one PTY pinned at the 512KiB per-PTY
high water plus a fresh terminal's 245-char prompt that was sent and never
consumed — while the user ran the snapshot over invoke from that same window.
Main→renderer push delivery (pty:data and every sibling channel) was dead;
renderer→main→renderer invoke was alive. Upstream precedent for
one-directional IPC death: electron#37067 (suspected Mojo pipe disconnect,
stalled as need-info). Every terminal goes blank, new terminals are born
blank, and only a renderer reload recovered.
The existing recovery layers cover the OTHER variants of this bug family and
structurally cannot reach this one:
- The xterm write-pipeline sync-throw guards, output-buffer caps, and
probe-certified replay-guard release (#7150 family) run only after bytes
arrive in the renderer — here they never do. (The pending cap did work as
designed in the field: ~2.1MB pendingDroppedChars, bounded main heap.)
- Cumulative ACKs self-heal lost ACK messages and the solicited delivery
resync reconciles verified totals (4647df86a; #7260 on main) — but the
resync probe, the powerMonitor wake relay, and the droppedOutput restore
markers all ride main→renderer push, the direction that is dead. The
probe's unanswered path deliberately only logs.
This adds the missing lane, renderer-initiated and ridden entirely over
invoke — the direction the field snapshot proved alive:
- terminal-delivery-watchdog.ts: 15s heartbeat, free while output flows.
Hot-path cost is one Map upsert per received chunk; a tick does no IPC
unless the terminal plane was silent for the whole interval and a PTY
still expects delivery. Two consecutive silent ticks with main reporting
ACK-starved in-flight confirm the wedge; heals are one-shot per 60s
cooldown so a persisting wedge cannot repaint-storm.
- pty:reportRendererDeliveryState (invoke): always max-merges the renderer's
cumulative processed totals (a free extra repair lane for the lost-ACK
variant); with heal:true — and only after main has itself seen ≥10s of ACK
silence — writes off bytes the renderer provably never received
(received ≤ acked < sent; a received-but-unparsed backpressure window is
never written off), drops that PTY's pendingData (snapshot covers
everything ≤ markerSeq, hidden-drop parity), credits provider flow
control, and returns restore markers in the reply.
- The renderer re-attaches all push listeners (cures a detached-listener
variant outright; a safe no-op against a dead channel) and routes the
pulled markers through the existing pty:modelRestoreNeeded machinery —
panes repaint from the main-owned buffer snapshot with zero push
delivery involved.
- Field discrimination built in: the heal warn logs
ipcRenderer.listenerCount('pty:data') (listener detached vs channel dead)
with the full delivery snapshot, so the next occurrence names the root
cause without asking the user to run anything in a console.
Repro harness: the exposeStore-gated __terminalDeliveryWatchdog hook
blackholes pty:data ahead of the dispatcher — the field failure in
miniature (no receive count, no ACK credit, no dispatch).
terminal-push-delivery-loss-recovery.spec.ts proves the wedged output
repaints while the blackhole is still engaged and live flow resumes after
release, with no reload. Unit suites pin the watchdog state machine
(zero IPC under flow, two-tick confirm, cooldown), the dispatcher reattach
seam, and the main-side write-off semantics.
Perf: nothing added to main's send/flush path; the renderer data path gains
one integer/Map update per chunk; idle cost is one ~100-byte invoke per 15s
only during total terminal silence. Terminal perf e2e suite (typing
latency, redraw freeze, output scheduler, hidden TUI restore, artificial
opencode load) passes on this change; no watchdog activity occurs under
ack-gate pressure scenarios because receive-progress gates the heartbeat.
Co-authored-by: Orca <help@stably.ai>
* Expose the hidden-yet-visible delivery-gate contradiction in the debug snapshot
The v1.4.124-rc.2.perf blank-terminal field snapshot showed a different
state than the v1.4.121 transport wedge: no delivery gating at all
(ackGatedFlushSkipCount 0, in-flight 38KB, far under every cap) but TWO
ptys hidden-delivery-gated with 78MB dropped as hidden. The aggregate
counters cannot say whether the pane the user was staring at was one of
the gated ones — the one number that separates "normal background
dropping" from "main is starving a visible pane because the reveal
unmark never fired".
Add hiddenDeliveryGatedVisiblePtyCount / hiddenDeliveryGatedActivePtyCount
(overlap of the gate's hidden set with the renderer's visible/active
reports — a contradiction that must be zero) to the delivery debug
snapshot, and a once-per-minute warn when hidden-gated bytes are dropped
for a pty the renderer reports visible or active, with the full snapshot
attached. Zero cost outside the debug read and the already-dropping path.
Co-authored-by: Orca <help@stably.ai>
* Unlatch the hidden-delivery gate when user input disproves a stuck document.visibilityState
macOS occlusion tracking can wedge document.visibilityState at 'hidden'
after display sleep and never fire another visibilitychange. The hidden-
delivery gate then keeps dropping renderer-bound bytes for panes the user
is looking at (field snapshot 2026-07-06, v1.4.124-rc.2.perf: 78MB dropped
across 2 pane-level-visible ptys with a fully healthy transport), and every
recovery path (window focus, system-resume relay, backlog recovery) re-ran
syncHiddenRendererPtyDelivery only to recompute the same stale predicate —
nothing could ever clear the gate. The user sees a frozen terminal; typing
echo is dropped in main; only a reload recovers.
Real user input while the document claims hidden is a physical
contradiction: keystrokes and clicks only reach a focused, on-screen
window. stale-document-visibility.ts latches that proof, runs each pane's
existing visibilitychange resync (gate unhide + hidden-output snapshot
restore), and hands authority back to the occlusion tracker on the next
genuine visibilitychange. No timers; the failure bias is safe — a wrong
latch can only restore pre-gate delivery cost, never drop bytes. Hot path
unchanged: the foreground predicate still returns on the same single
comparison while the document is visible.
tests/e2e/terminal-stuck-occlusion-recovery.spec.ts pins the wedge
(visibilityState pinned hidden -> output dropped, not painted; the
hiddenDeliveryGatedVisiblePtyCount field discriminator reads >0) and the
recovery (one Shift keypress repaints the missed output from the main-owned
snapshot, no reload, while visibilityState still reads hidden). Negative
control verified: the spec fails without this fix. Typing-latency perf
gate passes; terminal-pane unit suites 366/366.
Co-authored-by: Orca <help@stably.ai>
* Add a one-paste terminal freeze report: __orcaTerminalFreezeReport()
Every field report of the frozen-terminal family so far has needed
follow-up asks (console output, main logs, second snapshots) because each
capture showed one process's counters at one instant. This makes a single
DevTools command sufficient: `await window.__orcaTerminalFreezeReport()`
returns renderer state (document.visibilityState + the stale-visibility
override, pty:data listener count, delivery-watchdog totals), main's debug
snapshot extended with a per-pty delivery table (sent/acked/pending, hidden
vs visible-set membership, last send/ACK ages, window focus flags, power
suspend/resume ages, app version), and bounded breadcrumb rings from BOTH
processes recording the transitions that matter: gate marks/unmarks,
visibilitychange and stale-visibility latches, watchdog stalls and heals,
restore markers, heal write-offs, and renderer lifecycle resets (so "user
already reloaded" is visible in the history).
Costs stay off the data path: breadcrumbs record only rare transitions into
a 100-entry ring with same-kind coalescing (a flood costs one slot per
second); the per-pty table is built only when the snapshot is read; the
per-send bookkeeping adds one Date.now() to existing accounting writes. Pty
ids are redacted to their `@@` suffix because daemon session ids embed
worktree paths. The report assembles over invoke IPC — the direction proven
alive in every observed wedge — and a failing invoke is captured as data
instead of sinking the report.
The stuck-occlusion e2e now also pins the report end-to-end: after the
wedge + keystroke recovery, the report must carry the stale-visibility
latch and gate transitions in the renderer ring, gate-mark/unmark in main's
ring, and a populated per-pty table. Suites: pty.test.ts 258, terminal-pane
1705, shared ring 5; typing-latency perf gate passes.
Co-authored-by: Orca <help@stably.ai>
* perf(daemon): keep-tail thin hidden panes' stream so agent floods never bury typing (STA multi-workspace lag)
Hidden panes are exempt from pendingData flow control (main gate-drops
their bytes after ingestion), so N background agents ran unbounded ahead
on the one shared daemon->main stream socket (measured 192MB user-space
backlog) and visible-pane echo waited FIFO behind it — typing appeared
seconds late whenever several agents burst on a loaded machine
(8x512KB/s + 12 CPU spinners: p50 293ms fix-off; 12x1MB/s: 6.1s).
Mechanism (replaces producer pacing — no reveal catch-up, ever):
- Shallow socket write gate (128KB) + per-session fairness bypass bounds
echo latency by construction; kernel-flush refill sentinel keeps held
bulk draining at full speed (drain-only refill capped at ~8MB/s).
- Backgrounded sessions' queued output is keep-tail dropped (newest
512KB kept, in-order dataGap replaces the middle); a ~2MB GLOBAL
budget shrinks per-session keep-tails (floor 64KB) so a worktree
switch never waits behind the aggregate. Reply-eliciting query bytes
(DSR/DA/OSC probes) are salvaged from dropped spans.
- Notifications are structurally lossless: the daemon runs the same
shared scanners main uses (bell/OSC 133/pr-link/2031) over every byte
BEFORE drop decisions and relays facts in byte order; ordered
background markers hand scan authority back and forth, seeded with the
emulator's partial escape tail so a sequence split across the handoff
neither phantom-fires nor goes missing. Titles/agent-status stay
main-side (kept-tail convergent).
- Main: background = hidden AND no remote view subscriber (a live
mobile/web view is never thinned); on dataGap main resets cross-chunk
parse carries, drops the headless mobile mirror, and reuses the
hidden-drop model-restore marker.
Wire: three new stream events, tolerated within protocol v19 (old mains
ignore unknown events; old daemons never see the trigger). Kill
switches: ORCA_DAEMON_BACKGROUND_STREAM_DROP=0,
ORCA_DAEMON_SHALLOW_SOCKET_GATE=0.
A/B (pnpm bench:multi-workspace-typing): 8x512KB/s + 12 CPU workers
p50 293ms/p90 647ms -> 15/21ms (= baseline); 12x1MB/s 6,146ms -> 20ms;
light loads unchanged; zero missing echoes. Latin hidden-restore e2e
green (probe-verified aggregate-drain root cause). New deterministic
repro harness: tests/e2e/terminal-multi-workspace-typing-latency.spec.ts
+ CPU pressure workers.
Co-authored-by: Orca <help@stably.ai>
* diag(terminal): breadcrumb WebGL context-loss/atlas + wake triggers into freeze report
Silent instrumentation (memory ring only, no new console lines) so the next
post-wake garble report attributes itself. Adds:
- shared/terminal-webgl-diagnostics.ts: lib-safe sink so pane-webgl-renderer
(lib) can record without importing the components-layer ring; wired to the
ring in terminal-freeze-breadcrumbs.
- webgl-context-loss crumb at onContextLoss, webgl-atlas-reset crumb at the
atlas registry reset — the pair that distinguishes 'atlas corrupted' from
'missed repaint'.
- wake-recovery:<source> crumb (focus/visibilitychange/system-resumed) with the
clearGlyphAtlases decision; source in the kind so distinct triggers don't
coalesce.
- per-pane WebGL state (getAllPaneRenderingDiagnostics) in the freeze report.
Gates: typecheck 0 errors; terminal suites 328 files pass; oxlint clean.
Co-authored-by: Orca <help@stably.ai>
* fix(lint): use Number.parseInt/parseFloat in terminal-view-attributes
oxlint unicorn(prefer-number-properties) flagged 24 global parseInt/parseFloat
calls in the terminal-view-attributes feature (
|
||
|
|
03a6737089 | Fix Windows terminal daemon death after workspace close (#8140) | ||
|
|
da0f03fc2a | Sign Windows inner binaries during release (fail-open two-request SignPath flow) (#7866) | ||
|
|
20832ef36f | ci(daemon): fail builds, packaging, and CI when the terminal daemon cannot start — and stop failing silently (#7849) | ||
|
|
c256a5a417 | Windows terminal update-survival (single consolidated PR) (#7538) | ||
|
|
f311307560 |
Add max-lines ratchet CI gate to block new line-limit bypasses (#7608)
oxlint already fails any file over max-lines that is not suppressed, so the only way to grow past the budget is to add an eslint/oxlint-disable max-lines comment or a per-file max-lines bump in mobile/.oxlintrc.json. This adds a CI gate that freezes the current set of suppressions (config/max-lines-baseline.txt, 355 grandfathered entries) and fails the build when a NEW one appears — with a loud, actionable message pointing at 'split the file'. Existing oversized files are untouched; the baseline may only shrink (pnpm check:max-lines-ratchet --prune). Wired into the root lint script and as a dedicated pr.yml step. Unit-tested (15 cases) and verified against all three failure paths + clean-tree pass. Co-authored-by: Orca <help@stably.ai> |
||
|
|
6a4b89785c |
revert: back out the Windows terminal update-survival chain (#7421→#7499) (#7505)
* Revert "Preload the daemon windowsHide shim via --require; wrap promisify custom (#7499)" This reverts commit |
||
|
|
f4faafa987 |
fix(daemon): relocate daemon host image out of the install-dir kill zone (#7473)
Co-authored-by: Neil <neil@stably.ai> Co-authored-by: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com> |
||
|
|
b3f86813b3 |
fix(ci): register PSGallery via PSResourceGet for Windows SignPath install (#7379)
Register-PSRepository -Default failed inside the legacy nuget.exe provider with "Missing option value for: '-source'" on the hosted windows-2022 image, leaving PSGallery unregistered so the subsequent Set-PSRepository threw "No repository with the name 'PSGallery'". Prefer Microsoft.PowerShell.PSResourceGet (bundled with PowerShell 7.4+), which has PSGallery registered by default and avoids that code path; fall back to PowerShellGet with an explicit NuGet-provider bootstrap only when PSResourceGet is unavailable. Co-authored-by: Orca <help@stably.ai> |
||
|
|
5b5ce6cef8 |
Harden reliability gate policy and register merged terminal regression tests as gates (#7295)
Co-authored-by: Orca <help@stably.ai> Co-authored-by: brennanb2025 <delta-eng@stably.ai> |