mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 08:02:32 +00:00
2eb3e11327a3188eaa2b6c43608e32661c4eec5a
1561
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2eb3e11327 | fix(terminal): make close and handles incarnation-stable (STA-4327) (#14590) | ||
|
|
9367169888 |
refactor(tests): split every oversized test file off the max-lines suppression list (#14728)
* refactor(tests): split oversized test files off the max-lines suppression list Every `*.test.ts`/`*.spec.ts` that carried an `eslint/oxlint-disable max-lines` directive is now split into focused, behavior-scoped suites that fit the 800-line test budget, with shared setup extracted into co-located `*-test-harness.ts` / `*-test-fixtures.ts` modules (300-line budget). 83 files became ~930; the largest output is 797 effective lines. `orca-runtime.test.ts` is intentionally untouched. Test bodies were moved by scripted line-range slicing rather than retyped, so assertions are byte-identical. The only permitted body edits were mechanical rebinding where a shared value moved into a harness (e.g. `tmpHome` -> `homes.tmpHome`). Registries that enumerate test files were updated in lockstep: - config/max-lines-baseline.txt: pruned 341 -> 258 entries (all 83 removed). - config/reliability-gates.jsonc: 33 gates repointed at the split files, with assertionRefs split per file where a gate's coverage now spans several. - .github/workflows/pr.yml: the real-zsh lane now lists the 4 split files that actually exercise zsh, so they keep running in the dedicated shell lane. Also renamed agent-hooks `server-test-fixtures.ts` to `server.test-fixtures.ts` so the global-fetch call-site audit keeps skipping it, and added `.js` extensions to the CLI suites' dynamic harness imports (node16 resolution) to unbreak `build:cli`. Verification: full suite 52,449 passing vs 52,448 at baseline with zero assertions lost; `pnpm lint`, `pnpm typecheck`, and `pnpm build:cli` all exit 0; the terminal-pane e2e spec runs 31/31 headless. * refactor(tests): split hook-idle arbitration suite that oxfmt pushed over budget The pre-commit oxfmt pass reflowed pty-connection-hook-idle-arbitration.test.ts to 811 effective lines, 11 over the test budget. Split the hook-completion side effect and replacement-agent veto cases into their own suite; both files now sit well under the cap and the 15 tests are unchanged. * test: port upstream test changes into the split files after rebase Rebasing onto main surfaced 27 tests that main had added to files this branch deleted, plus edits to tests that had already moved. Taking the deletion side of those modify/delete conflicts would have dropped that coverage silently, so each upstream change is ported into the split file that now owns the behavior — for example main's six orchestration mailbox tests land across orchestration-runs, -send, and -check. Also repoints `orchestration.notification-mailbox-consistency`, a gate main added after this branch's gate remap, at those same three split files, and re-prunes the max-lines baseline against main's (257 entries). Verified: all 27 upstream test titles present; full suite 52,761 passing with the only diff vs baseline being 12 tests main itself removed and 3 that moved from skipped to passing; lint and typecheck exit 0. * fix(test): flush pending continuations before tearing down terminal test globals CI shard 5/16 failed on both Node 24 and 26 with `ReferenceError: window is not defined` from pty-connection.ts, surfacing through pty-connection-daemon-snapshot-replay.test.ts. The reattach/settle chains `await` a real promise and then touch `window.api`. Under fake timers those continuations cannot run, so they only become schedulable once restoreTerminalTestGlobals() switches back to real timers — which previously happened immediately before `delete globalThis.window`, so a late continuation threw and failed the whole file. Flush async ticks in that window instead. This is latent in the source rather than new: the pre-split 25k-line file kept running other tests after these, which gave the chains time to settle before teardown. Splitting the file moved teardown directly behind them. * fix(test): keep an inert window after terminal test teardown instead of deleting it The async-tick flush was not enough: the reattach/settle chain can resolve after teardown regardless of how long we drain, so CI shard 5/16 still failed with `ReferenceError: window is not defined` from pty-connection.ts. A real renderer never loses `window`, so deleting it was the artificial part. Swap in an inert proxy whose properties resolve to callables and whose calls resolve to undefined, making a late `window.api.pty.*` call a harmless no-op. The next test replaces it wholesale via installTerminalTestGlobals(), and no test asserts that `window` is absent. |
||
|
|
375b735e9c |
fix(agent-launch): preserve cold Codex startup drafts (#14688)
* fix(agent-launch): preserve cold Codex startup drafts * fix(agent-launch): honor startup draft readiness budgets |
||
|
|
ab9d1a29a9 |
fix(worktree): never reissue a generated workspace name (#14350)
* fix(worktree): never reissue a generated workspace name
Generated workspace names were deduped only against currently-live
worktrees, so deleting a workspace returned its name to the pool. A later
workspace could draw the same name, land on the same directory path, and
inherit the previous occupant's agent conversation history — coding-agent
CLIs key their prompt history and transcripts by cwd.
Names are now retired permanently per repo. The registry is written in
main with the name Git actually used (the create loop can advance past a
requested name on collision), and seeded once per run from workspace
directories and surviving agent transcript buckets so already-spent names
are excluded from the start. Suggestions degrade to -2, -3 variants
instead of recycling, and those variants retire too.
User-typed names are untouched: retirement filters suggestions only.
* fix(mobile): honor retired workspace names, on one shared implementation
Mobile hand-duplicated the desktop name-suggestion algorithm and deduped
only against live workspaces, so a phone could still be offered a name
whose deleted workspace left agent conversation state behind at that path.
Both platforms now call one shared selector in src/shared, so the two can
no longer drift. The host publishes retired names as an optional field on
the existing worktree.list response, and mobile fetches them per selected
repo while the create sheet is open — mirroring the desktop hook.
Mobile never calls worktree.list for its catalog (it uses worktree.ps,
which carries rows only), so this is a targeted request rather than a
change to the catalog or its cache. Hosts predating the field omit it and
mobile falls back to live-only dedupe, which is the pre-change behavior.
* fix(worktree): close retirement consistency gaps
* test(worktree): cover retirement runtime contracts
* fix(worktree): retire generated collision names
* fix(worktree): enforce retired names at creation
* refactor(ai-vault): extract the Claude project-dir encoder
The bucket-name encoder and its scope-boundary check were private to the
session scanner, so a second consumer had to reimplement them — and got the
per-character encoding wrong. Move both to a shared module with direct tests.
* fix(worktree): make the retirement seed scan actually match buckets
The bucket encoder collapsed runs of non-alphanumerics while the real one
emits a dash per character, so every dot-path bucket missed and the Windows
default workspace root (C:\...) matched nothing at all. Reuse the shared
encoder and its boundary check, which also stops a repo absorbing a sibling
whose path merely shares its prefix.
Also:
- Derive the workspace leaf by stripping the known encoded parent instead of
guessing from trailing dash segments, which retired the parent directory's
name whenever a workspace was named numerically.
- Reuse isAutoGeneratedCreatureBranchName so the -10 and -100 tiers retire.
- Drop the .codex/sessions root: Codex keeps the cwd inside the transcript
rather than in a directory name, so the scan could only ever see a year
folder. Reading transcript contents is not a trade this feature justifies,
so the gap is documented instead.
- Honor CLAUDE_CONFIG_DIR, which relocates the bucket root.
- Delete the unused retirableLeafName export.
Tests write buckets with the real per-character encoding against a fake home,
covering POSIX, dot-directory, Windows drive and WSL UNC roots; all three
platform cases fail against the previous encoder.
* fix(worktree): retire only generated names, keyed by cwd namespace
Two problems in the host-side registry.
Retirement fired for every create, including names the user typed. The
creature pool contains ordinary words — orca, runner, sole, molly, oscar — so
typing a retired 'nautilus' silently produced directory and branch
'nautilus-2' and burned the name for good. Creates now carry an explicit
nameWasGenerated flag; both the skip and the retire are gated on it, and it
defaults to false so CLI and automation callers are unaffected.
The registry was keyed by repo id, but both readers already discarded the id
and unioned by the cwd collision key, because the collision this prevents is
on the path. Keying by that namespace directly fixes several things at once:
entries no longer orphan when a repo is removed, remove/re-add no longer loses
every retirement for an unchanged path, the missing removeProject prune is
moot, and the backfill promise no longer merges into only the first repo id it
saw. The feature is unreleased, so no migration is needed.
Also:
- Memoize the collision key. It runs computeWorktreePath, which for a WSL repo
is a blocking execFileSync('wsl.exe') whose failure path is uncached, and
the previous code recomputed it once per repo on every create and every
listRetiredNames call.
- Drop retiredNamesByRepo from the worktree list result. It had no readers and
leaked onto 'orca worktree list --json', and its awaited backfill sat on CLI
selector resolution. The dedicated listRetiredNames RPC keeps its consumers.
- Make the three RuntimeStore methods required. RuntimeStore is file-private
with two constructors, so the 'older embedders' the optionality protected do
not exist, and the optional chain silently returned no retirements.
- Revert the unrelated forceDeleteBranch rewrite, and make room under the
file's line budget by extracting the create-args mapping instead.
* fix(worktree): send name provenance and stop gating Create on the fetch
Desktop and mobile now mark a create as generated-name only when the user
typed nothing and the composer fell back to the suggestion, so the host knows
which names it may retire.
Remove the retired-names loading gate from every create path. The host already
skips retired candidates before doing any git work, so the client gate bought
nothing while it could disable Create for the length of a full mobile
reconnect ladder (the wait had no timeout) and blank the desktop button
between queued creates. The suggestion still waits; the button never does.
Also make the web client call worktree.listRetiredNames instead of hardcoding
an empty list — the method is registered and mobile-allowlisted, so the
comment claiming no wire call existed was wrong — and filter the mobile
response to strings so a malformed row cannot throw during normalization.
* fix(worktree): key retirement by repo id and prune it with the repo
Reverts the collision-key storage key. It was a function of workspaceDir,
nestWorkspaces, worktreeBasePath and repo.path, so toggling any one of those
orphaned every retirement for every affected repo at once — trading a rare
churn (remove/re-add) for a common one. The read path already unions by cwd
namespace at query time, so cross-repo sharing never depended on the storage
key.
Instead, address the growth and orphaning directly:
- Drop the registry in removeProject, and in removeProjectForHost once the last
host's copy of the repo id is gone, alongside the sparse-preset deletes that
already follow this convention.
- Bound each repo's registry. The cap sits far above the 552-name pool because
evicting inside it would reissue a name whose agent state is still on disk;
only -2/-3 tier accumulation can ever reach it.
- Carry retirements through profile transfer, re-keyed to the destination repo
id and dropped from the source, mirroring sparsePresetsByRepo.
Separately, fix the backfill merge: the scan promise is cached per cwd
namespace, but it closed over the first repo id that triggered it, so a second
repo in the same namespace received nothing. The scan stays shared; the merge
moves out of the cached promise and runs for whichever repo asked.
Local repos re-seed on re-add through that backfill. SSH repos do not — the
scan cannot see the execution host — which is now stated in the module.
* docs(worktree): spell out why the retirement bound sits above the pool
Names the trap directly: the neighbouring 50/200 bounds cap histories, so
lowering this one to match them would silently start reissuing names whose
agent state is still on disk. Also states that oldest-first eviction is a
deliberate least-bad choice rather than a neutral one.
* fix(worktree): send name provenance from the web runtime client
This client hand-enumerates worktree.create params, so the new optional field
was silently dropped and typecheck could not see it. On web and paired-desktop
the host therefore never received it: generated names were never retired, and
the host-side skip that backstops a stale suggestion was disabled too. The same
client does fetch retired names for suggestions, so it was filtering against a
registry nothing ever wrote to.
The test asserts both directions, and fails without the fix.
* fix(worktree): retire names that took more than one collision suffix
isAutoGeneratedCreatureBranchName strips exactly one trailing -N, which is
right for auto-rename eligibility but wrong here. Once the pool is spent the
suggester emits nautilus-2, and a collision on that yields nautilus-2-3 —
which a single strip leaves as nautilus-2, not a pool name, so retirement
no-opped at exactly the tier where every base name is already gone. Strip
repeated suffixes locally rather than moving the auto-rename predicate.
* perf(worktree): keep the retirement backfill off the blocking WSL probe
The backfill runs on composer repo-select, not just at create time, and it
derived the probe path synchronously — which for a WSL repo with a mirrored
workspace dir reaches getWslHome and its blocking execFileSync('wsl.exe').
A stopped distro froze the main process for up to 5s on composer open.
Adds an async twin of computeWorktreePath and uses it for the probe. Resolving
the home there also warms the shared cache, so later sync callers are free.
Also stops memoizing the collision key when the WSL home is still unresolved:
only the success path is cached upstream, so caching the fallback namespace
would strand the repo there for the rest of the session.
* fix(worktree): hold retired names across a refresh instead of blanking
refreshKey changes on every workspace-list mutation, so create-multiple
refetches after each create and the hook returned an empty list until the
refetch landed — precisely the window in which resetForNextCreate clears the
name field and a fresh suggestion is drawn. Keep the previous answer while
revalidating and reset only when the repo changes; a failed refresh keeps what
was already loaded rather than un-retiring everything.
Also makes the returned array referentially stable, so the suggestion memo
downstream stops rerunning on every refetch.
* refactor(worktree): put the retired-name cache rules on one implementation
The desktop and mobile hooks that fetch retired names had already drifted
four ways. The transports genuinely differ (IPC vs RPC), but the caching
rules must not, and mobile's copy reset to [] on any error -- which
un-retires every name for the rest of the sheet session, the one outcome
retirement exists to prevent.
Moves the rules into src/shared/worktree/retired-name-cache: response
normalization, the never-leak-across-repos rule, and the hold-previous-on-
failure rule. Pure, no React, because src/shared is on the main process's
import graph. Each platform keeps its own transport and effect.
Mobile moves up to desktop's behavior: it now holds the previous answer
through a failed refresh, and refetches when the workspace list changes
instead of never refetching after mount.
Also drops the unused `loading` return. Neither platform consumed it; its
only consumer was the Create-button gate reviewed out earlier, and removing
it makes that regression unexpressible.
* fix(worktree): import shared types from their real modules
Main dropped the src/shared/types barrel, so the retirement module's import
resolved locally but not against the PR's merge base.
* refactor(worktree): bound the retirement registry by tier compaction, not eviction
Retirement is a correctness guarantee — a spent name's directory may still hold
agent conversation state keyed by that cwd — so the 2000-entry cap was the wrong
shape: reaching it handed a name back. At the owner's measured rate (~6.6 pool
names retired per day in one repo) the cap was ~9 months out.
Names come from a fixed 552-entry pool and the suggester only reaches tier N+1
once every tier-N name is taken, so a completed tier is exactly a set that no
longer needs listing. A row is now a watermark plus the names above it: reads
answer at-or-below the watermark with no lookup, and compaction drops the 552
entries the watermark now covers. Bounded at one pool per repo forever, with no
eviction and nothing un-retired.
Tiers can complete out of order (a create-time collision can spend `nautilus-2`
while tier 1 is open), so compaction loops and higher-tier names simply wait.
The RPC result carries the watermark beside the names as a new field; a client
predating it reads the names only and under-retires the compacted tiers, which
degrades to the pre-retirement behavior rather than breaking.
* fix(worktree): preserve generated name retirement across failures
|
||
|
|
a275f5ad85 |
fix(skills): cancel and bound abandoned skill discovery scans (#14670)
* fix(skills): cancel and bound abandoned skill discovery scans A root on a stalled network mount never settles its readdir, so after 30s the coalescer starts a replacement walk. The abandoned walk kept running with no cancellation, and nothing counted it, so live filesystem work accumulated for the life of the process. Superseding a scan for age now aborts it, and `findSkillFiles` plus the candidate tasks bail on the signal — that cannot unblock a syscall already in the kernel, but it stops an abandoned walk issuing more of them. A budget caps how many abandoned scans may be live at once; past it a replacement is shed rather than started, and the stalled entry is left in place so the root recovers on its own once the mount answers. The budget counts abandoned scans rather than live ones on purpose: one discovery legitimately walks a dozen-plus roots at once, so a cap on live scans would shed healthy roots and empty the picker. A `refresh` still does not abort what it supersedes — on a healthy root that scan is about to finish and its callers want an answer, and the existing publish fence already stops it writing a pre-mutation result. Shed roots report the new `unavailable` skipped reason so they stay distinct from roots that genuinely are not there. * fix(skills): propagate a walk abort thrown through a symlinked directory The broken-link `catch` around the symlink branch also wrapped the nested `visit`, so an abort thrown from inside a symlinked subtree was swallowed. When that link was the last entry, nothing afterwards re-checked the signal and the walk returned a truncated listing as success — the exact outcome the abort path exists to prevent. Only the `stat` is guarded now. The test pins the narrow window by aborting inside the symlink's stat, after the entry loop's check and before the nested visit's. * fix(skills): degrade an aborted root instead of failing the whole discovery Aborting a scan abandoned for age gave the coalescer a way to reject that it did not have before: every error inside the walk and the candidate tasks was caught locally, so the task effectively could not fail. Callers already waiting on that scan now see it reject, and `scanRootShared` re-threw anything that was not a shed — so one slow root failed the entire discovery and emptied the picker for every healthy root beside it, which is exactly what the shed path was written to avoid. Both endings now mean the same thing to a caller that can degrade a single root, behind one predicate: shed before the walk began, or aborted after it was abandoned. `discoverSkillsOnTarget` runs a second coalescer over the whole target, where there is no partial answer to degrade to, so it converts both into a retryable error rather than leaking the internal class to IPC/RPC. It must not answer with an empty result — zero skills reads as "nothing installed" and re-offers installs for skills that are present. Also drops the scan key from the shed message, which carried an absolute workspace path to a paired client and the renderer's error string, and corrects the budget comment: it is global and per-abandonment, so one wedged root can spend it alone, one replacement per 30s window. * fix(skills): keep the original error as the cause of a stalled-target error The target layer replaces the internal error with a user-facing one, which discarded what actually went wrong. Attaching it as `cause` keeps that in logs while the message stays free of the host path. Also records why name-matching is narrow rather than broad: the walk and the candidate tasks catch every filesystem error locally, so the only AbortError that can escape a scan is the one its own signal raised. * test(skills): pin that a real abort produces the name the predicate matches `isSkillRootUnavailableError` decides on `error.name`, and the discovery and target tests fabricate that shape rather than driving a real abort. So nothing pinned the linkage: if `throwIfAborted()` stopped producing an `AbortError` that is an `instanceof Error`, every test would still pass while a stalled root began failing whole discoveries again. The walk test already drives a genuine abort, so it asserts the name, the Error subclassing DOMException relies on, and the predicate itself. |
||
|
|
e12b22c435 |
fix(terminal): bracket agent-pane pastes so a pasted newline can't submit the draft (#14456)
* fix(terminal): bracket agent-pane pastes so a pasted newline can't submit A paste of <=64KB is handed to xterm's term.paste(), which rewrites newlines to CR and wraps in ESC[200~/ESC[201~ only when its parser observed DECSET 2004. Windows ConPTY never forwards that mode, so an unbracketed pasted newline reaches a TUI agent as Enter and submits the draft parked in its composer. The existing force-bracket guard keyed on isWindowsUserAgent() - the client's platform - but the ConPTY can be on a remote host, so the guard was off in exactly the configuration that needs it. Gate on the pane's own TUI agent instead, applied to all four paste entry points; middle-click primary selection had no force flag at all. The agent status row is retained deliberately and is not lifecycle-driven, so it can outlive the agent. Veto on shell-confirmed foreground (OSC 133;D) and on a row rehydrated across an app restart. The freshness TTL and a state check are both unusable here - an idle-but-live agent sits at done and still needs bracketing. Refs STA-4294 * fix(terminal): key agent-paste bracketing on live evidence, not shellForeground Measured against a real pane: shellForeground is republished only at OSC 133 boundaries, so a shell without 133 integration leaves it latched true while an agent owns the foreground. Vetoing on it silently reinstated the submit bug the parent commit fixes - the parked draft was sent on paste with the gate in place. Prefer process-confirmed agent identity when present, keep the restart-rehydrated veto, and drop the shellForeground veto. Erring toward bracketing costs a literal ESC[200~ in a non-2004 program; erring the other way sends the user's draft. Refs STA-4294 * docs(terminal): record why paste bracketing keys on the agent, not the mode bit Measured in real ptys before taking the obvious alternative. A tri-state on DECSET 2004 (observed-on / observed-off / never-observed) does not work: zsh 5.9, fish 4.8.1 and bash >= 5.1 announce and withdraw the mode cleanly, but macOS /bin/bash 3.2, /bin/sh, bash 4.4 and any shell with bracketed paste disabled emit nothing at all, byte-identical to a bare `cat`. Silence cannot be read as consent. Agent identity disambiguates it in the one direction that matters: agents always enable the mode, so silence on an agent pane means the announcement was lost in transit, never an opt-out. Also measured: bracketing a program that never negotiated is worse than useless - the markers land as literal payload bytes and ICRNL still turns the CR into a submit - so the gate stays narrow. Refs STA-4294 * fix(terminal): guard the paste pane key and document the evidence policy Readiness review follow-ups, none behaviour-changing for reachable inputs. makePaneKey throws on a malformed leaf/tab id. It is unreachable today (pane.leafId is a minted UUID and the same pair is already called unguarded from a hotter site), but the failure mode was bad: the throw escapes before the paste helper's catch is attached, so the paste would be a silent no-op with no error surface. Degrade to the pre-fix path instead, with a test. Also record two things a future reader needs: the process-confirmed branch is dead for remote-runtime and SSH panes because foreground tracking is disabled there, so a remote pane's status row is its only evidence; and why this resolver deliberately omits the shellForeground/routingRevoked/routingTrusted gates its two siblings enforce - they route input bytes, this only wraps a paste whose payload is ESC-sanitized downstream. Refs STA-4294 * fix(terminal): encode Windows agent paste newlines as input records * fix protected paste handling in dashboard previews |
||
|
|
0e8d52912c |
fix(source-control-ai): stop duplicating singleton CLI flags in agent argv (#14585)
* fix(source-control-ai): stop duplicating singleton CLI flags in agent argv Recipe CLI arguments and agent command overrides were appended on top of Orca's generated flags, so a user-supplied --model produced a repeated flag. yargs collapses a repeated flag into an array, crashing OpenCode with "j.split is not a function"; clap rejects it outright for Codex. Declare the at-most-once option groups per agent spec and fold every user-supplied occurrence into the generated slot, with recipe args outranking a command-override prefix. Fixes #12305 * fix(source-control-ai): preserve command override argv order |
||
|
|
500b72d8ef |
fix(vm): harden provisioned root ownership and cleanup (#14477)
* fix(vm): verify provisioned root ownership * test(vm): retry transient removal menu * test(vm): stabilize provisioned root teardown * fix(vm): clarify recipe-owned cleanup * fix(vm): pin provisioned root source commit * fix(vm): make runtime cleanup user-cancellable |
||
|
|
32f46f9a24 | fix(vm): keep failed cleanup retryable (#14476) | ||
|
|
8b22f044f5 |
fix(vm): preserve runtime sidecar rollback compatibility (#14444)
* test(vm): reproduce runtime store rollback poisoning * fix(vm): keep runtime sidecar rollback-readable * fix(vm): harden rollback-compatible runtime persistence * fix(vm): publish rollback lifecycle authority first * test(vm): harden rollback compatibility coverage |
||
|
|
3a212584ec |
feat(native-chat): Extra high grok effort per model (#14577)
* feat(native-chat): offer Extra high grok effort per model Slice grok's reasoning-effort menu by each model's advertised ceiling so 4.6 can reach xhigh while 4.5 stays at high, and keep the untouched default at high so launch argv does not silently escalate. * fix(grok): parse dashed rows in grok models listing Grok stars only the default model and dashes the rest. A star-only bullet dropped 4.5 from the picker once discovered models were authoritative. |
||
|
|
9bb8836bb6 |
fix(agent-launch): wait longer for cold-boot Codex composer before dropping prompt (STA-3367) (#12853)
* fix(agent-launch): wait longer for cold-boot Codex composer before dropping prompt (STA-3367)
Continue-in-new-session pastes the handoff prompt once Codex renders its
composer glyph, gated on an 8s readiness budget. A cold/first-run Codex can
take longer than 8s to mount its composer, so the wait timed out and the
prompt was silently dropped into an empty terminal.
Marker-gated ready signals (Codex glyph, opencode show-cursor) are positive
proofs: the paste fires only when the marker actually renders, so a longer
budget can never paste prematurely — it only tolerates slow cold boots. Give
those signals a 20s budget while the markerless quiet-window signal keeps 8s.
* fix(agent-launch): share the composer-readiness budget across all three delivery owners (STA-3367)
The cold-boot fix was correct but landed as a single-path exception, and it
double-spent its own budget. Three follow-ups so the behavior is a system rule:
1. Split the PTY-spawn wait from the composer wait in pasteDraftWhenAgentReady.
Both were handed the same budget, so a codex tab took up to 41s to report a
dropped prompt. "Tab has a PTY" and "composer accepts input" are separate
states: spawn keeps a fixed 8s, and the readiness budget now starts once the
PTY exists, so a slow spawn can't shorten a cold composer's window.
2. Move the per-signal budget to draftPasteReadyBudgetMs() beside the shared
readiness scanner. The budget is a property of the ready signal — only that
module knows which signals are marker-gated — so all three delivery owners
(renderer tab paste, renderer startup paste, main runtime startup paste)
consume one policy instead of three hardcoded 8s constants.
3. Give the main-runtime startup paste the process-ownership fallback both
renderer paths already have. It resolved null on budget expiry, silently
dropping the prompt on worktree-create / CLI / remote-host delivery — the
same STA-3367 failure, on the path the original fix didn't reach.
Adds coverage for the main-runtime waiter, which had none.
Test: vitest src/main/runtime src/shared src/renderer/src/lib
src/renderer/src/components/terminal-pane — all green; tsc clean.
* test(agent-launch): consume the shared readiness budget instead of restating it
Hardcoding 20000 in the runtime waiter test meant it would keep passing if
OrcaRuntimeService stopped consuming draftPasteReadyBudgetMs — the exact drift
this PR exists to prevent. The literal values stay pinned once, in the scanner
test.
* refactor(agent-launch): collapse the readiness budget to one flat timeout
The per-signal budget (marker 20s / quiet-window 8s) tied the timeout to how
readiness is DETECTED. The budget is really a property of how slowly an agent
can boot — a marker, a quiet window, and a process check all wait out the same
cold start — so one number covers all three signals.
Replaces draftPasteReadyBudgetMs() with DRAFT_PASTE_READY_TIMEOUT_MS: drops a
constant, a branch, and two tests, and removes the only reason a delivery path
needed to know which signal class it was using.
Cost: a launch that never emits DECSET 2004 now surfaces its 'prompt not sent'
toast at 20s instead of 8s. That is the failed-launch path only; successful
markerless delivery still resolves on the 1.5s quiet window as before.
* fix(agent-launch): constrain cold Codex readiness budget
* fix(agent-launch): observe Codex readiness from PTY bind
* fix(agent-launch): anchor early Codex prompt to TUI screen
|
||
|
|
83e2123582 |
Add global worktree visibility source defaults (#14276)
* Add global external worktree visibility defaults * Expand global worktree visibility source defaults * Fix host-scoped visibility settings races * Fix global worktree visibility integration * Enable source visibility defaults on mobile * Polish external worktree settings navigation * Clarify inherited worktree visibility settings * feat(sidebar): replace the inherited-visibility switch with a Show/Hide picker Each source row now shows a two-segment Show / Hide control preselected to the global setting, and explains itself only where the project actually disagrees: an "Overriding global setting: <value>" card names the value being ignored. Picking the segment global already holds drops the override instead of pinning a duplicate, so the same control both overrides and reverts, retiring the separate "Use global" link. The dialog footer now lists every inheritable source with its global value. * fix(sidebar): preserve reset for matching visibility overrides |
||
|
|
266b5ae8f5 |
fix(mobile): match desktop project and run target picker (#14457)
* fix(mobile): disambiguate repository locations * fix(mobile): preserve explicit repository ownership * test(mobile): use explicit renderer type * refactor(mobile): match desktop project targets |
||
|
|
87bafb5e6a |
fix(terminal): ground the emulator to the snapshot's baseline after a byte gap (#14374)
A hidden-delivery byte gap can strand more than the SGR pen, and the reset #14241 added to the split alt-screen replay is undone before any content is painted: xterm answers `?1049l` with restoreCursor(), which reloads the pen, all four G-set designations, GL, origin mode and wraparound from the register saved at `?1049h`. - Bracket the buffer switch with the baseline: before, so `?1049h` banks grounded state rather than the gap's; after, so `?1049l`'s restore cannot reapply it. - Ground everything a serialized payload is diffed against, not just the pen: SGR, GL, all four G-sets, origin, autowrap, insert, the per-buffer scroll region, and the saved-cursor register. - Switch buffers only when the pane is actually on the other one. `?1049` is not a no-op otherwise — it still swaps the kitty flag registers, which would park the flags of an agent that negotiated them on the normal screen. - Return to the normal buffer when the gap ate the TUI's exit sequence; the restored history was painting into the alt buffer with scrollback left empty. - Restore the CAN #14241 dropped, so a control string the gap truncated is discarded instead of committed by the next ESC. - Ground the abandon path exactly once instead of twice. - Derive the parity/fuzz preambles from the same builder; they had drifted and were asserting against bytes production no longer emits. |
||
|
|
92b6ffd17d |
Terminate renderer graph reload generations and contain disposed-frame notifications (#14070)
* fix(runtime): terminate renderer graph reload generations * fix(runtime): harden renderer reload teardown * fix(runtime): fence renderer graph publication ownership * test(runtime): register renderer graph reload gate * test(runtime): record live reload validation * fix(runtime): ignore cancelled renderer navigations * chore: preserve main formatting during branch sync * chore: satisfy changed-code quality gate * fix(runtime): restore cancelled renderer reloads * fix(runtime): preserve committed reload fencing * test(runtime): prove cancelled reload timeout * docs(reliability): record reload cancellation oracle --------- Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> |
||
|
|
b82a8791f5 |
perf(runtime): resolve an explicit worktree id without scanning every repo (#14399)
`resolveWorktreeSelector` resolved every selector kind from the whole-fleet snapshot, so a targeted `id:<repoId>::<path>` lookup fanned `git worktree list` across every registered repo to answer a question about one of them. With a cold scan cache -- app startup, or the first lookup after a mutation clears the snapshot -- that is one subprocess per repo, ~17ms each, to find a worktree whose owning repo the id already names. Measured on a ten-repo fleet: one `id:` lookup scans 10 repos before and 1 after. Scope only `id:`. Every other selector kind is matched across the fleet and its `selector_ambiguous` contract is defined over all repos, so scoping `branch:`, `name:`, `issue:`, or a bare selector would silently pick a winner where they correctly refuse today. A test pins that: `branch:main` across ten repos still throws `selector_ambiguous` and still scans all ten. Lineage stays correct because edges are intra-repo by construction. The scoped path returns null and falls back whenever that does not hold: a repo id registered on several execution hosts, an unknown repo id, or a worktree the scoped scan does not contain. A warm fleet snapshot always wins. Row resolution moves out of orca-runtime.ts into repo-worktree-row-resolution.ts, which owns no state -- the cache-aware scan and folder-workspace stamping are injected. orca-runtime.ts ends up 65 lines shorter than before despite the added feature. |
||
|
|
2100fb2553 |
fix(runtime): cap remote git.diff and file previews at the transport budget (#14160)
* fix(runtime): cap remote git.diff and file previews at the transport budget A remote or mobile user who opens the diff of a large image loses their whole WebSocket, not just that request: the E2EE channel closes with 1013 when a reply exceeds the 4 MiB outbound envelope. Two producers can exceed it unaided. git.diff/branchDiff/commitDiff cap text with MAX_RENDERED_DIFF_COMBINED_CHARACTERS (6M chars) -- a *renderer* budget that sits above the transport limit -- and return base64 for previewable binaries bounded only by MAX_GIT_SHOW_BYTES, so a 10 MiB PNG changed in place is ~26.7 MiB in one envelope. files.readPreview inlines base64 up to 10 MiB, and mobile calls it for every image tab. Both now measure against a budget derived from the outbound limit. The check sits in orca-runtime-git.ts, downstream of the dedupe and of both the SSH-provider and local branches, so a payload forwarded verbatim by an old relay is covered by the same code and src/relay needs no change. Local and in-process callers pass no budget and keep full fidelity. Measuring raw bytes would not work, which is the whole reason this needs a module. JSON escaping turns one control byte into six (\u00XX), and binary-buffer.ts sniffs only for NUL in the first 8 KiB -- so a NUL-free file of 0x01-0x1f bytes is classified as *text*, would pass a raw-byte cap, and would then blow the envelope. The budget is escape-aware, with a three-branch fast path that keeps normal diffs at two native byteLength calls and scans only the ambiguous band. The SSH branch of readFileExplorerPreview had the same raw-vs-escaped gap: its stat gate sizes base64 binaries, but text crossed unbounded. It now honours the same decoded-text limit the local branch already enforced. No wire change: GitDiffResult is untouched -- no third kind, no new field. Old clients see an error for one request instead of a dropped connection. diff_too_large joins the structured passthrough codes and lands on an existing error arm in both mobile consumers and the desktop remote path; file_too_large was already handled on both. Instruments the 1013 close, which nothing measured before, so the incidence this cap is meant to drive to zero is finally observable. `emitter` separates a producer size bug from a wedged link. Known regression: remote image previews between ~3.096 and ~3.146 MB now return file_too_large. They only intermittently worked before -- above ~3.0 MB they killed the socket -- so this trades intermittent connection loss for a consistent error. Test: 10281 passed in src/main/runtime + src/shared + src/main/git; mobile 3427 passed. Each of the six budget-enforcement sites is independently mutation-killed. Escaping fixtures cover newline-dense, control-char, CJK, lone-surrogate and base64 content against native JSON.stringify. tsc clean for node, web and cli; oxlint clean. Co-authored-by: Orca <help@stably.ai> * fix(runtime): harden remote reply transport budgets * test(runtime): cover desktop remote preview budgets * test(runtime): close telemetry review gaps * chore(shared): repoint budget imports after the shared/types barrel removal Upstream #14447 dropped the shared/types barrel; GitDiffResult now lives in git-diff-compare-types and GlobalSettings in global-settings-types. Co-authored-by: Orca <help@stably.ai> * fix(ssh): surface an over-cap preview read as file_too_large The stream reader aborts an over-cap read with StreamProtocolError, whose numeric code falls through mapRuntimeError to a generic runtime_error carrying the raw "Reported totalSize N exceeds client cap M" string. Neither preview client recognizes that: runtime-file-client.ts and mobile-file-preview-response.ts both key on file_too_large. It also made the two file_too_large guards directly below the read unreachable on the streaming path. Gives the cap its own error type so the caller can translate it, keeping the bandwidth saving the cap exists for. A genuine protocol fault still propagates unmasked. Found by the readiness review. Mutation-verified: removing the translation fails exactly the new test. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
d137bb93e1 |
fix(agent-status): stop start-less child stops from minting phantom working (#14375)
* fix(agent-status): stop start-less child stops from minting phantom working buildClaudeCachedLeadStatusPayload fell back to 'working' whenever the pane had no cached lead-turn state. That default is right for a spawn or a child tool call, but the same helper serves SubagentStop and TeammateIdle, which end work and prove the opposite. claudeLeadStateByPaneKey is in-memory only, so every app restart empties it. A Claude session that outlives the restart reports its next child event into an empty map and the pane latches 'working' with an empty roster -- no Stop ever clears it, and the 30-minute window only decays the sidebar dot, never the stored state. Fall back by the event's evidence: terminating child events resolve to 'done', which still gates up through resolveClaudePaneState when the roster or background work proves the pane is busy. * fix(agent-status): require evidence for child completion * fix(agent-status): publish matched teammate idle * fix(agent-status): preserve confirmed child work * fix(agent-status): retain live restored teammates * fix(agent-status): reap unconfirmed siblings after child drain * fix(agent-status): preserve unmatched restored children * fix(agent-status): wait for lead completion after child stop * fix(agent-status): persist restored child transitions --------- Co-authored-by: Brennan Benson <brennan@stably.ai> |
||
|
|
9e5ee5ef8e |
feat(workspaces): rework cleanup discovery and dialog (#13413)
* fix(workspaces): support full cleanup scans * feat(workspaces): persist cleanup snapshots * feat(workspaces): add cleanup filter model * refactor(workspaces): remove cleanup presets * feat(workspaces): rework cleanup dialog * fix(workspaces): keep cleanup row ordering render-pure * refactor(workspaces): simplify cleanup browsing * refactor(workspaces): show cleanup facts * refactor(workspaces): surface cleanup row facts * fix(workspaces): remove misleading cleanup count * fix(workspaces): preserve full scan semantics * fix(workspaces): scope snapshot persistence * fix(workspaces): preserve cleanup browse compatibility * fix(workspaces): reconcile cleanup dialog state * test(workspaces): update snapshot store fixtures * test(workspaces): preserve cleanup scan modes * perf(workspace-cleanup): stream scan progress and size results * fix(workspace-cleanup): address review feedback * fix(workspace-cleanup): preserve host-scoped cleanup metadata * fix(workspace-cleanup): declare review source dependencies * fix(workspace-cleanup): align size scan banner * fix(workspace-cleanup): shorten scan action * perf(workspace-cleanup): avoid redundant scan IO * perf(workspace-cleanup): bound restarted evidence scans * fix(workspace-cleanup): satisfy scan queue lint * perf(workspace-cleanup): bound scan and snapshot work * perf(workspace-cleanup): serialize final enrichment * test(workspace-cleanup): assert final enrichment drain * fix(workspace-cleanup): stop progress after renderer teardown * perf: batch workspace cleanup git evidence scans * perf(workspace-cleanup): stop redundant snapshot and scan work * fix(workspace-cleanup): resolve review findings across scan, store, and dialog Correctness: - Chunk git-evidence dispatches at the shared 500-target limit and exclude queued/in-flight ids from target selection, so fleets past the limit can no longer strand rows permanently mislabeled as checked-but-unknown. - Key destructive selection pruning on the user's filter state instead of the per-tick matched-set identity; streaming reclassification no longer silently deselects rows. - Clamp the facet clock to max(scannedAt, open time): a stale hydrated snapshot no longer misbuckets idle thresholds or keeps dead agents fresh; row labels use the same clock. - Supersede and cancel the previous broad scan when a new one starts (renderer registry and same-sender guard in main) instead of racing two fleet scans. - Gate snapshot persistence on hasTargetedWorkspaceCleanupScan so worktreeIds: [] can never persist an empty fleet snapshot. - Re-apply dismissals at set-time in progress application so a dismissal landing mid-enrichment is not clobbered. - Record a one-off local snapshot prune for single (unbatched) remote deletes so removed workspaces cannot resurrect from cache. - Strip .exe when normalizing foreground process names so Windows agent processes match. Performance: - Cache per-candidate facet and review-info objects on candidate identity; no-op streaming ticks reuse the previous rows array and skip every downstream pass; matched-set identity is stable under equal membership. - Compute facet counts/options only while the filter popover is open. - Equality-bail git-evidence publishes; structural (non-stringify) facet-group comparison memoized in the toolbar. - Identity-token fast path for the enrichment cache (cache hits skip both JSON.stringify signatures); prune viewed/dismissal records on removal and expiry; bound the superseded-scan-id set. - Restore the no-op bail in removeWorkspaceSpaceWorktrees (regression). - Abort main-side scans when the renderer is destroyed; module-scope controller maps survive handler re-registration. - Batch removal preflight into one targeted scan (with refreshActivity) per 500 ids instead of one scan per row. - Scan repos at concurrency 2, report discovered counts upfront for honest progress, share fs-activity probes per path (folder workspaces), read only the reflog tail, and skip the snapshot read-before-write via a remembered scannedAt. Split workspace-cleanup-worktree-listing, workspace-cleanup-facet-row-caches, and workspace-cleanup-selection-model out of files that crossed max-lines. * fix(workspace-cleanup): address verifier findings - Fall back to a full reflog read when the newest record exceeds the 8KB tail window, so an oversized subject cannot hide recent ref activity. - Bound the single-removal snapshot prune batch id with a UUID; embedding the unbounded worktreeId silently failed main's 128-char validation and skipped the prune for long remote ids. - Key the main-side broad-scan supersession by sender AND scan mode so legacy suggestion-only and full-workspace scans stay isolated, matching the renderer registry. * fix(workspace-cleanup): own facet caches with useMemo instead of render-time ref writes React Doctor (CI changed-lines gate) correctly flagged the three cache refs written during render. Each per-candidate cache now lives in one memo with the derived context it is keyed on, so the memo deps are the invalidation and interior fills stay content-addressed; the matched-set identity stabilization is dropped since its only consumer reads through a useEffectEvent and never keys on identity. |
||
|
|
77f23b013f |
refactor(shared): drop the shared/types barrel and import from the real modules (#14447)
#14397 split `shared/types.ts` into 46 per-domain modules but kept the path as a re-export barrel so the import sites did not have to change. This removes the barrel: every consumer now imports from the module that actually declares the type, and `src/shared/types.ts` is deleted. Barrels hide where a type lives, make every consumer look like it depends on the whole domain, and let an unrelated edit invalidate a module that ~2,000 files transitively import. 2,323 import declarations across 2,321 files. Rewritten mechanically: each specifier was resolved to an absolute path via the TypeScript AST and recomputed, rather than string-substituted, so alias forms (`@/../../shared/ types`) and per-specifier `type` modifiers survive. Four cases the mechanical pass had to handle, each found by a gate rather than by reading the diff: - Modules inside `src/shared` import the barrel as `./types`, not `shared/types`. A pre-filter on the latter string skipped 176 of them and left imports dangling at a deleted file, which surfaced as confusing `Property 'x' is optional in type 'Repo' but required in Pick<Repo, ...>` errors rather than "module not found". - The barrel RENAMED one type on the way through (`WorkspaceSource as WorkspaceCreateTelemetrySource`), so the original name in the owning module has to be re-aliased at each consumer. - Three test files put `;(globalThis as ...)` on the line after the import. TypeScript parses that `;` as the import statement's terminator, so replacing through `statement.getEnd()` deletes it and breaks ASI. The rewrite now stops at the module specifier. - A file that already imported directly from a module got a SECOND import from it, because the barrel re-exported those same names — which trips `import/no-duplicates` under `--deny-warnings`. A post-pass merges declarations sharing a specifier and type-only-ness; the `import type` plus `import` pair from one module is left alone, since that form is allowed. Splitting one barrel import into several genuinely adds lines, which pushed `terminal-layout-pty-ownership.ts` to 301 counted lines: its 107-character import must wrap, and neither local type collapses onto one line (101 and 116 characters). Rather than contort a type declaration to fit a line budget, `collectLeafIds` and `pruneLeaves` move to `terminal-pane-layout-tree.ts` — they are pure structural operations on the layout tree and independent of PTY ownership. `visible-worktrees.ts` similarly loses its own mini-barrel re-export of `isDefaultBranchWorkspace`, with the four real consumers repointed at the declaring module. No `max-lines` bypass added. Verified: cold `tsc --noEmit` green on node, cli, and web (buildinfo deleted first — these projects are `composite: true` and reuse stale caches); the full `pnpm lint` green, not just bare oxlint — the narrower local check is what let the duplicate imports reach CI; max-lines ratchet OK at 344. |
||
|
|
cd6114ab7e | fix(browser): acknowledge paired tab before navigation (#14402) | ||
|
|
5a6837fce4 |
refactor(store): unify the duplicated catalog equality and identity-key helpers (#13804)
* refactor(store): unify the catalog structural-equality walks Three near-identical structural deep-equality walks had landed independently in the same window: areValuesEqual (#13744, repo-identity-reconcile.ts), areCatalogEntriesEqual (#13770, repos.ts — already folded into the first on this branch's base) and catalogValuesEqual (#13662, worktree-catalog-reconciliation.ts). All three walk plain records and arrays and fall back to reference equality for anything exotic. They are not interchangeable. Two axes genuinely differ, and each caller depends on its own side: - Own-key set. #13744/#13770 require equal own-key counts plus hasOwnProperty, so an absent key differs from a key present and holding `undefined`. #13662 compares the union of both sides' keys, so those are equal. The strict side is load-bearing: the repo/project merges branch on `'localWindowsRuntimePreference' in project` (repos-project-runtime.test.ts "clears stale local runtime preferences"), and projects are now reconciled with this comparator. The loose side is test-pinned by worktree-catalog-reconciliation.test.ts "reuses rows with equivalent nested catalog data", where a locally built row carries `optional: undefined` that the host omits. - Leaf comparison. #13744/#13770 use `===` (NaN never equal, 0 equals -0); #13662 uses `Object.is` (the reverse). So instead of picking a winner, src/shared/structural-value-equality.ts holds one walk parameterised by those two axes and exports the two policies as `structuralValuesEqual` and `structuralValuesEqualIgnoringUndefined`. Every caller keeps its exact current semantics; the ~40 duplicated lines and the silent divergence go away. src/shared/persisted-ui-equality.ts (a fourth copy with a Set branch and no plain-object guard) is deliberately left alone: it gates a disk write in main with no direct test coverage. Also folded, all provably behaviour-identical: - The `${hostId}\0${repoId}` composite key had three copies (getRepoHostIdentityForParts, repoOwnerKey, getEntryKey) that must agree or repos silently stop reconciling. Moved to src/shared/repo-host-identity.ts because one of them lives in src/shared; the renderer module re-exports it. - mergeFetchedReposForHost's hand-inlined upsert loop now calls mergeByIdentity. mergeByIdentity additionally skips replacing a structurally equal row, which cannot change the result here: reconcileFetchedRepos runs immediately after over the same identities in the same order and restores exactly those rows. - Renamed repos.ts's `catalogRowsUnchanged` to `arrayElementsUnchanged`. It is a pure element-identity compare, two files away from `catalogRowsEqual`, which is a full structural compare. src/shared/structural-value-equality.test.ts pins both policies over arrays, nested records, null-prototype records, absent-vs-undefined keys, symbol keys, and non-plain objects (Date/Map/Set/class) falling back to reference equality. * fix(store): keep merged sourceRepoIds order host-independent Prefixing the cross-host remainder made a cross-host project's sourceRepoIds order a function of the refreshing host, so the projects reconcile never reused the row. Also pins the repo-derived host-id contribution the new per-project slice feeds the host-id resolvers. Co-authored-by: Orca <help@stably.ai> * refactor(store): migrate call sites that landed after this branch github.ts and ai-vault-session-identity.ts began using areValuesEqual on main while this branch was stale, and repo-identity-reconcile's record reconciler still called its own deleted walker. All three now use structuralValuesEqual; reuseEqualCatalogRows keeps its duplicate-id cap and calls the ignoring-undefined variant, which is the key-union semantics catalogValuesEqual had. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
583ab1601b |
refactor(shared): group worktree, github, and linear modules into folders (#14437)
`src/shared` is a flat directory of ~1,150 entries. The worktree, github, and
linear domains accounted for 71 of them, so finding the module you wanted meant
scanning a wall of same-prefixed filenames.
Move each domain into its own folder and drop the now-redundant prefix:
src/shared/github-pr-types.ts -> src/shared/github/pull-request-types.ts
src/shared/worktree-id.ts -> src/shared/worktree/id.ts
src/shared/linear-links.ts -> src/shared/linear/links.ts
This follows the existing `network/` and `new-workspace/` convention in the
same directory, which also drop the prefix inside the folder.
Whole clusters move, including tests. Foldering only part of a domain would be
worse than flat: a reader would have to check both `github/` and the flat
directory, and `github-auth-types.ts` / `github-project-types.ts` are type
modules that belong with the rest. No files with these prefixes remain flat.
Import specifiers were rewritten by resolving each one to an absolute path and
recomputing it, not by string substitution, so the `@/../../shared/...` alias
forms are handled correctly. 501 specifiers across 298 files.
Two things `tsc` cannot catch, handled explicitly:
- `github-project-types.ts` carries its own `max-lines` bypass, so its baseline
entry is REPOINTED to the new path rather than pruned. Pruning would drop the
bypass and then flag the new path as a fresh violation. Ratchet stays at 345.
- `mobile/` is outside `pnpm typecheck` and cannot be typechecked here
(`mobile/node_modules` is empty). Instead every relative specifier in the repo
was resolved against the filesystem: 174 unresolved before this change and 174
after — identical, so nothing broke in mobile either.
The pinned `tests/e2e/.cross-version-checkouts` fixtures are deliberately NOT
rewritten; they are a snapshot of an older release and still reference the old
paths.
Verified: cold `tsc --noEmit` green on node, cli, and web (buildinfo deleted
first — these projects are `composite: true` and reuse stale caches).
|
||
|
|
070572bfd7 |
refactor(shared): split shared/types.ts into per-domain type modules (#14397)
`src/shared/types.ts` was 3,981 raw lines (2,825 counted, 9.4x the 300-line budget) behind an `eslint-disable max-lines`, and is imported by 2,092 files — the single widest contract surface in the repo. Move all 320 top-level declarations into 46 per-domain modules (`repo-types.ts`, `worktree-types.ts`, `github-pr-types.ts`, ...) and reduce `types.ts` to an explicit re-export barrel, so the 2,092 import sites are untouched. `HostSettingOverrides` moves into the pre-existing `host-setting-overrides.ts` alongside the accessors that operate on it, which also removes that module's circular import back into `types.ts`. Named re-exports only, never `export type *`: with star re-exports a name exported by two modules is silently dropped, which would surface as a confusing "has no exported member" at a random call site. Verified lossless mechanically, not by inspection: - export parity — the module's resolved export set through the TS checker is identical before and after (396 names, no additions, no removals) - declaration parity — all 320 declarations compare character-identical modulo comments and whitespace, so no optionality, union order, or generic parameter drifted - `tsc --noEmit` green on the node, cli, and web projects - `oxfmt --write` is byte-identical, so the barrel is format-stable Drops the `max-lines` bypass and its baseline entry (ratchet 346 -> 345). |
||
|
|
400321edcd |
fix(workspaces): gate the GitHub palette number match on repo remote identity (#14413)
* fix(workspaces): gate the GitHub palette number match on repo identity `repoMatchesGitHubSlug` returned the permissive `'unknown'` whenever the repo displayName was not in `owner/repo` form and no upstream metadata existed — the common basename-named non-fork case. The caller only rejects on `false`, so a pasted issue/PR URL could activate a workspace in a different repo that happened to share the number, since issue/PR numbers are per-repo. Mirror the GitLab gate from #14381: fall back to the probed `gitRemoteIdentity.canonicalKey` before giving up, comparing host and owner/repo after normalizing port, `www.`, and case. An `upstream`-derived identity stays `'unknown'` because `deriveGitRemoteIdentity` ranks `upstream` above `origin`, so a fork's own origin is invisible and rejecting would drop URLs from the fork the user actually checked out. The canonicalKey compare runs after the displayName branch: displayName is compared host-agnostically, so mirrors and host aliases of the same owner/repo keep matching as they do today, and the probed remote only fills in where no name evidence exists. Refs STA-4237 * fix(workspaces): keep SSH host aliases matching in the palette identity gate `git remote -v` reports ssh.github.com, www., and ~/.ssh/config `Host` aliases verbatim, so comparing a probed canonicalKey against a pasted URL host rejected legitimate GitHub/GitLab remotes. Normalize the alias hosts both sides can fold offline, and downgrade a host-only mismatch to 'unknown' when the probed host is dotless (an unexpandable OpenSSH alias); dotted hosts like ghe.example.com still lose. Lifts the GitHub host normalizer into shared instead of a third copy. * fix(repos): keep the www host fold out of the derived project identity getProjectIdentityKey feeds the persisted Project id, so folding www. there re-keyed existing projects on upgrade and dropped localWindowsRuntimePreference. Restrict the fold to the palette's URL-vs-remote comparison, and pin the derived id for a www. remote so it cannot drift silently again. |
||
|
|
17690d49ea |
fix(repos): re-probe git remote identity so a stale snapshot stops misjudging identity gates (#14414)
* fix(repo-identity): re-probe resolved git remote identities on a long TTL A resolved gitRemoteIdentity was written once and frozen for the life of the repo record, so adding an `upstream` remote later — or a project rename or transfer — left identity gates judging against the path the repo had when it was added. Re-probe resolved repos on a 6h TTL, seeded 5 minutes after a repo is first seen in a process and capped at 4 refreshes per sweep so a restart cannot fan out a subprocess per repo. Only a successful probe that yields a different canonicalKey overwrites; failures and no-remote answers leave the existing identity alone. Also explain why the worktree-scan admin fingerprint timeout deliberately exceeds its caller budget, and log when that probe expires — expiry was silent and indistinguishable from "fingerprint unavailable". Refs STA-4247 * fix(projects): carry project state across derived project id changes A project id is derived from repo identity, so a remote re-probe (or a repo:->git:->github: promotion) rewrites it. The compatibility merge matched prior rows by id only, dropping the user's localWindowsRuntimePreference and leaving a ghost project row that independent host setups still pointed at. Both merge sites now fall back to the prior row whose sourceRepoIds overlap and re-point independent setups at the surviving project. |
||
|
|
eb22e497bb | Revert "fix(ssh): reapply the reattach-identity work and stop the fallback fence stranding moved panes" (#14395) | ||
|
|
d16092e503 |
fix(updater): recover renderer shutdown checkpoint (#14373)
* fix(updater): recover renderer shutdown checkpoint * test(updater): cover checkpoint recovery in Electron * fix(updater): keep staging failures blocking |
||
|
|
537864a248 |
Fix Codex hook trust before manual shell launches (#14326)
* fix codex hook trust before shell launch * fix packaged cli preflight dependency * fix codex shell preflight safety * fix Codex shell preflight settings and startup safety |
||
|
|
6a0c8fa541 |
fix(ssh): reapply the reattach-identity work and stop the fallback fence stranding moved panes (#14384)
* Reapply #13326 and #13928 (un-revert #14361) Restores the SSH reattach-identity and daemon-occupancy fixes. Reverting them reintroduced their P0s, filed as STA-4224, STA-4225, STA-4227, STA-4230, STA-4232, STA-4233 and STA-4234 against #14361. The tab loss that motivated the revert is fixed in the commits that follow, so this reapplication is not a straight redo. * fix(relay): stop the fallback attach fence refusing a pane that moved tabs The primary fence was moved to the shell's own incarnation precisely because paneKey/tabId froze the pane's LOCATION at spawn and refused panes that had merely moved. The fallback that older clients fall into kept the old rule, so the correction never reached it — the same 'the rule exists, but this path does not ask it' leak this work has hit repeatedly. A refusal here is not recoverable: an identity mismatch never grounds a respawn, so the pane keeps a live shell it can no longer reach and renders blank. Narrowed to paneKey, which is the identity; the tab is a location. Restoring the tabId comparison reddens the new test. |
||
|
|
f9f55075f6 | fix(vm): adopt provisioned SSH checkout roots (#14353) | ||
|
|
99d19d4635 | feat(vm): add provisioned root recipe contract (#14352) | ||
|
|
953cfab635 |
fix(terminal): make the bold font weight its own setting (#14368)
* fix(terminal): make the bold font weight its own setting
Deriving bold as max(700, regular + 200) silently destroyed bold. A family
exposes only a few real faces: the monospace the default chain resolves to on
macOS has exactly two, splitting at 600. Measured by rasterizing each weight to
a canvas — 100-500 are byte-identical (ink 3023) and 600-900 are byte-identical
(ink 3855), at every weight the same advance. So any base weight at or above 600
put both values in the same face and bold stopped existing, on 4 of the 9
positions the slider offers, with no error and nothing the user could do.
Arithmetic cannot fix it — on a two-face family there is no heavier face to
escape to. So bold is now user-owned: a new terminalFontWeightBold setting with
its own control, defaulting to 700. The default pair (500/700) straddles the
boundary, so existing profiles render exactly as before; a collision is now a
choice the user can see and undo.
The old test asserted 800 -> {800, 900} as 'keeps bold heavier', which is where
this hid: numerically heavier, identically rendered.
* fix(terminal): surface bold face collisions accurately
|
||
|
|
11cd2b4310 |
revert(ssh): back out #13326 and #13928 — reconnect loses every tab (#14361)
* Revert "fix(daemon): stop killing live coding agents when the daemon can't report its sessions (#13928)" This reverts commit |
||
|
|
281cc77e79 |
fix(ai-vault): make the merged scan stamp independent of leg order (#14270)
* fix(ai-vault): make the merged scan stamp independent of leg order
The all-host merge picked its stamp with a strict `stampMs > latestMs` and
echoed the winning leg's verbatim string. Two legs reporting the same instant
in different legal ISO shapes ("...:05Z" vs "...:05.000Z") therefore resolved
by position in the results array, i.e. by host-enumeration order (local, then
SSH, then runtime). The prior lexicographic max was order-independent, so this
was a regression with no test covering it.
A merge has no single scan instant, so its stamp is derived data rather than
any one leg's string: return the canonical ISO form of the newest accepted
instant. That is order-independent and format-independent, and drops a
variable instead of adding a tie-break branch.
Also share one request resolver between main and the renderer so the
renderer's merged-scope predicate is equivalent to main's routing by
construction, rather than by a comment that overclaimed it.
* docs(ai-vault): scope the merged-predicate comment to the desktop IPC path
The replacement comment still asserted the result is always several hosts'
legs. The paired web transport drops executionHostScope and serves one host,
so 'all' there is a single scan. State that the predicate is deliberately
over-inclusive and why erring the other way would be unsafe.
* test(ai-vault): pin the merged-stamp Date range boundary
new Date(ms).toISOString() throws RangeError outside +/-8.64e15. That is
unreachable only because Date.parse applies TimeClip, so the NaN guard alone
constrains the argument. Nothing pinned that. Dropping the guard now fails
these two cases with the RangeError they exist to prevent.
* refactor(ai-vault): route session-title scope through the shared resolver
The last character-for-character copy of the request-scope default. Leaving
it would make the shared resolver the single source of truth for two of three
sites, which is the drift this change exists to remove. No behavior change.
|
||
|
|
2f0c33757d |
fix(worker-start): match Codex effort ceilings (#14281)
Honor the advertised reasoning-effort ceilings for Codex models, preserve conservative unknown-model handling, and localize the new ultra effort label. |
||
|
|
0824ed39ea |
fix(terminal): clear the SGR pen on hidden-output restore and abandon (#14241)
* fix(terminal): clear the SGR pen on hidden-output restore and abandon The hidden-delivery gate drops renderer-bound PTY bytes while a pane has no visible view. The renderer's xterm is a separate emulator from the daemon model, so when the dropped span contains the sequence closing an attribute run (e.g. the ESC[22m ending a bold run) the renderer's pen stays latched while the daemon model stays correct. Neither recovery path cleared it: - buildMainModelSnapshotReplayWrites reset the pen on the two alt-screen branches but not on the normal-buffer branch, and replayed scrollbackAnsi ahead of the reset it did emit, so replayed content inherited the stale pen. - abandonHiddenOutputRestoreAndDrainPendingForeground declares the dropped bytes unrecoverable (it writes a user-visible warning) and then drained the queued foreground chunks straight into xterm under that same unknown pen. Add RESET_GRAPHIC_RENDITION and emit it ahead of replayed content in every branch, and on both abandon exits. The existing profiles all clear DEC mode bits and none touched SGR. * fix(terminal): also restore charset designation after a dropped-byte gap A gap can strand more than the pen: a dropped `ESC(B` leaves line-drawing selected and ordinary text renders as box characters. Route both recovery paths through one RESET_AFTER_BYTE_GAP profile covering SGR + charset. Deliberately not a soft reset (DECSTR): xterm's DECSTR wipes kitty flags and stacks (terminal-kitty-keyboard-mode-tracker applySoftReset), which would silence Option chords for a live agent that negotiates them only at startup. Reset what a gap strands and no running TUI re-asserts on its own; leave the rest to its next repaint. * fix(terminal): close the emulator state gap where the drop is announced The restore-needed marker is the single point where "renderer-bound bytes were dropped" is known. The handler already resets the transport's cross-chunk parser state there for exactly this reason — a partial escape spanning the gap would corrupt the next chunk. The emulator carries state across chunks in the same way, so reset it in the same place. That makes restore, abandon and overflow all start from a known pen by construction, instead of each recovery path having to remember. * fix(terminal): fully ground byte-gap recovery state * fix(terminal): reset state when remote restore re-arms * fix(terminal): keep the gap reset on the warning abandon path The reset had been folded into an else of the unavailable-warning branch, so the primary abandon path relied on the marker's earlier reset still standing. It does not always: this function captures a replayingSnapshot, so it can run after a partially-applied replay has already moved the pen, and the warning itself is plain text carrying no SGR. Restore the unconditional write, guarded only against the remote re-arm which writes its own. * fix(terminal): scope the byte-gap reset to the pen and skip it under flood Two regression risks in the widened recovery reset, both removed: - The profile had grown to cancel partial escapes, close OSC 8 and re-designate all four ISO 2022 registers. Each changes what a live TUI sees on a path that runs in production, and none has a reported symptom behind it — a legitimately line-drawing TUI that does not re-designate after recovery would render box characters as ASCII. Scope back to SGR, which is what the field reports show. - The marker-time reset ran before the flood-backpressure guard, so a flood wrote one reset per marker in exactly the case that guard exists to damp. Move it after; the flood path repaints through buildMainModelSnapshotReplayWrites, which grounds the pen itself, so no coverage is lost. Coverage verified non-vacuous: blanking RESET_AFTER_BYTE_GAP fails 5 tests across all four paths (replay branches, marker, abandon-with-warning, remote re-arm). |
||
|
|
b908b55f6d |
feat(worktrees): add per-source visibility controls (#14189)
* feat(worktrees): add per-source visibility controls * fix(worktrees): explain unsupported visibility hosts * fix(worktrees): keep add location form inline * fix(worktrees): align source visibility across runtimes * test(worktrees): cover Windows drive-relative roots |
||
|
|
4882eeb8ac |
rm git shim: neutralize stale wrappers without a host gate (#14255)
* Revert "fix terminal attribution shim removal edge cases (#14187)"
This reverts
|
||
|
|
4c5f818187 |
refactor(skills): remove the unreachable Skills page and the file count it rendered (#14259)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
3984023375 |
feat(workspaces): make workspace board shortcut a toggle (#14240)
* feat(workspaces): make workspace board shortcut a toggle The workspace.openBoard command only opened the board; pressing the bound shortcut again was a no-op, so closing required Escape, the toolbar button, or collapsing the sidebar. Bind the shortcut bridge event to the existing toggleWorkspaceBoard so one shortcut both opens and closes. Rename the bridge event to TOGGLE_WORKSPACE_BOARD_EVENT and retitle the command "Toggle Workspace Board". The action id stays workspace.openBoard to preserve users' stored keybinding overrides. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(keybindings): assert new toggle/open/close search keywords Cover the search-keyword additions from the toggle rename, per CodeRabbit review on #14240. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * style(keybindings): wrap workspace board search keywords for oxfmt --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> |
||
|
|
3ab8b6a117 |
fix(ssh): stop SSH reconnect from multiplying terminals and resuming agents twice (STA-3077) (#13326)
* fix(ssh): stop reconnect from grafting panes and stacking remote leases Reconnecting an SSH-backed workspace added terminal panes the user never opened, and the remote host accumulated shells nobody was using — one report went from 2 to 19 to 20 relay PTYs across three reconnects (STA-3077). Two root causes, both in the store. Reattach could create UI. `persistPtyBinding` has four creating branches — mint a tab, mint a root leaf, split the root and graft a leaf, mint a layout. All four are load-bearing for `pty:spawn`, which can beat the renderer's debounced layout writer, but none of them is appropriate on reattach, where the pane either already exists or is gone for good. Add `mayCreate`, defaulting true so the spawn path is untouched; every creating branch already sets `terminalMembershipChanged`, so refusing is a check rather than a new code path. Lease identity had no pane key. `upsertSshRemotePtyLease` matched on `(targetId, ptyId)` alone, so a pane that re-leased under a new relay id left its predecessor live with nothing to retire it, and the next reattach fanned out over both. One pane now keeps at most one live lease. Superseded leases are marked `expired` rather than terminated: losing a lease is not proof the shell died, so the remote process is deliberately left running. Tests assert observable behavior rather than mechanism, so they stay valid under any implementation that fixes this. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the terminal session behavior contract Properties stated as observable behavior rather than mechanism, so an oracle written against them survives a change of implementation. Records the weaker, correct form of the timer rule — a timer may never be the sole cause of a destructive action — because recovery budgets and scratch-file age gates are correct code that an absolute ban would condemn. Also notes which mechanisms are deliberately not required, so each has to earn its place rather than arrive with an architecture. Co-authored-by: Orca <help@stably.ai> * fix(ssh): heal duplicate pane leases that predate pane-keyed supersession Pane-keyed supersession stops new duplicates, but it does nothing for installs that already carry the ones STA-3077 accumulated — the report behind this reached 20 live leases across a handful of panes, and every reconnect fanned out over all of them. Retire the stale duplicates once per reattach pass, keeping the newest lease for each pane under a total order so two hosts resolve a tie the same way. As with supersession, retired leases are marked `expired` rather than terminated: their remote shells are deliberately left running, because a lease we chose not to revive is not evidence the shell died. The relay-session store stubs gain the new method. Note the gap this leaves open: those shells keep running and are no longer reachable from the app, so the "accumulates unused shells" half of the report needs a visible recovery surface rather than a silent kill. Co-authored-by: Orca <help@stably.ai> * fix(terminal): stop respawning a shell that is still running A pane that failed to reattach spawned a fresh shell. Because the restored session id came along, the replacement resumed the same agent session, and two processes appended to one transcript — reported repeatedly, up to five concurrent resumes of a single session. Two defects fed it. The relay reported a source that merely needed re-establishing as `SSH_SESSION_EXPIRED`. The shell was still running; only its output source was gone. Give that outcome its own error so it stops reading as "the session no longer exists". The reattach failure handler then treated every error as proof of death. It checked for expiry and, in the else branch, took the identical action — so the check bought nothing and a transport fault, a timed-out call, or a wedged relay all respawned. Respawn now requires proof: an explicit host expiry or a not-found PTY. Anything else, including an error we have never seen before, is unresolved, leaves the shell running, and keeps the binding for a later reattach. Two existing tests asserted the old behavior. One threw a bare error as scaffolding to reach the spawn-adoption door; it now throws proof, which is what it meant. The other pinned the expiry mapping itself, and now asserts the outcome fails closed *without* being reported as expiry. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record what makes a retention bound safe Shortening a grace period is the wrong lever. Measuring process time and gating reclamation on an independent observation are what make one safe, and they are what deployed systems actually do. Also records that lifecycle belongs in the attach reply rather than a delivered event — that is what removes the need for a durable per-consumer cursor to guarantee an exit is never lost. Co-authored-by: Orca <help@stably.ai> * test(terminal): assert the empty-failure case without an empty Error A thrown empty value exercises the same property — a failure carrying no usable message is not proof the session is gone — and does not trip the empty-error-message lint. Co-authored-by: Orca <help@stably.ai> * fix(ssh): let the durable pane binding outrank recency when retiring leases Choosing the newest lease for a pane is wrong whenever a newer lease exists that no pane is bound to: it retires the lease the pane is actually attached to, detaching a live terminal instead of healing it. Two changes. Arbitration now prefers the lease matching the pane's durable binding, across both the SSH-target and local partitions, falling back to recency only when no binding names either candidate. And supersession at upsert time now defers rather than expiring a bound predecessor. When a lease arrives for a pane that is still bound to a different PTY, the binding has not caught up yet, so both stay live and reattach arbitrates once the binding is available. Co-authored-by: Orca <help@stably.ai> * fix(ssh): roll back a lease retirement whose durable write fails `flush()` logs and swallows write errors, so a failed write left these leases retired in memory while disk still called them attached — and the pane bindings scrubbed alongside them stayed scrubbed. Use `flushOrThrow` and restore both the lease states and the affected session partitions when it throws, reporting nothing retired. Co-authored-by: Orca <help@stably.ai> * test(ssh): prove pane and remote PTY cardinality across reconnects Counts the shells the relay actually hosts, on the container, rather than inferring them from app state — that is the census the report was based on. Asserts the PIDs are unchanged, not merely the count, so a kill-and-respawn cannot pass. Every pane streams before the transport is severed: an idle pane sends no recovery checkpoint, so only a live source comes back needing re-establishment, which is the outcome that used to read as expiry. Co-authored-by: Orca <help@stably.ai> * fix(ssh): actually pass mayCreate:false from the reattach binding write The `mayCreate` guard was correct and had no production caller, so the reattach path still went through the creating branches and grafted panes back. `restoreReattachedPtyRuntime` is that call site — RC3 in the original diagnosis — and it now refuses to create. Binding moves ahead of runtime registration, because registering first would surface a pane the user never opened before the refusal landed. A refusal leaves the remote shell running and reattachable; a *thrown* write stays unknown and still registers, so a failed disk write cannot detach a live pane. Adds an oracle over the call site itself. The store-level tests all passed while the fix was inert, because they called the store directly — only pinning the wiring catches that. Co-authored-by: Orca <help@stably.ai> * fix(terminal): apply the respawn-requires-proof rule to both reattach paths connectPanePty has two near-verbatim reattach blocks — one keyed on the deferred SSH session, one on the restored session — and only the second was fixed. The first still checked for expiry and then respawned unconditionally anyway, so a transport fault there resumed the same agent session a second time. Also keep the wire token out of the pane. The main-process bridge only special-cases expiry, so a source-restore failure crossed IPC as raw `SSH_SOURCE_RESTORE_REQUIRED: <id>` text and surfaced to the user. It correctly does not respawn; it just should not read like that. Co-authored-by: Orca <help@stably.ai> * test(ssh): state plainly that the reconnect spec is a forward guard It was run against an unfixed tree and passed, so it does not prove the STA-3077 fixes and should not be read as if it does. A clean severed transport does not reproduce the field conditions — accumulated duplicate leases, or a source returning needing re-establishment. It keeps its place as a forward guard: it counts the shells the relay actually hosts and pins their PIDs, so a later change that grafts a pane or respawns a shell fails here. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record that a guard must be pinned at its call site A refusal that exists and is never passed is indistinguishable from no refusal, and store-level tests cannot tell the difference — they call the store directly. Learned from `mayCreate`, which was correct and had no production caller for several commits. Co-authored-by: Orca <help@stably.ai> * fix(ssh): park one PTY's exhausted delivery recovery instead of dropping the channel A per-PTY recovery budget running out disposed the whole relay channel, so one PTY that could not re-prove its delivery aborted every in-flight filesystem and git request on that host and stalled every sibling pane. A retry count is not proof of anything, and it certainly is not proof about the other sessions sharing the channel. Exhaustion now parks that PTY's delivery. The remote shell keeps running, its lease stands, and the next relay open reattaches it with a fresh delivery generation — the parked state is cleared on teardown and the generation changes on reconnect, so a reconnect recovers it. The consecutive-attempt ceiling goes away entirely; the per-generation one is what bounds the retry cost, and the second ceiling only existed to reach the channel drop sooner. Tradeoff worth stating: the failing pane used to self-heal within seconds because the forced reconnect wiped all rejection state, and it now stays frozen until the next relay open. That is a worse outcome for that one pane and a much better one for every other session on the host, and reconnecting is user-reachable. Co-authored-by: Orca <help@stably.ai> * fix(pty): let liveness say unknown instead of forcing it to say dead `IPtyProvider.hasPty` returned a boolean, so a provider whose inventory was empty for reasons that have nothing to do with the session — socket down, cache never hydrated, provider generation just constructed — had no way to say so and answered "absent". Its own siblings already knew better: `probePtyLiveness` and the runtime's `PtyController.hasPty` were both already `boolean | null`, with consumers branching on null correctly. The lie was injected at exactly one interface. Now three-valued, and each provider answers unknown where it cannot prove absence: the daemon adapter off-socket, the SSH provider before a completed listing, the router when any adapter cannot answer, and the degraded provider rather than fabricating a verdict. `terminal_gone` requires unanimous proven absence. Also fixes a real cold-start bug this surfaced: `pty:hasPty` never awaited the daemon-swap startup promise, though the sibling `probePtyLiveness` bridge already did, so before the swap the local provider answered an authoritative false for every daemon-owned id. Net +27 production lines. The plan behind this predicted -92 on the strength of deleting the renderer's dead-session reconcile path; that code is live (`pty-connection.ts` imports it), so nothing was deleted. Expressing a third value where there were two costs lines, and a deletion that is not real is not worth manufacturing. Co-authored-by: Orca <help@stably.ai> * docs(terminal): track the terminal-session correctness handoff package The package was untracked under a gitignored `docs/**`, with the un-ignore rules living only in an uncommitted .gitignore edit — a single `git clean -xdf` would have destroyed the authoritative plan. The 814-path construction snapshot is now pushed as `nwparker/react185-authority-snapshot` too; it had no remote ref. Co-authored-by: Orca <help@stably.ai> * test(ssh): make the reconnect settle window actually wait The settle poll reused a matcher the assertion 15 lines above had already satisfied, and Playwright's poll engine probes immediately and returns as soon as the matcher passes — so it observed the same state twice and elapsed 0ms. A shell grafted a second or two after reattach reported ready slipped through into the next cycle. Reviewer was right on #13111. Test-only; no production change. Co-authored-by: Orca <help@stably.ai> * test(ssh): census both durable session partitions on reconnect Adds a second reconnect scenario and a helper that reads pane records from the local partition as well as the ssh host partition. That split matters: the reattach binding call passes no hostId, so a grafted pane lands in the LOCAL partition and an oracle reading only the host partition passes whether or not the guard is present. Both tests remain forward guards. The second one was reported as discriminating and did not reproduce: with `mayCreate: false` removed from the call site and the app rebuilt, both still passed. Its induction races `pty:kill` against a severed transport, so when the kill lands the lease is cleaned up and there is nothing left to graft. The handoff README is corrected to say so rather than claim a journey. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the user decision relaxing G6 G6 becomes minimise-and-justify rather than strictly net-negative. The deletion budget the plan assumed does not exist: an entrypoint-rooted import graph found 51 of 53 candidate files reachable and instantiated on live paths, leaving 263 deletable LOC against roughly +1,021 to offset. Correctness may still not be traded for line count. Co-authored-by: Orca <help@stably.ai> * test(terminal): add discriminating oracles for restart, daemon, skew and namespaces Six parallel streams, each required to fail with its guard removed rather than merely pass. Local restart proves the OS process itself survives, by reading `ps -o lstart=` for the shell's own pid. That matters: with the quit path made destructive, the tab, leaf and pty ids all came back byte-identical while the shell underneath was a new process — every existing restart spec would have stayed green. Two separate guards were removed to redden it, and the second reddens only the stale-operation case. Daemon restart discriminates by reverting three-valued `hasPty`; version skew now covers publication semantics and confirms the new `SSH_SOURCE_RESTORE_REQUIRED` token mutates nothing on an old client; two-host isolation censuses both containers. Deletes `src/relay/pty-source-replay-index.ts` — 201 production lines with no importer outside its own test, verified against an entrypoint-rooted import graph rather than a name grep. Five namespace tests are skipped, not passing: they reproduce a defect still live on main where folder-workspace ids compare equal with the instance suffix stripped. PR #12474 fixes it; they are its oracle. Co-authored-by: Orca <help@stably.ai> * test(ssh): induce the reattach graft deterministically instead of racing a kill The previous induction closed a pane while the transport was severed and relied on `pty:kill` FAILING so the lease outlived the pane record. It does not fail: with the provider already torn down, `pty:kill` takes its tombstone branch and marks the lease terminated, and `reattachKnownPtys` filters terminated leases out of the fan-out — so the reconnect never visited the PTY the test was about. It passed on both trees. Seed the precondition instead. Spawn a real remote PTY on a leaf that never becomes a pane, then roll the host partition back to its pre-spawn snapshot, leaving a live lease and a live remote shell that no durable pane owns. No failure races a success. Adds a vacuity guard that is independent of the tree under test: the lease's own `lastAttachedAt` must advance, proving the fan-out actually visited this lease before the pane census is trusted. Verified on this machine under an isolated TMPDIR, since the e2e harness keys its seeded-repo pointer on a machine-global tmpdir path: guard present passes, guard removed fails with the phantom leaf grafted into the local partition, guard restored passes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): propose one authoritative binding identity Every defect this program has touched is the same defect: identity compared with the wrong key, or not compared at all. Lease keyed without the pane, reattach using a creating write, folder-workspace ids compared with the instance suffix stripped, local mutating IPC carrying only an id, a live shell classified as expired, liveness unable to say unknown. Proposal: one branded binding type built from fields that already exist and are already persisted, constructible only from an authoritative source, carried by mutating operations, compared by one shared function. Makes a wrong-key comparison a type error rather than the next incident. Under adversarial review, including against the open issue corpus. Not accepted. Co-authored-by: Orca <help@stably.ai> * fix(pty): refuse mutating operations aimed at a superseded PTY `pty:write`, `pty:writeAccepted` and `pty:resize` accepted any id. The renderer queues input, so a keystroke buffered before a reattach landed on whatever PTY had since taken the pane — and a resize reshaped the successor's shell. Main already tracks `ptyPaneKey` and `paneKeyPtyId` in lock-step, so their disagreement is proof the caller's id was superseded. No wire change, no renderer change, nothing added to the input payload. An id with no recorded pane stays permitted: unowned and orphaned PTYs are unknown, not stale, and unknown never authorizes refusing an explicit operation. That is also what keeps orphan cleanup working — those ids have no pane by construction. The tests pin the CALL SITES, not the predicate. A capability that exists and is never called is indistinguishable from no capability, which is exactly how `mayCreate` sat inert here for several commits with every test green. Co-authored-by: Orca <help@stably.ai> * fix(pty): fence signals at a superseded PTY, and pin why kill is exempt A signal means "interrupt my pane", so delivering one to a PTY the pane has already replaced is a misdirected interrupt. Fence it with the same lock-step proof used for write and resize. `pty:kill` stays deliberately unfenced and a test now pins that: a superseded PTY is orphaned, and reclaiming it is exactly what the orphan-cleanup callers ask for. Refusing there would break the operation that reclaims leaked shells — the opposite of the intent. The fence sits at the IPC boundary, above `tryGetProviderForPty`, so it covers local, daemon and SSH rather than the local path alone. Co-authored-by: Orca <help@stably.ai> * test(terminal): poll the pane binding read so a slower host cannot flake it `readPaneBinding` took a single unpolled read of a DOM dataset attribute immediately after a renderer reload, while its sibling helper polls the same data for 15s. On a native Linux host both tests failed every run with 'No bound terminal pane is mounted' while the app was demonstrably healthy — the screenshot showed the terminal restored with a live prompt and the boot PID echoed. The assertion is unchanged; it is only awaited. Nothing is weakened. Found by running this spec on native Linux rather than assuming macOS behaviour generalises. Co-authored-by: Orca <help@stably.ai> * test(terminal): make the restart identity spec run on Windows too Both probes were POSIX-only and unconditional: `echo ...=\$\$` for the shell's own pid, and `ps -o lstart=` for its start time. Running the spec on a real Windows host proved it dies before reaching either guard, so Journey 1's Windows half was unprovable rather than merely unproven. PowerShell exposes the same two facts as `$PID` and `Get-Process` StartTime. The start time still matters on both platforms for the same reason: a PID alone cannot separate a survivor from a reused number. Still green on macOS. The Windows path is written from the host probe and has not itself been executed end to end — that is the next thing to run there, not a claim being made here. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the fence's real gap and what peer designs taught Marks the client-constructed binding proposal as rejected with the three false claims that sank it, and records what shipped instead. States the shipped fence's actual limitation rather than leaving it implied: it compares a binding, not an incarnation, so a respawn under a reused ptyId passes. The obvious remedy is wrong here — the agent-create id is deterministic by design so a replayed create stays idempotent, and randomising it would trade this narrow gap for a duplicate-spawn bug. Also records the ranked lessons from four comparable agent IDEs, chiefly that a typed end-reason at end time is what stops a user quit from looking like a resume candidate. Co-authored-by: Orca <help@stably.ai> * docs(terminal): promote Journey 1 to proven on all three platforms The oracle now runs natively on macOS, Linux and Windows, and its discrimination was watched on each: a mutation reddens it, a restore greens it. On Linux and Windows both mutations were run, and the second reddens only the stale-operation test — so the journey's two clauses are proved independently rather than jointly. Windows is the new evidence. The PowerShell branches added blind at ebffb85a848 executed correctly on their first run: `$PID` expanded to real integers, which also proves the pane shell there is PowerShell-family rather than Git Bash, and `Get-Process StartTime` returned kernel start times 5.4s apart — so a recycled pid could not have passed as a survivor. First journey promoted in this program. The other twelve are unchanged, and the residual limit on "every stale exact operation" is recorded rather than glossed. Co-authored-by: Orca <help@stably.ai> * test(terminal): add discriminating oracles for the daemon, skew and multi-host journeys Daemon: replaces a spec that modelled only a client restart and never crossed the daemon boundary, whose successor generation owned nothing so "the live successor is neither killed nor replaced" was vacuous. The PTY leader is now a real login shell reporting `$$` back through the production write path, resolved to a kernel start time. Two mutations each redden exactly one of the three clauses, on macOS and Linux: reverting three-valued `hasPty` reddens only the unknown-not-dead clause; widening the sole-provider fallback reddens only the stale generation clause. Skew: reverting the restore-required publication to expiry reddens 4 of 5 new tests while the legacy control stays green — the regression this branch fixed is now caught if reintroduced. Multi-host: restoring `mux.dispose('connection_lost')` reddens sibling isolation on one host. It does NOT redden across hosts, and that is recorded rather than glossed: a mux belongs to one relay session per target, so its dispose cannot cross a host boundary. Journey 4's cross-host clause rests on isolation-by-construction, not on a mutation. No production code changes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record journey evidence that falls short of promotion Four journeys now have discriminating oracles but none meets its full stated scope, and each shortfall is named rather than rounded up. Journey 2 is one WSL run from promotion. Journey 12's tests are in-process, so they do not close the live-skew gap the original ledger named. Journey 4's cross-host clause cannot be proven by mutation at all — a mux is per target, so its dispose cannot cross hosts, and the cross-host test stayed green under the mutation that reddens siblings. Journey 13 measured one dimension of ten, on lifted predicates rather than through real IPC. Co-authored-by: Orca <help@stably.ai> * docs(terminal): promote Journey 2 to proven on macOS, Linux and physical WSL The oracle runs on every environment the journey names, and is clause-selective on all three: reverting three-valued `hasPty` reddens only the unknown-not-dead clause, and widening the sole-provider fallback reddens only the stale-generation clause. Selectivity in WSL was established rather than assumed. The spec runs serially, so a red first test reports the others as "did not run" — they were re-run alone under the same mutation and stayed green. Also records that an Orca WSL-mode terminal now starts on that host at all, which it could not before: the distro had no provisioned default Unix user, so every interactive launch blocked on first-run setup. One diagnosis from the WSL run is corrected here rather than repeated: the unrelated `local-pty-shell-ready` failure was attributed to bash 5.3.9, but macOS runs the same bash version and passes 67/67. The trigger is environmental to that distro, and the underlying defect is that the spec pins an absolute count of OSC markers it does not own. Co-authored-by: Orca <help@stably.ai> * docs(terminal): correct the WSL provider-suite diagnosis The WSL run blamed bash 5.3.9 for the unrelated `local-pty-shell-ready` failure. macOS runs the same bash version and passes 67/67, so the version is not the cause — the trigger is environmental to that distro, and the underlying defect is that the spec asserts an absolute count of OSC markers it does not own. Co-authored-by: Orca <help@stably.ai> * test(runtime): unskip the workspace-namespace oracles now their fix has merged These five reproduced a defect that was live on main: folder-workspace ids were compared with the instance suffix stripped, so two workspaces sharing a directory read as the same namespace. They were committed skipped, pointing at the PR that fixes it. That PR is merged, and they pass. Verified they still bite: restoring the suffix-stripping comparison reddens exactly these five and leaves the other four green. An oracle written before its fix, held skipped, and confirmed against the fix after the merge — rather than deleted and rewritten from the answer. Co-authored-by: Orca <help@stably.ai> * test(ssh): add MaxSessions, lazy-discovery and paired-skew oracles Three journeys attempted; none promoted, and the reasons are recorded in the ledger rather than rounded up. MaxSessions=1 against real OpenSSH, with the cap read back from `sshd -T` rather than assumed, and remote pids read on the container two independent ways that must agree, each carrying its kernel start time. Two disjoint mutations discriminate — one reddens only the reconnect clause, the other only the two restart clauses. But the disconnect clause is a forward guard: four separate guard removals left it green, so nothing shipped is load-bearing for it. Lazy discovery samples sshd's own accept log and live session census across a 22s window with the in-use host as a positive control. No mutation reddens its third clause alone — the real cross-host lease scoping is load-bearing, but removing it breaks the sibling host during setup, so the failure carries no clause information. The paired-runtime skew spec pairs two real processes at different versions and refuses to run rather than degrade into a same-version pairing that would look green and prove nothing. No production code changes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record why the duplicate-resume fix was not built I recommended adding a typed end-reason so a user quit stops looking like a resume candidate, then went to implement it and stopped. `SleepingAgentSessionRecord` already carries three fields that each exist to stop something resuming that should not have — `origin`, `restoreOnTabOpenOnly`, and `automaticResumeBlockedBy` — each traceable to its own incident, consulted at 22 non-test sites. A fourth predicate, however well typed, is the fifth containment cycle. The designs without this bug do not have a better flag; they resume only on an explicit action, into a new terminal id, and make two agents in one terminal unrepresentable in the schema. The first of those is a product decision about whether automatic resume stays a feature, so it is the user's call rather than mine. Co-authored-by: Orca <help@stably.ai> * docs(terminal): reconcile G6 with the recorded decision and assess its clauses G6's body still demanded strictly-negative production LOC after the user relaxed it to minimise-and-justify, so the gate had two conflicting pass conditions and no single truth value. Its body now points at that decision. Assessed the remaining clauses against the branch rather than assuming. Two fail structurally: more than one identity comparison and mutation admission path still exist, and `terminal-input-quarantine.ts` is still reachable from two production files. Records why the quarantine is not subsumed by the superseded-PTY fence, which I had assumed and checked. The fence refuses writes aimed at a stale ptyId; the quarantine guards the user's next keystrokes landing on the successor under its current, correct id — a case the fence never sees. Removing it needs the recovery path to surface a different shell as unresolved, not a deletion. Co-authored-by: Orca <help@stably.ai> * docs(terminal): the input quarantine is load-bearing, not superseded G6 lists "no superseded quarantine remains reachable" and this module was assumed to be one. Disabling its single call site reproduces the hazard it exists for — `cho hi; rm -rf x` reaching the shell — so deleting it without a replacement re-opens command execution. The replacement was costed by building it rather than estimated: +26 production LOC to thread the incarnation, ~+33 complete, and the cross-remount state it needs outlives the destroyed pane so it becomes a module about the size of the one deleted. Floor is roughly +140 to delete 88, and it would add a second identity comparison to a gate already failing for having more than one. The decisive part is that the route is not uniformly available: remote runtime results carry no incarnation, old hosts cannot be made to publish one, and mixed versions are the normal state. A paired client reads unknown, which this program's own rule says is not proof — so either every remote reattach surfaces unresolved, or a fallback is needed and the only correct fallback is this module. Whether to amend the clause or accept something weaker on remote hosts is a user decision, so the clause verdict is left as failing rather than quietly reclassified. Co-authored-by: Orca <help@stably.ai> * refactor(runtime): collapse duplicate identity comparisons G6 requires one identity comparison; five implementations existed across two concepts. Worktree-namespace identity had two: `runtimeWorktreeIdsEqual` and `runtimeWorktreeIdentityKey` independently re-derived repoId plus normalized path. Equality now derives from the key, so the comparison and the sleep / mutation-queue keying cannot drift into two different rules — which is exactly how the suffix-stripping bug reached production once. Pane identity had three byte-identical leaf-UUID comparisons, in orchestration `db.ts`, `lifecycle-reconciliation.ts`, and `orchestration-legacy-process-identity.ts`. One copy moved to `stable-pane-id.ts`, which already owns `PaneKey`, `parsePaneKey` and `makePaneKey` and which all three already imported. No new module, no branded type, no parallel comparison. Net -14 production lines. The namespace oracle still bites: restoring the filesystem parser inside the identity key reddens exactly its five cases. The raw counts are not the actionable set, and the classification is worth recording: of 409 non-test `worktreeId` comparisons, 71 are typeof guards and 81 are sentinel tag checks. Most of the remainder are renderer predicates over store rows where both operands are the same main-minted id, so normalizing there would widen equality rather than correct it. Co-authored-by: Orca <help@stably.ai> * refactor(terminal): finish a half-done fixture move and audit the rest `xterm-bypass-event-fixture.ts` and `__fixtures__/xterm-bypass-event.ts` were byte-identical apart from an import path. The `__fixtures__` copy had zero importers and the live copy compiled as production — someone started the move and left both. Dead copy deleted, live one moved, its three test importers updated. Audited the wider G6 clause by importer rather than filename: 32 test-only files, roughly 3,300 LOC, currently compile as production; 4 of the 36 candidates have real production importers and are correctly placed. The list is recorded in the goalposts. Those 32 are almost all older than this program and outside the terminal surface, so sweeping them belongs in its own change rather than inside a terminal PR. The clause stays failing, with the remaining files named. Co-authored-by: Orca <help@stably.ai> * docs(terminal): the fixture clause already holds where it matters Checked what the build emits rather than reasoning from file paths. None of the 32 test-only fixtures appears in `out/` — Rollup drops them because no production entrypoint reaches them. On "compiles into the shipped product", this clause holds today. On the other reading it cannot be closed by moving files at all: both production tsconfigs use bare `include` globs with no `exclude`, so a `__tests__/` directory matches exactly like any other path, as does every `*.test.ts` in the repo. Relocating 32 fixtures would remove nothing from typecheck scope. A sweep was started and stopped once this was verified, rather than landing 32 moves across areas this program does not own for no gain. If the intent is that typecheck scope should exclude test code, that is a repo-wide tsconfig change with a different owner. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add plain-language design and test overviews Two reviewable documents with diagrams, written so someone with no prior context can follow what breaks, why, and what changed. The design overview explains the five things stacked behind one terminal rectangle, the 2 -> 19 -> 20 report, the three root causes, and the rule underneath all of them: unknown is not dead. The test overview explains why a green test proves nothing on its own, the four-step mutation proof we adopted, and — the part worth reviewing hardest — an honest account of what could not be proven and why, including the properties that are true by construction and therefore have no guard to remove. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add a self-contained visual report of the design and its evidence Pre-renders every diagram to inline SVG in both themes so the report opens offline and stays sharp when zoomed. States the gate/journey score and the retractions alongside the fixes, so the unproven half is as visible as the proven half. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record the finalized two-plane architecture decision Adopts the data-plane proposal and adds the control-plane track it does not cover: re-key ownership by pane, split orphan inventory out, then delete the compensating code. Records that the host-authority alternative was refuted and that the shipped keystroke fence is inert on the reattach path. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add the design brief the review counsel works from Separates verified code facts from unverified leads so reviewers attack the design rather than a reconstruction of it, and records which simpler alternatives were already refuted and why. Co-authored-by: Orca <help@stably.ai> * docs(terminal): report the design counsel's outcome and the live respawn bug it found Three review rounds across two models replaced the two-record split with one leaf-keyed record, deleted attach-time pane identity, and made orphans a connect-time projection. Records that a shipped gesture still turns a healthy remote shell into a duplicate agent resume, and that the renderer classifier in that chain treats an error-message shape as proof of death. Co-authored-by: Orca <help@stably.ai> * docs(terminal): correct the report — the respawn proof gate guards a minority path A final review traced every auto-respawn route. The primary one converts the reattach failure into a boolean before any classifier sees it, so the shipped proof gate never runs there. Records that two of the six shipped changes are narrower than claimed, and why their tests could not have caught it. Co-authored-by: Orca <help@stably.ai> * docs(terminal): explain the landed design on its own terms One leaf-keyed ownership record, orphans computed at connect, and replacement shells only on positive proof — with the shipping order and the one product trade the design asks the owner to accept. Co-authored-by: Orca <help@stably.ai> * docs(terminal): rewrite the design explainer in plain English The first version assumed the reader knew the codebase. Reframed around two bugs, two fixes and one decision, with the jargon replaced by pane / program / note / helper and a five-word glossary for what could not be avoided. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop reading an identity mismatch as a dead shell The relay reports a pane-identity mismatch by saying the pty was not found, but it found it — comparing identity is how it noticed. Publishing that as expiry made the renderer clear the binding and cold-restore with agent resume, so a live shell gained a second agent on one transcript. Reachable today by detaching a pane into a new tab, which changes the tab the relay froze at spawn. Mismatch now carries its own token and the classifier refuses it as proof. Genuine absence still expires, so a shell that really went away is not stranded. The three failure tokens move to src/shared: main published them and the renderer decided respawn on them, from two copies that had drifted apart. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop sending pane identity on reattach The relay froze pane identity at spawn, so moving a pane to another tab made it refuse a live shell — and refuse by saying 'not found'. The comparison is presence-guarded, so not sending the fields disarms it on every relay version including ones already installed on hosts: no wire change, no redeploy. Nothing is lost. It existed to catch a relay restart recycling pty-N for a new shell, and in exactly that case pane and tab both still match, so it accepted the wrong shell anyway. The incarnation the attach returns is what distinguishes those, and it already crosses the wire. Removes the whole client-side apparatus: the expected-identity type, its per-lease derivation, its map, and the parameter threaded through four layers. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add tracked goalposts for the new design Each goalpost is a behaviour with an oracle and the mutation that must redden it, so 'proven' cannot be claimed from a green test. Records the anti-inert rule as a first-class goalpost, since three guards in this program passed their tests while sitting off the route production takes. Co-authored-by: Orca <help@stably.ai> * docs(terminal): record that the recovery grant is dead code, deleting a design step The lease stores a relay-native pty id and the caller passes the app form, with a raw equality comparison between them, so the 30s grant cannot fire for a real SSH pane. The death rule that existed to referee it is deleted rather than built, and the dead path itself becomes a removal. Co-authored-by: Orca <help@stably.ai> * docs(terminal): keep the full design detail in the repo It only existed in an ephemeral job directory, so the plain-English explainer had no durable source for its specifics — record shape, death rule, reattach algorithm, migration order and the 25 oracles. Co-authored-by: Orca <help@stably.ai> * docs(terminal): add a resume prompt for a clean session Points at the goalposts as the contract, names the three goalposts whose oracles are already written and red, and carries the process rules that were learned the expensive way — prove guards reachable, verify mutations land, commit per step, and never let a subagent write production files in a shared worktree. Co-authored-by: Orca <help@stably.ai> * test(ssh): add the failing oracles for goalposts S3, S4 and S5 Intentionally RED: 14 clauses that fail against current behaviour and go green under the changes named in new-design-goalposts.md. The branch is held unmerged, so red here means unimplemented, not broken. Each was verified to fail for the right reason and to flip green under the identified fix, which was then reverted. Each pins the producer as well as the consumer, so no clause can pass vacuously if its route is ever severed — the failure mode that let three earlier guards ship inert. Co-authored-by: Orca <help@stably.ai> * fix(ssh): stop fabricating an exit when a reattach fails A failed attach never proves the shell exited. The relay answers not-found for a pane-identity mismatch and for any id it merely cannot hand back, so treating it as death sent the pane a synthetic `pty:exit { code: -1 }`, cleared provider state, deleted ownership and expired the lease — four claims about a process we know nothing about, on a shell that is usually still running. Collapse every failure into the non-destructive branch that already existed a few lines above (`restoreRequired = 'reattachAttemptsExhausted'` + wakeRecovery). A branch collapse, not a new mechanism: goalpost S3. Two tests pinned the deleted premise and are INVERTED rather than patched, so the new intent stays covered: - ssh-relay-orphan-abandon-paths: "retires the lease without a kill when the relay proves the PTY is gone" -> "leaves the shell running when the relay only reports the PTY as not found". Its comment claimed attach verifies liveness before answering not-found; it does not. - ssh-relay-session: "invalidates and broadcasts remote PTYs that cannot reattach" -> "leaves an unreattachable remote PTY alone while its sibling reattaches". Also repairs two clauses left red by |
||
|
|
ede69ffc7f |
perf(skills): bound and share skill discovery scans (#14204)
Skill discovery re-walked every skill root on every window focus, pane mount, and connected client. The root set was already bounded; what was not bounded was how often and how redundantly it was walked. - Focus called refresh(true), bypassing every cache down to a disk walk. - The process that owns the disk had no cache and no in-flight dedup. - Panes with different cwds each re-walked the same 12 home roots. - Fan-out inside a scan was unbounded, and every package was walked twice (once to find SKILL.md, once to count its files, node_modules included). Adds one coalescing primitive — in-flight dedup plus a short TTL behind a bounded LRU — used for per-target dedup below both the IPC and RPC entry points, per-root sharing on the native path, and whole-result reuse on the WSL path. A scan may publish only while it still owns its pending slot, so a scan begun before an invalidation can never re-cache a pre-mutation result. Bounds per-skill fan-out to the existing candidate concurrency limit, and bounds the package file walk by depth with a node_modules prune. Focus now reads through a 15s freshness window; explicit signals (install completed, Settings Refresh, native-chat Retry, terminal exit) set a new optional `refresh` wire field that bypasses every cache, including on remote runtimes. Measured on a 32-concurrent-scan burst across 8 workspaces: 134,880 -> 2,956 filesystem calls and 1080ms -> 45ms, same 31 skills returned. |
||
|
|
0f51d0b3bb |
Fix native chat image marker position handling (#14162)
* fix(native-chat): handle image markers in any position * fix(native-chat): preserve image caption whitespace * test(native-chat): cover marker boundary spacing * fix(mobile): normalize image echo reconciliation * fix(mobile): use idiomatic tail access * refactor(native-chat): share image echo matching * perf(native-chat): avoid unchanged block copies |
||
|
|
0ed6db77cf |
fix(mobile): open agent-cited external chat files (#14166)
* fix(mobile): open agent-cited external chat files * fix(mobile): keep cited external files read-only * refactor(mobile): derive cited-file mode from provenance * fix(mobile): accept sentence-final cited paths * fix(mobile): preserve cited SSH grant scope * refactor(file-links): share location suffix parsing |
||
|
|
af7dcdc196 |
feat(dashboard): identify SSH and remote hosts (#14177)
* feat(dashboard): identify SSH and remote hosts * fix(dashboard): resolve host labels consistently * fix(dashboard): reuse host server icon * test(dashboard): guard host label refresh cost * test(dashboard): satisfy runtime environment shape |
||
|
|
7c93aed6dc |
Fix fsync of read-only files on POSIX (#14235)
* fix(files): fsync read-only files on POSIX * test(e2e): add golden E2E tests for POSIX profile index fsync Validates that profile index files are properly persisted on POSIX systems, including with restrictive umask settings. These are release-blocking golden tests for Linux and macOS. * test(terminal): wait for fish child ownership before stdin write Fish 4.8 withdraws DECSET 2031 before spawning the child, so the shell-contracts harness could send hello into an intermediate prompt and hang waiting for CHILD-READ. Wait for the child's CHILD-READY marker and answer split DA1/CPR/OSC queries across chunk boundaries. * test(e2e): verify profile index persists to disk with restrictive umask Strengthen the POSIX fsync test to verify the rebuilt index is actually written to disk and has correct permissions under a restrictive umask, not just cached in memory. |
||
|
|
a73f122c61 |
fix(store): keep project catalog identity when repo.addedAt is 0
* fix(store): keep project catalog identity when repo.addedAt is 0 projectHostSetupProjectionFromRepos used `repo.addedAt || now`, so a missing or zero addedAt stamped Date.now() into createdAt/updatedAt on every refresh. #13803's reconcile then treated the project as changed and never reused it. Use a finite check with a stable 0 fallback, and treat 0 as unknown in merge so a persisted createdAt is not wiped. Refresh-identity tests use production-shaped nested fixtures plus structuredClone and go red if the fallback is reverted. Co-authored-by: Orca <help@stably.ai> * type(shared): return readonly setups from getProjectHostSetupsForProject The helper already takes a readonly catalog and returns a filter subset. Mark the return readonly so callers cannot mutate a live setups array. Co-authored-by: Orca <help@stably.ai> * style: oxfmt projection files and drop stale addedAt comment oxfmt --check failed on the addedAt identity commit. Also stop claiming the projection still restamps Date.now() when addedAt is 0. Co-authored-by: Orca <help@stably.ai> * fix(store): treat createdAt 0 as unknown when merging sibling repos Repo order decided a project's createdAt: a zero-addedAt repo seeded the accumulator with 0, and the merge only treated the *incoming* addedAt as unknown, so min(0, 100) kept 0 when the unknown sibling came first. Share unknown-aware mergeCatalogCreatedAt/mergeCatalogUpdatedAt helpers and apply them on both sides of the projection merge and of the renderer's cross-host mergeProjectCompatibilityProject, which had the same 0-poisoning via Math.min(base.createdAt, overlay.createdAt). Co-authored-by: Orca <help@stably.ai> * test(store): pass a valid updateProject payload in createdAt merge case updateProject only accepts localWindowsRuntimePreference. The new unknown-vs-known createdAt test used displayName and failed typecheck. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
d41cb21e94 |
Sta 4062 folder note rollback (#14232)
* fix(persistence): keep folder-workspace notes across a build rollback normalizeFolderWorkspaces rebuilds each FolderWorkspace field-by-field, so the inline diffComments field #14112 added is dropped by any build that predates it — and the next full-state write makes the loss durable with no user edit. Move the on-disk home to an optional top-level PersistedState.folderWorkspaceDiffComments, which older builds round-trip untouched through their {...defaults, ...parsed} load spread and omit-style getDurableState(). load() hydrates it onto the records and deletes it from Store state; buildStateToSave() is the only producer. The in-memory FolderWorkspace shape, and therefore every IPC/RPC/renderer/mobile path, is unchanged. Co-authored-by: Orca <help@stably.ai> * fix(persistence): prefer inline folder notes over a stale map entry Hydrate preferred a non-empty folderWorkspaceDiffComments entry over non-empty inline notes. A rollback to a notes-capable #14112 build writes notes inline and leaves the older map untouched, so re-upgrading deleted everything authored while rolled back. Inline now wins when present; the map only fills a stripped record. Co-authored-by: Orca <help@stably.ai> * Extract folder workspace diff comments to dedicated module Moves normalizeFolderWorkspaceDiffComments and collectFolderWorkspaceDiffComments from persistence.ts to a new folder-workspace-diff-comments.ts module for better code organization. --------- Co-authored-by: Orca <help@stably.ai> |