Commit Graph
17 Commits
Author SHA1 Message Date
Neil 8d8b9dad78 fix: keep macOS shell ownership proof within recovery budget (#18932)
* fix: keep macOS shell ownership proof within recovery budget

* fix: parse the shell-proof column set with its own anchored parser

The narrower macOS capture (`pid ppid pgid tpgid stat command`) was fed to the
shared lenient parser, whose optional tty/start pair has no `tty=` column left
to absorb it. It then eats the head of any argv shaped `python 3 app.py`
(parsing command as `app.py`, tty as `/usr/bin/python`), and turns a
command-less row into a garbage pid/stat pair. Either can flip a shell
ownership verdict, which is what gates dead-TUI recovery.

Give the column set a named constant and a parser anchored to exactly those
six columns, beside its `CHEAP_PS_ARGS` sibling. A capture that yields no rows
now raises `empty_capture` rather than reading as a machine with no processes.

Update the `confirmShellForegroundProcess` fixtures from the 4-column legacy
shape to the 6 columns the darwin reader actually emits; that describe block
already forces `platform=darwin`, so the stale fixtures were failing.
2026-09-06 18:28:54 -07:00
Neil 0bbf86bb7a fix(renderer): stop a Node-only process-table module blanking the app at boot (#18814)
Every Electron E2E spec that boots the app has been failing on
`workspaceSessionReady did not become true`, and the app itself has been
launching to a blank white window: the renderer threw
`ReferenceError: process is not defined` while evaluating a shared chunk,
so React never mounted and no startup step ever ran.

`agent-completion-poll-interval.ts` (renderer) imported one constant,
`PROCESS_TABLE_SNAPSHOT_MAX_STALENESS_MS`, out of
`shared/process-table-snapshot-reader.ts` — a `node:child_process` /
`node:fs/promises` module whose dependency evaluates `process.platform` at
module scope to pick `ps` columns. The renderer runs sandboxed with
contextIsolation, where `process` is undefined, so that module-scope read
threw and took the whole chunk with it. Introduced by #18742; #18780 added a
second module-scope read next to the first.

The constant now lives in `shared/process-table-snapshot.ts`, the
environment-neutral half of the pair, and the reader re-exports it so host
callers are unchanged. The two `ps` column sets read the platform behind a
`typeof process` guard, which defuses the same landmine for any future
renderer import of that module — only hosts ever run the argv.

The regression test walks the renderer import graph (lazy routes included)
from all three entries and refuses any module that reaches a `node:` builtin.
It fails on the pre-fix import with the full 10-hop chain from `main.tsx`.
2026-09-05 02:07:37 -07:00
Neil e95d247be1 perf(terminal): cheap-tier process inspection for anchored local agent panes (#18780)
* perf(terminal): cheap-tier process inspection for anchored local agent panes

Every idle local pane's completion cadence ran a full whole-host `ps` (with
`tty=` and `command=`, 0.34-0.50s on a 1,900-process Mac, 1.15s on Linux)
purely to build `foregroundProcessEvidence` that the renderer then discards
for local ids. Add a cheap tier (same job-control columns, no tty/command,
0.03s) gated so that it introduces no user-facing trade-off:

- Only a pane whose last FULL capture proved a recognized agent may take the
  cheap tier. Panes with no anchor always take the full capture, so start
  discovery keeps today's exact behaviour.
- The cheap tick compares a per-pane fingerprint (root shell pid+start, tpgid,
  every descendant's pid+start+pgid+job-control state). Any change, a changed
  node-pty foreground name, an unreadable capture, or an incarnation mismatch
  escalates to the full capture. A recognized agent's exit is always a pid
  vanishing, which the fingerprint always sees.
- A cheap answer OMITS evidence rather than fabricating a tty-less fence.
  Remote/restore consumers never send `steadyState`, so they keep the full
  capture unchanged.
- `steadyState` is a new optional request field; an old daemon ignores it and
  answers with the full capture.

Measured (8 idle panes, 60s, idle cadence, forks counted by column set):
30 full -> 1 full + 29 cheap.

* fix(terminal): route the cheap ps capture through runProcess

The cheap-tier reader imported node:child_process directly, which the
child-process import-boundary and windowsHide ratchet tests reject (CI shards
1/8 and 3/8). Use Orca's single spawn entry point instead; it pins windowsHide
and encodes argv. Map its result onto the capture-error vocabulary:
outputTruncated -> capture_truncated, timedOut -> capture_timeout, non-zero
exit -> ps_exit_<code>. Tests mock at the runProcess seam.

* fix(perf): refuse a pane fingerprint when any descendant start marker is missing

`buildPaneProcessFingerprint` rejected only a missing root start marker; a missing descendant
marker was stamped as `?`. Two captures that both failed to read the same descendant therefore
compared equal, which removes the pid-reuse protection the fingerprint exists to provide: a
recycled pid could make a vanished agent look unchanged, and the cheap tier would keep serving
its name instead of escalating.

Reachable on Linux, where `readLinuxProcStartTime` legitimately returns null when a process
exits between the `ps` capture and the `/proc/<pid>/stat` read.

Every subtree member now needs a start marker or the fingerprint is refused, which sends the
caller to the full capture — the same conservative default every other uncertain path takes.

Reported by CodeRabbit on #18780. The two new tests fail against the previous code with
`expected '4242@2400#4300:|4300@?:4300:+' to be null`.
2026-09-05 00:57:02 -07:00
Brennan BensonandMerge Sim d5803bdbc4 feat(ssh): host-stamped remote foreground identity (#18078)
* docs: add SSH agent identity implementation plan

* feat(ssh): host-stamped remote foreground identity

* fix(runtime): preserve unfenced inspect call shape

* perf(ssh): traverse foreground descendants linearly

* fix(ssh): bound retired PTY evidence records

* test(ssh): cover retired incarnation retention

* fix(ssh): make remote process inspection total

* Split SSH identity build hot spots

* Fix process table snapshot module split

* test(ssh): update process inspection expectations

* docs: drop the SSH identity plan from the PR

The design doc does not belong in the product repo; it stays out of the
shipped tree while the implementation carries its own comments.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-02 23:32:41 -07:00
Neil 31007c0d86 fix(ssh): reclaim relay PTYs the client has provably lost, on host attestation only (#17831)
* fix(ssh): reclaim relay PTYs the host attests this client orphaned (#9819)

Orca could lose track of terminals running on an SSH relay until the
50-slot cap refused to open any more. This reclaims them, and the whole
design is built around the fact that getting it wrong destroys a user's
running process on their remote machine: the failure mode is leak, never
kill.

A stop requires all nine of:

1. the relay published an `ownerClientInstanceId` read from the live
   authenticated consumer grant of the connection that requested the
   spawn — never from a spawn parameter, since an echoed claim is no
   evidence; absent means skip
2. that id equals this client's persisted consumer identity
3. this connection holds the negotiated `session-owner` grant
4. `paneBound === true`, host-published
5. no `agentSessionOwners` — the host still advertises it as adoptable
6. `hostAgeMs >= 30s`, measured on the host's clock
7. this client has no route: not reattached, no lease outside
   terminated/expired, no pending kill, and no `expired` lease either —
   an expired lease is the record of a process deliberately left
   running, never a licence to kill it
8. every stop is fenced on the incarnation the same listing published,
   and on the owner identity, both re-checked by the host
9. a pass wanting to stop more than 8 refuses entirely

Absence from a client-side set is `unverifiable` by construction
(docs/reference/ssh-execution-boundary.md): a second machine attaches to
the same relay and displaces the session owner, and its live agents are
missing from this client's store for exactly the reason a genuine orphan
is. So the host has to attest ownership, and the host has to attest that
nothing is running.

That second attestation is measured over the pane's whole tty, not its
foreground process group. `tpgid == pgid` is foreground-only: on a real
`bash -i` on a real pty, a shell holding `sleep 300 &` and a shell
holding a Ctrl-Z'd job both read `pgid == tpgid`, `Ss+` — byte-identical
to an idle prompt, with only the job's own row differing. A
foreground-only gate therefore attests `pnpm build &` and a suspended
editor as idle, and the stop that follows SIGKILLs every process group
on the tty. `shellOwnsEveryTtyProcessGroup` is measured over that same
set of groups, so the evidence and the kill describe the same thing. No
new probe: `tpgid` already identifies the terminal, because a process
group belongs to one session and a session to at most one controlling
terminal.

The freshness field is real rather than decorative. `capturedAgeMs` is
stamped from when the capture was taken, deliberately as an upper bound
since the process table is TTL-shared, and the sweep refuses an
observation older than its own pass budget, counting its own elapsed
time since the listing arrived. Stale evidence degrades to "do not
sweep", never to "sweep". The display consumer of the same measurement
keeps no age budget, as a stated decision: a stale pane title costs a
redraw and self-corrects.

`pty.shutdown` is authorized on the host that owns the process.
`pty.spawn` and `pty.attach` both take a request context and check it;
the one irreversible call took none, so the rule above lived entirely on
the client that decided to make the call. It gains an optional
`expectedOwnerClientInstanceId` and refuses unless the connection still
authenticates as that identity AND this host recorded it at spawn.

Finally, a reattach refusal now says whether it observed the process.
Three refusals carry the same `SSH_SESSION_EXPIRED` text and only one is
absence; `restoreRequired` means the PTY is live and only its source
stream is not. Testing that text with `.includes()` expired the lease
and deleted ownership for a running process, erasing this client's only
record of it — and a PTY with no record is one the sweep may stop.

Wire compatibility: four new optional fields and one new optional param
on existing methods, no new method and no new stream opcode (Rule 1, and
Rule 2 does not apply). Rule 1's caveat is discharged explicitly — no
reader requires any of them, each absence is a named skip reason, and an
ordinary pane teardown must omit the owner fence because a revived PTY
carries no attested owner at all. New client plus old relay stops zero
PTYs; old client plus new relay never reads the fields. Windows relay
hosts publish no evidence and therefore never sweep.

Verified by joining the real publisher to the real client reader over
`ps` captured verbatim from a Linux container, and by driving a real
group-for-group SIGKILL against a real pty: backgrounded and suspended
jobs survive by pid, and an idle shell is still reclaimed, so the
narrowed predicate is not a silent no-op.

Squashed deliberately. The sweep is unsafe at every intermediate commit
of its own history — before the foreground gate it reaps a hand-launched
`claude`, and with a foreground-only gate it reaps a backgrounded build
— so this ships as one commit with no bisectable state that kills live
work.

Refs #9819. Folds in #17939.

* fix(i18n): restore the activity-options key the rebase dropped

* fix(i18n): union en.json with main so the rebase cannot drop keys
2026-09-02 15:14:14 -07:00
Neil 0886db2b90 refactor(process-table): extract the correlation indexes into their own module (#18246)
`src/shared/process-table-snapshot.ts` is 308 code lines against the 300
cap for `**/*.ts`, so `static analysis` is red on `main` and every open PR
inherits it.

Neither PR that grew the file crossed the cap alone. #18151 took it to 427
raw lines; #18166 added ~35 more. #18166's branch predated #18151, so the
head CI linted was 428 raw lines and passed, while the squash onto main is
463 -> 308 code lines. The gate lints the PR head, not the merge result, so
nothing linted the sum until it was on main.

Pure move, no behaviour change: the generic index machinery
(ProcessIdentityRow, ProcessTableIndexOf, buildProcessTableIndex,
collectDescendantsFromIndex, lookupProcessTableIndex, getProcessTableIndex
and its WeakMap) moves to process-table-index.ts. `ProcessTableIndex` and
`scoreForegroundCandidateRow` stay behind because they need
`ProcessTableRow`, which keeps the new module free of any import back and
so introduces no cycle.
2026-09-02 14:01:24 -07:00
Neil 6b1cbe54a1 fix(process-table): fail a short ps capture loudly, and stop a resume spending 49 of them (#18166)
The POSIX process-table capture ran `execFile('ps', ...)` with no `maxBuffer`,
inheriting Node's 1MB default. Measured at 1,460 processes the capture is 326KB
with a 5,116-char longest row — ~3x headroom, which a busy host clears.

Two separate defects follow, fixed here:

1. `parseProcessTableRows` drops unparseable lines, so any short capture reads
   as a COMPLETE table whose missing processes simply are not running. Verified:
   a capture cut at 4KB parses to 59 of 1,463 rows, and an empty capture parses
   to `[]`, both with no error — and `resolveAgentForegroundProcessWithAvailability`
   then answers `available: true`. That is the `unverifiable` -> `exited` collapse
   the execution boundary forbids. The capture now rejects with
   `ProcessTableCaptureError` on a ceiling-length or row-less capture, so both the
   lenient and strict views fail loudly and callers report unavailable.

2. `maxBuffer` is now an explicit 32MB, matching the sibling reader in
   `pty-descendant-termination.ts` and its stated reasoning. Without it a 4,000-
   process host fails EVERY capture, degrading the whole subsystem permanently.

Separately, `readStructuredTuiProcessIdentity` polled a fresh whole-machine `ps`
every 50ms for up to 5s. Each capture costs ~0.065 CPU-s, and the 5s ceiling is
only reached when the child never appears — where the tight interval buys
nothing. The interval now holds at 50ms for the first second, then doubles to a
500ms cap. Identification latency is unchanged for any child appearing inside
that window, and the 5s ceiling is unchanged.
2026-09-02 13:22:25 -07:00
Neil 53f105827b perf(windows): stop asking the process table for memory, and share one projection per snapshot (#18151)
Two costs on the Windows process-table hot path, plus the EDR doc that
described neither of them accurately.

1. The snapshot set `ProcessDataFlag.Memory` and surfaced `memoryBytes`,
   which nothing read. The addon serves that flag with a second
   `OpenProcess(PROCESS_QUERY_INFORMATION | PROCESS_VM_READ)` and a
   `GetProcessMemoryInfo` per process (process.cc:47-63), so the flag was
   one wasted handle per process per snapshot.

2. The shared TTL cache gave every pane the same native rows array, but
   each pane still ran `native.map(toProcessRow)` over the whole table,
   rebuilt a `childrenByPpid` Map from scratch, and did two linear scans.
   The `.map()` also handed `getProcessTableIndex` a new array each call,
   defeating the POSIX memo by construction. Both now cache per snapshot
   identity, and the POSIX resolver drops its duplicate descendant walk.

`getProcessTableIndex` / `buildProcessTableIndex` are generic over the row
shape so the Windows rows reuse the existing pass instead of a parallel one.

No behavior change: same rows in, same rows out, same descendant ordering
and same has-children answers.
2026-09-02 13:08:01 -07:00
Neil 406bd0e378 perf(relay): cache process-table descendant indexes (#17646)
* perf(relay): cache process-table descendant indexes

* fix(relay): keep the process-table index first-wins and narrow

Two defects in the memoized index this PR introduced.

- Restore the first-wins duplicate-pid tie-break the relay had as
  `rows.find()`. A process whose argv contains a newline makes `ps` print a
  continuation line that the lenient parser can accept as a spurious row
  duplicating a real pid; that row always FOLLOWS the real one, so last-wins let
  it capture the pane's foreground. The rule now lives in
  `buildProcessTableIndex`, so the batched evidence resolver's `byPid.get(rootPid)`
  root lookup gets the same semantics the subsystem had before indexing.
- Build only the two indexes a resolver reads. `byPgid`/`byTpgid` have no readers
  repo-wide, and delegating to a four-map build made a one-pane relay pay more
  per 500ms capture than the single `childrenByParent` map it replaced --
  a regression in the majority topology, in a PR whose point is relay CPU.

Matches the same deletion in #17763 line for line so whichever merges second
resolves trivially.
2026-08-31 18:38:03 -07:00
Neil 704167197a perf(relay): serve one ps capture per window and pin the batched inventory path (#17763)
Follow-up defect fixes for the batched PTY-inventory evidence path (#17525),
now on main.

- One memoized `ps` capture serves both the lenient and strict views. The two
  readers ran byte-identical argv behind separate caches, so a relay serving
  both forked `ps` twice per 500ms window — the doubling issue #6288 removed.
- Drop the `byPgid`/`byTpgid` indexes no resolver reads, plus the zero-caller
  `parseProcessTableRowsStrict` and `getFreshStrictProcessTableSnapshot`; the
  batch resolver now reuses the shared index lookup and candidate score instead
  of private copies.
- Restore `getForegroundProcessName`'s ladder contract: the extracted table scan
  answers null again, so an unconfirmed wrapper fallback publishes the
  recognized (normalized) name rather than node-pty's raw one.
- Pin the SHIPPED `pty.listProcesses` path: one capture and one linear row pass
  for N panes, and node-pty's own name (never "shell") when the capture cannot
  disambiguate a `node`/`python` wrapper.
- Pin the hidden-pane cadence gate in the production option shape, and move the
  strict-parser coverage next to the parser it tests.
2026-08-31 17:45:55 -07:00
Brennan BensonandMerge Sim 9477b5fcbb feat(ssh): batch process evidence in PTY inventory (#17525)
* feat(ssh): batch process evidence in PTY inventory

* fix(ssh): accept Linux kernel process rows and make no-evidence polling push-driven

* fix(ssh): preserve process evidence polling semantics

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 12:44:41 -07:00
b902f19b0d fix(terminal): send CSI-u Shift+Enter to Droid on Windows (#7620) (#7668)
* fix(terminal): send CSI-u Shift+Enter to kitty TUIs (droid) on Windows (#7620)

On Windows, Shift+Enter was always sent as the Alt+Enter byte ESC+CR (added in
#2418 for Codex, which reads win32-input-mode and ignores CSI-u). droid speaks
the kitty keyboard protocol, parses CSI-u directly, and treats ESC+CR as a plain
Enter — so Shift+Enter SUBMITTED the message instead of inserting a newline.
droid works in other terminals (Windows Terminal, Warp) because those honor
win32-input-mode / kitty; Orca (xterm.js) withholds kitty from local Windows
ConPTY panes and emits neither.

Make the Windows Shift+Enter byte pane-aware: latch whether a pane's program
advertised the kitty keyboard protocol (query CSI ? u, push CSI > .. u, or set
CSI = .. u) and send CSI-u (\x1b[13;2u) to those panes, keeping the
Codex-compatible ESC+CR for win32-input-mode-only TUIs. Non-Windows is unchanged
(always CSI-u).

Verified end-to-end against the real droid and Codex CLIs through the actual
production functions: droid now newlines, Codex still newlines.

* fix(terminal): route Windows Shift+Enter safely for Droid

---------

Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-07-10 21:00:31 -07:00
NeilandOrca 88640ef4c6 docs(runtime): fix stale ps-snapshot factory comment after #7518 (#7529)
Co-authored-by: Orca <help@stably.ai>
2026-07-06 00:44:21 -07:00
NeilandOrca 27ba95cf31 perf(runtime): cache parsed ps rows on POSIX so panes share one parse (#7518)
getProcessTableSnapshot deduped the ps fork (#6288/#6667) but cached only the
raw stdout string on POSIX, so every concurrent agent pane re-ran parsePsRows
over the identical output within each 500ms TTL window — O(M*P) redundant
tokenization + row allocation. The Windows reader already caches parsed rows;
this makes the POSIX default reader do the same by parsing inside the deduped
scan and returning ProcessTableRow[]. Collapses the duplicate parsePsRows in
the main and relay foreground resolvers into one shared parseProcessTableRows.

Co-authored-by: Orca <help@stably.ai>
2026-07-06 00:03:54 -07:00
NeilandOrca 91323cf925 perf(windows): dedupe per-pane process-table scans in agent inspection (#7384)
* perf(windows): dedupe per-pane process-table scans in agent inspection

Windows agent foreground-process inspection forks a whole-process-table
PowerShell/CIM scan per pane on the same 750ms/2000ms cadence the POSIX path
uses. The POSIX side routes through getProcessTableSnapshot (500ms TTL + single
in-flight, #6288/#6667), collapsing N concurrent panes to ~2 scans/sec. The
Windows path (queryWindowsProcessDescendants) had no such dedup: K concurrent
agent panes forked K powershell.exe cold-starts, each enumerating the ENTIRE
process table then filtering per-pid in JS — ~10-40x heavier than `ps` (a
powershell cold start is ~150-400ms CPU + tens of MB RSS). The degraded/local
PTY provider path calls it with no per-pane throttle at all. This is the
Windows analogue of the idle-CPU churn #6288 fixed for POSIX.

Generalize the existing createProcessTableSnapshotReader factory to be generic
over its scan result (default T = string, so the POSIX path and its test are
byte-identical) and add a Windows singleton reader that caches parsed
WindowsProcessRow[]. queryWindowsProcessDescendants now reads the shared
snapshot and runs its own descendant walk; runWindowsProcessRows throws on total
enumeration failure so the miss is not cached and the prior null-fallback
contract (callers fall through to node-pty's name) is preserved.

Windows scan-volume regression test (mirrors the POSIX #6288 guard) drives
PANE_COUNT concurrent panes over the cadence window and asserts powershell.exe
spawns are bounded by ticks, not pane count, while every pane still resolves its
descendant. Reverting the dedup fails both cases. POSIX snapshot + volume tests
unchanged and green; node/web/cli typecheck clean.

Co-authored-by: Orca <help@stably.ai>

* test: reset windows process-rows snapshot between agent-foreground cases

The new module-level Windows rows reader caches for 500ms with real
Date.now(), so one case's mocked process table was served to the next
case's assertions (7 CI failures in agent-foreground-process.test.ts).
Mirror the suite's existing POSIX resetProcessTableSnapshotForTests()
with the Windows reset in beforeEach.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-07-04 18:10:45 -07:00
NeilandOrca 46646d7ff1 chore(lint): upgrade oxlint to 1.71 + enable 7 new rules (autofixed backlog) (#6841)
* chore(lint): upgrade oxlint to 1.71 and enable 7 new rules

Upgrade oxlint 1.67.0 -> 1.71.0 (1.72 was blocked by the repo's 3-day
minimum-release-age supply-chain guard; nothing here needs it). The
bump is a no-op on the existing config.

Enable 3 error rules (backlog autofixed to zero in this commit) and
4 warn rules (surface signal without gating CI):

error (autofixed, behavior-preserving):
- unicorn/prefer-node-protocol        (~1531 sites: bare builtin -> node:)
- typescript/no-import-type-side-effects (~36: all-inline-type -> import type)
- unicorn/no-array-reverse            (19: copy-then-reverse -> toReversed)

warn (real signal, current fires are test-only/correct):
- unicorn/no-array-fill-with-reference-type  (aliasing footgun guard)
- typescript/no-unsafe-function-type         (bans bare Function type)
- unicorn/prefer-array-flat-map              (map().flat() -> flatMap())
- unicorn/prefer-regexp-test                 (.match() in bool ctx -> .test())

mobile/.oxlintrc.json extends root, so it inherits all 7; the autofix
ran from root and covered mobile/ too.

Verification (all green): oxlint 0 errors (root+mobile+aux configs),
oxfmt clean, typecheck (node+cli+web), vitest 22795 passed / 0 failed,
builds (electron-vite + web + cli) succeed. node: rewrites confirmed to
skip embedded SSH/CLI string payloads (AST-only); all toReversed sites
verified to operate on fresh copies or write-once locals.

* chore(lint): bump mobile oxlint to 1.71 so inherited rules parse

mobile/ is a standalone pnpm project pinning its own oxlint@1.67, which
lacks unicorn/no-array-fill-with-reference-type (needs >=1.70). Since
mobile/.oxlintrc.json extends the root config, mobile CI's 'cd mobile &&
oxlint' failed to parse the new rule. Bump mobile to match root (1.71).

Verified in mobile/: oxlint 0 errors, oxfmt --check clean, tsc --noEmit
pass, vitest 978 passed / 0 failed.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-06-29 22:38:29 -07:00
NeilandOrca 06392a4523 perf: dedupe relay process-table scans behind a short-TTL cache (#6288) (#6667)
* perf: dedupe relay process-table scans behind a short-TTL cache (#6288)

Agent foreground-process inspection runs `ps -axo pid=,ppid=,stat=,command=`
(a full system process-table scan) on a 750ms/2000ms per-pane cadence. On a
shared SSH relay every tracked agent terminal drives it, so concurrent panes
each forked their own `ps` — sustaining up to the per-second inspection cap of
full-table scans for as long as agents are open, pinning idle relay CPU and
amplified by AV process scanning. This is the CPU half of #6288 (PR #6564
covers the memory-leak half).

Memoize the scan behind a single in-flight promise + 500ms TTL shared by the
relay and local main-process call sites. 500ms sits below the active poll's
minimum inter-poll gap (~675ms after jitter), so a single pane never reuses a
snapshot older than it would have scanned itself — same data and freshness,
just deduplicated within the cadence window (worst case ~8 scans/sec -> ~2).

Failures are never cached (in-flight cleared on settle) so a transient `ps`
error retries and the existing best-effort fall-through is preserved. Windows
branches are untouched; no git-provider implications.

Co-authored-by: Orca <help@stably.ai>

* test: regression guard for #6288 ps-scan volume (repro + measurement)

Drives the real local foreground-inspection call site under the documented
750ms agent-completion cadence across 6 concurrently-inspecting agent panes
over a 30s window, counting actual `ps -axo pid=,ppid=,stat=,command=`
full-table scans.

Reproduces the waste on `main` (240 inspections -> 240 scans, 1.0/inspection;
the test fails there) and proves the fix (240 inspections -> 40 scans,
0.167/inspection — bounded by poll ticks, not pane count) while every pane
still resolves its foreground agent. Guards against regressing the cache back
to a per-call scan.

Co-authored-by: Orca <help@stably.ai>

* docs: clarify 500ms TTL covers cadence floor + tolerated event-driven staleness (#6288)

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-06-28 17:28:37 -07:00