Commit Graph
10415 Commits
Author SHA1 Message Date
Neil 274ea37183 test: validate original captured OpenCode redraw replay 2026-09-06 20:20:06 -07:00
Neil 2dd3958339 test: distinguish external retention from owned worker recovery (#19190) 2026-09-06 20:16:38 -07:00
Jinwoo Hong 3160b54c69 feat: real background push notifications for the mobile app (#8129) (#18554)
* feat(cloud): add the mobile push gateway and its contract package (#8129)

A small open-source service that holds the APNs key and FCM credentials and
sends background push to paired phones on the desktop's behalf. Hosts
authenticate with a box challenge and HMAC proof on their pairing key, the
same shape the relay uses, so signed-in and accountless desktops share one
path. Tokens are stored; alert text is held only for the coalescing window.

The contract doc in docs/reference is the source of truth for every wire
shape. The interop test runs the real desktop answerer against a real
gateway-issued challenge so transcript drift fails in CI.

* feat(push): register phones and send background push from the desktop (#8129)

Adds the notifications.remote-push.v1 capability, the registerPush and
unregisterPush RPCs on the mobile allowlist, a gateway client with a cached
session and 401 re-auth, a durable unregister outbox, and a dispatcher that
offers every mobile notification to the gateway after the socket fan-out.
The dispatcher is fire-and-forget with one retry and drops registrations the
gateway reports dead.

Puts agentState on the mobile frame and fixes the #4375 wording so a working
agent is never announced as finished. The relay host-proof code moves onto a
shared envelope module with no behaviour change.

* feat(mobile): background push registration, receive, and settings (#8129)

Fetches the native APNs or FCM token, registers it with every paired host
that advertises the capability, and re-registers on token change. Foreground
pushes are suppressed inside handleNotification against the same seen set
the socket path uses, so nothing shows twice. Taps route by host fingerprint.
One Background notifications switch, off by default, with the disclaimer and
needs-input / finished sub-switches; hidden until a paired desktop is new
enough. Adds google-services.json and the expo-notifications plugin.

* chore(cloud): Terraform and deploy workflow for the push gateway (#8129)

Declares the Cloud Run service, runtime account, secrets, and orca_push
database behind push_gateway_enabled, true only in production. The deploy
workflow is gated like the relay's, deploys with no traffic, probes /ready
and a validate-only FCM send, then shifts traffic. It runs as the shared
production deploy account because the Cloud SQL rollout lease grant is
foundation-owned; its extra authority is three bindings on the push service.
docs/push-gateway.md carries the import commands for the resources created
by hand and the APNs key rotation procedure.

* docs: describe background notifications on the phone (#8129)

* docs: check in the mobile push contract (#8129)

Seven committed files cite it as the source of truth for every wire shape;
docs/reference is allowlisted per file, so add the entry.

* test(push): replay one checked-in host-proof vector on both sides (#8129)

Cloud Verify installs only the cloud workspace, so the gateway suite cannot
import the desktop answerer. Replace the cross-workspace import with a fixed
challenge vector generated from the contract package; the gateway fixture and
the desktop answerer each replay it and must produce the same HMAC. A
transcript drift on either side now fails in that side's own suite.

* fix(cloud): open the push gateway with invoker_iam_disabled, not an allUsers binding (#8129)

The production domain-restricted-sharing policy rejects an allUsers
run.invoker member, which the runbook anticipated. Opt the service out of
invoker IAM the way the relay director already does; the host proof is the
authentication either way.

* docs(cloud): the push.onorca.dev record exists and is hand-managed (#8129)

* fix(push): close review findings in the gateway (#8129)

- Quota reservation takes a per-host advisory lock; READ COMMITTED admitted
  a whole burst past the cap (80/80 without, 60/80 with, against Postgres 16).
- Challenge issuance no longer writes push_hosts; the row lands on proof
  verification. Stale hosts prune after 30 days. Per-IP token bucket on the
  two unauthenticated routes.
- Streaming body limit via hono bodyLimit; a chunked body bypassed the
  Content-Length check.
- registrationIds deduped in the schema; per-host device cap of 64; list
  bounded to its schema.
- Gateway-side challenge TTL is the specified 10 s, not 40 s.
- APNs stream settles on close as well as end/error.

* fix(push): close review findings in the desktop client (#8129)

- A gateway registration the registry cannot persist is enqueued for delete
  instead of leaking a live token.
- Unregister outbox re-reads pending per pass, honours enqueues during a
  drain, and retries with backoff instead of waiting for the next launch.
- Dispatcher batches registrations by 20 rather than starving the rest.
- 401 compare-and-clear; a 401 after re-auth is unreachable; refused
  handshakes and 429s are cached briefly instead of re-handshaking per event.
- Service is stopped on quit.

* fix(mobile): close review findings in push registration and receive (#8129)

- Consent generation guards a register that finishes after the switch went
  off; the host is re-queued for unregister instead of recorded live.
- Foreground pushes seed the watermark before adopting the epoch, so a push
  on a never-connected session cannot wipe a valid watermark.
- aps-environment follows the build via app.config.js; the iOS release
  workflow sets it to production. A bare plugin entry wrote development.
- Pushes the OS showed while closed are marked seen before catch-up replay.
- Token null result is not cached; failed capability probes are retried and
  never block an unregister; coalesced summaries are shown but not marked.
- Unresolvable fingerprint routes nowhere and is suppressed in foreground.
- Android channel ensured at boot; capability hook diffs clients by identity.

* fix(cloud): harden the push deploy workflow and size the gateway to the budget (#8129)

- Roll traffic back on a failed post-shift check; delete a candidate that
  never took traffic; retry the origin probe and the FCM probe.
- Assert Terraform-owned scaling instead of mutating it from the workflow.
- Build before taking the Cloud SQL rollout lease.
- Declare the database pool in Terraform (2 per instance, max 2 instances)
  and add the gateway to the connection budget; the previous default put the
  shared instance 65 connections over its ceiling.
- State plainly that the shared deploy identity's relay authority is inherited.

* fix(push): read the runtime from shared state at push startup (#8129)

Threading the runtime through launchDesktopMode put the launch module one
line over the 300-line lint budget after the rebase.

* fix(push): key the unauthenticated rate limit on the hop Cloud Run wrote (#8129)

Cloud Run appends the connecting peer to x-forwarded-for; the limiter read
the left-most value, which the caller controls, so a forged first hop earned
a fresh bucket per request.

* fix(push): close the final security review findings in the gateway and infra (#8129)

- app.onError logs only the error name and answers a bare 500; hono's default
  handler printed the whole error, and a pg error carries the row in detail
- a second per-IP bucket (240/min) runs ahead of the bearer lookup on every
  authenticated route, so forged bearers cannot spend the two-connection pool
- one live session per host: minting deletes the host's earlier row
- device-less hosts are pruned after 1 h, not 30 d; any keypair mints one free
- notificationId is printable ASCII, since it becomes the APNs collapse header
- the impersonated FCM probe token is masked in the workflow log
- prevent_destroy on the Apple secrets and the orca_push database

* fix(push): close the final security review findings in the desktop client (#8129)

- fetch never follows a redirect: a 307 would replay the host proof and the
  phone's token to whatever origin the redirect named
- registerPush params are strict and the paired identity is spread last
- a per-device bucket (10/min) bounds a phone looping registerPush, which
  costs a gateway write and a synchronous registry write each time

* fix(mobile): close the final security review findings in push receive (#8129)

- a push with no epoch can no longer claim a seq-derived dedup key, in the
  foreground or from the tray; a forged seq:N could otherwise swallow the
  real bell at that seq
- a provider-delivered push with no host catalog, or no fingerprint at all,
  stays unrouted instead of falling back to the hostId its raw data carries

* docs(push): record the ip buckets, session and host retention, and the token-ownership limit (#8129)

* fix(push): apply the schema on an untimed pool and retry statement-timeout aborts (#8129)

Ports the relay's #18722 pattern to the gateway: DDL runs on a one-connection
pool with statement_timeout 0 that is closed before the serving pool opens, and
SQLSTATE 57014 joins the bounded transaction retry path.

* fix: harden mobile push delivery and deployment recovery

* feat: align mobile notification preferences with desktop delivery

* fix: accept variable-length APNs device tokens

* fix: deduplicate native APNs and background socket notifications
2026-09-06 23:16:29 -04:00
Neil 79eb66608a test: retain paired browser value from successful poll (#19189) 2026-09-06 20:09:39 -07:00
Neil 6c8ce54ad8 test: publish restored snapshot before draining its held FIFO (#19186) 2026-09-06 20:00:08 -07:00
Neil 1848855515 test: enable software WebGL for Linux CI headful specs (#19001)
* test: enable CI WebGL and route GPU-dependent regressions

* test: retain headful atlas cases in terminal rendering goldens

* test: reuse golden command in project coverage assertions
2026-09-06 19:27:12 -07:00
TimothyVangandNeil 4934920f06 fix(rate-limits): stop reporting Grok usage as 0% when the API omits the percent (#17936)
mapWeeklyCredits treated an absent creditUsagePercent as a confirmed protobuf
zero whenever the weekly period matched billing bounds, so unified-billing
accounts whose credits view never reports the percent showed a confident 0%
and short-circuited the monthly fallback (#15740). Those payloads emit
onDemandUsed/prepaidBalance zeros, which disproves the "encoder drops zeros"
premise.

Resolution order is now: reported percent → monthly used/monthlyLimit pair as
a monthly window → synthetic 0 only when the payload emits no usage scalars at
all and the weekly period is confirmed → unavailable with an explicit reason
the Accounts pane surfaces.

Rebased onto current main from nwparker/grok-usage-percent-fallback (#15878).

Fixes #15740

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-06 19:25:57 -07:00
weekbinandNeil e48d83a5e1 Fix MiniMax China usage routing and credential handling (#14929)
* feat(minimax): endpoint selector, API key auth, weekly usage window (#14264)

The MiniMax (MiniMax) Coding Plan usage fetch was hardcoded to the
overseas platform (platform.minimax.io) and a single 5h session
window, so users on the CN endpoint (www.minimaxi.com) got nothing.

Three changes:

- Add `minimaxEndpoint` (`overseas`|`cn`) and
  `minimaxApiKeyConfigured` settings fields with sensible defaults
  that preserve current behavior. The CN endpoint also accepts an
  API key (safeStorage-encrypted via a new
  `minimax-api-key-store.ts` + IPC pair) for users without a
  browser session cookie. Status-bar visibility now OR's both
  credential flags.
- Cookie-jar origin now tracks the active endpoint. Previously
  cookies were stored under the overseas origin and silently
  dropped when the user picked CN — fixed by threading
  `endpointMode` through the request context, the manual cookie
  header path, and the cookie-jar clear.
- Parse the weekly window in addition to the 5h session and
  surface both as per-window chips (`5h [bar] 10%   wk [bar] 20%`).
  The status bar's compact section prefers the session window; the
  popover keeps the existing `Session` / `Weekly` labels. The
  MiniMax fetcher is split into three files (data / parse / main)
  to stay under the 300-line cap.

i18n is scoped to the Settings-page text (en + zh only); the 5H/7D
duration shorthands stay English across locales by project convention.

Tests: 9 new/updated files; cookies + API key exercised end-to-end
via the rate-limit service with the upstream-refactored test files
(`service-minimax-usage.test.ts`,
`web-preload-api-settings.test.ts`,
`web-preload-api-agent-providers.test.ts`,
`service-test-harness.ts`, and the runtime-home / reset-credit
fixtures).

Refs #14264

* Keep merge formatting scoped to MiniMax

* Keep MiniMax credential status in rate-limit test fixtures

* Use the China console origin for MiniMax request referer

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-06 19:20:50 -07:00
Neil e9af947035 test: confirm running-command prompts when closing tabs (#18965)
* test: wait for rendered tabs and handle busy close confirmation

* test: wait for create-menu item click actionability

* test: settle initial terminal focus before create-menu actions

* test: capture menu focus events for Linux CI diagnosis

* test: remove menu diagnostics after identifying deferred layout focus

* test: check Markdown menu dismissal after editor readiness
2026-09-06 19:15:57 -07:00
Neil b3acef218a test: verify imported projects through the virtualized sidebar (#19003) 2026-09-06 19:14:42 -07:00
Neil 4be1c01c42 test: await rendered remote agent placement before checking mirrors (#18983) 2026-09-06 19:14:39 -07:00
Neil 357a4d4920 test(e2e): scope paired preview link checks to confirmation (#18924) 2026-09-06 19:14:37 -07:00
Neil 1e301ab1df test: cover native Wayland Hangul in isolated CI (#19174)
* test: exercise native Wayland Hangul in isolated CI session

* test: wait for nested compositor socket before selecting IBus

* test: align Wayland IBus discovery with GNOME environment filtering

* test: assert Wayland launch and register native Hangul evidence
2026-09-06 19:02:54 -07:00
Neil af5918a254 test: use host-qualified paired palette row identities (#19175) 2026-09-06 18:55:26 -07:00
Neil e7563c63f1 fix: fence browser recovery to attach inventory placements (#18910)
* fix: fence browser recovery to attach inventory placements

* refactor: name the attach-inventory fence and make its test deterministic

Extract the placement check into isPlacedAsObservedAtAttach so the recovery
filter stays a flat list of named predicates, and document that omitting
pagePlacementsAtAttach recovers against unfenced live state.

Replace the 30-microtask drain in the post-attach regression with the handler's
own completion: attach only settles after recovery returns, so awaiting the
dispatch orders the assertions instead of guessing at a microtask count.

Verified by forcing the fence open: both regressions fail (the post-attach one
in 60ms on a retired placement) and the other 27 still pass.

* test: settle the attach handler even when the regression fails early

The barrier ran inline, so a waitFor timeout or the placement guard left the
attach handler parked on a promise nothing awaited. Hoist it into settleAttach
and call it from a finally as well; cleanup is guarded and the dispatch promise
is already settled, so the second call is a no-op.
2026-09-06 18:50:15 -07:00
Neil fc37958b45 fix: release floating terminal WebGL contexts while closed (#19000)
* fix: release floating terminal WebGL contexts while closed

* test: pin retention polarity through a real PaneManager

Replace the prototype-surgery fake with a constructed PaneManager so the
suspend path exercises real constructor state, and add the retain-branch
case so an inverted default cannot pass silently.

De-shadow `window` in the system-resume e2e main-process callback.
2026-09-06 18:48:40 -07:00
Andrey 14e0d40e06 fix: recognize the kimi-code process as the kimi agent (#18634) 2026-09-06 18:43:39 -07:00
Neil 463cab2f71 fix(cmd-j): remove duplicate browser ownership inputs (#19172) 2026-09-06 18:40:43 -07:00
Neil 373514ef26 perf(worktrees): stop worktree teardown replacing arrays and maps it never touched (#19145)
* perf(worktrees): stop worktree teardown replacing arrays and maps it never touched

Removing a worktree fires three store writes through
removed-worktree-renderer-teardown.ts, and each handed back a fresh reference
even when it removed nothing:

- remove-worktree-store-cleanup filtered openFiles unconditionally. #19058 gave
  the ~50 record maps in this file identity preservation and missed the one plain
  array; the sibling purge path already had the guard this copies.
  openFiles is selected whole by the editor panel, file explorer and git-status
  polling.
- shutdownWorktreeBrowsers spread-then-deleted browserTabsByWorktree and
  activeBrowserTabIdByWorktree; both now go through omitRecordKeys.
- markShutdownPending rebuilt suppressedPtyExitIds and pendingPtyShutdownIds even
  with no guard ids at all, which is the normal case when the panes already
  exited. It now returns early, and skips the suppressed map when every id is
  already true.

Same contents, same keys removed; only the reference is reused when nothing
changed.

* fix(test): use AppState['openFiles'][number] instead of a nonexistent module

The test imported OpenFile from shared/editor-types, which does not exist. Vitest
passed because a type-only import is erased at runtime; CI typecheck caught it.
I had run tsc before adding this file and never re-ran it.

* refactor(terminals): reuse copyOnWriteRecord in markShutdownPending and pin its identity contract
2026-09-06 18:33:55 -07:00
Neil be10e5455e perf(store): keep recentlyRetiredAgentStatusPaneKeys identity on no-op retirement (#19142)
boundRecentlyRetiredAgentStatusPaneKeys always rebuilt the record, replacing
its reference even when nothing changed; a probe counted 1,099 such writes
across the store suite. Return the existing record when no key would be
evicted and the additions are already its tail in the same relative order.
Key-set equality is deliberately NOT enough: re-adding a key must move it to
the tail because that LRU order decides which key the cap evicts next.

Share the LRU bound with boundRecentlyClosedAgentStatusTabIds, which had the
same always-rebuild shape.
2026-09-06 18:33:44 -07:00
Neil 0cba706b01 fix(ports): coalesce advertised URL refresh bursts (#19150) 2026-09-06 18:33:19 -07:00
Neil deb0be1c52 fix: recognize working WSL1 without a WSL2 kernel (#19061)
* fix: recognize working WSL1 without a WSL2 kernel

* fix: recognize unsigned Windows missing-kernel status

* fix(wsl): fold the missing-kernel guest probe into wsl-availability

The separate wsl-missing-kernel-probe module failed three CI gates: it was
not in the web typecheck project (TS6307), it added a new direct wsl.exe
spawn outside wsl-runner, and its `catch { return false }` tripped the
probe-failure-semantics ratchet.

wsl-availability.ts already owns the answer and is already on the invocation
allowlist, so the probe lives there now. A guest probe that cannot spawn
keeps the real --status failure instead of minting a fresh negative, which
is what the ratchet exists to prevent -- and is the more correct semantics.
2026-09-06 18:33:12 -07:00
Neil e28b15928a fix: avoid credit deadlock during large SSH PTY recovery (#19026)
* fix: avoid credit deadlock during large SSH PTY recovery

* test: restore bounded SSH flood recovery coverage

* test(relay): pin the recovery fence to the accepted checkpoint

The oversized-tail cases asserted that the drain completes, but not that
recoveryEndSu lands on the checkpoint, so passing the pre-rotation snapshot
(which carries the old client's window and a stale creditedEndSu) fenced
below the checkpoint and still passed. Assert the fence value, narrow
boundedPtyRecoveryEnd to the three fields it reads, and cover the exact
one-window boundary that separates a live drain from an ordinary fence.
2026-09-06 18:29:31 -07:00
Neil f88cbb4fc9 fix: keep paired tab updates live after runtime terminal fallback (#19022)
* fix: preserve session publication during runtime terminal fallback

* refactor(runtime): align fallback epoch comment and test preamble

Match the file's `// Why:` comment convention on the inherited
publication epoch, and drop a redundant duplicate mocks import in the
lineage regression test while keeping the required side-effect order.

No behavior change.
2026-09-06 18:29:28 -07:00
Neil 0d973c1505 fix: honor remote terminal insertion in the calling client (#18995)
* fix: settle remote terminal insertion in the calling client

* refactor: share one anchor insertion path for local and remote terminals

Extract the created-tab-after-anchor reorder that the local terminal IPC
bridge already carried into insertUnifiedTabAfterAnchor, and settle the
remote placement through it instead of a second copy.

Also repairs two anchor-resolution gaps in the settlement:
- keep an exact unified tab id (legacy leaf-keyed anchors, browser and
  editor tabs) instead of collapsing every anchor to a terminal parent,
  which could mint a `web-terminal-<browser tab>` id that matches nothing
- fall back to the anchor's own group when the requested group was closed
  while the mirrored tab was still in flight
2026-09-06 18:29:25 -07:00
Neil b497f15b53 fix: preserve renderer browser publication during client-hosted page updates (#18961)
* fix: preserve renderer browser publication during client-hosted page updates

* refactor: drop the now-dead publicationEpoch selection argument

applyBrowserSessionTabSelection took a publicationEpoch and wrote it over
the epoch the spread snapshot already carried. Its only production caller
now passes snapshot.publicationEpoch, so the parameter is a no-op whose
only remaining power is to reintroduce the epoch rotation this PR fixes.

Remove it, and collapse the repeated prototype-cast boilerplate in the new
reconciliation test into one helper.

No behavior change.

* fix: keep reconcile from publishing a browser row twice

The retention filter partitioned existing rows by placement kind, so its
disjointness from the live build relied on a non-local invariant: that the
page registry only ever stores client placements and that server tabs are
empty while no offscreen backend exists. Drop ids the live build already
published instead, so a duplicate row is impossible by construction rather
than by coincidence.

* fix: stop the browser reconcile republishing on a pure reordering

headlessBrowserTabsUnchanged compares by array index, so rebuilding the live
list renderer-first read an interleaved snapshot as changed and republished
with a bumped version and rebuilt tab groups for no semantic change - the
same churn this branch exists to remove.

Key the live set by id and emit it in the order the snapshot already had.
Keying also makes uniqueness unconditional rather than resting on the page
registry only ever storing client placements.
2026-09-06 18:29:22 -07:00
Neil 225a47533d fix: preserve paired host sessions during startup residue cleanup (#18922)
* fix: preserve paired host sessions during startup residue cleanup

* refactor(persistence): tighten the paired-host retention pass

Dedupe the owner-key -> repo-id extraction the retention and seeding
passes both needed, and name the `runtime:*` check instead of repeating
the parse three times.

Reach the session walker directly by exporting
`addWorkspaceSessionWorktreeOwners` rather than fabricating a
`{ workspaceSession }` state slice to get at it.

Correct the docstrings: `runtime:*` also covers a serving host's own
partition, and the "authoritative removal" they promised has no product
caller on a paired client today, so say what the exemption actually
costs.

Add a survived-load assertion to the explicit-removal test, which
otherwise passed against the pre-fix sweep -- the partition was already
empty before the removal ran.

No behavior change beyond the docs and the test assertion.
2026-09-06 18:29:19 -07:00
Neil a272a1eeaf fix: preserve terminal command probes across control frames (#19006)
* fix: preserve terminal command probes across control frames

* refactor(terminal): make the command-probe output flag explicit

Hoist the duplicated Output/OutputSpan predicate in the binary frame
handler, and require carriesOutput on recordInbound so no future call
site can silently disarm the command-response probe by omitting it.

Rework the control-frame regression into a named table so the
fit-override and driver-changed cases send valid event payloads instead
of stubs that returned before dispatch.
2026-09-06 18:29:08 -07:00
Neil 8dad5958c8 fix: preserve overlay focus during terminal mounting and layout (#18982)
* fix: preserve overlays during terminal mounting and layout

* fix(terminal): stop a dismissed overlay from blocking pane focus

Overlay primitives animate out (data-[state=closed]:animate-out, up to
300ms on sheets), so a dismissed dialog stays mounted and painted well
past the point it should stop owning focus. The rAF-deferred focus in
activateTabAndFocusPane lands inside that window, so revealing an agent
from the dashboard drawer or a menu left the terminal unfocused.

Treat data-state="closed" as gone, matching the [data-state="open"]
convention already used by AgentDashboardDrawer and useWorkspaceBoardPanel.

Also revert unrelated comment churn on scheduleRevealRepaint and note the
new focus consumer in the hasVisibleOverlay doc comment.

* refactor(terminal): scope the dismissed-overlay rule to pane focus

Gate the data-state="closed" exclusion behind an ignoreDismissed option
that only focusPanePreservingOverlays passes, leaving Escape semantics
for the four existing hasVisibleOverlay callers unchanged.

The focus race this fixes is specific to deferred focus (activateTabAndFocusPane
defers by one rAF, landing inside the overlay's exit animation). Escape is
synchronous and does not need the rule: Radix's useEscapeKeydown is capture
phase, so every Escape caller runs while data-state is still "open".

Avoids any behavior change on the Settings Escape path, which unlike the
other three callers is bubble phase on document with no ordering guarantee.
2026-09-06 18:29:05 -07:00
Neil 4ba8ddce48 fix: prefer retained provider snapshots during hidden terminal recovery (#18972) 2026-09-06 18:29:03 -07:00
Neil 1ef75d79d7 fix: avoid starting browser helpers just to reset absent sessions (#18952)
* fix: avoid starting browser helpers just to reset absent sessions

* refactor(browser): tighten the session-reset skip guard and its tests

Drop the platform and absolute-path guards: ownsSocketDirectory is already
false on Windows and for inherited directories, and an Orca-derived directory
is always absolute. Fold the empty-name and traversal checks into agent-browser's
own session-name rule.

Stop lstat state leaking between lifecycle tests, and pin the probed socket path
so the skip test cannot pass on an unwired mock.
2026-09-06 18:29:00 -07:00
Neil deebe05ff0 fix: open editor rename after context menu releases focus (#18934)
* fix: open editor rename after context menu releases focus

* refactor(editor): tighten rename focus-handoff comments and test setup

Correct the rename-input focus comment that still credited the animation
frame with outrunning menu teardown, clarify why the rename now runs from
onCloseAutoFocus, and fold the repeated menu-close invocation in the tab
tests into one helper.
2026-09-06 18:28:57 -07:00
Neil 8d8b9dad78 fix: keep macOS shell ownership proof within recovery budget (#18932)
* fix: keep macOS shell ownership proof within recovery budget

* fix: parse the shell-proof column set with its own anchored parser

The narrower macOS capture (`pid ppid pgid tpgid stat command`) was fed to the
shared lenient parser, whose optional tty/start pair has no `tty=` column left
to absorb it. It then eats the head of any argv shaped `python 3 app.py`
(parsing command as `app.py`, tty as `/usr/bin/python`), and turns a
command-less row into a garbage pid/stat pair. Either can flip a shell
ownership verdict, which is what gates dead-TUI recovery.

Give the column set a named constant and a parser anchored to exactly those
six columns, beside its `CHEAP_PS_ARGS` sibling. A capture that yields no rows
now raises `empty_capture` rather than reading as a machine with no processes.

Update the `confirmShellForegroundProcess` fixtures from the 4-column legacy
shape to the 6 columns the darwin reader actually emits; that describe block
already forces `platform=darwin`, so the stale fixtures were failing.
2026-09-06 18:28:54 -07:00
Neil e8496f810a fix(cmd-j): pass browser tab ownership into palette search (#18925)
* fix(cmd-j): pass browser tab ownership into palette search

* test(cmd-j): cover restored browser recency in the ownership regression

The same unifiedTabsByWorktree map that establishes host ownership also
feeds lastActiveAt, which orders Open Tabs and renders the row's session
age. That half of the fix had no coverage, so assert it alongside the
execution host.
2026-09-06 18:28:51 -07:00
Neil c13d37036a fix(terminal): preserve ordinary foreground command names (#18882)
* fix(terminal): preserve ordinary foreground command names

* refactor(terminal): reuse the non-shell foreground check in inspection

Fold the duplicated isShellProcess call into one binding shared by the
ordinary-name fallback and hasChildProcesses. No behavior change; the
focused daemon inspection suites still pass.
2026-09-06 18:28:48 -07:00
Neil 57d4f63ac3 test: refresh palette identities and structured-session journal fixtures (#19165)
* test: persist palette fixture names across inventory refresh

* test: locate palette workspaces by host-qualified identity

* test: supply journal activity clocks in branch-rename fixtures
2026-09-06 18:20:49 -07:00
github-actions[bot] 8b197ffdc2 Update README downloads badge 2026-09-07 01:04:26 +00:00
Neil 2e8fa3fe9b test: exercise packaged browser compatibility in scheduled CI (#19157)
* test: exercise packaged browser compatibility in scheduled CI

* test: record final packaged workflow participation evidence

* test: expose manual packaged revision and simplify executable check

* test: reject missing package checksum assertion
2026-09-06 17:42:49 -07:00
Brennan BensonandMerge Sim c49345d358 Fix native chat completion sorting and restored activity timestamps (#19144)
* Fix structured native chat completion sorting and timestamps

* Preserve native chat activity across settled updates and host upgrades

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 17:41:56 -07:00
Brennan BensonandMerge Sim b7b6ea3942 fix(native-chat): auto-rename the workspace on a structured chat's first turn (#19138)
* fix(native-chat): auto-rename the workspace on a structured chat's first turn

Structured native chat (Claude and Codex) never reached the first-work
workspace rename. The orchestrator has a single production caller, the
agent-hook server listener, and structured sessions never set
ORCA_PANE_KEY, so no hook event could ever be attributed to one. The
renderer knew this and suppressed pendingFirstAgentMessageRename for
structured launches at three sites, which also closed the gate the
folder-workspace title rename depends on.

The host's status feed already computes the exact edge: status 'working'
with a latestPrompt normalized the same way the hook payload is, and a
workspaceId that IS the worktree id. Publish that projection to the host,
thread it out to the runtime, and hand it to the same orchestrator the
hook path uses.

Re-projections of state the host already knew (restore, an arriving
subscriber) are flagged as replays and map to the orchestrator's existing
isReplay gate, so a host restart cannot rename off a stale journal.

One host and one journal serve both providers, so this covers Claude and
Codex together.

Verified in a live Electron instance, worktrees created through the real
composer and prompts sent through the real chat composer:
  Codex  langouste -> retry-helper-exponential-backoff
  Claude prowfish  -> parse-csv-headers

* fix(native-chat): preserve first-work rename across runtime and queued turns

* fix(native-chat): skip branch rename for folder projects

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 17:38:06 -07:00
Brennan BensonandMerge Sim 41934759ea fix(windows): reject stale parent PID links in shutdown snapshots (#19149)
* fix(windows): reject stale parent PID links in exit snapshots

* refactor(windows): share the walk's pid index in the stale-link filter

Resolve parent links through the same index the descendant walk builds, so a
table that repeats a pid answers both the same way, and drop the non-null
assertion on the walk by keeping the "cannot see" null contract.

Pin the two filter branches nothing exercised: the root surviving its own
recycled ppid, and the root's start bounding a link whose claimed parent
denied its creation time.

* test(windows): pin the root creation-time floor and its tie

The floor clause survived deletion: for a chain of timestamped rows the
per-parent check already enforces order transitively, so it only does work
below a row that denied its creation time -- admitted unchecked, and its
children then find no parent time to compare against either. Cover that chain
with a child at the root's exact timestamp, which a same-millisecond spawn
produces routinely, and one that predates the root.

Also pin that pruning a link drops the unidentified rows beneath it from the
count, since a retained one would cap the verdict at unverifiable over a
process the root never owned.

Record why ties pass, what the floor is for, and the clock monotonicity the
filter assumes.

* docs(windows): say why the pid index is shared with the walk

The index is not reused across the two calls -- the walk indexes the filtered
array -- so name the actual reason: a repeated pid must resolve first-wins, the
way the walk resolves it, rather than last-wins as a Map over the rows would.

* docs(windows): describe why both pid lookups share one index

* docs(windows): put each pruning rationale on the code it justifies

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 17:36:29 -07:00
Neil 6fd03a74ef fix(ui): ignore the persistent workspace list when detecting overlays (#18881) 2026-09-06 16:44:52 -07:00
Brennan BensonandMerge Sim f7d5216016 Show provider activity in chat turn tails (#19055)
* feat(chat): show turn-scoped activity tail

* fix(chat): keep turn activity broad

* feat(chat): surface provider activity in turn tail

* fix(chat): keep reasoning headline as activity and widen redaction

A Codex reasoning summary streams as a bold headline followed by body text.
Folding the whole summary into the tail leaked literal ** markers and body
prose; only the first non-empty line is activity copy, and an unterminated
bold header mid-stream is unwrapped too.

Redaction used a hyphen for GitHub token prefixes (they use an underscore),
and missed fine-grained GitHub tokens, AWS access key ids, JWTs, URL
userinfo passwords, and bare token= values.

* fix(chat): wait for a complete reasoning headline

A bold headline still streaming has no closing marker yet; holding the
previous activity copy until it lands avoids flashing a half word.

* refactor(chat): drop bespoke secret redaction from activity copy

Reference agent hosts render provider-derived status text unredacted;
this table was the only one of its kind and its GitHub pattern matched
no real token. Bounding and the reasoning-headline extraction stay.

* Bound provider headline updates and clear activity on reconnect

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 16:42:18 -07:00
Brennan BensonandMerge Sim ad4dc353f3 fix(native-chat): settle a structured send the provider proves it received after the ack window (#19140)
* fix(native-chat): settle a structured send the provider proves it received after the ack window

A send waits a bounded window for the provider to echo the message it was given.
On timeout the dispatch resolves `unknown`. The echo that arrives later IS matched
— `recoverLateIdentity` uses it to repair the session's turn identity — but nothing
tells the journal, and `unknown` is terminal there. The submission stays unknown for
the life of the session.

Two consequences, both reachable on any ordinary session:

- The composer renders "Message delivery is unconfirmed." with a Retry, forever,
  for a message that was delivered and answered.
- Retry redispatches, because the host only replays a recorded outcome unless
  `retryUnknown` is set, which that button is the only thing that sets. So the
  banner is a duplicate delivery armed and waiting for a click — and a user who
  believes the banner and resends is doing exactly that by hand.

Every send made while a turn is already running takes this path: the provider does
not echo a queued message until the running turn ends, which is far past the 10s
ack window. Sends made while idle are unaffected, which is why this reads as
intermittent.

Carry the `clientMessageId` on the dispatch waiter and settle the journal
submission `accepted` when the late echo proves delivery. Deliberately unfenced
against the dispatch sequence: that fence decides which turn owns the identity,
while delivery is settled either way. Already-terminal rows are untouched.

* fix(native-chat): persist late dispatch receipts before session close

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 16:41:09 -07:00
Brennan BensonandMerge Sim ade9718557 fix(native-chat): suppress provider user echoes in Claude and Codex (#19136)
* fix(native-chat): keep provider user echoes out of the conversation

* fix(native-chat): retain input beside Codex skill context

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 16:28:10 -07:00
Brennan BensonandMerge Sim c7bcfa750a fix: restore the full sidebar agent row for structured native chat (#19137)
* fix: restore the full sidebar agent row for structured native chat

The host status feed projected only state, prompt, and agent type, so a
structured Claude/Codex row fell back to the tab title and the agent-type
label where a hook-reported row shows the running tool, the agent's last
message, and the model.

Project the tool line and the newest assistant prose from the journal, and
take the model from the session record's acknowledged options. The tool scan
stops at the live turn's lifecycle row and only runs while a turn is running,
so an abandoned call from a crashed turn is never reported as live work. The
assistant line is bounded to the shared preview cap rather than the hook
field's 8 KB body: a streamed reply re-projects on every journal checkpoint,
and the row renders one line of it.

* fix: keep structured session status current

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 16:20:52 -07:00
Neil 75c1f32f81 fix(gh): log when gh/glab is killed at its deadline (#18555) 2026-09-06 16:17:54 -07:00
Neil 51a17db7e3 fix(ui): keep source control headers readable in narrow sidebars (#19146)
* fix(ui): contain source control header actions in narrow sidebars

* fix(ui): preserve source control headings and conflict status at narrow widths

* chore(ui): rely on shared section toggle padding
2026-09-06 16:08:03 -07:00
Brennan BensonandMerge Sim 20eea184cc feat(native-chat): offer the link-action popover for chat links (#19130)
* feat(native-chat): offer the link-action popover for chat links

A plain click on an http(s) link in a native chat transcript opened the
system browser outright, ignoring the link-routing preference the same
link honors in the terminal. Chat now shows the terminal's destination
popover, with the modifier chords routing straight to a destination.

The popover, its request type, the destination policy and the routed open
move out of terminal-pane so both surfaces share one implementation; the
catalog keys keep their original namespace because they carry shipped
translations. Chat resolves its link owner from the session workspace
(runtime, then SSH, unresolved stays unknown) so a remote transcript only
offers Orca Browser when that host's managed browser route is eligible.

The existing toggle now governs both surfaces, so it is retitled; with it
off a chat link still opens on a plain click instead of going dead.

* Fix native chat link popover lifecycle and keyboard anchoring

* test(native-chat): use one store mock for link actions

* fix: update reliability gate for shared link popover tests

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 15:52:50 -07:00
Jinjing a224e2da74 Improve cmd j ranking (#19005)
* Refactor Cmd+J ranking to semantic-first ordering with activity bucketin

Replaces the old score-based ranking with a semantic-first contract that
compares destination, recovery, word match, coverage, strength, and placement
before using age buckets and recency to break ties. Adds explicit field roles
(primary, secondary, alias, container), identity encoding, and activity-based
bucketing so recent activity never overrides semantic relevance. Removes the
substring-elision deduplication of secondary fields. This fixes the fixture
where titles like "atlas-follow-up.md" beat recently active "Clarify Atlas
action items".

* Encode palette IDs and display secondary matches as badge

- Structured identity encoding for consistent ID handling
- Badge+tooltip reduces clutter of additional secondary matches
- Reorder activation to refocus group after state updates

* Encode tab palette identities to resolve collisions across hosts and wor

- Use composite keys (executionHostId, worktreeId, tabId) to uniquely identify tabs
- Validate tab accessibility before activation to prevent mutation on invalid state
- Extract getActivatableBrowserWorkspaceTab for consistent browser workspace validation
- Refactor workspace tab validation with stricter collision and ownership checks
- Remove unused comparePaletteActivity and mergeCandidateSummaries functions

* Update palette identity tests to use encodePaletteIdentity

Replace manual command-item ID construction with encodePaletteIdentity()
to include host and worktree context, ensuring tests match the encoding
scheme. Also adjust component styling (flex-1→flex-auto) and make
HighlightedText highlight class customizable for secondary match badges.

* rm design doc

* Use stable field identity and field objects for ranking optimization

- Add proofIdentity field to enable consistent tiebreaking in matches
- Pass field objects in FieldHit instead of fieldId strings
- Encode metric keys as numbers via bitwise operations
- Eliminate document lookups for field coverage calculation

* Reject hostless tabs when worktree IDs are ambiguous

When worktree IDs collide across hosts, hostless tabs cannot be safely
attributed. Refuse activation to prevent accidental host switching.

Improve badge accessibility by keeping it out of tab order and
exposing secondary matches through screen reader text only.

* Improve cmd-j palette ranking with token-count tiebreakers and identity

Add containerOnlyTokenCount and recoveryTokenCount fields to distinguish
entities when match quality is equal, enabling better ranking of results
that rely on container fields or recovery mechanisms. Cache paletteIdentity
in search results to avoid repeated encoding during sorting. Extract omnibox
field filtering and open-tab capping into reusable functions. Optimize
evidence-unit iteration to only process matched units. Strengthen worktree
ambiguity checks to reject hostless tabs when IDs collide across hosts.

* Add clarifying comments to palette ranking retention logic

- Document why capPaletteSection retains the selected match
- Explain retainedResultId's role in keeping keyboard selection visible
- Clarify secondaryMatches exposes additional match offsets

* Centralize palette identity and unify host ownership resolution

- Compute palette identity at search result level instead of constructing ad-hoc
- Include folder workspaces in palette ownership via getPaletteOwnershipWorktreeIds
- Add duplicate detection to filter colliding tab, page, and file IDs
- Refine ranking with containerOnly metric and source-order tiebreakers
- Improve secondary matches badge accessibility for keyboard users

* Route same-target SSH worktrees through paired runtime owners

- Centralize worktree palette identity resolution via getPaletteWorktreeIdentity
  and getPaletteWorktreeExecutionHostId, which use runtimeOwnerEnvironmentId when
  present instead of physical hostId
- Deduplicate worktrees by palette identity to keep same-target SSH worktrees
  distinct when paired with different runtime environments
- Replace scattered getWorktreeHostIdentity calls with new palette-specific
  resolution functions across palette components and search logic
- Fix accessibility: move badge out of tab order, expose extra matches through
  row text instead of interactive tooltip

* fix static analysis
2026-09-06 15:45:41 -07:00