Commit Graph
11951 Commits
Author SHA1 Message Date
Brennan Benson 89cf55dfc8 fix(agent-launch): report a launch that failed before spawning as failed, with its cause (#22913)
* fix(agent-launch): settle a launch that failed before spawning as failed, with its cause

A launch into an existing workspace whose terminal create threw before the
spawn request left this process (agent disabled, no launch command, runtime
unavailable) created nothing, yet agent.launchReplay recorded and answered it
as agent_session_operation_unknown. createTerminal now reports when it hands
the spawn to the pty controller; a failure before that point settles the
ledger row as failed and returns the original error. After the request leaves,
the outcome stays unknown: an SSH or daemon spawn whose reply was lost may
still have started.

* test(agent-launch): expect the spawn-dispatch hook on the launch's terminal create

* test(agent-launch): drive the pre-spawn failure with a missing launch command

A disabled agent is moving to a check made before either launch route runs,
so the tests use a failure that stays inside the terminal build.

* fix(agent-launch): keep the not-started verdict on the launch, not the shared error

A failed pane spawn rejects the same error object into the spawner (after its
request left) and into a concurrent create waiting on that pane (before its
own). Marking the error object globally let the waiting launch's verdict clear
the spawner's, recording a launch that may have started an agent as failed.
The launch now owns its dispatch tracker and carries the decision on its
execution error instead of re-deriving it from the error.

Also names the test's launch parameter type for the anti-slop audit.

* refactor(agent-launch): move launch failure classification into its own module

Rebasing onto the caller-selection change took agent-launch.ts past its line
limit; the failure-code helpers are a self-contained concern.
2026-09-27 16:37:04 -07:00
Brennan Benson da57f47353 fix(native-chat): say a refused chat write in plain words, and keep it with the write it describes (#22999)
* fix(native-chat): say a refused chat write in plain words, and keep it with the write it describes

A send's refusal now lives on the queued message and goes when that message is sent again or delivered, so a resend that succeeds after an agent restart no longer leaves a red line under the composer. Stop, answers, option and goal changes report a refusal once as a toast; a conversation command answers inline. One shared table turns every refusal code into copy with a next step, on desktop and mobile, instead of showing the host's diagnostic text.

* test(native-chat): pin the refused-then-restarted send in the order the live app sees it

* fix(native-chat): tell the phone to send a refused message again, not to press Retry

The phone puts a refused message back in the composer and has no Retry
control, so "Retry to send it again." named something that is not there. A
phone send now reads "Your message was not sent. Send it again."

Also types the rejection-cause test's empty submissions without a cast.

* fix(native-chat): keep why a message failed as a fact, and say only what is true

A queued message saved the words of its failure to local storage, so the copy
lived in users' data. It now saves the fact: the refusal code (with the host's
words only for a failed restart, which the host writes for people), the
provider's rejection reason, or that the host could not be reached. The Retry
row chooses the words when it shows the message. A saved failure this build
cannot read is dropped, and the row says only that the message was not sent.

The words are true for every host path behind each code:
- A refusal whose code does not say why (a cleared conversation, a pending
  question, a provider's own rejection) says only what did not happen, with no
  next step that could repeat the refusal.
- An unsettled owner no longer claims the agent was restarting; it says Orca
  could not confirm which agent process owns the chat.
- A code from a newer host says only what did not happen.

Desktop translates each sentence whole, with the shared English as the
fallback the phone shows as is, so the two never say it differently. The phone
no longer shows transport or host text when a request fails without a
refusal.

* fix(native-chat): never show a host refusal message, including a failed restart's

Every refusal code has at least one host path that writes its message for a
log or carries a marker. A replayed refusal, for one, says "Operation <id> was
already refused: <code>." because the ledger stores no message. A failed
restart's message also embeds the resume's own refusal text, which can be the
ledger's. So no code's message is safe to show, and the failed-restart
exception goes. The Retry row now says "The agent couldn't restart. Your
message was not sent." with no next step, because some restarts need a new
chat. The cause is still in the chat's own status row. The saved failure keeps
only the code.

The test lists one such emitter per code, so a code that later becomes safe to
show has to be argued against that list.

* fix(native-chat): offer only a next step that works, on the phone too

A message the agent turned away on the phone said "Retry to send it again", but the phone has no Retry control; it now says "Send it again." like every other refused phone message, built from the same sentences as the other notices.

The capacity refusal said Orca was handling too many requests "for this chat" and offered a retry. The limit is counted across every chat over a day, so an immediate retry would likely be refused again. It now says so, with no next step.

* chore: restore pnpm-lock.yaml

A local install rewrote it and it was committed by mistake.

* fix(native-chat): say only what is true for every host path behind a refusal

No refusal notice says how to try again any more: the control that sent
the write already does that, and a sentence per code was false for some
of the host situations the code covers. The phone keeps "Send it again."
where a resend under the same operation id can go through, and an older
host still gets "Update Orca, then try again."

A cause is named only for codes whose every emitter means it. The owner
codes and identity_required now say only what did not happen, the
outcome-unknown sentence no longer claims the write itself is in doubt
(a send behind an unconfirmed rewind is refused outright), and a request
that failed without a refusal is recorded as `failed` rather than
`unreachable`, since most such failures are host or compatibility errors.
The census test now pins the cause allowlist and the absence of retry
sentences, and the unused sentences leave every catalog.

* fix(native-chat): don't say a chat write failed when its request may have run

A desktop Stop, answer, option, goal or conversation command whose request threw
said the write did not happen, for every error. A timeout or a lost connection can
come after the host ran the write, so a timed-out /compact said "The command didn't
run." while it ran. Only an RPC error the host answers before running the method
proves that; any other failure now says Orca couldn't confirm what happened.

The phone already drew this line; both clients now read it from one shared rule.

* fix(native-chat): say a refused background-task stop is about the task, not the agent

A background-task stop goes through agentSession.cancel, so a refusal of it
said "The agent wasn't stopped." although the agent was never asked to stop.
The write kind now reads the cancel's scope and names the task or tasks.

* fix(native-chat): say a Stop refused over a moved-on question did not stop the agent

A Stop pressed while a question or approval is pending names that prompt, so the host can refuse it
as already resolved or stale. The notice then spoke only about the question and never said the agent
kept running.

* test(native-chat): check every refusal notice cell against one spec

Replaces the per-rule loops with one table of every failure (each refusal
code, a failed or unconfirmed request, and a code from a newer host) by
every write kind, and five rules each cell must keep: a certain refusal
says that write did not happen, once; a write that may have run never
says it did not; only "Update Orca" and the phone's resend where it can
work say to try again; a cause is named only for allowlisted codes; no
cell is empty or shows the host's text.

* fix(native-chat): the launch path drops a message's old failure when it sends it again

The first message of a new chat is sent by the launch path, which staged it
without removing the failure saved by an earlier attempt. A message that failed
through the outbox and was then refused again through the launch path showed the
earlier reason in its Retry row. Both paths now stage through one shared step
that removes it.
2026-09-27 16:34:42 -07:00
NeilandClaude b61797dce3 fix(agent-trust): bound the remaining Codex trust writes a launch waits on (#23380)
* fix(agent-trust): bound the remaining Codex trust writes a launch waits on

#23148 capped the trust write on the agentTrust:markTrusted IPC handler and the
worktree-remote startup path, but three call sites still awaited
markCodexProjectTrusted with no bound: the main-process worktree startup used by
orchestration worker-start, Codex quick-launch preparation, and Codex session
resume preparation. All three share the per-config.toml lane with hook installs
and app-server trust grants, which has no cap on queue depth, and the SSH writer
adds a resolveHome round trip plus unbounded SFTP over a possibly half-open
link. A wedged lane left the user clicking "start agent" with nothing happening
and no error.

Each site now routes through the existing awaitAgentTrustWriteWithinDeadline
helper with unchanged semantics for its two current callers. Abandoning at the
deadline never cancels the write, writes nothing else, and never substitutes a
local fallback, so the workspace stays untrusted and Codex raises its own trust
prompt. The await still sits ahead of the runtime-home resolution and hook
repair that the PTY spawn waits on, so the ordering #23148 established is kept.

Reliability only; there is no speed gain. The win is that a launch cannot hang.

Verified with 2,755 agent-trust/Codex/startup/runtime tests, the 53-test
reliability-gate command, the gate manifest check, typecheck:node, changed-code
quality, oxlint and oxfmt. Red-green: with the three sites reverted to a bare
await the new bounding tests hang to the 30s vitest timeout (3 fail/14 pass);
restored, 17 pass.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(reliability-gates): stop the trust-preflight gate claiming a bounded launch

The gate's performance budget said a stuck predecessor "degrades to the
agent's own trust prompt instead of an indefinite caller wait." That is false
at two of this branch's three new sites. In startup/codex-launch-preparation.ts
the next awaited call after the abandoned write is
codexHookService.prepareRuntimeHomeForLaunch; in
startup/codex-session-resume-launch.ts it is installForLaunchPrep or
refreshRuntimeUserHooksForLaunchPrep. Both reach
runExclusivelyForRuntimeAndSystemTrustConfig, which takes the same
runExclusivelyForCodexTrustConfig queue - FIFO per config.toml, no depth cap
and no timeout - on the runtime home and on ~/.codex/config.toml that
markCodexProjectTrusted itself takes. A wedged lane still stalls those two
launches one step later, with the abandoned write holding its queue slot ahead
of the hook step. Only markLocalWorktreeTrusted has nothing after its Codex
branch and so is bounded end to end.

The budget, invariant, oracle and coverage notes now say what is true: all five
markCodexProjectTrusted call sites are bounded, so no trust write hangs a
caller indefinitely, but the bound is on the write and not on the launch. A new
knownGaps entry records the residual stall and that the new launch-prep tests
mock codexHookService wholesale, so nothing in these suites can catch it. The
already-complete-write case stays labelled a timer control, not a red.

Wording only; no production code, test or count changed. The cited command
still reports 6 files / 53 tests, and the gate manifest check passes for 138
gates.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(reliability-gates): name the first unbounded re-entry, not the hook-service call

The trust-preflight gate cited codexHookService.prepareRuntimeHomeForLaunch as
the immediate next step after the abandoned write in
startup/codex-launch-preparation.ts. Tracing the module, the first awaited step
is ensureRealHomeHooksIfSelected. It is conditional: only when the target is not
WSL and runtimeHome.isHostSystemDefaultRealHomeSelected(launchEnv) is true does
it call ensureRealHomeCodexHookState, which chains behind that module's own
serial ensure promise and then acquires runExclusivelyForCodexTrustConfig
directly on ~/.codex/config.toml - the same lane markCodexProjectTrusted's inner
acquire takes. When the real home is not selected that call returns without
awaiting the lane, runtimeHome.prepareForCodexLaunchAsync is synchronous on the
non-WSL path, and only then is prepareRuntimeHomeForLaunch the first re-entry.
So the gate could miss a launch stall one step earlier than the one it named.

startup/codex-session-resume-launch.ts had the same omission: its hook branch
takes ensureRealHomeCodexHookState when the resume home is ~/.codex, and only
otherwise installForLaunchPrep or refreshRuntimeUserHooksForLaunchPrep. The
coverage notes also understated the mocking - codex-launch-trust-write-deadline
mocks ensureRealHomeCodexHookState as well as codexHookService and defaults the
real-home selection to false, so neither lane is driven by any suite here.

The performance budget, coverage notes and the residual-stall knownGap now name
the first re-entry on each path and say when it applies, rather than only the
later hook-service call.

Wording only; no production code, test or count changed. The cited command
still reports 6 files / 53 tests, and the gate manifest check passes for 138
gates.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(reliability-gates): scope the trust-preflight invariant to the local writes

The invariant claimed every Codex trust write a launch waits on is bounded or
abandoned at a deadline. The entry's own knownGaps contradicted that: the remote
write reached through markRemoteWorktreeTrusted in
runtime/runtime-worktree-agent-startup.ts is awaited with no timer, and it is a
Codex write, not an adjacent one - it reads TUI_AGENT_CONFIG[agent].preflightTrust,
which is 'codex' for the codex agent, and markRemoteAgentWorkspaceTrusted then
branches into markRemoteCodexProjectTrusted. Its one production caller,
markRemoteWorkspaceTrustedForAgent, is reached with a connectionId from seven
launch entry points; five await it directly before the createTerminal that spawns
the agent, and two gate the launch command they hand back for their caller to
spawn. It chains a session.resolveHome round trip plus SFTP realpath/read/mkdir/
write, none with a timeout, over a possibly half-open link - the same hang this
branch bounds locally.

The invariant now claims only the five local markCodexProjectTrusted sites and
names the remote exception inline instead of leaving it to knownGaps. Audited the
rest of the entry against the code in the same pass: performanceBudget said "no
trust write can hold its caller indefinitely" (now scoped to those five); the
"all five awaited Codex trust writes" tally is now "all five awaited local" and
points at the unbounded sixth; the remote gap no longer calls that path merely
"preset-agnostic and out of Codex scope"; oracle and coverageNotes now record
that nothing here drives markRemoteWorktreeTrusted, since the remote-preset
suite only checks what markRemoteAgentWorkspaceTrusted writes, never how long it
may take; and surfaces gains the remote writer that suite actually covers.

Re-verified the rest rather than assuming it: the launch-prep re-entry chain,
the FIFO no-cap no-timeout trust-config queue, the IPC handler capping both its
remote and local writes with no cross-branch fallback, and that
upsertProjectTrustLevel and upsertProjectTrustLevelInContent remain the only
producers of a project trust_level, so no further writer needs covering. The
already-complete-write case stays labelled a timer control, not a red.

Wording only; no production code, test or count changed. The cited command still
reports 6 files / 53 tests, and the gate manifest check passes for 140 gates.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 16:30:30 -07:00
Jinjing 29847641ab test(e2e): fix failures and improve stability (#23480)
- Narrow toolbar width and set window minimums for consistent testing
- Add node_modules symlink to fixture for ESM import resolution
- Exercise manual paging and fix button selector
- Preserve repo filters in reveal workflow
- Adjust timing strategy and increase test timeout
2026-09-27 16:17:06 -07:00
Neil fb51e6cb4c Speed up Markdown document discovery with bundled ripgrep (#23306) 2026-09-27 16:16:20 -07:00
Jinwoo Hong 433986fa3b fix(runtime): read Codex readiness from the live screen (#23475)
* fix(runtime): read Codex readiness from the live screen

Codex 0.157 runs in embedded mode when Orca passes `-c model_reasoning_effort=…`,
shows a startup warning, and repaints its header by cell diff
(`ESC[5;3Hdir ESC[5;7Hctory:`). The tui-idle body check read the line-folded
tail, which drops those cursor moves and reads `dirctory:`, so worker-start
timed out at agent_readiness while Codex sat idle at its prompt.

For Codex (or unknown) panes whose live emulator screen shows the Codex header,
the screen now decides readiness: `model:` and `directory:` present and neither
still `loading`. Otherwise today's text rules apply unchanged. All six tui-idle
satisfaction sites share one helper so they cannot disagree.

Fixes STA-8628 / #23241.

* refactor(runtime): read the Codex screen lazily and only for Codex panes

Pass the screen as a thunk so non-Codex panes never build the grid, hoist the
pane agent lookup, and share the unblocked-ready check with the Muse rule.

* fix(runtime): let the Codex screen only add readiness

A grid out of step with the PTY (size mismatch or a resize mid-paint) can garble
Codex's header. Keep every verdict the text rules give today and consult the
screen only when they say not ready, so no flow that settles today can stop.

* fix(runtime): read Codex readiness from the header box only

Chat below the header can mention "OpenAI Codex" or "model: loading", so the
screen rule now reads model/directory/loading only inside the header box.

The serialize round-trip sweep picks up every runtime fixture; record the
pre-existing header-border attribute divergence the new Codex 0.157 recordings
expose, which this change does not touch.
2026-09-27 18:48:23 -04:00
NeilandClaude 161bdf93c3 fix(bench): load benchmark modules under test through jiti (#23482)
`pnpm run bench:terminal-partial-escape-tail` and
`pnpm run bench:worktree-refresh-churn` both died at startup with
ERR_MODULE_NOT_FOUND. Each entrypoint static-imported a `src/` module with an
explicit `.ts` extension, but bare `node` type-stripping cannot resolve the
extensionless relative specifiers *inside* that module's graph
(`terminal-partial-escape-tail.ts` -> `./terminal-escape-introducer`,
`worktree-catalog-reconciliation.ts` -> `../../../../shared/structural-value-equality`).

Routes both through jiti, matching the four benchmarks that already load `src/`
TypeScript that way (`pty-source-ack-boundary`, `locale-collator-sort`,
`worktree-base-pending-marker`, `wsl-git-shell`). Adding the extension at each
import site was the alternative, but `config/tsconfig.node.json` does not set
`allowImportingTsExtensions`, so a `.ts` specifier in `src/shared` fails
typecheck with TS5097 -- and no file under `src/` uses that shape today.

Developer tooling only; no production code changed.

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 15:34:38 -07:00
Brennan Benson 993183afd7 fix(terminal): every explicit terminal close commits through one main transaction (#22929)
* fix(terminal): every explicit terminal close commits through one main transaction

A renderer save cannot shrink terminal membership once main owns a repo's
topology, so desktop tab and pane closes, CLI split-pane closes and mobile
split-pane closes only became durable when the killed process's exit retired
the surface. A close whose kill failed or threw, or whose exit was never
certified, came back after a reload.

Every close now reaches closeTerminalSurface: the renderer sends an explicit
intent for user and cleanup closes, the CLI and mobile split-pane closes commit
the pane after their stop, and the headless and relayed mobile closes reuse the
same commit. A failed flush keeps the in-memory removal and no longer cancels
the kill. Exit retirement is unchanged.

* fix(terminal): tell the desktop renderer to drop a split pane main closed

A CLI or mobile close of one pane in a split commits the pane in main, but the
desktop kept showing it until reload when no exit arrived to remove it. The
close now sends a leaf-addressed notice: a mounted pane closes by leaf id, and
a parked tab collapses its stored layout. Addressing by leaf makes the notice
and the renderer's exit handling no-ops after each other, which replaces the
numeric pane-id notice that could close the whole tab when the exit won.

* fix(terminal): a pane close never widens into a whole-tab close

A leaf-addressed close fell through to the whole-tab close whenever main's layout no longer
held that leaf as one of several. Main's exit handling retires an exited split pane from the
saved layout, so closing that pane afterwards (the exited-pane overlay's Close, or a CLI close
whose stop delivers the exit first) removed the whole tab, live sibling included, and the
next renderer save could not restore it. A pane close is now a no-op unless its leaf is in a
multi-pane layout.

Also updates two mobile split-close assertions to expect the leaf-addressed notice, and adds a
test that a relayed mobile close of a renderer-listed tab still reaches the renderer's pin guard.

* test(terminal): cover the PTY-handle branch of a CLI split-pane close

The existing CLI split test resolves its handle through the renderer graph, so the branch
that closes a runtime-owned pane by its PTY handle had no test failing without its commit.

* fix(terminal): a CLI pane close with an unconfirmed stop closes only that pane

`orca terminal close <handle>` on one pane of a split used to close the
whole tab, live sibling included, whenever that pane's stop could not be
confirmed (for example an unreachable SSH host). An unconfirmed stop is
unverifiable, not a reason to drop siblings: the close now commits only
that leaf, tells the renderer to drop that leaf, and leaves the owed kill
to the controller's existing SSH pending-kill path.

On a host where no renderer lists the tab, main now also removes the
closed pane from the paired-client snapshot (with its retirement proof),
since no exit may arrive to do it.

* refactor(terminal): one resolver decides whether a pane close becomes a tab close

Every explicit close now states its target as `{kind:'tab'}` or `{kind:'pane', leafId}`; no
optional leaf id silently means the whole tab. Main resolves a close it started in exactly one
place, reading the copy of the tab's panes its layout owner holds (the renderer-published layout
for tabs the desktop renderer lists, main's session layout otherwise). Only `last-pane` escalates,
through the existing tab path so the renderer's pin guard still runs; an unknown pane never widens.

- The CLI and phone paths drop their per-site sibling counts for the resolver.
- The notifier splits into a tab-only close and a leaf-addressed pane close.
- The headless tab closer takes a parent tab id, so a pane row cannot reach it.
- A phone close of one pane on a host with no desktop window now stops and closes only that pane.
- A phone close of one pane with no live process record closes that pane, not its tab.

* fix(cli): an unverifiable stop says the close happened

`orca terminal close` still exits 1 when the process stop cannot be verified, but its message now
says the terminal was closed and names the host's reason, instead of "close failed". It promises
that the kill retries on reconnect only when the SSH relay itself never answered the stop, the one
case a recorded kill order backs.

* fix(terminal): a phone pane close commits even when its kill fails

A paired client's close of one pane threw `terminal_close_failed` before committing anything when
the controller reported the kill failed, so the pane stayed. The kill is now best-effort, as it is
for a whole-tab close: the pane's removal always commits and the failure stays on the PTY's
liveness verdict.

* fix(terminal): a pane close widens only when a copy shows it is the last pane

The close resolver read an owner copy that records no panes as "the tab has
one pane", so a CLI close of one pane of a split, addressed while the
renderer listed the tab before publishing its panes, closed the whole tab.

Every copy now counts only if it records at least one pane, read in the
owner's order with the published rows as the last fallback, and a pane
close widens only when a copy lists that pane as the tab's only one. An
unsplit tab whose saved layout predates its pane still closes: its
published row names the pane.

* fix(cli): promise a kill retry only when the host recorded the kill

The close receipt inferred "the kill retries when the host reconnects" from
the stop reason's text, which a new transport message or a reworded error
would silently break.

An explicit close now records the replayable kill order when its stop goes
unconfirmed, before sending the follow-up kill (whose own failure is
recorded only once its RPC settles), and reports that on the receipt as an
optional `pendingKillRecorded`. The CLI promises the retry only from that
field, so an older host, which never sends it, gets no promise.

* test(pty): justify the controller cast the recorded-stop tests extend

* fix(terminal): parse the close target with typed narrowing

The low-evidence lint gate rejects Reflect.get and broad object parameters,
which failed static analysis. Narrow with 'in' checks instead and cover the
boundary parser's accept and reject cases.

* fix(terminal): a desktop tab close is not refused by a split that bound while it waited

The renderer has already removed and killed a tab it closes, so its close intent now skips the
owner fence phone and CLI closes use. Before, a split pane whose binding was admitted between the
close request and its durable write made main refuse the close, and the tab came back on the next
launch whenever its processes did not exit.

* chore(terminal): note that closedByLayoutOwner goes away once main owns the terminal layout

* test(terminal): reload the close-intent fixture through the SQLite profile store

Main now requires a SQLite profile-state authority for a writable Store, so the save-and-reload
close tests build and reopen their store through the shared SQLite test harness.
2026-09-27 15:29:22 -07:00
Brennan Benson f72079bd21 fix(mobile): send typed question answers as structured answers (#23458)
* fix(mobile): send typed question answers as structured answers

The phone packed a typed answer into agentSession.respondToQuestion's
`optionId`, a field capped at 1024 characters, so a long typed answer was
refused and never reached the agent.

The phone now builds per-question answers for single and grouped questions
and sends them as `answers` when the host advertises
agent-session.question-answers.v1, falling back to the packed `optionId`
for older hosts. An answer too long to pack is sent as `answers` while the
host's support is still unknown, and on a host known to predate the field
the phone asks the user to update the computer instead of sending an
answer the host must refuse. The host features the phone negotiates for
structured sessions (prompt cancel, question answers) travel as one
required object from the shared capability probe instead of one optional
boolean each.

* test(mobile): cover a long typed answer on a grouped question and the default question id

* fix(mobile): keep the chat controller's host support argument optional

* fix(mobile): refuse a grouped answer too long for an older host on the step that overflows

Packed grouped answers only grow, so an overflow on an earlier step could never reach an older host. Refusing it only at the final submit left the long text folded into the draft with no way to shorten it short of cancelling the question.

* fix(mobile): refuse a grouped step that leaves no room for the rest on an older host

* refactor(mobile): check the packed answer length once, at send
2026-09-27 15:17:38 -07:00
Jinwoo Hong 618a8b0758 fix(terminal): keep a dead app's input modes recoverable after a refuted proof (#23474)
A command's armed input modes were demoted at the first OSC 133;D whether or
not the foreground proof confirmed, and the alternate-screen trigger was spent
the same way. A full-screen agent's nested command shells leak their own D onto
the main PTY; the refuted proof for that stray D used up the only trigger, so
the app's real death later went unrecovered.

Every D now re-asks while a command's mode (or the alternate screen) is still
up; only the confirmed ground or the app's own disable clears it. Proofs that
can never succeed (wsl.exe, no Windows job reads) already refute immediately
without spawning anything. The accepted cost is one process read plus a bounded
output hold per D while a command-owned mode stays up and the proof keeps
refuting: a live agent leaking D, a stopped job, a nested subshell, or an rc
that execs another shell.
2026-09-27 18:17:04 -04:00
Jinwoo Hong 10807d8b45 test(e2e): stabilize terminal shortcut Kitty and Ctrl+C tests (#23472)
Shift+Enter: the host now grounds keyboard modes a finished command left
armed, so the test's one-shot `printf '\033[>1u'` was reset to 0 before the
poll. Keep the command alive via `read` while asserting, then let it pop
its own flags on Enter.

Ctrl+C: each test opens a fresh tab, and SIGINT during shell startup can
kill the shell and close the tab. Wait for a ready prompt before
interrupting. The existing post-interrupt assertions are unchanged.
2026-09-27 18:16:49 -04:00
Jinwoo Hong 0547b32224 test(usage): add long-context token fields to the Codex fixture (#23478)
#18759's test fixture predates #22360 adding required long-context token
fields to Codex session and location breakdowns, so main's typecheck fails.
2026-09-27 17:52:43 -04:00
leilei3167 838769d73e fix(linear): accept team keys that start with a digit (#23423)
Fixes #23422
2026-09-27 14:30:45 -07:00
Mr. ZengandNeil ef987e42d7 feat(analytics): persist local usage session identities (#18759)
Co-authored-by: Neil <neil@stably.ai>
2026-09-27 14:29:19 -07:00
Neil 32668c3417 perf: preserve matching Accounts settings sections during search (#23192)
* perf: preserve matching Accounts settings sections during search

* refactor(settings): route Accounts' section stack through SettingsSectionStack

Replaces the pane-local index-keyed wrapper with the shared SettingsSectionStack so the
search-stable identity rule lives in one place, and restates the read/subscription-count
oracles as node identity plus non-growing counts so a future freshness fix cannot read as
a regression. Applied on the branch head; the rebase onto main is still outstanding.
2026-09-27 13:51:56 -07:00
Jinjing eaa6fb71b0 Tab opening ranking (#23435)
* Add literal-basename file matching for tab opening

- Distinguish literal filename matches from fuzzy matches
- Prioritize search for short tokens (1-2 Latin characters)
- Sort literal matches naturally instead of by fuzzy score
- Reserve search option even when literal matches fill results

* Keep literal-basename file matches in tab entry options

Ensure literal-basename file matches appear in tab creation suggestions
as fallback options when exact-basename matches exist, improving file
discovery for users.
2026-09-27 13:50:12 -07:00
Neil a4b60c2ea3 test(windows): stop the NSIS capability probe failing on a slow PowerShell cold start (#23381) 2026-09-27 13:24:48 -07:00
Neil 3f1bf68c84 perf: preserve matching General settings sections during search (#23177) 2026-09-27 13:21:54 -07:00
Brennan Benson bdaf9a3ed3 fix(agent-launch): move only the requesting client's view to a launched tab (#22914)
* fix(agent-launch): move only the requesting client's view to a launched tab

A launch into an existing workspace from a paired client (a phone, or a
desktop client of a remote server) activated a new chat for every client and
the host, and recorded no selection for the caller's terminal. It now
publishes the chat without activating it for everyone and records the new tab
as the caller's own selection, the way that client's own tab tap does
(session.tabs.activate with caller navigation). In-process callers and
workspace-creating launches are unchanged. Selecting the tab is bookkeeping:
a failure there never fails the launch.

* fix(agent-launch): select a launched tab for its caller the way a create does

The launch recorded the caller's selection through the tab-tap path, which
refreshes the PTY inventory (twice for a chat) on the reply path and treats a
non-ready pane as a wake gesture that respawns its agent. Record it through
the same caller-navigation step session.tabs.createTerminal uses after its own
create: no host focus, no refresh, no materialize.

* test(agent-launch): prove a launched tab's selection spares subscribed clients and never respawns

* test(agent-launch): pin that the host's own desktop window still activates a launched chat

The desktop renderer reaches agent.launch as a runtime client with no paired device, so it must keep the host-wide activation rather than be treated as a paired caller.
2026-09-27 13:17:18 -07:00
Brennan Benson aeb96ab0e2 fix(native-chat): close only the reopened chat when main closes its tab (#22922)
After /clear the current chat keeps the window tab id derived from its first
session, so reopening that first session from history lands in a suffixed tab.
Main closed chat tabs in the window by re-deriving that id from the session,
which named the current chat instead of the reopened one: closing the reopened
tab (on the desktop or from a phone) also closed the current chat and stopped
its Claude.

Main now sends its own tab id, as it already does for every other tab kind, and
the window translates it through the host-to-window tab mapping it keeps for
mirrored chat tabs. The same translation lets a host-directed focus reach the
reopened tab. The close handoff marker is keyed on the id actually sent.
2026-09-27 13:10:59 -07:00
Brennan Benson 9795563695 fix(runtime): register phone terminal subscriptions when the request arrives (#22948)
* fix(runtime): register phone terminal subscriptions when the request arrives

* fix(runtime): don't let an already-aborted terminal subscribe take the stream slot
2026-09-27 12:35:52 -07:00
Brennan Benson 484598b18e fix(terminal): retry first-terminal seeding until a decision applies (#22919)
* fix(terminal): retry first-terminal seeding until a decision applies

The seeding effect marked a workspace as attempted before its async
activation check finished, so a StrictMode double run, an input change
mid-check, or a 'blocked' result left the workspace marked with no
terminal. Mark it only once an uncancelled, unblocked result is applied;
a rerun shares the gate's in-flight check.

* fix(terminal): read the closed-last-terminal row when the seed decision applies

The seeding callback runs after an async activation check, so a row captured
at render time can be stale by then. Read it from the store at decision time
and drop it from the effect dependencies, which no longer need to restart the
check when the row appears. Add tests for two separate empty checks and for a
last terminal closed while the check runs.
2026-09-27 11:32:41 -07:00
Brennan Benson d972bac38a fix(workspace): a restored workspace paints before its terminals reconnect (#22810)
* fix(workspace): a restored workspace paints before its terminals reconnect

After a restart the workspace content area stayed blank — no tab strip, no
pane — until the whole startup chain (SSH reconnect, PTY reconnect, legacy
worker recovery) had finished, even though the restored tab model had been in
the store for seconds. Only terminal panes need that chain, because a pane
binds a PTY on mount.

The workbench now mounts the active workspace from its hydrated tab model, and
holds its terminal tabs unadmitted (an empty admitted-tab restriction with no
deferral entry) until startup restoration has published PTY ownership; the
activation plan then replaces the hold. Chat, browser and editor panes mount
immediately.

Carried from #22293 (922c1c2626), renderer terminal files only.

* fix(workspace): track the startup terminal hold explicitly and keep its tabs watched

The startup hold was inferred from the restriction map's shape and released only
when a new hold was installed, so a workspace left for no workspace mid-startup
stayed mounted with no terminal panes and no background watchers. Track the held
workspace explicitly, release it on any pass where it is no longer held, apply it
after prune, and count held tabs as parked so their bells, titles, and
completions are still observed during the window.

Carried from #22293 (3fd421c41f), renderer terminal files only.

* test(workspace): cover the split surfaces' held-tab watcher wiring and fold the held branch

A held worktree now falls through to the existing else, which already resets the activation marker.

* fix(workspace): browser and editor panes wait for a connecting SSH host

A restored SSH workspace now mounts its panes before the host reconnects.
Derive one per-worktree host phase from the published SSH state, treating a
target startup restoration has not dialed yet as connecting. The browser route
holds its prepare until the host connects and re-derives a failed route on
connect or a new connection generation; the editor shows the connecting state
instead of a dropped connection and reloads once the host connects.

* fix(workspace): a held terminal tab shows that it is restoring

While startup restoration holds terminal panes unmounted, the visible tab's
slot was blank. Fill the slot of a visible held tab, including one created
during the hold, with a restoring state until its pane mounts.

* refactor(workspace): every pane reads one shared host-connection signal

The terminal host state now reads its SSH status from the same per-worktree
signal as the browser route and the editor, instead of resolving it on its own.
The signal carries the resolved status (an undialed target during startup
restoration reads as connecting), the owning environment, and a connection
epoch that changes on every reconnect, so no pane re-derives the no-entry case
or the reconnect generation. A nested target whose runtime cannot be seen is
reported as unverifiable rather than unavailable.

* fix(workspace): the terminal names the published SSH status from the shared signal

The shared host signal now carries the published status separately from its
phase. Only the phase reads a target startup restoration has not dialed yet
as connecting, so the terminal's reconnect overlay and error ownership keep
reporting what the host published, while the browser and editor still wait.
Tests pin that window for the terminal and pin that a host the client cannot
verify never holds the browser or editor.

* fix(startup): a degraded boot releases terminal startup restoration after reconnect

The degraded recovery path's successful reconnect set workspaceSessionReady
but never terminalStartupRestorationReady; only its failure branches did. Every
reader waiting on restoration then waited for the whole session: no fallback
terminal, no structured tab sync, the activation gate timing out, and an SSH
target never dialed at startup reading as connecting forever. Release the flag
once the reconnect finishes, unless a newer startup pass has taken over.
2026-09-27 11:31:11 -07:00
github-actions[bot] 6c75837750 Update README downloads badge 2026-09-27 18:30:40 +00:00
Brennan Benson b8ef5fcbc2 fix(native-chat): let the "Asked" row unfold a long question (#22941)
* fix(native-chat): let the "Asked" row unfold a long question

The question row clipped long questions to one line with no way to read
the rest. The row is now a toggle that wraps the full question in place.

* fix(native-chat): open a clipped question below its toggle, selectable

The first cut put the whole question inside the toggle button, where
Chromium will not select text: dragging or double-clicking in it selected
nothing and toggled the row instead. The open state also lived in the row,
so scrolling the row out of the windowed transcript, or answering the
question, folded it again.

- The full question now opens in a paragraph below the button, outside it,
  so it selects and copies like any other prose.
- The toggle is offered only when the question is actually clipped; a
  question that fits stays plain, selectable text as before.
- Open state is kept in the transcript's disclosure store, keyed by the
  message, so it survives windowing and the pending-to-answered remount.
- The chevron also shows on keyboard focus and on touch screens.
2026-09-27 11:18:13 -07:00
Jinjing 0a876b2f5a feat(activity): reveal threads in floating terminal workspace (#23239)
* feat(activity): reveal threads in floating terminal workspace

When selecting an activity thread from the floating terminal workspace, automatically open the floating panel if closed. Improves split pane focus handling to preserve active panes during thread reveals.

* fix(activity): enable floating terminal before revealing threads

Floating tabs persist even after disabling the feature, but the panel
ignores toggle events while disabled. Enable the setting first so that
subsequent toggle commands take effect.

* Use useLayoutEffect for floating terminal toggle binding

Enable-then-toggle callers dispatch events immediately, which passive useEffect rebind can miss. useLayoutEffect ensures the handler is bound synchronously before the next paint.
2026-09-27 10:48:48 -07:00
OrcaWinandm4air 27b823f934 ci: compile the E2E CLI once for all consumers (#23384)
* ci: share compiled CLI output across E2E consumers

* ci: preserve CLI setup and old-ref fallback for shared artifacts

* docs: record shared E2E CLI benchmark evidence

* docs: include final CLI reuse timing range

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 02:02:29 -07:00
Neil 069701446b fix(persistence): report a discarded cross-host repo update instead of failing silently (#23379)
`Store.updateRepo` takes an optional `hostId` that must be the row's own
`executionHostId` stamp. When it is anything else the lookup matches no row and
the method returns `null` — the same value a deleted row returns — so the write
is dropped with no error, no log, and no failing test. #22421 shipped exactly
that: identity enrichment addressed the probe host (`local`) instead of a
`runtime:` row's stamp, and every write was discarded for a full release.

Resolve the row through `findRepoRowForHostScopedWrite`, which logs once when the
id exists but only under other host stamps, naming the requested host and the
stored ones. The guard itself is unchanged: a matching host still writes, and a
client-local probe still cannot repair a peer's row.
2026-09-27 01:31:24 -07:00
OrcaWinandm4air c15f082031 ci: build independent Electron targets together for E2E (#23378)
* ci: reuse parallel Electron targets for E2E builds and guard cache action setup

* test: recognize the top-level cache repository preload

* docs: record E2E build timings and exact output parity

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 01:17:27 -07:00
Neil 17690e6b9a style: settle oxfmt 0.70 drift and stop formatting vendored licences (#23377)
The oxfmt 0.65 -> 0.70 bump landed without a repo-wide reformat, so 36 files
already in the tree no longer matched what the new version emits. Anyone running
`pnpm format` picked all of them up alongside their own change.

Also excludes `resources/licenses/**`: `oxfmt --write .` was rewriting the
vendored PCRE2 licence, turning its `*` redistribution bullets into `-`. Third
party licence text has to be reproduced verbatim, so formatting must not touch it.
2026-09-27 01:14:53 -07:00
93d8b1f042 fix(ssh): complete keyboard-interactive MFA prompt handling (#15588)
Honor SSH keyboard-interactive prompt echo and empty responses, reuse login
passwords without replaying rejected values, and stop cancelled or stale
credential requests from continuing authentication or restoring the cache.

Original implementation: Junho Kim (#8750).
Port and follow-up work: Allen (#15588).

Verified with 2,773 SSH/credential tests, full typecheck, changed-code quality,
a production Electron build, real-socket MFA fixtures, and rendered UI checks.

Fixes #8622

Co-authored-by: Junho Kim <arkimjh@illinois.edu>
Co-authored-by: microdaery <microdaery@gapp.nthu.edu.tw>
2026-09-27 01:12:29 -07:00
OrcaWinandm4air 47cebbf5d2 ci: use ARM unit runners, overlap web builds, and reuse verifier fixtures (#23376)
* test: reuse isolated mobile bundle fixtures for verifier checks

* ci: run PR unit shards on ARM and overlap independent web builds

* docs: record controlled CI overlap and runner measurements

* test: observe WebRTC packets with the host clock

* ci: isolate Windows installer CIM probe from native test load

* docs: record native probe scheduling validation

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 01:06:54 -07:00
NeilandClaude 983250a12e fix: await queued Codex trust preflight completion (#23148)
* fix: await queued Codex trust preflight completion

* fix(agent-trust): bound the trust preflight wait and order the worktree-startup write

The awaited trust write had no cap. The local Codex writer queues on the
per-config.toml lane it shares with hook installs and app-server trust grants,
and the SSH writer chains a resolveHome round trip plus SFTP calls over a link
that may be half-open - so "start agent" could hang with no error. Cap the wait
at a single deadline and continue untrusted, which only loses the ordering
optimisation: the agent then raises its own trust prompt, so giving up fails
closed.

Also await the discarded write in worktree startup, where the PTY spawns Codex
on the next line and the synchronous catch could not see its rejection.

Update the reliability gate: its oracle asserted the absence of a deadline, and
its manifest was left unformatted.

* fix(reliability-gates): record the trust-preflight counts the cited command actually reports

The gate's evidence still described the pre-deadline suite: "32 author tests",
"New11pass", "original4fail/7pass". The four deadline tests this branch adds make
the cited command run 15 tests in `agent-trust-completion.unit.test.ts` and 36
across the four suites, so the manifest asserted counts its own command no longer
produces. Re-ran the command and recorded what it printed.

Re-ran the red/green as well, against the same 15-test suite rather than the 11 it
was first measured on. Pre-await: 5 fail/10 pass. Unbounded-await: 3 fail/12 pass,
where the fourth deadline test — the already-complete write settling on microtasks
— passes unbounded too, so it is a timer control and not a red; saying so beats
counting it as evidence for the cap.

Two claims the commit left stale: only the IPC handler is under test, yet the same
commit also bounds the worktree-startup write, and the cap is per call site while
three other awaited Codex trust writes are still unbounded. Both are now gaps
instead of silence.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 00:47:18 -07:00
Neil bc78acc43e fix(editor): detect all bundled Monaco language associations (#23371) 2026-09-27 00:32:55 -07:00
Neil 268f0e9e8d fix(projects): enrich git remote identity for runtime-addressed repo rows (#23300)
A repo row stamped `executionHostId: runtime:<env>` is how a paired client
addresses a project registered in *this* process, so its files are local
(`runtime-repository-registration-controller.ts`, and the helper #23184 added in
`repo-execution-host.ts`). #22421 reclassified those rows as peer-owned and
dropped them from the identity sweep, so `gitRemoteIdentity` never settled on
remote-runtime and paired-web setups: the project stayed pending, fell back to a
host-local `repo:<id>` that never grouped across hosts, and a stale automatic
avatar never repaired.

Reuse `getStoredRepoExecutionHostId` instead of a second host classifier, and
keep skipping only a `runtime:` row that also names a nested SSH target — that
target lives in the peer's dispatch table and the same path here is a different
checkout. The store matches a write against the row's own stamp, so address
`updateRepo` by that stamp rather than the probe host, which is `local` for these
rows and would match no row at all.
2026-09-27 00:31:46 -07:00
NeilandClaude 8da0ca52e7 perf(relay): release rejected first-frame connections [trade-off] (#23011)
* perf(relay): release rejected first-frame connections [trade-off]

* fix(relay): bound the director's rejected and redirected first-frame closes too

The parent PR routed four cell-side first-frame rejections through
closeRelayWebSocket but left two raw socket.close() calls in the same
handler. A real-socket probe shows both still pin a connection unit for
ws's full 30s close timer when the peer ignores the close frame:

- 'invalid invite' is reachable by an unauthenticated peer with a
  well-formed but bogus credential, so the exhaustion the parent PR
  claims to prevent stayed reachable on the director;
- 'connect to assigned cell' is the happy path for every phone's first
  director contact, so it is the highest-volume unbounded close here.

closeWithDrain gets the same treatment; host-session-registry already
closes the identical drain through the helper.

Also records that the bounded close is not a user-facing trade-off: the
close frame is written before the force-close timer can fire and TCP
delivers it ahead of the FIN, so an abandoned peer still reads code and
reason over a graceful close. The new regression asserts that, plus
exactly-once release across concurrent bursts and rejection racing the
peer's own disconnect (the ledger does not clamp at zero, so a double
release would surface as a negative count).

* refactor(relay): drop the unused closeWithDrain helper

`closeWithDrain` has no callers anywhere in the repo, and its `graceMs`
parameter promised a caller-supplied drain window that the body no longer
honours: routing it through `closeRelayWebSocket` force-terminates after 1s
regardless, so a future caller passing `graceMs: 30_000` would have had its
drain silently cut short while the signature still claimed otherwise.

The real drain path is `host-session-registry`, which sends the same
`resolve-director` drain and closes it there. Delete the dead duplicate
rather than bound a helper whose contract says "graceful".

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(relay): call the bounded close's rejection delivery best effort, not guaranteed

The helper claimed the forced terminate costs an abandoned peer nothing it can
observe, and the test presented its fast-peer assertion as proof. Neither holds in
general: `ws` writes the close frame to the socket and `terminate()` destroys that
socket a second later, so under backpressure the frame — and any `relay-moved`
message queued ahead of it — can go unsent even to a peer that never stopped
reading.

Qualifies both comments to describe delivery as best effort and name the
backpressure case. The fast-peer assertion is valid and stays exactly as it was;
only its stated scope narrows, and the test is renamed to say which peer it speaks
for. What the bound actually buys — a stalled peer cannot hold admission — is now
stated on its own rather than resting on a delivery claim.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 00:28:52 -07:00
NeilandClaude 8e6c07178e perf: skip store notifications for unchanged project and folder catalogs (#23104)
* perf: skip store notifications for unchanged project and folder catalogs

* perf(store): stop the all-host folder refresh from republishing equal restored owners

The catalog gate landed in this branch suppressed the two catalog publications of
an unchanged all-host folder refresh but left the trailing restored-session-owner
cleanup, which rebuilt `restoredRuntimeHostIdByWorkspaceSessionKey` into a fresh
object on every refresh and so always replaced the store root. Reference-equality
readers (the live dashboard selector, the popout bridge, the tab-create entry gate)
re-ran on that fresh-but-equal identity, so an unchanged all-host refresh still cost
a full publication: 3 -> 1 rather than 3 -> 0.

Reuse `reuseEqualRecordMap` to keep the previous record when the cleanup produces an
equal one, and return the current state when it does. A cleanup that really retires
an owner still publishes.

Also seed the session writer from the current state at creation. `prev === null` is
what bootstraps its first full write, so a writer created when the session gate was
already open owed that write to whatever unrelated store tick arrived next. With
equal catalogs no longer publishing, that incidental wake-up is no longer guaranteed;
evaluating once at creation matches what editor-autosave-controller already does.

* test(store): pin the session writer's creation-time seed

The seed this branch added is the compensating fix for a real regression — it
removed the equal-catalog tick that used to boot the writer's first full write —
but nothing held it in place. Every existing subscriber case creates the writer
with the gate closed and opens it afterwards, so the opening `setState` is itself
the tick that produces that first write; deleting the seed left all five suites
green.

Cover the case the seed exists for: open `workspaceSessionReady` and
`hydrationSucceeded` *before* creating the subscriber, then assert `persist`
fires with a store-tick spy proving nothing woke it. Verified it fails
(`persist` called 0 times) with the seed line removed and is the only failure.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 00:28:39 -07:00
NeilandClaude d301658208 fix: exclude canonical SSH roots from local file access (#23196)
* fix: exclude canonical SSH roots from local file access

* fix(filesystem-auth): fail closed on an unplaceable host stamp and an unanchored workspace dir

Two gaps in the canonical-SSH exclusion:

- `getSshTargetIdForExecutionHost` answers null for every host id the shared
  parser rejects (`ssh:`, `ssh:%zz`, `ssh:host|alias`, `SSH:host`, an unknown
  prefix), so those rows read as local and their remote paths were added to the
  local allow-list. Classify by the parsed host instead: remote unless the stamp
  is absent, `local`, or `runtime:`.
- Filtering more repos out makes the `localRepos.length === 0` fallback reachable
  in configurations where it was not, and that branch `resolve`s a repo-relative
  `workspaceDir` against the main-process cwd — an unrelated tree. Only grant it
  when the configured value is absolute on its own path flavor.

Also name each denied fixture case instead of asserting a count, and cover the
local-path-resembling-an-SSH-root and workspace-dir fallback deltas.

* fix(filesystem-auth): share the unplaceable-host denial with worktree-root registration

Main's worktree-root owner rework landed its own local-repo filter built on
`getSshTargetIdForExecutionHost`, which answers null for a stamp the parser
rejects — the same fail-open shape this branch closed on the repo-path side.
Move the hardened predicate into `remote-filesystem-owner.ts` and use it for
both, so `ssh:`, `ssh:%zz`, `ssh:host|alias`, `SSH:host` and `relay:host` can
no longer register a linked worktree root for local access.

Also updates the fixture matrix for main's two-argument
`isRegisteredWorktreePath`, and widens the fixture stamp type past
`Repo['executionHostId']` so the matrix can build the malformed stamps a
persisted catalog actually carries.

* fix(filesystem-auth): decide an unanchored workspaceDir with the host's own path predicate

`settings:set` takes `workspaceDir` unvalidated, and the relative-fallback guard
used `isRuntimePathAbsolute`, which is syntax-based and accepts either flavour.
On POSIX it calls `C:\workspaces` absolute while the imported `node:path.resolve`
reads the same string as a relative name, so the allowed root landed at
`<main-process cwd>/C:\workspaces` and `isPathAllowed` then authorized every
descendant of that unintended local tree. A UNC-shaped value did the same.

`resolveUnanchoredWorkspaceRoot` now pairs the predicate with the resolver in one
place, so whichever `node:path` runs decides absoluteness and produces the root.
Genuine Windows drive and UNC values still grant on Windows, and POSIX absolute
values still grant on POSIX.

Tests inject `path.posix` and `path.win32` to cover all four flavour
combinations without branching on `process.platform`, plus a POSIX-guarded
end-to-end case proving the foreign-flavour string no longer reaches `resolve`.

Also records why folder-workspace authorization stays stricter than
`resolveFolderWorkspaceHost`: there the workspace's own `executionHostId` pin
wins, so a workspace pinned `local` under a group carrying only a legacy
`connectionId` dispatches locally but is denied here. Agreeing would grant a root
the store refuses today, so the fail-closed read stays and a test pins both
answers, including the mirrored row where the two already agree.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 00:28:26 -07:00
OrcaWinandm4air 25c3ac400b ci: overlap shell setup, localization extraction, and mobile route preparation (#23368)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 00:28:02 -07:00
Neil fbeaec6b69 Avoid store updates for unchanged SSH repository catalogs (#23034)
* Avoid store updates for unchanged SSH repository catalogs

* test(ssh): pin that the catalog gate fails closed on a Repo field it never enumerated

The gate's safety rests on `structuralValuesEqual` walking own keys generically rather
than on a hand-written per-field comparison, so a `Repo` field added after this landed
is compared without anyone revisiting the gate. Nothing asserted that at the gate
itself — only in the comparator's own unit tests — which is exactly the shape of
regression that would silently stop an update from reaching the sidebar.

Cover the two shapes that matter: an optional scalar key no previous row carried, and
a nested record no previous row carried. `issueSourcePreference: 'auto'` is the sharpest
case, since it is documented as semantically identical to the absent key yet must still
publish under the strict own-key policy.

* test(ssh): wait across decoder turns for the rollback cancel request

The SSH create-operation rollback test polled for `pty.cancelDelivery` with
`await Promise.resolve()`, so it could only observe work that landed in
microtasks. FrameDecoder stops decoding after FRAME_DECODER_MAX_TURN_MS and
finishes the fed bytes from `setImmediate`, and `feed()` will not drain while
that continuation is pending — so on a loaded runner the delivered
`pty.shutdown` response is decoded a macrotask later and the request the test
waits for cannot appear inside its budget. It failed exactly that way on
`tests node 24 8/8`.

Poll across macrotasks instead, matching `waitForRequestCount` in
`ssh-git-provider-test-harness`, and wait for the replacement's source data the
same way. Same assertions; verified by forcing the decoder to yield on every
frame (an ever-advancing `Date.now`), which reproduces the CI error before the
change and passes after it.
2026-09-27 00:27:43 -07:00
AnaandNeil 90bae01db9 fix(editor): highlight Solidity files (#20928)
Map .sol to Monaco's built-in 'sol' language id. The grammar already
ships with monaco-editor; only the extension lookup was missing, so
.sol fell through to plaintext.

Closes #13835

Co-authored-by: Neil <neil@stably.ai>
2026-09-27 00:10:40 -07:00
Neil fc69ec11c6 perf(github): coalesce stronger refreshes after pending requests (#22970)
* perf(github): coalesce stronger refreshes after pending requests

* fix(github): bound the refresh-upgrade wait so a strict caller cannot starve

The upgrade loop retried forever: a caller wanting a stronger refresh waited
for each weaker in-flight request, rechecked, and waited again. A repeating
weaker refresh (the quiet-refresh interval) could therefore pin a forced
noCache caller for the life of the process with no timeout or escape hatch.

Wait out at most one weaker request — enough for peers to share the upgrade —
then issue our own. Call counts are unchanged; progress is now guaranteed.

Retargets the coordination test at that invariant instead of asserting that
strict callers block until the weaker replacement finishes.

* fix(github): let only a dedupe key's current request write its cache

Bounding the upgrade wait fixed the starvation but opened a window the
unbounded loop never had: a stronger request can now run beside a weaker one
for the same key. Nothing fenced the cache writes, so whichever settled last
won. A force-only work-item request does not pass noCache, so gh's own cache
can answer it; settling after the noCache request buried the fresher rows
under a new fetchedAt and isFresh then served them for the rest of the TTL.
Checks rewound run state the same way, and a superseded project request could
stamp its failure over a newer table at the known view key.

Stamp each request with a monotonic id on its inflight entry and recheck
ownership after the provider call, immediately before every cache write — the
work-items entry, both project-view branches, and checksCache plus the PR
status syncPRChecksStatus derives from it. That is the same ownership question
the cleanup guard already asked, so the cleanup now reads the stamp too; the
promise itself cannot be compared from inside its own initializer.

The wait stays bounded and the twenty-one-caller upgrade still collapses to two
provider calls.
2026-09-26 23:57:43 -07:00
c3ff93fd70 fix(editor): highlight Twig templates in files and diffs (#23366)
Select Monaco's bundled Twig language for .twig files, including compound template names. Cover case variants, Windows and UNC paths, and misleading suffixes without changing the recently merged Typst mapping.

Adapted from robbdavis's proposal #22357.

Co-authored-by: Robb Davis <robb@affinitybridge.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-26 23:55:36 -07:00
OrcaWinandm4air d8e2a694f6 ci: overlap package preparation and security scans; share localization parsing (#23364)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 23:53:19 -07:00
487bad2596 fix(homebrew): remove deprecated URL verification parameter (#23365)
Remove the deprecated verified URL option from stable and RC Homebrew cask templates. Cross-reviewed against the identical original fix in #18848.

Fixes #18849
Fixes #19580

Co-authored-by: Xavier Xiqués <xavier.xiques@gmail.com>
Co-authored-by: Abdul Munim <423078+munim@users.noreply.github.com>
2026-09-26 23:51:02 -07:00
NeilandClaude 5da18e154f fix: preserve newer hosted review lookup ownership (#23126)
* fix: preserve newer hosted review lookup ownership

* refactor: reuse the shared lookup generation sequence

Drops this coordinator's private counter for the identical allocator added on
the pull-request branch, so the "never reuse a generation id" invariant lives in
one place. The file is byte-identical on both branches, so either may merge
first.

* docs(store): describe the lookup generation sequence by its contract, not its callers

The JSDoc claimed the sequence was "shared by the pull-request and hosted-review
request coordinators", but the pull-request call site arrives in a sibling
change, so on this branch alone the comment named a caller that does not exist.

Describe what the module provides instead: one process-lifetime monotonic
allocator for lookup-ownership stamps, safe to share across every cache because
callers compare stamps for equality only, never order or magnitude. That stays
accurate whether one coordinator draws from it or several, so it needs no edit
when the second call site lands.

Applied identically on the sibling branch so the file stays byte-identical and
the two add/add introductions keep merging without conflict.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-26 23:37:45 -07:00
NeilandClaude 6ccfc1246c fix: preserve newer pull request lookup ownership (#23119)
* fix: preserve newer pull request lookup ownership

* refactor: share one lookup generation sequence between review coordinators

The sibling hosted-review fix needs the same never-reused generation id, so
move the allocator into src/renderer/src/store/lookup-generation-sequence.ts
instead of keeping a second private counter per coordinator.

Also assert in the lifetime test that a stale lookup cannot publish into
prCache and that the live lookup's answer is what lands there.

* docs(store): describe the lookup generation sequence by its contract, not its callers

The JSDoc claimed the sequence was "shared by the pull-request and hosted-review
request coordinators", but the hosted-review call site arrives in a sibling
change, so on this branch alone the comment named a caller that does not exist.

Describe what the module provides instead: one process-lifetime monotonic
allocator for lookup-ownership stamps, safe to share across every cache because
callers compare stamps for equality only, never order or magnitude. That stays
accurate whether one coordinator draws from it or several, so it needs no edit
when the second call site lands.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-26 23:37:28 -07:00
Neil 172083f304 fix: remove PDF export temp files after setup failures (#23130)
* fix: remove PDF export temp files after setup failures

* fix(export): keep PDF temp cleanup to files this export created, and time out a hung load

Two holes in the temp-document cleanup:

- The write was not an exclusive create, so anything already sitting at the
  generated temp path (a symlink planted in the shared temp dir) would be
  written through, and the cleanup would then unlink an entry this export did
  not create. The write now uses `flag: 'wx'`, and the one failure that means
  "this path is not ours" (EEXIST) skips cleanup entirely.
- The 60s timeout only covered render-and-print, started after the load had
  already finished. An export document whose script never yields fires neither
  `did-finish-load` nor `did-fail-load`, so the load await hung forever and the
  hidden window plus its temp file leaked for the life of the app. The timer now
  starts before `loadFile` and races the whole load-render-print sequence.

Tests cover the preserved foreign file, a never-settling load, a load that
resolves without either event, `did-fail-load`, and that one export's cleanup
cannot touch a concurrent export's in-flight temp file.
2026-09-26 23:19:45 -07:00
OrcaWinandm4air ccd1e87287 Overlap independent CI checks with native Actions background steps (#23351)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 23:09:40 -07:00
Neil 7ac9c8d675 perf: skip macOS DNS probes for unrelated errors (#23121)
* perf: skip macOS DNS probes for unrelated errors

* review(dns-probe-admission): single-source the probe gate and widen resolution-failure coverage

The admission guard duplicated the hint predicate at a second call site, so a gate
that drifted stricter than the hint test would silently drop the DNS diagnostic for
a real lookup failure. Route both through isMacTailscaleDnsHintCandidate, evaluated
once per call, and add a parity test asserting admission never changes the message
the ungated hint decision produces.

Also: widen the predicate to the resolution-failure wordings it missed (EAI_NONAME /
EAI_FAIL / ENODATA / getaddrinfo / "could not resolve" / "name or service not known" /
ERR_NAME_RESOLUTION_FAILED), gate the probe on process.platform === 'darwin' explicitly
rather than relying on the reader's internal check, make the hint idempotent so a
re-wrapped hinted message neither duplicates copy nor re-probes, and rewrite the
cache-window test to assert the resolver state the next relevant error reports instead
of pinning an exact probe count.
2026-09-26 23:07:24 -07:00