* fix(native-chat): repair a structured tab fenced out by a returning publisher
A publication epoch is retired whenever another publisher takes over a worktree, and a
retired epoch is then rejected forever. But a live publisher can return after transient
interlopers - a `removed:` retraction, then a headless rebuild whose version restarts at
1 - and the structured tab publish inherits the worktree's existing epoch rather than
minting one, so it arrives under the blacklisted epoch and is dropped. The chat tab never
reaches the tab bar.
The fence is right to reject the frame: it cannot tell a returning publisher apart from a
delayed frame queued by a dead generation, whose version can outrank the live cursor. So
the drop is no longer final - it schedules one bounded, debounced authoritative
`session.tabs.listAll`, and only that census may revive an epoch, and only the one it
names current. Subscription frames stay fenced exactly as before.
* fix(native-chat): decay the structured tab repair cap and prune its state
The attempt cap latched: three transient RPC failures left `exhausted` set for the
renderer's lifetime, permanently hiding a chat tab behind a single console warning. It
now decays, so a worktree that has been quiet for a minute gets its full budget back.
The repair map was also missing from the sweep that drops publisher cursors for vanished
worktrees, leaking an entry per deleted worktree. Pruning it there required inverting the
repair lane's dependency on the inventory refresh, which is now injected.
---------
Co-authored-by: Merge Sim <sim@local>
The adhoc mac and dev-channel Windows builds vet the requested ref by
mirroring every branch and tag of this repo into a scratch bare repo and
proving the commit is reachable. Both runner disks are case-insensitive,
and the repo now has two branches differing only in casing, so the files
backend refuses the fetch outright — the whole job dies before checkout.
reftable keys refs in a table rather than as file paths, so both refs
store and every ref stays in the reachability set.
* fix(hooks): register the Claude hook script directly on Windows (#18875)
The Windows Claude Code lifecycle hook was registered as
`powershell.exe -NoProfile -EncodedCommand <...>` whose entire decoded payload
was a `Test-Path` and a call to `~/.orca/agent-hooks/claude-hook.cmd`. Every
hook event paid a full PowerShell start-up to reach a script that exits at its
first `ORCA_PANE_KEY` guard, so sessions outside Orca paid it to do nothing.
Register the script path itself instead, with `|| echo {}` for the
neutral-JSON-when-missing contract (#14818). Measured on Windows 11, invoked as
Claude Code invokes it (`printf payload | bash -c -l "<command>"`):
idle (n=12) baseline 177ms | before 471ms | after 213ms
10-way conc (n=40) -- | before 656ms | after 296ms
p95 under load -- | before 696ms | after 337ms
It also drops an interpreter from the chain the hook's timeout kill must tear
down. Killing the hook does not kill its PowerShell grandchild, which still
holds the stdout handle the agent reads to EOF -- measured, EOF arrived 352ms
AFTER the kill, when the orphan exited by itself. msys2 creates children
suspended and resumes them after, so a kill landing in that window strands one
that never exits and EOF never comes; that is the reported frozen session.
The encoded launcher stays as the fallback for profile paths the shells cannot
carry bare (space, `%`, `^`, `&`, non-ASCII) and for hosts where Git Bash is not
resolvable, because PowerShell 5.1 rejects `||`. Every other agent's hook is
untouched, as is the remote/SSH path.
Not adopted from the report: `cmd.exe /d /c <path>` (MSYS rewrites the `/c`
under Git Bash -- measured, the invocation fails), and raising the 10s timeout
(the orphan survives the kill regardless; the fast path puts the hook 30x under
the budget so the kill effectively stops firing).
* fix(build): list the new hook launcher modules in the CLI tsconfig project
config/tsconfig.cli.json enumerates its files explicitly, so the two new
imports reached by src/main/claude/hook-settings.ts failed tc:cli with TS6307.
src/main/git-bash.ts pulls in only node:fs, node:path and a shared constant,
so it adds nothing heavy to the CLI project.
* fix(hooks): address review of the direct Windows Claude hook launcher
- Make the Windows hook suites host-independent. A box with a cmd.exe AutoRun
(HKCU\...\Command Processor\AutoRun) failed them at HEAD too: the tests
redirect USERPROFILE, the AutoRun target vanishes, and MSYS spawns a .cmd
without /d so AutoRun runs and lands on the hook's stderr. Seed an empty
target, including under the deliberately-absent profile.
- Note in managed-hook-stdin-lifecycle why the "missing managed script" case no
longer exercises the fallback for the direct shape (it carries an absolute
path, so a redirected profile changes nothing); that path is covered live in
windows-direct-cmd-hook-command.test.ts.
- Keep the direct shape off UNC profiles: WINDOWS_CMD_SAFE_PATH admits them, but
//server/share/... is not a command cmd.exe reliably starts.
- Correct the comments: `|| echo {}` also fires when cmd.exe itself exits
non-zero (failing AutoRun), printing {} twice. The encoded launcher exited 1
on that same box, so neither shape is clean there.
- Test the contract that replaced runtime %USERPROFILE% resolution (STA-3348): a
stale absolute path reports not_installed and is rewritten on install.
- Record the standing unmeasured assumption in windows-edr-posture.md: `||` does
not parse in Windows PowerShell 5.1, so a compat consumer that hosts hook
strings there would fail closed. Measure before widening to another agent.
- Trim the launcher comments per AGENTS.md; the numbers live in the doc.
* test(win32): register the new Windows-gated hook test in the CI lane
win32-test-lane-registration guards against exactly this: a Windows-gated file
that self-skips on ubuntu and reports success, so it runs on no machine. The new
windows-direct-cmd-hook-command.test.ts needs both entries — WINDOWS_PACKAGE_TESTS
decides whether package_windows runs for a diff, and the workflow argv decides
whether the file runs once that job started.
* test(win32): remove the hook temp tree through the retrying helper
windows-lane-tree-removal-boundary scans exactly the specs in the Windows CI
lane, so registering windows-direct-cmd-hook-command.test.ts subjected it to the
rule: cmd.exe and bash have just exited in that tree, and a raw recursive rm
throws EPERM on Windows while their handles drain, turning a green spec into a
lane failure. Use removeTreeSync, which carries the repo's maxRetries policy.
---------
Co-authored-by: Orca Worker <orca-worker@localhost>
* fix(relay): abandon a client accept once the phone hangs up; jitter the control lease
The accept runs several serialized Postgres calls behind the contended
cell-inventory lock, and phones bound their dial. Finishing that work for a
phone that had already left acquired (and leaked for 90s) an activity lease and
then failed at bind with host_data_reservation_already_bound. Check the client
socket between the DB steps and unwind what was taken, reporting the stage on
orca_relay_client_accept_abandoned.
Jitter the control lease grant so a cohort that reconnected in the same minute
(a cell recreate dumps hundreds at once) walks apart instead of rebinding
together every cycle.
On the phone, treat a probe session that enters 'reconnecting' as a failed
probe: it is the direct client's own backoff after a dead-LAN 1006, and waiting
it out held the supervisor's operation mutex for the full 12s bound.
* perf(relay): lengthen the control lease to 6h
The lease bounds how long a host lingers on a cell after a missed drain, and
rebinding it is the only passive rebalancing we have, so it stays finite. 6h
keeps both properties while cutting control-activation traffic on the contended
cell-inventory lock ~6x. The relay JWT (5 min, refreshed by the desktop) and the
75s silence watchdog are enforced separately, so the longer grant authorizes
nothing extra. The jitter widens with it, to +/-30 min.
* fix(relay): let one flap recover the direct probe; correct the leak window
'reconnecting' is published on any socket close, so rejecting on it outright
turned a single access-point flap into a booked direct failure and a 60s
cooldown. Give the first 'reconnecting' a 2s grace in which a 'connected'
transition still resolves; a dead LAN still fails in ~2s rather than holding the
supervisor's operation mutex for the 12s bound.
The abandoned accept held its activity lease for the 10s attach deadline, not
90s -- the attach timer is armed before bind throws and already unwinds it.
Also cover the assignment-stage check that guards reserveCredential, and drop a
spread assertion the two exact-value assertions above already imply.
* fix(relay): extend the probe grace once on a handshake; pin the lease band top
The redial fires at 500ms but 'connected' waits on the Noise handshake and a
capability RPC, so one 2s window is too tight for real work. A 'handshaking'
transition is evidence the peer answered, so extend the grace once; a stalled
handshake still fails at ~3.5s, far inside the 12s bound.
The longest-lease case only had an upper bound, which a jitter clamped to one
side would satisfy. Pin it to the exact top of the band instead, and assert the
assignment resolve ran so the third-guard test cannot pass vacuously.
* fix(orchestration): typed error codes for dispatch and worker-start refusals
orchestration dispatch (and worker-start, which composes it) surfaced task
not found, task not ready, and inject rejected as the same bare
runtime_error, so an agent reading the receipt could not choose between
creating the task, waiting on dependencies, or picking another terminal.
Add task_not_found (data.taskId), task_not_ready (data.status,
data.unmetDependencies), and inject_rejected (data.terminal, data.reason),
each carrying data.nextSteps so every shipped CLI already prints the
recovery. worker-start's not-ready refusal moves from task_not_startable
to task_not_ready with the same detail. runtime_error stays for genuinely
unexpected failures.
Proven red-first from RpcDispatcher through the CLI's own failure
formatting, plus an SSH bridge test that the host CLI's typed refusal
relays unchanged.
* test(orchestration): load CLI formatter at runtime in the dispatch-code test
The composite node typecheck (config/tsconfig.node.json without
--composite false, as CI runs it) rejects a static import of src/cli from
a main test with TS6307. Load the formatter and error class dynamically
behind narrow structural types, as the CLI/runtime boundary test does.
* fix(orchestration): keep task_not_startable and split the CLI-format proof
Review on #18902:
- Drop task_not_ready. worker-start already published task_not_startable
for a not-ready Task, so renaming it would change an existing receipt
value under old clients. dispatch now emits task_not_startable too (it was
a bare runtime_error before, so this is purely additive), with the new
data.status / data.unmetDependencies / data.nextSteps.
- Move the refusal receipts (code, message, data) into
src/shared/orchestration-dispatch-refusal-contract.ts so the runtime
emits them and the CLI test formats the identical envelope. The RPC test
under src/main asserts toEqual against the contract; the new
src/cli/orchestration-dispatch-refusal-format.test.ts feeds those same
receipts to formatCliError / reportCliError. Neither tsconfig widens and
the composite typecheck CI runs is clean.
* fix(orchestration): keep published refusal messages and type the DB claim guards
Codex review of #18902:
- Every call site keeps the exact message it published on main
("Task not found: <id>", "only a ready Task can start.", "cannot retry
from Dispatch"); the shared contract now takes the message per site and
only owns the code and data. Baseline strings are pinned as literals.
- createDispatchContext's own missing/non-ready guards, including the
atomic-claim loser, now emit the same typed receipt instead of a bare
Error, so a dispatch that races a status change no longer flattens to
runtime_error. Covered by a dispatcher-level race test.
- Invalid --retry-of keeps task_not_startable but now carries status,
unmetDependencies, retryOf, and a retry-specific next step.
- Dependency recovery text distinguishes waiting on running deps from
retrying/unblocking failed ones.
- CLI test adds an unknown-code case so the old-client claim rests on an
assertion, not a comment; SSH test asserts exact stdout.
- Guide table narrowed to the covered preflight cases; occupancy stays
runtime_error and is named as such.
Operator record for the 2026-09-04 relay reconnect incident and the Roll 1
same-cap cell image roll (complete 2026-09-05, selector gen 148), plus the
follow-up checklist, roadmap, and the Roll 2 implementation plan.
Docs only; split out of #18565 so the record merges independently of the code.