Commit Graph
584 Commits
Author SHA1 Message Date
Kelvin Amoabaandmmarabel 1168e0f8c8 fix(ssh): let a placed worktree seed while its host is in conflict (#23213)
One host tab on a folder this client lacks marked every worktree on the host unverifiable, so reopening an emptied one never got a terminal.

Fixes #22015

Co-authored-by: mmarabel <166927047+mmarabel@users.noreply.github.com>
2026-10-05 16:22:17 -07:00
Jinwoo Hong f1ed355d06 feat(relay): drain pace window as a reviewed same-cap input, with drain-aware 503 gates (#25639)
* feat(relay): drain pace window as a reviewed same-cap input, with drain-aware 503 gates

The same-cap roll drained every cell over a fixed 300 s window, so a US roll
re-placed hosts at ~2/s and spent ~10 minutes draining and waiting for quiet.
The window is now a dispatch input from a closed set (300000, 60000, 30000),
defaulting to today's 300000.

- Below the default is refused for anything but US general cells; Asia drains
  are bound by their targets' accept rate, and migration-only cells hold no hosts.
- A non-default window must be named in the confirmation, the canary authority
  records it (v2), and a batch may run at its canary's window or slower only.
- Each cell job re-checks the window, scales the restart-safe timeout with it
  (15-min lease + window, unchanged at the default), and records what the cell
  applied and when it settled.
- The report-only shadow gate takes the director's drain-return deferrals out
  of the 503 count (window and baselines), adds the rung's 5-min non-drain 503
  budget and a Retry-After check, and reports the measured re-placement rate.
- relay-workflows.md documents the pace ladder and what each rung records.

* fix(relay): judge paced drains on counted 503s against the pre-drain minutes, and seal the canary's pace verdict

Review of #25639 replayed the shadow gate: it read 10-01 c29 as unverified (the
log read stopped at 20k entries), false-blocked 10-02 c22, and was blind on four
cells whose 24 h/48 h baseline held an incident.

- Director 503s now come from Cloud Run's request_count, aligned per minute by
  Cloud Monitoring, so volume cannot truncate the count.
- The background is the median of the 10 same-day minutes before the drain;
  the 24 h/48 h baselines are gone.
- Scheduled 503s come out: drain-return deferrals and sticky/placement answers
  to a host's own early retry (host-rate-limited, host-in-flight), each split
  across the minutes its 30 s sample covers.
- The rung budget counts only sustained breaches: two straight minutes over
  max(1.5x, +20) warn, over max(2x, +40) would-block.
- The report carries a paceVerdict over the three pace checks. seal_canary
  downloads the canary cell's report and seals that verdict; a batch below the
  default pace needs PASS, from a report on the same cell that drained at that
  pace.
- Docs: the step-down rule reads paceVerdict, 30 s waits for the lane service
  time (#25645), and the staging step is dropped since staging drains unpaced.

Replayed read-only: 10-01 c29 would-block (9 minutes over 41.5/min); all nine
10-02 cells and 10-01 c25 paceVerdict PASS.

* fix(relay): a partial count already past a block line blocks, in the shadow gate's Cloud SQL and pool checks

A truncated FATAL count is a floor, and one runtime sample over the SQL-failure
line is a fact, so neither waits for a complete read. The waiter-run rule still
needs a complete run, since holes can join two runs into one.

* fix(relay): a canary pace PASS needs a real cohort and whole telemetry; one median-based 503 check

From the final review of #25639:
- canaryPaceVerdict seals PASS only from a report that drained at least 400
  hosts (about half a 10-02 US cell), so a near-empty canary cannot authorize
  a fast batch.
- A director-metrics sub-window with fewer samples than one instance emits
  is unverified, so an empty or short Logging answer is never a calm drain.
- seal_canary names the shadow artifact, report path and cell from the gate's
  normalized cell list, as cell_1 uploads it.
- director503 folds into nonDrain503Budget as a single-minute spike rule,
  max(10x median, 200), dropping the pre-drain peak statistic. All 30 replayed
  windows keep their verdicts.
2026-10-05 17:31:29 -04:00
Jinwoo Hong b94cfa7fd0 feat(relay): declare Asia spare cell c34 as migration-only (#25336)
Adds a sixth asia-east2 cell at the C31 shape (cap 3000, 6000 request
units, pool 16, e2-standard-4) in asia-east2-c, pinned to the f30b5cb1
cell image. It gets its own topology and registration wave but no
promotion wave, so the admission script and workflow refuse to promote
it; it stays a migration-only landing zone and out of the fleet pool list.

Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041
2026-10-05 00:26:33 -04:00
Neil 1bec53ceb2 Reduce repeated CI setup and overlap mobile typechecks (#25359)
* Measure remaining CI import, diagnostic and checkout savings

* Qualify remaining CI candidates on hosted runners

* Qualify independent mobile typecheck overlap on Actions

* Keep explicit RPC test registries from loading unused methods

* Qualify complete RPC registry cohort and mobile cancellation

* Promote measured CI setup and typecheck savings

* Recognize the shared RPC test guard in lint policy

* Align the mobile barrier contract with independent typechecks
2026-10-04 19:58:35 -07:00
822fc5bed4 Add repository OpenCode permission defaults (#25326)
* test(config): reproduce rejected repository OpenCode config

* Add repository OpenCode permissions and allow its reviewed root config

* fix: preserve sensitive OpenCode confirmation prompts

---------

Co-authored-by: Orca campaign recovery <campaign-recovery@example.invalid>
Co-authored-by: Orca OpenCode Campaign <opencode-campaign@local.invalid>
2026-10-04 19:34:47 -07:00
keiandsetodeve cecb62158a fix(ui): restore IME Enter protection in workspace details (#24099)
Restore IME Enter protection in workspace details by reusing the existing composition tracker. Reset Notes ownership at textarea detachment and preserve sizing behavior. Repair isolated native test-window delivery without changing the production foreground policy or original native input assertions.

Fixes #24097

Related contributor history: #10711, #11067, #13128, #13282.
Original implementation and macOS recordings: @setodeve, commit b30f095.
Verified on required stock Linux X11/Wayland checks and independent frozen-source review.

Co-authored-by: setodeve <keinick11@outlook.com>
2026-10-04 16:21:04 -07:00
PM 75344e5850 docs: align contributor guidance with the PR template (#25034) 2026-10-04 16:19:53 -07:00
Neil 0c761a7610 Admit short required auxiliary checks after PR preflight (#25317) 2026-10-04 16:01:22 -07:00
Neil 8e8efb1947 Reduce avoidable work in PR checks and SSH test setup (#25309) 2026-10-04 15:26:53 -07:00
Neil b32462f246 Replace patched JSON parser with stream-json (#25202)
* Replace patched JSON parser with stream-json

* Isolate dependencies for historical server compatibility builds
2026-10-04 13:05:30 -07:00
Neil d77c57022e Verify shared preflight selection and record full unit timings (#25239)
* Strengthen shared preflight contracts and record unit timing results

* Record rejected shard-weight holdouts
2026-10-04 06:24:30 -07:00
d9173ffbdb Keep Orca CLI first after shell startup (#25130)
* Restore the owning Orca CLI path after shell profiles

* Use a literal marker for the Bash lookup regression

* Preserve plain panes and initialize zsh after prompt hook replacement

* Preserve user line-editor dispatchers during deferred startup

* fix: retain CLI startup when global Zsh replaces prompt hooks

* test: replay global Zsh hook replacement after host startup

* test: isolate controlled Zsh widgets from distro keyboard setup

* fix(shell): preserve user hooks during deferred zsh initialization

* Keep completed Zsh startup hooks retired when the wrapper is sourced again

---------

Co-authored-by: Codex <codex@openai.com>
Co-authored-by: Orca maintenance <orca-maintenance@users.noreply.github.com>
Co-authored-by: Orca campaign <orca-campaign@local.invalid>
2026-10-04 02:55:59 -07:00
Neil c4e8735f45 Share PR preflight setup to reduce runner demand (#25150)
* Share PR static analysis and compiler runner

* Preserve evidence document final newline for concurrent merges

* Keep readiness reuse contracts aligned with the physical preflight gate
2026-10-04 02:05:05 -07:00
Aashish Mahato e5ba5975df Recover working local forge CLIs behind broken PATH launchers
Recover a working local forge CLI when an earlier PATH launcher is broken. Bound executable probes and reuse the verified selection for native operations without replaying authentication or user requests.

Fixes #22975

Co-authored-by: Aashish <145881415+aashish254@users.noreply.github.com>
2026-10-03 20:45:52 -07:00
8cd9751963 fix(updater): keep macOS Orca open when background instances block updates (#24952)
* fix(updater): guard macOS installs against running app instances

* fix(updater): match native app blockers and preserve quit lifecycle

* fix(updater): keep ordinary macOS quit on Squirrel's install-on-exit path

Converting every quit with a staged update into quitAndInstall made Cmd+Q
relaunch Orca, refused the quit when background instances existed, and
hijacked app.relaunch()+app.quit() restart flows (profile switch, admin
restart) into an update install racing the relaunched old app. Only
Update & Restart runs the running-instance preflight now; the
quit-without-install allowance is no longer reachable and is removed.

* fix(updater): preserve quit intent through macOS staging

* test(native-chat): explicitly model legacy published tab ownership

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-03 16:27:02 -07:00
Neil 62451920ed fix(git): preserve SSH review context and worktree ownership (#24945)
* fix(git): preserve SSH arguments and guard background writes

* test(terminal): settle fish fixture startup readiness

* fix(git): preserve bare UNC SSH paths

* fix(git): unescape shell operators in Windows SSH paths

* test(runtime): settle removal writes before fixture cleanup

* test(git): skip optional OpenSSH probe when unavailable

* perf(git): skip equal-tip reads and bound relay discovery

* test(processes): ratchet the removed relay Git spawn

* test(shells): wait for initial zsh output before sending input

* Fix SSH review context and recover Git maintenance cleanup safely

* Keep relay Git compatibility fixtures outside shared client projects

* Fence superseded maintenance and preserve mixed-version session search

* test: model Git child termination and search catalogs

* test: retain catalog authority over history flags
2026-10-03 05:24:35 -07:00
aca2d51e0e fix(jcode): harden Windows hooks and negotiate remote history (#24998)
Redirect the managed Windows payload file into curl instead of starting
pipeline shells, register native Windows delivery coverage, and document
Jcode v0.89.0+ as the upstream launcher requirement for invisible hooks.

Negotiate Jcode history in both directions with mixed-version Orca hosts,
preserving supported search filters and old-client response compatibility.

Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
Co-authored-by: JianJia2018 <39438074+JianJia2018@users.noreply.github.com>
2026-10-03 04:27:17 -07:00
Neil 668d6d45c4 Reuse measured Electron preparation for current Terminal Perf refs (#24968) 2026-10-03 01:39:58 -07:00
Neil f93808dfff Collect test-selection evidence when full unit tests fail (#24955)
* Collect advisory unit-selection evidence from failed full runs

* Trigger checks after retargeting the evidence fix to main
2026-10-03 00:09:15 -07:00
Neil 8f26bfad22 Use the measured pnpm lookup policy automatically in hosted root CI (#24951)
* Select lookup automatically for the measured hosted root-install profile

* Record hosted automatic-mode cold cache publication proof
2026-10-02 22:32:04 -07:00
NeilandOrca Integration Recovery 77ad467ebb fix(ci): prevent concurrent pnpm refresh during mobile typechecks (#24776)
* fix(ci): run mobile typechecks without concurrent dependency refresh

* test(ci): check effective Linux E2E package list

* test(ci): preserve the mobile production compiler barrier

---------

Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
2026-10-02 21:52:36 -07:00
Neil 6e6f651380 Avoid repeated pnpm archive downloads in cache producers (#24927)
* Let measured cache producers keep stores without downloading them

* Check that restore-only callers do not publish a producer path

* Enable the measured producer mode and record hosted comparisons
2026-10-02 20:37:54 -07:00
Neil 328caa2160 fix(git): reduce queries and preserve data across execution hosts (#24602)
* fix(git): reduce queries and preserve data across execution hosts

* fix(ci): exercise pinned Git and serialize mobile dependency entrypoints

* fix(relay): preserve fresh diff retries after hung shared reads

* test(git): wait for fetch barrier before canceling preparation

* fix(i18n): describe index-preserving discard in every locale

* fix(git): retain clone diagnostics and allow WSL policy startup

* test(git): refresh default-base and branch-safety fixtures
2026-10-02 18:05:37 -07:00
Neil 75f8f34ce3 Overlap independent Linux headless runtime builds (#24910) 2026-10-02 17:18:44 -07:00
Neil 1241ce1b49 Skip slower root package-store restores in macOS PR jobs (#24908)
* Skip slower root package-store restores in macOS PR jobs

* Update the companion cache-policy contract
2026-10-02 17:16:49 -07:00
Neil a824fb74ab Reuse the headless detector compiler without installing full dependencies (#24895)
* Reuse the headless detector compiler without full dependency setup

* Keep optional compiler-cache saves from failing cache warming
2026-10-02 16:56:08 -07:00
Neil 58cf72d48e Skip slower root package-store restores in Linux PR jobs (#24896) 2026-10-02 16:09:59 -07:00
Neil add1c55590 Skip slower Windows root package-store restores in CI (#24885)
* Skip slower Windows root package-store restores in CI

* Update reviewed mobile dependency-store cache expression
2026-10-02 15:24:30 -07:00
Neil 75be95fd8c ci: move ARM Mac qualification to macOS 15 (#24760) 2026-10-02 14:37:48 -07:00
Neil 61836f6026 Reduce scheduled CI cache warming to every six hours (#24881)
* Reduce scheduled CI cache warming to every six hours

* Document cache warmer recovery interval and measured tradeoff
2026-10-02 14:35:06 -07:00
Jinwoo Hong d3a406fcbe ci(release): publish after a skipped orcad template (#24882)
* ci(release): publish after a skipped orcad template

#24872 skips orcad-template for tags that predate it, but a skipped ancestor
skips every job that keeps the implicit success(), so publish-release and the
post-release jobs never ran for v1.4.219.

* test: brace-free filter in the orcad downstream contract
2026-10-02 17:32:33 -04:00
Neil cc73c8e1a7 ci: overlap ARM SSH setup and independent observation waits (#24714) 2026-10-02 14:00:49 -07:00
Neil b94c75c4bd Reuse pnpm verification records in Alpine CI (#24817)
* ci: reuse pnpm verification records in Alpine builders

* ci: qualify consumers of the verification restore action

* ci: match Linux verification cache archive paths
2026-10-02 13:54:04 -07:00
Neil ac28e8c85e Skip dependency installation for known headless build inputs (#24716)
* ci: defer headless dependency installation until graph analysis is needed

* docs: align headless CI rollout with platform and cache policy

* test: isolate headless detector output from the parent CI step
2026-10-02 13:53:56 -07:00
Jinwoo Hong d3e592365e ci(release): skip the orcad template for tags that predate it (#24872)
A patch cut from a base older than #24155 has no orcad template source, so
the template job could never pass and every desktop build waited on it.
2026-10-02 16:05:00 -04:00
Jinwoo Hong 564f4d021a feat: live updates for agent state rules (#24387)
Orca downloads a newer agent-state-rules.json from a fixed GitHub release (stable or next channel), validates it like the bundled rules, and applies it without a restart; a local override wins over the download, which wins over the bundled rules. A hand-started workflow from main is the only publisher; merging publishes nothing.
2026-10-02 14:59:24 -04:00
Neil e2f2b707d2 ci: skip installed glibc tools in SSH host qualification (#24733) 2026-10-02 07:17:13 -07:00
Neil 76b1a90ff6 chore(deps): update reviewed dependencies across Orca (#24561)
* chore(deps): update reviewed desktop dependencies and tooling

* chore(deps): update compatible mobile packages and Fastlane

* chore(deps): update cloud transports and enforce release age

* chore(deps): patch documentation dependencies and record review

* chore: remove dependency review reports

* test(linear): smoke-load resolved SDK through CommonJS loader

* fix(deps): keep native rebuilds from reinstalling addon dependencies

* fix(native): invoke installed node-gyp directly for Node rebuilds

* test(cloud): exclude observer probes from row-lock timing budget

* test(mobile): preserve the CSS writer receiver in viewport spy

* test(native): remove obsolete batch-shim fixture exception

* Stream native rebuild output through the process wrapper
2026-10-02 05:05:43 -07:00
Neil ff452661e7 Stop stalled unit jobs after an hour (#24583)
* ci: bound unit jobs to one hour of execution

* docs: keep CI budget notes clear of the headless follow-up

* docs: keep CI deadline evidence in the pull request
2026-10-02 05:01:18 -07:00
Neil 302526411d ci: share bounded apt setup with the E2E native cache job (#24758) 2026-10-02 04:52:50 -07:00
Neil 13ecf051c3 Reuse prepared Windows native builds in SSH CI (#24555)
* ci: reuse qualified Windows server slots for SSH host tests

* ci: reuse prepared relay addons after an exact native cache hit
2026-10-02 03:31:57 -07:00
Neil efbf651c7b Reduce CI setup costs and fixture failures (#24537)
* Let scheduled CI warmers wait and measure WebRTC startup

* Measure a smaller daemon shutdown fixture image

* Counterbalance WebRTC startup and verify retained fixture files

* Record CI fixture measurements and remove temporary pilots

* Clarify fixture build dependency cleanup evidence

* Make coalesced snapshot fixture delivery deterministic

* test: type the PTY write delay observer
2026-10-02 02:46:02 -07:00
Neil 6153fbcfe4 Reduce redundant headless server CI work (#24527)
* ci: avoid unrelated headless server qualification

* ci: skip headless detection for ineligible draft PRs

* ci: preserve cross-host qualification and skip supplied prerequisites

* ci: include Windows server cache validation in change detection
2026-10-02 01:42:41 -07:00
Neil 5a56636f66 Bound E2E package setup and retain cancellation traces (#24617)
* test: align source-control fixtures with current store contracts

* Bound E2E package setup and retain cancelled-job traces
2026-10-02 01:03:48 -07:00
Neil 8ff6296bc7 Speed up serializer checks and keep native caches stable (#24476)
* Reuse serializer oracle cells and isolate native cache policy

* Preserve native cache post-save paths and record hosted oracle gain

* Record native cache reuse and separate cancel-test startup budget
2026-10-01 21:43:56 -07:00
Brennan Benson 444f1952c7 ci: run every cross-version wire test, picked up by folder so new ones can't be skipped (#24499)
* ci(cross-version-wire): run the whole directory so no compatibility test is left out

Three cross-version tests ran in no CI job because the job named its files by hand.
Run the directory instead, ratchet that every file kept out of the unit shards
runs in some PR job, and re-run the job when the modules the newly running
tests guard change.

* test(cross-version): give the orchestration downgrade test its siblings' 120 s budget

* ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it

The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are
left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable
workflows those jobs call.

It also only proved that some step names each excluded file, not that the job runs when the file
changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path
trigger matched neither it, its harness nor its subject, so a PR touching only those ran it
nowhere. The check now asserts a change to each excluded file fires a gating job that names it,
and the shell trigger gains those three paths.

* ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver

A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to
orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the
job whose tests guard exactly those contracts. Also corrects the publish/read direction in the
turn-end comment.

* test(cross-version): state why the orchestration downgrade test needs 120 s

* test(ci): glob the unit tree once for the unit-exclusion coverage checks
2026-10-01 20:53:51 -07:00
Neil f69052e113 Reuse qualified Windows server builds and dependency verification records (#24448) 2026-10-01 16:19:40 -07:00
Brennan Benson 976dc00337 fix(native-chat): Stop's pause is worked out from the chat's history, so a steered message is never re-sent (#24072)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused

Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.

* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered

A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.

* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget

* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it

The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.

* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card

* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows

Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.

One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.

The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.

Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.

* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones

A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.

* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction

The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.

* fix(native-chat): stop creating the unused queue pause table

The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.

* fix(native-chat): a Stop's pause never hides the restart pause

A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.

Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.

* test(native-chat): pin the Stop's no-resend, lift and held-card rules

- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
  again" at one instant, before a queue ignoring the pause re-sends. They
  now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
  whether or not a person's turn lifts it; it now reads the Stop's pause
  before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
  queued before a rewind.

* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller

The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.

* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event

* test(native-chat): pin that Stop and Resume rows never reach apps or count as history

* test(native-chat): only a person's Stop event pauses the queue

* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop

Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.

* test(native-chat): a card held at a starting agent is checked before the Stop's timing

Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.

* test(native-chat): a released build keeps and folds a journal holding Stop events

Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.

* style(native-chat): format the Stop event changes

* test(native-chat): type the released build's exports through one checked helper

* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only

* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade

The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.

Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.

* fix(native-chat): a Stop that stops nothing new writes no Stop event

A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.

It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.

* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop

* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled

* fix(native-chat): any later Stop event ends a person's Stop pause

A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.

An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.

* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed

A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.

A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.

Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.

* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes

A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.

The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.

* refactor(native-chat): one reading of a Stop's turn for its event and its note

A Stop's event and its note each worked out the same two facts on their own:
which turn the Stop is about (the one it named, else the one running), and
whether a named turn is the one the journal shows running. The event decides
before the interrupt; the note and whether the session ends decide after the
provider's answer, so those decisions stay separate, but the facts they read
are now one helper each in structured-agent-session-turn-stop-notes.ts:
structuredAgentSessionStoppedTurnId and
structuredAgentSessionStopNamesTurnNotLive. The event's turn, the note's key,
the session-ending condition, the running-command check and the repeat check
all read them. No behavior change.

Tests: a Stop naming no turn records the running turn on its event, and
rewrites that turn's note as a Stop naming it does.

* refactor(native-chat): a failed-interrupt Stop reads its turn through the shared helper

The new branch that ends a Codex child after a failed interrupt asked
whether the Stop's turn still runs with `turnId ?? liveTurnId`, a third
copy of "the turn a Stop is about". It now reads
structuredAgentSessionStoppedTurnId, the value the note key already uses,
read at the same point before the cancel. No behavior change.

Test: a Codex Stop whose interrupt failed ends the child, holds the card
queued before it with the queue paused, and writes its Stop event before
the turn's end.
2026-10-01 13:21:58 -07:00
Jinwoo HongandClaude Opus 5.5 ae41eb414a fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup (#24284)
* fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup

A `codex` typed into a plain fish tab ran without --no-daemon because only
wrapped fish tabs (startup command / ready marker) got Orca's codex function.

Plain fish spawns now prepend an Orca data dir to XDG_DATA_DIRS and record the
exact prefix in ORCA_FISH_XDG_DATA_DIRS_PREFIX. Fish sources the dir's
fish/vendor_conf.d snippet, which first restores XDG_DATA_DIRS (unset again if it
was unset), erases the marker, drops its dir from fish's derived vendor/function/
completion paths, then defines the shared fish codex function at the first prompt
so the user's config.fish still wins. fish argv is unchanged; wrapped tabs keep
their existing -C path. A local fallback to another shell restores the user's
XDG_DATA_DIRS instead of deleting it.

Bumps the terminal daemon protocol to v39 so new tabs move to a daemon that
injects the env; v38 owners stay attachable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(fish): skip the XDG handoff for -N/--no-config and empty XDG_DATA_DIRS

fish never reads vendor_conf.d under -N/--no-config (also abbreviated or
clustered), so the snippet could not undo the prefix; and the restore cannot
tell an empty XDG_DATA_DIRS from an unset one. Both now launch untouched.
Run the real-fish handoff tests in the shell contracts job, where fish is
required, so they no longer skip in CI.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(fish): compare the unset-restore case against a fish without Orca

Ubuntu runners ship snapd's fish vendor snippet, which sets XDG_DATA_DIRS on
every fish start, so "unset" was never the right oracle there.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(fish): treat an empty XDG_DATA_DIRS like unset so the tab still gets the codex hook

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(pty): put back the user's own launch env on a shell fallback

The primary shell's launch config now records the pre-launch value of each
key it writes. A fallback shell restores those values (unsetting keys that
had none) instead of deleting the keys, which hands back an inherited
XDG_DATA_DIRS after a fish fallback and an inherited ZDOTDIR after a
zsh->bash fallback, with no per-shell special case.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(fish): drop the Node restore twin and simplify the vendor snippet

- Remove restoreFishXdgDataDirs; the generic fallback restore covers it.
- Snippet: read ":$XDG_DATA_DIRS:" directly and filter Orca's vendor dirs
  with one string match per variable.
- Require inheritedXdgDataDirs in both getShellLaunchConfig option shapes.
- Drop the test-only FISH_XDG_DATA_DIRS_HANDOFF_DAEMON_PROTOCOL_VERSION.
- Fix stale fish comments and trim redundant -N launch cases.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(fish): stop scrubbing fish's lookup paths after the handoff

Only XDG_DATA_DIRS is restored, by exact prefix; Orca's dir holds nothing but this snippet, so leaving it on fish's derived paths loads nothing else and drops the glob match.

* docs(fish): drop the comment for the removed vendor-dir cleanup

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:38:53 -04:00
Jinwoo Hong 07dad6739a refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run (#24443)
* refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run

A same-cap wave no longer consumes a 15-minute monitor dry-run and its sealed,
single-use, five-minute-fresh evidence. Each apply wave now samples fleet health
itself right before isolation, with the monitor's evaluator, thresholds, and
tolerances, for a window sized to the cell's host count (3/5/8 min), plus three
lookback rules: no cell container exit in 10 min, no minute over 500 director
503s in 10 min, and director concurrency p99 within the monitor bar over 4 min.

Removes the monitor-run inputs, the gate's consume/authorize steps, the
break-glass override, and the same-cap-only authorization shapes in
relay-monitor-evidence.mjs. The monitor workflow and the rehome enable path are
unchanged.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): bound the pre-drain sample overrun and keep the drain token fresh

Review follow-ups: alternating tolerated readings could hold the sample open
until its step timeout, so cap the overrun at three samples past the window;
record why a read failed; mint a fresh admin ID token for the drain after the
sample; raise the job timeout to 90 min so a long sample cannot cancel the
job past the failsafe.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* feat(relay): exempt the rolled cell and existing-only cells from the pre-drain crash rule

The exit rule counted every relay container exit fleet-wide, so a cell that
crashes every few hours (c25, 12 a week) blocked the very roll that fixes it,
and existing-only legacy cells (c5, 15 a week) blocked rolls they take no part
in. Exits are now grouped by instance, each instance is named by its own newest
runtime-metrics log line, and only exits on general or migration-only cells
other than the target count. An exit no configured cell can be named for trips
the rule; a failed lookup is a failed read. relay-observability.tf joins the
evidence-code set because the rule depends on its filter.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): cover re-asking for an unnamed exiting instance; note the boot-exit risk

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 15:36:17 -04:00