One host tab on a folder this client lacks marked every worktree on the host unverifiable, so reopening an emptied one never got a terminal.
Fixes#22015
Co-authored-by: mmarabel <166927047+mmarabel@users.noreply.github.com>
* feat(relay): drain pace window as a reviewed same-cap input, with drain-aware 503 gates
The same-cap roll drained every cell over a fixed 300 s window, so a US roll
re-placed hosts at ~2/s and spent ~10 minutes draining and waiting for quiet.
The window is now a dispatch input from a closed set (300000, 60000, 30000),
defaulting to today's 300000.
- Below the default is refused for anything but US general cells; Asia drains
are bound by their targets' accept rate, and migration-only cells hold no hosts.
- A non-default window must be named in the confirmation, the canary authority
records it (v2), and a batch may run at its canary's window or slower only.
- Each cell job re-checks the window, scales the restart-safe timeout with it
(15-min lease + window, unchanged at the default), and records what the cell
applied and when it settled.
- The report-only shadow gate takes the director's drain-return deferrals out
of the 503 count (window and baselines), adds the rung's 5-min non-drain 503
budget and a Retry-After check, and reports the measured re-placement rate.
- relay-workflows.md documents the pace ladder and what each rung records.
* fix(relay): judge paced drains on counted 503s against the pre-drain minutes, and seal the canary's pace verdict
Review of #25639 replayed the shadow gate: it read 10-01 c29 as unverified (the
log read stopped at 20k entries), false-blocked 10-02 c22, and was blind on four
cells whose 24 h/48 h baseline held an incident.
- Director 503s now come from Cloud Run's request_count, aligned per minute by
Cloud Monitoring, so volume cannot truncate the count.
- The background is the median of the 10 same-day minutes before the drain;
the 24 h/48 h baselines are gone.
- Scheduled 503s come out: drain-return deferrals and sticky/placement answers
to a host's own early retry (host-rate-limited, host-in-flight), each split
across the minutes its 30 s sample covers.
- The rung budget counts only sustained breaches: two straight minutes over
max(1.5x, +20) warn, over max(2x, +40) would-block.
- The report carries a paceVerdict over the three pace checks. seal_canary
downloads the canary cell's report and seals that verdict; a batch below the
default pace needs PASS, from a report on the same cell that drained at that
pace.
- Docs: the step-down rule reads paceVerdict, 30 s waits for the lane service
time (#25645), and the staging step is dropped since staging drains unpaced.
Replayed read-only: 10-01 c29 would-block (9 minutes over 41.5/min); all nine
10-02 cells and 10-01 c25 paceVerdict PASS.
* fix(relay): a partial count already past a block line blocks, in the shadow gate's Cloud SQL and pool checks
A truncated FATAL count is a floor, and one runtime sample over the SQL-failure
line is a fact, so neither waits for a complete read. The waiter-run rule still
needs a complete run, since holes can join two runs into one.
* fix(relay): a canary pace PASS needs a real cohort and whole telemetry; one median-based 503 check
From the final review of #25639:
- canaryPaceVerdict seals PASS only from a report that drained at least 400
hosts (about half a 10-02 US cell), so a near-empty canary cannot authorize
a fast batch.
- A director-metrics sub-window with fewer samples than one instance emits
is unverified, so an empty or short Logging answer is never a calm drain.
- seal_canary names the shadow artifact, report path and cell from the gate's
normalized cell list, as cell_1 uploads it.
- director503 folds into nonDrain503Budget as a single-minute spike rule,
max(10x median, 200), dropping the pre-drain peak statistic. All 30 replayed
windows keep their verdicts.
Adds a sixth asia-east2 cell at the C31 shape (cap 3000, 6000 request
units, pool 16, e2-standard-4) in asia-east2-c, pinned to the f30b5cb1
cell image. It gets its own topology and registration wave but no
promotion wave, so the admission script and workflow refuse to promote
it; it stays a migration-only landing zone and out of the fleet pool list.
Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041
* Measure remaining CI import, diagnostic and checkout savings
* Qualify remaining CI candidates on hosted runners
* Qualify independent mobile typecheck overlap on Actions
* Keep explicit RPC test registries from loading unused methods
* Qualify complete RPC registry cohort and mobile cancellation
* Promote measured CI setup and typecheck savings
* Recognize the shared RPC test guard in lint policy
* Align the mobile barrier contract with independent typechecks
Restore IME Enter protection in workspace details by reusing the existing composition tracker. Reset Notes ownership at textarea detachment and preserve sizing behavior. Repair isolated native test-window delivery without changing the production foreground policy or original native input assertions.
Fixes#24097
Related contributor history: #10711, #11067, #13128, #13282.
Original implementation and macOS recordings: @setodeve, commit b30f095.
Verified on required stock Linux X11/Wayland checks and independent frozen-source review.
Co-authored-by: setodeve <keinick11@outlook.com>
* Restore the owning Orca CLI path after shell profiles
* Use a literal marker for the Bash lookup regression
* Preserve plain panes and initialize zsh after prompt hook replacement
* Preserve user line-editor dispatchers during deferred startup
* fix: retain CLI startup when global Zsh replaces prompt hooks
* test: replay global Zsh hook replacement after host startup
* test: isolate controlled Zsh widgets from distro keyboard setup
* fix(shell): preserve user hooks during deferred zsh initialization
* Keep completed Zsh startup hooks retired when the wrapper is sourced again
---------
Co-authored-by: Codex <codex@openai.com>
Co-authored-by: Orca maintenance <orca-maintenance@users.noreply.github.com>
Co-authored-by: Orca campaign <orca-campaign@local.invalid>
Recover a working local forge CLI when an earlier PATH launcher is broken. Bound executable probes and reuse the verified selection for native operations without replaying authentication or user requests.
Fixes#22975
Co-authored-by: Aashish <145881415+aashish254@users.noreply.github.com>
* fix(updater): guard macOS installs against running app instances
* fix(updater): match native app blockers and preserve quit lifecycle
* fix(updater): keep ordinary macOS quit on Squirrel's install-on-exit path
Converting every quit with a staged update into quitAndInstall made Cmd+Q
relaunch Orca, refused the quit when background instances existed, and
hijacked app.relaunch()+app.quit() restart flows (profile switch, admin
restart) into an update install racing the relaunched old app. Only
Update & Restart runs the running-instance preflight now; the
quit-without-install allowance is no longer reachable and is removed.
* fix(updater): preserve quit intent through macOS staging
* test(native-chat): explicitly model legacy published tab ownership
---------
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Redirect the managed Windows payload file into curl instead of starting
pipeline shells, register native Windows delivery coverage, and document
Jcode v0.89.0+ as the upstream launcher requirement for invisible hooks.
Negotiate Jcode history in both directions with mixed-version Orca hosts,
preserving supported search filters and old-client response compatibility.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
Co-authored-by: JianJia2018 <39438074+JianJia2018@users.noreply.github.com>
* fix(ci): run mobile typechecks without concurrent dependency refresh
* test(ci): check effective Linux E2E package list
* test(ci): preserve the mobile production compiler barrier
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
* Let measured cache producers keep stores without downloading them
* Check that restore-only callers do not publish a producer path
* Enable the measured producer mode and record hosted comparisons
* ci(release): publish after a skipped orcad template
#24872 skips orcad-template for tags that predate it, but a skipped ancestor
skips every job that keeps the implicit success(), so publish-release and the
post-release jobs never ran for v1.4.219.
* test: brace-free filter in the orcad downstream contract
* ci: reuse pnpm verification records in Alpine builders
* ci: qualify consumers of the verification restore action
* ci: match Linux verification cache archive paths
* ci: defer headless dependency installation until graph analysis is needed
* docs: align headless CI rollout with platform and cache policy
* test: isolate headless detector output from the parent CI step
Orca downloads a newer agent-state-rules.json from a fixed GitHub release (stable or next channel), validates it like the bundled rules, and applies it without a restart; a local override wins over the download, which wins over the bundled rules. A hand-started workflow from main is the only publisher; merging publishes nothing.
* ci: bound unit jobs to one hour of execution
* docs: keep CI budget notes clear of the headless follow-up
* docs: keep CI deadline evidence in the pull request
* Let scheduled CI warmers wait and measure WebRTC startup
* Measure a smaller daemon shutdown fixture image
* Counterbalance WebRTC startup and verify retained fixture files
* Record CI fixture measurements and remove temporary pilots
* Clarify fixture build dependency cleanup evidence
* Make coalesced snapshot fixture delivery deterministic
* test: type the PTY write delay observer
* ci: avoid unrelated headless server qualification
* ci: skip headless detection for ineligible draft PRs
* ci: preserve cross-host qualification and skip supplied prerequisites
* ci: include Windows server cache validation in change detection
* Reuse serializer oracle cells and isolate native cache policy
* Preserve native cache post-save paths and record hosted oracle gain
* Record native cache reuse and separate cancel-test startup budget
* ci(cross-version-wire): run the whole directory so no compatibility test is left out
Three cross-version tests ran in no CI job because the job named its files by hand.
Run the directory instead, ratchet that every file kept out of the unit shards
runs in some PR job, and re-run the job when the modules the newly running
tests guard change.
* test(cross-version): give the orchestration downgrade test its siblings' 120 s budget
* ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it
The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are
left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable
workflows those jobs call.
It also only proved that some step names each excluded file, not that the job runs when the file
changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path
trigger matched neither it, its harness nor its subject, so a PR touching only those ran it
nowhere. The check now asserts a change to each excluded file fires a gating job that names it,
and the shell trigger gains those three paths.
* ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver
A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to
orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the
job whose tests guard exactly those contracts. Also corrects the publish/read direction in the
turn-end comment.
* test(cross-version): state why the orchestration downgrade test needs 120 s
* test(ci): glob the unit tree once for the unit-exclusion coverage checks
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused
Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.
* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered
A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.
* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget
* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it
The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.
* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card
* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows
Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.
One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.
The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.
Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.
* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones
A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.
* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction
The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.
* fix(native-chat): stop creating the unused queue pause table
The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.
* fix(native-chat): a Stop's pause never hides the restart pause
A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.
Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.
* test(native-chat): pin the Stop's no-resend, lift and held-card rules
- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
again" at one instant, before a queue ignoring the pause re-sends. They
now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
whether or not a person's turn lifts it; it now reads the Stop's pause
before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
queued before a rewind.
* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller
The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.
* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event
* test(native-chat): pin that Stop and Resume rows never reach apps or count as history
* test(native-chat): only a person's Stop event pauses the queue
* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop
Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.
* test(native-chat): a card held at a starting agent is checked before the Stop's timing
Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.
* test(native-chat): a released build keeps and folds a journal holding Stop events
Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.
* style(native-chat): format the Stop event changes
* test(native-chat): type the released build's exports through one checked helper
* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only
* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade
The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.
Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.
* fix(native-chat): a Stop that stops nothing new writes no Stop event
A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.
It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.
* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop
* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled
* fix(native-chat): any later Stop event ends a person's Stop pause
A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.
An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.
* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed
A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.
A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.
Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.
* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes
A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.
The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.
* refactor(native-chat): one reading of a Stop's turn for its event and its note
A Stop's event and its note each worked out the same two facts on their own:
which turn the Stop is about (the one it named, else the one running), and
whether a named turn is the one the journal shows running. The event decides
before the interrupt; the note and whether the session ends decide after the
provider's answer, so those decisions stay separate, but the facts they read
are now one helper each in structured-agent-session-turn-stop-notes.ts:
structuredAgentSessionStoppedTurnId and
structuredAgentSessionStopNamesTurnNotLive. The event's turn, the note's key,
the session-ending condition, the running-command check and the repeat check
all read them. No behavior change.
Tests: a Stop naming no turn records the running turn on its event, and
rewrites that turn's note as a Stop naming it does.
* refactor(native-chat): a failed-interrupt Stop reads its turn through the shared helper
The new branch that ends a Codex child after a failed interrupt asked
whether the Stop's turn still runs with `turnId ?? liveTurnId`, a third
copy of "the turn a Stop is about". It now reads
structuredAgentSessionStoppedTurnId, the value the note key already uses,
read at the same point before the cancel. No behavior change.
Test: a Codex Stop whose interrupt failed ends the child, holds the card
queued before it with the queue paused, and writes its Stop event before
the turn's end.
* fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup
A `codex` typed into a plain fish tab ran without --no-daemon because only
wrapped fish tabs (startup command / ready marker) got Orca's codex function.
Plain fish spawns now prepend an Orca data dir to XDG_DATA_DIRS and record the
exact prefix in ORCA_FISH_XDG_DATA_DIRS_PREFIX. Fish sources the dir's
fish/vendor_conf.d snippet, which first restores XDG_DATA_DIRS (unset again if it
was unset), erases the marker, drops its dir from fish's derived vendor/function/
completion paths, then defines the shared fish codex function at the first prompt
so the user's config.fish still wins. fish argv is unchanged; wrapped tabs keep
their existing -C path. A local fallback to another shell restores the user's
XDG_DATA_DIRS instead of deleting it.
Bumps the terminal daemon protocol to v39 so new tabs move to a daemon that
injects the env; v38 owners stay attachable.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(fish): skip the XDG handoff for -N/--no-config and empty XDG_DATA_DIRS
fish never reads vendor_conf.d under -N/--no-config (also abbreviated or
clustered), so the snippet could not undo the prefix; and the restore cannot
tell an empty XDG_DATA_DIRS from an unset one. Both now launch untouched.
Run the real-fish handoff tests in the shell contracts job, where fish is
required, so they no longer skip in CI.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(fish): compare the unset-restore case against a fish without Orca
Ubuntu runners ship snapd's fish vendor snippet, which sets XDG_DATA_DIRS on
every fish start, so "unset" was never the right oracle there.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(fish): treat an empty XDG_DATA_DIRS like unset so the tab still gets the codex hook
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(pty): put back the user's own launch env on a shell fallback
The primary shell's launch config now records the pre-launch value of each
key it writes. A fallback shell restores those values (unsetting keys that
had none) instead of deleting the keys, which hands back an inherited
XDG_DATA_DIRS after a fish fallback and an inherited ZDOTDIR after a
zsh->bash fallback, with no per-shell special case.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* refactor(fish): drop the Node restore twin and simplify the vendor snippet
- Remove restoreFishXdgDataDirs; the generic fallback restore covers it.
- Snippet: read ":$XDG_DATA_DIRS:" directly and filter Orca's vendor dirs
with one string match per variable.
- Require inheritedXdgDataDirs in both getShellLaunchConfig option shapes.
- Drop the test-only FISH_XDG_DATA_DIRS_HANDOFF_DAEMON_PROTOCOL_VERSION.
- Fix stale fish comments and trim redundant -N launch cases.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* refactor(fish): stop scrubbing fish's lookup paths after the handoff
Only XDG_DATA_DIRS is restored, by exact prefix; Orca's dir holds nothing but this snippet, so leaving it on fish's derived paths loads nothing else and drops the glob match.
* docs(fish): drop the comment for the removed vendor-dir cleanup
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>