Jinwoo HongandClaude cdfdadf9ea fix(runtime): settle tui-idle on hook state for agents whose hooks cover the whole turn (#24388)
* fix(codex): install Codex's Interrupt hook so an Esc-cancelled turn settles

Codex 0.150+ fires an Interrupt hook when the user presses Esc on an
approval prompt or mid-tool, and nothing else. Orca did not install it, so
the pane stayed blocked/working until the next prompt.

- Add Interrupt to the managed Codex events and label maps, written with
  Codex's 3s cap (a larger value triggers a startup clamp warning).
- Hash the timeout Codex hashes (Interrupt is clamped to [1,3], default 1)
  so self-computed trust matches Codex; pinned against a real 0.159.3 hash.
- Map a root Interrupt to the existing cancelled-turn record
  (markCodexLeadTurnInterrupted), keeping child work in the fold; a
  child-scoped Interrupt is ignored. Relayed rows take the same path.

* test(runtime): add a readiness census pinning every tui-idle verdict

Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.

Refs STA-9098

* test(runtime): pin the census quiet probes to literal windows

A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.

Refs STA-9098

* test(runtime): say which census probe writes runtime state

Refs STA-9098

* refactor(codex): let the hook builder own Codex's per-event timeout

The managed hook's timeout is now Codex's own normalization of the shared
budget, and every installer derives its trust entry from the hook it wrote,
so no installer repeats the Interrupt special case.

Claude-Session: codex-interrupt-hook review

* refactor(codex): route Interrupt through the Stop lead update with an outcome

Interrupt now writes the lead record through the same setCodexMainAgentTurnState
call as Stop, so markCodexLeadTurnInterrupted keeps its original signature.
Drops the child-scoped Interrupt guard: Codex never runs Interrupt hooks for
subagents and its input schema has no agent_id.

Claude-Session: codex-interrupt-hook review

* test(runtime): observe the census through settled panes and caller-visible waits

- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
  of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
  coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
  work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
  directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.

* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix

Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.

* test(runtime): read the census baseline field without Reflect.get

The anti-slop lint rejects Reflect.get on parsed input.

* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files

Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.

The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.

Refs STA-9098

* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes

Refs STA-9098

* fix(runtime): refuse rule patterns that repeat an optional or alternating group

The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.

* refactor(runtime): give agent state rules and text anchors one when/answer shape

Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.

- Cursor's prompt is two anchors answering working and idle; the one-off
  workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
  matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.

* docs: point the readiness evidence docs at the agent state rule files

* refactor(runtime): read Codex, Claude, OpenCode, Pi, OMP and Gemini readiness from rule files

The rule engine gains the regions and answers these agents need, as closed-list entries:
- rule regions `title` (the classified title status) and `text` (one of the file's idle text
  anchors, settled), and a `predicate` form of the screen region for named engine scans;
- `withoutClock: skip` for strong quiet rules a clockless pane must not believe;
- anchors (renamed from textAnchors) gain a `title` region, and `live` and `hold` answers;
- `profile.screenSource` (trusted grid or live screen), and an `unknown-pane` file for panes
  with no known agent.

Codex's header, composer and provisional-startup checks become named predicates referenced
from codex.json; its ready header, header and startup hold become shared text anchors. Native
idle title markers become shared title anchors; name-only title handling becomes each agent's
idle-title rule. The agent-specific branches in terminal-wait-detection.ts and
tui-idle-evidence.ts are deleted, and the "later live prompt cancels a blocker" rule now reads
only rule-file anchors (plus Muse, which moves in part b2).

No behaviour change: the readiness census baselines are untouched and pass.

Refs STA-9098

* test(runtime): cover the rule engine's title, text and predicate regions and the bundled anchors

Refs STA-9098

* fix(runtime): reject a rule file that repeats an anchor or rule id

A text rule names its anchor by id, so a repeated id let a file pass validation and then throw
while compiling. Also states that engineVersion bumps once a version ships; version 1 is still
being defined.

* refactor(runtime): fold the working anchor answer into live

The engine treated an anchor's working and live answers identically: both mark a live prompt
that cancels an earlier blocker and settles nothing. Cursor's busy prompt now answers live, so
anchors have one non-settling prompt answer.

Refs STA-9098

* refactor(runtime): read the shared π title anchor from pi.json alone

Pi and OMP paint the same `π - <session>` rest title, and title anchors apply to every pane,
so one copy covers both.

Refs STA-9098

* refactor(runtime): key every rule file and read the trusted screen from screenSource alone

readsTrustedScreen no longer also asks for a screen rule (every trusted file has one, and the
schema requires screenSource where it matters), so rule-less files need no filter. A rule's
match is a plain boolean, and compileTitleAnchors is module-private.

Refs STA-9098

* test(runtime): pin that a clocked Codex pane takes no other agent's ready text

No test failed when holdsReadyTextToQuiet was removed; this one does.

Refs STA-9098

* fix(agent-hooks): keep an OMP approval wait until omp resolves it

omp posts tool_execution_start a few milliseconds after
tool_approval_requested, while its Approve/Deny select still holds the
human. Both mapped onto the pane row, so the working event overwrote the
blocked one and the pane read as busy for the whole prompt.

A working event now leaves an OMP approval wait in place; only
tool_approval_resolved or a new turn ends it. An ask row is unchanged:
its own tool_execution_end ends it. The test replays the order a live
omp 17 run posted for a denied bash call.

Refs STA-9100

* refactor(runtime): select the fresh hook row on any of a terminal's handles or pane keys

selectFreshExplicitAgentStatus matched one handle and one pane key and
returned only the mapped status. The row selection now takes sets of
handles and pane keys, an optional received-at floor, and returns the
row itself, so a reader can see the main agent's own state. The old
function keeps its signature and result on top of it.

Refs STA-9100

* feat(runtime): let tui-idle read hook state for agents whose hooks cover the whole turn

tui-idle read no hook state. Hook state reached readiness only through
the `<Agent> ready` titles the window writes, so a headless `orca serve`
never saw it (#16095), and Codex settled only once its screen had been
quiet for three seconds.

Rule files gain `profile.hooks: "authoritative" | "identity-only"`,
defaulting to identity-only. Codex (with its Interrupt hook), OpenCode,
OpenCode 2, Pi and OMP are authoritative. For them a fresh hook-store
row decides ahead of every other lane:

- the main agent's turn decides (`mainAgent.state` when published), so a
  subagent's Stop does not end the lead turn: done settles strong,
  working holds, a permission wait never settles;
- the tail's blocked text goes through the existing permission arbiter
  with the turn as its explicit status, so a denied prompt's dialog left
  in the tail no longer blocks a turn the hook says ended;
- the row joins on every pane key and terminal handle the PTY owns.

No row, a stale, restored or other agent's row, a session-start done,
and a row from before a PTY respawn all fall back to today's lanes. That
keeps startup on the screen and text rules: Codex posts SessionStart
only with the first prompt. Claude, Cursor, Gemini and the rest stay
identity-only.

The readiness census has no hook server, so its frames are unchanged.

Refs STA-9100

* docs(agent-status): record readiness as a reader of the hook store

Refs STA-9100

* fix(runtime): ignore a hook done older than the latest input Orca wrote

A finished turn leaves a fresh `done` row. A caller that sends the next
prompt and waits at once could settle on it before the new turn's first
hook arrives, so the wait returned while the agent was starting work.

Orca's own input writes (terminal send, agent prompts, mailbox pointers)
now stamp a per-PTY input clock, and the hook lane reads no `done`
received before it; the pane falls back to the screen and text rules
until the agent reports again. A `working` row is unaffected.

Refs STA-9100

* docs(agent-status): note the input floor on the hook lane's done

Refs STA-9100

* fix(runtime): take the hook lane's input floor from the PTY run's input record

The hook lane ignored a done older than Orca's latest write to the pane, kept in a
new per-PTY map stamped by a wrapper threaded through four write sites. The PTY
run register already sits on both write funnels, so it now records the last
input (launch writes included, terminal replies not) and the lane reads it.
Keys the user types now count too, which closes the restart-in-the-same-shell
gap: typing `codex` to relaunch no longer lets the previous process's done read
ready while the new one boots.

The respawn floor moves from the shared row join into the lane, beside the
input floor; the freshest row predates a floor exactly when every row does.

* test(runtime): drop runtime hook-lane cases the unit suite already proves

Working over a ready title, a permission wait, and an identity-only agent are
decided inside evaluateTuiIdle and covered there; the runtime suite keeps the
wiring: the join, both floors, Pi's own OSC 133 markers and the arbiter.

* fix(runtime): record a PTY's last input even when main adopted it without a spawn commit

A materialized pane re-adopted by the renderer returns before the spawn-commit
site, so it had no run record and its input never moved the hook lane's floor.
The last input now lives beside the run records: any PTY's input counts, and a
new process's commit still clears it.

* fix(runtime): keep a running process's input time when main reattaches or adopts it

A reattach or adoption commit without an incarnation id cleared the PTY's
last-input time, so a prompt sent just before an SSH adoption was forgotten
and the hook lane could accept the previous turn's done as ready. Only a new
process (or a reattach naming a different incarnation) now starts clean; the
first-input fact follows the same rule.

* docs(runtime): say why a lead turn that ended reads ready while a subagent runs

* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail

A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.

* refactor(runtime): state Codex's provisional startup and title anchors as plain rules

The provisional-startup hold becomes a lastOf anchor with an all/none test, so
its TypeScript scan goes. Title anchors drop their status field (every caller
already gates on an idle title), and withoutClock keeps only the value a rule
can set.

* fix(runtime): leave Codex readiness to its title and screen rules

Codex before its Interrupt hook posts nothing for an Esc mid-turn, so its hook
row stays working and a hook-authoritative tui-idle wait hangs until the row
goes stale. Current Codex already settles fast through its ready title.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-02 01:36:30 -04:00
2026-09-26 20:50:46 +00:00
2026-05-04 20:42:03 -07:00
2026-03-16 22:27:51 -07:00
2026-03-28 10:19:14 -07:00

Orca Orca

GitHub stars Total downloads across all releases License: MIT Join the Orca Discord Follow Orca on X Supported platforms: macOS, Windows, and Linux

中文 · 日本語 · 한국어 · Español · Français · Português

The AI Orchestrator for 100x builders.
Run Codex, ClaudeCode, OpenCode or Pi side-by-side — each in its own worktree, tracked in one place.

Download Orca

Orca desktop app running agents in parallel worktrees, with the Orca mobile companion app in the corner

Features

Mobile Companion

Monitor and steer your agents from your phone — get notified when an agent finishes and send follow-ups from anywhere.

iOS App Store · Android APK 0.0.50 · Docs →

Orca desktop with the mobile companion app

Parallel Worktrees

Fan one prompt across five agents, each in its own isolated git worktree — compare the results and merge the winner.

Docs →

Parallel worktree orchestration

Terminal Splits

Ghostty-class terminals with WebGL rendering, infinite splits, and scrollback that survives restarts.

Docs →

Terminal splits

Design Mode

Click any UI element in a real Chromium window to send its HTML, CSS, and a cropped screenshot straight into your agent's prompt.

Docs →

Embedded browser and Design Mode

GitHub & Linear, Native

Browse PRs, issues, and project boards in-app — open a worktree from any task and review without a context switch.

Docs →

GitHub and Linear task workflows in Orca

SSH Worktrees

Run agents on a beefy remote box with full file editing, git, and terminals — auto-reconnect and port forwarding included.

Docs →

Remote worktrees over SSH

Annotate AI Diffs

Drop comments on any diff line and ship them back to the agent — review, edit, and commit without leaving Orca.

Docs →

Annotate AI-generated diffs

Drag Files to Agents

VS Code's editor with autosave everywhere — drag files or images straight into an agent prompt.

Docs →

Drag files and images into an agent prompt

Orca CLI

Agents drive Orca too — script every workflow with orca worktree create, snapshot, click, and fill.

Docs →

Script Orca from the CLI

Also in the box:

  • Quick open — Search across worktrees, files, agents, commands, and repo context without leaving your flow.
  • Account switcher & usage tracking — See Claude and Codex usage and rate-limit resets, and hot-swap accounts without re-logging in.
  • Rich repo previews — Preview Markdown, images, PDFs, and repo docs in the workspace.
  • Computer Use — Let agents operate desktop apps and visible UI when a workflow needs real interaction.
  • Notifications and unread state — Know when an agent finishes or needs attention, then mark threads unread to come back later.
  • And many, many more — we ship daily, so this list is perpetually behind. The changelog is the real feature list.

Supported Agents

Works with any CLI agent — if it runs in a terminal, it runs in Orca.

Claude Code logo Claude Code   Codex logo Codex   Grok logo Grok   Cursor logo Cursor   GitHub Copilot logo GitHub Copilot   Muse logo Muse   DeepSeek Harness logo DeepSeek Harness   ZCode logo ZCode   OpenCode logo OpenCode   MiMo Code logo MiMo Code   Amp logo Amp   OpenClaude logo OpenClaude   Antigravity logo Antigravity   Pi logo Pi   oh-my-pi logo oh-my-pi   Hermes Agent logo Hermes Agent   Devin logo Devin   Goose logo Goose   Auggie logo Auggie   Autohand Code logo Autohand Code   Charm logo Charm   Cline logo Cline   CodeBuddy logo CodeBuddy   Codebuff logo Codebuff   Freebuff logo Freebuff   Command Code logo Command Code   Continue logo Continue   Droid logo Droid   Kilocode logo Kilocode   Kimi logo Kimi   Kiro logo Kiro   Mistral Vibe logo Mistral Vibe   Qwen Code logo Qwen Code   Rovo Dev logo Rovo Dev   + any CLI agent


Install

Desktop — macOS, Windows, Linux

Or via a package manager:

# macOS (Homebrew)
brew install --cask stablyai/orca/orca

# Arch Linux (AUR) — or stably-orca-git to build from source
yay -S stably-orca-bin

Mobile Companion — iOS, Android

Pair with your desktop app to monitor and steer your agents from your phone.


Community & Support

  • Discord: Join the community on Discord.

  • Twitter / X: Follow @orca_build for updates and announcements.

  • WeChat: Scan to join the Orca community WeChat group 11.

    WeChat group 11 QR code for the Orca community
  • Feedback & Ideas: We ship fast. Missing something? Request a new feature.

  • Privacy: See the privacy & telemetry docs for what anonymous usage data Orca collects and how to opt out.

  • Show Support: Star this repo to follow along with our daily ships.


Developing

Want to contribute or run locally? See our CONTRIBUTING.md guide.

The relay that pairs the mobile app with a desktop host is also in this repository under cloud/, with a separate pnpm workspace and setup guide.

Orca contributors

GitHub star history chart for stablyai/orca

Signed Builds

Windows code signing sponored/provided by SignPath.io, certificate by SignPath Foundation.

License

Orca is free and open source under the MIT License.

S
Description
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.
Readme MIT
1.7 GiB
Languages
TypeScript 95.1%
JavaScript 4.2%
Swift 0.2%
HCL 0.1%
CSS 0.1%