* feat(agent-launch): host-side prompt delivery for agent.launch The host's agent.launch typed any launch prompt into the shell as part of the launch command. A long or multi-line prompt then ran line by line in the shell, and an agent that never showed readiness or crashed at startup had nothing guarding where its text went. agent.launch now carries a prompt on the typed line only when the line stays one line, control-free and at most 512 bytes; otherwise the agent starts clean and the host pastes the prompt once the agent's own ready signal fires (bracketed paste plus its composer marker or a quiet render, read only after the shell's last hand-off, never while the pane's own shell is proven in front), with main's draft-paste bytes and an Enter 50 ms later. Orchestration worker starts wait on tui-idle as before. A replay-safe launch admits and claims its ledger row in one write, Qwen Code gets a second Enter, the desktop and phone share one launch-refusal classifier, and hosts advertise agent.launch.prompt-carry.v1. Split out of #23748, which moves the desktop source-control buttons onto this path. * fix(agent-launch): keep a short-lined multi-line prompt on a local zsh launch line, as main did #24257 moved every multi-line or over-512-byte prompt off the typed launch line and pasted it after readiness. The phone's AI buttons and review notes, whose multi-line prompts main typed whole into zsh, then reached Claude 0.5-3 s later and their RPC reply waited for the paste. The host now names the shell a local macOS or Linux line is typed into, the way the spawn picks it, and a multi-line prompt rides a zsh line when every line is at most 512 bytes and the whole line at most 8 KB. A real-zsh test types such a line through Orca's own ready barrier and startup write, including when a slow user config makes the write land early. Elsewhere the measured unsafe cases keep the paste: bash 3.2 runs multi-line lines piecemeal, fish drops an early multi-line write, and any shell loses a line over 1 KB written early. * test(agent-launch): keep the real-zsh launch-line test out of the Windows lane's gate scan The Windows lane registration check read `const ZSH_PATH = process.platform === 'win32'` (the head of a multi-line ternary) as a Windows-true flag, so `describe.skipIf(!ZSH_PATH)` looked like a Windows-only suite. The file is POSIX-only; the zsh lookup is now a function. * refactor(protocol): move the agent.launch capabilities into their own module Main's protocol-version.ts sits at the 300-line cap, so the prompt-carry capability pushed it over. The four agent.launch capabilities and their doc move to agent-launch-runtime-capability.ts, re-exported by name and spread into RUNTIME_CAPABILITIES at the same position; the advertised lists and every export are unchanged. * refactor(protocol): import the agent.launch capabilities from their own module `export *` from protocol-version left the four names undefined under the mobile recording loader, which resolves a relative import through a Proxy with no own keys, so 37 phone recordings lost agent.launchReplay. Importers now name agent-launch-runtime-capability directly; protocol-version only spreads its list. * refactor(agent-launch): drop the unshipped viewMode field and trusted local caller id Both were inert in step 1 and existed only for step 2. agent.launch will become a public plugin API, so every wire field is permanent once shipped; a top-level viewMode reads as "choose terminal vs chat", which the host decides. Step 2 introduces placement and view intent under a placement object instead. * fix(agent-launch): read Codex's provisional startup from the rule files' hold anchor Main (#24375) moved Codex's provisional-header check into codex.json's provisional_startup hold anchor and deleted codex-terminal-readiness.ts, so the launch readiness hold now asks showsHoldAnchor, as main's own settled check does. * fix(agent-launch): hold rule-file name titles to quiet for a launch, and census the zsh fixture Main (#24375) answers a name-only title from each agent's rule file ahead of the sustained-title lane, so gemini.json's name_title settled a launch readiness wait on the shell's auto-title while Gemini was still booting. A launch now asks quiet of every weak idle verdict, as that lane did. Main's readiness census requires a recorder for every runtime fixture; the zsh prompt recording is a non-agent control. Gemini's synthetic baseline is regenerated for this PR's stated change: a bare gemini title is no longer its rest mark, so name-only rows settle weak, and a fresh working or blocked status is no longer overridden. * fix(agent-launch): paste a launch prompt only when the launched agent is proven in front A launch pasted its prompt unless a shell was proven in the terminal's foreground, so any read that could not prove one let the prompt through. After an agent exited at startup, its shell turned bracketed paste on at the next prompt, readiness fired on it, and the prompt was typed into the shell: - macOS: a pane runs its shell under login, so the process-group fence's root was never the shell's group and never proved it; the cached foreground name could also still name the exited process. - Windows Git Bash and WSL: the shell-alone-in-its-job check never answers. Now one fresh read of the terminal's foreground decides: agent, shell or unknown. Only 'agent' lets a write through (paste, Enter, second Enter, reused panes too); 'shell' still drops a ready signal. A Windows host never proves the agent, so there the launch line carries the prompt at any size, as on main. * test(agent-launch): cover the Windows QA stub, a grok override that exits at once * fix(agent-launch): keep the local socket alive while a prompted launch waits for its agent A launch with a prompt now waits up to 60 s for the terminal agent to be ready before it writes the prompt, and reports not-delivered when the agent never is. The local runtime socket closes a connection idle for 30 s unless the request is a long poll, so a launch whose agent exited at startup lost its reply and the caller saw 'runtime closed the connection' instead of not-delivered. Classify a prompted agent.launch and agent.launchReplay as a long poll, as orchestration.workerStart already is for the same wait. * refactor(agent-launch): narrow the launch params by 'in' instead of a cast * fix(agent-launch): find a launched agent behind a wrapper that leads its process group A tcsh or nu launch line runs the agent from /bin/sh '<script>', and a wrapper script that does not exec its agent does the same: the wrapper leads the terminal's foreground process group and the agent is a member of it. The fresh foreground read names the group's leader, sh, so a prompted launch was refused or pasted late (M4Air tcsh: 2 of 4 not delivered, 2 pasted ~9 s late). Before that read, take the host's process-group observation as positive proof when it names the launched agent among the foreground group's members and is younger than a ready signal's quiet window. It never proves a shell. * fix(agent-launch): judge the foreground-group proof by when its capture began, not how long ps took The age the host stamps on a process-group observation runs from the start of its whole-machine ps, so on a loaded Mac a capture begun after the read was asked for still read as older than 1 s and the proof was dropped. Count an observation whose capture began after the read was asked for, less the window a shared capture is reused across. * test(agent-launch): keep the crash-guard live test out of the Windows lane's gate scan The Windows-lane registration scan read the const assigned from a platform check as a Windows-only gate, though the suite runs everywhere but Windows; find zsh in a function instead, as the real-zsh typed-line test does. Under load the fresh foreground scan can fail to answer, which lets the shell's prompt settle readiness (2 of 4 paired runs). The guard still refuses that write, so assert the refused write, the property that must always hold. * perf(agent-launch): read a local pane's foreground from its own terminal, not the whole process table The foreground read that gates every launch paste ran the daemon's inspectProcess capture and then a fresh scan, each a whole-machine ps; the fresh one also waits for any capture already running before it starts its own. Measured here at load 5: 1.2 s a read (M4Air QA: 3.4-5.0 s, and worker starts 17.6-32 s against main's 9-12 s at load 25-84). On a local macOS or Linux host, take the pane's root pid from the provider's session inventory and run one ps limited to that pane's terminal. Its foreground process group decides: the launched agent or any non-shell member is the agent (a wrapper that did not exec its agent leads the group), a group of shells alone is the shell. Same pane, same verdict: 2.7 ms a read. SSH hosts keep the relay's observation and name. * test(mobile): re-measure the web app's script sweep after agent.launch's capabilities moved out of protocol-version The mobile web bundle check failed at 124 assets against a ceiling of 123. Main already sat exactly on that ceiling: its sweep table read 69 scripts at 16 routes while the tree builds 73, the whole margin of 4. This branch imports the agent.launch capabilities from their own module, so protocol-version is no longer pulled into the root layout and four other routes. That moves which routes share which modules, and the Qoder capability module, imported by protocol-version and the AI-vault resume path, no longer shares an importer set with anything, so it gets a chunk of its own: 74 scripts. The fence says to re-derive the bound rather than raise it, so the sweep is re-measured on this head (every prefix of the sorted route list). The worst route now adds 10 scripts (session), not 9, which moves the pinned shell crossing from 32 to 30 routes; main re-measured on its own lands on the same crossing. * fix(agent-launch): a worker's brief needs its agent found in front, and Grok's start answers on its composer A paired-server worker start whose agent exited at startup typed its brief into the server's shell, which ran it: the idle wait can settle on a shell back at its prompt, and the brief was written with no foreground read. Both worker-start paths now check before each brief write, as a launch prompt is checked: on a host that can find the agent in front it must be there; on one that cannot (Windows) a shell proven in front still refuses, and anything else writes as before. A Grok worker start waited ~10 s more than main: its only rest signal is its bare name, which a launch holds to quiet output, and Grok animates its logo for ten seconds after its composer glyph. A worker start for an agent whose rest signal is its bare name and whose composer draws a marker (Grok, DSH, mimo-code) now also answers on that marker, whichever comes first.
Orca Mobile
React Native companion app for Orca. Monitor worktrees, view terminal output, and send commands from your phone.
Local development uses two processes:
- Orca desktop/Electron from the repo root. This hosts the mobile WebSocket RPC server on port
6768. - Expo Metro from
mobile/. This serves the React Native app on port8081.
Unless a command says otherwise, run mobile app commands from the mobile/ directory.
Prerequisites
- Node.js 24+
- pnpm
- Xcode and/or Android Studio tooling for simulator or device builds
- Expo Go on your phone, or a development client build when native modules are needed
- Phone and desktop on the same LAN when testing a physical phone
Start Desktop Orca
From the repository root:
pnpm install
pnpm dev
Confirm the mobile RPC server is listening:
lsof -nP -iTCP:6768 -sTCP:LISTEN
Restart pnpm dev after changing Electron main-process code. Metro hot reload only applies to the mobile JavaScript bundle.
Start The Mobile App
cd mobile
pnpm install
pnpm start
Scan the Expo QR code with your phone's camera on iOS, or Expo Go on Android.
For a native dev-client build:
pnpm exec expo run:android
pnpm exec expo run:ios
pnpm start --dev-client
Pair With Desktop Orca
- Open Orca desktop.
- Go to Settings > Mobile.
- Scan the pairing QR code from the mobile app.
- Confirm the mobile host endpoint is
ws://<desktop-ip>:6768.
For the Android emulator, use ws://10.0.2.2:6768. For a physical phone, use the desktop LAN IP, for example ws://192.168.0.179:6768.
If the phone has a stale host entry, remove it from the app and pair again.
Development Paths
Android Phone
- Install Expo Go from Google Play
- Run
pnpm start, scan QR with Expo Go - For native modules:
pnpm exec expo run:android - Run with
pnpm start --dev-client
iOS Simulator
- Install Xcode from the App Store
- Run
pnpm start --iosto open in iOS Simulator
Physical Phone Debugging
The phone can be inspected through the connected device tooling:
orca snapshot --json
orca click --element @e3 --json
orca fill --element @e1 --value "ls" --json
orca screenshot --json
Use snapshot first to find the current element refs, then click/fill those refs. After mobile file edits, Metro usually hot reloads automatically, but navigating out of and back into the session screen can be useful because it re-runs terminal.subscribe.
Terminal Streaming Repro Without A Phone
Use this when terminal output does not render on device and you need to split server streaming bugs from WebView/UI bugs:
cd mobile
ORCA_MOBILE_WS_URL=ws://127.0.0.1:6768 pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64>
You can pass a worktree selector as the third argument:
pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "id:<worktreeId>"
pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "path:/absolute/worktree/path"
pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "name:my-worktree"
The expected result includes:
streamSawMarker: true
readSawMarker: true
If this repro fails, debug the desktop runtime/PTY path before the mobile WebView. If it passes but the phone is blank, debug the session screen or TerminalWebView readiness/queueing path.
Terminal Color Repro Without A Phone
Use this when terminal colors disappear after switching tabs. Open a Claude Code terminal and at least one other terminal in the target worktree, then run:
cd mobile
ORCA_MOBILE_WS_URL=ws://127.0.0.1:6768 pnpm exec tsx scripts/repro-terminal-colors.ts \
<deviceToken> <serverPublicKeyB64> "id:<worktreeId>"
The script captures terminal.subscribe snapshots in an A → B → A sequence and writes raw snapshots to mobile/terminal-color-repro/. If the two A snapshots have different sgrColor counts, the desktop snapshot changed during the switch. If they match, the ANSI color data is still present and the bug is in mobile replay/rendering.
Validation
Run these checks before committing mobile terminal changes:
cd mobile
pnpm exec tsc --noEmit
pnpm run check:tests-typecheck
pnpm lint
cd ..
pnpm typecheck:node
tsc --noEmit reads tsconfig.json, which excludes test files so Metro never bundles them.
tsconfig.test.json puts them back, and pnpm run typecheck:tests shows their errors in full.
check:tests-typecheck is the gate over it: a ratchet against tests-typecheck-baseline.txt, the
127 test files that do not typecheck yet. It fails when a file that checks today stops checking,
and when a baseline entry starts checking (prune it with
node scripts/check-tests-typecheck-ratchet.mjs --prune). The list may only shrink.
The same gate censuses the program first: every *.test.ts(x) on disk must be in it, or named in
the script's TESTS_OUTSIDE_PROGRAM with a reason. Without that, a test excluded from
tsconfig.test.json — or a Foo.test.tsx shadowed by a Foo.test.ts beside it, which a wildcard
include drops for the higher-priority extension — would leave the ratchet silently.
Protocol Version Compatibility
Mobile and desktop talk over a versioned protocol. Because mobile updates lag desktop by 24-48h via the App Store, both sides exchange version numbers on status.get so a genuinely incompatible combo can hard-block instead of silently misbehaving.
Constants live in two files (Metro can't resolve outside mobile/):
src/shared/protocol-version.ts—DESKTOP_PROTOCOL_VERSION,MIN_COMPATIBLE_MOBILE_VERSIONmobile/src/transport/protocol-version.ts—MOBILE_PROTOCOL_VERSION,MIN_COMPATIBLE_DESKTOP_VERSION
Today all four are set so evaluateCompat always returns { kind: 'ok' } — nothing blocks. The wire format is in place to flip a switch when needed.
When to bump
Bump DESKTOP_PROTOCOL_VERSION (and the mobile mirror MOBILE_PROTOCOL_VERSION when relevant) for breaking changes:
- Removed RPC method or required parameter that mobile uses
- Changed meaning (units, nullability) of an existing field mobile reads
- Changed encryption, framing, or auth handshake
Do not bump for additive changes:
- New RPC methods
- New optional fields on existing methods
- New event types in
terminal.subscribe
Set MIN_COMPATIBLE_MOBILE_VERSION (kill-switch) when desktop ships a change that requires a minimum mobile version to function safely. Same for MIN_COMPATIBLE_DESKTOP_VERSION from the mobile side.
When a verdict is blocked, mobile/src/components/ProtocolBlockScreen.tsx renders a screen pointing the user at the update that clears it. When mobile is too old it opens the newest release if the installed app's update check knows one, otherwise the App Store (iOS) or GitHub Releases (Android). When desktop is too old it opens GitHub Releases.
To exercise the block screen locally: set MIN_COMPATIBLE_DESKTOP_VERSION = 999 in mobile/src/transport/protocol-version.ts, rebuild, pair to any desktop. Revert before merging.
Mock Server
Develop the mobile app without a running Orca desktop instance:
pnpm mock-server # starts mock WebSocket server on port 6768
Connect from the app using endpoint ws://localhost:6768 and token mock-device-token.
Environment variables
MOCK_NATIVE_CHAT=1— serve the native-chat scenario (one live agent tab, empty transcript, image upload) instead of the default terminal fixtures.MOCK_CHAT_AGENT=omp— withMOCK_NATIVE_CHAT=1, present an OMP tab and four decoded transcript messages, including a tool call and result, instead of the default Claude scenario. It deliberately omitstranscriptPathto exercise legacy-hook readability discovery; current OMP hooks may report a path.MOCK_SERVER_KEY_FILE— persist the server keypair across restarts so a paired device keeps its public-key pin. A missing or invalid file is re-keyed with a warning, which forces a re-pair.
Scenario control files
Read on every request, so behaviour can be flipped mid-session without a restart (a restart would re-key E2EE and force a re-pair). Write the mode into the file, or delete it for the default.
MOCK_SEND_MODE_FILE(defaultorca-mock-send-modein the system temporary directory) —accept(default) accepts the send,errorfails it withmobile_input_floor_unavailable, anything else reports the send as rejected.MOCK_TERMINAL_LIST_MODE_FILE(defaultorca-mock-terminal-list-modein the system temporary directory) —omitreturns an empty terminal list,otherreturns a list that omits the chat handle, anything else lists it.MOCK_TERMINAL_STREAM_MODE_FILE(defaultorca-mock-terminal-stream-modein the system temporary directory) —deadanswers a subscribe withsubscribedthenend(a gone PTY), which is what exercises the rearm bound and terminal prune; anything else streams normally.
Connecting to Real Orca
- Start Orca desktop with WebSocket transport enabled
- In Orca, go to Settings > Mobile and scan the QR code with this app
- The QR encodes the connection endpoint, device token, and TLS fingerprint
Project Structure
mobile/
├── app/ # Expo Router screens (file-based routing)
│ ├── _layout.tsx # Root layout with navigation stack
│ ├── index.tsx # Home screen — paired hosts list
│ └── pair-scan.tsx # QR code scanning screen
├── src/
│ ├── terminal/ # Terminal WebView and xterm bridge
│ └── transport/ # WebSocket RPC client
├── scripts/
│ ├── test-subscribe.ts # Desktop streaming repro without a phone
│ └── mock-server.ts # Standalone mock WebSocket server
└── assets/ # App icons and splash screen