Files
orca/mobile
Brennan Benson 29c7d5d983 fix(mobile): start AI-button agents through agent.launch, never a bare shell (#22762)
* fix(mobile): start AI-button agents through agent.launch, never a bare shell

"Fix checks with AI", "Resolve conflicts with AI", commit-failure recovery and
diff review's "New Agent Session" created a terminal with no agent and typed the
multi-line prompt into the shell, so each line ran as a shell command.

They now call agent.launchReplay into the existing workspace with the prompt;
the host picks chat or terminal from the user's default and delivers the prompt.
Hosts without the launch capabilities get the buttons disabled with update copy.

The agent comes from the desktop's own resolution (moved to src/shared). The
replay loop and capability read are shared with the workspace-create launch.

* test(mobile): repin bridged-parity tallies for the AI-button launch goldens

The corpus goes from 787 to 790 goldens: five shell-path goldens are removed and
eight agent.launch ones added; one lands in identical and two in
result-absent-settlement.

* test(mobile): re-record goldens for AI-button launches through agent.launch

Repinned baseline to 514ab7f868 and re-recorded all goldens. Against the branch
point: 781 header-only (baseline on all; adapterSha256 on the 48 goldens whose
adapter module changed; scenarioSha256 on 3), one body moved
(pr-triage-launch: createTerminal + terminal.send becomes agent.launchReplay),
eight added (the new launch outcomes and their reply matrices) and five deleted
(the shell-path scenarios and their matrices).

* fix(mobile): show review notes' agent launch progress and failures, once

"New Agent Session" left the sheet open with no progress for the whole launch
(up to a minute while a terminal agent readies), so a second tap started a
second agent, and a launch that never started or could not be confirmed
rejected an unobserved promise, showing nothing. The sheet now closes on tap,
the review screen says "Starting an agent...", one launch runs at a time, and
every outcome lands in the review screen's status line.

Marking notes sent now reads the screen state when the launch settles, so a
note written during the wait is not dropped by the whole-list save.

* test(mobile): re-record goldens for review notes' launch outcome on the review screen

Repins the corpus to 6f3018576d. One golden body moves:
review-create-agent-refused now fulfils with "Workspace not found" in the
review screen's status line and the sheet closed, where it previously
rejected an unobserved promise and left the status line empty. The other
789 goldens move only their baseline header.

* test(agent-status): drop the retired PR-triage terminal send from the identity inventory

The phone's AI buttons no longer create a terminal and send the prompt into it
(`createTerminalAndSendPrompt` is gone); the host's agent launch delivers it.
There is no terminal action consumer left in that file to pin.

* fix(runtime): publish saved source-control launch recipes to paired clients

settings.get is an allowlist and omitted sourceControlAi, so the phone never
saw an agent saved globally for "Fix checks", "Resolve conflicts" or commit
recovery and always fell back to the default agent. The host now publishes
the launch actions' recipes (agent, prompt template, agent args), normalized
so legacy saved defaults are already migrated. A new optional reply field:
older clients ignore it, and a client talking to an older host sees none and
keeps using the default agent.

* fix(mobile): ask to update Orca only when the host answered without agent launch

An unread or failed status read settles with no capabilities, which the AI
buttons read as an old host and showed "Update Orca on your computer". The
update copy now needs a status the host actually returned; an unread one keeps
the buttons disabled without blaming the desktop's version.

* fix(mobile): send an AI button's saved agent arguments with its launch

The desktop's direct launches for "Fix checks", "Resolve conflicts" and commit
recovery pass the action's saved agent arguments to agent.launch; the phone
honoured the saved agent but dropped its arguments. It now sends them the same
way: absent when none are saved, so the host keeps the user's configured
defaults. A host that predates the field ignores it.

* test(mobile): repin the RPC recording corpus after the launch recipe and availability fixes

Repins to fe85e0346f. All 790 goldens move only their baseline header: no
scenario saves agent arguments or reads an unreadable status, so no recorded
behaviour changes.

* fix(mobile): say the host status is unreadable instead of nothing when it is

With the update copy now reserved for a host that answered without agent
launch, an unread status left the AI buttons disabled with no explanation.
They now say "Could not read this host's status. Go back and reopen it.", the
words the mobile web shell already uses for the same failure; leaving the host
re-reads its status.

* fix(mobile): wrap an AI button's prompt in the action's saved template

The desktop renders every source-control launch's prompt through the action's
saved template (buildSourceControlRecoveryAgentCommandInput); the phone sent
its built-in prompt as is. Now that the host publishes the recipes, the phone
renders through the same shared function, refuses an empty result as the
desktop does, and offers the rendered text when it could not be delivered.
Review notes have no recipe and are unchanged.

* test(mobile): re-record goldens for the templated AI-button prompt

Repins to 7fd1555d20. One golden body moves: pr-triage-prompt-not-delivered
now carries the prompt as sent (rendered through the action's template) on its
prompt-not-sent result, which is what Copy prompt offers. The other 789
goldens move only their baseline header.

* fix(mobile): re-read a host status that failed while the connection stayed up

A status.get that timed out or was cut over settled the host's gates closed
and was never asked again until the connection state changed, so the phone's
AI buttons stayed disabled behind "Could not read this host's status" on a
link that was working. The gate still settles closed at once, so a failed
read never holds the host screen, but it now re-asks in the background with
the same backoff the runtime capability probe uses, and opens once a status
lands. A reply this app cannot decode is not re-asked.

* test(mobile): repin the RPC recording corpus after the host status re-read

The status gate change moves no recorded behavior: every golden's body is
unchanged and only its baseline header moves to the new pin.

* fix(mobile): show a launch's host warning as a note, not an error

A launch that went ahead can carry a host warning (a structured chat ignores saved agent
arguments, including the '' a template-only save writes). The AI buttons rendered it in the red
error line beside a success haptic. The notice now carries it separately as secondary text, and
review notes keep saying they were sent. Commit recovery also takes the synchronous in-flight lock
the PR triage buttons use, so two taps before a re-render start one agent.

* chore(mobile): record the host status re-read timer for React Doctor

The status re-read arms one retry timer from inside its read and clears it in the effect's
cleanup. React Doctor reports that self-rescheduling shape even in its minimal form, which failed
both changed-lines gates. Suppressed the same way as the session startup timers.

* fix(mobile): say review notes are waiting for the desktop instead of doing nothing

With no live connection, New Agent Session threw from a handler whose promise the sheet drops, so
the tap did nothing visible while the button stayed enabled (proven capabilities survive a drop).
It now closes the sheet and shows "Waiting for desktop..." as the other AI buttons do.

* fix(mobile): stop sending an AI button's saved agent arguments

Whether saved arguments apply depends on the route and shell the host settles after the request
(a chat ignores them and warns; malformed ones fail after admission), and the desktop sends them
only when they apply. The phone cannot know that, so it now leaves them out and the agent's default
arguments apply, as before this series. The saved agent and prompt template still apply.

* fix(mobile): mark review notes sent through the latest save

The sent marks after an agent launch went through the save callback captured at tap time, whose
rollback restores the screen from that moment, so a failed save could drop notes written during
the launch. It now uses the latest render's save, as it already did for the screen state.

* refactor(mobile): own the host status re-read outside the effect

The re-read loop lived inside the effect body, so React Doctor could not see its cleanup and
needed an inline suppression plus a config allowlist entry. The loop is now a plain function that
returns its stop handle, and the effect returns that handle, the same shape every caller of the
runtime capability probe uses. Both suppressions are removed; behaviour is unchanged.

* test(mobile): record the host descriptor from a background status re-read

Pins that the status read records the host descriptor when a re-read succeeds after a failed first
read, not only on the first answer.

* fix(mobile): show a PR AI launch notice only under the button that launched it

Fix checks and Resolve conflicts shared one error, warning and undelivered prompt, so a Fix checks
launch whose prompt was not sent also offered "Copy prompt" under Resolve conflicts, copying the
fix-checks prompt. Notices are now kept per button. The host availability notice stays under each
disabled button, since it explains why that button cannot be tapped.

* fix(mobile): say the host status is being retried instead of asking to reopen it

The host status gate now re-reads a failed status in the background, so "Go back and reopen it"
asked the user for a step that is no longer needed. The review sheet hint uses the same words.
The mobile web shell keeps its own copy.

* refactor(mobile): run the host status gate on the shared status probe

The gate had its own copy of the status probe's retry loop (same delays, same cutover and backoff
split, same stop on an undecodable status). The probe now takes an optional callback for each
failed attempt, which the gate uses to settle closed on the first failure, and the duplicate loop
and its now-unused reader are removed. Existing probe callers are unchanged.

* test(mobile): repin the RPC recording corpus after merging main

Re-records every golden against the merge commit and drops the three goldens whose
scenarios this branch removed, which the merge had restored from main.

* test(mobile): re-record the RPC goldens on the merge with main

Conflicted goldens were seeded from main and re-recorded against the merged
tree; every value either side recorded survives except main's terminal.send in
the PR triage launch, which this branch removes. Drops three goldens main still
had for scenarios this branch deleted.

* feat(mobile): confirm an AI button's agent started, naming the workspace

Fix checks, Resolve conflicts and commit recovery now show "Agent started in
<workspace>" under the button once the host started the agent with its prompt,
so a tap is no longer silent. The workspace label comes from the Source Control
panel and falls back to the branch.

* test(mobile): repin the recording baseline to the success-confirmation commit (header-only)

* fix(mobile): name the workspace in the diff review's AI-button confirmation

The diff review screen mounted the PR sidebar without a workspace label, so
"Agent started in ..." under Fix checks and Resolve conflicts named the branch
while the screen header named the workspace. The sidebar now requires the label
so no screen can drop it, and the diff review passes the one its header shows.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit, the last commit to touch a fenced
path. Against this branch before the merge, only header fields move:
baseline on every golden, and adapterSha256 on the 14 review-action goldens
whose adapter main now drives through the review sheet state. No recorded
body changed.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit, the last commit to touch a fenced
path. Against this branch before the merge only the baseline header moves,
on every golden; no recorded body changed.

* test(mobile): give the send-sheet stacking test the review controller's host status inputs

The merge with main brought in #22951's stacking test, which builds the review
controller without the host capability and status inputs this branch made
required, so the mobile test typecheck ratchet failed.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit 03995ae29d, the last commit to touch a
fenced path. Against this branch before the merge only headers move: baseline
on every golden, and adapterSha256 on the 14 goldens recorded through the
terminal adapter main changed in #23080. No recorded body changed, and the
merged corpus differs from main exactly as this branch did before.
2026-09-28 14:28:09 -07:00
..

Orca Mobile

React Native companion app for Orca. Monitor worktrees, view terminal output, and send commands from your phone.

Local development uses two processes:

  • Orca desktop/Electron from the repo root. This hosts the mobile WebSocket RPC server on port 6768.
  • Expo Metro from mobile/. This serves the React Native app on port 8081.

Unless a command says otherwise, run mobile app commands from the mobile/ directory.

Prerequisites

  • Node.js 24+
  • pnpm
  • Xcode and/or Android Studio tooling for simulator or device builds
  • Expo Go on your phone, or a development client build when native modules are needed
  • Phone and desktop on the same LAN when testing a physical phone

Start Desktop Orca

From the repository root:

pnpm install
pnpm dev

Confirm the mobile RPC server is listening:

lsof -nP -iTCP:6768 -sTCP:LISTEN

Restart pnpm dev after changing Electron main-process code. Metro hot reload only applies to the mobile JavaScript bundle.

Start The Mobile App

cd mobile
pnpm install
pnpm start

Scan the Expo QR code with your phone's camera on iOS, or Expo Go on Android.

For a native dev-client build:

pnpm exec expo run:android
pnpm exec expo run:ios
pnpm start --dev-client

Pair With Desktop Orca

  1. Open Orca desktop.
  2. Go to Settings > Mobile.
  3. Scan the pairing QR code from the mobile app.
  4. Confirm the mobile host endpoint is ws://<desktop-ip>:6768.

For the Android emulator, use ws://10.0.2.2:6768. For a physical phone, use the desktop LAN IP, for example ws://192.168.0.179:6768.

If the phone has a stale host entry, remove it from the app and pair again.

Development Paths

Android Phone

  1. Install Expo Go from Google Play
  2. Run pnpm start, scan QR with Expo Go
  3. For native modules: pnpm exec expo run:android
  4. Run with pnpm start --dev-client

iOS Simulator

  1. Install Xcode from the App Store
  2. Run pnpm start --ios to open in iOS Simulator

Physical Phone Debugging

The phone can be inspected through the connected device tooling:

orca snapshot --json
orca click --element @e3 --json
orca fill --element @e1 --value "ls" --json
orca screenshot --json

Use snapshot first to find the current element refs, then click/fill those refs. After mobile file edits, Metro usually hot reloads automatically, but navigating out of and back into the session screen can be useful because it re-runs terminal.subscribe.

Terminal Streaming Repro Without A Phone

Use this when terminal output does not render on device and you need to split server streaming bugs from WebView/UI bugs:

cd mobile
ORCA_MOBILE_WS_URL=ws://127.0.0.1:6768 pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64>

You can pass a worktree selector as the third argument:

pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "id:<worktreeId>"
pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "path:/absolute/worktree/path"
pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "name:my-worktree"

The expected result includes:

streamSawMarker: true
readSawMarker: true

If this repro fails, debug the desktop runtime/PTY path before the mobile WebView. If it passes but the phone is blank, debug the session screen or TerminalWebView readiness/queueing path.

Terminal Color Repro Without A Phone

Use this when terminal colors disappear after switching tabs. Open a Claude Code terminal and at least one other terminal in the target worktree, then run:

cd mobile
ORCA_MOBILE_WS_URL=ws://127.0.0.1:6768 pnpm exec tsx scripts/repro-terminal-colors.ts \
  <deviceToken> <serverPublicKeyB64> "id:<worktreeId>"

The script captures terminal.subscribe snapshots in an A → B → A sequence and writes raw snapshots to mobile/terminal-color-repro/. If the two A snapshots have different sgrColor counts, the desktop snapshot changed during the switch. If they match, the ANSI color data is still present and the bug is in mobile replay/rendering.

Validation

Run these checks before committing mobile terminal changes:

cd mobile
pnpm exec tsc --noEmit
pnpm run check:tests-typecheck
pnpm lint
cd ..
pnpm typecheck:node

tsc --noEmit reads tsconfig.json, which excludes test files so Metro never bundles them. tsconfig.test.json puts them back, and pnpm run typecheck:tests shows their errors in full. check:tests-typecheck is the gate over it: a ratchet against tests-typecheck-baseline.txt, the 127 test files that do not typecheck yet. It fails when a file that checks today stops checking, and when a baseline entry starts checking (prune it with node scripts/check-tests-typecheck-ratchet.mjs --prune). The list may only shrink.

The same gate censuses the program first: every *.test.ts(x) on disk must be in it, or named in the script's TESTS_OUTSIDE_PROGRAM with a reason. Without that, a test excluded from tsconfig.test.json — or a Foo.test.tsx shadowed by a Foo.test.ts beside it, which a wildcard include drops for the higher-priority extension — would leave the ratchet silently.

Protocol Version Compatibility

Mobile and desktop talk over a versioned protocol. Because mobile updates lag desktop by 24-48h via the App Store, both sides exchange version numbers on status.get so a genuinely incompatible combo can hard-block instead of silently misbehaving.

Constants live in two files (Metro can't resolve outside mobile/):

  • src/shared/protocol-version.ts — DESKTOP_PROTOCOL_VERSION, MIN_COMPATIBLE_MOBILE_VERSION
  • mobile/src/transport/protocol-version.ts — MOBILE_PROTOCOL_VERSION, MIN_COMPATIBLE_DESKTOP_VERSION

Today all four are set so evaluateCompat always returns { kind: 'ok' } — nothing blocks. The wire format is in place to flip a switch when needed.

When to bump

Bump DESKTOP_PROTOCOL_VERSION (and the mobile mirror MOBILE_PROTOCOL_VERSION when relevant) for breaking changes:

  • Removed RPC method or required parameter that mobile uses
  • Changed meaning (units, nullability) of an existing field mobile reads
  • Changed encryption, framing, or auth handshake

Do not bump for additive changes:

  • New RPC methods
  • New optional fields on existing methods
  • New event types in terminal.subscribe

Set MIN_COMPATIBLE_MOBILE_VERSION (kill-switch) when desktop ships a change that requires a minimum mobile version to function safely. Same for MIN_COMPATIBLE_DESKTOP_VERSION from the mobile side.

When a verdict is blocked, mobile/src/components/ProtocolBlockScreen.tsx renders a screen pointing the user at either the App Store (mobile too old) or GitHub Releases (desktop too old).

To exercise the block screen locally: set MIN_COMPATIBLE_DESKTOP_VERSION = 999 in mobile/src/transport/protocol-version.ts, rebuild, pair to any desktop. Revert before merging.

Mock Server

Develop the mobile app without a running Orca desktop instance:

pnpm mock-server           # starts mock WebSocket server on port 6768

Connect from the app using endpoint ws://localhost:6768 and token mock-device-token.

Environment variables

  • MOCK_NATIVE_CHAT=1 — serve the native-chat scenario (one live agent tab, empty transcript, image upload) instead of the default terminal fixtures.
  • MOCK_CHAT_AGENT=omp — with MOCK_NATIVE_CHAT=1, present an OMP tab and four decoded transcript messages, including a tool call and result, instead of the default Claude scenario. It deliberately omits transcriptPath to exercise legacy-hook readability discovery; current OMP hooks may report a path.
  • MOCK_SERVER_KEY_FILE — persist the server keypair across restarts so a paired device keeps its public-key pin. A missing or invalid file is re-keyed with a warning, which forces a re-pair.

Scenario control files

Read on every request, so behaviour can be flipped mid-session without a restart (a restart would re-key E2EE and force a re-pair). Write the mode into the file, or delete it for the default.

  • MOCK_SEND_MODE_FILE (default orca-mock-send-mode in the system temporary directory) — accept (default) accepts the send, error fails it with mobile_input_floor_unavailable, anything else reports the send as rejected.
  • MOCK_TERMINAL_LIST_MODE_FILE (default orca-mock-terminal-list-mode in the system temporary directory) — omit returns an empty terminal list, other returns a list that omits the chat handle, anything else lists it.
  • MOCK_TERMINAL_STREAM_MODE_FILE (default orca-mock-terminal-stream-mode in the system temporary directory) — dead answers a subscribe with subscribed then end (a gone PTY), which is what exercises the rearm bound and terminal prune; anything else streams normally.

Connecting to Real Orca

  1. Start Orca desktop with WebSocket transport enabled
  2. In Orca, go to Settings > Mobile and scan the QR code with this app
  3. The QR encodes the connection endpoint, device token, and TLS fingerprint

Project Structure

mobile/
├── app/                   # Expo Router screens (file-based routing)
│   ├── _layout.tsx        # Root layout with navigation stack
│   ├── index.tsx          # Home screen — paired hosts list
│   └── pair-scan.tsx      # QR code scanning screen
├── src/
│   ├── terminal/          # Terminal WebView and xterm bridge
│   └── transport/         # WebSocket RPC client
├── scripts/
│   ├── test-subscribe.ts  # Desktop streaming repro without a phone
│   └── mock-server.ts     # Standalone mock WebSocket server
└── assets/                # App icons and splash screen