Files
orca/mobile
Jinwoo Hong 895f2cf477 test(mobile): the web app's script fence is re-derived from a measured sweep (OTA phase C, C7.8) (#22152)
* fix(mobile): re-derive the web app script fence from the measured spread (OTA phase C, C7.8)

`4r + 16` was a guess at break-even, and its slack ran from 17 scripts at one
route to 4 at thirteen -- loosest where nothing is and tightest where the tree
actually sits. Re-measured by building every prefix of the sorted route key
list: 15 routes emit 67 scripts, and the marginal cost of a route runs 1 to 9
depending on what it shares, so no line through the route count is both an
upper bound and a budget.

The sweep is now the fence's only input. The envelope is the measurement plus
one margin at the swept tree, growing by the worst route the sweep saw for
every route past it, so a new route breaches it only by costing more than any
route measured. The margin is four, which is the most the count has been seen
to move at a fixed route count with no route added: the head that wrote the old
fence read 32, 43, 61 and 69 at 8, 10, 12 and 14 routes where this one reads
34, 44, 57 and 65.

Two-sided in the test, which is what stops the next bump: a build more than the
margin under the envelope fails there too, so the fence has to be re-measured
rather than raised. The route count stays the only term and the mermaid control
still tells one artifact from 172 chunks.

The shell-fit crossing comes in from 50 routes to 31 with it, because the
envelope grants the worst swept route where `4r + 16` granted four. Byte budgets
re-read from this build and unchanged: 8,053,438 of 9,437,184 total, 1,612,006
before the first route of the 3 MiB allowed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): spell the session route's grant counts off the table (OTA phase C, C7.8)

#22072 deleted `native.wakelock.set` from the session route and left both prose
counts behind: "Fourteen grants" for a list of thirteen, and "the four audio
verbs" for the three that remain. The all-or-none reason went stale with them --
it argued from a screen free to lock, which is the verb that was removed, and
the device side has owned that lock since. It now argues from the microphone a
route granted two of the three cannot close.

The census beside the route-declaration test is what stops the next one. It
reads which number words appear before "grants", "audio verbs", "media verbs"
and "or none" anywhere in the file and compares them with the list itself, so a
second spelling left in place fails rather than passing on the first correct
hit, and a comment rewrapped at a different column still matches.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the four sibling comments spell their grant counts off the tables too (round 1)

#22072 removed `native.wakelock.set` from the session route and left the count
standing in four more files than the route table: "fourteen grants" in the
call-site census and in the hop census, "the four audio grants" in the census
and its test, and "the four audio verbs" on the bridge schemas. Each is now
counted off the table it describes rather than copied beside it.

The six-versus-eight split in the call-site census could not be re-derived --
its own "those eight" never summed to the six rows plus the audio grants, so it
was wrong before #22072 too. It is replaced by what the tables say today: six
rows pin eight of the session route's thirteen grants, and the other five have
censuses of their own. Both numbers come from `PAGE_GRANT_CALL_SITES` and the
route list, and the existing case that names six rows covering eight grants is
what holds them.

The census the route table got is now `spelled-count-census.mjs`, driven from
three files instead of one. A phrase restated in a file collapses to one claim,
so a header and a test name may spell the same count; a phrase with no number
before it reads as the empty list, which no table count matches.

The bridge test is the one site with no count left to pin: it names
`BRIDGE_NATIVE_VERB_NAMES` instead, and a new assertion holds its own
`AUDIO_VERBS` equal to that table's audio rows, so the cases below cannot pin
one set while reading as coverage of another.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): count the session sentence off the session intersection (round 2)

`'of the session'` was counted off every grant in `PAGE_GRANT_CALL_SITES`, which
is not what the sentence it pins is about. A row pinning a grant no route
declares -- the shell can serve a verb before a screen asks for it -- would have
made the census demand the comment overstate what the session route has.

Proved before fixing: adding `native.share.send` to the navigate row moved
`pinnedHere` from eight to nine while the session intersection stayed at eight,
and the census failed asking for "nine of the session route's thirteen grants"
against a sentence that was right to say eight. With the intersection it passes
under the same probe, and the failure moves to `'grants this file pins'`, which
is about this file's rows and does correctly demand nine.

The other three rows are left as they were, for the same reason read the other
way: `'grants this file pins'` and `'did not'` are about the rows here, so they
keep the full list, and `'have censuses of their own'` and
`'are not repeated here'` already count the session grants no row pins.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the two counts in this file's it titles too (round 3)

The census pinned the JSDoc counts and stopped there, so "names six rows
covering eight grants" and "reaches every one of the eight through the session
route" were free to go stale. Shown rather than argued: with a ninth grant on a
row and the header corrected to nine the way the old census forced, the suite
went green with both titles still saying eight.

`'grants this file pins'` becomes `'grants'`, which reads the header and that
title as one claim -- they are the same number, and a row per site would have
let them disagree while both passed. The session title takes the intersection,
for the reason round 2 gave.

Every other spelled number in the file is not a count of a table: "the two
tables" is how many sources the census reads, and "the one it reads", "a new one
cannot be missed", "any one grant went missing" and "every one of" are
quantifiers with no table behind them.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): count the session-route title off the rows it asserts over (round 4)

`'through the session route'` took the intersection, but the title it pins heads
an assertion that compares `grantsNeeded` with every row's grants. The title
names the session route and is not a claim about it: it says the rows here are
all reached through that route, so it moves when the rows move.

In the divergence the JSDoc already described, the two parted. With a ninth
grant on a row the assertion compares nine while the census held the title at
eight and passed, leaving a title reading below the assertion under it.

Exactly one row takes the intersection now, the `.mjs` sentence for how many of
the session route's grants these rows cover, and the JSDoc says so rather than
describing a rule with two members.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): two wording fixes, and the byte readings re-measured on the merge (round 5)

pullfrog on the census JSDoc, both correct. "make the census demand overstates"
was a finite verb where the bare infinitive belongs. "Every other row counts the
rows here" was contradicted by its own table: the rows reading the route's list
and the rows counting grants no row pins take neither count, so the sentence is
scoped to the rows that choose between the two and says what the rest read.

The two byte readings in the fence doc are re-measured on the merged tree, since
they name a head: 8,055,568 of 9,437,184 total and 1,612,253 before the first
route. C8.1 added no route, so the sweep and the envelope are untouched and the
tree still builds 15 routes into 67 scripts.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 02:43:35 -04:00
..

Orca Mobile

React Native companion app for Orca. Monitor worktrees, view terminal output, and send commands from your phone.

Local development uses two processes:

  • Orca desktop/Electron from the repo root. This hosts the mobile WebSocket RPC server on port 6768.
  • Expo Metro from mobile/. This serves the React Native app on port 8081.

Unless a command says otherwise, run mobile app commands from the mobile/ directory.

Prerequisites

  • Node.js 24+
  • pnpm
  • Xcode and/or Android Studio tooling for simulator or device builds
  • Expo Go on your phone, or a development client build when native modules are needed
  • Phone and desktop on the same LAN when testing a physical phone

Start Desktop Orca

From the repository root:

pnpm install
pnpm dev

Confirm the mobile RPC server is listening:

lsof -nP -iTCP:6768 -sTCP:LISTEN

Restart pnpm dev after changing Electron main-process code. Metro hot reload only applies to the mobile JavaScript bundle.

Start The Mobile App

cd mobile
pnpm install
pnpm start

Scan the Expo QR code with your phone's camera on iOS, or Expo Go on Android.

For a native dev-client build:

pnpm exec expo run:android
pnpm exec expo run:ios
pnpm start --dev-client

Pair With Desktop Orca

  1. Open Orca desktop.
  2. Go to Settings > Mobile.
  3. Scan the pairing QR code from the mobile app.
  4. Confirm the mobile host endpoint is ws://<desktop-ip>:6768.

For the Android emulator, use ws://10.0.2.2:6768. For a physical phone, use the desktop LAN IP, for example ws://192.168.0.179:6768.

If the phone has a stale host entry, remove it from the app and pair again.

Development Paths

Android Phone

  1. Install Expo Go from Google Play
  2. Run pnpm start, scan QR with Expo Go
  3. For native modules: pnpm exec expo run:android
  4. Run with pnpm start --dev-client

iOS Simulator

  1. Install Xcode from the App Store
  2. Run pnpm start --ios to open in iOS Simulator

Physical Phone Debugging

The phone can be inspected through the connected device tooling:

orca snapshot --json
orca click --element @e3 --json
orca fill --element @e1 --value "ls" --json
orca screenshot --json

Use snapshot first to find the current element refs, then click/fill those refs. After mobile file edits, Metro usually hot reloads automatically, but navigating out of and back into the session screen can be useful because it re-runs terminal.subscribe.

Terminal Streaming Repro Without A Phone

Use this when terminal output does not render on device and you need to split server streaming bugs from WebView/UI bugs:

cd mobile
ORCA_MOBILE_WS_URL=ws://127.0.0.1:6768 pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64>

You can pass a worktree selector as the third argument:

pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "id:<worktreeId>"
pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "path:/absolute/worktree/path"
pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "name:my-worktree"

The expected result includes:

streamSawMarker: true
readSawMarker: true

If this repro fails, debug the desktop runtime/PTY path before the mobile WebView. If it passes but the phone is blank, debug the session screen or TerminalWebView readiness/queueing path.

Terminal Color Repro Without A Phone

Use this when terminal colors disappear after switching tabs. Open a Claude Code terminal and at least one other terminal in the target worktree, then run:

cd mobile
ORCA_MOBILE_WS_URL=ws://127.0.0.1:6768 pnpm exec tsx scripts/repro-terminal-colors.ts \
  <deviceToken> <serverPublicKeyB64> "id:<worktreeId>"

The script captures terminal.subscribe snapshots in an A → B → A sequence and writes raw snapshots to mobile/terminal-color-repro/. If the two A snapshots have different sgrColor counts, the desktop snapshot changed during the switch. If they match, the ANSI color data is still present and the bug is in mobile replay/rendering.

Validation

Run these checks before committing mobile terminal changes:

cd mobile
pnpm exec tsc --noEmit
pnpm run check:tests-typecheck
pnpm lint
cd ..
pnpm typecheck:node

tsc --noEmit reads tsconfig.json, which excludes test files so Metro never bundles them. tsconfig.test.json puts them back, and pnpm run typecheck:tests shows their errors in full. check:tests-typecheck is the gate over it: a ratchet against tests-typecheck-baseline.txt, the 127 test files that do not typecheck yet. It fails when a file that checks today stops checking, and when a baseline entry starts checking (prune it with node scripts/check-tests-typecheck-ratchet.mjs --prune). The list may only shrink.

The same gate censuses the program first: every *.test.ts(x) on disk must be in it, or named in the script's TESTS_OUTSIDE_PROGRAM with a reason. Without that, a test excluded from tsconfig.test.json — or a Foo.test.tsx shadowed by a Foo.test.ts beside it, which a wildcard include drops for the higher-priority extension — would leave the ratchet silently.

Protocol Version Compatibility

Mobile and desktop talk over a versioned protocol. Because mobile updates lag desktop by 24-48h via the App Store, both sides exchange version numbers on status.get so a genuinely incompatible combo can hard-block instead of silently misbehaving.

Constants live in two files (Metro can't resolve outside mobile/):

  • src/shared/protocol-version.ts — DESKTOP_PROTOCOL_VERSION, MIN_COMPATIBLE_MOBILE_VERSION
  • mobile/src/transport/protocol-version.ts — MOBILE_PROTOCOL_VERSION, MIN_COMPATIBLE_DESKTOP_VERSION

Today all four are set so evaluateCompat always returns { kind: 'ok' } — nothing blocks. The wire format is in place to flip a switch when needed.

When to bump

Bump DESKTOP_PROTOCOL_VERSION (and the mobile mirror MOBILE_PROTOCOL_VERSION when relevant) for breaking changes:

  • Removed RPC method or required parameter that mobile uses
  • Changed meaning (units, nullability) of an existing field mobile reads
  • Changed encryption, framing, or auth handshake

Do not bump for additive changes:

  • New RPC methods
  • New optional fields on existing methods
  • New event types in terminal.subscribe

Set MIN_COMPATIBLE_MOBILE_VERSION (kill-switch) when desktop ships a change that requires a minimum mobile version to function safely. Same for MIN_COMPATIBLE_DESKTOP_VERSION from the mobile side.

When a verdict is blocked, mobile/src/components/ProtocolBlockScreen.tsx renders a screen pointing the user at either the App Store (mobile too old) or GitHub Releases (desktop too old).

To exercise the block screen locally: set MIN_COMPATIBLE_DESKTOP_VERSION = 999 in mobile/src/transport/protocol-version.ts, rebuild, pair to any desktop. Revert before merging.

Mock Server

Develop the mobile app without a running Orca desktop instance:

pnpm mock-server           # starts mock WebSocket server on port 6768

Connect from the app using endpoint ws://localhost:6768 and token mock-device-token.

Environment variables

  • MOCK_NATIVE_CHAT=1 — serve the native-chat scenario (one live agent tab, empty transcript, image upload) instead of the default terminal fixtures.
  • MOCK_CHAT_AGENT=omp — with MOCK_NATIVE_CHAT=1, present an OMP tab and four decoded transcript messages, including a tool call and result, instead of the default Claude scenario. It deliberately omits transcriptPath to exercise legacy-hook readability discovery; current OMP hooks may report a path.
  • MOCK_SERVER_KEY_FILE — persist the server keypair across restarts so a paired device keeps its public-key pin. A missing or invalid file is re-keyed with a warning, which forces a re-pair.

Scenario control files

Read on every request, so behaviour can be flipped mid-session without a restart (a restart would re-key E2EE and force a re-pair). Write the mode into the file, or delete it for the default.

  • MOCK_SEND_MODE_FILE (default orca-mock-send-mode in the system temporary directory) — accept (default) accepts the send, error fails it with mobile_input_floor_unavailable, anything else reports the send as rejected.
  • MOCK_TERMINAL_LIST_MODE_FILE (default orca-mock-terminal-list-mode in the system temporary directory) — omit returns an empty terminal list, other returns a list that omits the chat handle, anything else lists it.
  • MOCK_TERMINAL_STREAM_MODE_FILE (default orca-mock-terminal-stream-mode in the system temporary directory) — dead answers a subscribe with subscribed then end (a gone PTY), which is what exercises the rearm bound and terminal prune; anything else streams normally.

Connecting to Real Orca

  1. Start Orca desktop with WebSocket transport enabled
  2. In Orca, go to Settings > Mobile and scan the QR code with this app
  3. The QR encodes the connection endpoint, device token, and TLS fingerprint

Project Structure

mobile/
├── app/                   # Expo Router screens (file-based routing)
│   ├── _layout.tsx        # Root layout with navigation stack
│   ├── index.tsx          # Home screen — paired hosts list
│   └── pair-scan.tsx      # QR code scanning screen
├── src/
│   ├── terminal/          # Terminal WebView and xterm bridge
│   └── transport/         # WebSocket RPC client
├── scripts/
│   ├── test-subscribe.ts  # Desktop streaming repro without a phone
│   └── mock-server.ts     # Standalone mock WebSocket server
└── assets/                # App icons and splash screen