Files
orca/mobile
Brennan Benson 21124db4d5 refactor(native-chat): a subagent's rows live in its own section, not in the parent's conversation (#23752)
* refactor(native-chat): a subagent's rows live with that subagent, not in the conversation

A subagent's rows were drawn in its parent's conversation, each captioned with
the subagent's name. They now belong to the subagent: the transcript projection
keeps the session's own rows as the conversation and each subagent's rows apart,
keyed by the agent id its roster entry already carries, folded on their own.

Desktop: a subagent's rows open in a section under the roster row that names it,
from that agent's roster entry, and are windowed like any other rows. A subagent
no loaded roster names opens where its first row happened, inside the section of
the agent that spawned it or in the conversation. Its edits still count in the
turn they were made, and revealing one opens the sections around it.

Mobile shows the conversation, with each spawn's roster line. Worker reads and
structured terminal reads serve the worker's own rows.

Removes what the move makes redundant: the per-row caption and its copy, the
producer check in the tool fold and the turn answer, the per-agent frontier
interleaved in the conversation, worker-text subagent tags, and the agent id on
worker-read messages.

* refactor(native-chat): a diff target names the sections its row sits in

Revealing a subagent's edit opens the sections around it from the target the
rollup already holds, instead of looking the row up at click time. The section
head keeps to the agent's name and dot; its state in words stays on the roster
entry. The worker page test stubs the host through its module rather than a cast.

* fix(native-chat): a working subagent's section is open; a worker page windows its own rows

A subagent's section is open while its agent works and closes once it settles,
the way the turn's own live run does; a section the reader opened or closed by
hand keeps that choice. A subagent another subagent spawned opens inside that
one's section, so a working grandchild shows inside its working parent. Openness
is derived from the roster's state and the reader's choices; nothing stores an
automatic open.

A worker page is now the newest page of the worker's own rows. The host windows
the read over them before the limit, so a subagent's burst can no longer crowd
the worker's rows off the page, and "older" still means older worker rows. The
scope is an in-process argument of the host's history read; no wire request
carries it.

* fix(native-chat): a subagent section head names the turn it sits in, for the outline rail

* fix(native-chat): a subagent section's rows sit in the turn the section is shown in, for the outline rail

A background subagent's rows written during a later turn carried that later
turn onto their slots, so scrolling through its section lit the later turn's
rail tick and then snapped back. The rollup still counts each edit in the turn
it was made; only the slot, which the rail reads, takes the shown turn.

* fix(mobile): Load earlier reads past pages that hold only a subagent's rows

Mobile draws only the session's own rows, so an older page made entirely of a
subagent's rows landed as nothing: the reader tapped Load earlier, saw the
spinner, and got the same transcript back. One load now reads on (up to 8 pages)
until a page holds a row of the session's own, then applies the pages in order.

* test(mobile): stub the RPC client the way the other structured-session hook tests do

* perf(native-chat): order subagent rows for the changed-files rollup once per change to them

The rollup flattened and re-sorted every subagent row on each update, including
every token the parent streamed. The ordering now keys on the projection's
subagent rows, which keep their identity while only the conversation changes.

* refactor(native-chat): order subagent rows in the sections hook, keeping the list under its line limit

* fix(mobile): a transcript whose newest page is only a subagent's rows reads back on its own

Opened while a subagent is busy, the newest page can hold nothing but that
subagent's rows. Mobile draws none of them, so the reader saw an empty chat with
a Load earlier button, and an empty list cannot be scrolled to page. The hook now
reads back once from each such head, and the read runs on to the session's own rows.

* fix(native-chat): count the live window in the session's own rows, so a subagent's burst keeps its roster

The live window kept the newest 1,024 rows of every agent. A subagent writing
more than that trimmed its own spawn's roster row and the prompt, and its
section fell back to a closed, unnamed header. The window now keeps the newest
1,024 of the session's own rows and everything after, with an 8,192-row cap on
every agent's rows as the memory backstop. A transcript with no subagent rows
trims exactly as before.

* fix(agent-session): window history pages by the session's own rows, with a subagent's rows riding along

A history page held the newest 200 rows of every agent, so a subagent's burst
could fill a page on its own: the phone opened on an empty chat and "Load
earlier" landed nothing. A page now starts at the oldest of the newest `limit`
rows of the session's own and serves every row from there, so the subagent's
rows come with the conversation they happened in. The page stays contiguous,
the cursor still names its first row, and the byte bound still applies. A
transcript with no subagent rows gets the same pages as before.

Clients already take a page larger than its limit: both reducers raise their
retained window to the page's size. The mobile read-on and read-back stay for
older hosts.

* test(agent-session): a page reaches back to the start rather than leaving a subagent-only page

* fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it

The live window trimmed to just after the own row it dropped, so a subagent
whose roster row went kept its rows at the top as an unnamed section until
the parent wrote again. Trim to the oldest own row kept instead; it still
fires only once an own row passes the limit, so a paged-in run of subagent
rows at the head stays until then. With no subagent rows nothing changes.

* perf(native-chat): cap the live window at 4,096 rows, bounding each delta's re-derivation

Every live batch re-derives the transcript over every retained row. On the
largest real window (7,374 rows) that cost 7-8 ms a delta on desktop against
0.6 ms at the old 1,024-row window, and held about 26 MB of row content.
4,096 halves both. The most rows any local journal puts between a roster and
its subagent's last row, with the parent inside its own-row limit, is 3,005,
so no observed subagent loses its roster to the lower cap.

* fix(native-chat): a subagent section opens only while its roster is the running scope's live frontier

A section used to open whenever its roster said the subagent was working, anywhere
in the transcript and whether or not the session was running, so a background
subagent's section stayed open and grew mid-transcript while the parent moved on.

It now opens by default only while the session runs and the roster row naming the
subagent is the newest thing the parent produced, user rows aside. Newer parent
output closes it even while the subagent still works; the roster row keeps
showing that live state. A subagent still working is a running scope of its own
for the sections it spawned; a settled one closes its scope. Derived every
render, no latch; the reader's own open or close still wins.

* fix(native-chat): name a subagent's section from a client roster the window never trims

A section took its name and state from a roster row in the loaded window. Once a
burst trimmed that row, or the row sat on an older page, the section fell back to
an unnamed, closed "Subagent" header.

The shared reducer now keeps a roster keyed by agent id, folded from every roster
row and revision the client receives: pages, older pages and live batches,
including revisions of roster rows outside the window, which live batches already
carry. The first roster naming an agent wins and its revisions update it; a
removed roster row drops its entries; it is rebuilt on every page that replaces
the window and bounded to 512 agents. Sections take their name, state and
live-frontier place from it; placement stays under the loaded roster row, else
at the section's first loaded row. Only a subagent no roster ever named stays
unnamed.

* feat(agent-session): a history page names the subagents whose roster row is older than it

A page is a contiguous run of the journal whose older-page cursor is its first
item, so it cannot pull an older roster row in without skipping the rows between.
When a page held a subagent's rows but not the roster row naming it (about 11% of
the moments a reader could open a session on local journals), that subagent drew
as an unnamed "Subagent" header.

History and hydration pages now carry an optional `subagentRoster`: the first
roster entry naming each subagent whose rows are on the page and whose roster row
is not, with the row's id, sequence and revision; bounded to 64 entries and
16 KB. Items and cursor are unchanged. The client seeds its roster from it.

Rule 1 in docs/reference/remote-wire-compatibility.md: an optional field on an
existing frame, no capability gate. An older client ignores it (the released
reducer reads a page with it exactly as one without); against an older host the
field is absent and the section falls back to an unnamed header.

* Revert "fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it"

This reverts commit 22078656b2.

Its only purpose was to stop a subagent whose roster row an own-row trim had
dropped from showing at the head of the window as an unnamed section. The client
roster now names that section whatever the window holds, so the cut is back at
just after the own row the limit passes. The retention test that pinned the
unnamed-section case now asserts the section at the head keeps its name.

* chore(native-chat): state the retention limits' own reasons, now that no name depends on the window

Own-row retention keeps the conversation a reader sees from being crowded out by
rows drawn as a one-row section on desktop and not at all on mobile; the
every-agent cap bounds memory and each live delta's re-derivation. Neither is
about keeping a roster row loaded any more.

* fix(native-chat): hold the roster fold's draft map where type narrowing can see closure writes

* fix(native-chat): a roster row's newer revision replaces it in the client roster too

A revision that stops naming an agent (the host drops an entry it learns is not a
subagent, or re-keys a provisional one) left the client roster holding the old
entry, often still "working", with nothing to re-derive it. The section then read
as working forever and could auto-open, while a fresh read of the same journal
left it unnamed. The fold now drops an entry when a newer revision of the row that
named it no longer does, before any roster takes it over.

* fix(native-chat): a parent's spawn and wait calls keep the subagent they name open

A subagent section auto-opened only while its roster row was the running session's
newest row, so any later row closed it: a Codex wait on the agent, or the parent's
text before its next spawn call. Now a row that is part of delegating to a subagent
keeps that subagent open:

- a Codex collab call (spawn, wait, resume, message, close) opens each agent its
  receiver thread ids name; one naming none is ordinary output;
- a Claude spawn call names no agent, so it counts toward the roster announcing it;
- a roster row at the frontier opens its most recently added agent, not all of them.

A roster or call naming only agents one subagent spawned is that subagent's output,
so a grandchild's roster, which the host journals as the session's row, no longer
closes the spawner's section.

* fix(native-chat): a parent's call right after the roster closes its subagent's section

A parent's tool calls after a roster row fold into the tool run drawn above
the roster, so the roster stayed the newest drawn row and its section stayed
open while the parent was already reading or running commands. The fold now
records the newest journal position among the rows it merged, and the live
frontier orders rows by that newest part. The layout is unchanged. A spawn
call folded there still counts as part of the roster announcing it.

* fix(native-chat): a Codex call naming several subagents delegates to the first

A Codex collab call that names several agents opened every one of their
sections. It now counts as delegating to the first agent it names, so one
section opens, the same as a call naming one agent.

* fix(native-chat): closing a roster's list closes the sections under it

Collapsing a roster row's list of subagents left their open sections drawn,
so the section's own head became the only way to close them. And the list's
open state lived in the row, so a row the window unmounted came back
collapsed.

The transcript now holds each roster list's open state beside the section
choices. A closed list hides every section it anchors; each section keeps
its own open or closed choice for when the list reopens. With no choice from
the reader, a list is open while a section under it is open. Closing a
section from its entry keeps the list open, and revealing a subagent's edit
opens the list it sits under.

* perf(native-chat): a reveal finds the roster lists it opens with one set lookup per entry

* fix(native-chat): a subagent's roster entry heads its own rows

An open section drew the agent's name twice: its entry in the roster's list,
then a separate section head above its rows. The entry is now the head. The
roster row draws its entries through the first open one, that agent's rows
follow, then the entries after it, each run in its own windowed slot. A
section no loaded roster row holds (an older page, a grandchild, an unnamed
agent) keeps its own head.

A roster list is open while the live frontier or a reader's choice is on
one of its agents, unless the reader closed the list, so closing an agent
from its entry no longer needs to pin the list open.

The section emitter moves to its own module, and the trailing-run
predicates it shares with the slot builder to theirs, to keep the slot
builder under its line limit.

* fix(native-chat): the entries after an open subagent's rows set in its roster's type

The roster row's list inherits the system row's small muted type; the entries that
follow an open section sit outside that row, so they now carry the same type.

* docs(native-chat): a current host can also serve a page of only a subagent's rows

A page is bounded by bytes after it is windowed by the session's own rows, so a
burst that fills the bound yields a page, or an opening page, with none of the
session's own rows. Mobile's read-on and read-back therefore serve current hosts
too, not only older ones; the comments said otherwise. The retention comment
still described a closed section as a row of its own; it now sits behind its
roster entry.

* fix(native-chat): a section's prose keeps its copy/timestamp controls inside the section

An assistant row's hover controls (copy, scroll-to-top, timestamp) hang 20px
below the row into the gap before the next one (`-mb-5`). Inside a subagent's
section that put them below the section's left border, and on the section's
last row they touched the parent's next row with no gap.

Inside a section the controls now stay in flow, so the border covers them and
the next row sits the normal gap below. The row-height estimate reserves the
same 20px for a section's prose so windowing does not jump on measure.

* test(agent-session): state each appended row's turn scope, as the journal now requires

* refactor(native-chat): the client's journal retention policy lives in its own module
2026-09-29 23:47:02 -07:00
..

Orca Mobile

React Native companion app for Orca. Monitor worktrees, view terminal output, and send commands from your phone.

Local development uses two processes:

  • Orca desktop/Electron from the repo root. This hosts the mobile WebSocket RPC server on port 6768.
  • Expo Metro from mobile/. This serves the React Native app on port 8081.

Unless a command says otherwise, run mobile app commands from the mobile/ directory.

Prerequisites

  • Node.js 24+
  • pnpm
  • Xcode and/or Android Studio tooling for simulator or device builds
  • Expo Go on your phone, or a development client build when native modules are needed
  • Phone and desktop on the same LAN when testing a physical phone

Start Desktop Orca

From the repository root:

pnpm install
pnpm dev

Confirm the mobile RPC server is listening:

lsof -nP -iTCP:6768 -sTCP:LISTEN

Restart pnpm dev after changing Electron main-process code. Metro hot reload only applies to the mobile JavaScript bundle.

Start The Mobile App

cd mobile
pnpm install
pnpm start

Scan the Expo QR code with your phone's camera on iOS, or Expo Go on Android.

For a native dev-client build:

pnpm exec expo run:android
pnpm exec expo run:ios
pnpm start --dev-client

Pair With Desktop Orca

  1. Open Orca desktop.
  2. Go to Settings > Mobile.
  3. Scan the pairing QR code from the mobile app.
  4. Confirm the mobile host endpoint is ws://<desktop-ip>:6768.

For the Android emulator, use ws://10.0.2.2:6768. For a physical phone, use the desktop LAN IP, for example ws://192.168.0.179:6768.

If the phone has a stale host entry, remove it from the app and pair again.

Development Paths

Android Phone

  1. Install Expo Go from Google Play
  2. Run pnpm start, scan QR with Expo Go
  3. For native modules: pnpm exec expo run:android
  4. Run with pnpm start --dev-client

iOS Simulator

  1. Install Xcode from the App Store
  2. Run pnpm start --ios to open in iOS Simulator

Physical Phone Debugging

The phone can be inspected through the connected device tooling:

orca snapshot --json
orca click --element @e3 --json
orca fill --element @e1 --value "ls" --json
orca screenshot --json

Use snapshot first to find the current element refs, then click/fill those refs. After mobile file edits, Metro usually hot reloads automatically, but navigating out of and back into the session screen can be useful because it re-runs terminal.subscribe.

Terminal Streaming Repro Without A Phone

Use this when terminal output does not render on device and you need to split server streaming bugs from WebView/UI bugs:

cd mobile
ORCA_MOBILE_WS_URL=ws://127.0.0.1:6768 pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64>

You can pass a worktree selector as the third argument:

pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "id:<worktreeId>"
pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "path:/absolute/worktree/path"
pnpm exec tsx scripts/test-subscribe.ts <deviceToken> <serverPublicKeyB64> "name:my-worktree"

The expected result includes:

streamSawMarker: true
readSawMarker: true

If this repro fails, debug the desktop runtime/PTY path before the mobile WebView. If it passes but the phone is blank, debug the session screen or TerminalWebView readiness/queueing path.

Terminal Color Repro Without A Phone

Use this when terminal colors disappear after switching tabs. Open a Claude Code terminal and at least one other terminal in the target worktree, then run:

cd mobile
ORCA_MOBILE_WS_URL=ws://127.0.0.1:6768 pnpm exec tsx scripts/repro-terminal-colors.ts \
  <deviceToken> <serverPublicKeyB64> "id:<worktreeId>"

The script captures terminal.subscribe snapshots in an A → B → A sequence and writes raw snapshots to mobile/terminal-color-repro/. If the two A snapshots have different sgrColor counts, the desktop snapshot changed during the switch. If they match, the ANSI color data is still present and the bug is in mobile replay/rendering.

Validation

Run these checks before committing mobile terminal changes:

cd mobile
pnpm exec tsc --noEmit
pnpm run check:tests-typecheck
pnpm lint
cd ..
pnpm typecheck:node

tsc --noEmit reads tsconfig.json, which excludes test files so Metro never bundles them. tsconfig.test.json puts them back, and pnpm run typecheck:tests shows their errors in full. check:tests-typecheck is the gate over it: a ratchet against tests-typecheck-baseline.txt, the 127 test files that do not typecheck yet. It fails when a file that checks today stops checking, and when a baseline entry starts checking (prune it with node scripts/check-tests-typecheck-ratchet.mjs --prune). The list may only shrink.

The same gate censuses the program first: every *.test.ts(x) on disk must be in it, or named in the script's TESTS_OUTSIDE_PROGRAM with a reason. Without that, a test excluded from tsconfig.test.json — or a Foo.test.tsx shadowed by a Foo.test.ts beside it, which a wildcard include drops for the higher-priority extension — would leave the ratchet silently.

Protocol Version Compatibility

Mobile and desktop talk over a versioned protocol. Because mobile updates lag desktop by 24-48h via the App Store, both sides exchange version numbers on status.get so a genuinely incompatible combo can hard-block instead of silently misbehaving.

Constants live in two files (Metro can't resolve outside mobile/):

  • src/shared/protocol-version.ts — DESKTOP_PROTOCOL_VERSION, MIN_COMPATIBLE_MOBILE_VERSION
  • mobile/src/transport/protocol-version.ts — MOBILE_PROTOCOL_VERSION, MIN_COMPATIBLE_DESKTOP_VERSION

Today all four are set so evaluateCompat always returns { kind: 'ok' } — nothing blocks. The wire format is in place to flip a switch when needed.

When to bump

Bump DESKTOP_PROTOCOL_VERSION (and the mobile mirror MOBILE_PROTOCOL_VERSION when relevant) for breaking changes:

  • Removed RPC method or required parameter that mobile uses
  • Changed meaning (units, nullability) of an existing field mobile reads
  • Changed encryption, framing, or auth handshake

Do not bump for additive changes:

  • New RPC methods
  • New optional fields on existing methods
  • New event types in terminal.subscribe

Set MIN_COMPATIBLE_MOBILE_VERSION (kill-switch) when desktop ships a change that requires a minimum mobile version to function safely. Same for MIN_COMPATIBLE_DESKTOP_VERSION from the mobile side.

When a verdict is blocked, mobile/src/components/ProtocolBlockScreen.tsx renders a screen pointing the user at the update that clears it. When mobile is too old it opens the newest release if the installed app's update check knows one, otherwise the App Store (iOS) or GitHub Releases (Android). When desktop is too old it opens GitHub Releases.

To exercise the block screen locally: set MIN_COMPATIBLE_DESKTOP_VERSION = 999 in mobile/src/transport/protocol-version.ts, rebuild, pair to any desktop. Revert before merging.

Mock Server

Develop the mobile app without a running Orca desktop instance:

pnpm mock-server           # starts mock WebSocket server on port 6768

Connect from the app using endpoint ws://localhost:6768 and token mock-device-token.

Environment variables

  • MOCK_NATIVE_CHAT=1 — serve the native-chat scenario (one live agent tab, empty transcript, image upload) instead of the default terminal fixtures.
  • MOCK_CHAT_AGENT=omp — with MOCK_NATIVE_CHAT=1, present an OMP tab and four decoded transcript messages, including a tool call and result, instead of the default Claude scenario. It deliberately omits transcriptPath to exercise legacy-hook readability discovery; current OMP hooks may report a path.
  • MOCK_SERVER_KEY_FILE — persist the server keypair across restarts so a paired device keeps its public-key pin. A missing or invalid file is re-keyed with a warning, which forces a re-pair.

Scenario control files

Read on every request, so behaviour can be flipped mid-session without a restart (a restart would re-key E2EE and force a re-pair). Write the mode into the file, or delete it for the default.

  • MOCK_SEND_MODE_FILE (default orca-mock-send-mode in the system temporary directory) — accept (default) accepts the send, error fails it with mobile_input_floor_unavailable, anything else reports the send as rejected.
  • MOCK_TERMINAL_LIST_MODE_FILE (default orca-mock-terminal-list-mode in the system temporary directory) — omit returns an empty terminal list, other returns a list that omits the chat handle, anything else lists it.
  • MOCK_TERMINAL_STREAM_MODE_FILE (default orca-mock-terminal-stream-mode in the system temporary directory) — dead answers a subscribe with subscribed then end (a gone PTY), which is what exercises the rearm bound and terminal prune; anything else streams normally.

Connecting to Real Orca

  1. Start Orca desktop with WebSocket transport enabled
  2. In Orca, go to Settings > Mobile and scan the QR code with this app
  3. The QR encodes the connection endpoint, device token, and TLS fingerprint

Project Structure

mobile/
├── app/                   # Expo Router screens (file-based routing)
│   ├── _layout.tsx        # Root layout with navigation stack
│   ├── index.tsx          # Home screen — paired hosts list
│   └── pair-scan.tsx      # QR code scanning screen
├── src/
│   ├── terminal/          # Terminal WebView and xterm bridge
│   └── transport/         # WebSocket RPC client
├── scripts/
│   ├── test-subscribe.ts  # Desktop streaming repro without a phone
│   └── mock-server.ts     # Standalone mock WebSocket server
└── assets/                # App icons and splash screen