Files
orca/mobile/src/session
Brennan BensonandMerge Sim 7faa9f7cd3 fix(native-chat): reasoning rows with a real open/finished state, and readable Claude thinking (#19221)
* fix(native-chat): render structured reasoning as collapsible messages

* fix(native-chat): align expanded reasoning with summary

* fix(native-chat): place reasoning chevron after summary

* De-emphasize reasoning headlines with observed timing labels

* Exclude later turn work from observed reasoning duration

* fix(native-chat): give reasoning rows a host-owned open/closed lifecycle

A reasoning row now says whether its block is still streaming (`state`)
and, when the host saw it end, when (`completedAt`). The row's own
start is its first-write time, so "Thought for N s" is measured on the
execution host instead of by whichever window happened to be watching.

Rows open with their first non-empty text and close on every path that
ends them: the block's final frame, a new message in the same stream,
every Claude turn end through one hook on the open turn, Codex
item/completed, turn and session settlement, active-item eviction, and
both host sweeps for a dead generation. Closing writes are lifecycle
writes so backpressure cannot leave a row open. Rows without the field
(older hosts, older journals) never read as live.

The reasoning row's headline reads the row's lifecycle and its own
turn's liveness: Thinking while open in a running turn, then Thought
for N s, Thought when no span was seen, Reasoning when the host kept no
lifecycle.

* feat(native-chat): ask Claude for readable thinking summaries

Under Orca's launch the Claude CLI streams thinking blocks with empty
text, so no reasoning row ever had anything to show. Pass only
`--thinking-display summarized`: it fills thinking blocks with the
API's summaries without turning thinking on, so a user who disabled
thinking keeps it off.

Orca runs the user's own binary, and a CLI older than the flag exits on
it before the session starts. The launch probes the binary it is about
to run, overlapping the rest of launch resolution, and passes the flag
only when that probe has already answered with 2.1.94 or newer. A slow
or failed probe never delays a launch and is never remembered; a
successful one is kept per binary until the binary changes.

* test(native-chat): type the superseding send in the reasoning lifecycle test

* fix(native-chat): stop reading an ended reasoning row as thinking

The spinner line infers "Thinking" from the newest root row being a
reasoning message. With summaries on, a closed reasoning row stays the
newest row while Claude streams a tool's input, so the line read
Thinking for the whole Write. A reasoning row now counts only while its
own state is running; a row from a host that keeps no state reads as
before.

* fix(native-chat): measure reasoning from the block's start, not its first text

A row opens with its first summary text, which trails the thinking
block's start by seconds, and the journal stamps a row with its queued
append time. So "Thought for N s" read 1 s for blocks that ran 4.65 s
and 12.95 s. The block registry now records the host time at each
block's content_block_start (else its first delta), and every write of
that block's row carries it as the row's observed time; the close keeps
the frame's own receipt time. Codex reasoning rows likewise carry their
item/started time.

* perf(native-chat): stop rewriting the reasoning row for every thinking token

Claude sends a thinking_tokens frame after every thinking delta. Every
frame the stream path did not consume forced the streamed text out
first, bypassing the checkpoint widening and the coalescing window, so
a single thinking block was rewritten and republished once per token:
212 full-row writes and 258 KB for one captured block. A frame that can
write no row (token tallies, stream deltas no registry carries, pings)
no longer forces that flush; every frame that can write still does.
The same block now takes 14 writes.

* fix(native-chat): time Codex reasoning by receipt and close evicted rows

Codex summary text lands 13-50 ms before item/completed, inside one
coalescing window, so a reasoning row's first write can be its
completion. Item boundaries are now stamped with their host receipt
time the way turn boundaries already were, so a retried or buffered
delivery keeps it, and the item translator falls back to the host
clock rather than Date.now(). The row carries its item/started time on
every first write, including the completion, so its span is
item/started to item/completed.

An evicted active item is now closed from the text streamed so far,
like both settle paths, and the eviction runs before the incoming item
is tracked: tracking first let the stream bound drop the evictee's text
before the eviction could close its row, stranding it running.

completedAt now has one meaning everywhere: the host time the message
was seen to end, or the end of the turn or stream that cut it off;
absent only when no end was seen live.

* refactor(native-chat): keep the Codex streaming body translation pure

The streaming translation preserves a reasoning body as-is again; the
stream writer, which is what knows the item has not completed, stamps
it running.

* test(native-chat): keep a re-collapsed reasoning row collapsed through a revision

* fix(native-chat): estimate a collapsed reasoning row as its trigger

A reasoning row renders collapsed, as one small button, but its height
was estimated from its full text: a 4,129-character summary reserved
about 950 px for a 24 px row, so long chats jumped as rows were
measured. It is now estimated at the trigger's height; opening the row
remeasures it.

* fix(native-chat): probe the CLI the launch will run, and learn from a refusal

The thinking-display gate probed `claude --version` with Orca's own cwd
and env, while the launch spawns with the workspace's cwd and the shell
env. Behind a version manager's shim those can pick different CLIs, so
the probe could approve a CLI the launch never ran, and an older CLI
exits on the unknown flag before the session starts. The probe also
only counted if it had already finished when resolution did, so a
first launch, or the first after a CLI update, usually went without
the flag.

The probe now runs with the launch's own resolved cwd and env, is keyed
by the binary and the workspace, and a launch waits up to 200 ms for it
(about 3x the probe's measured p95) before going without the flag. Only
answers are kept, so a slow or failed probe is asked again next launch.
A child that exits with commander's "unknown option '--thinking-display'"
marks that binary in that workspace so the next launch skips the flag;
that one start fails exactly as any CLI startup failure does today.

* test(native-chat): type the thinking-display probe mock with both of its parameters

* fix(native-chat): recheck the account switch after the probe, last as before

Moving the invocation ahead of the probe, the transcript check and the
permission mode put its account-switch recheck before those awaits, so a
switch that began during them launched unchecked. The invocation is the
last await again. The probe gets its own env from the same sources the
launch uses, the inherited env and the overlay with the CLI's runtime on
PATH, built by the same code, minus every credential: asking a CLI its
version needs none.

* fix(native-chat): wait up to 1.5 s for a cold probe, and remember every outcome

200 ms only covered a warm CLI; a cold disk, a node install or an
antivirus scan exceeds it, and that chat's child then ran its whole
life without summaries. A launch now waits up to 1.5 s, once per binary
per workspace, measured from when that binary's probe began, so a later
launch never waits again on a probe already past it. The probe gets its
own 10 s kill timeout, and every outcome, including no version printed,
a failure or a kill, is kept for the binary's life, so a probe that
hangs costs one launch rather than every one. A refusal seen while a
probe still runs wins over its late answer.

* fix(native-chat): leave Thinking to the activity line while reasoning runs

A reasoning row still being written drew a pulsing "Thinking…" header
right under the turn's activity line, which already says Thinking: two
live indicators for one fact. In a running turn an open reasoning row
now draws nothing and reserves no height; it appears when it closes, as
"Thought for N s". A row from a host that keeps no state, a closed row,
and a row left open by a turn that ended draw as before. The row's
Thinking headline is gone with its catalog key.

* fix(native-chat): end every unfinished Codex item through one rule

A reasoning completion with no text of its own left the row its stream
wrote running for good: the completion translated to nothing and the
item left the active set, so no settle could reach it. It now closes
from the text streamed so far.

Settlement, eviction and that completion now build an unfinished item's
row through one choice (the streamed text when there is any, else the
item as it started) and end it through one rule. An evicted file change
with streamed tool output no longer keeps that output as its patch; it
reads as interrupted, as a settled one does. A completion whose start
was never recorded claims no span, so it reads "Thought".

* fix(native-chat): catalogue Claude's stream keep-alive as benign

An uncatalogued `ping` stream frame classified as substantive, so the
fallback wrote a visible "claude · message:stream_event:ping" row, and
since such frames no longer force streamed text out first, a ping
inside a coalescing window landed above the open reasoning row. A ping
is now benign: it writes no row.

* fix(native-chat): bound finding the binary by the probe budget, and keep it LRU

Resolving the command's real path and its mtime was awaited before the
budgeted wait, so a slow filesystem could hold a launch indefinitely;
it now counts against the same budget, and running out caches nothing.
The cache is least recently used rather than first written, and holds
32 binary-and-workspace entries rather than 16.

* feat(mobile): collapse reasoning rows the way desktop does

With summaries on, every Claude turn now carries reasoning text, and the
phone drew all of it inline, dimmed, between the prompt and the answer.
Mobile now draws a reasoning row as desktop does: collapsed to "Thought
for N s", "Thought" or "Reasoning", its text mounted only once opened,
and nothing at all while the row is still being written in the live
turn or has no text. The headline and the visibility rule live in one
shared module both clients read, so they cannot drift.

* fix(native-chat): let a failed CLI probe heal instead of latching

A probe killed at its timeout, failing to spawn under a loaded boot, or
printing no version was cached as "no flag" for the binary's life, so
that workspace never got summaries again in that run. Only a version
(either side of the floor) or the CLI's own refusal is kept for good
now; a probe that gave no version is kept for 10 minutes, so a hung CLI
still costs one wait per stretch and a boot-time failure heals.

* docs(native-chat): say exactly what the version probe's env leaves out

* fix(native-chat): record a Codex item's start whatever its first frame carried

The start was recorded only for an item/started that wrote no row, so a
reasoning item that started with text lost it and its completion
claimed no span. Every tracked item/started now records its receipt
time, and the started write carries it too.

* fix(mobile): label the reasoning toggle and give it a full touch target

The toggle now tells a screen reader what it is, "Reasoning: Thought
for 3s", as desktop's prefix does, and reaches a 44 pt target. The
shared English copy stays private to the module that formats it.

* refactor(native-chat): build reasoning rows from one provider-neutral helper

Claude, Codex and the terminal sweeps each built the reasoning row body and
its running/ended stamp themselves. They now share journal-reasoning-row:
blank text journals no row, text is bounded the same way, and an end carries
completedAt only when the host saw it.

* refactor(codex): move the active journal item type into the contracts file

codex-unfinished-item-body imported the type from the settlement module,
which imports values from it.

* test(claude): read the launch PATH the way Windows spells it

* test(claude): compare the probe's PATH to the launch's without Orca's CLI dir

When the CLI's directory also holds node (Linux CI's /usr/local/bin), the
runtime pairing puts that directory first, ahead of the Orca CLI directory the
launch adds, so the launch PATH no longer ends with the probe's. Both still
resolve the same claude and shims. The test now checks that exactly, for a CLI
with and without a sibling node.

* feat(native-chat): lead the reasoning row with a brain glyph in the tool-row column

* fix(native-chat): route the reasoning glyph through the shared icon names, keep its chevron findable, and match it on mobile

* fix(mobile): keep the long-press actions sheet on reasoning rows for Android

* fix(native-chat): forward every exit argument through the thinking-display connection wrapper

* refactor(codex): keep the receipt-timed notification methods with the event they stamp

* feat(native-chat): read an open reasoning block through the one live "Thinking" line

While the agent's open reasoning block has text, the turn's live activity line is its
disclosure: collapsed by default, expandable to the live text (capped and scrollable), and
the block's row draws nothing meanwhile. When the block ends, its row appears in place,
open if the reader opened it live, because the line and the row read one disclosure key.
Which block the line discloses is derived from the line's own render condition, so a row
is never hidden while nothing on screen shows it; any other open block (a subagent's, or
one a prompt pushed off the line) draws as "Reasoning". Desktop and mobile alike; no host
or wire change.

* fix(native-chat): a slot kept for its turn bar or diff rollup no longer draws its message

The transcript row drew the message of every message slot, while the slot builder pushes a slot
for a row it declined to draw whenever that row also carries its turn's bar or diff rollup. So the
open reasoning block the live line discloses still drew as a "Reasoning" row when it was a
provider-opened turn's first row or the last row of a turn that changed files, and one click
opened both. The builder's decision now travels on the slot (`drawsMessage`) and the row draws
only the bar and rollup when it is false; the row-level `folded` guard it made redundant is gone.

* refactor(native-chat): draw the live line from one shared value, with one live region

Desktop and mobile now render the live activity line from one pure function,
`nativeChatLiveLine`: whether it draws, what it says, and the open reasoning block it
discloses with the text it has so far. The open block's row is hidden from that same value,
so desktop no longer restates the line's render condition beside it, and the lines no longer
re-derive the text.

The line keeps one element, and so one live region, through every state; only its trigger
and body come and go, so a screen reader hears "Thinking" and the label after it. On mobile
the live text gets the finished row's Android long press (copy or select through the message
actions sheet), the line's touch target is the row's 44 pt, and its label and body sit in
the finished row's column so nothing moves when the row takes over.

* fix(mobile): keep the live text's actions sheet on the block it was opened for

On Android the sheet opened by a long press on the live reasoning was a flag gated on a live
block: it vanished when the block ended, mid Select text, and the stale flag reopened it
unprompted on the next block. The sheet now holds the message it was opened for.

* test(native-chat): the reasoning body owns its tone, live and once landed

Rendered QA on a pre-merge build showed the live line's open reasoning in full foreground and the
landed row's in muted text, so it dimmed as the block landed: the body set no colour of its own and
inherited one from wherever it was mounted. Since the main merge (017ad743fa) the body carries the
chat's faint tone itself; this pins that, and that the line and the row draw the same body.

---------

Co-authored-by: Merge Sim <sim@local>
2026-10-06 07:19:24 -07:00
..