Files
orca/docs/reference/agent-pty-transcript-capture.md
T
Jinwoo Hong daf63e659c fix(runtime): read Antigravity, Cline and Prime Agent readiness from the live screen (#24222)
* fix(runtime): decide Antigravity readiness from the live screen

agy paints its composer with cursor addressing, so the line-folded wait
text misses the 1.2.14 accept-edits and plan composers and an ended turn,
while the grid keeps the bare `>` caret painted mid-turn and behind the
/model picker. Read the screen's bottom rows instead: rule, caret, rule,
`? for shortcuts`. A clocked pane is held to quiescence (tier 1b) because
the submit repaint reads ready for a moment; a clockless restored pane
settles from the screen alone. When a trustworthy screen exists it
decides, so the name-only title lane no longer settles an open picker.

Retires the visible-read probe's Antigravity branch: the probe now runs
the shared screen rule for any screen-ruled agent without an output
clock, and keeps its generic empty-pane read for everyone else.

Adds twelve agy 1.2.14 recordings and a replay suite shared by
screen-ruled agents. STA-8741.

* fix(runtime): decide Cline readiness from the live screen

Cline paints its composer box with cursor addressing on the alternate
screen, so no text rule saw it and worker-start timed out at
agent_readiness (#23268). Read the box off the grid: rule, an empty
composer with one of the captured placeholders, rule, the Plan/Act row
and the auto-approve row, with no braille spinner above it.

A streaming reply repaints the same box once its spinner has scrolled
away, so Cline is tier 1b only: a clocked pane waits for quiet and a
clockless one never settles from the screen. The screen now decides for
a Cline pane, which shuts the quiet-process lane that would have settled
its unworded tool-approval prompt and the Cline Desktop promo.

readLiveTerminalScreenLines now returns raw rows: the read projection
blanks a composer it takes for a draft, and it takes Cline's
placeholder for one, so a typed draft and an empty composer looked the
same.

Adds nine cline 3.0.66 recordings (macOS) and the 3.0.65 Windows capture
from #23269. STA-8741.

* fix(runtime): decide Prime Agent readiness from the live screen

Prime redraws its composer on the alternate screen, so the text tail
never showed a settled prompt and tui-idle timed out (#22153). Read the
grid instead: a bare `>` directly over the `<- manage` footer, with no
braille status row (`Writing - 6s`) above it. The footer and caret alone
stay painted for a whole turn.

Replayed chunk by chunk, Prime erases that status row before redrawing
it, and on first launch paints the idle composer just before the
trace-sharing question covers it. Both keep repainting, so a clocked
pane is held to quiescence (tier 1b); a clockless restored pane settles
from the screen alone.

Adds nine prime-agent 0.9.8 recordings (isolated HOME, OpenRouter) and
the two 0.9.5 captures from #22154. STA-8741.

* refactor(runtime): drop Cline-only readiness branches

Cline now follows the same pattern as Antigravity and Prime: a screen
rule plus table entries.

- Drop MID_TURN_COMPOSER_AGENTS. onPtyData stamps lastOutputAt on every
  chunk, so a re-attached streaming pane has an output clock from its
  first byte; the exception only guarded a pane that printed nothing
  since attach. A clockless Cline pane now settles from its screen like
  the other two.
- Drop the 'ready-body' rest-signal entries for all three agents. The
  rest signal is read only by quietForegroundLane, and a readable screen
  already shuts that lane and the title lane (isReadinessDecidedByScreen),
  so the entries only removed the quiet-process fallback for a pane with
  no trustworthy grid. The census now checks that screen-shut instead.
- Drop the Cline rule's auto-approve row check; no recorded verdict
  depends on it.

Kept: raw rows from readLiveTerminalScreenLines. Every frame of every
codex-* and qoder-* capture at 120x40, 80x24 and 100x32 gives the same
isKnownReadyPromptBody (with and without a clock) and
isQuietReadyScreenBody verdict through both readers.

Serializer known-failures for the new captures are pre-existing
serializer behaviour, not this branch: row-0 cells restore with a
true-colour background where the source has the default (the DSH
class), and Prime's cursor restores at column 119 instead of the pending
wrap at 120 (the qoder class). STA-8741.

* fix(runtime): trust a screen rule only on the PTY's own grid

Review findings on the screen-ruled readiness (STA-8741):

- A grid out of step with the PTY garbles cursor-addressed chrome, and a
  model resize does not make the TUI repaint. readLiveTerminalScreenLines
  now returns null unless the emulator's grid matches the PTY's reported
  size and was never reflowed without a repaint (a re-attach that learned
  the real size late), so the pre-existing lanes decide there instead of
  timing out.
- The visible-read probe reads the draft-blanking projection, which
  turns Cline's `❯ Ask anything...` into a bare `❯`. It now restores the
  blanked composer row before the rule reads it; `terminal read --screen`
  output is unchanged.
- The quiet lane no longer ORs the text rules over a trustworthy screen
  that refused; without one, tier 1 already ran them. No recorded
  verdict changes.

Tests: ready recordings on a mismatched and on a reflowed grid settle
through the old lanes; the restored-pane probe runs every ready
recording through the real projection; the rest-signal census checks
the lane verdict with and without a screen.

* test(runtime): trim STA-8741 recordings to the screens they prove

* refactor(runtime): one screen verdict for every screen-ruled lane

readScreenRuledReady, readScreenRuledQuietReady and isReadinessDecidedByScreen
each re-derived the same thing: the agent's rule applied to a trustworthy live
screen. They collapse into readScreenRuledVerdict (true / false / null), which
tier 1, tier 1b and the lane gate read.

This also makes a refusal final in tier 1: a clockless pane whose trustworthy
screen refused fell through to the text rules, so retained ready text could
settle over an open picker (Greptile review). The quiet tier already refused
there; now both do.

The tier-1b agent set derives the screen-ruled agents from the rule table
instead of listing them again, and the lane test that repeated the census
case is dropped.

* refactor(runtime): let the visible-read probe read its own output clock

The probe's clock was captured at start and threaded through the wait
dependencies as a one-off parameter. The probe now reads it from the live
record when its screen read returns, which is also the fresher answer.

* fix(runtime): trust a reflowed grid again once a PTY resize repaints it

The reattach-reflow flag was never cleared, so a pane stayed on the old lanes
for the rest of its life even after a real resize made the TUI repaint
(Greptile review). The record now keeps the reflowed grid, and a PTY resize
off that grid clears it; an echo of the same size sends no SIGWINCH and keeps
it.

Tests: the reflow case in every screen-ruled suite now includes a same-size
echo, and an Antigravity recording only the screen reads ready settles after a
resize and repaint.

* refactor(runtime): keep screen-rule trust and raw rows to screen-ruled agents

Two shared changes reached agents this PR does not target: the live
screen reader returned raw rows, and it refused a grid that did not
match the PTY. Both now live in readScreenRuledLines, which only the
screen-ruled agents read (screenReader picks it from the rule table);
readLiveTerminalScreenLines is main's again. The probe keeps main's
Antigravity-banner trigger, so a Codex or unknown pane is probed exactly
as before.

Proof: the non-screen-ruled suites give identical pass sets on this
branch and its base (1,781 tests), and replaying every other recording
frame by frame through the readiness and blocked verdicts, for its
agent and for an unknown pane, gives identical results (93 pairs). A new
test keeps a Codex pane reading its screen when the PTY reports another
grid; it fails if the trust check moves back into the shared reader.
2026-10-01 00:02:39 -04:00

7.7 KiB

Capturing an agent PTY transcript

Orca's readiness and blocked-prompt rules are text rules over what an agent CLI paints on a terminal. They are only as good as the screens they were written against. This is how to record one, byte for byte, so a rule can be pinned to evidence instead of to a remembered screen.

Related: antigravity-readiness-evidence.md names the specific Antigravity transcripts that are still missing and what each one decides.

The recorder

node config/scripts/capture-agent-pty-transcript.mjs --name <fixture-name> [options] -- <command> [args...]

It allocates a real PTY, spawns the agent inside it, mirrors the session to your terminal so you can drive it by hand, and appends every byte it receives to src/main/runtime/__fixtures__/<fixture-name>.txt. It does not strip escapes, fold \r, rewrap lines, or normalise anything — the file is what the terminal received.

  • Ending a capture: press Ctrl+]. The recorder consumes that key and never forwards it, which is the only way to end a capture while a dialog still owns the screen. Quitting the agent instead would first dismiss the dialog you came to record.
  • --cols N --rows M pin the PTY size (default: your terminal's). Wrapping is part of the evidence, so record the size — the sidecar does it for you.
  • --duration S stops unattended after S seconds, for a screen that needs no interaction.
  • --send "<ms>:<text>" types into the PTY at a fixed offset, repeatable, with \r \n \t \e escapes. A dialog capture has to be driven, and an unattended run (CI, or an agent) has no TTY to type into; the keystrokes ride the same PTY a human's would. For example, the committed antigravity-dialog-model-picker.txt was recorded with --duration 24 --send "14000:/model" --send "16000:\r", which leaves the picker owning the screen when the capture stops.
  • --note "<text>" records the account type, plan, model and CLI version in the sidecar.
  • --out <path> writes outside the fixture directory (use it for a first dry run).

Each capture also writes <fixture-name>.meta.json with the timestamp, platform, command, PTY size, note and exit code. Commit it with the transcript; the version and account type behind a screen are not recoverable from the bytes.

Prerequisite: node-pty must be built for plain Node:

node config/scripts/ensure-native-runtime.mjs --runtime=node

Orca itself does not need to be running, and the recorder never touches Orca state.

Platform notes

  • macOS / Linux: nothing special. TERM=xterm-256color is set for the child.
  • Windows: run it from Windows Terminal / PowerShell, not a Git Bash (MSYS) pane — MSYS rewrites arguments that start with /, which mangles the cmd.exe /c hand-off. A .cmd or .bat agent shim cannot be spawned by node-pty directly, so the recorder routes those through cmd.exe for you.
  • WSL: capture inside the distro (run the recorder from the distro's checkout). Recording wsl.exe from the Windows side adds the login-shell banner to the transcript.
  • SSH: record on the execution host. A transcript recorded locally is not evidence about what a remote agent prints.

Privacy: scrub before committing

A live agent screen routinely contains things that must not enter git history:

Scrub Why
Account email / sign-in identifier The account row on a ready screen prints it verbatim
Org, tenant or team name Identifies a customer
Machine hostname and OS username Appear in prompts, paths and the OSC title
Absolute home paths (/Users/<you>, C:\Users\<you>) Contain the username
JWTs, AIza… keys, 1//… refresh tokens, Bearer …, sk-…, ghp_… Live credentials; a sign-in screen can echo one
Private repo, branch and ticket names Leak roadmap detail
Anything you pasted into the agent during the capture You typed it; it is in the transcript

The recorder scans the file as soon as the capture ends and prints every hit with a line and column. To scrub:

node config/scripts/capture-agent-pty-transcript.mjs --scan src/main/runtime/__fixtures__/<name>.txt --redact

Redaction replaces each finding with a same-length placeholder (u…u@example.com, XXXX…). Length matters: a transcript's value is its exact wrapping and column alignment, and a shorter replacement reflows the screen and destroys the evidence.

Verify it is gone

  1. node config/scripts/capture-agent-pty-transcript.mjs --scan src/main/runtime/__fixtures__/<name>.txt must print clean and exit 0. It recognises its own placeholders, so a scrubbed file passes.
  2. Grep for the specifics the scanner cannot know: rg -n -i -- "$(whoami)|<your-email>|<your-org>|<your-hostname>" src/main/runtime/__fixtures__/<name>.txt
  3. Read it once with escapes visible: LC_ALL=C cat -v src/main/runtime/__fixtures__/<name>.txt. The scanner matches shapes; only a human catches a project name.
  4. Check the sidecar too — --note text is free-form and is committed.

config/scripts/pty-transcript-secret-scan.test.mjs re-scans every committed __fixtures__/*.txt, so a transcript that skips step 1 fails the suite.

Consuming a transcript in a test

Feed the raw bytes through the runtime rather than into a matcher directly: escape handling, tail retention and title tracking all live in onPtyData, and a rule tested on pre-normalised text is tested on something no pane ever sees.

src/main/runtime/agent-transcript-pane-test-harness.ts builds the pane; src/main/runtime/terminal-interactive-wait-visibility.test.ts (cursor-agent) and src/main/runtime/antigravity-readiness-transcripts.test.ts (Antigravity) are examples. An agent whose readiness is read off the live screen gets its suite from src/main/runtime/screen-ruled-agent-transcript-suite.ts.

Worked example: the Antigravity captures

The six committed antigravity-*.txt fixtures were recorded this way on macOS against agy 1.1.25. Two points generalise:

  • Reach a state without mutating the operator's config. The ready-screen captures ran in a directory the CLI already trusted, so no trust answer was written. Where a dialog could only be reached by signing the operator out or deleting their settings, it was left uncaptured and recorded as such rather than forced.
  • An environment variable is a legitimate capture knob where a setting is not. AGY_CLI_HIDE_ACCOUNT_INFO=1 produced a second ready screen with no account row, which is evidence no amount of reasoning about the first screen could have supplied. It changes nothing on disk.

Known gap in the existing captures

The three cursor-agent-*.txt fixtures contain no escape bytes and no carriage returns. Whatever produced them went through a renderer and a clipboard, so they preserve wording and box-drawing glyphs but not the caret, the cursor moves, the repaints, or whether the CLI uses the alternate screen buffer. They are good enough for the wording-based rules built on them and are not evidence for anything else. New captures made with this recorder keep those bytes; the Antigravity scaffold asserts their presence so a pasted screen cannot pass as a capture.