mirror of
https://github.com/stablyai/orca.git
synced 2026-10-03 08:02:12 +00:00
* fix(runtime): decide Antigravity readiness from the live screen agy paints its composer with cursor addressing, so the line-folded wait text misses the 1.2.14 accept-edits and plan composers and an ended turn, while the grid keeps the bare `>` caret painted mid-turn and behind the /model picker. Read the screen's bottom rows instead: rule, caret, rule, `? for shortcuts`. A clocked pane is held to quiescence (tier 1b) because the submit repaint reads ready for a moment; a clockless restored pane settles from the screen alone. When a trustworthy screen exists it decides, so the name-only title lane no longer settles an open picker. Retires the visible-read probe's Antigravity branch: the probe now runs the shared screen rule for any screen-ruled agent without an output clock, and keeps its generic empty-pane read for everyone else. Adds twelve agy 1.2.14 recordings and a replay suite shared by screen-ruled agents. STA-8741. * fix(runtime): decide Cline readiness from the live screen Cline paints its composer box with cursor addressing on the alternate screen, so no text rule saw it and worker-start timed out at agent_readiness (#23268). Read the box off the grid: rule, an empty composer with one of the captured placeholders, rule, the Plan/Act row and the auto-approve row, with no braille spinner above it. A streaming reply repaints the same box once its spinner has scrolled away, so Cline is tier 1b only: a clocked pane waits for quiet and a clockless one never settles from the screen. The screen now decides for a Cline pane, which shuts the quiet-process lane that would have settled its unworded tool-approval prompt and the Cline Desktop promo. readLiveTerminalScreenLines now returns raw rows: the read projection blanks a composer it takes for a draft, and it takes Cline's placeholder for one, so a typed draft and an empty composer looked the same. Adds nine cline 3.0.66 recordings (macOS) and the 3.0.65 Windows capture from #23269. STA-8741. * fix(runtime): decide Prime Agent readiness from the live screen Prime redraws its composer on the alternate screen, so the text tail never showed a settled prompt and tui-idle timed out (#22153). Read the grid instead: a bare `>` directly over the `<- manage` footer, with no braille status row (`Writing - 6s`) above it. The footer and caret alone stay painted for a whole turn. Replayed chunk by chunk, Prime erases that status row before redrawing it, and on first launch paints the idle composer just before the trace-sharing question covers it. Both keep repainting, so a clocked pane is held to quiescence (tier 1b); a clockless restored pane settles from the screen alone. Adds nine prime-agent 0.9.8 recordings (isolated HOME, OpenRouter) and the two 0.9.5 captures from #22154. STA-8741. * refactor(runtime): drop Cline-only readiness branches Cline now follows the same pattern as Antigravity and Prime: a screen rule plus table entries. - Drop MID_TURN_COMPOSER_AGENTS. onPtyData stamps lastOutputAt on every chunk, so a re-attached streaming pane has an output clock from its first byte; the exception only guarded a pane that printed nothing since attach. A clockless Cline pane now settles from its screen like the other two. - Drop the 'ready-body' rest-signal entries for all three agents. The rest signal is read only by quietForegroundLane, and a readable screen already shuts that lane and the title lane (isReadinessDecidedByScreen), so the entries only removed the quiet-process fallback for a pane with no trustworthy grid. The census now checks that screen-shut instead. - Drop the Cline rule's auto-approve row check; no recorded verdict depends on it. Kept: raw rows from readLiveTerminalScreenLines. Every frame of every codex-* and qoder-* capture at 120x40, 80x24 and 100x32 gives the same isKnownReadyPromptBody (with and without a clock) and isQuietReadyScreenBody verdict through both readers. Serializer known-failures for the new captures are pre-existing serializer behaviour, not this branch: row-0 cells restore with a true-colour background where the source has the default (the DSH class), and Prime's cursor restores at column 119 instead of the pending wrap at 120 (the qoder class). STA-8741. * fix(runtime): trust a screen rule only on the PTY's own grid Review findings on the screen-ruled readiness (STA-8741): - A grid out of step with the PTY garbles cursor-addressed chrome, and a model resize does not make the TUI repaint. readLiveTerminalScreenLines now returns null unless the emulator's grid matches the PTY's reported size and was never reflowed without a repaint (a re-attach that learned the real size late), so the pre-existing lanes decide there instead of timing out. - The visible-read probe reads the draft-blanking projection, which turns Cline's `❯ Ask anything...` into a bare `❯`. It now restores the blanked composer row before the rule reads it; `terminal read --screen` output is unchanged. - The quiet lane no longer ORs the text rules over a trustworthy screen that refused; without one, tier 1 already ran them. No recorded verdict changes. Tests: ready recordings on a mismatched and on a reflowed grid settle through the old lanes; the restored-pane probe runs every ready recording through the real projection; the rest-signal census checks the lane verdict with and without a screen. * test(runtime): trim STA-8741 recordings to the screens they prove * refactor(runtime): one screen verdict for every screen-ruled lane readScreenRuledReady, readScreenRuledQuietReady and isReadinessDecidedByScreen each re-derived the same thing: the agent's rule applied to a trustworthy live screen. They collapse into readScreenRuledVerdict (true / false / null), which tier 1, tier 1b and the lane gate read. This also makes a refusal final in tier 1: a clockless pane whose trustworthy screen refused fell through to the text rules, so retained ready text could settle over an open picker (Greptile review). The quiet tier already refused there; now both do. The tier-1b agent set derives the screen-ruled agents from the rule table instead of listing them again, and the lane test that repeated the census case is dropped. * refactor(runtime): let the visible-read probe read its own output clock The probe's clock was captured at start and threaded through the wait dependencies as a one-off parameter. The probe now reads it from the live record when its screen read returns, which is also the fresher answer. * fix(runtime): trust a reflowed grid again once a PTY resize repaints it The reattach-reflow flag was never cleared, so a pane stayed on the old lanes for the rest of its life even after a real resize made the TUI repaint (Greptile review). The record now keeps the reflowed grid, and a PTY resize off that grid clears it; an echo of the same size sends no SIGWINCH and keeps it. Tests: the reflow case in every screen-ruled suite now includes a same-size echo, and an Antigravity recording only the screen reads ready settles after a resize and repaint. * refactor(runtime): keep screen-rule trust and raw rows to screen-ruled agents Two shared changes reached agents this PR does not target: the live screen reader returned raw rows, and it refused a grid that did not match the PTY. Both now live in readScreenRuledLines, which only the screen-ruled agents read (screenReader picks it from the rule table); readLiveTerminalScreenLines is main's again. The probe keeps main's Antigravity-banner trigger, so a Codex or unknown pane is probed exactly as before. Proof: the non-screen-ruled suites give identical pass sets on this branch and its base (1,781 tests), and replaying every other recording frame by frame through the readiness and blocked verdicts, for its agent and for an unknown pane, gives identical results (93 pairs). A new test keeps a Codex pane reading its screen when the PTY reports another grid; it fails if the trust check moves back into the shared reader.