* fix(terminal): fill DOM block glyphs only in repainted rows
Adapt the block-fill approach from PR #15955 and bound painting to xterm render ranges without observer or animation-frame rescans.
Co-authored-by: mmarabel <mmarabel@users.noreply.github.com>
* fix(terminal): preserve block fills across DOM row replacements
---------
Co-authored-by: mmarabel <mmarabel@users.noreply.github.com>
* fix(cursor): preserve Windows hook input across profile paths
Adapt the reviewed direct-command approach from PR #24381, and set UTF-8 input/output encoding for profiles requiring PowerShell.
Co-authored-by: Vladimir Kurgansky <vladimir.kurgansky@gmail.com>
* fix(cursor): preserve missing-script permission replies
Keep the established encoded launcher for every event and explicitly use UTF-8 input/output. Preserve original missing-script acceptance and verify ordinary/spaced profiles rather than adding shell-specific direct guards.
---------
Co-authored-by: Vladimir Kurgansky <vladimir.kurgansky@gmail.com>
* Let scheduled CI warmers wait and measure WebRTC startup
* Measure a smaller daemon shutdown fixture image
* Counterbalance WebRTC startup and verify retained fixture files
* Record CI fixture measurements and remove temporary pilots
* Clarify fixture build dependency cleanup evidence
* Make coalesced snapshot fixture delivery deterministic
* test: type the PTY write delay observer
* test: align source-control fixtures with current store contracts
* Bound E2E package setup and retain cancelled-job traces
* Remove empty passing sentinels from opt-in socket tests
* Make SSH typing pressure fixture readiness and replies observable
* Advertise browser support for anchored terminal placement
* fix(recovery): back off a launch-failed renderer instead of tripping the crash breaker
A renderer that the OS refused to spawn (macOS exit 1003 = LAUNCH_RESULT_FAILURE; field
cause: per-user process limit, posix_spawn EAGAIN) burned the 3-reload crash-loop budget
in ~750ms and raised a "graphics driver" prompt, while the condition lasted minutes.
- launch-failed retries in place on a 250ms..60s backoff (~2 min), outside the breaker;
a loaded document resets it. Other crash reasons keep the breaker.
- Each launch failure records renderer_launch_failed_probe {spawnError} from a cheap
spawn probe, so bundles name EAGAIN/EACCES/ENOENT directly.
- The exhausted prompt says the process limit was hit (probe EAGAIN), drops the
graphics-driver wording, keeps Try Again as default, and offers no Restart:
app.relaunch also needs a free process slot and silently fails without one.
* fix(recovery): skip the launch probe on Windows and probe the prompt once
- Re-check quitting after the prompt's probe; don't re-probe on Copy Commands.
- recordRendererLaunchFailureProbe never rejects (breadcrumb write guarded).
- Windows: no spawn probe; a child per failed launch is the per-operation burst EDR scores.
* test: cover quitting and duplicate renderer launch failures
* test: use typed access in PTY delay regression fixture
* fix: scope extended launch retries to POSIX hosts
* test: cover launch probe behavior on native Windows
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
* ci: avoid unrelated headless server qualification
* ci: skip headless detection for ineligible draft PRs
* ci: preserve cross-host qualification and skip supplied prerequisites
* ci: include Windows server cache validation in change detection
* fix(recovery): ask instead of reloading into a repeat Windows OOM with exhausted commit
When another program exhausts Windows commit (RAM + page file), the renderer
OOMs, Orca auto-reloads 250 ms later, and the new renderer OOMs again within
seconds (launch 13084: 3.5 s after the reload; launch 22912: 34 s). The crash-loop
breaker (3 in 60 s) never opens for this cadence, so the user is never told the
machine is out of memory.
Keep the first automatic reload, but when a win32 reason=oom death follows
another OOM within 5 minutes and the pre-gone host sample shows under 512 MB of
available commit, escalate to the existing recovery prompt with a new
'low-commit' cause that names the MB left and suggests closing apps or growing
the page file. Records renderer_recovery_low_commit_prompt. No-op on
macOS/Linux and when commit is healthy.
* fix(recovery): gate low-commit prompt on post-OOM readings and recovered deaths only
- Reject pre-gone samples taken at or before the previous OOM; they miss the commit that corpse released.
- Record an OOM for the repeat window only once recovery actually runs, so skipped teardown OOMs cannot suppress the next first reload.
- Skip the install-ACL diagnosis on the low-commit prompt, whose text would not explain Copy Commands.
* fix(recovery): read commit at gone time when no sampler tick followed the previous OOM
The 10 s pre-gone sampler lands between ~3.5 s repeat OOMs only ~35% of the time, so the gate usually fell back to a silent reload. A gone-time read can only over-report free commit (the corpse already released its pages), so it can miss a prompt but never raise a false one.
* fix: reject invalid low-commit readings and clarify recovery advice
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
* fix(linux): move Chromium shared memory off a tiny /dev/shm
Containers such as GitHub Codespaces mount a 64 MB /dev/shm by default.
When it fills, Chromium aborts the renderer (IMMEDIATE_CRASH, SIGILL/SIGTRAP)
on every reload, trapping Orca in a renderer crash loop. On Linux, stat
/dev/shm before app ready and append --disable-dev-shm-usage when it is under
512 MB or unreadable (ORCA_DEV_SHM=off|force overrides). Records a
dev_shm_policy crash breadcrumb so future container crashes are diagnosable.
* docs(linux): correct dev-shm hasSwitch rationale
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
* fix(markdown): cap rendered markdown size in preview tab and rich override
The Open Preview tab rendered any file synchronously through react-markdown
with no size check, and 'Open anyway' removed the rich-editor cap entirely.
A 2 MB file blocked the renderer 7.8 s at 2.3 GB (5 MB: 38 s, 4 GB), matching
scan29 crash reports where users killed a frozen Orca after opening a large .md.
- Preview tab: over 600 KB shows a 'Render anyway' gate instead of rendering.
- Both overrides stop at a 1 MiB hard cap; above it no render is offered.
* fix(markdown): gate diff-tab markdown preview by size and localize gate copy
The single-file diff Preview toggle sent the whole modified side to
MarkdownPreview with only the 6 MB large-diff limit, so a multi-MB .md diff
still froze the renderer. Wrap it in MarkdownPreviewSizeGate keyed by the
diff tab id, and add the gate's strings to all locale catalogs.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
* fix(terminal): stop parking pass queueing no-op renders during pane-close bursts
Removing an active worktree with many terminal panes closed each pane in its
own commit. Every commit re-ran the parking pass, whose functional setters
queued a render even when the sets were unchanged, so each commit left work
pending and React's nested-update counter climbed past 50 (#185), taking down
the terminal workbench boundary. Dispatch only when a set actually changes.
* fix: initialize parking state mirror lazily and verify transitions
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
* fix(native-chat): an older Orca skips and keeps a journal row of a kind it does not know
* test(native-chat): a newer build's journal row kind survives reads, writes, rewinds and reopens
* fix(native-chat): an older Orca keeps an unknown journal row kind read-only unless its writer declared it skippable
A row of a kind this build does not know, in a well-formed envelope, now latches the chat
read-only with every row kept, the same way a newer row version does. It is read past only
when its writer declared `ifUnknown` on the row: `skip` (a rewind drops it) or `carry` (a
rewind carries it after the rebuilt history, epoch, seq and fence restamped). Every existing
kind changes queue or turn state, so skipping by default would let an older build write from
a wrong fold.
- journal-row-kind-compatibility.ts: each kind states how older builds read it, typed over
every row kind, so a new kind cannot be added without a declaration.
- Rewind restates the Resume and Stop as before, then carries `carry` rows in source order;
the restatement goes back to { lifted, liveStop }.
- Replay treats a row whose body names another sequence than its stored key as malformed at
the key, so the next write never collides with it; catch-up reads stop there too.
* refactor(native-chat): drop the writer opt-in; an unknown journal row kind only latches read-only
An older Orca now treats a row of a kind it does not know exactly like a row from a newer
schema version: every row stays on disk and the chat opens read-only until an update. The
writer-declared skip/carry opt-in, its in-memory placeholder, the carry through rewinds and
the per-kind registry are removed: no current or planned kind could use them, and they can
come with the first kind that may safely be read past.
Kept: an unknown kind needs the envelope every row keeps (epoch, sequence, fence, timestamp),
else it is damage as before; a row whose body names another sequence than its stored key is
malformed at the key; the epoch row's validator names its kind. The schema header states the
rule for adding a kind: keep the envelope, and either ship the reader first or bump `v`.
* refactor(native-chat): derive the journal's known row kinds from the row union
Each kind's own-field check now lives in one table keyed by every kind JournalRow holds, and
the set of kinds this build knows is derived from that table. A kind added to the union without
a check fails to compile, rather than latching this build's own chats read-only as a newer
build's kind. A test reads one valid row of every kind.
* fix(worktrees): remove repeated scans and keep prepared checkouts fresh
* fix(worktrees): reclaim unlocked fallback preparations safely
* refactor(worktrees): simplify creation ownership and idle maintenance
* fix(git): keep ref maintenance armed after an index-only pass
An idle attempt that found the pack index due but refs still cooling down
returned without rescheduling, so loose refs from the arming fetch waited
for the next write instead of the ref cooldown.
On an SSH host without a C/C++ compiler, the relay installs without node-pty, and opening a terminal used to say "could not establish why … reconnect to retry", which never helped. The relay now treats a missing node-pty folder ("not found" only) as not installed, runs its existing build-tools check, and names the missing tools with the install command for the host's package manager. It diagnoses the node-pty install the relay's import actually resolved, and keeps any other error as "can't tell".
Part of #20386. Removing the need for a compiler on the host is #1693.
Read-only export of a direct-SSH target's catalog and dormant state into the signed T6-7 manifest: repositories, folder workspaces and their project groups, worktree metadata and lineage, sparse presets, retired worktree names, the workspace session with bounded scrollback snapshots, automations, and client routing. Reads go through the profile-state Store via a read-only OrcadSourceExportPersistence domain; nothing retires the source. Adds the export-aware migration preflight on top of T6-5's dependents census, a resumable snapshot transfer driver with injected destination operations, and destination-side chunk staging keyed to a caller-supplied staged manifest. Lands the P7/P9 holds: session-owner projection hooks, syncDirectoryDurablySync and the durable-write mode, scrollback path and stored-bytes exports, retained refs, and dormant-tab buffer preservation. Inert until T6-10.
Co-authored-by: m4air <m4air@Mac.localdomain>
* fix(codex): install Codex's Interrupt hook so an Esc-cancelled turn settles
Codex 0.150+ fires an Interrupt hook when the user presses Esc on an
approval prompt or mid-tool, and nothing else. Orca did not install it, so
the pane stayed blocked/working until the next prompt.
- Add Interrupt to the managed Codex events and label maps, written with
Codex's 3s cap (a larger value triggers a startup clamp warning).
- Hash the timeout Codex hashes (Interrupt is clamped to [1,3], default 1)
so self-computed trust matches Codex; pinned against a real 0.159.3 hash.
- Map a root Interrupt to the existing cancelled-turn record
(markCodexLeadTurnInterrupted), keeping child work in the fold; a
child-scoped Interrupt is ignored. Relayed rows take the same path.
* test(runtime): add a readiness census pinning every tui-idle verdict
Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.
Refs STA-9098
* test(runtime): pin the census quiet probes to literal windows
A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.
Refs STA-9098
* test(runtime): say which census probe writes runtime state
Refs STA-9098
* refactor(codex): let the hook builder own Codex's per-event timeout
The managed hook's timeout is now Codex's own normalization of the shared
budget, and every installer derives its trust entry from the hook it wrote,
so no installer repeats the Interrupt special case.
Claude-Session: codex-interrupt-hook review
* refactor(codex): route Interrupt through the Stop lead update with an outcome
Interrupt now writes the lead record through the same setCodexMainAgentTurnState
call as Stop, so markCodexLeadTurnInterrupted keeps its original signature.
Drops the child-scoped Interrupt guard: Codex never runs Interrupt hooks for
subagents and its input schema has no agent_id.
Claude-Session: codex-interrupt-hook review
* test(runtime): observe the census through settled panes and caller-visible waits
- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.
* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix
Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.
* test(runtime): read the census baseline field without Reflect.get
The anti-slop lint rejects Reflect.get on parsed input.
* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files
Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.
The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes
Refs STA-9098
* fix(runtime): refuse rule patterns that repeat an optional or alternating group
The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.
* refactor(runtime): give agent state rules and text anchors one when/answer shape
Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.
- Cursor's prompt is two anchors answering working and idle; the one-off
workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.
* docs: point the readiness evidence docs at the agent state rule files
* refactor(runtime): read Codex, Claude, OpenCode, Pi, OMP and Gemini readiness from rule files
The rule engine gains the regions and answers these agents need, as closed-list entries:
- rule regions `title` (the classified title status) and `text` (one of the file's idle text
anchors, settled), and a `predicate` form of the screen region for named engine scans;
- `withoutClock: skip` for strong quiet rules a clockless pane must not believe;
- anchors (renamed from textAnchors) gain a `title` region, and `live` and `hold` answers;
- `profile.screenSource` (trusted grid or live screen), and an `unknown-pane` file for panes
with no known agent.
Codex's header, composer and provisional-startup checks become named predicates referenced
from codex.json; its ready header, header and startup hold become shared text anchors. Native
idle title markers become shared title anchors; name-only title handling becomes each agent's
idle-title rule. The agent-specific branches in terminal-wait-detection.ts and
tui-idle-evidence.ts are deleted, and the "later live prompt cancels a blocker" rule now reads
only rule-file anchors (plus Muse, which moves in part b2).
No behaviour change: the readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the rule engine's title, text and predicate regions and the bundled anchors
Refs STA-9098
* fix(runtime): reject a rule file that repeats an anchor or rule id
A text rule names its anchor by id, so a repeated id let a file pass validation and then throw
while compiling. Also states that engineVersion bumps once a version ships; version 1 is still
being defined.
* refactor(runtime): fold the working anchor answer into live
The engine treated an anchor's working and live answers identically: both mark a live prompt
that cancels an earlier blocker and settles nothing. Cursor's busy prompt now answers live, so
anchors have one non-settling prompt answer.
Refs STA-9098
* refactor(runtime): read the shared π title anchor from pi.json alone
Pi and OMP paint the same `π - <session>` rest title, and title anchors apply to every pane,
so one copy covers both.
Refs STA-9098
* refactor(runtime): key every rule file and read the trusted screen from screenSource alone
readsTrustedScreen no longer also asks for a screen rule (every trusted file has one, and the
schema requires screenSource where it matters), so rule-less files need no filter. A rule's
match is a plain boolean, and compileTitleAnchors is module-private.
Refs STA-9098
* test(runtime): pin that a clocked Codex pane takes no other agent's ready text
No test failed when holdsReadyTextToQuiet was removed; this one does.
Refs STA-9098
* fix(agent-hooks): keep an OMP approval wait until omp resolves it
omp posts tool_execution_start a few milliseconds after
tool_approval_requested, while its Approve/Deny select still holds the
human. Both mapped onto the pane row, so the working event overwrote the
blocked one and the pane read as busy for the whole prompt.
A working event now leaves an OMP approval wait in place; only
tool_approval_resolved or a new turn ends it. An ask row is unchanged:
its own tool_execution_end ends it. The test replays the order a live
omp 17 run posted for a denied bash call.
Refs STA-9100
* refactor(runtime): select the fresh hook row on any of a terminal's handles or pane keys
selectFreshExplicitAgentStatus matched one handle and one pane key and
returned only the mapped status. The row selection now takes sets of
handles and pane keys, an optional received-at floor, and returns the
row itself, so a reader can see the main agent's own state. The old
function keeps its signature and result on top of it.
Refs STA-9100
* feat(runtime): let tui-idle read hook state for agents whose hooks cover the whole turn
tui-idle read no hook state. Hook state reached readiness only through
the `<Agent> ready` titles the window writes, so a headless `orca serve`
never saw it (#16095), and Codex settled only once its screen had been
quiet for three seconds.
Rule files gain `profile.hooks: "authoritative" | "identity-only"`,
defaulting to identity-only. Codex (with its Interrupt hook), OpenCode,
OpenCode 2, Pi and OMP are authoritative. For them a fresh hook-store
row decides ahead of every other lane:
- the main agent's turn decides (`mainAgent.state` when published), so a
subagent's Stop does not end the lead turn: done settles strong,
working holds, a permission wait never settles;
- the tail's blocked text goes through the existing permission arbiter
with the turn as its explicit status, so a denied prompt's dialog left
in the tail no longer blocks a turn the hook says ended;
- the row joins on every pane key and terminal handle the PTY owns.
No row, a stale, restored or other agent's row, a session-start done,
and a row from before a PTY respawn all fall back to today's lanes. That
keeps startup on the screen and text rules: Codex posts SessionStart
only with the first prompt. Claude, Cursor, Gemini and the rest stay
identity-only.
The readiness census has no hook server, so its frames are unchanged.
Refs STA-9100
* docs(agent-status): record readiness as a reader of the hook store
Refs STA-9100
* fix(runtime): ignore a hook done older than the latest input Orca wrote
A finished turn leaves a fresh `done` row. A caller that sends the next
prompt and waits at once could settle on it before the new turn's first
hook arrives, so the wait returned while the agent was starting work.
Orca's own input writes (terminal send, agent prompts, mailbox pointers)
now stamp a per-PTY input clock, and the hook lane reads no `done`
received before it; the pane falls back to the screen and text rules
until the agent reports again. A `working` row is unaffected.
Refs STA-9100
* docs(agent-status): note the input floor on the hook lane's done
Refs STA-9100
* fix(runtime): take the hook lane's input floor from the PTY run's input record
The hook lane ignored a done older than Orca's latest write to the pane, kept in a
new per-PTY map stamped by a wrapper threaded through four write sites. The PTY
run register already sits on both write funnels, so it now records the last
input (launch writes included, terminal replies not) and the lane reads it.
Keys the user types now count too, which closes the restart-in-the-same-shell
gap: typing `codex` to relaunch no longer lets the previous process's done read
ready while the new one boots.
The respawn floor moves from the shared row join into the lane, beside the
input floor; the freshest row predates a floor exactly when every row does.
* test(runtime): drop runtime hook-lane cases the unit suite already proves
Working over a ready title, a permission wait, and an identity-only agent are
decided inside evaluateTuiIdle and covered there; the runtime suite keeps the
wiring: the join, both floors, Pi's own OSC 133 markers and the arbiter.
* fix(runtime): record a PTY's last input even when main adopted it without a spawn commit
A materialized pane re-adopted by the renderer returns before the spawn-commit
site, so it had no run record and its input never moved the hook lane's floor.
The last input now lives beside the run records: any PTY's input counts, and a
new process's commit still clears it.
* fix(runtime): keep a running process's input time when main reattaches or adopts it
A reattach or adoption commit without an incarnation id cleared the PTY's
last-input time, so a prompt sent just before an SSH adoption was forgotten
and the hook lane could accept the previous turn's done as ready. Only a new
process (or a reattach naming a different incarnation) now starts clean; the
first-input fact follows the same rule.
* docs(runtime): say why a lead turn that ended reads ready while a subagent runs
* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail
A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.
* refactor(runtime): state Codex's provisional startup and title anchors as plain rules
The provisional-startup hold becomes a lastOf anchor with an all/none test, so
its TypeScript scan goes. Title anchors drop their status field (every caller
already gates on an idle title), and withoutClock keeps only the value a rule
can set.
* fix(runtime): leave Codex readiness to its title and screen rules
Codex before its Interrupt hook posts nothing for an Esc mid-turn, so its hook
row stays working and a hook-authoritative tui-idle wait hangs until the row
goes stale. Current Codex already settles fast through its ready title.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* test(runtime): add a readiness census pinning every tui-idle verdict
Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.
Refs STA-9098
* test(runtime): pin the census quiet probes to literal windows
A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.
Refs STA-9098
* test(runtime): say which census probe writes runtime state
Refs STA-9098
* test(runtime): observe the census through settled panes and caller-visible waits
- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.
* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix
Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.
* test(runtime): read the census baseline field without Reflect.get
The anti-slop lint rejects Reflect.get on parsed input.
* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files
Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.
The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes
Refs STA-9098
* fix(runtime): refuse rule patterns that repeat an optional or alternating group
The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.
* refactor(runtime): give agent state rules and text anchors one when/answer shape
Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.
- Cursor's prompt is two anchors answering working and idle; the one-off
workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.
* docs: point the readiness evidence docs at the agent state rule files
* refactor(runtime): read Codex, Claude, OpenCode, Pi, OMP and Gemini readiness from rule files
The rule engine gains the regions and answers these agents need, as closed-list entries:
- rule regions `title` (the classified title status) and `text` (one of the file's idle text
anchors, settled), and a `predicate` form of the screen region for named engine scans;
- `withoutClock: skip` for strong quiet rules a clockless pane must not believe;
- anchors (renamed from textAnchors) gain a `title` region, and `live` and `hold` answers;
- `profile.screenSource` (trusted grid or live screen), and an `unknown-pane` file for panes
with no known agent.
Codex's header, composer and provisional-startup checks become named predicates referenced
from codex.json; its ready header, header and startup hold become shared text anchors. Native
idle title markers become shared title anchors; name-only title handling becomes each agent's
idle-title rule. The agent-specific branches in terminal-wait-detection.ts and
tui-idle-evidence.ts are deleted, and the "later live prompt cancels a blocker" rule now reads
only rule-file anchors (plus Muse, which moves in part b2).
No behaviour change: the readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the rule engine's title, text and predicate regions and the bundled anchors
Refs STA-9098
* fix(runtime): reject a rule file that repeats an anchor or rule id
A text rule names its anchor by id, so a repeated id let a file pass validation and then throw
while compiling. Also states that engineVersion bumps once a version ships; version 1 is still
being defined.
* refactor(runtime): fold the working anchor answer into live
The engine treated an anchor's working and live answers identically: both mark a live prompt
that cancels an earlier blocker and settles nothing. Cursor's busy prompt now answers live, so
anchors have one non-settling prompt answer.
Refs STA-9098
* refactor(runtime): read the shared π title anchor from pi.json alone
Pi and OMP paint the same `π - <session>` rest title, and title anchors apply to every pane,
so one copy covers both.
Refs STA-9098
* refactor(runtime): key every rule file and read the trusted screen from screenSource alone
readsTrustedScreen no longer also asks for a screen rule (every trusted file has one, and the
schema requires screenSource where it matters), so rule-less files need no filter. A rule's
match is a plain boolean, and compileTitleAnchors is module-private.
Refs STA-9098
* test(runtime): pin that a clocked Codex pane takes no other agent's ready text
No test failed when holdsReadyTextToQuiet was removed; this one does.
Refs STA-9098
* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail
A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.
* refactor(runtime): state Codex's provisional startup and title anchors as plain rules
The provisional-startup hold becomes a lastOf anchor with an all/none test, so
its TypeScript scan goes. Title anchors drop their status field (every caller
already gates on an idle title), and withoutClock keeps only the value a rule
can set.
* Reuse serializer oracle cells and isolate native cache policy
* Preserve native cache post-save paths and record hosted oracle gain
* Record native cache reuse and separate cancel-test startup budget
* ci(cross-version-wire): run the whole directory so no compatibility test is left out
Three cross-version tests ran in no CI job because the job named its files by hand.
Run the directory instead, ratchet that every file kept out of the unit shards
runs in some PR job, and re-run the job when the modules the newly running
tests guard change.
* test(cross-version): give the orchestration downgrade test its siblings' 120 s budget
* ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it
The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are
left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable
workflows those jobs call.
It also only proved that some step names each excluded file, not that the job runs when the file
changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path
trigger matched neither it, its harness nor its subject, so a PR touching only those ran it
nowhere. The check now asserts a change to each excluded file fires a gating job that names it,
and the shell trigger gains those three paths.
* ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver
A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to
orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the
job whose tests guard exactly those contracts. Also corrects the publish/read direction in the
turn-end comment.
* test(cross-version): state why the orchestration downgrade test needs 120 s
* test(ci): glob the unit tree once for the unit-exclusion coverage checks
* fix(claude): an informational note is a warning row in its own words, never a raw frame row
Claude Code shows info, notice and suggestion notes as transcript chrome; only a warning earns a
row. The frame used to fall through to the provider fallback and print its opcode.
* fix(claude): name the real source of an informational warning in comments and tests
The host file reached 302 lines after #24203 and #24311, over the 300-line limit,
so the static-analysis job failed on every commit to main. The five chat-tab
members now come from createStructuredAgentSessionTabSurface in the existing
structured-agent-session-host-tabs module; behaviour is unchanged.
#24203 moved ELECTRON_REMOTE_RUNTIME_CLIENT_CAPABILITIES out of protocol-version.ts into
electron-remote-runtime-client-capabilities.ts. Three files that landed on main while it was in
review (SSH access links and managed orcad server work) still import it from protocol-version.ts,
and a test #24203 added predates main making the structured host's logger option required, so
main's node typecheck fails. Point the three imports at the new module and pass the logger.
* test(runtime): add a readiness census pinning every tui-idle verdict
Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.
Refs STA-9098
* test(runtime): pin the census quiet probes to literal windows
A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.
Refs STA-9098
* test(runtime): say which census probe writes runtime state
Refs STA-9098
* test(runtime): observe the census through settled panes and caller-visible waits
- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.
* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix
Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.
* test(runtime): read the census baseline field without Reflect.get
The anti-slop lint rejects Reflect.get on parsed input.
* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files
Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.
The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes
Refs STA-9098
* fix(runtime): refuse rule patterns that repeat an optional or alternating group
The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.
* refactor(runtime): give agent state rules and text anchors one when/answer shape
Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.
- Cursor's prompt is two anchors answering working and idle; the one-off
workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.
* docs: point the readiness evidence docs at the agent state rule files
* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail
A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.
* fix(native-chat): a host admits structured sessions by client capability, not its own chat setting
A host's experimentalStructuredNativeChat decided whether any paired client could reach
agentSession.* at all, and whether session.tabs.* showed it structured tabs. That setting is the
host user's own launch preference: whether a new agent opens as a chat or a terminal is decided by
whoever launches it. Using it as admission control meant a client whose own preference was
"structured chat" was refused on a host whose preference was "terminal", and chats opened while
the setting was on were withheld from mobile once it was turned off.
The gate now asks one thing: did the client advertise agent-session.structured.v1 (in-process
callers negotiate nothing and are always admitted). Tab projection and restore follow the same
rule. With the setting no longer gating anything, the separate cleanup gate (close, cancel,
unsubscribe, release), which existed only so those kept working after the setting was switched
off, is identical to the main gate and is folded into it. The settings listener that republished
tabs when the setting changed is removed, since projection no longer depends on it.
The host setting still picks the default for launches that start on the host itself
(agent.launch from mobile, orchestration worker-start).
* fix(native-chat): negotiate client-chosen launch mode so released phones and old servers keep terminals
Hosts advertise agent-session.structured.client-launch-mode.v1: they admit
structured sessions by client capability alone. A remote client that does
not advertise it (phones released before agent.launch) asks createSupport
to pick the launch mode, so the host keeps answering that with its own
setting, exactly as before. Cleanup methods keep their own named gate so a
future admission condition cannot make close or cancel refusable.
* chore(native-chat): justify the two type assertions this change's lines touch
* fix(native-chat): chats that already exist keep showing whatever the chat setting says
The structured chat setting decides only what new agents open as. With it
off, this machine's structured chats used to be hidden while the host,
which no longer reads the setting, still reported them to the workspace
activation gate, so a workspace holding only a chat opened empty. The
local chat mirror and its startup restore now run whatever the setting
says, the continue-after-restart offer follows the chats that exist, and
the setting's copy says it applies to new agents.
* test(native-chat): pin that a host advertises the client-chosen launch mode
* fix(native-chat): mirror this machine's chats only where it holds them
Round 1 ran the local chat mirror for everyone so existing chats show
whatever the setting says. That gave every desktop a permanent
session-tabs listener, which turns on the runtime's phone replication
paths, plus two full session-tab censuses at startup, and made the
browser client mirror its remote host a second time.
The runtime now says whether it holds structured chats: its structured
host is built only when saved chats were restored at startup or a client
created one here, and it announces the moment one is built. The mirror,
the startup restore and the continue-after-restart offer run only when
the setting launches chats or the host holds some, and never in the
browser client. A chat a paired client creates here with the setting off
still appears at once. The chat behaviour settings show wherever chats
exist, and the setting's copy says it picks what new agents open as. The
toggle-off teardown this made dead is removed.
* test(native-chat): record install listeners without a cast
* fix(native-chat): mirror this machine's chats only once it holds one, not once its host is built
Session history, resume preparation, terminal resume commands and replay-safe phone launches all
build the structured host for users who never had a chat, which turned on the chat mirror and the
structured-only settings rows until the next restart. The signal is now derived from the host's
records (or a records file still owed its import) and pushed when the first chat is restored or
created. A throwing listener no longer fails the install that fired it.
* feat(native-chat): createSupport reports the saved selection a new chat on this host starts with
A chat on a paired server starts with the server's saved model and options, which the desktop could
not read, so its picker showed a guess. createSupport's answer, which the desktop already waits for
before a paired launch, now also carries that seed as a new optional field (older clients ignore it).
Create and createSupport read it through one resolver so they cannot drift.
* refactor(protocol): move the Electron remote client capability list into its own module
Merging main left protocol-version.ts one line over the max-lines limit on this branch. The list of
capabilities the desktop advertises to a paired host moves, unchanged, into
electron-remote-runtime-client-capabilities.ts, the module the next PR in the stack already uses
for it; importers point there.
* test(cross-version): stub the launch seed resolver createSupport now reads
* fix(native-chat): the desktop tells its own host it picks each launch mode, so retrying an existing chat works with the setting off
* docs(native-chat): name the real exit for the released-phone createSupport rule
* test(cross-version): a released client still gets the host-setting createSupport answer; a launch-mode client gets supported plus the seed
expect.poll does not retry a thrown read, so one spurious 'Resulting promise was
garbage collected' from Electron's main evaluate failed the crash-recovery test.