mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 16:02:29 +00:00
c39f5f25fcc9b2a25fa6e189eee19e93cf98fde2
881
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6486624b44 |
docs: reference local orca.yaml and .worktreeinclude
Document the existing workspace configuration, include/copy rules and sharing behavior. Co-authored-by: Neil <neil@stably.ai> |
||
|
|
97fd7edb39 |
fix(editor): recognize Ruby task and configuration files
Add Ruby task/configuration filenames to the existing generated language associations. Co-authored-by: ggbdpq <ggbdpq@gmail.com> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
c792386cdc |
docs(terminal): explain local macOS/Linux shell startup files
Document the actual login/interactive shell startup-file order and existing shell setup behavior. Co-authored-by: brynnclaw <261708852+brynnclaw@users.noreply.github.com> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
f93808dfff |
Collect test-selection evidence when full unit tests fail (#24955)
* Collect advisory unit-selection evidence from failed full runs * Trigger checks after retargeting the evidence fix to main |
||
|
|
30e0ccaba6 | Let retired-cache GC observations settle across the existing six-turn budget (#24967) | ||
|
|
36d008688a | Update README downloads badge | ||
|
|
87f5a98a61 |
Reuse documentation search excerpts during result navigation (#24574)
Memoize the existing pure excerpt renderer by complete raw text for the current result array, retaining original rendering as a miss fallback. |
||
|
|
7539d19416 |
Skip impossible link and tag matches in docs search results (#24646)
* Skip impossible HTML matches when formatting docs search results After the existing link-removal phase, skip the unchanged HTML regex only when the current string has no closing delimiter. * Skip impossible link and tag matches in docs search results Skip the five unchanged link regexes when their protected-code input lacks ]; skip the unchanged HTML regex when its post-link input lacks >. |
||
|
|
8f26bfad22 |
Use the measured pnpm lookup policy automatically in hosted root CI (#24951)
* Select lookup automatically for the measured hosted root-install profile * Record hosted automatic-mode cold cache publication proof |
||
|
|
6e6f651380 |
Avoid repeated pnpm archive downloads in cache producers (#24927)
* Let measured cache producers keep stores without downloading them * Check that restore-only callers do not publish a producer path * Enable the measured producer mode and record hosted comparisons |
||
|
|
9e27050955 |
Add verified native Antigravity Accounts on the owning runtime (#24691)
* Add verified native Antigravity accounts on the owning runtime * Keep Antigravity usage tied to its observed native account * Refuse oversized encrypted Antigravity snapshots before writing * fix(antigravity): localize account heading and search terms |
||
|
|
1d5d02e396 |
Resolve explicitly configured command aliases for workers (#24648)
* feat(orchestration): resolve explicitly configured command aliases * docs(orchestration): explain configured command aliases * fix(orchestration): validate configured aliases with target shell grammar * fix(orchestration): use actual shell and refuse assignment-only aliases * test: preserve typed calls in configured worker target checks Replace Reflect.apply with the existing typed prototype call pattern so the unchanged regression cases pass the anti-slop lint gate. |
||
|
|
75f8f34ce3 | Overlap independent Linux headless runtime builds (#24910) | ||
|
|
1241ce1b49 |
Skip slower root package-store restores in macOS PR jobs (#24908)
* Skip slower root package-store restores in macOS PR jobs * Update the companion cache-policy contract |
||
|
|
533446dde6 |
Stop mocked renderer imports from qualifying headless CI (#24902)
* Decouple headless running-work tests from the renderer * Keep the shared running-work probe contract documented |
||
|
|
a824fb74ab |
Reuse the headless detector compiler without installing full dependencies (#24895)
* Reuse the headless detector compiler without full dependency setup * Keep optional compiler-cache saves from failing cache warming |
||
|
|
58cf72d48e | Skip slower root package-store restores in Linux PR jobs (#24896) | ||
|
|
add1c55590 |
Skip slower Windows root package-store restores in CI (#24885)
* Skip slower Windows root package-store restores in CI * Update reviewed mobile dependency-store cache expression |
||
|
|
61836f6026 |
Reduce scheduled CI cache warming to every six hours (#24881)
* Reduce scheduled CI cache warming to every six hours * Document cache warmer recovery interval and measured tradeoff |
||
|
|
cc73c8e1a7 | ci: overlap ARM SSH setup and independent observation waits (#24714) | ||
|
|
ac28e8c85e |
Skip dependency installation for known headless build inputs (#24716)
* ci: defer headless dependency installation until graph analysis is needed * docs: align headless CI rollout with platform and cache policy * test: isolate headless detector output from the parent CI step |
||
|
|
d6d2795da5 |
chore(worktree): include create timing and spare outcome in the workspace create events (#24483)
* chore(worktree): include create timing and spare outcome in the workspace-created event The workspace_created and workspace_create_failed events gain optional, numbers-and-enums-only fields built from what the create already measured: total and per-phase durations, the prepared-checkout hit/miss and miss reason, the execution host (local/WSL/SSH), a worktree count bucket, how many other creates were in flight, whether the repo has a post-checkout hook (file existence only, probed after the create returns), and for a failure the phase it died in plus elapsed time. No new git process runs; consent and opt-out are unchanged. * fix(worktree): attribute failed_phase by error, label WSL-path repos, skip the hook check with telemetry off - failed_phase now names the outermost timed step the thrown error (or its cause) left, so a caught failure or a concurrent sibling step can no longer be misattributed; the old-relay SSH error keeps its cause so it still reads as git_worktree_add. - execution_host follows the same rule Git routing uses, so a \\wsl.localhost repo reads wsl. - The post-checkout hook check does not read the repo when telemetry is disabled. - Privacy page mentions the miss reason code and the failed step. * test(worktree): pin the old-relay SSH add error to git_worktree_add through its cause * fix(worktree): name the create event field sets for their role, and type the old-relay test's caught error * fix(worktree): send create events from runtime creates and record what the spare checkout did Runtime creates (CLI, agents, phone app, paired clients, orchestration, server automations) reuse prepared checkouts like the app's own creates, but recorded no timing and sent no events. Both entry points now start one shared sender (workspace-create-telemetry.ts), so every create sends exactly one event with the same fields, plus create_entry_point (app | runtime). Spare-checkout fields: - concurrent_preparations: peak prepared-checkout builds and background discards running during the create, excluding the one it used; the window closes before the create's own re-arm starts. - prepared_checkout_claim / prepared_checkout_discard phases, so on a miss git_worktree_add minus the prepared_checkout_* phases is the plain checkout. - prepared_checkout_reset (none | base_moved | retargeted) replaces the retargeted flag; prepared_checkout_origin (prefetch | rearm) on hits. - workspace_create_failed carries the spare outcome and its wait. - repo_index_size_bucket from one stat of .git/index in the existing post-create probe (telemetry on, local/WSL only, 2 s cap). * fix(worktree): add spare build and idle time, the re-arm prefetch origin, and a tracked-file count - prepared_checkout_build_ms / prepared_checkout_idle_ms on hits: from arming the spare to ready, and how long it sat ready before the claim (0 when the create waited). readyAt is recorded in the pool's existing ready handler. - prepared_checkout_origin gains rearm_then_prefetch: an automatic re-arm that the dialog prefetch then asked for too, so rearm means the re-arm alone. - repo_file_count_bucket replaces the index byte size: the entry count from the 12-byte index header, which is the same in every index version; left out for a split or sparse index. - The shared sender never lets a failed send change the create's result or error; it logs instead and still ends the create's concurrency membership. - Tests pin the runtime SSH create's timing hand-off and the throwing-send cases on both entry points. * fix(worktree): leave out the file count under any sparse checkout and time spare builds monotonically - The repo probe also reads .git/config.worktree, where git sparse-checkout --sparse-index writes index.sparse, and omits the file count whenever sparse checkout or a sparse index is on in either file, with Git's boolean spellings. core.hooksPath there is honoured too. - prepared_checkout_build_ms / _idle_ms use performance.now(), like every other duration; the build is timed from its own start (buildStartedAt). - The origin field comment names all three values. * fix(worktree): keep the spare's build time on its first build and count worktrees by lock reason - prepared_checkout_build_ms runs from the first build's start (including any wait for the base fetch it is built on) to its first ready; a later tip refresh no longer restarts it, though it still counts as new preparation work. prepared_checkout_idle_ms runs from the latest ready (build or refresh) to the claim. - The worktree count reads each .git/worktrees entry's locked file and leaves out entries whose lock reason names an Orca preparation, the way the listing does, instead of subtracting this process's spares. That covers spares from other processes, crash leftovers and spares being discarded, and cannot run one low while a spare's admin dir does not exist yet. It has its own 1.5 s cap inside the probe. |
||
|
|
9503f9de34 |
fix(runtime): settle Codex tui-idle waits on its hook done, leaving working rows to the rules (#24541)
A headless orca serve has no window to write the Codex ready title, so a Codex tui-idle wait settled only after three quiet seconds. Rule files gain profile.hooks: "turn-end": a fresh hook done settles the wait, while a working or permission row leaves the decision to the rules, so a Codex whose Esc posts no event (before its Interrupt hook) cannot hang the wait. Codex moves from identity-only to turn-end. |
||
|
|
1fc24d3114 | Update README downloads badge | ||
|
|
76b1a90ff6 |
chore(deps): update reviewed dependencies across Orca (#24561)
* chore(deps): update reviewed desktop dependencies and tooling * chore(deps): update compatible mobile packages and Fastlane * chore(deps): update cloud transports and enforce release age * chore(deps): patch documentation dependencies and record review * chore: remove dependency review reports * test(linear): smoke-load resolved SDK through CommonJS loader * fix(deps): keep native rebuilds from reinstalling addon dependencies * fix(native): invoke installed node-gyp directly for Node rebuilds * test(cloud): exclude observer probes from row-lock timing budget * test(mobile): preserve the CSS writer receiver in viewport spy * test(native): remove obsolete batch-shim fixture exception * Stream native rebuild output through the process wrapper |
||
|
|
cd8d03bc06 | fix(dsh): recognize 0.2 profiles and open workspace composer (#24589) | ||
|
|
13ecf051c3 |
Reuse prepared Windows native builds in SSH CI (#24555)
* ci: reuse qualified Windows server slots for SSH host tests * ci: reuse prepared relay addons after an exact native cache hit |
||
|
|
efbf651c7b |
Reduce CI setup costs and fixture failures (#24537)
* Let scheduled CI warmers wait and measure WebRTC startup * Measure a smaller daemon shutdown fixture image * Counterbalance WebRTC startup and verify retained fixture files * Record CI fixture measurements and remove temporary pilots * Clarify fixture build dependency cleanup evidence * Make coalesced snapshot fixture delivery deterministic * test: type the PTY write delay observer |
||
|
|
6153fbcfe4 |
Reduce redundant headless server CI work (#24527)
* ci: avoid unrelated headless server qualification * ci: skip headless detection for ineligible draft PRs * ci: preserve cross-host qualification and skip supplied prerequisites * ci: include Windows server cache validation in change detection |
||
|
|
5f308bfa9c |
revert: take the 26 Phase 3 (#16741 port) PRs back out of main (#24559)
* Revert "feat(orcad): source-side dormant export of a relay-hosted SSH target (#16741 T6-8) (#24519)" This reverts commit |
||
|
|
34ae0933e4 |
fix(worktrees): keep creation fast in large repositories (#24346)
* fix(worktrees): remove repeated scans and keep prepared checkouts fresh * fix(worktrees): reclaim unlocked fallback preparations safely * refactor(worktrees): simplify creation ownership and idle maintenance * fix(git): keep ref maintenance armed after an index-only pass An idle attempt that found the pack index due but refs still cooling down returned without rescheduling, so loose refs from the arming fetch waited for the next write instead of the ref cooldown. |
||
|
|
cdfdadf9ea |
fix(runtime): settle tui-idle on hook state for agents whose hooks cover the whole turn (#24388)
* fix(codex): install Codex's Interrupt hook so an Esc-cancelled turn settles
Codex 0.150+ fires an Interrupt hook when the user presses Esc on an
approval prompt or mid-tool, and nothing else. Orca did not install it, so
the pane stayed blocked/working until the next prompt.
- Add Interrupt to the managed Codex events and label maps, written with
Codex's 3s cap (a larger value triggers a startup clamp warning).
- Hash the timeout Codex hashes (Interrupt is clamped to [1,3], default 1)
so self-computed trust matches Codex; pinned against a real 0.159.3 hash.
- Map a root Interrupt to the existing cancelled-turn record
(markCodexLeadTurnInterrupted), keeping child work in the fold; a
child-scoped Interrupt is ignored. Relayed rows take the same path.
* test(runtime): add a readiness census pinning every tui-idle verdict
Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.
Refs STA-9098
* test(runtime): pin the census quiet probes to literal windows
A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.
Refs STA-9098
* test(runtime): say which census probe writes runtime state
Refs STA-9098
* refactor(codex): let the hook builder own Codex's per-event timeout
The managed hook's timeout is now Codex's own normalization of the shared
budget, and every installer derives its trust entry from the hook it wrote,
so no installer repeats the Interrupt special case.
Claude-Session: codex-interrupt-hook review
* refactor(codex): route Interrupt through the Stop lead update with an outcome
Interrupt now writes the lead record through the same setCodexMainAgentTurnState
call as Stop, so markCodexLeadTurnInterrupted keeps its original signature.
Drops the child-scoped Interrupt guard: Codex never runs Interrupt hooks for
subagents and its input schema has no agent_id.
Claude-Session: codex-interrupt-hook review
* test(runtime): observe the census through settled panes and caller-visible waits
- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.
* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix
Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.
* test(runtime): read the census baseline field without Reflect.get
The anti-slop lint rejects Reflect.get on parsed input.
* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files
Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.
The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes
Refs STA-9098
* fix(runtime): refuse rule patterns that repeat an optional or alternating group
The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.
* refactor(runtime): give agent state rules and text anchors one when/answer shape
Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.
- Cursor's prompt is two anchors answering working and idle; the one-off
workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.
* docs: point the readiness evidence docs at the agent state rule files
* refactor(runtime): read Codex, Claude, OpenCode, Pi, OMP and Gemini readiness from rule files
The rule engine gains the regions and answers these agents need, as closed-list entries:
- rule regions `title` (the classified title status) and `text` (one of the file's idle text
anchors, settled), and a `predicate` form of the screen region for named engine scans;
- `withoutClock: skip` for strong quiet rules a clockless pane must not believe;
- anchors (renamed from textAnchors) gain a `title` region, and `live` and `hold` answers;
- `profile.screenSource` (trusted grid or live screen), and an `unknown-pane` file for panes
with no known agent.
Codex's header, composer and provisional-startup checks become named predicates referenced
from codex.json; its ready header, header and startup hold become shared text anchors. Native
idle title markers become shared title anchors; name-only title handling becomes each agent's
idle-title rule. The agent-specific branches in terminal-wait-detection.ts and
tui-idle-evidence.ts are deleted, and the "later live prompt cancels a blocker" rule now reads
only rule-file anchors (plus Muse, which moves in part b2).
No behaviour change: the readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the rule engine's title, text and predicate regions and the bundled anchors
Refs STA-9098
* fix(runtime): reject a rule file that repeats an anchor or rule id
A text rule names its anchor by id, so a repeated id let a file pass validation and then throw
while compiling. Also states that engineVersion bumps once a version ships; version 1 is still
being defined.
* refactor(runtime): fold the working anchor answer into live
The engine treated an anchor's working and live answers identically: both mark a live prompt
that cancels an earlier blocker and settles nothing. Cursor's busy prompt now answers live, so
anchors have one non-settling prompt answer.
Refs STA-9098
* refactor(runtime): read the shared π title anchor from pi.json alone
Pi and OMP paint the same `π - <session>` rest title, and title anchors apply to every pane,
so one copy covers both.
Refs STA-9098
* refactor(runtime): key every rule file and read the trusted screen from screenSource alone
readsTrustedScreen no longer also asks for a screen rule (every trusted file has one, and the
schema requires screenSource where it matters), so rule-less files need no filter. A rule's
match is a plain boolean, and compileTitleAnchors is module-private.
Refs STA-9098
* test(runtime): pin that a clocked Codex pane takes no other agent's ready text
No test failed when holdsReadyTextToQuiet was removed; this one does.
Refs STA-9098
* fix(agent-hooks): keep an OMP approval wait until omp resolves it
omp posts tool_execution_start a few milliseconds after
tool_approval_requested, while its Approve/Deny select still holds the
human. Both mapped onto the pane row, so the working event overwrote the
blocked one and the pane read as busy for the whole prompt.
A working event now leaves an OMP approval wait in place; only
tool_approval_resolved or a new turn ends it. An ask row is unchanged:
its own tool_execution_end ends it. The test replays the order a live
omp 17 run posted for a denied bash call.
Refs STA-9100
* refactor(runtime): select the fresh hook row on any of a terminal's handles or pane keys
selectFreshExplicitAgentStatus matched one handle and one pane key and
returned only the mapped status. The row selection now takes sets of
handles and pane keys, an optional received-at floor, and returns the
row itself, so a reader can see the main agent's own state. The old
function keeps its signature and result on top of it.
Refs STA-9100
* feat(runtime): let tui-idle read hook state for agents whose hooks cover the whole turn
tui-idle read no hook state. Hook state reached readiness only through
the `<Agent> ready` titles the window writes, so a headless `orca serve`
never saw it (#16095), and Codex settled only once its screen had been
quiet for three seconds.
Rule files gain `profile.hooks: "authoritative" | "identity-only"`,
defaulting to identity-only. Codex (with its Interrupt hook), OpenCode,
OpenCode 2, Pi and OMP are authoritative. For them a fresh hook-store
row decides ahead of every other lane:
- the main agent's turn decides (`mainAgent.state` when published), so a
subagent's Stop does not end the lead turn: done settles strong,
working holds, a permission wait never settles;
- the tail's blocked text goes through the existing permission arbiter
with the turn as its explicit status, so a denied prompt's dialog left
in the tail no longer blocks a turn the hook says ended;
- the row joins on every pane key and terminal handle the PTY owns.
No row, a stale, restored or other agent's row, a session-start done,
and a row from before a PTY respawn all fall back to today's lanes. That
keeps startup on the screen and text rules: Codex posts SessionStart
only with the first prompt. Claude, Cursor, Gemini and the rest stay
identity-only.
The readiness census has no hook server, so its frames are unchanged.
Refs STA-9100
* docs(agent-status): record readiness as a reader of the hook store
Refs STA-9100
* fix(runtime): ignore a hook done older than the latest input Orca wrote
A finished turn leaves a fresh `done` row. A caller that sends the next
prompt and waits at once could settle on it before the new turn's first
hook arrives, so the wait returned while the agent was starting work.
Orca's own input writes (terminal send, agent prompts, mailbox pointers)
now stamp a per-PTY input clock, and the hook lane reads no `done`
received before it; the pane falls back to the screen and text rules
until the agent reports again. A `working` row is unaffected.
Refs STA-9100
* docs(agent-status): note the input floor on the hook lane's done
Refs STA-9100
* fix(runtime): take the hook lane's input floor from the PTY run's input record
The hook lane ignored a done older than Orca's latest write to the pane, kept in a
new per-PTY map stamped by a wrapper threaded through four write sites. The PTY
run register already sits on both write funnels, so it now records the last
input (launch writes included, terminal replies not) and the lane reads it.
Keys the user types now count too, which closes the restart-in-the-same-shell
gap: typing `codex` to relaunch no longer lets the previous process's done read
ready while the new one boots.
The respawn floor moves from the shared row join into the lane, beside the
input floor; the freshest row predates a floor exactly when every row does.
* test(runtime): drop runtime hook-lane cases the unit suite already proves
Working over a ready title, a permission wait, and an identity-only agent are
decided inside evaluateTuiIdle and covered there; the runtime suite keeps the
wiring: the join, both floors, Pi's own OSC 133 markers and the arbiter.
* fix(runtime): record a PTY's last input even when main adopted it without a spawn commit
A materialized pane re-adopted by the renderer returns before the spawn-commit
site, so it had no run record and its input never moved the hook lane's floor.
The last input now lives beside the run records: any PTY's input counts, and a
new process's commit still clears it.
* fix(runtime): keep a running process's input time when main reattaches or adopts it
A reattach or adoption commit without an incarnation id cleared the PTY's
last-input time, so a prompt sent just before an SSH adoption was forgotten
and the hook lane could accept the previous turn's done as ready. Only a new
process (or a reattach naming a different incarnation) now starts clean; the
first-input fact follows the same rule.
* docs(runtime): say why a lead turn that ended reads ready while a subagent runs
* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail
A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.
* refactor(runtime): state Codex's provisional startup and title anchors as plain rules
The provisional-startup hold becomes a lastOf anchor with an all/none test, so
its TypeScript scan goes. Title anchors drop their status field (every caller
already gates on an idle title), and withoutClock keeps only the value a rule
can set.
* fix(runtime): leave Codex readiness to its title and screen rules
Codex before its Interrupt hook posts nothing for an Esc mid-turn, so its hook
row stays working and a hook-authoritative tui-idle wait hangs until the row
goes stale. Current Codex already settles fast through its ready title.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
|
||
|
|
8ff6296bc7 |
Speed up serializer checks and keep native caches stable (#24476)
* Reuse serializer oracle cells and isolate native cache policy * Preserve native cache post-save paths and record hosted oracle gain * Record native cache reuse and separate cancel-test startup budget |
||
|
|
e3cb32791e |
refactor(runtime): read four agents' readiness from JSON rule files through one engine (#24348)
* test(runtime): add a readiness census pinning every tui-idle verdict
Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.
Refs STA-9098
* test(runtime): pin the census quiet probes to literal windows
A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.
Refs STA-9098
* test(runtime): say which census probe writes runtime state
Refs STA-9098
* test(runtime): observe the census through settled panes and caller-visible waits
- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.
* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix
Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.
* test(runtime): read the census baseline field without Reflect.get
The anti-slop lint rejects Reflect.get on parsed input.
* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files
Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.
The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes
Refs STA-9098
* fix(runtime): refuse rule patterns that repeat an optional or alternating group
The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.
* refactor(runtime): give agent state rules and text anchors one when/answer shape
Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.
- Cursor's prompt is two anchors answering working and idle; the one-off
workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.
* docs: point the readiness evidence docs at the agent state rule files
* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail
A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.
|
||
|
|
f69052e113 | Reuse qualified Windows server builds and dependency verification records (#24448) | ||
|
|
43d9b43d3f |
feat(ssh): remote orcad stop by request file and journaled decommission (#16741 T6-4) (#24449)
Clients stop an orcad that advertises health.stopRequests through its slot-local request file and keep SIGTERM for older builds. Decommission runs through the activation journal and fence: it refuses while the terminal census is live or uncounted, stops the instance with an instance-bound managed request, cancels a stop orcad never acted on, and deactivates the record only on proven exit. orcad gains --cancel-managed-stop and an exclusive per-transaction decision file so a cancel can never race a dispatched stop. POSIX-only and inert: no production caller. Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
b093d3ab20 |
feat(orcad): supervisable server: stop requests, managed stop receipts and a lifetime that keeps its lock on failed teardown (#16741 T6-3) (#24433)
orcad stops through slot-local and instance-bound request files, so a reused PID is never signalled. A managed stop is proven by its completion command and recorded as a receipt. Optional daemon retirement is best effort: an idle daemon retires, while a busy or unverifiable one stays up with its admission fence released. Runtime teardown runs in reverse order and keeps the instance lock and profile admission when any writer fails to stop. Browser discovery no longer delays readiness. Legacy worker recovery and watcher children are drained before the final flush. Headless terminal close no longer waits on a renderer tab that does not exist. No production deployment. Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
197ea3a3b3 |
Free PR CI capacity by avoiding repeated setup and real-time test waits (#24355)
* Reduce repeated PR setup and transcript timing waits; add hosted comparisons * Align parallelism contract with Node-only external rebuild toolchain * Record hosted coverage and launch package, store, and cancellation comparisons * Apply hosted Windows setup savings and remove measured test waits * Keep measured PR package gains and remove completed comparison jobs * Report measured test counts with precise units |
||
|
|
198fe72066 | Update README downloads badge | ||
|
|
c6cfcc034e |
refactor(native-chat): structured chat failures always reach the diagnostics log (#24312)
* refactor(native-chat): give the structured chat host one required logger The structured chat runtime took an optional onError callback that the desktop never passed, so a late dispatch settlement, an unanswered-dispatch release, a journal event-sink write and a provider lifecycle delivery that failed were dropped with no trace. Other host failures went to scattered console.warn calls, which reach nothing in a packaged desktop build. The runtime and host now take one required logger (warn/error with a scope and fields). The production logger writes each entry as a failed span to <userData>/logs/main.trace.ndjson, which the diagnostic bundle collects, and to the console (stderr under a supervised headless host). The runtime and the host wrap it so a logger that throws never fails what it reports, and the install refuses without one. Sites that deliberately kept a recovery-capsule error out of the log still log no error object. * refactor(native-chat): hand the chat host's collaborators the logger, and give orcad its trace file The delivery loop, idle sweep, queued-message drain, lease renewer, event sink, conversation map and provider start/exit settlement each took an internal error callback that the host mapped onto the logger. They now take the logger itself and log under their own scope. The event sink keeps one onFailed hook, which decides whether to stop the provider, not whether to report. The dead-generation settlement returns its failure so each caller logs it under its own scope. orcad now installs the desktop's local trace sink under its own data root, so a headless host's chat failures reach <data-root>/logs/main.trace.ndjson as well as stderr. Also passes the logger in the test fixtures the first commit missed, which tc:node caught. * fix(native-chat): keep repeated chat failures from flooding the trace file, and record their causes - The production structured-chat logger writes a repeated failure (same level, scope, session, message and error text) once per 5 minutes, carrying how many repeats it swallowed; the tracked set is capped at 256. - Trace entries now carry the error's code (and SQLite errcode) and up to three causes by name and message. - A chat read whose conversation will not open is logged through the host's logger (open-for-read), and so are the runtime's chat-tab bookkeeping failures that already hold the host. - orcad writes its own orcad.trace.ndjson, closes it after every quit handler, and flushes it on process exit; a trace file that cannot be opened leaves tracing off instead of stopping the app or orcad. - Tests: the desktop wiring test proves the logger reaches the trace sink, and the privacy tests read every level the logger received. * fix(native-chat): log a created chat's tab-publication and launch-prompt failures through the host's logger * fix(native-chat): key a repeated chat failure on everything its entry writes The repeat suppression keyed on the message and the error's text, so two refusals with the same code but different causes, a plain error and a refusal of one code, or two object-valued errors shared a key and the second was swallowed for five minutes. The key is now the entry's whole written content (fields, code, errcode, refusal reason, cause chain, a stable rendering of a non-error value) plus the error's name and message; a refusal's reason is also written. * test(native-chat): pin that an error's name keeps two repeated failures apart * test(native-chat): build the refusal in the repeat-key test as the wire does * fix(native-chat): read an error's code and a refusal's reason by narrowing, not Reflect.get |
||
|
|
6d1a97ef98 |
fix(ssh): launch the Windows relay outside sshd's job so standard users work (#24224)
* fix(ssh): launch the Windows relay outside sshd's job without WMI Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js gains a one-shot launcher mode that starts the detached relay with CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a relay without the addon, and a refusal there is named. The Windows SSH-host lanes drop their WMI grant and assert the breakaway route and adoption. * fix(ssh): find runtime holds without WMI on a standard-user Windows host The store GC read held runtimes through Get-CimInstance Win32_Process, which WMI refuses to a standard user's SSH logon, so the pass kept every runtime. On a refusal it now reads this account's own process image paths through Get-Process. * build(relay): ship the Windows relay launcher addon in every desktop package macOS and Linux packages carried Windows relays without windows-process-tree.node, so a legacy-runtime relay they uploaded to a Windows SSH host could not launch outside sshd's job and fell back to WMI, which a standard user is refused. A reusable Windows job now compiles the x64 and arm64 addons once and uploads them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds download them before build:release and require both arches. Staging now rejects a binary with the wrong PE machine, the ReadProcessMemory import, or no spawnOutsideJob export, so a stale pre-launcher build cannot ship. * ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change The staging and gyp-rebuild scripts decide which windows-process-tree addon the relay ships, so a change to either must re-prove the Windows host cells. * test(ci): find the mac orcad-template download by artifact name The release mac job now also downloads the relay Windows process-tree addons, so the first download-artifact step is no longer the template's. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
14d4bb2e2a |
fix(ssh): Windows hosts without Add-Type staging; runtime-store GC on Windows (#24149)
* fix(ssh): collect the pinned-Node runtime store on Windows hosts
Windows SSH hosts now run runtime-store GC instead of skipping it: one
PowerShell inventory reads .runtime-ref-node-<sha> and .runtime-node refs from
every version dir, and one Get-CimInstance Win32_Process query filtered on an
image path under runtimes\ adds process holds (never by image name; a failed
query keeps everything). Stale upload stages are swept with the same rule as
POSIX. Promotion and the post-upload hold check now take the store lock on
Windows too, and the lock's own commands run unwrapped there.
Windows relay version-dir liveness now honours .relay-pid (design D5): a live
PID answers ALIVE before any pipe is touched, a dead one (ESRCH) plus refusing
pipes is exited, anything else is unverifiable. The runtime probe adopts a
pinned node.exe an earlier vault reader left without a .verified marker after
running it.
* fix(ssh): Windows stage fencing and vault runtime go through the verified node.exe
Upload-stage file identity on Windows no longer compiles an Add-Type P/Invoke
helper when the relay runs on Orca's verified pinned node.exe: the stage
commands run a fixed fs.lstatSync(..., {bigint:true}) script through it. It
prints the legacy helper's vol:high:low lowercase hex, and identity files are
compared after normalising hex spelling, so old and new clients recover each
other's stages. Host-Node relays keep the legacy helper; the choice is
documented in windows-edr-posture.md.
The Windows OpenCode vault reader now installs the pinned runtime through
ensureRemoteOrcadNodeRuntime (official zip, host-side extraction, .verified,
store lock) instead of uploading a client-extracted node.exe, and the relay dir
gains a .runtime-ref-node-<sha> so store GC keeps the runtime the vault uses.
* test(ssh): run the Windows stage-identity and store-GC tests on the Windows lane
The legacy/node.exe identity compatibility test and the Win32_Process hold path
were gated to win32 but no CI lane ran them. Add both files to the Windows
package lane and a real running-node.exe hold test.
* test(ssh): tear down Windows-lane temp trees through removeTreeSync
* test(ssh): grant the store lock to the Windows OpenCode runtime setup test
The Windows promote now runs under runtimes/.store-lock, so the mocked host
must answer the lock's CreateNew step.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
|
||
|
|
bd90da7a5b | ci: share PR planning setup and reuse the static native cache (#24329) | ||
|
|
53fd2dea0b |
feat(ssh): relay runtime fallback ladder, telemetry and host runtime setting (#24133)
* feat(ssh): complete the relay runtime fallback ladder (D6 rungs B slot, C, D) Rung C runs the relay on the host's Node >= 18 with Orca's prebuilt N-API addons and no npm (addon-only probe mode). Rung B is a data-driven slot chosen only when a compat runtime is listed. Rung D fails the connect with a classified reason carried as a TerminalUnavailableCause. The ladder steps down only on classified refusals; unanswered probes throw. The rung decision is persisted per host keyed by (glibc, runtime hash, Orca major), and ssh_remote_runtime_resolved reports it once per host per session. * feat(settings): SSH host runtime choice (Auto | Orca-managed Node | Host Node) * docs(telemetry): describe ssh_remote_runtime_resolved * fix(ssh): let a passing rung C disprove a remembered noexec; allow glibc-less compat runtimes A remembered rung A noexec was re-persisted even after rung C self-tested addons from the same ~/.orca-remote tree, so rung A stayed skipped until the key changed. Rung B's evaluator also could never match a musl compat runtime. * test(ssh): import node:fs once in the host-node addon test --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
ddd4927a0b |
build(orcad): server node-pty slots at glibc 2.28, plus a glibc 2.17 compat slot (#24134)
* build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot
Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in
manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25
through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps
its Ubuntu 20.04 (2.31) default.
Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS
pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside
SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014
with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in
DT_NEEDED, and smokes it under the glibc-217 Node.
* refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
|
||
|
|
6593d7d194 |
feat(orcad): run orcad on the pinned Node instead of Bun (#24110)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working tree must attach the newest release tag's daemon. Rollback crossing is reported only. Runs in the cross-version-wire job, which already has full tags; tag selection moves to config/scripts/stable-release-tags.mjs so both use one rule. * feat(persistence): run profile backups in the worker whenever its entry is bundled * refactor(orcad): make profile and native preflight runtime-neutral The profile preflight parser now takes the expected runtime identity from the caller (shipped callers pass the pinned Bun identity), and the native preflight is renamed to orcad-runtime-native-preflight with neutral wording. * feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS, NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball), generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no network, that the pin tracks the locked Electron, matches engines.node's major, and covers exactly SERVER_TARGETS; it runs in the static analysis job. ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target list; orcad's Bun runtime and build output are unchanged. * test(persistence): skip plain-Node backup selection tests in the Bun profile suite * fix(runtime): reject a pinned archive that belongs to another target * ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change, so one PR must not do both. The launcher file list lives in the check script; the allow-runtime-launcher-protocol-bump label overrides it. * feat(orcad): select pinned-Node slots by a .runtime-node marker D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through .runtime-node instead of .build-target, so Bun-era clients read it as a legacy slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet. * fix(runtime): load the Node pin without the typeless-module warning check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from its own module, so it no longer loads the update script's build graph. * fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout * feat(orcad): 8-slot node-pty prebuilds against the pinned Node headers at N-API 8 - build-orcad-prebuilds.mjs adds win32-x64/arm64 (conpty.node, the vendored conpty.dll/OpenConsole.exe, upstream's N-API conpty_console_list.node), compiles in a scratch copy against the hash-verified pinned headers (node.lib pinned per Windows arch) with NAPI_VERSION=8, rejects post-8 node_api_* imports, and writes a schema 2 manifest with per-file sha256, N-API level and the glibc need. - --require-slots [slots] verifies files against hashes; --smoke loads the slot under the pinned Node and spawns a PTY; --print-slot names the host slot. - The slot installer gates on N-API, libc, arch, glibc and file hashes instead of the exact NODE_MODULE_VERSION, and installs nested files (conpty/). - bun-profile-tests.yml builds, verifies and smokes each runner's slot. * fix(orcad): scope node-pty's glibc .symver pins to glibc on musl prebuild slots musl's unversioned libc cannot satisfy openpty@GLIBC_* references at link time, so the Alpine slot compile would fail. Pin the staged pty.cc guard to __GLIBC__ and assert both musl transforms against the installed patch. * feat(orcad): run orcad on the pinned Node instead of Bun A packaged orcad slot now references the pinned Node 24.21.0 by its executableSha256 (`.runtime-node`, `.server-target`) instead of carrying bun-runtime, and ships node-pty from the slot's prebuild, only its own ripgrep, and no Windows Bun PTY gate. The runtime lives beside the slots at runtimes/node-<sha>/bin/node (node.exe on Windows, upstream name). - build:orcad (build-orcad-node.mjs) builds the host slot's prebuild when missing and places the pinned runtime; the template is schema 3 with per-target files. - handoffToBundledOrcad() resolves the slot's runtime reference and checks process.versions.node against the pin; a host Node >= 18 still hands off. Startup preflight keys on running as that runtime; callers expect 'node'. - orcad and its daemon use node-pty (ConPTY + windows-pty-job on Windows); the Bun PTY sources, gate entry and canUseBunPty branches are removed. - SSH deploy uploads the official archive once per pin, extracts and hash-checks it on the host, and self-tests it before publishing. Bun slots stay launchable for rollback; Node slots never use host Node. - The runtime materializer is generic over pinned assets; the Bun wrapper remains only for the OpenCode vault reader (design Phase 2). - Cross-runtime test: a profile DB written by Bun 1.4.2 (WAL left by SIGKILL) opens and backs up under the pinned Node, and the reverse. No daemon PROTOCOL_VERSION change (design D7.1 R3). * docs(ci): name the headless lanes after the pinned Node, drop Bun shard timings Design D10: ci-demand-rollout.md and ci-runner-efficiency.md follow the bun-profile-tests.yml -> node-server-tests.yml rename; shard timings drop the deleted Bun PTY tests and follow the renamed ones. * chore(ci): count the runtime archive download as a runtime launcher path * fix(orcad): pin the macOS C++ standard for node-pty prebuilds The official Node headers' config.gypi sets clang: 0, so common.gypi skips its gnu++20 xcode_settings and Apple clang 15 (macos-14 runners) compiles node-addon-api as C++98. * fix(orcad): resolve the preflight's slot through realpath, as the handoff does A symlinked orcad.js handed off to its real slot's pinned Node, but the startup and profile preflights read the symlink's directory, found no runtime marker there, and silently skipped the readiness check. * refactor(ssh): drop materializeCachedNodeRuntime, which nothing calls Deploys upload the verified official archive (design D5); no client path needs an extracted Node executable cached by digest. * test(orcad): gate the Bun-to-Node upgrade and Node-to-Bun rollback with live terminals Design D7.1 R1/R3/R4 and D7.2. The last Bun orcad and this checkout's Node slot are installed side by side under ~/.orca-remote, launched and stopped with the client's own deploy commands, and share one data root. Each direction proves the incoming orcad adopts the outgoing runtime's daemon (same PID, same shell, output continues), opens its profile database and backs it up with its own shipped worker, and that GC keeps the slot the live daemon was forked from. The node-server Linux lanes provide Bun 1.4.2 and build that Bun orcad from main, and run with --cross-runtime. --artifact and --cross-runtime now make their tests fail on a missing input instead of skipping. * ci(node-server): pin node:24.21.0-alpine by its multi-arch index digest * test(ssh): name the runtime archive fixture after its role * test(node-server): load node-pty from the packaged slot in artifact runs The node-server lane installs dependencies without building node-pty, and Linux has no upstream prebuild, so the real-PTY failed-I/O teardown test (picked up by the pty-subprocess selector) could not load pty.node. In --artifact runs, alias node-pty to out/orcad's shipped slot so the test exercises the addon orcad actually runs under the pinned Node. * fix(orcad): let the Windows profile preflight exit after its PTY probe On Windows, node-pty keeps the conout worker thread and pseudoconsole alive until kill(), even after the shell exits. The PTY health probe never killed a cleanly exited probe, so the packaged preflight printed its readiness line and then hung until the build's 30s timeout, reported with an empty stderr. - The probe kills its PTY on Windows after exit and uses the bundled ConPTY the daemon spawns with. - The preflight exits once stdout is flushed; its owner reads to EOF. - Preflight failures now report code, signal, timeout, stdout and stderr. * test(node-server): load the slot's node-pty in the real-PTY test, not by alias A vite alias redirected only ESM imports of node-pty; windows-pty-job and local-pty-utils resolve it through require, so Windows loaded two conpty.node copies and the Git Bash job-membership proof read an empty job. The failed-I/O teardown test now loads node-pty through a fixture that picks the packaged slot in artifact lanes. The pty-subprocess selector was a prefix that also pulled in its POSIX-host sibling unit tests, which pr.yml runs and which were never qualified on Windows. Select the directory plus the two sibling files that belong here. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
d4ae6a904b | Update README downloads badge | ||
|
|
daf63e659c |
fix(runtime): read Antigravity, Cline and Prime Agent readiness from the live screen (#24222)
* fix(runtime): decide Antigravity readiness from the live screen agy paints its composer with cursor addressing, so the line-folded wait text misses the 1.2.14 accept-edits and plan composers and an ended turn, while the grid keeps the bare `>` caret painted mid-turn and behind the /model picker. Read the screen's bottom rows instead: rule, caret, rule, `? for shortcuts`. A clocked pane is held to quiescence (tier 1b) because the submit repaint reads ready for a moment; a clockless restored pane settles from the screen alone. When a trustworthy screen exists it decides, so the name-only title lane no longer settles an open picker. Retires the visible-read probe's Antigravity branch: the probe now runs the shared screen rule for any screen-ruled agent without an output clock, and keeps its generic empty-pane read for everyone else. Adds twelve agy 1.2.14 recordings and a replay suite shared by screen-ruled agents. STA-8741. * fix(runtime): decide Cline readiness from the live screen Cline paints its composer box with cursor addressing on the alternate screen, so no text rule saw it and worker-start timed out at agent_readiness (#23268). Read the box off the grid: rule, an empty composer with one of the captured placeholders, rule, the Plan/Act row and the auto-approve row, with no braille spinner above it. A streaming reply repaints the same box once its spinner has scrolled away, so Cline is tier 1b only: a clocked pane waits for quiet and a clockless one never settles from the screen. The screen now decides for a Cline pane, which shuts the quiet-process lane that would have settled its unworded tool-approval prompt and the Cline Desktop promo. readLiveTerminalScreenLines now returns raw rows: the read projection blanks a composer it takes for a draft, and it takes Cline's placeholder for one, so a typed draft and an empty composer looked the same. Adds nine cline 3.0.66 recordings (macOS) and the 3.0.65 Windows capture from #23269. STA-8741. * fix(runtime): decide Prime Agent readiness from the live screen Prime redraws its composer on the alternate screen, so the text tail never showed a settled prompt and tui-idle timed out (#22153). Read the grid instead: a bare `>` directly over the `<- manage` footer, with no braille status row (`Writing - 6s`) above it. The footer and caret alone stay painted for a whole turn. Replayed chunk by chunk, Prime erases that status row before redrawing it, and on first launch paints the idle composer just before the trace-sharing question covers it. Both keep repainting, so a clocked pane is held to quiescence (tier 1b); a clockless restored pane settles from the screen alone. Adds nine prime-agent 0.9.8 recordings (isolated HOME, OpenRouter) and the two 0.9.5 captures from #22154. STA-8741. * refactor(runtime): drop Cline-only readiness branches Cline now follows the same pattern as Antigravity and Prime: a screen rule plus table entries. - Drop MID_TURN_COMPOSER_AGENTS. onPtyData stamps lastOutputAt on every chunk, so a re-attached streaming pane has an output clock from its first byte; the exception only guarded a pane that printed nothing since attach. A clockless Cline pane now settles from its screen like the other two. - Drop the 'ready-body' rest-signal entries for all three agents. The rest signal is read only by quietForegroundLane, and a readable screen already shuts that lane and the title lane (isReadinessDecidedByScreen), so the entries only removed the quiet-process fallback for a pane with no trustworthy grid. The census now checks that screen-shut instead. - Drop the Cline rule's auto-approve row check; no recorded verdict depends on it. Kept: raw rows from readLiveTerminalScreenLines. Every frame of every codex-* and qoder-* capture at 120x40, 80x24 and 100x32 gives the same isKnownReadyPromptBody (with and without a clock) and isQuietReadyScreenBody verdict through both readers. Serializer known-failures for the new captures are pre-existing serializer behaviour, not this branch: row-0 cells restore with a true-colour background where the source has the default (the DSH class), and Prime's cursor restores at column 119 instead of the pending wrap at 120 (the qoder class). STA-8741. * fix(runtime): trust a screen rule only on the PTY's own grid Review findings on the screen-ruled readiness (STA-8741): - A grid out of step with the PTY garbles cursor-addressed chrome, and a model resize does not make the TUI repaint. readLiveTerminalScreenLines now returns null unless the emulator's grid matches the PTY's reported size and was never reflowed without a repaint (a re-attach that learned the real size late), so the pre-existing lanes decide there instead of timing out. - The visible-read probe reads the draft-blanking projection, which turns Cline's `❯ Ask anything...` into a bare `❯`. It now restores the blanked composer row before the rule reads it; `terminal read --screen` output is unchanged. - The quiet lane no longer ORs the text rules over a trustworthy screen that refused; without one, tier 1 already ran them. No recorded verdict changes. Tests: ready recordings on a mismatched and on a reflowed grid settle through the old lanes; the restored-pane probe runs every ready recording through the real projection; the rest-signal census checks the lane verdict with and without a screen. * test(runtime): trim STA-8741 recordings to the screens they prove * refactor(runtime): one screen verdict for every screen-ruled lane readScreenRuledReady, readScreenRuledQuietReady and isReadinessDecidedByScreen each re-derived the same thing: the agent's rule applied to a trustworthy live screen. They collapse into readScreenRuledVerdict (true / false / null), which tier 1, tier 1b and the lane gate read. This also makes a refusal final in tier 1: a clockless pane whose trustworthy screen refused fell through to the text rules, so retained ready text could settle over an open picker (Greptile review). The quiet tier already refused there; now both do. The tier-1b agent set derives the screen-ruled agents from the rule table instead of listing them again, and the lane test that repeated the census case is dropped. * refactor(runtime): let the visible-read probe read its own output clock The probe's clock was captured at start and threaded through the wait dependencies as a one-off parameter. The probe now reads it from the live record when its screen read returns, which is also the fresher answer. * fix(runtime): trust a reflowed grid again once a PTY resize repaints it The reattach-reflow flag was never cleared, so a pane stayed on the old lanes for the rest of its life even after a real resize made the TUI repaint (Greptile review). The record now keeps the reflowed grid, and a PTY resize off that grid clears it; an echo of the same size sends no SIGWINCH and keeps it. Tests: the reflow case in every screen-ruled suite now includes a same-size echo, and an Antigravity recording only the screen reads ready settles after a resize and repaint. * refactor(runtime): keep screen-rule trust and raw rows to screen-ruled agents Two shared changes reached agents this PR does not target: the live screen reader returned raw rows, and it refused a grid that did not match the PTY. Both now live in readScreenRuledLines, which only the screen-ruled agents read (screenReader picks it from the rule table); readLiveTerminalScreenLines is main's again. The probe keeps main's Antigravity-banner trigger, so a Codex or unknown pane is probed exactly as before. Proof: the non-screen-ruled suites give identical pass sets on this branch and its base (1,781 tests), and replaying every other recording frame by frame through the readiness and blocked verdicts, for its agent and for an unknown pane, gives identical results (93 pairs). A new test keeps a Codex pane reading its screen when the PTY reports another grid; it fails if the trust check moves back into the shared reader. |
||
|
|
a4606ccae3 |
fix(cli): orca file open no longer moves your view unless you pass --focus (#24244)
* docs(cli): file open/diff/open-changed say they switch the user's view and are for user requests only Refs #9944 * fix(cli): file open/diff/open-changed leave the user's view alone unless --focus `orca file open`, `file diff` and `file open-changed` always switched the desktop to the target worktree, selected the tab and revealed it in the sidebar. An agent skill that opens its answer pulled the user out of whatever they were typing in (#9944), and a phone opening a file moved the desktop too. The commands now add the tab in its worktree without changing anything on screen, including when that worktree is the one being viewed: the new tab is added to the tab bar but the active tab, tab type and focus stay put. In a worktree the user is not viewing, the tab becomes that worktree's selection so it is in front when they go there. `--focus` keeps today's behavior. files.open / files.openDiff take an optional `navigation` target (the existing RUNTIME_NAVIGATION_TARGETS vocabulary); the CLI sends 'all' for --focus, like `worktree create --activate`, and nothing otherwise. The renderer moves the host view only when the target reaches the host; a missing field (phones, older CLIs) leaves it still. Editor opens for a worktree other than the on-screen one no longer write the global activeFileId/activeTabType. Refs #9944 * test(cli): justify the window and runtime stubs in the file-open notification test * fix(cli): keep phone file opens switching the desktop; the CLI asks for 'caller' Phone opens send no `navigation` field, and the phone's diff-review "Open in session" relies on the desktop selecting the diff it opened. A missing field now keeps the original switch exactly; the CLI says what it wants instead: 'caller' (no host move) by default and 'all' for --focus. Older CLIs, which send nothing, keep switching as they always have. Refs #9944 * fix(cli): background file opens select the tab without counting as a visit A CLI open into a worktree the user is not viewing selected the new tab with the same activation a user click uses, which stamps lastFocusedAt and the group's recency list. The worktree jump palette sorts recent tabs by that time, so every agent `orca file open` into another worktree jumped to the top of the user's recent tabs. Editor opens now take a selection mode: 'focus' (default, unchanged), 'background' (select within its worktree without recording focus or recency) and 'none' (add only). createUnifiedTab and activateTab gain recordFocus:false for the background case. Also: tests for reopening an already-open file or diff without --focus, a comment that file opens move only the host window ('all' acts as 'host'), root help lines back under 100 columns, and an accurate remote test title. Refs #9944 * fix(tabs): a background-selected tab still joins its group's tab history recordFocus:false skipped both the focus-time stamp and the group's recentTabIds append while still making the tab the group's active tab. Ctrl+Tab looks the active tab up in that history, so after a background CLI open it did nothing (or went to the wrong tab) once the user switched to that worktree, and hydrate kept the broken history across a restart. Only the focus-time stamp is skipped now; the jump palette's recent rows sort by that alone, so the palette fix stands. Refs #9944 * fix(cli): file open/diff/open-changed --focus help says it brings the user to the file The three commands borrowed the shared --focus line written for terminal create ("Reveal the created terminal session in Orca"). They now use the per-command flag help table; terminal create's line is unchanged. Refs #9944 |
||
|
|
0b79720c2e |
feat(native-chat): the chat strip and the sidebar read the host's child records (#22614)
* feat(native-chat): publish the host's child records to the status summary and the chat strip
The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.
* feat(native-chat): the chat strip reads the host's child records with its parent's verdict
The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.
Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.
* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open
- The view decoder ignores unknown keys, degrades unknown kinds, states,
outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
roster of finished children and never the views themselves; a stop-only
reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.
* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered
* test(native-chat): type the switch tests' mocks instead of asserting them
* test: remote clients advertise reading child views
* docs(agent-status): the structured row folds the store's child records
* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary
The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.
* refactor(native-chat): the status summary's broadcast equality gets its own module
The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.
* fix(native-chat): command admission reads the strip's child records
A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.
Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.
* refactor(native-chat): command admission takes only what it reads of a turn
* fix(native-chat): the session list drops a session's children when the store does
A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.
The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.
* test(native-chat): write the Codex frame script's parent row out step by step
Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.
* fix(native-chat): the idle sweep and the restart snapshot read the host's child records
The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.
The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.
* test(native-chat): the child-record tests follow the merged command lifecycle
A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.
Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.
* refactor(native-chat): the status feed's journal projection cache gets its own module
The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.
* test(native-chat): the admission test's compaction resolves with a real outcome
Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.
* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished
The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.
This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.
* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source
`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.
A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.
* test(native-chat): the switch test passes the startup child key main's status bar takes
* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own
Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.
Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.
* fix(native-chat): a background Stop reaches the tasks the child records show
The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.
The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.
* fix(native-chat): one rule for a finished child that still owns live work, at any depth
The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.
* fix(native-chat): an older client sees a Codex child's shell as it did before views
Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.
* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives
The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.
* fix(native-chat): the strip channel forgets a closed conversation's roster
It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.
* docs(native-chat): rewrap the retention comment
* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays
Two lifecycle gaps from the round-1 fixes.
A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.
A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.
Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.
* fix(native-chat): the strip keeps one empty list for a roster that omits one
A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.
* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent
The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.
* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once
A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.
The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.
Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.
* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader
CI on
|